Analyze Datasets and Train ML Models using AutoML

所在平台: Coursera

课程主页: https://www.coursera.org/learn/automl-datasets-ml-models

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:使用 AutoML 分析数据集并训练机器学习模型 课程概述:在实践数据科学专项课程的第一门课中,您将学习探索性数据分析(EDA)、自动化机器学习(AutoML)和文本分类算法的基础概念。通过使用 Amazon SageMaker Clarify 和 Amazon SageMaker Data Wrangler,您将分析数据集中的统计偏差,将数据集转换为机器可读的特征,并选择最重要的特征来训练多类文本分类器。随后,您将利用 Amazon SageMaker Autopilot 进行自动化机器学习,自动训练、调优并部署适用于给定数据集的最佳文本分类算法。接着,您将与 Amazon SageMaker BlazingText 一起使用,这是一种高度优化且可扩展的流行 FastText 算法实现,您将用极少的代码训练文本分类器。 实践数据科学旨在处理不适合您本地硬件的大型数据集,这些数据集可能来自多个来源。在云中开发和运行数据科学项目的最大好处之一是云提供的敏捷性和弹性,可以以最低的成本进行大规模扩展。 实践数据科学专项课程帮助您开发有效部署数据科学项目的实际技能,克服机器学习工作流各个步骤中的挑战,借助 Amazon SageMaker。本专项课程面向熟悉 Python 和 SQL 编程语言的数据驱动开发者、科学家和分析师,旨在教会他们如何在 AWS 云中构建、训练和部署可扩展的端到端机器学习管道,包括自动化管道和人机协作管道。 课程大纲: 1. 第1周:探索用例并分析数据集 - 描述:获取、探索和可视化用于多类文本分类的产品评论数据集。 2. 第2周:数据偏差与特征重要性 - 描述:确定数据集中最重要的特征并检测统计偏差。 3. 第3周:使用自动化机器学习训练文本分类器 - 描述:检查并比较使用自动化机器学习(AutoML)生成的模型。 4. 第4周:内置算法 - 描述:使用 BlazingText 训练文本分类器,并将其作为实时推断端点进行部署以提供预测。

课程大纲

Part: 1

Title:Week 1: Explore the Use Case and Analyze the Dataset

Description:Ingest, explore, and visualize a product review data set for multi-class text classification.

Part: 2

Title:Week 2: Data Bias and Feature Importance

Description:Determine the most important features in a data set and detect statistical biases.

Part: 3

Title:Week 3: Use Automated Machine Learning to train a Text Classifier

Description: Inspect and compare models generated with automated machine learning (AutoML).

Part: 4

Title:Week 4: Built-in algorithms

Description:Train a text classifier with BlazingText and deploy the classifier as a real-time inference endpoint to serve predictions.

课程评论(0条)

课程详情

In the first course of the Practical Data Science Specialization, you will learn foundational concepts for exploratory data analysis (EDA), automated machine learning (AutoML), and text classification algorithms. With Amazon SageMaker Clarify and Amazon SageMaker Data Wrangler, you will analyze a dataset for statistical bias, transform the dataset into machine-readable features, and select the most important features to train a multi-class text classifier. You will then perform automated machine learning (AutoML) to automatically train, tune, and deploy the best text-classification algorithm for the given dataset using Amazon SageMaker Autopilot. Next, you will work with Amazon SageMaker BlazingText, a highly optimized and scalable implementation of the popular FastText algorithm, to train a text classifier with very little code. Practical data science is geared towards handling massive datasets that do not fit in your local hardware and could originate from multiple sources. One of the biggest benefits of developing and running data science projects in the cloud is the agility and elasticity that the cloud offers to scale up and out at a minimum cost. The Practical Data Science Specialization helps you develop the practical skills to effectively deploy your data science projects and overcome challenges at each step of the ML workflow using Amazon SageMaker. This Specialization is designed for data-focused developers, scientists, and analysts familiar with the Python and SQL programming languages and want to learn how to build, train, and deploy scalable, end-to-end ML pipelines - both automated and human-in-the-loop - in the AWS cloud.

课程标签

0人关注该课程

主题相关的课程