Perform data science with Azure Databricks

所在平台: Coursera

课程主页: https://www.coursera.org/learn/perform-data-science-with-azure-databricks

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:在Azure Databricks上执行数据科学 课程概述:本课程将教你如何利用Apache Spark的强大功能和在Azure Databricks平台上运行的强大集群,以便在云中执行数据科学工作负载。这是一个五门课程计划中的第四门,旨在为您准备DP-100认证考试:在Azure上设计和实施数据科学解决方案。 该认证考试是一个证明您在使用Azure机器学习以云规模运行机器学习解决方案方面的知识和专业技能的机会。本课程特别适合对Python和机器学习有一定基础的数据科学家,帮助他们管理数据摄取和准备、模型训练和部署,以及机器学习解决方案的监控。每门课程都教授与考试相关的概念和技能。 该专业化课程旨在帮助那些希望在云端构建和操作机器学习解决方案的数据科学家,教学内容包括如何在Microsoft Azure中创建端到端解决方案。学员将学习如何管理Azure的机器学习资源;运行实验和训练模型;部署和实现机器学习解决方案,并实施负责任的机器学习。同时,他们还将学习如何使用Azure Databricks进行数据探索、准备和建模,并将Databricks的机器学习流程与Azure机器学习集成。 课程大纲: 1. **Azure Databricks简介**:发现Azure Databricks的能力和Apache Spark笔记本,用于处理大型文件,了解Azure Databricks平台及其工作流程。 2. **在Azure Databricks中处理数据**:掌握日常数据处理功能,处理来自多源和不同原始格式的大量数据,应用DataFrame列类进行数据转换。 3. **在Azure Databricks中处理数据**:学习如何注册和调用用户自定义函数(UDF),利用Delta Lake创建、追加和更新Apache Spark表中的数据。 4. **开始使用Databricks和机器学习**:使用PySpark机器学习包构建机器学习工作流的关键组件,包括探索性数据分析、模型训练和评估。 5. **管理机器学习生命周期并微调模型**:学会使用MLflow追踪机器学习实验,利用Spark的机器学习库模块进行超参数调整和模型选择。 6. **训练分布式神经网络并使用Azure机器学习提供模型服务**:使用Uber的Horovod框架和Petastorm库在Spark上运行分布式深度学习训练作业,学习如何注册、打包和部署训练模型。 通过本课程的学习,学员将掌握在Azure环境中构建和部署机器学习解决方案所需的关键技能。

课程大纲

Name:Introduction to Azure Databricks

Description:In this module, you will discover the capabilities of Azure Databricks and the Apache Spark notebook for processing huge files. You will come to understand the Azure Databricks platform and identify the types of tasks well-suited for Apache Spark. You will also be introduced to the architecture of an Azure Databricks Spark Cluster and Spark Jobs.

Name:Working with data in Azure Databricks

Description:Azure Databricks supports day-to-day data-handling functions, such as reads, writes, and queries. In this module, you will work with large amounts of data from multiple sources in different raw formats. You will also learn to use the DataFrame Column Class Azure Databricks to apply column-level transformations, such as sorts, filters and aggregations. You will also use advanced DataFrame functions operations to manipulate data, apply aggregates, and perform date and time operations in Azure Databricks.

Name:Processing data in Azure Databricks

Description:Azure Databricks supports a range of built in SQL functions, however, sometimes you have to write custom function, known as User-Defined Function (UDF). In this module, you will learn how to register and invoke UDFs. You will also learn how to use Delta Lake to create, append, and upsert data to Apache Spark tables, taking advantage of built-in reliability and optimizations.

Name:Get started with Databricks and machine learning

Description:In this module, you will learn how to use PySpark’s machine learning package to build key components of the machine learning workflows that include exploratory data analysis, model training, and model evaluation. You will also learn how to build pipelines for common data featurization tasks.

Name:Manage machine learning lifecycles and fine tune models

Description:In this module, you will learn how to use MLflow to track machine learning experiments and how to use modules from the Spark’s machine learning library for hyperparameter tuning and model selection.

Name:Train a distributed neural network and serve models with Azure Machine Learning

Description:In this module, you will learn how to use the Uber’s Horovod framework along with the Petastorm library to run distributed, deep learning training jobs on Spark using training datasets in the Apache Parquet format. You will also learn how to use MLflow and Azure Machine Learning service register, package, and deploy a trained model to both Azure Container Instance, and Azure Kubernetes Service as a scoring web service.

课程评论(0条)

课程详情

In this course, you will learn how to harness the power of Apache Spark and powerful clusters running on the Azure Databricks platform to run data science workloads in the cloud. This is the fourth course in a five-course program that prepares you to take the DP-100: Designing and Implementing a Data Science Solution on Azurec ertification exam. The certification exam is an opportunity to prove knowledge and expertise operate machine learning solutions at a cloud-scale using Azure Machine Learning. This specialization teaches you to leverage your existing knowledge of Python and machine learning to manage data ingestion and preparation, model training and deployment, and machine learning solution monitoring in Microsoft Azure. Each course teaches you the concepts and skills that are measured by the exam. This Specialization is intended for data scientists with existing knowledge of Python and machine learning frameworks like Scikit-Learn, PyTorch, and Tensorflow, who want to build and operate machine learning solutions in the cloud. It teaches data scientists how to create end-to-end solutions in Microsoft Azure. Students will learn how to manage Azure resources for machine learning; run experiments and train models; deploy and operationalize machine learning solutions, and implement responsible machine learning. They will also learn to use Azure Databricks to explore, prepare, and model data; and integrate Databricks machine learning processes with Azure Machine Learning.

课程标签

0人关注该课程

主题相关的课程