|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/data-engineering-and-machine-learning-using-spark
课程评论:没有评论
课程名称:使用Spark的数据工程与机器学习 课程概述:在当今大数据时代,组织需要能够应用商业和技术技能处理非结构化数据(如推文、帖子、图片、音频文件、视频、传感器数据和卫星图像等)的高技能大数据从业者。该短期课程将使您掌握实用技能,通过学习如何利用Apache Spark来进行数据工程和机器学习应用。您将通过实践,使用Spark MLlib、Spark结构化流处理等工具执行抽取、转换和加载(ETL)任务,以及回归、分类和聚类分析。 课程以一个项目为结束,您将在项目中应用所学的Spark技能设计一个用于机器学习的ETL工作流程案例。 注意:本课程要求您具备Apache Spark和Jupyter Notebooks的基础技能。建议您在开始本课程之前完成IBM的《大数据简介与Spark和Hadoop》课程,以便获得所需技能。 课程大纲:
第一部分:Spark用于数据工程
描述:在本模块的开始,您将学习流数据的概念,以及如何使用Spark进行结构化流处理的基本知识。了解数据源、流输出模式和支持的数据目标,研究数据操作的注意事项,并发现Spark结构化流处理中的侦听器和检查点在处理流数据时的优势。此外,探索图论与流数据的关系,了解Apache Spark GraphFrames的优点,以及哪些特性使数据适合GraphFrames的处理。进一步探讨ETL的概念,并实践使用Apache Spark进行数据的提取、转换和加载,从而在机器学习管道的实践实验室中获得实际的、现实的技能。
第二部分:SparkML
描述:本模块将揭秘使用SparkML和Spark机器学习库的相关概念和实践。探索监督学习和无监督学习,学习如何支持分类和回归任务的SparkML。深入了解无监督学习,重点关注聚类,并探索如何使用Spark MLlib应用k-means聚类算法。课程最后的实验室将巩固您的学习,帮助您获得实际的Spark ML经验。
第三部分:期末项目
描述:这个期末项目将提供实战经验,您将创建自己的Apache Spark应用程序。该应用程序将作为一个端到端的用例,遵循提取、转换和加载(ETL)流程,包括数据获取、转换、模型训练和使用IBM Watson机器学习的部署。
Part: 1
Title:Spark for Data Engineering
Description:In this first of two modules, learn what streaming data is and get the essential knowledge to use Spark for Structured Streaming. Learn about data sources, streaming output modes, and supported data destinations. Learn about data operations considerations and discover how Spark Structured streaming listeners and checkpointing benefit streaming data processing. Discover how graph theory works with streaming data. You’ll gain insights into the advantages that Apache Spark GraphFrames offers and learn what qualities make data suitable for GraphFrames processing. Then, explore ETL and learn how to use Apache Spark for data extraction, transformation, and loading, put your newfound knowledge to practice, and gain practical, real-world skills in the ETL for Machine Learning Pipelines hands-on lab.
Part: 2
Title:SparkML
Description:This module demystifies the concepts and practices related to machine learning using SparkML and the Spark Machine learning library. Explore both supervised and unsupervised machine learning. Explore classification and regression tasks and learn how SparkML supports these machine learning tasks. Gain insights into unsupervised learning, with a focus on clustering, and discover how to apply the k-means clustering algorithm using the Spark MLlib. Complete this learning with the lab that solidifies your learning and gain real-world experience with Spark ML.
Part: 3
Title:Final Project
Description:This final project provides real-world experience where you'll create your own Apache Spark application. You will create this Spark application as an end-to-end use-case that follows the Extract, Transform and Load processes (ETL) including data acquisition, transformation, model training, and deployment using IBM Watson Machine Learning.
Organizations need skilled, forward-thinking Big Data practitioners who can apply their business and technical skills to unstructured data such as tweets, posts, pictures, audio files, videos, sensor data, and satellite imagery and more to identify behaviors and preferences of prospects, clients, competitors, and others. In this short course you'll gain practical skills when you learn how to work with Apache Spark for Data Engineering and Machine Learning (ML) applications. You will work hands-on with Spark MLlib, Spark Structured Streaming, and more to perform extract, transform and load (ETL) tasks as well as Regression, Classification, and Clustering. The course culminates in a project where you will apply your Spark skills to an ETL for ML workflow use-case. NOTE: This course requires that you have foundational skills for working with Apache Spark and Jupyter Notebooks. The Introduction to Big Data with Spark and Hadoop course from IBM will equip you with these skills and it is recommended that you have completed that course or similar prior to starting this one.