|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/spark-streaming-using-python/
课程评论:没有评论
课程名称:Spark Streaming - Lakehouse中的流处理 - PySpark 课程概述: 本课程旨在教授使用Python语言和PySpark API在Apache Spark和Databricks上进行流处理的知识。通过此课程,您将了解如何使用Apache Spark和Databricks Cloud进行实时流处理,并应用所学知识构建实时流处理解决方案。课程采用示例驱动的方式,模拟实际工作场景进行讲解,并采取实时编码的方法阐述所需的概念。 结束项目: 本课程还包括一个端到端的Capstone项目,该项目将帮助您理解现实生活中的项目设计、编码、实施、测试和持续集成/持续交付(CI/CD)方法。 适合人群: 该课程是为希望开发基于Apache Spark的实时流处理管道和应用程序的软件工程师设计的。同时,数据架构师和数据工程师也能从中受益,他们负责设计和构建组织的数据基础设施。此外,即使是与Spark实现无直接关联的管理者和架构师,参与该课程也对他们理解基层实施Apache Spark的工作大有裨益。 使用的Spark版本: 本课程使用Apache Spark 3.5版本。所有课程中使用的源代码和示例已在Azure Databricks Cloud的Databricks Runtime 14.1上进行了测试。
About the CourseI am creating Apache Spark and Databricks - Stream Processing in Lakehouse using the Python Language and PySpark API. This course will help you understand Real-time Stream processing using Apache Spark and Databricks Cloud and apply that knowledge to build real-time stream processing solutions. This course is example-driven and follows a working session-like approach. We will take a live coding approach and explain all the needed concepts.Capstone ProjectThis course also includes an End-To-End Capstone project. The project will help you understand the real-life project design, coding, implementation, testing, and CI/CD approach. Who should take this Course?I designed this course for software engineers willing to develop a Real-time Stream Processing Pipeline and application using Apache Spark. I am also creating this course for data architects and data engineers who are responsible for designing and building the organization's data-centric infrastructure. Another group of people is the managers and architects who do not directly work with Spark implementation. Still, they work with those implementing Apache Spark at the ground level.Spark Version used in the Course.This Course is using the Apache Spark 3.5. I have tested all the source code and examples used in this Course on Azure Databricks Cloud using Databricks Runtime 14.1.