|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-engineering-projects-with-pyspark-2025/
课程评论:没有评论
**课程名称:** PySpark 数据工程全栈训练营 (2025) **课程概述:** 本课程旨在帮助您掌握使用 PySpark 成为一名数据工程师的实用技能,避免空泛的理论和过时的工具。通过真实的生产环境工具、结构和工作流程,您将学习到专业数据工程师的实际工作内容。 **学习内容与项目实践:** * **搭建完整数据工程栈:** 使用 Docker、Spark、Airflow、HDFS 和 Jupyter 搭建一套完整的数据工程环境。 * **构建生产级 PySpark ETL 作业:** 熟练运用 DataFrame API 和 Spark SQL 编写和部署生产就绪的 PySpark ETL 作业。 * **自动化与调度管道:** 通过 cron、Airflow DAGs 实现管道的自动化和调度,并利用 Spark UI 进行监控。 * **像真实数据工程师一样工作:** * 掌握 Git 分支、合并及实际版本控制工作流程。 * 学习专业项目结构:scripts/, configs/, env shell, 可重用模块。 * 无缝切换开发与生产环境。 * 模拟基于工单的部署和团队协作。 **课程特色:** * **超越语法:** 不同于多数仅教授语法的 PySpark 课程,本课程侧重于真实世界数据管道的构建。 * **生产环境定位:** 理解 Spark 在生产数据工作流中的具体作用。 * **模块化代码:** 构建模块化、生产就绪的代码库。 * **部署与监控:** 使用 spark-submit、cron 和 Airflow 进行作业部署,并通过 Spark UI、日志、缓存和调优技巧进行监控、调试和优化。 **具体学习要点:** * 搭建基于 Docker 的数据工程环境(Spark, Airflow, HDFS, Jupyter)。 * 利用 DataFrame 和 Spark SQL 构建可靠的 PySpark ETL 作业。 * 使用 spark-submit、Airflow DAGs 和 cron 调度实现管道自动化。 * 采用真实项目结构和 Git 工作流程组织代码。 * 完成两个完整的真实世界数据工程项目,模仿数据工程团队的工作方式。 **课程目标:** 完成本课程后,您将掌握数据工程师日常使用的实用、生产级技能。
Want to become a data engineer using PySpark - without wasting time on abstract theory or outdated tools?This course shows you exactly what professional data engineers do, using the tools, structures, and workflows used in real production environments.What You'll Learn Through Real Projects:Set up a complete data engineering stack with Docker, Spark, Airflow, HDFS, and Jupyter.Write and deploy production-ready PySpark ETL jobs using DataFrame API and Spark SQL.Automate and schedule pipelines using cron, Airflow DAGs, and monitor them with Spark UI.From Day 1, You'll Work Like a Real Data Engineer:Master Git branching, merging, and real-world version control workflows.Structure your projects professionally: scripts/, configs/, env shell, and reusable modules.Seamlessly switch between development and production environments.Simulate ticket-based deployments and team collaboration - just like real companies.What Makes This Course Different?Most PySpark courses teach only syntax. This course prepares you for real-world data pipelines:Understand exactly where Spark fits in production data workflows.Build modular, production-ready codebases.Deploy jobs using spark-submit, cron, and Airflow.Monitor, debug, and optimize pipelines using Spark UI, logs, caching, and tuning techniques.This course is a practical guide to building and deploying real data pipelines - like a professional data engineer.You Will Specifically Learn:Set up a Docker-based data engineering environment with Spark, Airflow, HDFS, and Jupyter.Build reliable PySpark ETL jobs using DataFrames and Spark SQL.Automate pipelines with spark-submit, Airflow DAGs, and cron scheduling.Organize your code with real-world project structures and Git workflows.Complete two full real-world data engineering projects - exactly how data engineering teams work.By the end of this course, you'll have practical, production-grade skills that real data engineers use daily.