|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-engineering-bootcamp-series-1/
课程评论:没有评论
课程名称:数据工程训练营 - 系列 1 课程概述:这一动手实践为基础的训练营是迈入数据工程领域的第一步,旨在帮助您提升职业技能。课程由一位拥有超过11年行业经验的数据架构师授课,结合理论与实践,特别设计给有志成为数据工程师、软件工程师、分析师以及任何渴望学习构建真实数据管道的人士。您将学习如何设计可扩展的数据湖、构建维度数据模型、实施数据质量框架,并使用Apache Airflow编排管道,课程通过真实的打车应用案例模拟企业级系统。 您将学习的内容: 第一部分:环境设置 构建现代数据栈的基础,了解OLTP系统,探索现实世界的数据平台架构。通过打车应用案例了解数据在数据驱动公司的流动,做好进入训练营的准备。 第二部分:数据湖基础 学习如何在AWS S3上构建和管理可扩展的数据湖。内容包括S3架构、分区、层次和模式演变、IAM、加密、存储类别、事件通知、生命周期管理及备份恢复。利用Boto3 S3 API进行实践。 第三部分:数据建模 掌握星型模式设计,实现SCD类型1和类型2维度。学习维度和事实建模、用于分析报告的ETL开发,并通过动手实验构建端到端模型和数据集市。 第四部分:数据质量 确保数据管道的可靠性和完整性。理解准确性、完整性和一致性,运用行业最佳实践实施数据质量检查,利用数据合同确保数据的可靠性。 第五部分:AWS Athena 使用AWS Athena的无服务器能力查询大量数据集。掌握DDL、Glue Catalog和工作组管理,通过Boto3 API自动化查询,比较Athena与Presto和Trino,并优化查询。 第六部分:Apache Spark 通过AWS EMR使用PySpark构建生产级数据管道。学习Spark架构和PySpark API,运用写-审核-发布(WAP)模式构建数据管道,在AWS EMR上运行可扩展的作业,应用UDF和数据质量检查于转换逻辑中。 第七部分:Apache Airflow 利用Airflow编排工作流并构建自定义插件,设计DAG,调度管道,管理依赖关系,使用自定义AWS EMR插件自动化Spark作业,进行数据摄取和转换DAG的实践,构建高效、可重用的编排解决方案。 所构建的内容:一个用于打车公司的生产级数据平台,包括: - AWS S3上的数据湖 - 具有SCD逻辑的维度数据模型 - 基于Spark的转换管道 - 使用Airflow进行自动化编排 - Athena查询层和内置数据质量验证
Take your first step into the world of data engineering and future-proof your career with this hands-on, project-based bootcamp built on the modern data stack. Taught by a seasoned data architect with over 11 years of industry experience, this course blends theory with practice, designed for aspiring data engineers, software engineers, analysts, and anyone eager to learn how to build real-world data pipelines.You will learn to design scalable data lakes, build dimensional data models, implement data quality frameworks, and orchestrate pipelines using Apache Airflow, all using a real-life ride-hailing application use case to simulate enterprise-scale systems.What You'll LearnSection 1: Context SetupBuild your foundation with the Modern Data Stack, understand OLTP systems, and explore real-world data platform architectures.Gain clarity on how data flows in data-driven companiesLearn using a ride-hailing app scenarioGet properly onboarded into the bootcamp journeySection 2: Data Lake EssentialsLearn how to build and manage scalable data lakes on AWS S3. S3 architecture, partitioning, layers, and schema evolutionIAM, encryption, storage classes, event notificationsLifecycle management, backup & recoveryHands-on with Boto3 S3 APIsSection 3: Data ModelingMaster star schema design and implement SCD Type 1 and Type 2 dimensions.Dimensional & Fact modelingETL development for analytical reportingBuild end-to-end models and data marts with hands-on labsSection 4: Data QualityEnsure trust and integrity in your data pipelines.Understand accuracy, completeness, and consistencyImplement DQ checks using industry best practicesUse data contracts for accountabilitySection 5: AWS AthenaQuery massive datasets with serverless power using AWS Athena.Learn DDL, Glue Catalog, and workgroup managementAutomate queries using Boto3 APIsCompare Athena vs Presto vs TrinoOptimize queries with best practicesSection 6: Apache SparkBuild production-grade data pipelines with PySpark on AWS EMR.Learn Spark architecture and PySpark APIsBuild data pipelines using the WAP (Write-Audit-Publish) patternRun scalable jobs on AWS EMRApply UDFs and data quality within transformation logicSection 7: Apache AirflowOrchestrate workflows using Airflow and build custom plugins:Design DAGs, schedule pipelines, manage dependenciesAutomate Spark jobs using custom AWS EMR pluginHands-on labs for ingestion and transformation DAGsBuild reliable, reusable orchestration solutionsWhat You'll BuildA production-style data platform for a ride-hailing company, including:Data lake on AWS S3Dimensional data model with SCD logicSpark-based transformation pipelinesAutomated orchestration with AirflowQuery layer with AthenaBuilt-in data quality validations