|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/ultimate-aws-data-engineering-bootcamp-with-real-world-labs/
课程评论:没有评论
**课程名称:** Ultimate AWS Data Engineering Bootcamp: 15 Real-World Labs **课程概述:** 本门课程是学习 AWS 数据工程的权威指南,旨在帮助学员从初阶成长为精通级专家。通过掌握强大的 AWS 服务和工具,学员将能够应对各种真实世界的数据挑战。 **主要学习内容:** 学员将深入学习数据工程的核心概念,重点关注批处理和实时数据处理。课程将提供实践经验,包括: * **AWS Glue 和 EMR 上的 PySpark 批处理 ETL:** 设计、实施和优化可扩展的 ETL 管道,将原始数据转化为可操作的见解。 * **PySpark Streaming 实时流处理:** 精通实时数据处理和分析,高效精确地处理流式数据。 * **ECS 上的容器化 Python 工作负载:** 使用 ECS 管理和部署容器化 Python 应用,实现可扩展性和可靠性。 * **Airflow 和 Step Functions 数据编排:** 使用一流的工具编排复杂的工作流程和自动化数据管道。 * **AWS Kinesis 事件驱动和实时处理:** 构建健壮的事件驱动架构,实时处理流式数据,确保数据管道始终是最新的。 * **Amazon Redshift 数据仓库:** 深入理解 Redshift,高效存储和分析海量数据集。 * **MySQL Aurora 和 DynamoDB 数据库管理:** 实际操作关系型和 NoSQL 数据库,针对不同用例优化数据存储和检索。 * **Lambda 函数无服务器数据处理:** 利用 AWS Lambda 实时处理数据,根据事件触发工作流程。 * **Glue Python Shell 作业:** 在托管环境中运行 Python 脚本,用于自定义数据处理任务。 * **Spark 上的 Delta Lake:** 理解 Delta Lake 和 Lakehouse 架构的概念,以及它如何增强 Spark 以构建可靠、可扩展的数据湖。 * **GitHub Actions CI/CD:** 通过 GitHub Actions 实现持续集成和持续交付,自动化数据工程工作流程。 **为何选择本课程:** 本训练营不仅仅是理论课程,更包含大量真实的实战实验,模拟数据工程师日常面临的挑战。学员将有机会构建、部署和管理数据管道和架构,并能直接应用于工作或项目中。无论您是初学者还是希望提升技能,本课程都将助您成为 AWS 数据工程专家。
Welcome to the most definitive course for mastering data engineering on AWS. This comprehensive bootcamp is designed to take you from a beginner to an expert, equipping you with the skills to tackle real-world data challenges using the most powerful AWS services and tools.What You'll Learn:In this course, you'll dive deep into the core aspects of data engineering, focusing on both batch and real-time data processing. You'll gain hands-on experience with:Batch ETL and Processing with PySpark on AWS Glue and EMR: Learn to design, implement, and optimize scalable ETL pipelines, transforming raw data into actionable insights.Real-Time Streaming with PySpark Streaming: Master real-time data processing and analytics to handle streaming data with precision and efficiency.Containerized Python Workloads with ECS: Discover how to manage and deploy containerized Python applications on AWS, leveraging ECS for scalability and reliability.Data Orchestration with Airflow and Step Functions: Orchestrate complex workflows and automate data pipelines using the best-in-class tools for data orchestration.Event-Driven and Real-Time Processing with AWS Kinesis: Build robust, event-driven architectures and process streaming data in real-time, ensuring that your data pipelines are always up to date.Data Warehousing with Amazon Redshift: Explore the intricacies of Redshift, AWS's powerful data warehouse, to store and analyze massive datasets efficiently.Database Management with MySQL Aurora and DynamoDB: Get hands-on with relational and NoSQL databases, optimizing data storage and retrieval for different use cases.Serverless Data Processing with Lambda Functions: Harness the power of AWS Lambda to process data in real-time, triggering workflows based on events.Glue Python Shell Jobs for Python Workloads: Utilize Glue's Python shell jobs to run Python scripts in a managed environment, perfect for custom data processing tasks.Delta Lake on Spark: Understand the concepts behind Delta Lake and a lakehouse architecture, and how it enhances Spark for building reliable, scalable data lakes.CI/CD with GitHub Actions: Implement continuous integration and continuous delivery pipelines, automating your data engineering workflows with GitHub Actions.Why This Course?This bootcamp is not just another theoretical course - it's packed with real-world labs that simulate the challenges data engineers face daily. You'll get to build, deploy, and manage data pipelines and architectures that you can directly apply in your work or projects. Whether you're just starting out or looking to level up your skills, this course provides everything you need to become an AWS data engineering expert.