Data Engineering on AWS Vol 1 - OLAP & Data Warehouse

所在平台: Udemy

课程主页: https://www.udemy.com/course/data-engineering-vol-1-aws/

课程评论:没有评论

第一个写评论        关注课程

课程简介

**AWS 数据工程 Vol.1 - OLAP 与数据仓库** 本课程是 AWS 数据工程系列课程的第一部分,专注于数据仓库和数据存储层。您将深入了解 AWS 的关键数据工程服务,包括: * **S3 (Simple Storage Service)**:作为大规模数据存储的基础。 * **Redshift**:AWS 的高性能并行数据仓库服务,我们将深入探讨其查询性能调优、分布式与排序键、WLM (Workload Management)、ACID 事务、COPY 命令以及行级和列级安全。 * **Athena**:一种交互式查询服务,用于直接分析 S3 中的数据,我们将学习其分区和 WLM 配置。 * **Hive**:用于大规模数据集处理的分布式 SQL 查询引擎。 * **Glue Data Catalog**:用于存储和管理元数据的数据目录。 * **Lake Formation**:一种服务,用于安全地构建、保护和管理数据湖。 **课程亮点:** * **数据建模实践**:涵盖 OLTP 系统的规范化和 ER 图设计,以及 OLAP/DWH 系统的维度建模,并提供相关实践操作。 * **大规模数据集实战**:有机会使用 100GB - 300GB 或更多的数据集进行练习。 * **真实场景模拟**:课程包含的动手练习将涵盖 Redshift 查询性能调优、窗口函数、ACID 事务、COPY 命令、分布式与排序键、WLM、行级和列级安全、Athena 分区和 Athena WLM 等真实业务场景。 * **其他相关技术**:课程还将涉及 EC2、EBS、VPC 和 IAM 等基础 AWS 服务。 **后续课程 (Vol.2) 预告**: 本课程是数据工程全系列的第一部分。在第二部分中,我们将深入学习数据处理(批处理和流处理)服务,包括:Spark (AWS EMR, AWS Glue ETL, GCP Dataproc)、 Kafka (AWS & GCP)、Flink、Apache Airflow、Apache Pinot、AWS Kinesis 等。

课程评论(0条)

课程详情

This is Volume 1 of Data Engineering course on AWS. This course will give you detailed explanations on AWS Data Engineering Services like S3 (Simple Storage Service), Redshift, Athena, Hive, Glue Data Catalog, Lake Formation. This course delves into the data warehouse or consumption and storage layer of Data Engineering pipeline. In Volume 2, I will showcase Data Processing (Batch and Streaming) Services. You will get opportunities to do hands-on using large datasets (100 GB - 300 GB or more of data). Moreover, this course will provide you hands-on exercises that match with real-time scenarios like Redshift query performance tuning, streaming ingestion, Window functions, ACID transactions, COPY command, Distributed & Sort key, WLM, Row level and column level security, Athena partitioning, Athena WLM etc. Some other highlights:Contains training of data modelling - Normalization & ER Diagram for OLTP systems. Dimensional modelling for OLAP/DWH systems.Data modelling hands-on.Other technologies covered - EC2, EBS, VPC and IAM.This is Part 1 (Volume 1) of the full data engineering course. In Part 2 (Volume 2), I will be covering the following Topics.Spark (Batch and Stream processing using AWS EMR, AWS Glue ETL, GCP Dataproc)Kafka (on AWS & GCP)FlinkApache AirflowApache PinotAWS Kinesis and more.

课程标签

0人关注该课程

主题相关的课程