|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/build-a-secure-data-lake-in-aws-using-aws-lake-formation/
课程评论:没有评论
**课程名称:** 使用 AWS Lake Formation 在 AWS 中构建安全的数据湖 **课程概述:** 本课程将指导您使用 AWS Lake Formation 构建数据湖,并通过 Amazon Redshift 将数据仓库功能融入其中,形成湖仓一体(Lakehouse)架构。您将学习如何利用 Lake Formation 收集和编目不同数据源的数据,将其迁移到 S3 数据湖,并进行数据清洗和分类。课程将遵循真实项目实施的逻辑流程,提供实践经验,包括设置数据湖、创建数据管道进行数据摄取、以及为分析和报告准备数据转换。 **课程内容:** * **第一章:利用 Lake Formation 设置数据湖** * 创建不同的数据源(MySQL RDS 和 Kinesis)。 * 通过设置 Lake Formation 的蓝图和工作流作业,从 MySQL RDS 数据源摄取数据到数据湖。 * 使用爬虫程序对数据库进行编目。 * 使用托管表(Governed Tables)进行访问控制和安全管理。 * 使用 Athena 查询数据湖中的数据。 * **第二章:数据探索与清洗** * 探索 AWS Glue DataBrew,在执行复杂 ETL 之前对数据进行分析和理解。 * 创建数据处理配方(Recipes),利用各种转换来操纵数据湖中的数据。 * 清洗和规范化数据。 * 运行作业以将配方应用于所有新数据或大型数据集。 * **第三章:ETL 作业实现** * 介绍 Glue Studio。 * 编写和监控 ETL 作业,以转换数据并将数据在数据湖的不同区域之间移动。 * 创建 DynamoDB 数据源,并通过 AWS Glue 将数据摄取到数据湖。 * **第四章:湖仓一体架构** * 介绍并创建 Redshift 集群,为数据湖引入数据仓库功能,形成湖仓一体架构。 * 创建 ETL 作业,将数据从数据湖迁移到数据仓库以进行分析。 * 使用 Redshift Spectrum 直接查询 S3 数据湖中的数据,无需复制数据或基础设施。 * **第五章:数据安全与隐私** * 介绍 Amazon Macie,用于管理数据安全和数据隐私。 * 确保随着数据湖的增长,能够持续大规模地识别敏感数据。
In this course, we will be creating a data lake using AWS Lake Formation and bring data warehouse capabilites to the data lake to form the lakehouse architecture using Amazon Redshift. Using Lake Formation, we also collect and catalog data from different data sources, move the data into our S3 data lake, and then clean and classify them.The course will follow a logical progression of a real world project implementation with hands on experience of setting up a data lake, creating data pipelines for ingestion and transforming your data in preparation for analytics and reporting.Chapter 1Setup the data lake using lake formationCreate different data sources (MySQL RDS and Kinesis)Ingest data from the MYSQL RDS data source into the data lake by setting up blueprint and workflow jobs in lake formationCatalog our Database using crawlersUse governed tables for managing access control and securityQuery our data lake using AthenaChapter 2,Explore the use of AWS Gluw DataBrew for profiling and understanding our data before we starting performing complex ETL jobs.Create Recipes for manipulating the data in our data lake using different transformationsClean and normalise dataRun jobs to apply the recipes on all new data or larger datasetsChapter 3Introduce Glue StudioAuthor and monitor ETL jobs for tranforming our data and moving them between different zone of our data lakeCreate a DynamoDB source and ingest data into our data lake using AWS GlueChapter 4Introduce and create a redshift cluster to bring datawarehouse capabilities to our data lake to form the lakehouse architectureCreate ETL jobs for moving data from our lake into the warehouse for analyticsUse redshift spectrum to query against data in our S3 data lake without the need for duplicating data or infrastructureChapter 5Introduce Amazon Macie for managing data security and data privacy and ensure we can continue to identify sensitive data at scale as our data lake grows