|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/introduction-to-designing-data-lakes-in-aws
课程评论:没有评论
课程名称:AWS数据湖设计入门 课程概述:在本课程中,我们将帮助您了解如何以安全和可扩展的方式创建和运营数据湖,而无需具备数据科学的先验知识!课程将从数据湖的必要性开始,探讨其价值主张、特点和组成部分。由于数据规模和增长的挑战,设计数据湖是复杂的,开发人员需要了解最佳实践,以避免难以纠正的常见错误。课程覆盖数据湖的基础知识、如何将数据摄取和组织到数据湖中,并深入讨论优化大规模数据消费时性能和成本的数据处理。 此课程适合需要设计和构建安全、可扩展的数据湖架构的专业人士(架构师、系统管理员和DevOps)。学生将了解数据湖的用例,并将其与传统的服务器和存储基础设施进行对比。 课程大纲: 1. 模块1:数据湖简介 - 概述数据湖的目的及其与数据仓库的区别,涵盖数据湖的组件和架构。 2. 模块2:数据摄取、目录管理和准备 - 关注数据摄取到数据湖的过程、数据目录管理和分析准备,涵盖存储、摄取方法、数据格式、分区、压缩及使用Amazon Athena查询数据等主题。 3. 模块3:使用AWS Lake Formation构建数据湖 - 介绍AWS Lake Formation服务的基本权限模型及其功能,帮助构建和管理AWS上的数据湖。 4. 模块4:数据处理与分析 - 涉及数据转换技术和工具(如AWS Glue)来处理和分析数据湖中的数据,包括实践演示和关于Glue和Athena联邦查询的技术讲座。 5. 模块5:AWS Lake Formation的额外配置和能力 - 探索AWS Lake Formation的高级功能和配置,包括蓝图、工作流和细粒度访问控制,同时介绍如何使用Amazon QuickSight进行数据可视化。 6. 模块6:AWS上的现代数据架构 - 介绍现代数据架构的概念及其在AWS上的实施,涵盖数据移动场景、数据共享模型和相关阅读材料。 这个课程将为您提供设计和管理数据湖所需的知识和技能,使您能够在AWS上构建高效、安全的数据解决方案。
Name:Module 1: Introduction to Data Lakes
Description:This module provides an overview of data lakes, their purpose, and how they differ from data warehouses. It also covers the components and architectures involved in data lakes.
Name:Module 2: Data ingestion, cataloging, and preparation
Description:This module focuses on the processes of ingesting data into a data lake, cataloging the data, and preparing it for analysis. It covers topics such as data lake storage, data ingestion methods, crawling and cataloging data, data formatting, partitioning, compression, and querying data with Amazon Athena.
Name:Module 3: Building a data lake with AWS Lake Formation
Description:This module introduces AWS Lake Formation, a service that helps build and manage data lakes on AWS. It covers the basic permission model, and provides an overview of the service’s features and capabilities.
Name:Module 4: Data processing and analytics
Description:This module covers data transformation techniques and tools like AWS Glue for processing and analyzing data in the data lake. It includes hands-on demos and a technical talk on Glue and Athena Federated Queries.
Name:Module 5: AWS Lake Formation additional configurations and capabilities
Description:This module explores advanced features and configurations of AWS Lake Formation, including blueprints, workflows, and fine-grained access control. It also covers data visualization with Amazon QuickSight.
Name:Module 6: Modern data architecture on AWS
Description:This module introduces the concept of modern data architecture and its implementation on AWS. It covers data movement scenarios, data sharing models, and relevant readings.
In this class, Introduction to Designing Data Lakes on AWS, we will help you understand how to create and operate a data lake in a secure and scalable way, without previous knowledge of data science! Starting with the "WHY" you may want a data lake, we will look at the Data-Lake value proposition, characteristics and components. Designing a data lake is challenging because of the scale and growth of data. Developers need to understand best practices to avoid common mistakes that could be hard to rectify. In this course we will cover the foundations of what a Data Lake is, how to ingest and organize data into the Data Lake, and dive into the data processing that can be done to optimize performance and costs when consuming the data at scale. This course is for professionals (Architects, System Administrators and DevOps) who need to design and build an architecture for secure and scalable Data Lake components. Students will learn about the use cases for a Data Lake and, contrast that with a traditional infrastructure of servers and storage.