|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-engineering-with-google-datafusion-and-big-query-cdap/
课程评论:没有评论
课程名称:使用 Google Datafusion 和 BigQuery 的数据工程 (CDAP) 课程概述: 本课程是关于 Google Cloud 的低代码数据摄取工具 Google Data Fusion 的入门课程。Google Data Fusion 是一个完全托管的数据集成平台,旨在帮助数据工程师高效创建、部署和管理数据管道。使用 Google Data Fusion 的主要原因之一是其易用性。通过直观的可视化界面,数据工程师能够在无需大量编码的情况下创建复杂的数据管道。其拖放界面简化了数据转化和清洗的过程,使专业人员可以专注于业务逻辑,而不必过多担心细节编码。 Google Data Fusion 的另一个显著优势是其可扩展性。该平台运行在 Google Cloud 上,可以处理大量数据以及高性能并行处理。数据工程师可以根据项目需要垂直或水平扩展处理能力,从而确保能够应对任何规模的数据需求。此外,Google Data Fusion 与 Google Cloud 生态系统中的其他服务和产品无缝集成,数据工程师可以轻松连接并集成数据管道与 BigQuery、Cloud Storage、Pub/Sub 等服务。这使得数据摄取、存储和分析跨多个平台变得更加统一和高效。 在本课程中,您将学习: - 理解 Google Data Fusion 的内部工作原理 - 掌握其优势 - 如何创建 Data Fusion 实例 - 使用 Google Cloud Storage 作为数据输入 - 使用 BigQuery 作为数据湖(铜层和银层) - BigQuery 的高级特性:分区表和 MERGE 命令 - 从不同来源获取数据 - 使用低代码工具 Wrangle 和查询来转换数据 - 创建数据 ETL(提取、转换、加载)和依赖关系的有向无环图(DAG) - 计划和管理不同 DAG 之间的依赖关系 通过本课程,您将深入了解数据工程的基础知识,并学习如何利用 Google Cloud 的强大工具来处理和分析数据。
This is an INTRODUCTORY course to Google Cloud's low-code ingestion tool, Datafusion. Google Data Fusion is a fully managed data integration platform that allows data engineers to efficiently create, deploy, and manage data pipelines.One of the main reasons to use Google Data Fusion is its ease of use. With an intuitive and visual interface, data engineers can create complex data pipelines without the need for extensive coding. The drag-and-drop interface simplifies the process of data transformation and cleansing, allowing professionals to focus on business logic rather than worrying about detailed coding.Another significant benefit of Google Data Fusion is its scalability. The platform runs on Google Cloud, which means it can handle large volumes of data and high-performance parallel processing. Data engineers can vertically or horizontally expand their processing capabilities according to project needs, ensuring they can handle any data demand at scale.Furthermore, Google Data Fusion seamlessly integrates with other services and products in the Google Cloud ecosystem. Data engineers can easily connect and integrate data pipelines with services such as BigQuery, Cloud Storage, Pub/Sub, and many others. This enables a cohesive and unified data architecture, facilitating data ingestion, storage, and analysis across multiple platforms.In this course, you will learn:Understanding its internal workings.What its benefits are.How to create a Datafusion instance.Using Google Cloud Storage as data input.Using BigQuery as a Data Lake (Bronze and Silver layers).Advanced features of BigQuery: Partitioned tables and MERGE command.Ingesting data from different sources.Transforming data with Wrangle (low code) and queries.Creating DAGs for data ETL (Extract, Transform, Load) and dependencies.Scheduling and inter-DAG dependencies.