|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/batch-data-pipelines-gcp
课程评论:没有评论
课程名称:在GCP上构建批量数据管道 课程概述:数据管道通常属于额外加载、提取-加载-转换或提取-转换-加载等范式。该课程将介绍哪些范式适用于批量数据,并讨论何时使用这些范式。此外,课程还涵盖了Google Cloud上用于数据转换的多种技术,包括BigQuery、在Dataproc上执行Spark、Cloud Data Fusion中的管道图和使用Dataflow进行无服务器数据处理。学习者将通过Qwiklabs获得在Google Cloud上构建数据管道组件的实践经验。 课程大纲: 1. **引言** - 描述:本模块介绍课程内容和日程安排。 2. **批量数据管道构建概论** - 描述:本模块回顾数据加载的不同方法:EL、ELT和ETL,以及何时使用它们。 3. **在Dataproc上执行Spark** - 描述:本模块展示如何在Dataproc上运行Hadoop,如何利用Cloud Storage,以及如何优化Dataproc作业。 4. **使用Dataflow进行无服务器数据处理** - 描述:本模块讲解如何使用Dataflow构建数据处理管道。 5. **使用Cloud Data Fusion和Cloud Composer管理数据管道** - 描述:本模块展示如何使用Cloud Data Fusion和Cloud Composer管理数据管道。 6. **课程总结** - 描述:课程总结 módulos para revisar todo el programa y los aprendizajes obtenidos.
Name:Introduction
Description:In this module, we introduce the course and agenda
Name:Introduction to Building Batch Data Pipelines
Description:This module reviews different methods of data loading: EL, ELT and ETL and when to use what
Name:Executing Spark on Dataproc
Description:This module shows how to run Hadoop on Dataproc, how to leverage Cloud Storage, and how to optimize your Dataproc jobs.
Name:Serverless Data Processing with Dataflow
Description:This module covers using Dataflow to build your data processing pipelines
Name:Manage Data Pipelines with Cloud Data Fusion and Cloud Composer
Description:This module shows how to manage data pipelines with Cloud Data Fusion and Cloud Composer.
Name:Course Summary
Description:Course Summary
Data pipelines typically fall under one of the Extra-Load, Extract-Load-Transform or Extract-Transform-Load paradigms. This course describes which paradigm should be used and when for batch data. Furthermore, this course covers several technologies on Google Cloud for data transformation including BigQuery, executing Spark on Dataproc, pipeline graphs in Cloud Data Fusion and serverless data processing with Dataflow. Learners will get hands-on experience building data pipeline components on Google Cloud using Qwiklabs.