Building Batch Data Pipelines on GCP

所在平台: CourseraArchive

课程类别: 其他类别

大学或机构: CourseraNew

课程主页: https://www.coursera.org/archive/batch-data-pipelines-gcp

课程评论:没有评论

第一个写评论        关注课程

课程大纲

Executing Spark on Cloud Dataproc
Summary

课程评论(0条)

课程详情

Data pipelines typically fall under one of the Extra-Load, Extract-Load-Transform or Extract-Transform-Load paradigms. This course describes which paradigm should be used and when for batch data. Furthermore, this course covers several technologies on Google Cloud Platform for data transformation including BigQuery, executing Spark on Cloud Dataproc, pipeline graphs in Cloud Data Fusion and serverless data processing with Cloud Dataflow. Learners will get hands-on experience building data pipeline components on Google Cloud Platform using QwikLabs.

在GCP上构建批处理数据管道:数据管道通常属于Extra-Load,Extract-Load-Transform或Extract-Transform-Load范式之一。本课程描述应该使用哪种范例以及何时使用批处理数据。此外,本课程涵盖了Google Cloud Platform上用于数据转换的多种技术,包括BigQuery,在Cloud Dataproc上执行Spark,Cloud Data Fusion中的管道图以及使用Cloud Dataflow进行无服务器数据处理。学习者将获得使用QwikLabs在Google Cloud Platform上构建数据管道组件的动手经验。

课程标签

0人关注该课程

主题相关的课程