Microsoft Azure Databricks for Data Engineering

所在平台: Coursera

课程主页: https://www.coursera.org/learn/microsoft-azure-databricks-for-data-engineering

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:Microsoft Azure Databricks数据工程 概述:本课程旨在让您学习如何利用Apache Spark和在Azure Databricks平台上运行的强大集群,在云端执行大规模数据工程工作负载。您将探索Azure Databricks的功能及其Apache Spark笔记本,以处理庞大的文件,了解Azure Databricks平台,并识别适合使用Apache Spark的任务类型。课程将介绍Azure Databricks Spark集群和Spark作业的架构,并通过处理来自多个源的不同原始格式的大量数据,掌握Azure Databricks对日常数据处理功能(如读取、写入和查询)的支持。 本课程是一个专门为数据工程师和开发人员设计的系列课程的一部分,旨在帮助他们展示在利用Microsoft Azure数据服务设计和实施数据解决方案方面的专业知识,同时为准备DP-203考试(Microsoft Azure上的数据工程)的人士提供支持。您将参加一场实践考试,该考试涵盖了认证考试所测量的关键技能。 这是为准备这种考试而设计的10门课程程序中的第八门课程,能够提升您在利用Microsoft Azure数据服务设计和实施数据解决方案方面的专业能力。Microsoft Azure上的数据工程考试是证明您在整合、转换和汇总来自各种结构化和非结构化数据系统的数据知识和专业性的机会,为构建使用Microsoft Azure数据服务的分析解决方案提供适合的数据结构。每门课程教授的概念和技能均与该考试相关。 课程结束时,您将具备参加DP-203考试(Microsoft Azure上的数据工程)的准备。 课程大纲: 1. Azure Databricks介绍:描述Azure Databricks及其Apache Spark笔记本处理庞大文件的能力,了解Azure Databricks平台及其适合的任务类型,以及Azure Databricks Spark集群和Spark作业的架构。 2. 在Azure Databricks中读取和写入数据:描述如何使用Azure Databricks支持日常数据处理功能。 3. 在Azure Databricks中处理数据:通过定义DataFrames来读取和处理数据,了解数据转换和执行操作的区别。 4. 在Azure Databricks中处理DataFrame:应用列级转换、聚合等高级DataFrame操作。 5. Azure Databricks平台架构、安全性和数据保护:描述Azure Databricks平台架构和安全性,使用Azure密钥保管库存储所需的秘密。 6. Delta Lake:使用Delta Lake创建、追加和更新Apache Spark表。 7. 分析流数据并创建生产工作负载:使用Azure Databricks结构化流处理流数据。 8. 创建数据架构:将笔记本纳入版本控制,构建部署管道,集成Azure Databricks与Azure Synapse Analytics。 9. 数据工程实践考试:为Microsoft认证的Azure数据工程师助理考试做准备。 通过本课程的学习,您将为DP-203考试做好充分准备。

课程大纲

Name:Introduction to Azure Databricks

Description:Describe the capabilities of Azure Databricks and the Apache Spark notebook for processing huge files. Describe the Azure Databricks platform and identify the types of tasks well-suited for Apache Spark. Describe the architecture of an Azure Databricks Spark Cluster and Spark Jobs.

Name:Read and write data in Azure Databricks

Description:Describe how to use Azure Databricks supports day-to-day data-handling functions, such as reads, writes, and queries.

Name:Data processing in Azure Databricks

Description:Process data in Azure Databricks by defining DataFrames to read and process the Data. Perform data transformations in DataFrames and execute actions to display the transformed data. Explain the difference between a transform and an action, lazy and eager evaluations, Wide and Narrow transformations, and other optimizations in Azure Databricks.

Name:Work with DataFrames in Azure Databricks

Description:Use the DataFrame Column Class Azure Databricks to apply column-level transformations, such as sorts, filters and aggregations. Use advanced DataFrame functions operations to manipulate data, apply aggregates, and perform date and time operations in Azure Databricks.

Name:Platform architecture, security, and data protection in Azure Databricks

Description:Describe the Azure Databricks platform architecture and how it is securedUse Azure Key Vault to store secrets used by Azure Databricks and other services. Access Azure Storage with Key Vault-based secrets

Name:Delta Lake

Description:Describe how to use Delta Lake to create, append, and upsert data to Apache Spark tables, taking advantage of built-in reliability and optimizations. Describe Azure Databricks Delta Lake architecture

Name:Analyze streaming data and create production workloads

Description:Process streaming data with Azure Databricks structured streaming. Create production workloads on Azure Databricks with Azure Data Factory.

Name:Create a data architecture

Description:Describe how to put Azure Databricks notebooks under version control in an Azure DevOps repo and build deployment pipelines to manage your release process. Describe how to integrate Azure Databricks with Azure Synapse Analytics as part of your data architecture. Describe best practices for workspace administration, security, tools, integration, databricks runtime, HA/DR, and clusters in Azure Databricks

Name:Practice Exam on Data engineering with Azure Databricks

Description:Prepare for the Microsoft Certified: Azure Data Engineer Associate exam

课程评论(0条)

课程详情

In this course, you will learn how to harness the power of Apache Spark and powerful clusters running on the Azure Databricks platform to run large data engineering workloads in the cloud. You will discover the capabilities of Azure Databricks and the Apache Spark notebook for processing huge files. You will come to understand the Azure Databricks platform and identify the types of tasks well-suited for Apache Spark. You will also be introduced to the architecture of an Azure Databricks Spark Cluster and Spark Jobs. You will work with large amounts of data from multiple sources in different raw formats. you will learn how Azure Databricks supports day-to-day data-handling functions, such as reads, writes, and queries. This course is part of a Specialization intended for Data engineers and developers who want to demonstrate their expertise in designing and implementing data solutions that use Microsoft Azure data services for anyone interested in preparing for the Exam DP-203: Data Engineering on Microsoft Azure (beta). You will take a practice exam that covers key skills measured by the certification exam. This is the eighth course in a program of 10 courses to help prepare you to take the exam so that you can have expertise in designing and implementing data solutions that use Microsoft Azure data services. The Data Engineering on Microsoft Azure exam is an opportunity to prove knowledge expertise in integrating, transforming, and consolidating data from various structured and unstructured data systems into structures that are suitable for building analytics solutions that use Microsoft Azure data services. Each course teaches you the concepts and skills that are measured by the exam. By the end of this Specialization, you will be ready to take and sign-up for the Exam DP-203: Data Engineering on Microsoft Azure (beta).

课程标签

0人关注该课程

主题相关的课程