Master Data Engineering using Azure Data Analytics

所在平台: Udemy

课程主页: https://www.udemy.com/course/data-engineering-using-azure-data-analytics/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:使用Azure数据分析掌握数据工程 课程概述:该课程主要关注数据工程,即构建数据管道以将数据从多个来源导入数据湖或数据仓库,并从数据湖或数据仓库将数据传输到下游系统。课程将带领学员使用Azure数据分析栈构建数据工程管道,涵盖Azure Storage(包括Blob和ADLS)、ADF数据流、ADF管道、Azure SQL、Azure Synapse、Azure Databricks等服务。 课程内容包括: 1. 环境设置:学员将学习如何在Windows和Mac上使用VS Code进行学习,并注册Azure Portal,提供注册说明和USD 200的信用额度(有效期一个月)。 2. 使用Azure Storage作为数据湖:学习如何管理Azure Storage中的文件,使用Azure Storage Explorer等工具。 3. ADF(Azure数据工厂)用于ETL和编排:学习如何使用ADF数据流进行ETL以及设置链接服务和数据集。 4. 创建管道和性能调优:构建ADF管道进行编排,了解参数化和基准加载,同时掌握使用ADF管道的性能调优技巧。 5. Azure SQL设置与查询:学习如何设置Azure SQL,创建所需的数据库表并运行查询。 6. 数据复制与Azure Synapse:学习如何使用ADF数据复制将数据导入数据库表,并概述Azure Synapse的无服务器及专用池设置,建立用于ETL的专用池。 7. Azure Databricks与大数据处理:学习如何设置Azure Databricks,与ADLS集成管理机密,了解Spark SQL和Pyspark Data Frame API,并构建基于Pyspark和Spark SQL的ELT管道。 8. ADF与Databricks笔记本的编排:掌握如何创建ADF管道以编排Databricks笔记本。 本课程为希望掌握Azure数据分析中的数据工程技能的学员提供了全面、系统的学习路径。

课程评论(0条)

课程详情

Data Engineering is all about building Data Pipelines to get data from multiple sources into Data Lakes or Data Warehouses and then from Data Lakes or Data Warehouses to downstream systems. As part of this course, I will walk you through how to build Data Engineering Pipelines using Azure Data Analytics Stack. It includes services such as Azure Storage (both Blob and ADLS), ADF Data Flow, ADF Pipeline, Azure SQL, Azure Synapse, Azure Databricks, and many more.As part of this course, first, you will go ahead and set up the environment to learn using VS Code on Windows and Mac.Once the environment is ready, you need to sign up for Azure Portal. We will provide all the instructions to sign up for Azure Portal Account including reviewing billing as well as getting USD 200 Credit valid for up to a month.We typically use Azure Storage as Data Lake. As part of this course, you will learn how to use Azure Storage as Data Lake along with how to manage the files in Azure Storage using tools such as Azure Storage Explorer. ADF (Azure Data Factory) is used for both ETL as well as Orchestration. First, you will understand how to perform ETL using ADF Data Flow. The source and target will be Files in Azure Storage Account. As part of this process, you will also learn how to set up Linked Services and Data Sets in ADF (Azure Data Factory).Once ADF Data Flow is ready, you will go ahead and build Pipeline for Orchestration using ADF Pipeline. You will also learn how to parameterize and also how to take care of baseline load.You will also understand key performance tuning techniques using ADF Pipeline such as controlling the number of partitions, custom integration runtimes (IR), etc.Azure provides RDBMS as different services for Postgres, SQL Server, etc. You will learn how to set up Azure SQL Once the Azure SQL is set up, you will also understand how to create required tables and run queries against them.ADF provides ADF Data Copy to copy data from different sources and different targets. Once the Database tables are ready you will use ADF Data Copy to copy data into the tables.Azure provides Synapse Analytics for Data Warehouse. You will get an overview of both serverless as well as dedicated pools. You will end up setting up a Dedicated Pool for ETL using ADF. Once Azure SQL and Azure Synapse are ready, you will build ETL Pipeline using ADF Data Flow and Orchestrate using ADF Pipeline.Azure Databricks is the service for Big Data Processing using Spark Engine. You will learn how to set up Azure Databricks, integrate with ADLS, and also managing secrets.You will also get an overview of Spark SQL and Pyspark Data Frame APIs using Azure Databricks.You will also build ELT Pipeline using Databricks Jobs and Workflows where tasks are defined based on Pyspark as well as Spark SQL.You will also understand how to build ADF Pipelines to orchestrate Databricks Notebooks.

课程标签

0人关注该课程

主题相关的课程