Azure Databricks & Spark For Data Engineers:Hands-on Project

所在平台: Udemy

课程主页: https://www.udemy.com/course/azure-databricks-spark-core-for-data-engineers/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:Azure Databricks与Spark数据工程师:实战项目 课程概述:本课程旨在帮助学生学习Azure Databricks和Spark核心的实用数据工程工具。课程从基本概念开始,通过分析和报告Formula 1赛车数据的真实项目,逐步引导学员实施数据工程解决方案。课程内容经过多次更新以确保与最新技术和推荐保持一致,主要包括以下几点更新: - **Unity Catalog的引入**:新增的25、26和27节详细介绍了Unity Catalog,它为数据湖屋提供统一的数据治理解决方案,并通过一个项目实施进行教学。 - **对Azure Data Lake的访问**:新增的6和7节及更新的第8节,反映了最新的Databricks关于访问Azure Data Lake的建议,适用于Azure学生或企业订阅用户的更好项目解决方案。 - **用户界面的更新**:第3、4和5节更新了关于Azure Databricks的用户界面变化,并加入了Databricks集群的新功能。 课程内容主要包括以下方面: 1. **Azure Databricks基础**:学习如何构建数据工程解决方案的架构,包括Azure Data Lake Gen2、Azure Data Factory和Power BI。 2. **工作与Databricks笔记本**:使用Databricks笔记本及实用工具,创建和配置Databricks集群。 3. **使用Delta Lake**:实现湖屋架构的解决方案,并创建可视化仪表盘。 4. **PySpark与Spark SQL**:处理数据源,执行数据转化,包括过滤、连接、聚合等操作。 5. **Unity Catalog**:掌握数据治理及Meta存储的创建与配置。 6. **Azure Data Factory**:学习如何设计和监控数据管道,包括调度和依赖管理。 尽管课程并不专门针对Azure数据工程师助理认证考试DP203,但可以帮助学生获取必要的技能,同时也为Databricks认证数据工程师的考试做准备。课程以快节奏和简洁明了的英语授课,让学员能够快速上手并最终成为Azure Databricks的熟练使用者。完成所有任务后,学生能够独立开展真实数据工程项目。

课程评论(0条)

课程详情

Major updates to the course since the launchUpdate 3 - New sections 25, 26 and 27 added to include Unity Catalog. Unity Catalog is a recent addition to Databricks which offers unified data governance solution for a Data Lakehouse. These sections cover all aspects of Unity Catalog and the implementation using a project. Update 2 - New sections 6 and 7 added. Section 8 Updated. These changes are to reflect latest Databricks recommendations around accessing Azure Data Lake. Also, this provides a better solution to complete the course project for students using Azure Student Subscription or Corporate Subscriptions with limited access to Azure Active Directory. Update 1 - Sections 3, 4 & 5 updated to reflect recent UI changes to Azure Databricks. Also included lessons on additional functionality included by Databricks recently to Databricks clusters.. Welcome! I am looking forward to helping you with learning one of the in-demand data engineering tools in the cloud, Azure Databricks! This course has been taught with implementing a data engineering solution using Azure Databricks and Spark core for a real world project of analysing and reporting on Formula1 motor racing data.This is like no other course in Udemy for Azure Databricks. Once you have completed the course including all the assignments, I strongly believe that you will be in a position to start a real world data engineering project on your own and also proficient on Azure Databricks. I have also included lessons on Azure Data Lake Storage Gen2, Azure Data Factory as well as PowerBI. The primary focus of the course is Azure Databricks and Spark core, but it also covers the relevant concepts and connectivity to the other technologies mentioned. Please note that the course doesn't cover other aspects of Spark such as Spark streaming and Spark ML. Also the course has been taught using PySpark as well as Spark SQL; It doesn't cover Scala or Java. The course follows a logical progression of a real world project implementation with technical concepts being explained and the Databricks notebooks being built at the same time. Even though this course is not specifically designed to teach you the skills required for passing the Azure Data Engineer Associate Certification Exam DP203, it can greatly help you get most of the necessary skills required for the exam. Similarly, the course teaches the skills required to pass the Databricks Certified Data Engineer Associate Certification. I value your time as much as I do mine. So, I have designed this course to be fast-paced and to the point. Also, the course has been taught with simple English and no jargons. I start the course from basics and by the end of the course you will be proficient in the technologies used. Currently the course teaches you the followingAzure DatabricksBuilding a solution architecture for a data engineering solution using Azure Databricks, Azure Data Lake Gen2, Azure Data Factory and Power BICreating and using Azure Databricks service and the architecture of Databricks within AzureWorking with Databricks notebooks as well as using Databricks utilities, magic commands etcPassing parameters between notebooks as well as creating notebook workflowsCreating, configuring and monitoring Databricks clusters, cluster pools and jobsMounting Azure Storage in Databricks using secrets stored in Azure Key VaultWorking with Databricks Tables, Databricks File System (DBFS) etcUsing Delta Lake to implement a solution using Lakehouse architectureCreating dashboards to visualise the outputsConnecting to the Azure Databricks tables from PowerBISpark (Only PySpark and SQL)Spark architecture, Data Sources API and Dataframe APIPySpark - Ingestion of CSV, simple and complex JSON files into the data lake as parquet files/ tables. PySpark - Transformations such as Filter, Join, Simple Aggregations, GroupBy, Window functions etc.PySpark - Creating local and temporary viewsSpark SQL - Creating databases, tables and viewsSpark SQL - Transformations such as Filter, Join, Simple Aggregations, GroupBy, Window functions etc.Spark SQL - Creating local and temporary viewsImplementing full refresh and incremental load patterns using partitionsDelta LakeEmergence of Data Lakehouse architecture and the role of delta lake.Read, Write, Update, Delete and Merge to delta lake using both PySpark as well as SQL History, Time Travel and VacuumConverting Parquet files to Delta filesImplementing incremental load pattern using delta lakeUnity CatalogOverview of Data Governance and Unity CatalogCreate Unity Catalog Metastore and enable a Databricks workspace with Unity CatalogOverview of 3 level namespace and creating Unity Catalog objectsConfiguring and accessing external data lakes via Unity CatalogDevelopment of mini project using unity catalog and seeing the key data governance capabilities offered by Unity Catalog such as Data Discovery, Data Audit, Data Lineage and Data Access Control.Azure Data FactoryCreating pipelines to execute Databricks notebooksDesigning robust pipelines to deal with unexpected scenarios such as missing filesCreating dependencies between activities as well as pipelinesScheduling the pipelines using data factory triggers to execute at regular intervalsMonitor the triggers/ pipelines to check for errors/ outputs.

课程标签

0人关注该课程

主题相关的课程