|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/mastering-databricks-apache-spark-build-etl-data-pipeline/
课程评论:没有评论
**课程名称:**精通 Databricks & Apache Spark - 构建 ETL 数据管道 **课程概述:** 本课程将深入介绍 Databricks 的湖屋架构,结合数据仓库和数据湖的优势。您将学习如何使用 Scala、Python 和 Spark SQL 进行各种数据操作,从而构建有价值的数据解决方案,并掌握在不同语言中构建批处理的能力。通过本课程,您将能够根据客户需求,使用不同的编程语言编写相同的命令,交付世界级的解决方案。我们将以 Azure Databricks 为平台,构建端到端的数据管道。 **核心学习要点:** * **集群构建与数据处理:** 学习如何构建自己的集群来处理数据。 * **数据加载:** 通过一键式操作,将不同来源的数据加载到 Azure SQL 和 Delta 表中。 * **仪表板构建与分析:** 利用 Databricks Notebook 准备仪表板,以回答业务问题。 * **Azure 云基础架构部署:** 根据需求,学习在 Azure 云上部署基础架构。 * **360 度云平台视角:** 通过这些实践场景,全面了解云平台以及如何配置各种资源。 * **Azure Databricks 全方位实践:** 所有活动均在 Azure Databricks 中完成。 **课程重点内容:** * **Databricks 基础:** 掌握 Databricks 的核心概念。 * **Delta 表:** 深入理解 Delta 表的特性,包括版本控制和 vacuum 操作。 * **Apache Spark SQL:** * 数据过滤 * DataFrame 重命名、删除、选择和类型转换 * 聚合操作(SUM, AVERAGE, MAX, MIN) * 窗口函数(Rank, Row Number, Dense Rank) * **仪表板构建与分析。** **目标学员:** 数据工程师、BI 架构师、数据分析师、ETL 开发人员、BI 经理。
Welcome to the course on Mastering Databricks & Apache spark -Build ETL data pipelineDatabricks combines the best of data warehouses and data lakes into a lakehouse architecture. In this course we will be learning how to perform various operations in Scala, Python and Spark SQL. This will help every student in building solutions which will create value and mindset to build batch process in any of the language. This course will help in writing same commands in different language and based on your client needs we can adopt and deliver world class solution. We will be building end to end solution in azure databricks.Key Learning PointsWe will be building our own cluster which will process our data and with one click operation we will load different sources data to Azure SQL and Delta tablesAfter that we will be leveraging databricks notebook to prepare dashboard to answer business questionsBased on the needs we will be deploying infrastructure on Azure cloudThese scenarios will give student 360 degree exposure on cloud platform and how to step up various resourcesAll activities are performed in Azure DatabricksFundamentalsDatabricksDelta tablesConcept of versions and vacuum on delta tablesApache Spark SQLFiltering DataframeRenaming, drop, Select, CastAggregation operations SUM, AVERAGE, MAX, MINRank, Row Number, Dense RankBuilding dashboardsAnalyticsThis course is suitable for Data engineers, BI architect, Data Analyst, ETL developer, BI Manager