Data Engineering with Spark Databricks Delta Lake Lakehouse

所在平台: Udemy

课程主页: https://www.udemy.com/course/data-engineering-with-spark-databricks-delta-lake-lakehouse/

课程评论:没有评论

第一个写评论        关注课程

课程简介

Coursera 数据工程课程:Spark、Databricks 和 Delta Lake 湖仓一体 本课程是为数据工程初学者设计的,旨在教授如何使用 Apache Spark 在 Databricks 的湖仓一体 (Lakehouse) 架构上构建数据管道。课程涵盖使用 Python 和 Scala 进行 Spark 分析,以及 Spark SQL 和 Databricks SQL 的应用。 **核心学习内容包括:** * **Databricks 社区版入门:** 熟悉 Databricks 平台并创建基本的 Spark 数据管道。 * **Spark 分析:** 学习使用 Python 和 Scala 进行 Spark 转换 (transformations)、动作 (actions)、连接 (joins),以及 Spark SQL 和 DataFrame API。 * **Delta Lake 操作:** 掌握 Delta 表的管理,包括访问版本历史、数据恢复和时间旅行 (time travel) 功能。 * **性能优化:** 学习使用 Delta Cache 优化查询性能。 * **数据管道构建:** 使用 Delta Live Tables 构建 ETL 数据管道。 * **AWS 集成(可选):** 学习如何在 AWS 上设置 Databricks 账户并运行 notebooks。 **课程结构:** 课程从 Databricks 基础操作开始,逐步深入到 Spark 分析和 Delta Lake 的高级特性。通过动手实践和真实场景示例,学员将获得构建可扩展、可靠数据管道的实践经验,为胜任数据工程师职位打下坚实基础。 **先修要求:** 无需 Python 和 Scala 知识,但具备数据库和 SQL 基础会更有帮助。 **课程目标:** 完成后,学员将能够理解 Spark 和湖仓一体概念,并利用 Databricks 湖仓一体架构,使用 Apache Spark 构建高效的数据管道。

课程评论(0条)

课程详情

Data Engineering is a vital component of modern data-driven businesses. The ability to process, manage, and analyze large-scale data sets is a core requirement for organizations that want to stay competitive. In this course, you will learn how to build a data pipeline using Apache Spark on Databricks' Lakehouse architecture. This will give you practical experience in working with Spark and Lakehouse concepts, as well as the skills needed to excel as a Data Engineer in a real-world environment.Throughout the Course, You Will Learn:Conducting analytics using Python and Scala with Spark.Applying Spark SQL and Databricks SQL for analytics.Developing a data pipeline with Apache Spark.Becoming proficient in Databricks' community edition.Managing a Delta table by accessing version history, restoring data, and utilizing time travel features.Optimizing query performance using Delta Cache.Working with Delta Tables and Databricks File System.Gaining insights into real-world scenarios from experienced instructors.Course Structure:Beginning with familiarizing yourself with Databricks' community edition and creating a basic pipeline using Spark.Progressing to more complex topics after gaining comfort with the platform.Learning analytics with Spark using Python and Scala, including Spark transformations, actions, joins, Spark SQL, and DataFrame APIs.Acquiring the knowledge and skills to operate a Delta table, including accessing its version history, restoring data, and utilizing time travel functionality using Spark and Databricks SQL.Understanding how to use Delta Cache to optimize query performance.Optional Lectures on AWS Integration:'Setting up Databricks Account on AWS' and 'Running Notebooks Within a Databricks AWS Account.'Building an ETL pipeline with Delta Live TablesProviding additional opportunities to explore Databricks within the AWS ecosystem.This course is designed for Data Engineering beginners with no prior knowledge of Python and Scala required. However, some familiarity with databases and SQL is necessary to succeed in this course. Upon completion, you will have the skills and knowledge required to succeed in a real-world Data Engineer role.Throughout the course, you will work with hands-on examples and real-world scenarios to apply the concepts you learn. By the end of the course, you will have the practical experience and skills required to understand Spark and Lakehouse concepts, and to build a scalable and reliable data pipeline using Spark on Databricks' Lakehouse architecture.

课程标签

0人关注该课程

主题相关的课程