|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/big-data-hadoop-spark-project/
课程评论:没有评论
课程名称:绝对初学者的大数据 Hadoop 和 Spark 项目(2025 版) 概述:本课程将为您准备进入实际的数据工程师角色!数据工程是数据驱动型组织的重要组成部分,它包括大规模数据集的处理、管理和分析,这是保持竞争力的关键。该课程提供了一个快速入门大数据的机会,通过使用免费的云集群来解决实际用例。您将学习 Hadoop、Hive 和 Spark 的基本概念,使用 Python 和 Scala 进行编码。课程旨在将您的 Spark Scala 和 PySpark 编码能力提升到专业开发者水平,同时介绍行业标准的编码实践,如日志记录、错误处理和配置管理。此外,您将了解 Databricks Lakehouse 平台,并学习如何使用 Python 和 Scala 与 Spark 进行分析,运用 Spark SQL 和 Databricks SQL 进行数据分析,开发 Apache Spark 数据管道,并通过访问版本历史、恢复数据和利用时间旅行功能管理 Delta 表。您还将学习如何使用 Delta Cache 优化查询性能,处理 Delta 表和 Databricks 文件系统,并从经验丰富的讲师那里获取现实场景的见解。 学习内容包括: - 大数据和 Hadoop 概念 - 如何使用 Google Dataproc 创建免费的 Hadoop 和 Spark 集群 - Hadoop 实操 - HDFS、Hive - Python 基础知识 - PySpark RDD 实操 - PySpark SQL、DataFrame 实操 - 使用 PySpark 和 Hive 的项目工作 - Scala 基础知识 - Spark Scala DataFrame - 使用 Spark Scala 的项目工作 - 通过实践理解 Databricks Delta Lake Lakehouse 概念 - 操作 Delta 表,包括访问版本历史、恢复数据和利用时间旅行功能 - 使用 Winutil、Maven 和 IntelliJ 进行 Spark Scala 的真实世界编码框架与开发 - 使用 PyCharm 进行 Python Spark Hadoop Hive 的编码框架与开发 - 使用 Hive、PostgreSQL、Spark 构建数据管道,日志记录、错误处理和 PySpark 与 Spark Scala 应用的单元测试 - 利用 Glue 在 AWS S3 上对数据应用 Spark 转换,并通过 Athena 查看数据 - 如何利用 ChatGPT 成为高效的数据工程师 先决条件:该课程是为没有 Python 和 Scala 背景的数据工程初学者设计的。但是,对数据库和 SQL 有一定的了解将有助于您成功完成此课程。完成课程后,您将具备在实际数据工程师角色中取得成功的技能和知识。
2025 EditionThis course will prepare you for a real world Data Engineer role! Data Engineering is a crucial component of data-driven organizations, as it encompasses the processing, management, and analysis of large-scale data sets, which is essential for staying competitive.This course provides an opportunity to quickly get started with Big Data through the use of a free cloud clusters, and solve a practical use case. You will learn the fundamental concepts of Hadoop, Hive, and Spark, using both Python and Scala. The course aims to develop your Spark Scala and PySpark coding abilities to that of a professional developer, by introducing you to industry-standard coding practices such as logging, error handling and configuration management.Additionally, you will understand the Databricks Lakehouse Platform and learn how to conduct analytics using Python and Scala with Spark, apply Spark SQL and Databricks SQL for analytics, develop a data pipeline with Apache Spark, and manage a Delta table by accessing version history, restoring data, and utilizing time travel features. You will also learn how to optimize query performance using Delta Cache, work with Delta Tables and Databricks File System, and gain insights into real-world scenarios from our experienced instructor.What you will learn:Big Data, Hadoop conceptsHow to create a free Hadoop and Spark cluster using Google DataprocHadoop hands-on - HDFS, HivePython basicsPySpark RDD - hands-onPySpark SQL, DataFrame - hands-onProject work using PySpark and HiveScala basicsSpark Scala DataFrameProject work using Spark ScalaDeveloping a practical comprehension of Databricks Delta Lake Lakehouse concepts through hands-on experience.Learning to operate a Delta table by accessing its version history, recovering data, and utilizing time travel functionalitySpark Scala Real world coding framework and development using Winutil, Maven and IntelliJ. Python Spark Hadoop Hive coding framework and development using PyCharmBuilding a data pipeline using Hive , PostgreSQL, Spark Logging , error handling and unit testing of PySpark and Spark Scala applicationsApplying spark transformation on data stored in AWS S3 using Glue and viewing data using AthenaHow to become a productive data engineer leveraging ChatGPTPrerequisites:This course is designed for Data Engineering beginners with no prior knowledge of Python and Scala required. However, some familiarity with databases and SQL is necessary to succeed in this course. Upon completion, you will have the skills and knowledge required to succeed in a real-world Data Engineer role.