|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/big-data-foundation-for-developers/
课程评论:没有评论
**课程名称:** 开发者大数据基础 **课程概述:** 本课程旨在为开发者提供在大数据领域的核心技能,重点讲解 Apache Hadoop、Yarn、Hive 和 Spark 这四大流行工具。通过一系列的实践活动,学员将能够: * **搭建个人大数据开发环境:** 熟悉并部署本地大数据开发环境。 * **掌握核心概念:** 深入理解 Hadoop、Yarn、Hive 和 Spark 的基本原理和工作机制。 * **数据交互:** 熟练地将数据导入和导出到大数据集群。 * **数据处理:** 掌握 Map/Reduce 编程模型,并在 Yarn 上运行 Map/Reduce 和 Spark 作业。 * **Scala 编程:** 学习使用 Scala 语言处理 Spark 中的大数据,包括 RDD(弹性分布式数据集)和 DataFrame 的使用。 * **数据存储:** 了解并应用 Parquet 格式存储大数据。 * **机器学习应用:** 利用 Spark 的机器学习库(MLlib)构建各类机器学习模型,如决策树、推荐引擎、线性回归和异常检测。 **课程特点:** * **实践导向:** 包含 50 多个动手实践活动,让学员在实践中学习。 * **技术栈全面:** 涵盖大数据处理和分析的关键工具。 * **语言支持:** 假设学员具备 Java 知识,并提供 Scala 基础教学,使学员能够高效地使用 Spark。 * **职业发展:** 为开发者融入大数据开发团队提供坚实的基础。
Apache Hadoop, Yarn, Hive and Spark are popular big data tools used by many organizations to develop big data analytics solutions. Through this course students can develop big data applications using these tools to process data and derive valuable insights from data. By the end of the course, students will be able to set up a personal big data development environment, master the fundamental concepts of Hadoop, Yarn, Hive and Spark, copy data into and from a big data cluster, process the data using the Map/Reduce paradigm, run Map/Reduce and Spark jobs on Yarn, Learn to process big data using Scala programming language in Spark, Use RDDs and dataframes to process big data, use Parquet format to store data, and finally use Machine Learning Libraries of Spark to develop Machine Learning solutions like decision trees, recommendation engine, Linear Regression and Anomaly detection. This is a hands on development course and you will practice more than 50 activities during this course. While Java knowledge is assumed, fundamentals of Scala are taught so that you can write Scala code to process data in Spark. The course provides a foundation for developers to join big data development teams in their organization.