|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/best-hands-on-big-data-practices-and-use-cases-using-pyspark/
课程评论:没有评论
**Coursera 课程总结:PySpark 与 Spark 性能调优的最佳实战** 本课程旨在通过深入的 PySpark 实操练习,结合学术界和工业界的真实案例,帮助学员掌握与海量数据进行交互的能力。课程将重点关注大数据处理中的分布式挑战,如数据倾斜(data skewness)和内存溢出(spill)。 **课程亮点:** * **实操驱动:** 学习将通过真实、复杂的用例驱动,让学员亲身体验大数据处理的每一个环节。 * **全面覆盖:** 课程将深入讲解 Spark 引擎的核心原理,并利用 Spark RDD、DataFrame (DF) 和 SQL 处理结构化、半结构化和非结构化数据。 * **解决痛点:** 专注于解决大数据处理中的实际问题,使学员能够从宏观视角转向微观细节,并学会如何应对数据倾斜等挑战。 * **工业导向:** 紧密结合行业需求,识别并培养学员在大数据分析领域最关键的技能。 * **成果导向:** 培养学员构建不同类型(体量、多样性、真实性)的大数据应用的能力,并提供 PySpark 处理大数据问题的最佳实践示例。 **适用人群:** 任何希望精通 Spark 和 PySpark,并希望通过真实、有挑战性的场景学习和传播大数据分析知识的学习者。 **学习目标:** 完成课程后,学员将能够: * 熟练运用 PySpark 处理海量数据。 * 理解并解决大数据处理中的关键分布式挑战。 * 构建针对不同类型数据的复杂大数据应用程序。 * 掌握 Spark RDD, DF, 和 SQL 的高级用法。 * 熟悉业界领先的 PySpark 大数据问题解决方案。
In this course, students will be provided with hands-on PySpark practices using real case studies from academia and industry to be able to work interactively with massive data. In addition, students will consider distributed processing challenges, such as data skewness and spill within big data processing. We designed this course for anyone seeking to master Spark and PySpark and Spread the knowledge of Big Data Analytics using real and challenging use cases.We will work with Spark RDD, DF, and SQL to process huge sized of data in the format of semi-structured, structured, and unstructured data. The learning outcomes and the teaching approach in this course will accelerate the learning by Identifying the most critical required skills in the industry and understanding the demands of Big Data analytics content.We will not only cover the details of the Spark engine for large-scale data processing, but also we will drill down big data problems that allow users to instantly shift from an overview of large-scale data to a more detailed and granular view using RDD, DF and SQL in real-life examples. We will walk through the Big Data case studies step by step to achieve the aim of this course.By the end of the course, you will be able to build Big Data applications for different types of data (volume, variety, veracity) and you will get acquainted with best-in-class examples of Big Data problems using PySpark.