|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/advanced-apache-spark-for-data-scientists-and-developers/
课程评论:没有评论
## Coursera 课程总结:面向数据科学家和开发者的 Apache Spark 高级应用 本课程名为“面向数据科学家和开发者的 Apache Spark 高级应用”,由 Adastra Academy 提供,旨在帮助学员深入掌握 Apache Spark 的强大功能,并能开发、构建、调优和调试 Spark 应用。 **课程概述:** Apache Spark 作为一个开源的数据处理引擎,以其快速处理海量数据集和高性能分析能力而著称。与 MapReduce 不同,Spark 支持内存集群计算,显著提升了迭代算法和交互式数据挖掘任务的速度。 本课程包含: * **精讲视频讲座:** 深入剖析 Spark 的核心概念和技术。 * **详实的应用程序示例:** 通过实际案例演示 Spark 的应用方法。 * **NetBeans IDE 安装指南:** 帮助学员配置开发环境。 * **测试题:** 巩固学习内容,检验学习效果。 **课程重点:** 学员将全面学习 Spark 的四大内置库: * **Spark Streaming:** 用于实时数据流处理。 * **DataFrames (SparkSQL):** 提供结构化数据处理接口,支持 SQL 查询。 * **MLlib:** Spark 的机器学习库,用于构建和部署机器学习模型。 * **GraphX:** 用于图计算和图分析。 通过课程中的实践练习,学员将能够熟练运用这些库,创建功能齐全的真实世界应用程序。本课程强调“从零开始”的引导式学习方法,帮助学员扎实掌握 Spark 知识,最终成为专家。
Apache Spark is an open source data processing engine. Spark is designed to provide fast processing of large datasets, and high performance for a wide range of analytics applications. Unlike MapReduce, Spark enables in-memory cluster computing which greatly improves the speed of iterative algorithms and interactive data mining tasks. Adastra Academy's Advanced Apache Spark includes illuminating video lectures, thorough application examples, a guide to install the NetBeans Integrated Development Environment, and quizzes. Through this course, you will learn about Spark's four built-in libraries - SparkStreaming, DataFrames (SparkSQL), MLlib and GraphX - and how to develop, build, tune, and debug Spark applications. The course exercises will enable you to become proficient at creating fully functional real-world applications using the Apache Spark libraries. Unlike other courses, we give you the guided and ground-up approach to learning Spark that you need in order to become an expert.