|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/learning-path-spark-data-science-with-apache-spark/
课程评论:没有评论
课程名称:学习路径:Spark:使用Apache Spark进行数据科学 概述:每年都会生成大量数据,这些数据需要存储和分析。Apache Spark能够高效处理这些大数据,其真实的强大之处在于快速处理和高效执行数据科学任务的平台。Spark结合了ETL(数据提取、转换和加载)、批量分析、实时流分析、机器学习、图处理和可视化,帮助数据科学家应对复杂的原始非结构化数据集。Spark的愿景是使从单机处理转向集群处理的过渡更加灵活,从而使数据科学任务的执行更加敏捷。如果你有兴趣学习大数据处理并高效执行数据科学任务,那么这个学习路径非常适合你。 此学习路径由Packt提供,包含一系列按照逻辑和步骤编排的独立视频产品,每个视频都在之前学习的基础上进行扩展。学习路径的亮点包括: - 探索Apache Spark架构,深入了解其API和关键特性 - 实现高效的大数据处理 - 编写可维护且易于测试的代码 - 探索Spark在数据科学中的各种应用 - 快速上手Apache Spark,轻松清理、分析和可视化数据 学习之旅一开始将详细解释Spark API和其架构的基础知识。接着,你将学习数据挖掘和数据清理,编写实际作业分析数据。此外,还将学习构建机器学习应用所需的步骤,深入了解机器学习算法和不同技术。进而将学习如何使用Spark Streaming收集、清理和可视化来自Twitter的数据。最后,你将掌握图处理等分析技术。通过这个学习路径,你将能够以视觉化、全面且吸引业务和其他利益相关者的方式,完成所有数据科学任务。 专家介绍:我们邀请了以下杰出的作者,确保你的学习之旅顺利进行: - Tomasz Lelek是一名软件工程师,主要使用Java和Scala编程。他喜欢微服务架构和函数式编程,并最近深入研究了Apache Spark和Hadoop等大数据技术。 - Eric Charles在数据科学领域拥有10年的经验,是Datalayer的创始人,一个数据科学家的社交网络。他热衷于使用软件和数学帮助公司从数据中获取洞察,日常工作包括构建使用先进机器学习算法、高效SQL、流处理与图分析的解决方案,并注重可视化和结果分享。他也很关注开源技术,并且是Apache的活跃成员。
Every year a large amount of data is generated which needs to be stored and analyzed. Apache Spark allows you to process such big data. The real power and value proposition of Apache Spark is its speed and platform to execute data science tasks. Spark's unique use case is that it combines ETL, batch analytic, real-time stream analysis, machine learning, graph processing, and visualizations to allow data scientists to tackle the complexities that come with raw unstructured data sets. Spark embraces this approach and has the vision to make the transition from working on a single machine to working on a cluster, something that makes data science tasks a lot more agile. So, if you're interested to learn big data processing and execute data science tasks efficiently, then go for this Learning Path. Packt's Video Learning Path is a series of individual video products put together in a logical and stepwise manner such that each video builds on the skills learned in the video before it. The highlights of this Learning Path are: Explore the Apache Spark architecture and delve into its API and key features Implement efficient big data processing Write code that is maintainable and easy to test Explore various facets of data science with Spark Get up and running with Apache Spark and clean, analyze, and visualize data with ease Let's take a quick look at your learning journey. This Learning Path starts off by explaining the basics of Spark API and its architecture in detail. You will then learn about data mining and data cleaning. You will also learn to analyze data by writing actual jobs. Next, you will learn the needed steps to build machine learning applications. You will also explore machine learning algorithms and different machine learning techniques. Further, you will learn to collect, clean, and visualize data coming from Twitter with Spark streaming. Finally, you will understand how to perform analysis including graph processing. By the end of this Learning Path, you will be able to do all your data science tasks in a very visual way, comprehensive and appealing for business and other stakeholders. Meet Your Experts: We have the best works of the following esteemed authors to ensure that your learning journey is smooth: Tomasz Lelek is a Software Engineer, programmer mostly in Java and Scala. He is a fan of microservices architecture, and functional programming. He recently dived into big data technologies such as Apache Spark and Hadoop. Eric Charles has 10 years of experience in the field of Data Science and is the founder of Datalayer, a social network for Data Scientists. He is passionate about using software and mathematics to help companies get insights from data. His typical day includes building efficient processing with advanced machine learning algorithms, easy SQL, streaming, and graph analytics. He also focuses a lot on visualization and result sharing. He is passionate about open-source technologies and is an active Apache Member.