|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/apache-spark-for-big-data-analytics-and-data-processing/
课程评论:没有评论
课程名称:Apache Spark 大数据分析与数据处理 课程概述: 在当今世界,数据每天以惊人的速度生成,组织机构越来越注重大数据处理,以实时、高效地处理大量数据。因此,Apache Spark 在大数据市场上迅速获得了广泛关注。如果您希望从这一流行的框架中最大限度地受益,那么这个学习路径将非常适合您。本课程旨在通过 Apache Spark 执行数据流和数据分析,包括从 JSON、Hive 和 Parquet 等多种结构化源加载数据,构建流应用程序,管理高速流和外部数据源的最佳实践,探索机器学习库及 GraphX 进行图形处理和分析,并通过项目实践巩固学习。 课程内容: 本培训程序包括三个完整课程,旨在提供尽可能全面的培训。 1. **实时数据处理的 Spark 分析**: 该课程首先介绍 Spark SQL 的使用,学习 Spark SQL API 和内置功能,进行互动分析及与 Java/Scala/Python 的集成。接下来,深入探讨 Spark Streaming 及其相关概念,例如 StreamingContext 和 DStreams,学习其在 Spark Core 之上的工作原理,最后掌握高速流和外部数据源管理的最佳实践。 2. **Apache Spark 的高级分析与实时数据处理**: 利用 Spark 框架的各个组件高效处理、分析和可视化数据,实施高速度流处理以进行实时数据分析,使用机器学习技术和图形进行数据分析,解决问题并了解 MLlib 工具包中的所有工具,运用 Spark MLlib 中的一些有用的机器学习算法,同时将 Spark 与 R 进行整合。 3. **基于 Apache Spark 的大数据分析项目**: 包含多个现实世界的项目示例。第一个项目是通过高效连接数据集来查找电子商务业务的畅销产品;接着进行市场篮子分析,识别可能一起购买的商品;然后通过概率逻辑回归为帖子找作者;构建基于内容的电影推荐系统,预测是否会发生某个动作;最后使用 MapReduce Spark 程序计算社交网络中的共同好友。 课程结束时,您将掌握 Spark 框架,有助于您实时分析和处理大数据。 专家介绍: 本课程的讲师包括经验丰富的 Nisant Garg 和 Tomasz Lelek。Nisant在多种技术上拥有超过17年的软件架构和开发经验,目前在大数据研发组担任技术架构师。Tomasz是一名软件工程师和InitLearn的联合创始人,专注于Java和Scala编程,对软件开发充满热情,曾在多个国际会议上发言。 总之,本课程提供了全面深入的 Apache Spark 培训,适合希望在大数据分析和处理领域提升技能的学习者。
Today's world witnesses a massive amount of data being generated everyday, everywhere. As a result, a number of organizations are focusing on Big Data processing to process large amounts of data in real-time with maximum efficiency. This has led to Apache Spark gaining popularity in the Big Data market rapidly. If you want to get the most out of the trending Big Data framework for all your data processing needs, then go for this Learning Path.This comprehensive 3-in-1 course focuses on performing data streaming and data analytics with Apache Spark. You will learn to load data from a variety of structured sources such as JSON, Hive, and Parquet using Spark SQL and schema RDDs. You will also build streaming applications and learn best practices for managing high-velocity streaming and external data sources. Next, you will explore Spark machine learning libraries and GraphX where you will perform graphical processing and analysis. Finally, you will build projects which will help you put your learnings into practice and get a strong hold of the topic.Contents and OverviewThis training program includes 3 complete courses, carefully chosen to give you the most comprehensive training possible.The first course, Spark Analytics for Real-Time Data Processing, starts off with explaining Spark SQL. You will learn how to use the Spark SQL API and built-in functions with Apache Spark. You will also go through some interactive analysis and look at some integrations between Spark and Java/Scala/Python. Next, you will explore Spark Streaming, streamingcontext, and DStreams. You will learn how Spark streaming works on top of the Spark core, thus inheriting its features. Finally, you will stream data and also learn best practices for managing high-velocity streaming and external data sources.In the second course, Advanced Analytics and Real-Time Data Processing in Apache Spark, you will leverage the features of various components of the Spark framework to efficiently process, analyze, and visualize your data. You will then learn how to implement the high velocity streaming operation for data processing in order to perform efficient analytics on your real-time data. You will also analyze data using machine learning techniques and graphs. Next, you will learn to solve problems using machine learning techniques and find out about all the tools available in the MLlib toolkit. Finally, you will see some useful machine learning algorithms with the help of Spark MLlib and will integrate Spark with R.The third course, Big Data Analytics Projects with Apache Spark, contains various projects that consist of real-world examples. The first project is to find top selling products for an e-commerce business by efficiently joining data sets in the Mapreduce paradigm. Next, a Market Basket Analysis will help you identify items likely to be purchased together and find correlations between items in a set of transactions. Moving on, you will learn about probabilistic logistic regression by finding an author for a post. Next, you will build a content-based recommendation system for movies to predict whether an action will happen, which you will do by building a trained model. Finally, you will use the Mapreduce Spark program to calculate mutual friends on social network.By the end of this course, you will have a sound understanding of the Spark framework, which will help you in analyzing and processing big data in real time.Meet Your Expert(s):We have the best work of the following esteemed author(s) to ensure that your learning journey is smooth:Nishant Garg has over 17 years of software architecture and development experience in various technologies, such as Java Enterprise Edition, SOA, Spring, Hadoop, Hive, Flume, Sqoop, Oozie, Spark, Shark, YARN, Impala, Kafka, Storm, Solr/Lucene, NoSQL databases (such as HBase, Cassandra, and MongoDB), and MPP databases (such as GreenPlum). He received his MS in software systems from the Birla Institute of Technology and Science, Pilani, India, and is currently working as a technical architect for the Big Data RandD Group with Impetus Infotech Pvt. Ltd. Previously, Nishant has enjoyed working with some of the most recognizable names in IT services and financial industries, employing full software life cycle methodologies such as Agile and SCRUM. Nishant has also undertaken many speaking engagements on big data technologies and is also the author of Apache Kafka and HBase Essentials, Packt Publishing. Tomasz Lelek is a Software Engineer and Co-Founder of InitLearn. He mostly does programming in Java and Scala. He dedicates his time and effort to get better at everything. He is currently diving into Big Data technologies. Tomasz is very passionate about everything associated with software development. He has been a speaker at a few conferences in Poland-Confitura and JDD, and at the Krakow Scala User Group. He has also conducted a live coding session at Geecon Conference. He was also a speaker at an international event in Dhaka. He is very enthusiastic and loves to share his knowledge.