Apache Spark Project for Beginners: A Complete Project Guide

所在平台: Udemy

课程主页: https://www.udemy.com/course/apache-spark-project-for-beginners/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:Apache Spark 项目初学者:完整的项目指南 课程概述:该课程旨在引导学员通过构建实时消息处理应用程序,深入了解 Apache Spark。课程中将开发一个 Meetup RSVP 流处理应用程序,使用技术包括 Apache Spark 的 Scala API、Spark Structured Streaming、Apache Kafka、Python、Python Dash、MongoDB 和 MySQL。学员将构建一个数据管道,从流数据源(Meetup Dot Com RSVP 流 API 数据)提取数据,直至使用 Apache Spark 和其他大数据框架进行数据可视化。 Apache Spark 是一个开源的统一分析引擎,专为大规模数据处理而设计。它提供了编程集群的接口,具备隐式数据并行性和容错性。Apache Kafka 是一个分布式事件存储和流处理平台,同样是由 Apache 软件基金会开发的开源系统,使用 Java 和 Scala 编写,旨在为处理实时数据提供统一的高吞吐量和低延迟的平台。 此外,Apache Hadoop 是一个开源软件工具集,帮助利用网络中的多台计算机来解决涉及大量数据和计算的问题,提供了一个基于 MapReduce 编程模型的分布式存储和处理大数据的软件框架。而 NoSQL(最初指“非 SQL”或“非关系型”)数据库则提供了一种存储和检索数据的机制,其模型与关系数据库中使用的表格关系不同。 通过本课程的学习,初学者将获得有关大数据处理、实时数据流处理及数据可视化的完整项目实践经验。

课程评论(0条)

课程详情

End to End Project Development of Real-Time Message Processing Application: In this Apache Spark Project, we are going to build Meetup RSVP Stream Processing Application using Apache Spark with Scala API, Spark Structured Streaming, Apache Kafka, Python, Python Dash, MongoDB and MySQL. And we are going to build a data pipeline which takes data from stream data source(Meetup Dot Com RSVP Stream API Data) to Data Visualisation using Apache Spark and other big data frameworks.Apache Spark is an open-source unified analytics engine for large-scale data processing. Spark provides an interface for programming clusters with implicit data parallelism and fault tolerance.Apache Kafka is a distributed event store and stream-processing platform. It is an open-source system developed by the Apache Software Foundation written in Java and Scala. The project aims to provide a unified, high-throughput, low-latency platform for handling real-time data feeds.Apache Hadoop is a collection of open-source software utilities that facilitates using a network of many computers to solve problems involving massive amounts of data and computation. It provides a software framework for distributed storage and processing of big data using the MapReduce programming model.A NoSQL (originally referring to "non-SQL" or "non-relational") database provides a mechanism for storage and retrieval of data that is modeled in means other than the tabular relations used in relational databases.

课程标签

0人关注该课程

主题相关的课程