Spark Project on Cloudera Hadoop(CDH) and GCP for Beginners

所在平台: Udemy

课程主页: https://www.udemy.com/course/spark-project-on-cloudera-hadoop-cdh-and-gcp-for-beginners/

课程评论:没有评论

第一个写评论        关注课程

课程简介

**课程名称:** 面向初学者的 Cloudera Hadoop (CDH) 和 GCP 上的 Spark 项目 **课程概述:** 本课程专注于构建一个强大的数据处理管线,以应对零售行业中海量实时数据的挑战。通过利用 Apache NiFi, Apache Kafka, Apache Spark, Apache Cassandra, MongoDB, Apache Hive 和 Apache Zeppelin 等一系列开源技术,学员将学习如何从实时生成的数据中提取有价值的洞察,从而帮助企业做出明智的业务决策,提升销售额并改善客户体验。 课程的核心是使用 Apache Spark,该技术在 Cloudera Hadoop (CDH 6.3) 集群上运行,而该集群部署在 Google Cloud Platform (GCP) 之上。Spark 是一个统一的、可扩展的分析引擎,为大规模数据处理提供了高容错性和并行处理能力。 课程还将深入介绍: * **Apache Kafka:** 一个为处理实时数据流而设计的分布式事件存储和流处理平台,提供高吞吐量和低延迟的数据传输。 * **Apache Hadoop:** 一个开源软件框架,用于处理海量数据和计算,提供分布式存储和 MapReduce 编程模型。 * **JSON 和 NoSQL 数据库 (Cassandra, MongoDB):** 用于存储和检索非关系型数据的技术,以应对复杂、多变的业务场景。 * **Apache Hive:** 用于大数据仓库的工具,可以对存储在 Hadoop 中的数据进行结构化查询。 * **Apache Zeppelin:** 一个交互式数据可视化和协作平台,用于数据探索和演示。 通过实践项目,学员将掌握使用 Scala 和 PySpark 在 Spark 中进行开发,并了解如何在 GCP 上搭建和优化大数据处理架构。本课程适合希望进入大数据领域,特别是对实时数据处理和分析感兴趣的初学者。

课程评论(0条)

课程详情

In retail business, retail stores and eCommerce websites generates large amount of data in real-time. There is always a need to process these data in real-time and generate insights which will be used by the business people and they make business decision to increase the sales in the retail market and provide better customer experience. Since the data is huge and coming in real-time, we need to choose the right architecture with scalable storage and computation frameworks/technologies.Hence we want to build the Data Processing Pipeline Using Apache NiFi, Apache Kafka, Apache Spark, Apache Cassandra, MongoDB, Apache Hive and Apache Zeppelin to generate insights out of this data.The Spark Project is built using Apache Spark with Scala and PySpark on Cloudera Hadoop(CDH 6.3) Cluster which is on top of Google Cloud Platform(GCP).Apache Spark is an open-source unified analytics engine for large-scale data processing. Spark provides an interface for programming clusters with implicit data parallelism and fault tolerance.Apache Kafka is a distributed event store and stream-processing platform. It is an open-source system developed by the Apache Software Foundation written in Java and Scala. The project aims to provide a unified, high-throughput, low-latency platform for handling real-time data feeds.Apache Hadoop is a collection of open-source software utilities that facilitates using a network of many computers to solve problems involving massive amounts of data and computation. It provides a software framework for distributed storage and processing of big data using the MapReduce programming model.A NoSQL (originally referring to "non-SQL" or "non-relational") database provides a mechanism for storage and retrieval of data that is modeled in means other than the tabular relations used in relational databases.

课程标签

0人关注该课程

主题相关的课程