|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-engineering-using-kafka-and-spark-structured-streaming/
课程评论:没有评论
在本 Coursera 课程“使用 Kafka 和 Spark Structured Streaming 进行数据工程”中,您将学习如何构建流式处理管道。 课程首先介绍如何设置一个包含 Hadoop、Hive、Spark 和 Kafka 的单节点 Linux 环境,这为您构建流式处理管道奠定了基础。 接着,您将学习 Kafka 的基本操作,包括创建主题、生产和消费消息。您还将了解如何使用 Kafka Connect 将数据从 Web 服务器日志导入 Kafka 主题,并将 Kafka 主题中的数据导出到 HDFS 作为数据接收端。 在掌握了 Kafka 的数据摄取能力后,课程将深入介绍 Spark Structured Streaming 的关键概念。 最后,您将整合 Kafka 和 Spark Structured Streaming,构建一个完整的流式处理管道:从 Kafka 主题消费数据,进行处理,然后将结果写入不同的目标。此外,您还将学习如何利用 Spark Structured Streaming 实现增量数据处理。 课程提纲包括: * 使用 AWS Cloud9 或 GCP 设置环境 * 设置单节点 Hadoop 集群 * 在单节点 Hadoop 集群上设置 Hive 和 Spark * 在单节点 Hadoop 集群上设置单节点 Kafka 集群 * Kafka 入门 * 使用 Kafka Connect 进行数据摄取(Web 服务器日志到 Kafka 主题) * 使用 Kafka Connect 进行数据摄取(Kafka 主题到 HDFS) * Spark Structured Streaming 概述 * Kafka 和 Spark Structured Streaming 集成 * 使用 Spark Structured Streaming 进行增量加载 如果您在学习过程中遇到技术问题,可以通过 Udemy Messenger 寻求支持,将在 48 小时内得到解决。
As part of this course, you will be learning to build streaming pipelines by integrating Kafka and Spark Structured Streaming. Let us go through the details about what is covered in the course.First of all, we need to have the proper environment to build streaming pipelines using Kafka and Spark Structured Streaming on top of Hadoop or any other distributed file system. As part of the course, you will start with setting up a self-support lab with all the key components such as Hadoop, Hive, Spark, and Kafka on a single node Linux-based system.Once the environment is set up you will go through the details related to getting started with Kafka. As part of that process, you will create a Kafka topic, produce messages into the topic as well as consume messages from the topic.You will also learn how to use Kafka Connect to ingest data from web server logs into Kafka topic as well as ingest data from Kafka topic into HDFS as a sink.Once you understand Kafka from the perspective of Data Ingestion, you will get an overview of some of the key concepts of related Spark Structured Streaming.After learning Kafka and Spark Structured streaming separately, you will build a streaming pipeline to consume data from Kafka topic using Spark Structured Streaming, then process and write to different targets.You will also learn how to take care of incremental data processing using Spark Structured Streaming.Course OutlineHere is a brief outline of the course. You can choose either Cloud9 or GCP to provision a server to set up the environment.Setting up Environment using AWS Cloud9 or GCPSetup Single Node Hadoop ClusterSetup Hive and Spark on top of Single Node Hadoop ClusterSetup Single Node Kafka Cluster on top of Single Node Hadoop ClusterGetting Started with KafkaData Ingestion using Kafka Connect - Web server log files as a source to Kafka TopicData Ingestion using Kafka Connect - Kafka Topic to HDFS a sinkOverview of Spark Structured StreamingKafka and Spark Structured Streaming IntegrationIncremental Loads using Spark Structured StreamingUdemy based supportIn case you run into technical challenges while taking the course, feel free to raise your concerns using Udemy Messenger. We will make sure that issue is resolved in 48 hours.