Data Engineering Master Course: Spark/Hadoop/Kafka/MongoDB

所在平台: Udemy

课程主页: https://www.udemy.com/course/big-data-ingestion-using-sqoop-and-flume-cca-and-hdpcd/

课程评论:没有评论

第一个写评论        关注课程

课程简介

Coursera 数据工程大师(Spark/Hadoop/Kafka/MongoDB)课程总结: 本课程是一门全面的数据工程入门与进阶课程,旨在教授学员使用行业领先的大数据技术栈构建和管理数据管道。 **核心技术栈涵盖:** * **Hadoop 生态系统:** * **HDFS (Hadoop Distributed File System):** 学习 HDFS 的基本概念、常用命令,以及如何在其上操作数据。 * **Sqoop:** 掌握 Sqoop 的导入和导出功能,实现 MySQL 与 HDFS、Hive 之间的数据迁移,并学习如何使用不同的文件格式、压缩、分隔符、WHERE 子句、queries、split-by、boundary queries 以及增量模式进行数据导入,以及如何将 HDFS 和 Hive 的数据导出到 MySQL。 * **Flume:** 学习 Flume 的架构、工作原理,并通过实例演示如何从 Twitter、netcat、exec 等不同来源采集数据并保存到 HDFS。还将学习 Flume 拦截器、多代理和数据整合。 * **数据仓库与分析:** * **Hive:** 了解 Hive 的基础知识,学习创建外部表和托管表,使用 Parquet、Avro 等不同文件格式,掌握压缩技术,并深入学习 Hive 的分析功能,包括字符串函数、日期函数、分区(Partitioning)和分桶(Bucketing)。 * **分布式计算框架:** * **Spark:** 深入理解 Spark 的架构,包括集群概览、RDD(Resilient Distributed Datasets)的 DAG、Stages 和 Tasks。学习 RDD 的 Actions 和 Transformations,并通过示例展示 Transformations 和 Action 的应用。重点讲解 Spark DataFrame,包括与不同文件格式和压缩的交互,DataFrame API 的使用,Spark SQL,以及 DataFrame 的实际应用示例。此外,还涵盖 Spark 与 Cassandra 的集成,以及在 IntelliJ IDEA 和 EMR 上运行 Spark 的实践。 * **实时数据流处理:** * **Kafka:** 学习 Kafka 的核心概念,包括架构、分区(Partitions)和偏移量(Offsets)。掌握 Kafka Producer 和 Consumer 的开发,理解 Kafka 的 SerDes(Serializer/Deserializer),学习 Kafka 消息的发送和接收,以及如何使用 Kafka Connector 进行数据摄取。 * **NoSQL 数据库:** * **MongoDB:** 了解 MongoDB 的使用场景,掌握 CRUD (Create, Read, Update, Delete) 操作,学习 MongoDB Operators,以及如何处理数组(Arrays)。还将学习 MongoDB 与 Spark 的集成。 **职业发展与面试准备:** 课程的最后部分专注于为数据工程面试做准备,涵盖了 Sqoop、Hive、Spark 的面试常见问题,以及通用数据工程问题和真实的行业项目问题,帮助学员提升面试竞争力。 **总体而言,本课程为学员提供了构建、管理和优化现代大数据系统的必需技能,覆盖了从数据采集、存储、处理到分析的完整流程,并注重实践操作和面试指导。**

课程评论(0条)

课程详情

In this course, you will start by learning what is hadoop distributed file system and most common hadoop commands required to work with Hadoop File system.Then you will be introduced to Sqoop Import Understand lifecycle of sqoop command.Use sqoop import command to migrate data from Mysql to HDFS.Use sqoop import command to migrate data from Mysql to Hive.Use various file formats, compressions, file delimeter,where clause and queries while importing the data.Understand split-by and boundary queries.Use incremental mode to migrate the data from Mysql to HDFS.Further, you will learn Sqoop Export to migrate data.What is sqoop exportUsing sqoop export, migrate data from HDFS to Mysql.Using sqoop export, migrate data from Hive to Mysql.Further, you will learn about Apache FlumeUnderstand Flume Architecture.Using flume, Ingest data from Twitter and save to HDFS.Using flume, Ingest data from netcat and save to HDFS.Using flume, Ingest data from exec and show on console.Describe flume interceptors and see examples of using interceptors.Flume multiple agents Flume Consolidation.In the next section, we will learn about Apache HiveHive IntroExternal & Managed TablesWorking with Different Files - Parquet,AvroCompressionsHive AnalysisHive String FunctionsHive Date FunctionsPartitioningBucketingYou will learn about Apache SparkSpark IntroCluster OverviewRDDDAG/Stages/TasksActions & TransformationsTransformation & Action ExamplesSpark Data framesSpark Data frames - working with diff File Formats & CompressionDataframes API'sSpark SQLDataframe ExamplesSpark with Cassandra IntegrationRunning Spark on Intellij IDERunning Spark on EMRYou will learn about Apache KafkaKafka ArchitecturePartitions and offsetsKafka Producers and ConsumersKafka SerDEsKafka MessagesKafka ConnectorIngesting Data using Kafka ConnectorYou will learn about MongoDBMongoDB UsecasesCRUD OperationsMongoDB OperatorsWorking with ArraysMongoDB with SparkData Engineering Interview PreparationSqoop Interview QuestionsHive Interview QuestionsSpark Interview QuestionsData Engineering common questionsData Engineering Real project questions.

课程标签

0人关注该课程

主题相关的课程