|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/spark-scala-cassandra/
课程评论:没有评论
**课程名称:** Apache Spark 和 Scala:Cassandra 专业技能提升 **课程概述:** 本课程专为 Apache Cassandra 数据库管理员设计,旨在教授如何利用 Apache Spark 的强大功能和 Scala 语言进行大规模数据处理和分析。课程内容录制于 2020 年 12 月至 2021 年 1 月,所有知识点仍然有效。 **核心内容:** * **Scala 基础入门:** * Scala REPL (Read-Evaluate-Print-Loop) 的理解与使用 * Scala 变量 (`var` 和 `val`)、数据类型 * 不可变对象的概念及其在现代技术中的重要性 * Scala 中的循环 (`for` 循环和 `foreach`) 以及基本的输出方法 (`print` 和 `println`) * **Apache Spark 基础与集群构建:** * Spark 的下载与安装(以 CentOS Linux 服务器为例) * Spark 集群的配置与搭建,包括 Spark Master 和 Spark Workers 的配置文件 * Spark 架构解析:Driver、Worker、Executor 和 Task * Spark Master 和 Spark Workers 的启动与停止 * **Spark 在 Cassandra 使用场景中的应用:** * 使用 `spark-shell` 读取 CSV 格式文件,并进行 `count`、`take`、`foreach`、`print` 和 `println` 等基本操作 * 学习 `filter`、`contains`、`map` 和 `reduce` 等核心 Spark 操作 * 利用大型数据集(百万行文件)进行高级分析 * 理解 API 差异:RDD、DataFrame 和 Dataset API * 使用 Spark SQL 和 DataFrame API * **Spark Cassandra Connector:** * 配置 Spark 与 Cassandra 的连接,支持 RDD、DataFrame 和 Dataset API * 使用 RDD API 在 `spark-shell` 中读写 Cassandra 表数据 * 利用 Spark SQL 和 DataFrame/Dataset API 读写 Cassandra 数据 * **解决复杂用例:** 学习如何使用 Spark 执行 Cassandra 本身难以实现或无法实现的复杂数据处理任务。 **课程特色:** * **专项培训:** 专为 Cassandra DBA 量身打造,填补技术鸿沟。 * **实践导向:** 涵盖从基础语法到实际应用场景的全面学习。 * **深度解析:** 深入讲解 Spark 的核心概念、架构及与 Cassandra 的集成。 * **独特视角:** 以师生对话形式录制,提供独一无二的学习体验,确保内容不易在网络上找到。 本课程旨在帮助 Cassandra 专业人士掌握 Spark 和 Scala,从而能够更有效地处理和分析海量数据,解决更复杂的业务挑战。
*** This training course was recorded between December 2020 & January 2021. All the content is still valid. ***Apache Cassandra is the most powerful NoSQL database.Apache Spark is the best analytics engine for large-scale data processing.It is very challenging for Apache Cassandra Database Administrators to learn a new language and a new analytics engine.In this course we will start with the basics of Scala Language.We will download and install Scala on CentOS Linux server.We will understand what is Scala REPL (Read-Evaluate-Print-Loop).We will discuss about Scala Variables, Data Types, var and val.We will understand why these new technologies are using Immutable objects.We will learn how to use for loop and foreach along with print and println.We will understand how to use Apache Spark for various use cases along with Apache Cassandra use cases.We will learn how to download and configure Apache Spark to build a cluster.We will understand the configuration files for Spark Master and Spark Workers.We will discuss about Spark Driver, Worker, Executor and Tasks.We will learn how to start and stop Spark Master and Spark Workers.We will use spark-shell to read data from CSV formatted files.We will use spark-shell for operations such as count, take, foreach, print & println.We will learn about filter, contains and map and reduce.We will use a very large million row file for advanced analytics.We will learn the differences between APIs such as RDD, DataFrame and Dataset APIs.We will learn how to use Spark SQL with DataFrame APIs.We will understand how to use Spark Cassandra Connector to use Spark Analytics on data stored in Cassandra.We will learn how to configure connectivity between Spark and Cassandra for various APIs such as RDD / DataFrame / Dataset APIs.We will use RDD APIs in spark-shell to read data from Cassandra tables and write data back to Cassandra tables.We will learn how to use Spark to perform the complicated tasks which are not possible in Cassandra.We will learn how to use Spark SQL with DataFrames API and Datasets API to read and write data from Cassandra.We will use Spark SQL to solve several complicated use cases which are not possible in Cassandra.This is a special of one of its kind training course for "Apache Spark and Scala for Cassandra DBAs".NOTE: This training was recorded as a conversation between 2 people. Instructor and the Student. During this course you will hear both of them speaking. I guarantee that you will not find this kind of course content anywhere else on whole internet. So please try this training course.