Spark SQL and Spark 3 using Scala Hands-On with Labs

所在平台: Udemy

课程主页: https://www.udemy.com/course/cca-175-spark-and-hadoop-developer-certification-scala/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:使用Scala的Spark SQL和Spark 3实战与实验 课程概述: 本课程旨在教授学员构建数据工程管道所需的关键技能,使用Spark SQL和Spark数据框API,以Scala作为编程语言。该课程原为CCA 175 Spark和Hadoop开发者认证考试的备考课程,现已更新以涵盖更广泛的行业相关主题。数据工程是根据下游需求处理数据的过程,涉及批处理管道和流处理管道等构建。所有与数据处理相关的角色都归类于数据工程,传统上称为ETL开发和数据仓库开发等。Apache Spark已成为大规模数据工程的领先技术。 本课程适合希望转型为数据工程师角色的学员,课程设计者是一名具备丰富经验的数据工程解决方案架构师。课程内容包括大量动手实践任务,让学员获得使用相关工具的实际操作经验,同时包含大量练习供自我评估。 课程内容亮点: 1. 单节点大数据集群的搭建,包括: - Ubuntu基础的AWS Cloud9实例配置 - Docker、Jupyter Lab的安装与设置 - Hadoop、Hive、YARN和Spark的设置与验证 - 提供两个月的实验室访问权限,帮助学员更好地练习。 2. Scala语言快速回顾: - 如果学员对Scala不熟悉,建议在课程开始前学习相关Scala课程,以便更好理解数据工程相关内容。 3. 使用Spark SQL进行数据工程: - 深入学习Spark SQL,构建数据工程管道的基本转化,管理Spark元存储表(基本DDL和DML),窗口函数等。 4. 使用Spark数据框API进行数据工程: - 介绍如何利用Spark数据框API进行大规模的数据工程开发,包括列数据处理、基本转化(过滤、聚合、排序)及数据集连接操作。 课程采用互动式学习环境,所有演示都在先进的大数据集群上进行。学员可通过提供Udemy的购票凭证,获取一个月的实验室免费访问权限。 通过该课程,学员将在数据工程领域秋水共长天一色,获取丰富的实战经验和专业知识。

课程评论(0条)

课程详情

As part of this course, you will learn all the key skills to build Data Engineering Pipelines using Spark SQL and Spark Data Frame APIs using Scala as a Programming language. This course used to be a CCA 175 Spark and Hadoop Developer course for the preparation of the Certification Exam. As of 10/31/2021, the exam is sunset and we have renamed it to Spark SQL and Spark 3 using Scala as it covers industry-relevant topics beyond the scope of certification.About Data EngineeringData Engineering is nothing but processing the data depending on our downstream needs. We need to build different pipelines such as Batch Pipelines, Streaming Pipelines, etc as part of Data Engineering. All roles related to Data Processing are consolidated under Data Engineering. Conventionally, they are known as ETL Development, Data Warehouse Development, etc. Apache Spark is evolved as a leading technology to take care of Data Engineering at scale.I have prepared this course for anyone who would like to transition into a Data Engineer role using Spark (Scala). I myself am a proven Data Engineering Solution Architect with proven experience in designing solutions using Apache Spark.Let us go through the details about what you will be learning in this course. Keep in mind that the course is created with a lot of hands-on tasks which will give you enough practice using the right tools. Also, there are tons of tasks and exercises to evaluate yourself.Setup of Single Node Big Data ClusterMany of you would like to transition to Big Data from Conventional Technologies such as Mainframes, Oracle PL/SQL, etc and you might not have access to Big Data Clusters. It is very important for you set up the environment in the right manner. Don't worry if you do not have the cluster handy, we will guide you through support via Udemy Q & A.Setup Ubuntu-based AWS Cloud9 Instance with the right configurationEnsure Docker is setupSetup Jupyter Lab and other key componentsSetup and Validate Hadoop, Hive, YARN, and SparkAre you feeling a bit overwhelmed about setting up the environment? Don't worry!!! We will provide complementary lab access for up to 2 months. Here are the details.Training using an interactive environment. You will get 2 weeks of lab access, to begin with. If you like the environment, and acknowledge it by providing a 5* rating and feedback, the lab access will be extended to additional 6 weeks (2 months). Feel free to send an email to support@itversity.com to get complementary lab access. Also, if your employer provides a multi-node environment, we will help you set up the material for the practice as part of the live session. On top of Q & A Support, we also provide required support via live sessions.A quick recap of ScalaThis course requires a decent knowledge of Scala. To make sure you understand Spark from a Data Engineering perspective, we added a module to quickly warm up with Scala. If you are not familiar with Scala, then we suggest you go through relevant courses on Scala as Programming Language.Data Engineering using Spark SQLLet us, deep-dive into Spark SQL to understand how it can be used to build Data Engineering Pipelines. Spark with SQL will provide us the ability to leverage distributed computing capabilities of Spark coupled with easy-to-use developer-friendly SQL-style syntax.Getting Started with Spark SQLBasic Transformations using Spark SQLManaging Spark Metastore Tables - Basic DDL and DMLManaging Spark Metastore Tables Tables - DML and PartitioningOverview of Spark SQL FunctionsWindowing Functions using Spark SQLData Engineering using Spark Data Frame APIsSpark Data Frame APIs are an alternative way of building Data Engineering applications at scale leveraging distributed computing capabilities of Spark. Data Engineers from application development backgrounds might prefer Data Frame APIs over Spark SQL to build Data Engineering applications.Data Processing Overview using Spark Data Frame APIs leveraging Scala as Programming LanguageProcessing Column Data using Spark Data Frame APIs leveraging Scala as Programming LanguageBasic Transformations using Spark Data Frame APIs leveraging Scala as Programming Language - Filtering, Aggregations, and SortingJoining Data Sets using Spark Data Frame APIs leveraging Scala as Programming LanguageAll the demos are given on our state-of-the-art Big Data cluster. You can avail of one-month complimentary lab access by reaching out to support@itversity.com with a Udemy receipt.

课程标签

0人关注该课程

主题相关的课程