|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/spark-sql-hadoop-for-data-scientists-big-data-analysts/
课程评论:没有评论
课程名称:[LEGACY-SUPPORT END] Spark SQL & Hadoop (数据科学专用) 课程概述: 该课程已经退休,不再提供支持。最初设计目的是帮助学生通过现已停用的Cloudera认证考试,但课程内容仍对想要在Spark和Hadoop集群上练习技能的学生有帮助。Apache Spark是当前流行的大数据处理系统,而Apache Hadoop则被许多希望在本地高效存储数据的组织广泛使用。随着数据科学、大数据分析和数据工程岗位的增加,对掌握Spark和Hadoop技术的人才需求也在增长。 本课程专为希望利用Hadoop和Apache Spark处理大数据的数据科学家、大数据分析师和数据工程师设计,适合那些希望进行交互式大数据分析或开始编写生产应用以准备数据进行进一步分析的学生。该课程也非常适合大学生和应届毕业生,帮助他们获取Spark和Hadoop的实际应用经验,或任何希望在大数据环境中应用SQL技能的人。 课程内容总结: - 本课程重点在于提供足够的理论知识,使学生能够有效使用Hadoop和Spark,而不会被过多的古老低级API(如RDDs)的理论所困扰。 - 课程包含近30个问题,涵盖HDFS命令、基本数据工程任务和数据分析。 - 提供所有问题的完整解决方案。 - 包含一个名为Verulam Blue的虚拟机环境,已安装Spark Hadoop集群,学生可以在此环境中练习。 - 虚拟机包括已加载数据集到HDFS,学生不需额外工作即可进行练习,并且安装了Apache Zeppelin,类似于Python的Jupyter Notebook。 实践内容包括: - 将存储在HDFS中的数据转换成新数据格式并写回HDFS。 - 从HDFS加载数据进行Spark应用,并将结果写回HDFS。 - 处理多种文件格式的读写。 - 使用Spark API进行标准的提取、转换和加载(ETL)过程。 - 使用元存储表作为Spark应用的输入源或输出目标。 - 理解Spark中数据集查询的基本原理。 - 使用Spark进行数据过滤、聚合统计计算、数据集连接和排序等操作。 本课程为希望在Spark Hadoop环境中获得实践经验的学生提供了良好的学习机会。
*Important Notice*This course has been retired and is no longer receiving support. Originally designed to help students pass the now-retired Cloudera Certification exams, the material remains useful for those wanting to practice their skills on Spark and Hadoop clusters. However, its primary focus was certification preparation, which many students successfully completed.Apache Spark is currently one of the most popular systems for processing big data.Apache Hadoop continues to be used by many organizations that look to store data locally on premises. Hadoop allows these organisations to efficiently store big datasets ranging in size from gigabytes to petabytes.As the number of vacancies for data science, big data analysis and data engineering roles continue to grow, so too will the demand for individuals that possess knowledge of Spark and Hadoop technologies to fill these vacancies.This course has been designed specifically for data scientists, big data analysts and data engineers looking to leverage the power of Hadoop and Apache Spark to make sense of big data.This course will help those individuals that are looking to interactively analyse big data or to begin writing production applications to prepare data for further analysis using Spark SQL in a Hadoop environment.The course is also well suited for university students and recent graduates that are keen to gain exposure to Spark & Hadoop or anyone who simply wants to apply their SQL skills in a big data environment using Spark-SQL.This course has been designed to be concise and to provide students with a necessary and sufficient amount of theory, enough for them to be able to use Hadoop & Spark without getting bogged down in too much theory about older low-level APIs such as RDDs.On solving the questions contained in this course students will begin to develop those skills & the confidence needed to handle real world scenarios that come their way in a production environment.(a) There are just under 30 problems in this course. These cover hdfs commands, basic data engineering tasks and data analysis.(b) Fully worked out solutions to all the problems.(c) Also included is the Verulam Blue virtual machine which is an environment that has a spark Hadoop cluster already installed so that you can practice working on the problems.The VM contains a Spark Hadoop environment which allows students to read and write data to & from the Hadoop file system as well as to store metastore tables on the Hive metastore.All the datasets students will need for the problems are already loaded onto HDFS, so there is no need for students to do any extra work.The VM also has Apache Zeppelin installed. This is a notebook specific to Spark and is similar to Python's Jupyter notebook.This course will allow students to get hands-on experience working in a Spark Hadoop environment as they practice:Converting a set of data values in a given format stored in HDFS into new data values or a new data format and writing them into HDFS.Loading data from HDFS for use in Spark applications & writing the results back into HDFS using Spark.Reading and writing files in a variety of file formats.Performing standard extract, transform, load (ETL) processes on data using the Spark API.Using metastore tables as an input source or an output sink for Spark applications.Applying the understanding of the fundamentals of querying datasets in Spark.Filtering data using Spark.Writing queries that calculate aggregate statistics.Joining disparate datasets using Spark.Producing ranked or sorted data.