PySpark & AWS: Master Big Data With PySpark and AWS

所在平台: Udemy

课程主页: https://www.udemy.com/course/pyspark-aws-master-big-data-with-pyspark-and-aws/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:PySpark & AWS:掌握大数据与PySpark和AWS 课程概述:本课程旨在帮助学习者从基础到高级水平学习大数据分析,特别专注于Python和Apache Spark的结合。学员将学习如何使用PySpark执行整个数据分析工作流,包括数据清洗、特征构建和机器学习模型的实现。课程内容涵盖Spark RDD、Dataframe及Spark SQL查询的使用,数据转换和操作,以及Spark和Hadoop的生态系统与架构。课程还将利用Databricks环境运行Spark脚本,并了解如何使用AWS云存储、数据库和计算服务,让Spark与AWS服务进行数据交互。 与其他课程的区别:本课程采用“边学边做”的方式,每个理论讲解后都会有实际操作,包含丰富的实战案例,帮助学员理解和掌握PySpark的核心概念与方法。课程内容易于理解,实用性强,并配有高质量的视频和详细的课程资料,学员在课程结束后将会接受作业和测验以巩固所学知识。课程共分为140多段简短视频,总时长约为16小时。 学习PySpark和AWS的原因:因为大数据处理领域对PySpark专业人才的需求持续增长,且薪资水平高,因此学习PySpark是非常有价值的。而AWS作为增长最快的公共云,掌握其技能也是顺应时代发展的必然选择。 课程内容:主要包括以下主题: 1. 大数据的应用与PySpark介绍 2. Hadoop及Spark生态系统与架构 3. Spark RDD的创建与操作 4. Spark DataFrame的使用 5. 协同过滤与推荐系统的开发 6. Spark实时数据流处理 7. ETL流程与数据加载 8. 项目:变更数据捕获与复制 本课程适合初学者、希望开发智能解决方案的学习者,以及对大数据领域有兴趣的人士。完成课程后,学员将能独立实施相关的PySpark项目,并掌握PySpark和AWS的理论与实践知识。 立即报名,提升大数据分析、数据处理及云计算的核心技能,开拓职业发展机会,成为PySpark与AWS的专家!

课程评论(0条)

课程详情

Comprehensive Course Description:The hottest buzzwords in the Big Data analytics industry are Python and Apache Spark. PySpark supports the collaboration of Python and Apache Spark. In this course, you'll start right from the basics and proceed to the advanced levels of data analysis. From cleaning data to building features and implementing machine learning (ML) models, you'll learn how to execute end-to-end workflows using PySpark.Right through the course, you'll be using PySpark for performing data analysis. You'll explore Spark RDDs, Dataframes, and a bit of Spark SQL queries. Also, you'll explore the transformations and actions that can be performed on the data using Spark RDDs and dataframes. You'll also explore the ecosystem of Spark and Hadoop and their underlying architecture. You'll use the Databricks environment for running the Spark scripts and explore it as well.Finally, you'll have a taste of Spark with AWS cloud. You'll see how we can leverage AWS storages, databases, computations, and how Spark can communicate with different AWS services and get its required data. How Is This Course Different? In this Learning by Doing course, every theoretical explanation is followed by practical implementation. The course ‘PySpark & AWS: Master Big Data With PySpark and AWS' is crafted to reflect the most in-demand workplace skills. This course will help you understand all the essential concepts and methodologies with regards to PySpark. The course is:• Easy to understand. • Expressive. • Exhaustive. • Practical with live coding. • Rich with the state of the art and latest knowledge of this field. As this course is a detailed compilation of all the basics, it will motivate you to make quick progress and experience much more than what you have learned. At the end of each concept, you will be assigned Homework/tasks/activities/quizzes along with solutions. This is to evaluate and promote your learning based on the previous concepts and methods you have learned. Most of these activities will be coding-based, as the aim is to get you up and running with implementations. High-quality video content, in-depth course material, evaluating questions, detailed course notes, and informative handouts are some of the perks of this course. You can approach our friendly team in case of any course-related queries, and we assure you of a fast response. The course tutorials are divided into 140+ brief videos. You'll learn the concepts and methodologies of PySpark and AWS along with a lot of practical implementation. The total runtime of the HD videos is around 16 hours.Why Should You Learn PySpark and AWS? PySpark is the Python library that makes the magic happen. PySpark is worth learning because of the huge demand for Spark professionals and the high salaries they command. The usage of PySpark in Big Data processing is increasing at a rapid pace compared to other Big Data tools. AWS, launched in 2006, is the fastest-growing public cloud. The right time to cash in on cloud computing skills-AWS skills, to be precise-is now.Course Content:The all-inclusive course consists of the following topics:1. Introduction:a. Why Big Data?b. Applications of PySparkc. Introduction to the Instructord. Introduction to the Coursee. Projects Overview2. Introduction to Hadoop, Spark EcoSystems, and Architectures:a. Hadoop EcoSystemb. Spark EcoSystemc. Hadoop Architectured. Spark Architecturee. PySpark Databricks setupf. PySpark local setup3. Spark RDDs:a. Introduction to PySpark RDDsb. Understanding underlying Partitionsc. RDD transformationsd. RDD actionse. Creating Spark RDDf. Running Spark Code Locallyg. RDD Map (Lambda)h. RDD Map (Simple Function)i. RDD FlatMapj. RDD Filterk. RDD Distinctl. RDD GroupByKeym. RDD ReduceByKeyn. RDD (Count and CountByValue)o. RDD (saveAsTextFile)p. RDD (Partition)q. Finding Averager. Finding Min and Maxs. Mini project on student data set analysist. Total Marks by Male and Female Studentu. Total Passed and Failed Studentsv. Total Enrollments per Coursew. Total Marks per Coursex. Average marks per Coursey. Finding Minimum and Maximum marksz. Average Age of Male and Female Students4. Spark DFs:a. Introduction to PySpark DFsb. Understanding underlying RDDsc. DFs transformationsd. DFs actionse. Creating Spark DFsf. Spark Infer Schemag. Spark Provide Schemah. Create DF from RDDi. Select DF Columnsj. Spark DF with Columnk. Spark DF with Column Renamed and Aliasl. Spark DF Filter rowsm. Spark DF (Count, Distinct, Duplicate)n. Spark DF (sort, order By)o. Spark DF (Group By)p. Spark DF (UDFs)q. Spark DF (DF to RDD)r. Spark DF (Spark SQL)s. Spark DF (Write DF)t. Mini project on Employees data set analysisu. Project Overviewv. Project (Count and Select)w. Project (Group By)x. Project (Group By, Aggregations, and Order By)y. Project (Filtering)z. Project (UDF and With Column)aa. Project (Write)5. Collaborative filtering:a. Understanding collaborative filteringb. Developing recommendation system using ALS modelc. Utility Matrixd. Explicit and Implicit Ratingse. Expected Resultsf. Datasetg. Joining Dataframesh. Train and Test Datai. ALS modelj. Hyperparameter tuning and cross-validationk. Best model and evaluate predictionsl. Recommendations6. Spark Streaming:a. Understanding the difference between batch and streaming analysis.b. Hands-on with spark streaming through word count examplec. Spark Streaming with RDDd. Spark Streaming Contexte. Spark Streaming Reading Dataf. Spark Streaming Cluster Restartg. Spark Streaming RDD Transformationsh. Spark Streaming DFi. Spark Streaming Displayj. Spark Streaming DF Aggregations7. ETL Pipelinea. Understanding the ETLb. ETL pipeline Flowc. Data setd. Extracting Datae. Transforming Dataf. Loading data (Creating RDS)g. Load data (Creating RDS)h. RDS Networkingi. Downloading Postgresj. Installing Postgresk. Connect to RDS through PgAdminl. Loading Data8. Project - Change Data Capture / Replication On Goinga. Introduction to Projectb. Project Architecturec. Creating RDS MySql Instanced. Creating S3 Buckete. Creating DMS Source Endpointf. Creating DMS Destination Endpointg. Creating DMS Instanceh. MySql WorkBenchi. Connecting with RDS and Dumping Dataj. Querying RDSk. DMS Full Loadl. DMS Replication Ongoingm. Stoping Instancesn. Glue Job (Full Load)o. Glue Job (Change Capture)p. Glue Job (CDC)q. Creating Lambda Function and Adding Triggerr. Checking Triggers. Getting S3 file name in Lambdat. Creating Glue Jobu. Adding Invoke for Glue Jobv. Testing Invokew. Writing Glue Shell Jobx. Full Load Pipeliney. Change Data Capture PipelineAfter the successful completion of this course, you will be able to:● Relate the concepts and practicals of Spark and AWS with real-world problems● Implement any project that requires PySpark knowledge from scratch● Know the theory and practical aspects of PySpark and AWSWho this course is for:● People who are beginners and know absolutely nothing about PySpark and AWS● People who want to develop intelligent solutions● People who want to learn PySpark and AWS● People who love to learn the theoretical concepts first before implementing them using Python● People who want to learn PySpark along with its implementation in realistic projects● Big Data Scientists● Big Data EngineersEnroll in this comprehensive PySpark and AWS course now to master the essential skills in Big Data analytics, data processing, and cloud computing. Whether you're a beginner or looking to expand your knowledge, this course offers a hands-on learning experience with practical projects. Don't miss this opportunity to advance your career and tackle real-world challenges in the world of data analytics and cloud computing. Join us today and start your journey towards becoming a Big Data expert with PySpark and AWS!List of keywords: Big Data analyticsData analysisData cleaningMachine learning (ML)Spark RDDsDataframesSpark SQL queriesSpark ecosystemHadoopDatabricksAWS cloudSpark scriptsAWS servicesPySpark and AWS collaborationPySpark tutorialPySpark hands-onPySpark projectsSpark architectureHadoop ecosystemPySpark Databricks setupSpark local setupSpark RDD transformationsSpark RDD actionsSpark DF transformationsSpark DF actionsSpark Infer SchemaSpark Provide SchemaSpark DF Filter rowsSpark DF (Count, Distinct, Duplicate)Spark DF (sort, order By)Spark DF (Group By)Spark DF (UDFs)Spark DF (Spark SQL)Collaborative filteringRecommendation systemALS modelSpark StreamingETL pipelineChange Data Capture (CDC)ReplicationAWS Glue JobLambda FunctionRDSS3 BucketMySql InstanceData Migration Service (DMS)PgAdminSpark Shell JobFull Load PipelineChange Data Capture Pipeline

课程标签

0人关注该课程

主题相关的课程