Taming Big Data with Apache Spark 4 and Python - Hands On!

所在平台: Udemy

课程主页: https://www.udemy.com/course/taming-big-data-with-apache-spark-hands-on/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:用Apache Spark 4和Python驾驭大数据 - 实践课程! 课程概述:全新更新,涵盖Spark 4的最新功能。本课程将教您大数据分析的热门技术——Apache Spark,特别是PySpark。包括亚马逊、eBay、NASA JPL和雅虎在内的众多雇主都利用Spark快速从庞大的数据集中提取有价值的信息。您将在家中的Windows系统上学习这些技术,掌握将数据分析问题转化为Spark问题的艺术,通过超过20个实践案例,进一步了解如何在云计算服务上扩展分析任务。 课程内容涵盖: - Spark的DataFrames和弹性分布式数据存储(RDD)的概念 - 使用Python和pyspark快速开发和运行Spark作业 - 将复杂分析问题转换为迭代或多阶段的Spark脚本 - 利用亚马逊的弹性MapReduce服务扩展大型数据集 - 理解Hadoop YARN如何在计算集群中分配Spark任务 - 学习其他Spark技术,如Spark SQL、Spark Streaming和GraphX - 实践使用Spark的最新功能,包括Pandas-On-Spark、Spark Connect和用户定义的表函数(UDTFs) 课程的实践性非常强,您将与讲师一同编写、分析和运行真实代码,不论是在您自己的系统上还是在亚马逊的弹性MapReduce服务中。课程包括8小时的视频内容,涵盖40多个复杂性逐步提升的真实案例,您可以根据自己的节奏学习。课程最后还将概述其他基于Spark的技术,包括Spark SQL、Spark结构化流和GraphX。 本课程非常适合希望在技术世界中掌握大数据处理能力的学习者,欢迎报名参与!

课程评论(0条)

课程详情

New! Updated for Spark 4's newest features"Big data" analysis is a hot and highly valuable skill - and this course will teach you the hottest technology in big data: Apache Spark and specifically PySpark. Employers including Amazon, EBay, NASA JPL, and Yahoo all use Spark to quickly extract meaning from massive data sets across a fault-tolerant Hadoop cluster. You'll learn those same techniques, using your own Windows system right at home. It's easier than you might think.Learn and master the art of framing data analysis problems as Spark problems through over 20 hands-on examples, and then scale them up to run on cloud computing services in this course. You'll be learning from an ex-engineer and senior manager from Amazon and IMDb.Learn the concepts of Spark's DataFrames and Resilient Distributed DatastoresDevelop and run Spark jobs quickly using Python and pysparkTranslate complex analysis problems into iterative or multi-stage Spark scriptsScale up to larger data sets using Amazon's Elastic MapReduce serviceUnderstand how Hadoop YARN distributes Spark across computing clustersLearn about other Spark technologies, like Spark SQL, Spark Streaming, and GraphXPractice using Spark's latest features, including Pandas-On-Spark, Spark Connect, and User-Defined Table Functions (UDTFs).By the end of this course, you'll be running code that analyzes gigabytes worth of information - in the cloud - in a matter of minutes. This course uses the familiar Python programming language; if you'd rather use Scala to get the best performance out of Spark, see my "Apache Spark with Scala - Hands On with Big Data" course instead.We'll have some fun along the way. You'll get warmed up with some simple examples of using Spark to analyze movie ratings data and text in a book. Once you've got the basics under your belt, we'll move to some more complex and interesting tasks. We'll use a million movie ratings to find movies that are similar to each other, and you might even discover some new movies you might like in the process! We'll analyze a social graph of superheroes, and learn who the most "popular" superhero is - and develop a system to find "degrees of separation" between superheroes. Are all Marvel superheroes within a few degrees of being connected to The Incredible Hulk? You'll find the answer.This course is very hands-on; you'll spend most of your time following along with the instructor as we write, analyze, and run real code together - both on your own system, and in the cloud using Amazon's Elastic MapReduce service. 8 hours of video content is included, with over 40 real examples of increasing complexity you can build, run and study yourself. Move through them at your own pace, on your own schedule. The course wraps up with an overview of other Spark-based technologies, including Spark SQL, Spark Structured Streaming, and GraphX.Wrangling big data with Apache Spark is an important skill in today's technical world. Enroll now!" I studied "Taming Big Data with Apache Spark and Python" with Frank Kane, and helped me build a great platform for Big Data as a Service for my company. I recommend the course! " - Cleuton Sampaio De Melo Jr."Awesome course on running big data jobs on Apache Spark using Python. As usual, Frank explains things very clearly and points out various items to watch out for and make sure you have set up correctly. There are many ways that a Spark job can fail or have issues, such as running out of memory, and Frank does a great job of pointing many of those out." -James Gershfiel"Easy steps so even a beginner should be able to install Spark and run the examples right away. Good examples and fun to do. Giving a nice set of useful examples as a toolbox." - HansEV"Great course to get you going with Apache Spark and Python! Frank's delivery is very thorough yet unpretentious; his explanations for each new concept that he introduces is down to earth and easy to follow." - Amiri McCain

课程标签

0人关注该课程

主题相关的课程