Scalable Machine Learning on Big Data using Apache Spark

所在平台: CourseraArchive

课程类别: 其他类别

大学或机构: CourseraNew

课程主页: https://www.coursera.org/archive/machine-learning-big-data-apache-spark

课程评论:没有评论

第一个写评论        关注课程

课程大纲

Week 1: Introduction
Week 2: Scaling Math for Statistics on Apache Spark
Week 3: Introduction to Apache SparkML
Week 4: Supervised and Unsupervised learning with SparkML

课程评论(0条)

课程详情

This course will empower you with the skills to scale data science and machine learning (ML) tasks on Big Data sets using Apache Spark. Most real world machine learning work involves very large data sets that go beyond the CPU, memory and storage limitations of a single computer. Apache Spark is an open source framework that leverages cluster computing and distributed storage to process extremely large data sets in an efficient and cost effective manner. Therefore an applied knowledge of working with Apache Spark is a great asset and potential differentiator for a Machine Learning engineer. After completing this course, you will be able to: - gain a practical understanding of Apache Spark, and apply it to solve machine learning problems involving both small and big data - understand how parallel code is written, capable of running on thousands of CPUs. - make use of large scale compute clusters to apply machine learning algorithms on Petabytes of data using Apache SparkML Pipelines. - eliminate out-of-memory errors generated by traditional machine learning frameworks when data doesn’t fit in a computer's main memory - test thousands of different ML models in parallel to find the best performing one – a technique used by many successful Kagglers - (Optional) run SQL statements on very large data sets using Apache SparkSQL and the Apache Spark DataFrame API. Enrol now to learn the machine learning techniques for working with Big Data that have been successfully applied by companies like Alibaba, Apple, Amazon, Baidu, eBay, IBM, NASA, Samsung, SAP, TripAdvisor, Yahoo!, Zalando and many others. NOTE: You will practice running machine learning tasks hands-on on an Apache Spark cluster provided by IBM at no charge during the course which you can continue to use afterwards. Prerequisites: - basic python programming - basic machine learning (optional introduction videos are provided in this course as well) - basic SQL skills for optional content The following courses are recommended before taking this class (unless you already have the skills) https://www.coursera.org/learn/python-for-applied-data-science or similar https://www.coursera.org/learn/machine-learning-with-python or similar https://www.coursera.org/learn/sql-data-science for optional lectures

使用Apache Spark在大数据上进行可扩展的机器学习:本课程将使您掌握使用Apache Spark在大数据集上扩展数据科学和机器学习(ML)任务的技能。大多数现实世界的机器学习工作都涉及非常大的数据集,这些数据集超出了单台计算机的CPU,内存和存储限制。 Apache Spark是一个开放源代码框架,该框架利用集群计算和分布式存储以高效且经济高效的方式处理超大型数据集。因此,与Apache Spark一起工作的应用知识对于机器学习工程师而言是一项巨大的资产,并且可能成为差异化因素。 完成本课程后,您将能够: -获得对Apache Spark的实践理解,并将其用于解决涉及小数据和大数据的机器学习问题 -了解如何编写可在数千个CPU上运行的并行代码。 -利用Apache SparkML Pipelines使用大型计算集群将机器学习算法应用于PB级数据。 -消除了传统机器学习框架在数据无法容纳在计算机主存储器中时产生的内存不足错误 -并行测试成千上万种不同的ML模型,以找到性能最佳的模型-许多成功的Kaggler都使用了该技术 -(可选)使用Apache SparkSQL和Apache Spark DataFrame API在非常大的数据集上运行SQL语句。 现在注册以学习使用大数据的机器学习技术,这些技术已被阿里巴巴,苹果,亚马逊,百度,eBay,IBM,NASA,三星,SAP,TripAdvisor,Yahoo!,Zalando等公司成功应用。 注意:您将在课程中免费练习由IBM免费提供的Apache Spark集群上的运行机器学习任务,以后您可以继续使用。 先决条件: -基本的python编程 -基本的机器学习(本课程还提供可选的介绍视频) -可选内容的基本SQL技能 建议您在上这堂课之前先学习以下课程(除非您已经具备该技能) https://www.coursera.org/learn/python-for-applied-data-science或类似名称 https://www.coursera.org/learn/machine-learning-with-python或类似内容 https://www.coursera.org/learn/sql-data-science进行可选讲座

课程标签

0人关注该课程

主题相关的课程