|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-science-hands-on-diabetes-prediction-with-pyspark-mllib/
课程评论:没有评论
课程名称:数据科学:Pyspark MLlib 实战糖尿病预测 课程概述:您是否想构建、训练、测试和评估一个能够使用逻辑回归检测糖尿病的机器学习模型?这是一个实践性的机器学习课程,您将与课堂内容同步进行练习。课程中会提供数据集,建议您在学习过程中与课堂一起实践,这将使您在短短一小时的实践中获得比数百小时理论课更扎实的知识。 在本课程中,您将学习到 Spark 机器学习(Spark MLlib)的最重要方面,包括 Pyspark 基础知识和实现 Spark 机器学习、导入和处理数据集以及使用 Spark MLlib 处理数据、构建和训练逻辑回归模型、测试和分析模型。整个课程被分为几个任务,每个任务都经过精心设计,以提供最佳的学习体验。 实践项目将包括以下任务: 1. 项目概述 2. 介绍 Colab 环境及安装 Spark 依赖 3. 克隆并探索糖尿病数据集 4. 数据清理 5. 相关性与特征选择 6. 使用 Spark MLlib 构建和训练逻辑回归模型 7. 性能评估与模型测试 8. 保存与加载模型 关于 Pyspark:Pyspark 是 Apache Spark 与 Python 的结合,是大数据分析中的一种工具。Apache Spark 是一个开源的集群计算框架,旨在提供高速、易用和流式分析,而 Python 是一种通用的高级编程语言,广泛用于机器学习和实时流式分析。Pyspark 使我们能够利用 Python 的简单性和 Apache Spark 的强大功能,来处理大数据。在这个项目中,我们将使用大数据工具,踏入数据科学的世界,通过这个 Spark MLlib 项目展示您的技能,为您的简历增添亮点。 立即点击“注册”按钮,开始学习。祝您学习愉快!
Would you like to build, train, test and evaluate a machine learning model that is able to detect diabetes using logistic regression?This is a Hands-on Machine Learning Course where you will practice alongside the classes. The dataset will be provided to you during the lectures. We highly recommend that for the best learning experience, you practice alongside the lectures. You will learn more in this one hour of Practice than hundreds of hours of unnecessary theoretical lectures.Learn the most important aspect of Spark Machine learning (Spark MLlib):Pyspark fundamentals and implementing spark machine learningImporting and Working with DatasetsProcess data using a Machine Learning model using spark MLlibBuild and train Logistic regression modelTest and analyze the modelThe entire course has been divided into tasks. Each task has been very carefully created and designed to give you the best learning experience. In this hands-on project, we will complete the following tasks:Task 1: Project overviewTask 2: Intro to Colab environment & install dependencies to run spark on ColabTask 3: Clone & explore the diabetes datasetTask 4: Data CleaningTask 5: Correlation & feature selectionTask 6: Build and train Logistic Regression Model using Spark MLlibTask 7: Performance evaluation & Test the modelTask 8: Save & load modelAbout Pyspark:Pyspark is the collaboration of Apache Spark and Python. PySpark is a tool used in Big Data Analytics.Apache Spark is an open-source cluster-computing framework, built around speed, ease of use, and streaming analytics whereas Python is a general-purpose, high-level programming language. It provides a wide range of libraries and is majorly used for Machine Learning and Real-Time Streaming Analytics.In other words, it is a Python API for Spark that lets you harness the simplicity of Python and the power of Apache Spark in order to tame Big Data. We will be using Big data tools in this project. Make a leap into Data science with this Spark MLlib project and showcase your skills on your resume.Click on the "ENROLL NOW" button and start learning. Happy Learning.