Real World Auto Machine Learning Bootcamp: Build 14 Projects

所在平台: Udemy

课程主页: https://www.udemy.com/course/real-world-automated-machine-learning-projects/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:真实世界自动化机器学习实训营:构建14个项目 课程概述:自动化机器学习(AutoML)标志着各类组织在机器学习和数据科学领域方法的根本性转变。传统的机器学习方法在解决实际业务问题时,通常耗时、资源密集且具有挑战性,需要多个学科的专家协同合作,而数据科学家是现今市场上最受追捧的专业人才之一。自动化机器学习则改变了这一局面,使构建和使用机器学习模型变得更加轻松高效,它通过对原始数据进行系统性处理,选择能够从数据中提取最相关信息的模型,常被称之为“噪声中的信号”。自动化机器学习融入了顶级数据科学家的最佳实践,使得数据科学在组织内部更为普及。 数据科学是通过数学和统计方法对数据进行转化,以获得有价值的见解、决策和产品。随着数据科学的发展及新工具的引入,核心商业目标依然是找到有用的模式并从数据中获得有价值的见解。目前,数据科学广泛应用于各行各业,并支持各种分析问题。例如,在营销领域,分析客户的年龄、性别、位置和行为,可以制定高度针对性的营销活动;在银行业,识别异常客户行为有助于防范欺诈;在医疗保健中,分析患者的病历可以预测疾病的可能性等。 数据科学领域涵盖多个相互关联的领域,利用不同的技术和工具。虽然数据挖掘和现在流行的机器学习有所不同,但机器学习专注于创建算法以提取有价值的见解,并强调在动态变化的环境中持续使用、调整、再训练和更新算法。机器学习的目标是不断适应新数据,发现新模式或规则,有时无需人类的引导或明确的重新编程。 机器学习是当今数据科学中发展最迅速的领域,这得益于多个理论和技术的突破。它催生了自然语言处理、图像识别,甚至是机器生成新图像、音乐和文本等应用。机器学习依然是构建人工智能的主要“工具”。 机器学习工作流程通常遵循以下简单步骤: 1. 收集数据:利用数字基础设施及其他来源收集尽可能多的有用记录,并将其整合为数据集。 2. 准备数据:对数据进行最优处理的准备,包括数据预处理和清洗,填补缺失值和纠正数据中的其他瑕疵(例如同一列中对同一值的不同表现)。 3. 划分数据:将数据划分为训练模型的子集以及进一步评估模型性能的新数据集。 4. 训练模型:利用历史数据的子集让算法识别其中的模式。 5. 测试和验证模型:利用历史数据的测试和验证子集评估模型性能,了解预测的准确性。 6. 部署模型:将经过测试的模型嵌入到决策框架中,作为分析解决方案的一部分,或让用户利用其能力(例如优化产品推荐)。 7. 迭代:在使用模型后收集新数据,逐步改进模型。 通过本课程,学员将深入理解自动化机器学习的基本原理和流程,并在实践中完成14个项目,提升在真实业务环境中应用机器学习的能力。

课程评论(0条)

课程详情

Automated machine learning (AutoML) represents a fundamental shift in the way organizations of all sizes approach machine learning and data science. Applying traditional machine learning methods to real-world business problems is time-consuming, resource-intensive, and challenging. It requires experts in several disciplines, including data scientists - some of the most sought-after professionals in the job market right now.Automated machine learning changes that, making it easier to build and use machine learning models in the real world by running systematic processes on raw data and selecting models that pull the most relevant information from the data - what is often referred to as "the signal in the noise." Automated machine learning incorporates machine learning best practices from top-ranked data scientists to make data science more accessible across the organization."Data science is the transformation of data using mathematics and statistics into valuable insights, decisions, and products"As data science evolves and gains new "instruments" over time, the core business goal remains focused on finding useful patterns and yielding valuable insights from data. Today, data science is employed across a broad range of industries and aids in various analytical problems. For example, in marketing, exploring customer age, gender, location, and behavior allows for making highly targeted campaigns, evaluating how much customers are prone to make a purchase or leave. In banking, finding outlying client actions aids in detecting fraud. In healthcare, analyzing patients' medical records can show the probability of having diseases, etc.The data science landscape encompasses multiple interconnected fields that leverage different techniques and tools.There's a difference between data mining and very popular machine learning. Still, machine learning is about creating algorithms to extract valuable insights, it's heavily focused on continuous use in dynamically changing environments and emphasizes adjustments, retraining, and updating of algorithms based on previous experiences. The goal of machine learning is to constantly adapt to new data and discover new patterns or rules in it. Sometimes it can be realized without human guidance and explicit reprogramming.Machine learning is the most dynamically developing field of data science today due to a number of recent theoretical and technological breakthroughs. They led to natural language processing, image recognition, or even the generation of new images, music, and texts by machines. Machine learning remains the main "instrument" of building artificial intelligence.Machine Learning WorkflowGenerally, the workflow follows these simple steps:Collect data. Use your digital infrastructure and other sources to gather as many useful records as possible and unite them into a dataset.Prepare data. Prepare your data to be processed in the best possible way. Data preprocessing and cleaning procedures can be quite sophisticated, but usually, they aim at filling the missing values and correcting other flaws in data, like different representations of the same values in a column (e.g. December 14, 2016 and 12.14.2016 won't be treated the same by the algorithm).Split data. Separate subsets of data to train a model and further evaluate how it performs against new data.Train a model. Use a subset of historic data to let the algorithm recognize the patterns in it.Test and validate a model. Evaluate the performance of a model using testing and validation subsets of historic data and understand how accurate the prediction is.Deploy a model. Embed the tested model into your decision-making framework as a part of an analytics solution or let users leverage its capabilities (e.g. better target your product recommendations).Iterate. Collect new data after using the model to incrementally improve it.

课程标签

0人关注该课程

主题相关的课程