|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/crisp-mlq-business-understanding-and-data-understanding/
课程评论:没有评论
课程名称:CRISP-ML(Q) - 商业理解与数据理解 课程概述:本课程旨在帮助您理解数据科学和探索性数据分析(EDA)的基本概念,使用Python作为工具。同时,我们将深入探讨项目管理方法论CRISP-ML(Q)(跨行业机器学习标准流程与质量保证)。数据科学在各个行业中无处不在,其目的在于通过各种技术从可用数据中发现趋势和模式。数据科学家还需在分析数据后提取洞见。数据科学是一个多学科领域,涉及数学、统计学、计算机科学、Python、机器学习等,数据科学家需要熟练掌握这些主题。 本课程将详细解释CRISP-ML(Q)的六个阶段: 1. 商业与数据理解 2. 数据准备 3. 模型构建 4. 评估 5. 模型部署 6. 监控与维护 在课程中,我们将深入了解商业目标与约束、商业成功标准、经济成功标准和项目章程的意义。同时,还将详细描述各种数据类型,包括连续数据、离散数据、定性数据、定量数据、结构化数据、半结构化数据、非结构化数据、大数据和非大数据、横截面数据、时间序列数据、面板数据、平衡数据和不平衡数据,以及离线数据与实时流数据。我们将探讨数据收集的各个方面,包括初级数据和次级数据、数据版本控制、描述、需求和验证。 数据准备部分将详细说明数据清洗、使用Python进行EDA或描述性统计以及特征工程。数据清洗涉及多种方法,如类型转换、重复处理、异常值处理、零和近零方差、缺失值、离散化、虚拟变量、转换、标准化和字符串处理。在Python EDA中,我们将学习集中趋势的测量(均值、中位数、众数)、离散程度的测量(方差、标准差、范围)、偏度和峭度等,使用条形图、Q-Q图、箱线图、直方图、散点图等进行可视化。 模型构建(也称为数据挖掘或机器学习)将得到详细讨论,涵盖监督学习、无监督学习和预测等内容,并将探索一些模型构建技术,如简单线性回归、多元线性回归、逻辑回归、决策树和朴素贝叶斯等。CRISP-ML(Q)的后续步骤,也是课程最后部分,将包括评估、模型部署及监控与维护。 通过本课程的学习,您将获得使用Python进行数据科学和EDA的全面理解,从而在数据科学领域建立职业生涯。
This course will help you understand the basics of Data Science and EDA using Python and we shall also dive deep into the Project Management Methodology, CRISP-ML(Q). Cross-Industry Standard Process for Machine Learning with Quality Assurance is abbreviated as CRISP-ML(Q). Data Science is omnipresent in every sector. The purpose of Data Science is to find trends and patterns with the data that is available through various techniques. Data Scientists are also responsible for drawing insights after analyzing data. Data Science is a multidisciplinary field that involves mathematics, statistics, computer science, Python, machine learning, etc. Data Scientists need to be adept in these topics. This course will provide you with an understanding of all the aforementioned topics.A detailed explanation of the 6 stages of CRISP-ML(Q) will be provided. These 6 stages are as follows:Business and Data UnderstandingData PreparationModel BuildingEvaluationModel DeploymentMonitoring & MaintenanceThe importance of Business objectives and constraints, Business success criteria, Economic success criteria, and Project charter will be thoroughly understood. Elaborate descriptions of various data types - continuous, discrete, qualitative, quantitative, structured, semi-structured, unstructured, big, and non-big data, cross-sectional, time series and panel data, balanced and unbalanced data, and finally, offline and live streaming data. Various aspects of data collection will be looked into. Primary, and secondary, data version control, description, requirements, and verification will be analyzed.Data Preparation involving data cleansing, EDA using Python or descriptive statistics, and feature engineering will be elaborately explained. Data cleansing involves numerous methods like typecasting, handling duplicates, outlier treatment, zero & near zero variance, missing values, discretization, dummy variables, transformation, standardization, and string manipulation. The realm of EDA using Python will be explored, This would include understanding measures of central tendency (mean, median, and mode), measures of dispersion (variance, standard deviation, and range), skewness, and kurtosis which are also termed first, second, third and fourth-moment business decisions. More about bar plots, Q-Q plots, box plots, histograms, scatter plots, etc., will be looked into in EDA using Python. Feature engineering, the last part of data cleansing, will also be given enough coverage.Further, the model building also known as data mining or machine learning will also be thoroughly talked about. Model building involves supervised learning, unsupervised learning, and, forecasting which will be explored. Several model-building techniques like Simple Linear regression, Multilinear regression, Logistic regression, Decision-Tree, Naive Bayes, etc.The last few steps of CRISP-ML(Q) are Evaluation, Model Deployment, and Monitoring & Maintenance.The learning journey will include CRISP-ML(Q) using Python & Data Science and EDA using Python. Having a thorough understanding of these topics will enable you to build a career in the field of data science.