|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-science-in-python-classification/
课程评论:没有评论
课程名称:Python数据科学:分类建模 课程概述:本课程是一个实践为主的项目课程,旨在帮助您掌握Python中分类建模和监督学习的基础知识。我们将首先回顾Python数据科学的工作流程,讨论分类算法的主要目标和类型,并深入探讨在整个课程中将使用的分类建模步骤。您将学习执行探索性数据分析(EDA),利用特征工程技术(如缩放、虚拟变量和分箱),并通过将数据拆分为训练、测试和验证数据集来为建模准备数据。接下来,我们将拟合K-近邻和逻辑回归模型,并培养解释其系数和使用混淆矩阵等工具以及准确率、精确率和召回率等指标评估模型性能的直觉。我们还将涵盖处理不平衡数据的建模技术,包括阈值调整、过采样和SMOTE等采样方法,以及调整模型成本函数中的类权重。 在整个课程中,您将作为Maven National Bank风险管理部门的数据科学家,利用您在课程中学到的技能,使用Python探索数据并构建分类模型,以准确判断客户的信用风险水平(高、中、低)。最后,您将学习构建和评估决策树模型,使用Python拟合、可视化和微调这些模型,然后将知识应用于更高级的集成模型,如随机森林和梯度提升机。 课程大纲: - Python数据科学简介:介绍数据科学与机器学习领域,回顾基本技能,并介绍数据科学工作流程的各个阶段。 - 分类基础:回顾分类的基本概念,包括关键术语、分类建模的类型和目标,以及建模工作流程。 - 数据准备与EDA:回顾进行建模所需的数据准备与探索性数据分析步骤,包括探索目标、特征及其关系的关键技术。 - K-近邻算法:学习K-近邻算法如何对数据点进行分类,并在Python中实践构建KNN模型。 - 逻辑回归:介绍逻辑回归,学习模型背后的数学原理,并实践拟合和调整正则化强度。 - 分类指标:学习何时使用几种重要指标来评估分类模型,如精确率、召回率、F1分数和ROC-AUC。 - 不平衡数据:理解建模不平衡数据的挑战,学习改善模型性能的策略。 - 决策树:构建和评估决策树模型,寻找最优数据分割以区分类别。 - 集成模型:熟悉集成模型的基本概念,深入特定模型如随机森林和梯度提升机。 该课程适合希望通过Python入门分类建模的商务智能专业人士或有志于成为数据科学家的学习者。立即加入,获取终身访问权限,课程包含9.5小时高质量视频、18个作业、9个测验、2个项目及大量额外资源与支持。 快乐学习!- Chris Bruehl (数据科学专家与Maven Analytics首席Python讲师)
This is a hands-on, project-based course designed to help you master the foundations for classification modeling and supervised machine learning in Python.We'll start by reviewing the Python data science workflow, discussing the primary goals & types of classification algorithms, and do a deep dive into the classification modeling steps we'll be using throughout the course.You'll learn to perform exploratory data analysis (EDA), leverage feature engineering techniques like scaling, dummy variables, and binning, and prepare data for modeling by splitting it into train, test, and validation datasets.From there, we'll fit K-Nearest Neighbors & Logistic Regression models, and build an intuition for interpreting their coefficients and evaluating their performance using tools like confusion matrices and metrics like accuracy, precision, and recall. We'll also cover techniques for modeling imbalanced data, including threshold tuning, sampling methods like oversampling & SMOTE, and adjusting class weights in the model cost function.Throughout the course, you'll play the role of Data Scientist for the risk management department at Maven National Bank. Using the skills you learn throughout the course, you'll use Python to explore their data and build classification models to accurately determine which customers have high, medium, and low credit risk based on their profiles.Last but not least, you'll learn to build and evaluate decision tree models for classification. You'll fit, visualize, and fine-tune these models using Python, then apply your knowledge to more advanced ensemble models like random forests and gradient boosted machines.COURSE OUTLINE:Intro to Data Science in PythonIntroduce the fields of data science and machine learning, review essential skills, and introduce each phase of the data science workflowClassification 101Review the basics of classification, including key terms, the types and goals of classification modeling, and the modeling workflowPre-Modeling Data Prep & EDARecap the data prep & EDA steps required to perform modeling, including key techniques to explore the target, features, and their relationshipsK-Nearest NeighborsLearn how the k-nearest neighbors (KNN) algorithm classifies data points and practice building KNN models in PythonLogistic RegressionIntroduce logistic regression, learn the math behind the model, and practice fitting them and tuning regularization strengthClassification MetricsLearn how and when to use several important metrics for evaluating classification models, such as precision, recall, F1 score, and ROC-AUCImbalanced DataUnderstand the challenges of modeling imbalanced data and learn strategies for improving model performance in these scenariosDecision TreesBuild and evaluate decision tree models, algorithms that look for the splits in your data that best separate your classesEnsemble ModelsGet familiar with the basics of ensemble models, then dive into specific models like random forests and gradient boosted machines__________Ready to dive in? Join today and get immediate, LIFETIME access to the following:9.5 hours of high-quality video18 homework assignments9 quizzes2 projectsPython Data Science: Classification ebook (250+ pages)Downloadable project files & solutionsExpert support and Q & A forum30-day Udemy satisfaction guaranteeIf you're a business intelligence professional or aspiring data scientist looking for an introduction to the world of classification modeling with Python, this is the course for you.Happy learning!-Chris Bruehl (Data Science Expert & Lead Python Instructor, Maven Analytics)__________Looking for our full business intelligence stack? Search for "Maven Analytics" to browse our full course library, including Excel, Power BI, MySQL, Tableau and Machine Learning courses!See why our courses are among the TOP-RATED on Udemy:"Some of the BEST courses I've ever taken. I've studied several programming languages, Excel, VBA and web dev, and Maven is among the very best I've seen!" Russ C."This is my fourth course from Maven Analytics and my fourth 5-star review, so I'm running out of things to say. I wish Maven was in my life earlier!" Tatsiana M."Maven Analytics should become the new standard for all courses taught on Udemy!" Jonah M.