|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/ibm-ai-workflow-feature-engineering-bias-detection
课程评论:没有评论
课程名称:AI 工作流程:特征工程与偏见检测 概述:本课程是IBM AI企业工作流程认证专业中的第三门课程。强烈建议按照顺序完成这些课程,因为它们并不是独立的课程,而是构成工作流程的一部分,每门课程都建立在之前的基础上。本课程将介绍我们假设媒体公司的工作流程的下一阶段。在这一阶段,您将学习特征工程的最佳实践、处理类别不平衡以及数据偏见检测等内容。类别不平衡可能严重影响机器学习模型的有效性,而减轻数据中的偏见对于降低偏见模型的风险至关重要。接下来将学习降维、离群点检测和无监督学习技术等最佳实践,以寻找数据中的模式。案例研究将集中在主题建模和数据可视化上。 完成本课程后,您将能够: 1. 使用工具解决类别和类别不平衡问题 2. 解释有关数据偏见的伦理考虑 3. 使用开源库AI Fairness 360检测模型中的偏见 4. 在探索性数据分析和转换阶段应用降维技术 5. 描述自然语言处理中的主题建模技术 6. 使用主题建模和可视化探索文本数据 7. 在高维数据中应用离群点处理最佳实践 8. 将离群点检测算法作为质量保证工具和建模工具 9. 在AI工作流程中使用管道的无监督学习技术 10. 采用基本聚类算法 适合人群:本课程面向已有数据科学从业经验的专业人士,他们希望深化在大型企业中构建和部署AI的技能。如果您是一个有抱负的数据科学家,那么本课程可能不适合您,因为您需要具备真实的专业知识才能从这些课程的内容中受益。 所需技能:假设您已经完成了IBM AI企业工作流程专业的第一和第二课程,并且在开始本课程之前,您对以下主题有扎实的理解:线性代数的基本知识;了解抽样、概率理论和概率分布;掌握描述统计和推断统计概念;对机器学习技术及最佳实践有一般理解;熟练掌握Python及数据科学常用包:NumPy、Pandas、matplotlib和scikit-learn;熟悉IBM Watson Studio;了解设计思维过程。 课程大纲: 第1部分:数据变换与特征工程 描述:本模块将介绍在现代企业中进行有效特征工程所需的技能,这些技能以最佳实践的形式展示,代表了多年的实践经验。 第2部分:模式识别与数据挖掘最佳实践 描述:本模块将继续讨论与特征工程相关的技能,重点关注离群点及使用无监督学习技术寻找模式的方法。
Part: 1
Title:Data transforms and feature engineering
Description:This module will introduce you to skills required for effective feature engineering in today's business enterprises. The skills are presented as a series of best practices representing years of practical experience.
Part: 2
Title:Pattern recognition and data mining best practices
Description:This module will continue the discussion of skill related to feature engineering for practicing data scientists, with a focus on outliers and the use of unsupervised learning techniques for finding patterns.
This is the third course in the IBM AI Enterprise Workflow Certification specialization. You are STRONGLY encouraged to complete these courses in order as they are not individual independent courses, but part of a workflow where each course builds on the previous ones. Course 3 introduces you to the next stage of the workflow for our hypothetical media company. In this stage of work you will learn best practices for feature engineering, handling class imbalances and detecting bias in the data. Class imbalances can seriously affect the validity of your machine learning models, and the mitigation of bias in data is essential to reducing the risk associated with biased models. These topics will be followed by sections on best practices for dimension reduction, outlier detection, and unsupervised learning techniques for finding patterns in your data. The case studies will focus on topic modeling and data visualization. By the end of this course you will be able to: 1. Employ the tools that help address class and class imbalance issues 2. Explain the ethical considerations regarding bias in data 3. Employ ai Fairness 360 open source libraries to detect bias in models 4. Employ dimension reduction techniques for both EDA and transformations stages 5. Describe topic modeling techniques in natural language processing 6. Use topic modeling and visualization to explore text data 7. Employ outlier handling best practices in high dimension data 8. Employ outlier detection algorithms as a quality assurance tool and a modeling tool 9. Employ unsupervised learning techniques using pipelines as part of the AI workflow 10. Employ basic clustering algorithms Who should take this course? This course targets existing data science practitioners that have expertise building machine learning models, who want to deepen their skills on building and deploying AI in large enterprises. If you are an aspiring Data Scientist, this course is NOT for you as you need real world expertise to benefit from the content of these courses. What skills should you have? It is assumed that you have completed Courses 1 and 2 of the IBM AI Enterprise Workflow specialization and you have a solid understanding of the following topics prior to starting this course: Fundamental understanding of Linear Algebra; Understand sampling, probability theory, and probability distributions; Knowledge of descriptive and inferential statistical concepts; General understanding of machine learning techniques and best practices; Practiced understanding of Python and the packages commonly used in data science: NumPy, Pandas, matplotlib, scikit-learn; Familiarity with IBM Watson Studio; Familiarity with the design thinking process.