|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/ibm-ai-workflow-data-analysis-hypothesis-testing
课程评论:没有评论
课程名称:AI工作流程:数据分析与假设检验 课程概述:这是IBM AI企业工作流程认证专项中的第二门课程。强烈建议按顺序完成这些课程,因为它们并不是独立的课程,而是一个逐步构建的工作流程。本课程将以一个假设的流媒体公司为背景,开始进行探索性数据分析(EDA)。课程中将介绍数据可视化、处理缺失数据和假设检验的最佳实践。您将学习使用概率分布进行估计的技术,并将这些估计扩展应用于零假设显著性检验。通过两个实际案例研究,您将运用所学内容:数据可视化和使用简单管道进行多重检验。 完成本课程后,您应能够: 1. 列出关于EDA和数据可视化的几项最佳实践 2. 在Watson Studio中创建一个简单的仪表板 3. 描述处理缺失数据的策略 4. 解释插补与多重插补之间的区别 5. 使用常见分布回答事件概率问题 6. 解释假设检验在EDA中的调查角色 7. 应用多重检验处理的几种方法 适合人群:本课程针对已有机器学习模型构建经验的数据科学从业者,旨在深化在大型企业中构建和部署AI的技能。如果您是有志于成为数据科学家的人士,本课程可能不适合您,因为您需要具备相关的实际经验才能从中受益。 应具备的技能:课程假设您已完成IBM AI企业工作流程专项的第一门课程,并且对以下主题有扎实的理解:线性代数的基本理解;了解采样、概率论和概率分布;掌握描述性和推断性统计概念;对机器学习技术及最佳实践有一般理解;熟练掌握Python及在数据科学中常用的包,如NumPy、Pandas、matplotlib、scikit-learn;熟悉IBM Watson Studio;了解设计思维过程。 课程大纲: 第一部分:数据分析 描述:探索性数据分析主要通过可视化和假设检验获取洞察。本单元关注EDA、数据可视化及缺失值处理。不同的缺失值处理策略可能适用于不同模型,并会影响预测性能。 第二部分:数据调查 描述:数据科学家使用多种统计工具分析数据并得出结论。本单元重点介绍使用概率分布进行估计的基本技术,以及如何将这些估计扩展应用于零假设显著性检验。
Part: 1
Title:Data Analysis
Description:Exploratory data analysis is mostly about gaining insight through visualization and hypothesis testing. This unit looks at EDA, data visualization, and missing values. One missing value strategy may be better for some models, but for others another strategy may show better predictive performance.
Part: 2
Title:Data Investigation
Description:Data scientists employ a broad range of statistical tools to analyze data and reach conclusions from data. This unit focuses on the foundational techniques of estimation with probability distributions and extending these estimates to apply null hypothesis significance tests.
This is the second course in the IBM AI Enterprise Workflow Certification specialization. You are STRONGLY encouraged to complete these courses in order as they are not individual independent courses, but part of a workflow where each course builds on the previous ones. In this course you will begin your work for a hypothetical streaming media company by doing exploratory data analysis (EDA). Best practices for data visualization, handling missing data, and hypothesis testing will be introduced to you as part of your work. You will learn techniques of estimation with probability distributions and extending these estimates to apply null hypothesis significance tests. You will apply what you learn through two hands on case studies: data visualization and multiple testing using a simple pipeline. By the end of this course you should be able to: 1. List several best practices concerning EDA and data visualization 2. Create a simple dashboard in Watson Studio 3. Describe strategies for dealing with missing data 4. Explain the difference between imputation and multiple imputation 5. Employ common distributions to answer questions about event probabilities 6. Explain the investigative role of hypothesis testing in EDA 7. Apply several methods for dealing with multiple testing Who should take this course? This course targets existing data science practitioners that have expertise building machine learning models, who want to deepen their skills on building and deploying AI in large enterprises. If you are an aspiring Data Scientist, this course is NOT for you as you need real world expertise to benefit from the content of these courses. What skills should you have? It is assumed that you have completed Course 1 of the IBM AI Enterprise Workflow specialization and have a solid understanding of the following topics prior to starting this course: Fundamental understanding of Linear Algebra; Understand sampling, probability theory, and probability distributions; Knowledge of descriptive and inferential statistical concepts; General understanding of machine learning techniques and best practices; Practiced understanding of Python and the packages commonly used in data science: NumPy, Pandas, matplotlib, scikit-learn; Familiarity with IBM Watson Studio; Familiarity with the design thinking process.