|
所在平台: Coursera |
课程评论:没有评论
课程名称:数据科学应用中的统计推断与假设检验 课程概述:该课程将重点关注假设检验的理论和实施,尤其是其在数据科学中的应用。学生将学习如何利用假设检验从数据中做出明智决策。课程特别关注假设检验的一般逻辑、错误及其发生率、检验的能力、模拟,以及p值的正确计算和解释。同时,也将关注假设检验概念的误用,特别是p值的误用及其伦理影响。 该课程可作为CU Boulder的数据科学硕士学位(MS-DS)学术学分的一部分在Coursera平台上修读。MS-DS是一个跨学科的学位项目,汇集了CU Boulder应用数学、计算机科学、信息科学等多个学科的教师。该项目采用基于绩效的招生方式,无需申请过程,适合拥有计算机科学、信息科学、数学和统计学广泛背景的本科教育和/或专业经验的个人。更多信息请访问:https://www.coursera.org/degrees/master-of-science-data-science-boulder。 课程大纲: 1. 开始这里!:欢迎来到课程!本模块包含帮助你开始学习的后勤信息。 2. 假设检验的基本概念:定义假设检验,并发展设计检验的直觉。学习假设检验的语言,包括原假设、备择假设和显著性水平的定义。 3. 复合检验、检验的功效与p值:扩展模块1的内容,介绍复合假设的一尾和双尾检验,定义“功效函数”并讨论其解释,以及“p值”的替代方法。 4. t检验与两样本检验:学习卡方分布和t分布及其与抽样分布的关系,识别何时适便使用这些分布的假设检验,推导“t检验”和第一项两样本检验并应用于实际数据。 5. 超越正态性:考虑假设底层正态分布不适合的情况,构建假设检验,定义“均匀最强”(UMP) 检验,讨论特定问题的检验存在与否。 6. 似然比检验与卡方检验:基于“似然比”开发形式化的假设检验方法,特别关注大量样本性质的似然比及Wilks定理,最后通过卡方检验测试课程中做的假设是否有效。 该课程适合希望加深统计推断和假设检验知识的学生,尤其是那些在数据科学领域工作的专业人士。
Name:Start Here!
Description:Welcome to the course! This module contains logistical information to get you started!
Name:Fundamental Concepts of Hypothesis Testing
Description:In this module, we will define a hypothesis test and develop the intuition behind designing a test. We will learn the language of hypothesis testing, which includes definitions of a null hypothesis, an alternative hypothesis, and the level of significance of a test. We will walk through a very simple test.
Name:Composite Tests, Power Functions, and P-Values
Description:In this module, we will expand the lessons of Module 1 to composite hypotheses for both one and two-tailed tests. We will define the “power function” for a test and discuss its interpretation and how it can lead to the idea of a “uniformly most powerful” test. We will discuss and interpret “p-values” as an alternate approach to hypothesis testing.
Name:t-Tests and Two-Sample Tests
Description:In this module, we will learn about the chi-squared and t distributions and their relationships to sampling distributions. We will learn to identify when hypothesis tests based on these distributions are appropriate. We will review the concept of sample variance and derive the “t-test”. Additionally, we will derive our first two-sample test and apply it to make some decisions about real data.
Name:Beyond Normality
Description:In this module, we will consider some problems where the assumption of an underlying normal distribution is not appropriate and will expand our ability to construct hypothesis tests for this case. We will define the concept of a “uniformly most powerful” (UMP) test, whether or not such a test exists for specific problems, and we will revisit some of our earlier tests from Modules 1 and 2 through the UMP lens. We will also introduce the F-distribution and its role in testing whether or not two population variances are equal.
Name:Likelihood Ratio Tests and Chi-Squared Tests
Description:In this module, we develop a formal approach to hypothesis testing, based on a “likelihood ratio” that can be more generally applied than any of the tests we have discussed so far. We will pay special attention to the large sample properties of the likelihood ratio, especially Wilks’ Theorem, that will allow us to come up with approximate (but easy) tests when we have a large sample size. We will close the course with two chi-squared tests that can be used to test whether the distributional assumptions we have been making throughout this course are valid.
This course will focus on theory and implementation of hypothesis testing, especially as it relates to applications in data science. Students will learn to use hypothesis tests to make informed decisions from data. Special attention will be given to the general logic of hypothesis testing, error and error rates, power, simulation, and the correct computation and interpretation of p-values. Attention will also be given to the misuse of testing concepts, especially p-values, and the ethical implications of such misuse. This course can be taken for academic credit as part of CU Boulder’s Master of Science in Data Science (MS-DS) degree offered on the Coursera platform. The MS-DS is an interdisciplinary degree that brings together faculty from CU Boulder’s departments of Applied Mathematics, Computer Science, Information Science, and others. With performance-based admissions and no application process, the MS-DS is ideal for individuals with a broad range of undergraduate education and/or professional experience in computer science, information science, mathematics, and statistics. Learn more about the MS-DS program at https://www.coursera.org/degrees/master-of-science-data-science-boulder.