|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/the-total-data-quality-framework
课程评论:没有评论
课程名称:全面数据质量框架 课程概述: 在全面数据质量专业化的第一门课程结束时,学习者将能够: 1. 识别设计数据和收集数据之间的基本差异,并总结全面数据质量(TDQ)框架的关键维度; 2. 定义全面数据质量框架的三种测量维度,并描述在这些维度上可能影响数据质量的威胁,适用于收集数据和设计数据; 3. 定义全面数据质量框架的三种呈现维度,并描述在这些维度上可能影响数据质量的威胁,适用于收集数据和设计数据; 4. 说明为什么数据分析是全面数据质量框架的重要维度,并总结设计和/或收集数据的分析计划中可能存在的威胁。 本专项课程旨在深入探讨全面数据质量框架,并为学习者提供关于在进行数据分析之前必须进行的全面数据质量评估的更多信息。目标是让学习者在其项目过程中将数据质量评估纳入关键组成部分。我们希望向所有学习者(例如数据科学家和定量分析师)传授全面数据质量的知识,这些学习者可能在数据收集和评估数据质量的初始步骤方面没有足够的培训。我们认为,如果所收集/收集的数据质量不足,即使具备广泛的数据科学技术和统计分析程序的知识,也无法使定量研究有效。 本专业化课程将重点关注任何数据科学研究的基本第一步:生成或收集数据、了解数据来源、评估数据质量,并在进行任何形式的统计分析或应用数据科学技术来解决研究问题之前采取措施来最大化数据质量。鉴于此,课程将很少涉及数据分析的内容,相关内容在其他现有的Coursera专业化课程中已有大量覆盖。该专业化课程的主要重点是理解和最大化数据质量,确保在分析之前做好的准备。 课程大纲: 第一部分:介绍、不同类型的数据与全面数据质量框架 描述:欢迎参加全面数据质量框架课程!本周,您将通过查看课程大纲和学习目标来认识您的讲师。接下来,我们将通过一系列视频讲座介绍全面数据质量(TDQ)框架的基本组成部分,包括设计数据、收集数据和混合数据。随后,我们将提供TDQ框架的高层概述,并融入全球TDQ专家的见解,进行讲座和访谈。最后,我们将通过小测验来结束本周的学习,内容包括测量和呈现概念。 第二部分:全面数据质量的测量维度:有效性、数据来源和数据处理 描述:在第二周中,我们将探讨有效性、数据来源和数据处理的概念。首先,我们将定义有效性,并讨论设计数据和收集数据的有效性威胁。我们还将通过访谈、实际应用和案例研究来进一步探讨有效性。在知识测试后,我们将进入数据来源模块,定义数据处理并通过一系列视频讲座和案例研究探讨设计数据和收集数据的数据来源威胁。第二周的最后将通过小测验总结数据处理模块。 第三部分:全面数据质量的呈现维度:数据访问、数据源和数据缺失 描述:本周,我们将探索TDQ框架的三个呈现维度及其潜在的数据质量威胁。首先,我们将定义和讨论数据访问及其对收集数据和设计数据的威胁,通过视频讲座、阅读材料和案例研究进行解释。在针对数据访问的小测验后,我们将定义数据源,并探讨针对设计和收集数据的数据威胁,以及两个案例研究。最后,我们将定义数据缺失,以及对设计数据和收集数据的缺失威胁,周末将通过小测验结束本周的学习。 第四部分:数据分析作为TDQ的重要方面 描述:本周,我们将结束全面数据质量框架课程。我们将讨论为什么数据分析是TDQ框架的关键维度,以及设计数据和收集数据的数据分析质量威胁。您还将回顾几个案例研究,并能够完成一个使用免费R软件的选修教程。在针对数据分析威胁的小测验后,我们将总结第一课程中的参考资料,并请求您完成课程调查。
Part: 1
Title:Introduction, Different Types of Data and the Total Data Quality Framework
Description:Welcome to the Total Data Quality Framework Course! This is the first course in the Total Data Quality Specialization. This week, you’ll get to know your instructors after reviewing the course syllabus and the learning goals. We will then introduce you to the basic components of the Total Data Quality (TDQ) Framework through a series of video lectures, including Designed Data, Gathered Data, and Hybrid Data. Next, we’ll provide a high-level overview of the TDQ Framework and incorporate the perspectives of global TDQ experts in both a lecture and an interview. We’ll then wrap up the week with a short quiz about measurement and representation concepts.
Part: 2
Title:Measurement Dimensions of Total Data Quality: Validity, Data Origin, and Data Processing
Description:In Week 2, we’ll explore the concepts of validity, data origin, and data processing. First, we’ll define validity and discuss threats to validity for designed data and gathered data. We’ll also explore validity through an interview, a real-world application, and a case study. After taking a short quiz to test your knowledge of validity, you’ll then move to the data origin module. We’ll define data processing and explore data origin threats for designed and gathered data through a series of video lectures and case studies. The data processing module will conclude with a short quiz. Week 2 will conclude with an exploration of data processing; data processing threats for designed and gathered data; case studies; and a quiz to check your understanding of data processing.
Part: 3
Title:Representation Dimensions of Total Data Quality: Data Access, Data Source, and Data Missingness
Description:This week, we’ll be exploring three representation dimensions of the TDQ framework along with potential threats to data quality. First, we’ll define and discuss data access - as well as data access threats for gathered and designed data - through a series of video lectures, readings, and case studies. After you complete a quiz on data access, we’ll then define data sources and explore data threats for designed and gathered data, along with two case studies. Lastly, we’ll define data missingness along with data missingness threats for designed and gathered data, and then conclude the week with a quiz.
Part: 4
Title:Data Analysis as an Important Aspect of TDQ
Description:We’ll be wrapping up the Total Data Quality Framework course this week. We’ll be discussing why data analysis is a critical dimension of the TDQ framework and threats to data analysis quality for designed and gathered data. You’ll also be reviewing several case studies and will be able to complete an optional tutorial using free R software. After a short quiz on data analysis threats, we’ll conclude the course with a list of references from across Course 1 and we’ll ask you to complete a course survey.
By the end of this first course in the Total Data Quality specialization, learners will be able to: 1. Identify the essential differences between designed and gathered data and summarize the key dimensions of the Total Data Quality (TDQ) Framework; 2. Define the three measurement dimensions of the Total Data Quality framework, and describe potential threats to data quality along each of these dimensions for both gathered and designed data; 3. Define the three representation dimensions of the Total Data Quality framework, and describe potential threats to data quality along each of these dimensions for both gathered and designed data; and 4. Describe why data analysis defines an important dimension of the Total Data Quality framework, and summarize potential threats to the overall quality of an analysis plan for designed and/or gathered data. This specialization as a whole aims to explore the Total Data Quality framework in depth and provide learners with more information about the detailed evaluation of total data quality that needs to happen prior to data analysis. The goal is for learners to incorporate evaluations of data quality into their process as a critical component for all projects. We sincerely hope to disseminate knowledge about total data quality to all learners, such as data scientists and quantitative analysts, who have not had sufficient training in the initial steps of the data science process that focus on data collection and evaluation of data quality. We feel that extensive knowledge of data science techniques and statistical analysis procedures will not help a quantitative research study if the data collected/gathered are not of sufficiently high quality. This specialization will focus on the essential first steps in any type of scientific investigation using data: either generating or gathering data, understanding where the data come from, evaluating the quality of the data, and taking steps to maximize the quality of the data prior to performing any kind of statistical analysis or applying data science techniques to answer research questions. Given this focus, there will be little material on the analysis of data, which is covered in myriad existing Coursera specializations. The primary focus of this specialization will be on understanding and maximizing data quality prior to analysis.