Visualization and Imputation of Missing Data

所在平台: Udemy

课程主页: https://www.udemy.com/course/visualization-and-imputation-of-missing-data/

课程评论:没有评论

第一个写评论        关注课程

课程简介

**课程名称:** 对缺失数据进行可视化及插补 **课程概述:** 该课程专注于处理数据集中存在的缺失数据问题。缺失数据在数据分析中会引发诸多挑战,但通过各种插补技术,可以根据数据的自身特征和缺失模式,用合理的估计值来“填充”缺失的数据。 传统的插补技术多适用于多元正态分布数据,而对于非多元正态分布的数据则效果不佳。本课程特别强调理解数据样本中“缺失模式”,尤其是针对非多元正态分布数据集。课程将教授多种适用于这些非正态数据的插补技术,以有效地“填充”缺失数据。 课程中将重点介绍如何使用 R 语言中的 VIM 和 VIMGUI 包,创建数十种独特的视觉化图表,以更深入地理解数据样本中现有数据和插补后数据的模式。 **核心内容包括:** * **理解缺失模式:** 重点关注非多元正态分布数据的缺失模式。 * **插补技术:** * Hot-Deck 插补(顺序和随机)。 * 基于距离的 K-近邻插补。 * 个体回归插补。 * 迭代模型化逐步回归插补(IRMI 算法),包括标准和鲁棒方法。 * **可视化工具:** 使用 VIMGUI 软件创建丰富的缺失数据模式可视化,帮助识别和理解缺失情况。 **适用人群:** 本课程适合任何需要分析可能包含缺失数据的数据集的人员,包括: * 从事实证研究的研究生和教职员工。 * 从事定量研究或数据分析的在职专业人士。 **特别说明:** 课程提供的 VIMGUI 工具栏(使用 R 的 RGtk2 语言开发)在 Mac 操作系统上可能存在启动问题。 **课程目标:** 通过本课程,学员将能够: 1. 理解缺失数据带来的问题及其对分析的影响。 2. 识别和理解不同类型(尤其是非多元正态分布)数据集中的缺失模式。 3. 掌握并应用多种先进的、适用于非多元正态分布数据的插补技术。 4. 熟练使用 VIM 和 VIMGUI 包进行数据可视化,以辅助理解和解决缺失数据问题。 5. 具备选择和应用最适合数据情况的缺失数据处理方法的能力。

课程评论(0条)

课程详情

There are many problems associated with analyzing data sets that contain missing data. However, there are various techniques to 'fill in,' or impute, missing data values with reasonable estimates based on the characteristics of the data itself and on the patterns of 'missingness.' Generally, techniques appropriate for imputing missing values in multivariate normal data and not as useful when applied to non-multivariate-normal data. This Visualization and Imputation of Missing Data course focuses on understanding patterns of 'missingness' in a data sample, especially non-multivariate-normal data sets, and teaches one to use various appropriate imputation techniques to "fill in" the missing data. Using the VIM and VIMGUI packages in R, the course also teaches how to create dozens of different and unique visualizations to better understand existing patterns of both the missing and imputed data in your samples. The course teaches both the concepts and provides software to apply the latest non-multivariate-normal-friendly data imputation techniques, including: (1) Hot-Deck imputation: the sequential and random hot-deck algorithm; (2) the distance-based, k-nearest neighbor imputation approach; (3) individual, regression-based imputation; and (4) the iterative, model-based, stepwise regression imputation technique with both standard and robust methods (the IRMI algorithm). Furthermore, the course trains one to recognize the patterns of missingness using many vibrant and varied visualizations of the missing data patterns created by the professional VIMGUI software included in the course materials and made available to all course participants.This course is useful to anyone who regularly analyzes large or small data sets that may contain missing data. This includes graduate students and faculty engaged in empirical research and working professionals who are engaged in quantitative research and/or data analysis. The visualizations that are taught are especially useful to understand the types of data missingness that may be present in your data and consequently, how best to deal with this missing data using imputation. The course includes the means to apply the appropriate imputation techniques, especially for non-multivariate-normal sets of data which tend to be most problematic to impute.The course author provides free-of-charge with the course materials his own unique VIMGUI toolbar developed in the RGtk2 visualization programming language in R. However, please note that both the R-provided VIMGUI package (developed in RGtk2), as well as the course author's provided VIMGUI toolbar application (also developed in RGtk2) may have some problems starting up properly on a Mac computer. So if you only have a Mac available to you, you may have some initial difficulties getting the applications to run properly.

课程标签

0人关注该课程

主题相关的课程