24h Pro data science in R

所在平台: Udemy

课程主页: https://www.udemy.com/course/24h-pro-data-science-in-r/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:24小时专业数据科学于R 课程概述:本课程探索了多种现代机器学习和数据科学技术,主要使用R语言。R是数据科学家最常用的工具之一。课程展示了多种统计和机器学习技术,具体内容包括:利用R的统计函数进行随机数生成、密度计算、直方图绘制等;使用CARET包处理监督学习问题;通过sqldf、caret等进行数据处理;应用PCA、DBSCAN、K均值等无监督技术;从R调用Keras(Python)中的深度学习模型;使用强大的XGBOOST方法进行回归和分类;制作有趣的图表,如地理热图和互动图表;使用caret训练多种机器学习方法的超参数;在R中进行线性回归,构建对数-对数模型,并进行方差分析;估计混合效应模型以明确建模观察值之间的协方差;训练健壮的模型,使用齿轮回归和分位数回归来识别异常值和新颖观察;估计ARIMA(时间序列)模型以预测时间变量。课程中的大多数实例来自互联网上收集的真实数据集,如Kaggle、美国人口普查局等。所有讲座均可下载,并附带相应的材料。教学方法是简要介绍每种技术,侧重于计算方面,尽量避免复杂的数学公式,专注于实际应用。本课程涵盖了成为数据科学家或参与Kaggle竞赛所需的大多数知识,假设学员已经对数据科学或统计学有一定的了解。

课程评论(0条)

课程详情

This course explores several modern machine learning and data science techniques in R. As you probably know, R is one of the most used tools among data scientists. We showcase a wide array of statistical and machine learning techniques. In particular: Using R's statistical functions for drawing random numbers, calculating densities, histograms, etc.Supervised ML problems using the CARET packageData processing using sqldf, caret, etc.Unsupervised techniques such as PCA, DBSCAN, K-meansCalling Deep Learning models in Keras(Python) from RUse the powerful XGBOOST method for both regression and classificationDoing interesting plots, such as geo-heatmaps and interactive plotsTrain ML train hyperparameters for several ML methods using caretDo linear regression in R, build log-log models, and do ANOVA analysisEstimate mixed effects models to explicitly model the covariances between observationsTrain outlier robust models using robust regression and quantile regressionIdentify outliers and novel observationsEstimate ARIMA (time series) models to predict temporal variables Most of the examples presented in this course come from real datasets collected from the web such as Kaggle, the US Census Bureau, etc. All the lectures can be downloaded and come with the corresponding material. The teaching approach is to briefly introduce each technique, and focus on the computational aspect. The mathematical formulas are avoided as much as possible, so as to concentrate on the practical implementations. This course covers most of what you would need to work as a data scientist, or compete in Kaggle competitions. It is assumed that you already have some exposure to data science / statistics.

课程标签

0人关注该课程

主题相关的课程