|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/databricks-machine-learning/
课程评论:没有评论
课程名称:Databricks认证机器学习协理考试指南 概述:欢迎参加我们的全面课程,旨在帮助您掌握成为认证Databricks机器学习工程师协理所需的技能。Databricks是一个基于云的数据分析平台,提供统一的数据处理、机器学习和分析方法。随着对数据工程师需求的不断增长,Databricks已成为业内最受追捧的技能之一。 本课程的最低合格要求包括: 1. 熟练使用Databricks机器学习及其在机器学习工作流中的能力,包括: - Databricks机器学习(集群、代码库、作业) - Databricks机器学习运行时(基础知识、库) - AutoML(分类、回归、预测) - 特征商店(基础知识) - MLflow(跟踪、模型、模型注册) 2. 在机器学习工作流中实施正确决策,包括: - 探索性数据分析(总结统计、异常值移除) - 特征工程(缺失值填补、独热编码) - 调优(超参数基础知识、超参数并行化) - 评估与选择(交叉验证、评估指标) 3. 使用Spark ML和其他工具大规模实施机器学习解决方案,包括: - 分布式机器学习概念 - Spark ML建模API(数据分割、训练、评估、估计器vs.转换器、管道) - Hyperopt - Spark上的Pandas API - Pandas UDFs与Pandas函数API 4. 理解经典机器学习模型的高级扩展特性,包括: - 分布式线性回归 - 分布式决策树 - 集成方法(套袋、提升) 通过本课程,您将掌握成为Databricks机器学习工程师所需的核心知识和技能,助您在瞬息万变的数据科学领域蓬勃发展。
Welcome to our comprehensive course on Databricks Certified Machine Learning Engineer Associate certification. This course is designed to help you master the skills required to become a certified Databricks ML engineer associate.Databricks is a cloud-based data analytics platform that offers a unified approach to data processing, machine learning, and analytics. With the growing demand for data engineers, Databricks has become one of the most sought-after skills in the industry.The minimally qualified candidate should be able to:Use Databricks Machine Learning and its capabilities within machine learning workflows, including:Databricks Machine Learning (clusters, Repos, Jobs)Databricks Runtime for Machine Learning (basics, libraries)AutoML (classification, regression, forecasting)Feature Store (basics)MLflow (Tracking, Models, Model Registry)Implement correct decisions in machine learning workflows, including:Exploratory data analysis (summary statistics, outlier removal)Feature engineering (missing value imputation, one-hot-encoding)Tuning (hyperparameter basics, hyperparameter parallelization)Evaluation and selection (cross-validation, evaluation metrics)Implement machine learning solutions at scale using Spark ML and other tools, including:Distributed ML ConceptsSpark ML Modeling APIs (data splitting, training, evaluation, estimators vs. transformers, pipelines)HyperoptPandas API on SparkPandas UDFs and Pandas Function APIsUnderstand advanced scaling characteristics of classical machine learning models, including:Distributed Linear RegressionDistributed Decision TreesEnsembling Methods (bagging, boosting)