|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/azure-databricks-build-data-engineering-and-aiml-pipeline/
课程评论:没有评论
课程名称:Azure Databricks - 构建数据工程和AI/ML管道 课程概述:该课程旨在帮助您掌握在Databricks中执行ETL操作的技能,构建无监督的异常检测模型,学习MLOPS,在Databricks中进行CI/CD操作,并将机器学习模型部署到生产环境中。 大数据工程:大数据工程师在大规模计算环境中与大量数据处理系统和数据库进行互动,为组织提供分析,帮助评估其绩效,识别市场人群,预测即将发生的变化和市场趋势。 Azure Databricks:Azure Databricks是一个数据分析平台,针对Microsoft Azure云服务平台进行了优化,提供三个开发数据密集型应用的环境:Databricks SQL、Databricks数据科学与工程及Databricks机器学习。 异常检测:异常检测是一种数据挖掘中的步骤,用于识别偏离数据集中正常行为的数据点、事件和/或观察。异常数据可能指示关键事件(如技术故障)或潜在机会(如消费者行为变化),机器学习正逐步被用于自动化异常检测。 数据湖仓:数据湖仓是一种数据解决方案概念,结合了数据仓库与数据湖的要素。数据湖仓为数据湖实施数据仓库的数据结构和管理特性,通常在数据存储方面更具成本效益。 可解释的AI:可解释的AI是指其结果可以被人类理解的人工智能。这与机器学习中的“黑箱”概念形成对比,后者即使是设计者也无法解释AI为何做出某个特定决定。 Spark结构化流处理:结构化流处理是建立在Spark SQL引擎上的可扩展且容错的流处理引擎。简而言之,结构化流处理提供快速、可扩展、容错的端到端精准一次流处理,而用户无需考虑流处理的复杂性。 CI/CD操作:CI和CD分别代表持续集成和持续交付/持续部署。简单来说,CI是一种现代软件开发实践,其中频繁地和可靠地进行增量代码更改。 该课程不仅提供实践技能的培训,还为学员在大数据和机器学习领域的职业发展奠定了坚实的基础。
This course is designed to help you develop the skill necessary to perform ETL operations in Databricks, build unsupervised anomaly detection models, learn MLOPS, perform CI/CD operations in databricks and Deploy machine learning models into production.Big Data engineering:Big data engineers interact with massive data processing systems and databases in large-scale computing environments. Big data engineers provide organizations with analyses that help them assess their performance, identify market demographics, and predict upcoming changes and market trends.Azure Databricks:Azure Databricks is a data analytics platform optimized for the Microsoft Azure cloud services platform. Azure Databricks offers three environments for developing data intensive applications: Databricks SQL, Databricks Data Science & Engineering, and Databricks Machine Learning.Anomlay detection:Anomaly detection (aka outlier analysis) is a step in data mining that identifies data points, events, and/or observations that deviate from a dataset's normal behavior. Anomalous data can indicate critical incidents, such as a technical glitch, or potential opportunities, for instance a change in consumer behavior. Machine learning is progressively being used to automate anomaly detection.Data Lake House:A data lakehouse is a data solution concept that combines elements of the data warehouse with those of the data lake. Data lakehouses implement data warehouses' data structures and management features for data lakes, which are typically more cost-effective for data storage.Explainable AI:Explainable AI is artificial intelligence in which the results of the solution can be understood by humans. It contrasts with the concept of the "black box" in machine learning where even its designers cannot explain why an AI arrived at a specific decision.Spark structured streaming:Structured Streaming is a scalable and fault-tolerant stream processing engine built on the Spark SQL engine..In short, Structured Streaming provides fast, scalable, fault-tolerant, end-to-end exactly-once stream processing without the user having to reason about streaming.CI/CD Operation:CI and CD stand for continuous integration and continuous delivery/continuous deployment. In very simple terms, CI is a modern software development practice in which incremental code changes are made frequently and reliably.