|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/ibm-data-ops-methodology
课程评论:没有评论
课程名称:数据操作(DataOps)方法论 概述:根据Gartner的定义,数据操作是一种协作数据管理实践,旨在改善数据管理者与组织内部数据消费者之间数据流的沟通、集成和自动化。与DevOps类似,数据操作并非一种严格的教条,而是一种基于原则的实践,影响数据提供和更新的方式,以满足组织数据消费者的需求。数据操作方法论旨在帮助组织利用可重复的流程来构建和部署分析及数据管道。通过遵循数据治理和模型管理实践,组织可以提供高质量的企业数据以支持人工智能的落地。成功实施这一方法论使组织能够了解、信任并利用数据来创造价值。 在数据操作方法论课程中,您将学习最佳实践,以定义可重复且面向业务的框架,确保可信数据的交付。该课程为数据工程专业化的一部分,帮助学习者获得成为数据工程师所需的基础技能。 课程大纲: 1. **建立数据操作 - 准备运营** - 学习数据操作方法的基础知识,了解参与定义数据和为多种数据消费群体进行数据策划的人员,以及他们如何合作为特定目的交付数据。 2. **建立数据操作 - 优化运营** - 学习数据操作团队如何共同努力定义工作业务的价值,以清晰表达他们对整个组织的贡献。 3. **迭代数据操作 - 了解您的数据** - 了解了解跨组织数据存储库的能力,运用数据发现技术识别和定位敏感或受监管的数据,并提高组织对数据的理解水平。 4. **迭代数据操作 - 信任您的数据** - 探讨数据语义理解对数据消费者的重要性,讨论数据源的可靠性、数据质量的常见维度以及确保数据质量所需的政策。 5. **迭代数据操作 - 使用您的数据** - 学习如何优化数据的添加及转换,以便于满足各种业务用例,规划并实施必要的数据移动与集成任务。 6. **改进数据操作** - 评估最近一次数据冲刺的效果,识别有效与无效之处,并提出改进建议。 7. **总结与期末考试** - 全面回顾课程内容,进行期末测验以巩固学习成果。 该课程通过深入的学习内容和实践案例,帮助学员掌握数据操作方法论,为应用于真实案例做好准备。
Name:Establish DataOps - Prepare for operation
Description:In this module you will learn the fundamentals of a DataOps approach. You will learn about the people who are involved in defining data, curating it for use by a wide variety of data consumers, and how they can work together to deliver data for a specific purpose:
Name:Establish DataOps – Optimize for operation
Description:In this lesson you will learn the fundamentals of a DataOps approach. You will learn about how the DataOps team works together in defining the business value of the work they undertake to be able to clearly articulate the value they bring to the wider organization:
Name:Iterate DataOps - Know your data
Description:In this lesson you will learn about the capabilities that you will need to use to understand the data in repositories across an organization. Data discovery is most appropriately employed when the scale of available data is too vast to devise a manual approach or where there has been institutional loss of data cataloging. It utilizes various techniques to programmatically recognize semantics and patterns in data. It is a key aspect of identifying and locating sensitive or regulated data to adequately protect it, although in general, knowing what stored data means unlocks its potential for use in analytics. Data Classification provides a higher level of semantic enrichment, enabling the organization to raise data understanding from technical metadata to a business understanding, further helping to discover the overlap between multiple sources of data according to the information that they contain:
Name:Iterate DataOps – Trust your data
Description:In this lesson you will learn that understanding data semantics helps data consumers to know what is available for consumption, but it does not provide any guidance on how good that data is. This module is all about trust, how reliable a data source can be in providing high fidelity data that can be used to drive key strategic decisions, and whether that data should be accessible to those who want to use it; whether the data consumer is permitted to see and use it. This module will address the common dimensions of data quality, how to both detect and remediate poor data quality. And it will look at enforcing the many policies that are needed around data quality, not least the need to respect an individual’s wishes and rights around how their data is used:
Name:Iterate DataOps – Use your data
Description:In this lesson you will learn that providing useful data in a catalog can often necessitate some transformation of that data. Modifying original data can optimize data ingestion in various use-cases, such as combining multiple data sets, consolidating multiple transaction summaries, or manipulating non-standard data to conform to international standards. This module will examine the choices for data preparation, how visualization can be used to facilitate the human understanding of the data and what needs to be changed, and the various options for single use, optimization of data workflows and ensuring the regular production of transformations for operational use. Furthermore, this module will show you how to plan and implement the data movement and integration tasks that are required to support a business use case. The module is based on a real-world data movement and integration project required to support implementation of an AI-based SaaS analytical system for supply chain management running in the Google cloud. The module will cover the major topics that need to be addressed to complete a data movement and integration project successfully:
Name:Improve DataOps
Description:In this lesson you will learn about evaluating the last data sprint, observe what worked and what did not, and make recommendations on how the next iteration could be improved.
Name:Summary & Final Exam
Description:
DataOps is defined by Gartner as "a collaborative data management practice focused on improving the communication, integration and automation of data flows between data managers and consumers across an organization. Much like DevOps, DataOps is not a rigid dogma, but a principles-based practice influencing how data can be provided and updated to meet the need of the organization’s data consumers.” The DataOps Methodology is designed to enable an organization to utilize a repeatable process to build and deploy analytics and data pipelines. By following data governance and model management practices they can deliver high-quality enterprise data to enable AI. Successful implementation of this methodology allows an organization to know, trust and use data to drive value. In the DataOps Methodology course you will learn about best practices for defining a repeatable and business-oriented framework to provide delivery of trusted data. This course is part of the Data Engineering Specialization which provides learners with the foundational skills required to be a Data Engineer.