AI Workflow: Business Priorities and Data Ingestion

所在平台: Coursera

课程主页: https://www.coursera.org/learn/ibm-ai-workflow-business-priorities-data-ingestion

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程概述:AI工作流程:商业优先事项与数据摄取 这是六部分专业化的第一门课程。强烈建议您按顺序完成这些课程,因为它们并不是独立的,而是形成一个工作流程,每门课程都建立在前面课程的基础上。 这门IBM AI企业工作流程认证专业化的第一门课程将带您了解专业化的范围和先决条件。具体来说,这些课程旨在为具备概率、统计、线性代数和数据科学及机器学习的Python工具的实践数据科学家设计。您将接触到一个假设的流媒体公司,并学习设计思维的概念,这是IBM组织大型企业AI项目的框架。此外,课程还介绍了科学思维的基础,因为将一名经验丰富的数据科学家与初学者区分开来的质量就是创造性的科学思维。最后,您将开始为假设的媒体公司工作,了解他们拥有的数据,并使用Python和Jupyter笔记本构建数据摄取管道。 课程结束时,您应该能够: 1. 理解使用结构化流程进行数据科学的优势。 2. 描述设计思维阶段与AI企业工作流程的对应关系。 3. 讨论多种优先考虑商业机会的策略。 4. 解释数据科学与数据工程在AI工作流程中重叠的地方。 5. 解释数据摄取中测试的目的。 6. 描述稀疏矩阵作为数据摄取目标位置的使用案例。 7. 了解实现数据摄取管道自动化的初步步骤。 适合人群: 本课程针对已经具备构建机器学习模型经验的数据科学从业者,旨在深化他们在大型企业中构建和部署AI的技能。如果您是一名有志成为数据科学家的初学者,建议您不要选修此课程,因为您需要现实世界的专业知识才能从课程内容中受益。 技能要求: 在开始本课程之前,建议您已具备以下主题的扎实理解:线性代数的基本理解;掌握抽样、概率论和概率分布;了解描述性和推论统计概念;对机器学习技术及最佳实践有一般了解;熟练掌握Python及数据科学常用的库(如NumPy、Pandas、matplotlib、scikit-learn);熟悉IBM Watson Studio;熟悉设计思维过程。 课程大纲: 第1部分:数据收集 在本模块中,您将学习或强化关于识别和阐述商业机会的知识。您将了解到将科学思维应用于理解商业用例的重要性。这个过程与调查员的工作有许多相似之处,并且您将对在这一阶段暂停、退一步并科学地思考主要过程的重要性产生深刻的敬畏。 第2部分:数据摄取 数据清洗、解析、组装及核对是数据科学家必须完成的最耗时的任务之一。数据清洗所花费的时间可能高达60%,并且会随着数据质量和项目需求的增加而增加。该模块关注数据摄取的过程,并呈现一个实际场景的案例研究。

课程大纲

Part: 1

Title:Data Collection

Description:Throughout this module you will learn or reinforce what you already know about identifying and articulating business opportunities. In this module you will learn the importance of applying a scientific thought process to the task of understanding the business use case. This process has many similarities to that of being an investigator. You will also generate a healthy respect for the need to pause, step back and think scientifically about the main processes in this stage.

Part: 2

Title:Data Ingestion

Description:Cleaning, parsing, assembling and gut-checking data is among the most time-consuming tasks that a data scientist has to perform. The time spent on data cleaning can start at 60% and increase depending on data quality and the project requirements. This module looks at the process of ingesting data and presents a case study working a real world scenario.

课程评论(0条)

课程详情

This is the first course of a six part specialization.  You are STRONGLY encouraged to complete these courses in order as they are not individual independent courses, but part of a workflow where each course builds on the previous ones. This first course in the IBM AI Enterprise Workflow Certification specialization introduces you to the scope of the specialization and prerequisites.  Specifically, the courses in this specialization are meant for practicing data scientists who are knowledgeable about probability, statistics, linear algebra, and Python tooling for data science and machine learning.  A hypothetical streaming media company will be introduced as your new client.  You will be introduced to the concept of design thinking, IBMs framework for organizing large enterprise AI projects.  You will also be introduced to the basics of scientific thinking, because the quality that distinguishes a seasoned data scientist from a beginner is creative, scientific thinking.  Finally you will start your work for the hypothetical media company by understanding the data they have, and by building a data ingestion pipeline using Python and Jupyter notebooks.   By the end of this course you should be able to: 1.  Know the advantages of carrying out data science using a structured process 2.  Describe how the stages of design thinking correspond to the AI enterprise workflow 3.  Discuss several strategies used to prioritize business opportunities 4.  Explain where data science and data engineering have the most overlap in the AI workflow 5.  Explain the purpose of testing in data ingestion  6.  Describe the use case for sparse matrices as a target destination for data ingestion  7.  Know the initial steps that can be taken towards automation of data ingestion pipelines   Who should take this course? This course targets existing data science practitioners that have expertise building machine learning models, who want to deepen their skills on building and deploying AI in large enterprises. If you are an aspiring Data Scientist, this course is NOT for you as you need real world expertise to benefit from the content of these courses.   What skills should you have? It is assumed you have a solid understanding of the following topics prior to starting this course: Fundamental understanding of Linear Algebra; Understand sampling, probability theory, and probability distributions; Knowledge of descriptive and inferential statistical concepts; General understanding of machine learning techniques and best practices; Practiced understanding of Python and the packages commonly used in data science: NumPy, Pandas, matplotlib, scikit-learn; Familiarity with IBM Watson Studio; Familiarity with the design thinking process.

课程标签

0人关注该课程

主题相关的课程