Data Engineering Fundamentals with Prefect Workflow

所在平台: Udemy

课程主页: https://www.udemy.com/course/data-engineering-fundamentals-with-prefect-workflow/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:使用Prefect工作流的数据工程基础 课程概述: 数据工程是设计和构建系统的过程,这些系统使人们能够从多个来源和格式收集和分析原始数据。这些系统使人们能够找到数据的实际应用,以帮助企业蓬勃发展。各类公司都有大量分散的数据需要筛选,以回答关键的商业问题。数据工程旨在支持这一过程,使数据消费者(如分析师、数据科学家和管理人员)能够可靠、快速和安全地审视所有可用数据。 十年前,数据分析主要集中在关系数据库或ERP系统中可用的结构化数据上,决策基于历史数据的分析,ETL(提取、转换和加载)工具用于数据仓库系统。然而,在这个动态变化的世界中,非关系型数据库信息需要用于快速分析。因此,除了数据库中的交易外,CSV、webhooks、HTTP和MQTT等网络信息源也需要适当处理。此外,ETL过程已经演变为数据管道。数据管道是一种方法,从各种数据源中获取原始数据,然后将其转移到数据存储(如数据湖或数据仓库)进行分析。在数据管道中,可以建立不同任务之间的依赖关系,这些任务还可以基于某些事件的发生(如订单预订或问题提出)触发。为此,使用了webhooks的概念。 Prefect是一个新兴的数据管道或工作流工具,可以构建不仅限于静态任务依赖关系的系统,这些任务依赖关系还可以基于某些事件的发生来建立。本课程使用云端版本的Prefect工作流工具,该工具可以从云基础虚拟机调用。需要具备Python和Shell脚本的知识。 课程内容涵盖以下主题: - 数据工程、数据分析与数据科学的区别 - 数据科学、机器学习与数据科学概述 - ETL与数据管道的比较 - 在Oracle云基础设施上配置Oracle Linux虚拟机 - Prefect Cloud数据管道与客户端虚拟机设置 - 文档参考 - Prefect工作流/数据管道 - Perfect Flow及任务依赖关系的实践演示 - 基于Python的Oracle数据库提取构建Prefect数据流管道 - Webhooks简介及与Prefect和Github的实践演示 - 数据工程师的职业发展路径 祝学习愉快!

课程评论(0条)

课程详情

Data engineering is the process of designing and building systems that let people collect and analyze raw data from multiple sources and formats. These systems empower people to find practical applications of the data, which businesses can use to thrive.Companies of all sizes have huge amounts of disparate data to comb through to answer critical business questions. Data engineering is designed to support the process, making it possible for consumers of data, such as analysts, data scientists and executives, to reliably, quickly and securely inspect all of the data available.About a decade back, the data analysis was merely on the structured data available on the a Relational data base or in ERP system and any decision was made based on analysis of the historic data and tools like ETL (extract, Tranform & load) was used for datawarehousing system. However in this dynamic ever changing world, non relational data base information need to used for quick analysis.So apart from transactions in database, the other source of web information from CSV, webhooks, http & MQTT need to taken care as appropriate.Further more, the process of ETL as evolved into Data pipelines. A data pipeline is a method in which raw data is ingested from various data sources and then ported to data store, like a data lake or data warehouse, for analysis. In data pipe line task dependency can be build with different task. These task can be also based on some events happening like Order booked or Issues raise which can trigger a task. For this concepts of Webhooks are used.Prefect is one such newly evolved data pipeline or workflow tool, in which one can build not only static task dependency, but these task dependency can be built based on some event happeningas well. This course uses the cloud version Prefect worflow tool which can be invoked from a cloud based virtual machine. Knowledge of Python & shell scripting is essential.This course covers following topic:•Difference between Data Engineering Vs Data Analysis Vs Data Science•An Overview about Data Science, Machine Learning & Data Science.•Extract, Transform, Load vs Data pipeline.•Provisioning Oracle Linux Virtual machine On Oracle Cloud Infrastructure.•Prefect Cloud Data pipeline and Client VM Set up.•Documentation reference - Prefect Workflow / Data pipelines.•Hands-on Demonstration of Perfect Flow with Tasks dependency.•Building Prefect dataflow pipeline for Oracle Database extract using Python.•Introduction to Webhooks and Hands-on Demonstration with Prefect & Github.•Career Path for Data EngineersHappy Learning!

课程标签

0人关注该课程

主题相关的课程