Data Engineering - SSIS/ETL/Pipelines/Python/Web Scraping

所在平台: Udemy

课程主页: https://www.udemy.com/course/data-engineering-ssisetlpipelinespythonweb-scraping/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:数据工程 - SSIS/ETL/管道/Python/网络爬虫 课程概述:数据工程师负责创建大数据ETL管道,能够将大量数据转化为有价值的洞察。他们专注于数据的生产就绪性,包括格式、弹性、扩展性和安全性等方面。SQL Server集成服务(SSIS)是微软SQL Server数据库软件的一部分,支持执行广泛的数据迁移任务。SSIS是一个用于数据集成和工作流应用的平台,提供了用于数据提取、转换和加载的数据仓库工具。ETL(提取、转换和加载)是一种数据集成过程,通过它可以将多个数据源中的数据整合为一个一致的数据存储,通常加载到数据仓库或其他目标系统中。ETL管道是将数据从一个或多个源移动到数据库(如数据仓库或目标数据库)的流程集合。 SQL Server集成服务(SSIS)为从不同源读取数据(提取)、进行聚合和转换(转换),然后将数据集成(加载)以用于数据仓库和分析提供了一种方便的统一方式。当需要处理大量数据(GB或TB级别)时,SSIS是处理此类工作负载的理想选择。网络爬虫、网络采集或网络数据提取是从网站提取数据的数据抓取技术。网络爬虫软件可以直接通过超文本传输协议或网页浏览器访问互联网。虽然网络爬虫可以由软件用户手动完成,但通常指的是使用机器人或网络爬虫实现的自动化过程。这是一种复制形式,特定数据从网络上收集并复制,通常存入中央本地数据库或电子表格,以便后续检索或分析。

课程评论(0条)

课程详情

A data engineer is someone who creates big data ETL pipelines, and makes it possible to take huge amounts of data and translate it into insights. They are focused on the production readiness of data and things like formats, resilience, scaling, and security. SQL Server Integration Services is a component of the Microsoft SQL Server database software that can be used to perform a broad range of data migration tasks. SSIS is a platform for data integration and workflow applications. It features a data warehousing tool used for data extraction, transformation, and loading.ETL, which stands for extract, transform and load, is a data integration process that combines data from multiple data sources into a single, consistent data store that is loaded into a data warehouse or other target system.An ETL pipeline is the set of processes used to move data from a source or multiple sources into a database such as a data warehouse or target databases.SQL Server Integration Service (SSIS) provides an convenient and unified way to read data from different sources (extract), perform aggregations and transformation (transform), and then integrate data (load) for data warehousing and analytics purpose. When you need to process large amount of data (GBs or TBs), SSIS becomes the ideal approach for such workload.Web scraping, web harvesting, or web data extraction is data scraping used for extracting data from websites. The web scraping software may directly access the World Wide Web using the Hypertext Transfer Protocol or a web browser. While web scraping can be done manually by a software user, the term typically refers to automated processes implemented using a bot or web crawler. It is a form of copying in which specific data is gathered and copied from the web, typically into a central local database or spreadsheet, for later retrieval or analysis.

课程标签

0人关注该课程

主题相关的课程