|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/data-engineering-serverless-elt-bi-on-amazon-cloud/
课程评论:没有评论
课程名称:亚马逊云上的数据工程、无服务器ETL与商业智能 课程概述: AWS云平台的庞大生态系统可能让许多人感到畏惧与困惑,但本课程旨在为希望掌握在Redshift中搭建数据仓库或从头开始建立商业智能基础设施的人提供实践经验。数据科学家、分析师和商业分析师将越来越被期望成为多面手,处理数据的摄取、工程和仓储的技术方面。有基础云知识的人均可从本课程中获益,因为: - 本课程设计考虑了典型数据工程项目的整个生命周期 - 提供针对实际应用场景的实用解决方案 课程内容包括: - 从零开始在AWS Redshift中搭建数据仓库 - 基础数据仓储概念 - 编写无服务器的AWS Glue作业(pyspark和python shell)进行ETL和批处理 - 使用AWS Athena进行临时分析(使用Athena的时机) - 利用AWS数据管道同步增量数据 - 使用Lambda函数触发和自动化ETL/数据同步过程 - QuickSight的设置、分析与仪表盘 课程先决条件: - 必须具备Python和SQL基础 - 应懂得编写一些基本的PySpark脚本 - 有探索、学习的意愿并愿意付出额外努力以取得成功 - 拥有一个活跃的AWS账户 注意事项: 本课程利用Redshift和RDS的免费套餐,除非超过免费套餐使用限制,否则不会收取费用,这样的使用额度足以让您在课程中获得丰富的实践。此外,本课程使用浏览器中的AWS UI创建集群和设置作业,无需涉及bash脚本,可在任何操作系统上完成实验。课程的编码量并不多,只有约35%涉及编程,其余为执行、理解和连接不同组件。整个课程的目的是让每位学员了解并熟悉课程所用的所有工具和功能。 小贴士: - 尝试以1.2倍速观看视频 - 每次使用新组件或功能时,进行一些研究,了解其他同类工具的区别和特性,例如Redshift/Athena vs Snowflake或BigQuery,以及QuickSight vs PowerBi vs Microstrategy。
AWS Cloud can seem intimidating and overwhelming to a lot of people due to its vast ecosystem, but this course will make it easier for anyone who wants a hands-on expertise in setting up a data-warehouse in Redshift or setup a BI infrastructure from scratch.Data Scientists/Analysts/Business Analysts will soon be expected to (if not already) become all-rounders and handle the technical aspect of data ingestion/engineering/warehousing. Anyone who has the basic understanding of how cloud works can benefit from this course because: - This course is designed keeping in mind end to end life cycle of a typical data engineering project - Provides a practical solution to real-world use-cases This Course covers: Setting up a data warehouse in AWS Redshift from scratch Basic Data Warehousing Concepts Writing server-less AWS Glue Jobs (pyspark and python shell) for ETL and batch processing AWS Athena for ad-hoc analysis (when to use Athena) AWS Data Pipeline to sync incremental data Lambda functions to trigger and automate ETL/Data Syncing processes QuickSight Setup , Analyses and Dashboards Prerequisites for this course are: Python / Sql (Absolute must)PySpark (should know how to write some basic Pyspark scripts)Willingness to explore ,learn and put in the extra effort to succeed An active AWS Account Important Note - This course makes use of the free tiers for Redshift and RDS , so you will not be billed for them unless you exceed the free tier usage which should be more than enough to get enough practice from this course . Also , this course makes use of AWS UI on the browser for creating clusters and setting up jobs , there is no bash scripting involved. One can use any operating system to perform the lab sessions in this course. This course is not code-intense or code-heavy ,there is only 35% coding involved , the rest is execution,understanding and chaining different component together. The whole purpose of this course is to make everyone aware of and feel comfortable with all the tools/features used in this course. Some Tips: Try to watch the videos at 1.2X speed Every time you work on a new component or feature , do some research on the other tools that are meant for the same purpose and see how they differ and in what aspects , For Eg Redshift/Athena vs Snowflake or Bigquery , QuickSight vs PowerBi vs Microstrategy