Apache Iceberg: The Beginner's Guide

所在平台: Udemy

课程主页: https://www.udemy.com/course/apache-iceberg-data-lakehouse-engineering/

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:Apache Iceberg:初学者指南 课程概述: 欢迎来到《数据湖屋工程与Apache Iceberg:从基础到最佳实践》课程,这是您掌握下一代开源分析表格式的完整指南。随着数据世界的发展,Apache Iceberg迅速成为现代数据架构的基石,适用于PB级数据集。Iceberg提供ACID事务、模式演化、时间旅行、分区修剪以及跨多个引擎的兼容性,所有功能都以开放、供应商无关的格式提供。 在这个实践课程中,您将超越基础,使用强大的工具构建真实世界的数据湖屋管道,包括: - PyIceberg:在Python中编程访问Iceberg表 - Polars:快速的数据框库,用于内存中的转换 - DuckDB:本地SQL强大工具,用于交互式开发 - Apache Spark:用于大规模批处理和流处理 - AWS S3:用于Iceberg表的云原生对象存储 - 以及更多工具:SQL、Parquet、Glue、Athena等现代开源工具 课程特色: - 实践与工具丰富:不仅限于Spark!学习如何使用Polars和DuckDB等现代引擎与Iceberg协作。 - 云端架构准备:学习如何在AWS S3上存储和管理Iceberg表,实现可扩展且成本效益高的部署。 - 概念与实践项目结合:理解表格式、目录管理、模式演变,并通过真实数据集应用这些知识。 - 开源导向:无供应商锁定。您将使用开放的社区驱动工具构建可互操作的管道。 您将学习: - Apache Iceberg及其在数据湖屋生态系统中的重要性 - 设计Iceberg表的模式演变、分区和元数据管理 - 使用Python(PyIceberg)、SQL和Spark查询和操作Iceberg表 - 与DuckDB和Polars的真实世界集成 - 使用S3对象存储实现云原生Iceberg表 - 执行时间旅行、增量读取和基于快照的回滚 - 通过文件压缩、统计和集群优化性能 - 构建可重现、可扩展和可维护的数据管道 适合人群: - 构建现代湖屋系统的数据工程师和架构师 - 处理大规模数据集和分析的Python开发者 - 使用AWS S3的数据湖的云专业人士 - 从Hive、Delta Lake或传统仓库转型的分析师或工程师 - 热爱数据工程、分析和开源创新的任何人 将使用的工具与技术: Apache Iceberg、PyIceberg、Spark、DuckDB、Polars、Pandas、SQL、AWS S3、Parquet,集成Metastore/目录(REST、Glue),以及使用Jupyter笔记本和CLI进行实践。 课程结束后,您将能够自信、高效地设计、部署和扩展基于Apache Iceberg的数据湖屋解决方案,并利用丰富的开源工具生态系统。

课程评论(0条)

课程详情

Welcome to Data Lakehouse Engineering with Apache Iceberg: From Basics to Best Practices - your complete guide to mastering the next generation of open table formats for analytics at scale.As the data world moves beyond traditional data lakes and expensive warehouses, Apache Iceberg is rapidly becoming the cornerstone of modern data architecture. Built for petabyte-scale datasets, Iceberg brings ACID transactions, schema evolution, time travel, partition pruning, and compatibility across multiple engines - all in an open, vendor-agnostic format.In this hands-on course, you'll go far beyond the basics. You'll build real-world data lakehouse pipelines using powerful tools like:PyIceberg - programmatic access to Iceberg tables in PythonPolars - lightning-fast DataFrame library for in-memory transformationsDuckDB - local SQL powerhouse for interactive developmentApache Spark - for large-scale batch and streaming processingAWS S3 - cloud-native object storage for Iceberg tablesAnd many more: SQL, Parquet, Glue, Athena, and modern open-source utilitiesWhat Makes This Course Special?Hands-on & Tool-rich: Not just Spark! Learn to use Iceberg with modern engines like Polars, DuckDB.Cloud-Ready Architecture: Learn how to store and manage your Iceberg tables on AWS S3, enabling scalable and cost-effective deployments.Concepts + Practical Projects: Understand table formats, catalog management, schema evolution, and then apply them using real datasets.Open-source Focused: No vendor lock-in. You'll build interoperable pipelines using open, community-driven tools.What You'll Learn:The why and how of Apache Iceberg and its role in the data lakehouse ecosystemDesigning Iceberg tables with schema evolution, partitioning, and metadata managementHow to query and manipulate Iceberg tables using Python (PyIceberg), SQL, and SparkReal-world integration with DuckDB, and PolarsUsing S3 object storage for cloud-native Iceberg tablesPerforming time travel, incremental reads, and snapshot-based rollbacksOptimizing performance with file compaction, statistics, and clusteringBuilding reproducible, scalable, and maintainable data pipelinesWho Is This Course For?Data Engineers and Architects building modern lakehouse systemsPython Developers working with large-scale datasets and analyticsCloud Professionals using AWS S3 for data lakesAnalysts or Engineers moving from Hive, Delta Lake, or traditional warehousesAnyone passionate about data engineering, analytics, and open-source innovationTools & Technologies You'll Use:Apache Iceberg, PyIceberg, Spark,DuckDB, Polars, Pandas, SQL, AWS S3, ParquetIntegration with Metastore/Catalogs (REST, Glue)Hands-on with Jupyter Notebooks, CLIBy the end of this course, you'll be able to design, deploy, and scale data lakehouse solutions using Apache Iceberg and a rich ecosystem of open-source tools - confidently and efficiently.

课程标签

0人关注该课程

主题相关的课程