|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/dbt-on-databricks/
课程评论:没有评论
课程名称:dbt在Databricks上的应用 课程概述: 准备好释放数据分析管道的全部潜力了吗?本课程“dbt在Databricks上”是为数据专业人士量身定制的一门综合课程,旨在帮助学员掌握如何在Databricks平台上使用dbt(数据构建工具)进行数据转换,充分利用Apache Spark进行可扩展高效的工作流程。尽管Databricks提供强大的数据处理能力,dbt则通过提供一个版本控制、模块化和可测试的SQL基础转换框架,增强了这一体验。此课程覆盖dbt Cloud和dbt Core,帮助学员掌握适用于任何环境的多样化技能。 课程内容: 1. dbt及关键概念介绍:深入了解dbt、Jinja模板和YAML配置工具,了解它们如何协同工作以简化数据转换。 2. 环境设置:逐步指导如何配置dbt Cloud与Databricks,介绍版本控制的基本要素及核心组件和数据管道。 3. 数据建模与转换:探索多层数据架构,包括铜层、银层和金层模型,学习如何使用dbt命令进行源引用、模式配置和高效数据管道构建。 4. 高级测试与验证:通过通用和单独测试实施强大的数据质量检查,掌握从tests:语法到data_tests:的转换,并集成dbt工具包如dbt_utils以增强功能。 5. Jinja、宏和自定义函数:掌握使用Jinja语法和宏编写可重用、可扩展代码的技巧,动态操作数据模型、变化模式并为特定用例开发自定义逻辑。 6. 材料化策略解读:理解各种材料化策略,包括表、视图、增量加载和快照,深入探讨像维度表的SCD类型2和事实表的增量更新等特定场景。 7. 部署与生产工作流程:在Databricks上建立生产就绪的dbt环境,管理作业,顺利部署模型,学习配置环境和目标变量以简化CI/CD工作流程。 8. 使用dbt Core进行开发:通过本地项目设置、GitHub集成和命令行导航体验dbt Core的灵活性,同时学习版本控制和协作的最佳实践。 9. 故障排除与高级技巧:掌握处理常见连接问题、优化项目性能和在Databricks上扩展工作负载的技巧。 目标受众: 本课程面向熟悉SQL的数据工程师、分析师和架构师,旨在提升其在Databricks平台上使用dbt进行数据转换的技能。建议具备基本的Python、Git和基于云的数据环境知识。 为何参加本课程? 通过实践项目、指导练习和可下载资源,本课程帮助学员建立可以应用于现实数据挑战的实用技能。课程结束时,学员将在构建、测试和部署强大数据管道方面掌握熟练技术,将他们塑造成能够应对复杂分析工作流程的数据专业人士。
Are you ready to unlock the full potential of your data analytics pipelines? dbt on Databricks is a comprehensive course tailored for data professionals aiming to master data transformation using dbt (data build tool) on the Databricks platform, harnessing the power of Apache Spark for scalable and efficient workflows. While Databricks offers robust data processing capabilities, dbt enhances the experience by providing a framework for version-controlled, modular, and testable SQL-based transformations. This combination leverages Apache Spark's power for scalable workflows while maintaining cleaner, more maintainable, and reusable code.The course covers both dbt Cloud and dbt Core, equipping learners with versatile skills for any environment.What This Course Covers:Introduction to dbt and Key Concepts: Begin with an in-depth overview of dbt, Jinja templating, and YAML for configuration. Understand how these tools come together to streamline data transformation.Setting Up the Environment: Follow step-by-step guidance on configuring dbt Cloud with Databricks, version control essentials, and an introduction to core components and data pipelines.Data Modeling and Transformations: Explore multi-layer data architecture, including Bronze, Silver, and Gold models. Learn practical approaches for source referencing, schema configuration, and building efficient data pipelines using dbt commands.Advanced Testing and Validation: Implement robust data quality checks through generic and singular tests, transitioning from tests: syntax to data_tests:, and integrate dbt packages like dbt_utils for enhanced functionality.Jinja, Macros, and Custom Functions: Master the art of reusable, scalable code with Jinja syntax and macros. Gain the skills to manipulate data models dynamically, change schemas, and develop custom logic for specific use cases.Materializations Explained: Understand various materialization strategies including tables, views, incremental loads, and snapshots. Delve into specific scenarios like SCD Type 2 for dimension tables and incremental updates for fact tables.Deployment and Production Workflows: Set up a production-ready dbt environment on Databricks, manage jobs, and deploy models seamlessly. Learn to configure environment and target variables for streamlined CI/CD workflows.Developing with dbt Core: Experience the flexibility of dbt Core through local project setups, GitHub integration, and command-line navigation, while learning best practices for version control and collaboration.Troubleshooting and Advanced Techniques: Gain insights into handling common connection issues, optimizing project performance, and scaling workloads on Databricks.Target Audience:This course is designed for data engineers, analysts, and architects who are already familiar with SQL and want to elevate their skills in data transformation using dbt on the Databricks platform. Basic knowledge of Python, Git, and cloud-based data environments is recommended.Why Take This Course?With hands-on projects, guided exercises, and downloadable resources, this course builds practical skills that can be applied to real-world data challenges. By the end of the course, proficiency in building, testing, and deploying robust data pipelines will set learners apart as skilled data professionals equipped to handle complex analytics workflows.