|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/dp-700-implementing-data-eng-solutions-using-ms-fabric/
课程评论:没有评论
**Microsoft Fabric 数据工程师 (DP-700) 课程总结** 本课程旨在培养具备使用 Microsoft Fabric 设计和实施可扩展数据解决方案能力的 Microsoft Fabric 数据工程师。课程涵盖了数据工程师的关键职责,包括: **1. 数据集成与注入:** * 连接和整合各类数据源(结构化、半结构化、非结构化)。 * 利用 Fabric 中的 Dataflows、Pipelines 和 Notebooks 进行数据注入和转换。 * 通过 Eventstreams、Data Factory 或 Apache Spark 处理实时和批量数据注入。 **2. 数据转换与处理:** * 实施数据清洗、规范化和数据丰富技术。 * 利用 Fabric 中的 Lakehouse 和 Data Warehouse 功能进行数据转换和存储。 * 使用 PySpark 或 SQL 创建和管理 Notebooks 进行数据转换。 **3. 数据建模与存储:** * 设计和实施星型模型和雪花型模型。 * 开发语义模型和数据集以供 Power BI 等工具使用。 * 优化 OneLake 中的存储,并管理 Delta Lake 表以实现高性能访问。 **4. 编排与自动化:** * 使用 Data Pipelines 和 Triggers 设计和调度数据工作流。 * 实施数据依赖链、重试、日志记录和监控。 **5. 安全与合规:** * 通过 Microsoft Purview 和 Fabric 工作区角色管理访问控制。 * 确保数据治理、血缘跟踪以及符合组织策略。 * 实施数据掩码、加密和审计程序。 **6. 性能优化:** * 优化 Fabric 中的查询性能、分区和缓存机制。 * 分析和调优 Spark 作业和 SQL 查询。 **7. 监控与故障排除:** * 利用 Fabric 的监控工具诊断和解决数据管道故障及瓶颈。 * 跟踪管道执行、数据流运行状况和作业性能。 **8. 协作与报告:** * 与数据分析师、BI 开发人员和业务利益相关者合作,理解数据需求。 * 在 Microsoft Fabric 的共享工作区和环境中协作。
Role: Microsoft Fabric Data Engineer (DP-700)Overview: A Microsoft Fabric Data Engineer is responsible for designing and implementing scalable data solutions using Microsoft Fabric. This includes transforming data from various sources into usable formats for analytics, managing data pipelines, optimizing data storage, and ensuring data security and compliance.Key Responsibilities:1. Data Integration & IngestionConnect and integrate various data sources including structured, semi-structured, and unstructured data.Use Dataflows, Pipelines, and Notebooks in Fabric to ingest and transform data.Handle real-time and batch data ingestion using Eventstreams, Data Factory, or Apache Spark.2. Data Transformation & ProcessingImplement data cleaning, normalization, and data enrichment techniques.Use Lakehouse and Data Warehouse features in Fabric for transforming and storing processed data.Create and manage Notebooks using PySpark or SQL for data transformation.3. Data Modeling & StorageDesign and implement star schemas and snowflake schemas.Develop semantic models and datasets for consumption by Power BI and other tools.Optimize storage in OneLake, and manage Delta Lake tables for high-performance access.4. Orchestration & AutomationDesign and schedule data workflows using Data Pipelines and Triggers.Implement data dependency chains, retries, logging, and monitoring.5. Security & ComplianceManage access controls using Microsoft Purview and Fabric workspace roles.Ensure data governance, lineage tracking, and compliance with organizational policies.Implement data masking, encryption, and auditing procedures.6. Performance OptimizationOptimize query performance, partitioning, and caching mechanisms in Fabric.Analyze and tune Spark jobs and SQL queries.7. Monitoring & TroubleshootingUse Fabric's monitoring tools to diagnose and resolve data pipeline failures and bottlenecks.Track pipeline execution, dataflow health, and job performance.8. Collaboration & ReportingWork with data analysts, BI developers, and business stakeholders to understand data requirements.Collaborate in shared workspaces and environments in Microsoft Fabric.