Microsoft DP-203 Exam Practice Questions *Updated Nov 2023*

所在平台: Udemy

课程主页: https://www.udemy.com/course/microsoft-dp-203-exam-practice-questions-updated-nov-2023/

课程评论:没有评论

第一个写评论        关注课程

课程简介

《Microsoft DP-203考试练习题(2023年11月更新)》课程内容总结: 本课程旨在帮助学员准备Microsoft DP-203考试,涵盖了数据工程相关的核心概念和技能。课程内容主要分为三个大方向: **1. 设计和实现数据存储(15-20%)** * **分区策略:** 学习如何为文件、分析工作负载、流式工作负载以及Azure Synapse Analytics设计和实施有效的分区策略,并了解在Azure Data Lake Storage Gen2中何时需要分区。 * **数据探索层:** 掌握使用SQL无服务器和Spark集群创建和执行查询,推荐和实施Azure Synapse Analytics数据库模板,以及将数据沿袭推送到Microsoft Purview,并在Microsoft Purview数据目录中浏览搜索元数据。 **2. 开发数据处理(40-45%)** * **数据摄取与转换:** 学习如何设计和实现增量加载,使用Apache Spark和Transact-SQL(T-SQL)在Azure Synapse Analytics中进行数据转换,以及利用Azure Synapse Pipelines或Azure Data Factory进行数据摄取和转换。 * **数据清洗:** 掌握使用Azure Stream Analytics处理重复数据、缺失数据、迟到数据,并实现“Exactly Once Delivery”概念。此外,还包括数据拆分、JSON解析、数据编码解码,以及配置错误处理、数据规范化和反规范化。 * **数据探索性分析:** 学习进行数据探索性分析。 * **批处理解决方案:** 学习使用Azure Data Lake Storage、Azure Databricks、Azure Synapse Analytics和Azure Data Factory开发批处理解决方案。内容包括使用PolyBase加载数据到SQL池,实现Azure Synapse Link并查询复制数据,创建数据管道,扩展资源,配置批处理大小,创建管道测试,集成Jupyter或Python notebooks into a data pipeline,upsert数据,数据回滚,配置异常处理,配置批处理保留,读写delta lake。 * **流处理解决方案:** 学习使用Stream Analytics和Azure Event Hubs创建流处理解决方案,使用Spark结构化流处理数据,创建窗口聚合,处理模式漂移,处理时序数据,跨分区及分区内数据处理。此外,还包括处理数据过程中配置检查点和水印,扩展资源,创建测试,优化管道,处理中断,配置异常处理,upsert数据,重播存档流数据。 * **批处理与管道管理:** 学习管理批处理和管道,包括触发批处理,处理失败的批加载,验证批加载,管理Azure Data Factory或Azure Synapse Pipelines中的数据管道,以及在Data Factory或Azure Synapse Pipelines中调度数据管道。同时,学习实现管道工件的版本控制,以及在管道中管理Spark作业。 **3. 安全、监控和优化数据存储与数据处理(30-35%)** * **数据安全:** 学习实施数据屏蔽,加密静态和动态数据,实现行级和列级安全,实施Azure基于角色的访问控制(RBAC),以及为Data Lake Storage Gen2实施类似POSIX的访问控制列表(ACLs)。还包括实施数据保留策略,以及实施和使用安全端点(私有和公共)。 * **敏感信息管理:** 学习在Azure Databricks中使用资源令牌,加载包含敏感信息的DataFrame,以及将加密数据写入表或Parquet文件,并管理敏感信息。 * **监控:** 学习监控数据存储和数据处理,包括实施Azure Monitor日志记录,配置监控服务,监控流处理,衡量数据移动性能,监控和更新系统内数据统计信息,监控数据管道性能,衡量查询性能,调度和监控管道测试,以及解释Azure Monitor指标和日志。 * **管道警报策略:** 学习实施管道警报策略。 * **优化与故障排除:** 学习优化和排除数据存储与数据处理的故障。内容包括压缩小文件,处理数据倾斜和数据溢出,优化资源管理,通过索引器和缓存调优查询。此外,还包括故障排除失败的Spark作业和失败的管道运行,包括外部服务执行的活动。 本课程内容详实,覆盖了数据工程师在Azure平台上的关键技能,旨在帮助学员掌握设计、开发、管理和优化数据解决方案的能力。

课程评论(0条)

课程详情

The course includes the below concepts for the exam: *Updated Nov 2023*Design and implement data storage (15-20%)Implement a partition strategyImplement a partition strategy for filesImplement a partition strategy for analytical workloadsImplement a partition strategy for streaming workloadsImplement a partition strategy for Azure Synapse AnalyticsIdentify when partitioning is needed in Azure Data Lake Storage Gen2Design and implement the data exploration layerCreate and execute queries by using a compute solution that leverages SQL serverless and Spark clusterRecommend and implement Azure Synapse Analytics database templatesPush new or updated data lineage to Microsoft PurviewBrowse and search metadata in Microsoft Purview Data CatalogDevelop data processing (40-45%)Ingest and transform dataDesign and implement incremental loadsTransform data by using Apache SparkTransform data by using Transact-SQL (T-SQL) in Azure Synapse AnalyticsIngest and transform data by using Azure Synapse Pipelines or Azure Data FactoryTransform data by using Azure Stream AnalyticsCleanse dataHandle duplicate dataAvoiding duplicate data by using Azure Stream Analytics Exactly Once DeliveryHandle missing dataHandle late-arriving dataSplit dataShred JSONEncode and decode dataConfigure error handling for a transformationNormalize and denormalize dataPerform data exploratory analysisDevelop a batch processing solutionDevelop batch processing solutions by using Azure Data Lake Storage, Azure Databricks, Azure Synapse Analytics, and Azure Data FactoryUse PolyBase to load data to a SQL poolImplement Azure Synapse Link and query the replicated dataCreate data pipelinesScale resourcesConfigure the batch sizeCreate tests for data pipelinesIntegrate Jupyter or Python notebooks into a data pipelineUpsert dataRevert data to a previous stateConfigure exception handlingConfigure batch retentionRead from and write to a delta lakeDevelop a stream processing solutionCreate a stream processing solution by using Stream Analytics and Azure Event HubsProcess data by using Spark structured streamingCreate windowed aggregatesHandle schema driftProcess time series dataProcess data across partitionsProcess within one partitionConfigure checkpoints and watermarking during processingScale resourcesCreate tests for data pipelinesOptimize pipelines for analytical or transactional purposesHandle interruptionsConfigure exception handlingUpsert dataReplay archived stream dataManage batches and pipelinesTrigger batchesHandle failed batch loadsValidate batch loadsManage data pipelines in Azure Data Factory or Azure Synapse PipelinesSchedule data pipelines in Data Factory or Azure Synapse PipelinesImplement version control for pipeline artifactsManage Spark jobs in a pipelineSecure, monitor, and optimize data storage and data processing (30-35%)Implement data securityImplement data maskingEncrypt data at rest and in motionImplement row-level and column-level securityImplement Azure role-based access control (RBAC)Implement POSIX-like access control lists (ACLs) for Data Lake Storage Gen2Implement a data retention policyImplement secure endpoints (private and public)Implement resource tokens in Azure DatabricksLoad a DataFrame with sensitive informationWrite encrypted data to tables or Parquet filesManage sensitive informationMonitor data storage and data processingImplement logging used by Azure MonitorConfigure monitoring servicesMonitor stream processingMeasure performance of data movementMonitor and update statistics about data across a systemMonitor data pipeline performanceMeasure query performanceSchedule and monitor pipeline testsInterpret Azure Monitor metrics and logsImplement a pipeline alert strategyOptimize and troubleshoot data storage and data processingCompact small filesHandle skew in dataHandle data spillOptimize resource managementTune queries by using indexersTune queries by using cacheTroubleshoot a failed Spark jobTroubleshoot a failed pipeline run, including activities executed in external servicesI hope you find the course useful,If I can help in any way please message me,Happy learning!Thanks,Neil

课程标签

0人关注该课程

主题相关的课程