|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/dp-203-data-engineering-on-microsoft-azure-2025/
课程评论:没有评论
课程名称:DP-203:微软Azure上的数据工程2025 课程概述:启动您的数据工程职业生涯,掌握使用微软Azure数据服务设计和实施数据解决方案的技能。本专业证书旨在为希望在微软Azure数据服务领域展示其专业知识的数据工程师和开发人员提供支持,适合想要准备参加DP-203考试的任何人。本证书将帮助您发展在使用微软Azure数据服务设计和实施数据解决方案方面的专业知识。您将学习如何将来自各种结构化和非结构化数据系统的数据集成、转换和整合为适合构建分析解决方案的结构。该项目由10门课程组成,旨在帮助您为DP-203考试做准备。每门课程教授考试所测量的概念和技能。完成本专业证书后,您准备好参加DP-203考试,并可进行注册。 应用学习项目:学习者将在整个项目中参与互动练习,提供实践和实施所学内容的机会。他们可以使用Microsoft Learn Sandbox,这是一个免费环境,允许学习者探索微软Azure并亲自操作实时的微软Azure资源和服务。 技能测量: 1. 设计和实施数据存储(40-45%) 2. 设计和开发数据处理(25-30%) 3. 设计和实施数据安全(10-15%) 4. 监控和优化数据存储和数据处理(10-15%) 考试将测量您完成以下技术任务的能力:设计和实施数据存储;设计和开发数据处理;设计和实施数据安全;监控和优化数据存储和数据处理。 课程内容涵盖多个功能组: - 数据存储设计与实施 - 数据处理的设计与开发 - 数据安全设计与实施 - 数据存储和数据处理的监控与优化 通过这一课程,学习者将具备应对数据工程行业需求的各项技能,为未来的职业发展打下坚实的基础。
Launch Your Career in Data Engineering. Master designing and implementing data solutions that use Microsoft Azure data servicesThis Professional Certificate is intended for data engineers and developers who want to demonstrate their expertise in designing and implementing data solutions that use Microsoft Azure data services anyone interested in preparing for the Exam DP-203: Data Engineering on Microsoft Azure. This Professional Certificate will help you develop expertise in designing and implementing data solutions that use Microsoft Azure data services. You will learn how to integrate, transform, and consolidate data from various structured and unstructured data systems into structures that are suitable for building analytics solutions that use Microsoft Azure data services. This program consists of 10 courses to help prepare you to take Exam DP-203: Data Engineering on Microsoft Azure. Each course teaches you the concepts and skills that are measured by the exam. By the end of this Professional Certificate, you will be ready to take and sign-up for the Exam DP-203: Data Engineering on Microsoft Azure.Applied Learning ProjectLearners will engage in interactive exercises throughout this program that offers opportunities to practice and implement what they are learning. They use the Microsoft Learn Sandbox. This is a free environment that allows learners to explore Microsoft Azure and get hands-on with live Microsoft Azure resources and services.Skills measured on Microsoft Azure DP-203 ExamDesign and Implement Data Storage (40-45%)Design and implement data storage (40-45%)Design and develop data processing (25-30%)Design and implement data security (10-15%)Monitor and optimize data storage and data processing (10-15%)The exam measures your ability to accomplish the following technical tasks: design and implement data storage; design and develop data processing; design and implement data security; and monitor and optimize data storage and data processing.Functional groupsDesign and implement data storage (40-45%)Design a data storage structureDesign an Azure Data Lake solutionRecommend file types for storageRecommend file types for analytical queriesDesign for efficient queryingDesign for data pruningDesign a folder structure that represents the levels of data transformationDesign a distribution strategyDesign a data archiving solutionDesign a partition strategyDesign a partition strategy for filesDesign a partition strategy for analytical workloadsDesign a partition strategy for efficiency/performanceDesign a partition strategy for Azure Synapse AnalyticsIdentify when partitioning is needed in Azure Data Lake Storage Gen2Design the serving layerDesign star schemasDesign slowly changing dimensionsDesign a dimensional hierarchyDesign a solution for temporal dataDesign for incremental loadingDesign analytical storesDesign metastores in Azure Synapse Analytics and Azure DatabricksImplement physical data storage structuresImplement compressionImplement partitioning Implement shardingImplement different table geometries with Azure Synapse Analytics poolsImplement data redundancyImplement distributionsImplement data archivingImplement logical data structuresBuild a temporal data solutionBuild a slowly changing dimensionBuild a logical folder structureBuild external tablesImplement file and folder structures for efficient querying and data pruningImplement the serving layerDeliver data in a relational starDeliver data in Parquet filesMaintain metadataImplement a dimensional hierarchyDesign and develop data processing (25-30%)Ingest and transform dataTransform data by using Apache SparkTransform data by using Transact-SQLTransform data by using Data FactoryTransform data by using Azure Synapse PipelinesTransform data by using Stream AnalyticsCleanse dataSplit dataShred JSONEncode and decode dataConfigure error handling for the transformationNormalize and denormalize valuesTransform data by using ScalaPerform data exploratory analysisDesign and develop a batch processing solutionDevelop batch processing solutions by using Data Factory, Data Lake, Spark, Azure Synapse Pipelines, PolyBase, and Azure DatabricksCreate data pipelinesDesign and implement incremental data loadsDesign and develop slowly changing dimensionsHandle security and compliance requirementsScale resourcesConfigure the batch sizeDesign and create tests for data pipelinesIntegrate Jupyter/Python notebooks into a data pipelineHandle duplicate dataHandle missing dataHandle late-arriving dataUpsert dataRegress to a previous stateDesign and configure exception handlingConfigure batch retentionDesign a batch processing solutionDebug Spark jobs by using the Spark UIDesign and develop a stream processing solutionDevelop a stream processing solution by using Stream Analytics, Azure Databricks, and Azure Event HubsProcess data by using Spark structured streamingMonitor for performance and functional regressionsDesign and create windowed aggregatesHandle schema driftProcess time series dataProcess across partitionsProcess within one partitionConfigure checkpoints/watermarking during processingScale resourcesDesign and create tests for data pipelinesOptimize pipelines for analytical or transactional purposesHandle interruptionsDesign and configure exception handlingUpsert dataReplay archived stream dataDesign a stream processing solutionManage batches and pipelinesTrigger batchesHandle failed batch loadsValidate batch loadsManage data pipelines in Data Factory/Synapse PipelinesSchedule data pipelines in Data Factory/Synapse PipelinesImplement version control for pipeline artifactsManage Spark jobs in a pipelineDesign and implement data security (10-15%)Design security for data policies and standardsDesign data encryption for data at rest and in transitDesign a data auditing strategyDesign a data masking strategyDesign for data privacyDesign a data retention policyDesign to purge data based on business requirementsDesign Azure role-based access control (Azure RBAC) and POSIX-like Access Control List (ACL) for Data Lake Storage Gen2Design row-level and column-level securityImplement data securityImplement data maskingEncrypt data at rest and in motionImplement row-level and column-level securityImplement Azure RBACImplement POSIX-like ACLs for Data Lake Storage Gen2Implement a data retention policyImplement a data auditing strategyManage identities, keys, and secrets across different data platform technologiesImplement secure endpoints (private and public)Implement resource tokens in Azure DatabricksLoad a DataFrame with sensitive informationWrite encrypted data to tables or Parquet filesManage sensitive informationMonitor and optimize data storage and data processing (10-15%)Monitor data storage and data processingImplement logging used by Azure MonitorConfigure monitoring servicesMeasure performance of data movementMonitor and update statistics about data across a systemMonitor data pipeline performanceMeasure query performanceMonitor cluster performanceUnderstand custom logging optionsSchedule and monitor pipeline testsInterpret Azure Monitor metrics and logsInterpret a Spark directed acyclic graph (DAG)Optimize and troubleshoot data storage and data processingCompact small filesRewrite user-defined functions (UDFs)Handle skew in dataHandle data spillTune shuffle partitionsFind shuffling in a pipelineOptimize resource managementTune queries by using indexersTune queries by using cacheOptimize pipelines for analytical or transactional purposesOptimize pipeline for descriptive versus analytical workloadsTroubleshoot a failed spark jobTroubleshoot a failed pipeline run