|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/practice-exams-microsoft-azure-dp-203-data-engineering/
课程评论:没有评论
课程名称:Practice Exams Microsoft Azure DP-203 Data Engineering 课程概述:本课程旨在帮助学员准备Microsoft Azure DP-203 Data Engineering认证考试。请注意,课程中的问题并不是官方考试中的正式题目,但涵盖了考试知识点。这些问题大多基于虚构场景,并在其中提出相关问题。所有问题经过定期审查,以确保内容符合最新考试要求,并附有详细的解答和参考资料链接,确保解答的准确性。每次测试时问题顺序会被打乱,学员必须理解每个答案的原因,而不仅仅是记住上次的正确选项。 重要提示:本课程不应作为备考官方考试的唯一学习材料,主要用作主题学习材料的补充。如果发现需要改进的内容,请提供相关截图,并及时反馈给课程负责人。 作为考试的候选人,学员应具备以下内容的专业知识:从各种结构化、非结构化及流式数据系统中整合、转换和汇总数据,以构建分析解决方案。Azure数据工程师需要帮助利益相关者理解数据,通过不同工具和技术构建和维护安全合规的数据处理管道。学员将学习使用多种Azure数据服务和框架,设计适合商业需求的数据储存架构,包括管理数据仓库、大数据和湖屋架构。 课程内容还包括确保数据管道和数据存储的高效性和可靠性,以及设计、实施、监控和优化数据平台。候选人需具备强大的数据处理语言知识,包括SQL、Python和Scala,并了解并行处理与数据架构模式。此外,学员需熟练使用Azure Data Factory、Azure Synapse Analytics、Azure Stream Analytics等工具来创建数据处理解决方案。 课程技能概览: 1. 数据存储设计与实施(15-20%) 2. 数据处理开发(40-45%) 3. 数据存储与数据处理的安全性、监控与优化(30-35%) 课程将为学员提供在工作中所需的实际技能,确保学员能够有效处理和优化数据流程,以应对商业需求与约束。
In order to set realistic expectations, please note: These questions are NOT official questions that you will find on the official exam. These questions DO cover all the material outlined in the knowledge sections below. Many of the questions are based on fictitious scenarios which have questions posed within them.The official knowledge requirements for the exam are reviewed routinely to ensure that the content has the latest requirements incorporated in the practice questions. Updates to content are often made without prior notification and are subject to change at any time.Each question has a detailed explanation and links to reference materials to support the answers which ensures accuracy of the problem solutions.The questions will be shuffled each time you repeat the tests so you will need to know why an answer is correct, not just that the correct answer was item "B" last time you went through the test.NOTE: This course should not be your only study material to prepare for the official exam. These practice tests are meant to supplement topic study material.Should you encounter content which needs attention, please send a message with a screenshot of the content that needs attention and I will be reviewed promptly. Providing the test and question number do not identify questions as the questions rotate each time they are run. The question numbers are different for everyone.As a candidate for this exam, you should have subject matter expertise in integrating, transforming, and consolidating data from various structured, unstructured, and streaming data systems into a suitable schema for building analytics solutions.As an Azure data engineer, you help stakeholders understand the data through exploration, and build and maintain secure and compliant data processing pipelines by using different tools and techniques. You use various Azure data services and frameworks to store and produce cleansed and enhanced datasets for analysis. This data store can be designed with different architecture patterns based on business requirements, including:Management data warehouse (MDW)Big dataLakehouse architectureAs an Azure data engineer, you also help to ensure that the operationalization of data pipelines and data stores are high-performing, efficient, organized, and reliable, given a set of business requirements and constraints. You help to identify and troubleshoot operational and data quality issues. You also design, implement, monitor, and optimize data platforms to meet the data pipelines.As a candidate for this exam, you must have solid knowledge of data processing languages, including:SQLPythonScalaYou need to understand parallel processing and data architecture patterns. You should be proficient in using the following to create data processing solutions:Azure Data FactoryAzure Synapse AnalyticsAzure Stream AnalyticsAzure Event HubsAzure Data Lake StorageAzure DatabricksSkills at a glanceDesign and implement data storage (15-20%)Develop data processing (40-45%)Secure, monitor, and optimize data storage and data processing (30-35%)Design and implement data storage (15-20%)Implement a partition strategyImplement a partition strategy for filesImplement a partition strategy for analytical workloadsImplement a partition strategy for streaming workloadsImplement a partition strategy for Azure Synapse AnalyticsIdentify when partitioning is needed in Azure Data Lake Storage Gen2Design and implement the data exploration layerCreate and execute queries by using a compute solution that leverages SQL serverless and Spark clusterRecommend and implement Azure Synapse Analytics database templatesPush new or updated data lineage to Microsoft PurviewBrowse and search metadata in Microsoft Purview Data CatalogDevelop data processing (40-45%)Ingest and transform dataDesign and implement incremental loadsTransform data by using Apache SparkTransform data by using Transact-SQL (T-SQL) in Azure Synapse AnalyticsIngest and transform data by using Azure Synapse Pipelines or Azure Data FactoryTransform data by using Azure Stream AnalyticsCleanse dataHandle duplicate dataAvoiding duplicate data by using Azure Stream Analytics Exactly Once DeliveryHandle missing dataHandle late-arriving dataSplit dataShred JSONEncode and decode dataConfigure error handling for a transformationNormalize and denormalize dataPerform data exploratory analysisDevelop a batch processing solutionDevelop batch processing solutions by using Azure Data Lake Storage, Azure Databricks, Azure Synapse Analytics, and Azure Data FactoryUse PolyBase to load data to a SQL poolImplement Azure Synapse Link and query the replicated dataCreate data pipelinesScale resourcesConfigure the batch sizeCreate tests for data pipelinesIntegrate Jupyter or Python notebooks into a data pipelineUpsert dataRevert data to a previous stateConfigure exception handlingConfigure batch retentionRead from and write to a delta lakeDevelop a stream processing solutionCreate a stream processing solution by using Stream Analytics and Azure Event HubsProcess data by using Spark structured streamingCreate windowed aggregatesHandle schema driftProcess time series dataProcess data across partitionsProcess within one partitionConfigure checkpoints and watermarking during processingScale resourcesCreate tests for data pipelinesOptimize pipelines for analytical or transactional purposesHandle interruptionsConfigure exception handlingUpsert dataReplay archived stream dataManage batches and pipelinesTrigger batchesHandle failed batch loadsValidate batch loadsManage data pipelines in Azure Data Factory or Azure Synapse PipelinesSchedule data pipelines in Data Factory or Azure Synapse PipelinesImplement version control for pipeline artifactsManage Spark jobs in a pipelineSecure, monitor, and optimize data storage and data processing (30-35%)Implement data securityImplement data maskingEncrypt data at rest and in motionImplement row-level and column-level securityImplement Azure role-based access control (RBAC)Implement POSIX-like access control lists (ACLs) for Data Lake Storage Gen2Implement a data retention policyImplement secure endpoints (private and public)Implement resource tokens in Azure DatabricksLoad a DataFrame with sensitive informationWrite encrypted data to tables or Parquet filesManage sensitive informationMonitor data storage and data processingImplement logging used by Azure MonitorConfigure monitoring servicesMonitor stream processingMeasure performance of data movementMonitor and update statistics about data across a systemMonitor data pipeline performanceMeasure query performanceSchedule and monitor pipeline testsInterpret Azure Monitor metrics and logsImplement a pipeline alert strategyOptimize and troubleshoot data storage and data processingCompact small filesHandle skew in dataHandle data spillOptimize resource managementTune queries by using indexersTune queries by using cacheTroubleshoot a failed Spark jobTroubleshoot a failed pipeline run, including activities executed in external services