|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/google-cloud-professional-data-engineer-test-2021/
课程评论:没有评论
课程名称:谷歌云专业数据工程师认证考试 课程概述: SkillPractical 提供的谷歌云专业数据工程师认证考试课程旨在帮助数据科学家、解决方案架构师、DevOps 工程师以及希望在谷歌环境中进入机器学习和数据工程领域的人员。学生应具备谷歌云平台(GCP)基础知识,如存储、计算和安全,具备一定的编码技能(如 Python)以及良好的数据库理解。虽然不要求具备数据工程或机器学习背景,但对 GCP 的一定经验是必需的。本课程为高级认证,强烈建议学生在参加此课程之前,先完成谷歌认证的助理云工程师考试。 课程学习目标: - 设计数据处理系统 - 构建和维护数据结构与数据库 - 分析数据并启用机器学习 - 优化数据表示、数据基础设施性能与成本 - 确保数据处理基础设施的可靠性 - 可视化数据 - 设计安全的数据处理系统 课程大纲简介: 1. 设计数据处理系统 - 选择适当的存储技术 - 设计数据管道 - 设计数据处理解决方案 - 数据仓库与数据处理的迁移 2. 建立和运营数据处理系统 - 构建和运营存储系统 - 构建和运营数据管道 - 构建和运营处理基础设施 3. 运用机器学习模型 - 利用预构建的ML服务模型 - 部署ML管道 - 选择合适的训练和服务基础设施 - 衡量、监控和故障排查机器学习模型 4. 确保解决方案质量 - 设计安全性与合规性 - 确保可扩展性与效率 - 确保可靠性与准确性 - 确保灵活性与可移植性 本课程为希望提升在谷歌云领域的数据工程技能的学员提供了全面的学习内容和实践指导,旨在帮助学生为获得谷歌云专业数据工程师认证做好充分准备。
SkillPractical Google Cloud Professional Data Engineer Certification Test is for data scientists, solution architects, DevOps engineers, and anyone wanting to move into machine learning and data engineering in the context of Google. Students will need to have some familiarity with the basics of GCP, such as storage, compute, and security; some basic coding skills (like Python); and a good understanding of databases. You do not need to have a background in data engineering or machine learning, but some experience with GCP is essential.This is an advanced certification and we strongly recommend that students take the SkillPractical Google Certified Associate Cloud Engineer exam before.FYI, 87% of Google Cloud certified users feel more confident in their cloud skills.Course Learning ObjectivesDesign a data processing systemBuild and maintain data structures and databasesAnalyze data and enable machine learningOptimize data representations, data infrastructure performance, and costEnsure reliability of data processing infrastructureVisualize dataDesign secure data processing systemsCourse syllabus description:1. Designing data processing systems1.1 Selecting the appropriate storage technologies. Considerations include:Mapping storage systems to business requirementsData modelingTradeoffs involving latency, throughput, transactionsDistributed systemsSchema design1.2 Designing data pipelines. Considerations include:Data publishing and visualization (e.g., BigQuery)Batch and streaming data (e.g., Cloud Dataflow, Cloud Dataproc, Apache Beam, Apache Spark and Hadoop ecosystem, Cloud Pub/Sub, Apache Kafka)Online (interactive) vs. batch predictionsJob automation and orchestration (e.g., Cloud Composer)1.3 Designing a data processing solution. Considerations include:Choice of infrastructureSystem availability and fault toleranceUse of distributed systemsCapacity planningHybrid cloud and edge computingArchitecture options (e.g., message brokers, message queues, middleware, service-oriented architecture, serverless functions)At least once, in-order, and exactly once, etc., event processing1.4 Migrating data warehousing and data processing. Considerations include:Awareness of current state and how to migrate a design to a future stateMigrating from on-premises to cloud (Data Transfer Service, Transfer Appliance, Cloud Networking)Validating a migration2. Building and operationalizing data processing systems2.1 Building and operationalizing storage systems. Considerations include:Effective use of managed services (Cloud Bigtable, Cloud Spanner, Cloud SQL, BigQuery, Cloud Storage, Cloud Datastore, Cloud Memorystore)Storage costs and performanceLifecycle management of data2.2 Building and operationalizing pipelines. Considerations include:Data cleansingBatch and streamingTransformationData acquisition and importIntegrating with new data sources2.3 Building and operationalizing processing infrastructure. Considerations include:Provisioning resourcesMonitoring pipelinesAdjusting pipelinesTesting and quality control3. Operationalizing machine learning models3.1 Leveraging pre-built ML models as a service. Considerations include:ML APIs (e.g., Vision API, Speech API)Customizing ML APIs (e.g., AutoML Vision, Auto ML text)Conversational experiences (e.g., Dialogflow)3.2 Deploying an ML pipeline. Considerations include:Ingesting appropriate dataRetraining of machine learning models (Cloud Machine Learning Engine, BigQuery ML, Kubeflow, Spark ML)Continuous evaluation3.3 Choosing the appropriate training and serving infrastructure. Considerations include:Distributed vs. single machineUse of edge computeHardware accelerators (e.g., GPU, TPU)3.4 Measuring, monitoring, and troubleshooting machine learning models. Considerations include:Machine learning terminology (e.g., features, labels, models, regression, classification, recommendation, supervised and unsupervised learning, evaluation metrics)Impact of dependencies of machine learning modelsCommon sources of error (e.g., assumptions about data)4. Ensuring solution quality4.1 Designing for security and compliance. Considerations include:Identity and access management (e.g., Cloud IAM)Data security (encryption, key management)Ensuring privacy (e.g., Data Loss Prevention API)Legal compliance (e.g., Health Insurance Portability and Accountability Act (HIPAA), Children's Online Privacy Protection Act (COPPA), FedRAMP, General Data Protection Regulation (GDPR))4.2 Ensuring scalability and efficiency. Considerations include:Building and running test suitesPipeline monitoring (e.g., Stackdriver)Assessing, troubleshooting, and improving data representations and data processing infrastructureResizing and autoscaling resources4.3 Ensuring reliability and fidelity. Considerations include:Performing data preparation and quality control (e.g., Cloud Dataprep)Verification and monitoringPlanning, executing, and stress testing data recovery (fault tolerance, rerunning failed jobs, performing retrospective re-analysis)Choosing between ACID, idempotent, eventually consistent requirements4.4 Ensuring flexibility and portability. Considerations include:Mapping to current and future business requirementsDesigning for data and application portability (e.g., multi-cloud, data residency requirements)Data staging, cataloging, and discovery