|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/logging-monitoring-observability-google-cloud
课程评论:没有评论
课程名称:Google Cloud中的日志记录、监控和可观察性 概述:本课程旨在教会您如何监控、排除故障,并改善基础设施和应用程序的性能。在网站可靠性工程(SRE)原则的指导下,本课程结合了讲座、演示、动手实验和真实案例研究。您将在课程中获得全面监控、实时日志管理与分析、生产中调试代码以及分析CPU和内存使用情况的经验。 课程大纲: 1. **引言**:欢迎来到Google Cloud中的日志和监控课程!我们将介绍先决条件、受众和课程目标。 2. **Google Cloud可观察性简介**:在此模块中,我们将对Google Cloud的日志、监控和可观察性产品进行高层次的概述。 3. **关键系统监控**:本模块主要关注如何跟踪Google Cloud内资源的实际情况。我们将讨论与监控项目架构相关的选项和最佳实践,明确核心Cloud IAM角色以决定谁可以进行哪些监控操作。此外,我们还将探讨默认仪表板的使用,并学习如何创建自定义仪表板以展示资源消耗和应用负载,以及定义正常运行检查以监控系统的可用性和延迟。 4. **警报策略**:本模块将教授如何制定警报策略,定义警报政策,添加通知渠道,识别警报类型及其常见用途,构建资源组的警报,并以编程方式管理警报策略,以便及时发现云应用中的问题并迅速解决。 5. **高级日志记录和分析**:您将学习Google Cloud的高级日志和分析功能,包括资源标签的方法、定义日志接收器、基于日志条目创建监控指标、将应用错误链接到日志记录及其他操作工具,以及将日志导出到BigQuery进行长期存储和SQL分析。 6. **使用审计日志**:本模块将介绍如何使用Cloud审计日志,以回答“谁做了什么,何时做的?”并讨论审计日志的最佳实践。 7. **课程总结**:我们将总结课程中涵盖的主题。 通过本课程,您将为维护和优化云基础设施与应用程序的性能提供所需的详细知识和技能。
Name:Introduction
Description:Welcome to Logging and Monitoring in Google Cloud! We will cover the pre-requisites, audience and the course objectives.
Name:Introduction to Google Cloud Observability
Description:In this module, we will take some time to do a high-level overview of the various products which comprise Google Cloud’s logging, monitoring, and observability suite.
Name:Monitoring Critical Systems
Description:Monitoring is all about keeping track of exactly what's happening with the resources we've spun up inside of Google's Cloud. In this module, we'll take a look at options and best practices as they relate to monitoring project architectures. We'll differentiate the core Cloud IAM roles needed to decide who can do what as it relates to monitoring. Just like architecture, this is another crucial early step. We will examine some of the Google created default dashboards, and see how to use them appropriately. We will create charts and use them to build custom dashboards to show resource consumption and application load. And, finally, we will define uptime checks to track liveliness and latency.
Name:Alerting Policies
Description:Alerting gives timely awareness to problems in your cloud applications so you can resolve the problems quickly. In this module, you will learn how to develop alerting strategies, define alerting policies, add notification channels, identify types of alerts and common uses for each, construct and alert on resource groups, and manage alerting policies programmatically.
Name:Advanced Logging and Analysis
Description:In this module, we will examine some of Google Cloud's advanced logging and analysis capabilities. Specifically, in this module you will learn to identify and choose among resource tagging approaches, define log sinks, create monitoring metrics based on log entries, link application errors to Logging and other operation tools using Error Reporting, and export logs to BigQuery for long term storage and SQL based analysis.
Name:Working with Audit Logs
Description:In this module, we will examine how to use Cloud Audit Logs. You will learn how to use Cloud Audit Logs to answer the question, “Who, did what, and when?” We will also cover best practices for Audit Logging.
Name:Course Summary
Description:We will summarize the topics covered in this couse.
Learn how to monitor, troubleshoot, and improve your infrastructure and application performance. Guided by the principles of Site Reliability Engineering (SRE), this course features a combination of lectures, demos, hands-on labs, and real-world case studies. In this course, you'll gain experience with full-stack monitoring, real-time log management and analysis, debugging code in production, and profiling CPU and memory usage.