|
所在平台: Coursera |
课程主页: https://www.coursera.org/learn/site-reliability-engineering-slos
课程评论:没有评论
课程名称:网站可靠性工程:测量与管理可靠性 课程概述:服务级别指标(SLIs)和服务级别目标(SLOs)是测量和管理可靠性的基本工具。本课程将教授学生如何制定适当的SLIs和SLOs,并通过使用错误预算来管理可靠性。 课程大纲: 1. **引言到SRE** 本模块旨在让您了解SRE、CRE和SLOs的基本概念。如果您已经熟悉这些概念,您仍然可以在此模块中找到新信息和视角,但完成此模块并非必要。 2. **目标可靠性** 在该模块中,我们将讨论如何测量服务的期望可靠性。我们将围绕设置组织内应用程序SLOs时需要考虑的因素进行讨论,并介绍测量服务可靠性的三个原则。 3. **以可靠性为目标的运营** 本模块将介绍一种量化不可靠性的机制——错误预算。我们将展示如何利用错误预算决定何时集中精力提高服务的可靠性,并学习一些工程和运营改进的方法。 4. **选择优良的SLI** 本模块将探讨监控指标的特性,使其作为SLI时更为有效,同时对比一些不太有用的指标。我们将讨论五种主要的SLI测量方式及其优缺点。 5. **开发SLOs和SLIs** 本模块将概述开发用户旅程的SLOs和SLIs的四步流程,介绍一个虚构的公司以及其示例移动游戏的基础设施,并应用四步流程于简单的用户旅程。 6. **量化SLO的风险** 在此模块中,我们将对示例服务的可用性风险进行批判性分析,回答问题:“我们的SLO目标和错误预算是否现实?” 7. **SLO未达标准的后果** 本模块将讨论记录SLO的最佳实践、正式错误预算政策的依据以及如何创建最佳实践,同时分析一个示例错误预算政策,以理解在制定错误预算政策时的权衡和激励。 通过本课程,学生将获得管理和增强服务可靠性的实用技能和工具。
Name:Introduction to SRE
Description:This module is intended to bring you up to speed on the concepts underpinning SRE, CRE, and SLOs. If you're already familiar with these concepts, you may still find new information and perspectives in this module, but it is not necessary to complete it.
Name:Targeting Reliability
Description:In this module we’re going to talk about how you measure the desired reliability of a service. We will address what to consider when setting SLOs for your application within your organization. We'll look at the three principles we use to measure the desired reliability of a service: figuring out what you want to promise and to whom, figuring out the metrics you care about that make your service reliability “good", and finally, deciding how much reliability is good enough.
Name:Operating for Reliability
Description:In this module, we’ll start by introducing a mechanism for quantifying unreliability using something called an error budget. We'll show how error budgets help you decide when to focus on making a service more reliable. And then we'll learn about some of the engineering and operational improvements that can help you do that.
Name:Choosing a Good SLI
Description:In this module we will start off by taking a look at some characteristics of monitoring metrics that can make them useful as SLIs and contrast these against other metrics that are less useful. Because the choice of where to measure an SLI is a key variable, we'll cover the five main ways you can measure an SLI and compare their pros and cons.
Name:Developing SLOs and SLIs
Description:In this module, we'll start off with an overview of our four step process for developing SLOs and SLIs for a user journey. We'll introduce the fictional company that created our example mobile game, the infrastructure that we'll be working with, and the simple user journey we'll be applying the four step process to.
Name:Quantifying Risks to SLOs
Description:In this module we'll be taking a critical look at the availability risks for our example service. We want to answer the question: "are our SLO targets and error budgets realistic?"
Name:Consequences of SLO Misses
Description:In this module, we'll cover best practices for documenting your SLOs, the rationale behind a formal error budget policy and how best to create one and finally, we'll look at an example error budget policy in order to understand the trade-offs and incentives that play out during negotiations when trying to write an error budget policy.
Service level indicators (SLIs) and service level objectives (SLOs) are fundamental tools for measuring and managing reliability. In this course, students learn approaches for devising appropriate SLIs and SLOs and managing reliability through the use of an error budget.