Decision Making and Reinforcement Learning

所在平台: Coursera

课程主页: https://www.coursera.org/learn/dmrol

课程评论:没有评论

第一个写评论        关注课程

课程简介

课程名称:决策制定与强化学习 课程概述:本课程是对序列决策制定和强化学习的介绍。我们将从效用理论开始,讨论如何将偏好表示和建模以支持决策制定。课程初期,我们将简单的决策问题建模为多臂老虎机问题,并讨论几种评估反馈的方法。接下来,我们会将决策问题建模为有限马尔可夫决策过程(MDPs),并通过动态规划算法讨论其解决方案。 课程大纲: 1. **决策与效用理论** 描述:欢迎来到决策制定与强化学习课程!在这一周,托尼·迪尔教授将对课程进行概述。您还将看到一些指导方针,以支持您在建模序列决策问题和实现强化学习算法方面的学习之旅。 2. **多臂老虎机问题** 描述:欢迎进入第二周!这一周,我们将学习多臂老虎机问题,这是一种优化问题,算法在探索与利用之间进行平衡,以最大化奖励。主题包括行动值和样本平均估计等内容。 本课程将帮助学员理解决策制定的基本理论,掌握强化学习的基本算法,并应用于实际的决策场景。

课程大纲

Name:Decision Making and Utility Theory

Description:Welcome to Decision Making and Reinforcement Learning! During this week, Professor Tony Dear provides an overview of the course. You will also view guidelines to support your learning journey towards modeling sequential decision problems and implementing reinforcement learning algorithms.

Name:Bandit Problems

Description:Welcome to week 2! This week, we will learn about multi-armed bandit problems, a type of optimization problem in which the algorithm balances exploration and exploitation to maximize rewards. Topics include action values and sample averaging estimation,

课程评论(0条)

课程详情

This course is an introduction to sequential decision making and reinforcement learning. We start with a discussion of utility theory to learn how preferences can be represented and modeled for decision making. We first model simple decision problems as multi-armed bandit problems in and discuss several approaches to evaluate feedback. We will then model decision problems as finite Markov decision processes (MDPs), and discuss their solutions via dynamic programming algorithms. We touch on the n

课程标签

0人关注该课程

主题相关的课程