メインコンテンツへスキップ

ホーム Python

コース

Pythonで学ぶGymnasiumによるReinforcement Learning

上級スキルレベル

更新日 2024/09

強化学習の旅を始めましょう！エージェントが相互作用を通じて環境を解決する方法を学びます。

コースを無料で開始

PythonArtificial Intelligence

4時間

15 ビデオ

52 演習

4,400 XP

12,915

修了証明書

何千もの企業の従業員が支持

チームのトレーニングを担当していますか？

Businessをお試しください

コース説明

前提条件

Supervised Learning with scikit-learn Python Toolbox Introduction to NumPy

1

Introduction to Reinforcement Learning

Dive into the exciting world of Reinforcement Learning (RL) by exploring its foundational concepts, roles, and applications. Navigate through the RL framework, uncovering the agent-environment interaction. You'll also learn how to use the Gymnasium library to create environments, visualize states, and perform actions, thus gaining a practical foundation in RL concepts and applications.

Fundamentals of reinforcement learning

What is Reinforcement Learning?

RL vs. other ML sub-domains

Scenarios for applying RL

Navigating the RL framework

RL interaction loop

Episodic and continuous RL tasks

Calculating discounted returns for agent strategies

Interacting with Gymnasium environments

Setting up a Mountain Car environment

Visualizing the Mountain Car Environment

Interacting with the Frozen Lake environment

チャプターを開始

2

Model-Based Learning

Delve deeper into the world of RL focusing on model-based learning. Unravel the complexities of Markov Decision Processes (MDPs), understanding their essential components. Enhance your skill set by learning about policies and value functions. Gain expertise in policy optimization with policy iteration and value Iteration techniques.

Markov Decision Processes

Custom Frozen Lake MDP components

Exploring state and action spaces

Transition probabilities and rewards

Policies and state-value functions

Defining a deterministic policy

Computing state-values for a policy

Comparing policies

Action-value functions

Computing Q-values

Improving a policy

Policy iteration and value iteration

Applying policy iteration for optimal policy

Implementing value iteration

チャプターを開始

3

Model-Free Learning

Embark on a journey through the dynamic realm of Model-Free Learning in RL. Get introduced to to the foundational Monte Carlo methods, and apply first-visit and every-visit Monte Carlo prediction algorithms. Transition into the world of Temporal Difference Learning, exploring the SARSA algorithm. Finally, dive into the depths of Q-Learning, and analyze its convergence in challenging environments.

Monte Carlo methods

Episode generation for Monte Carlo methods

Implementing first-visit Monte Carlo

Implementing every-visit Monte Carlo

Temporal difference learning

Implementing the SARSA update rule

Solving 8x8 Frozen Lake with SARSA

Implementing Q-learning update rule

Solving 8x8 Frozen Lake with Q-learning

Evaluating policy on a slippery Frozen Lake

チャプターを開始

4

Advanced Strategies in Model-Free RL

Dive into advanced strategies in Model-Free RL, focusing on enhancing decision-making algorithms. Learn about Expected SARSA for more accurate policy updates and Double Q-learning to mitigate overestimation bias. Explore the Exploration-Exploitation Tradeoff, mastering epsilon-greedy and epsilon-decay strategies for optimal action selection. Tackle the Multi-Armed Bandit Problem, applying strategies to solve decision-making challenges under uncertainty.

Expected SARSA

Expected SARSA update rule

Applying Expected SARSA

Double Q-learning

Implementing double Q-learning update rule

Applying double Q-learning

Balancing exploration and exploitation

Defining epsilon-greedy function

Solving CliffWalking with epsilon greedy strategy

Solving CliffWalking with decayed epsilon-greedy strategy

Multi-armed bandits

Creating a multi-armed bandit

Solving a multi-armed bandit

Assessing convergence in a multi-armed bandit

Congratulations!

チャプターを開始

Pythonで学ぶGymnasiumによるReinforcement Learning

コース完了

修了証明書を取得

この修了書をLinkedInや履歴書、CVに追加しましょう
ソーシャルメディアや人事評価で共有しましょう今すぐ登録

19百万人を超える学習者と共にPythonで学ぶGymnasiumによるReinforcement Learningを始めましょう！

DataCamp for Mobileでデータスキルを磨きましょう

モバイルコースと毎日の 5 分間のコーディングチャレンジで、外出先でも進歩できます。