Reinforcement Learning Basics Cheat Sheet
A quick reference to core RL concepts, algorithms like Q-learning and policy gradients, and how to implement a simple agent with OpenAI Gym/Gymnasium.
Gymnasium Environment Loop
Interact with a standard RL environment.
import gymnasium as gymenv = gym.make("CartPole-v1")obs, info = env.reset(seed=42)for _ in range(1000): action = env.action_space.sample() # random policy obs, reward, terminated, truncated, info = env.step(action) if terminated or truncated: obs, info = env.reset()env.close()
Tabular Q-Learning
Update rule for learning an action-value table.
import numpy as npQ = np.zeros((n_states, n_actions))alpha = 0.1 # learning rategamma = 0.99 # discount factorepsilon = 0.1 # exploration ratefor episode in range(num_episodes): state = env.reset() done = False while not done: # epsilon-greedy action selection if np.random.rand() < epsilon: action = env.action_space.sample() else: action = np.argmax(Q[state]) next_state, reward, done, _ = env.step(action) # Bellman update best_next = np.max(Q[next_state]) Q[state, action] += alpha * (reward + gamma * best_next - Q[state, action]) state = next_state
Core Concepts
Key vocabulary used across all RL algorithms.
- Agent- the learner/decision-maker that interacts with the environment
- Environment- everything the agent interacts with; returns next state and reward
- Policy (π)- the agent's strategy mapping states to actions
- Reward (r)- scalar feedback signal from the environment after each action
- Value function V(s)- expected cumulative future reward from state s under a policy
- Q-function Q(s,a)- expected cumulative reward from taking action a in state s, then following the policy
- Discount factor (γ)- weights future rewards; 0 = myopic, close to 1 = far-sighted
- Episode- one full sequence from initial state to terminal state
Common Algorithms
Widely used RL algorithm families.
- Q-Learning- off-policy, tabular method that learns the optimal action-value function
- SARSA- on-policy variant that updates using the action actually taken next
- DQN- Q-learning with a neural network function approximator and experience replay
- REINFORCE- Monte Carlo policy gradient method that directly optimizes the policy
- Actor-Critic- combines a policy (actor) with a value function (critic) to reduce variance
- PPO- Proximal Policy Optimization; clips policy updates for stable on-policy training
REINFORCE Policy Gradient
The Monte Carlo policy-gradient update: nudge log-probabilities of actions in proportion to the return they earned.
import torchimport torch.nn.functional as Fdef reinforce_update(policy_net, optimizer, log_probs, rewards, gamma=0.99): # Compute discounted returns G_t for each time step, back to front returns = [] G = 0.0 for r in reversed(rewards): G = r + gamma * G returns.insert(0, G) returns = torch.tensor(returns) returns = (returns - returns.mean()) / (returns.std() + 1e-8) # baseline via normalization loss = -torch.stack([lp * G for lp, G in zip(log_probs, returns)]).sum() optimizer.zero_grad() loss.backward() optimizer.step()
DQN With Experience Replay & Target Network
The two stabilization tricks that made deep Q-learning actually converge.
import randomfrom collections import dequeimport torchimport torch.nn.functional as Freplay_buffer = deque(maxlen=100_000)target_net.load_state_dict(policy_net.state_dict()) # sync target with policy netdef train_step(batch_size=64, gamma=0.99): batch = random.sample(replay_buffer, batch_size) states, actions, rewards, next_states, dones = zip(*batch) states, next_states = torch.stack(states), torch.stack(next_states) actions, rewards, dones = torch.tensor(actions), torch.tensor(rewards), torch.tensor(dones) q_values = policy_net(states).gather(1, actions.unsqueeze(1)).squeeze(1) with torch.no_grad(): next_q = target_net(next_states).max(1).values target = rewards + gamma * next_q * (1 - dones.float()) loss = F.smooth_l1_loss(q_values, target) # Huber loss, robust to outlier TD errors optimizer.zero_grad(); loss.backward(); optimizer.step()# Periodically: target_net.load_state_dict(policy_net.state_dict())
PPO Clipped Surrogate Objective
The core loss term behind Proximal Policy Optimization, the default choice for stable on-policy training.
import torchdef ppo_clip_loss(new_log_probs, old_log_probs, advantages, epsilon=0.2): ratio = torch.exp(new_log_probs - old_log_probs) # pi_new(a|s) / pi_old(a|s) unclipped = ratio * advantages clipped = torch.clamp(ratio, 1 - epsilon, 1 + epsilon) * advantages return -torch.min(unclipped, clipped).mean() # pessimistic (lower) bound# advantages are typically computed via Generalized Advantage Estimation (GAE):# A_t = sum_{l=0}^{T-t} (gamma * lambda)^l * delta_{t+l}, delta_t = r_t + gamma*V(s_{t+1}) - V(s_t)
Stable-Baselines3 Training Loop
Production-grade RL algorithms without hand-rolling the update rules.
from stable_baselines3 import PPOfrom stable_baselines3.common.vec_env import make_vec_envvec_env = make_vec_env("CartPole-v1", n_envs=8) # parallel envs for faster rolloutsmodel = PPO("MlpPolicy", vec_env, verbose=1, n_steps=2048, batch_size=64, gamma=0.99, gae_lambda=0.95, clip_range=0.2)model.learn(total_timesteps=200_000)model.save("ppo_cartpole")obs = vec_env.reset()action, _states = model.predict(obs, deterministic=True)
Advanced Concepts
Ideas that separate toy tabular RL from methods used in real agents.
- Advantage function A(s,a)- Q(s,a) - V(s); measures how much better an action is than the policy's average, reduces variance versus using raw returns
- Generalized Advantage Estimation (GAE)- exponentially-weighted blend of n-step advantage estimates controlled by lambda, trading bias for variance between TD(0) and Monte Carlo
- Reward shaping / sparse rewards- dense intermediate rewards speed up learning but risk reward hacking; sparse terminal-only rewards are 'correct' but slow, often paired with curiosity-driven exploration bonuses
- On-policy vs. off-policy- on-policy methods (REINFORCE, PPO, A2C) must discard data after each policy update; off-policy methods (Q-learning, DQN, SAC) can reuse old experience via a replay buffer, improving sample efficiency
- Exploration strategies beyond epsilon-greedy- entropy bonuses (maximize policy entropy alongside reward), Boltzmann/softmax exploration, and intrinsic curiosity (novelty-based bonus rewards) all address epsilon-greedy's inefficiency in large state spaces
- Credit assignment problem- determining which of many past actions caused a delayed reward; discounting, eligibility traces, and advantage estimation all exist to attack this
- Sim-to-real gap- policies trained in simulation often fail on real hardware due to unmodeled dynamics; domain randomization (varying simulator physics/visuals during training) is the standard mitigation
Always normalize rewards and use a replay buffer with DQN-style methods — raw sparse rewards and correlated sequential samples are the most common causes of unstable training.