Actor-Critic Methods
When to Use This Skill
Invoke this skill when you encounter:
- Algorithm Selection: "Should I use actor-critic for my continuous control problem?"
- SAC Implementation: User implementing SAC and needs guidance on entropy coefficient tuning
- TD3 Confusion: "Why does TD3 have twin critics and delayed updates?"
- Training Instability: "My actor-critic diverges. How do I stabilize it?"
- A2C/A3C Questions: "What's the difference between A2C and A3C?"
- Continuous Control: User has continuous action space and needs appropriate algorithm
- Critic Issues: "My critic loss isn't decreasing" or "Advantage estimates are wrong"
- SAC vs TD3: "Which algorithm should I use for my problem?"
- Entropy Tuning: "How do I set the entropy coefficient α in SAC?"
- Policy Gradient Variance: "My policy gradients are too noisy. How do I reduce variance?"
- Implementation Bugs: Critic divergence, actor-critic synchronization, target network staleness
- Continuous Action Handling: Tanh squashing, log determinant Jacobian, action scaling
This skill provides practical guidance for continuous action space RL using actor-critic methods.
Do NOT use this skill for:
- Discrete action spaces (route to value-based-methods for Q-learning/DQN)
- Pure policy gradient without value baseline (route to policy-gradient-methods)
- Model-based RL (route to model-based-rl)
- Offline RL (route to offline-rl-methods)
- Theory foundations (route to rl-foundations)
Core Principle
Actor-critic methods achieve the best of both worlds: a policy (actor) for action selection guided by a value function (critic) for stable learning. They dominate continuous control because they're designed for infinite action spaces and provide sample-efficient learning through variance reduction.
Key insight: Continuous control has infinite actions to explore. Value-based methods (compare all action values) are infeasible. Policy gradient methods (directly optimize policy) have high variance. Actor-critic solves this: policy directly outputs action distribution (actor), value function provides stable baseline (critic) to reduce variance.
Use them for:
- Continuous control (robot arms, locomotion, vehicle control)
- High-dimensional action spaces (continuous angles, forces, velocities)
- Sample-efficient learning from sparse experiences
- Problems requiring exploration via stochastic policies
- Continuous state/action MDPs (deterministic or stochastic environments)
Do not use for:
- Discrete small action spaces (too slow compared to DQN)
- Imitation learning focused on behavior cloning (use behavior cloning directly)
- Very high-dimensional continuous spaces without careful design (curse of dimensionality)
- Planning-focused problems (route to model-based methods)
Part 1: Actor-Critic Foundations
From Policy Gradient to Actor-Critic
You understand policy gradient from policy-gradient-methods. Actor-critic extends it with a value baseline to reduce variance.
Pure Policy Gradient (REINFORCE):
∇J = E_τ[∇log π(a|s) * G_t]
Problem: G_t (cumulative future reward) has high variance. All rollouts pulled toward average. Noisy gradients = slow learning.
Actor-Critic Solution:
∇J = E_τ[∇log π(a|s) * (G_t - V(s))]
= E_τ[∇log π(a|s) * A(s,a)]
where:
- Actor: π(a|s) = policy (action distribution)
- Critic: V(s) = value function (baseline)
- Advantage: A(s,a) = G_t - V(s) = "how much better than average"
Why baseline helps:
Without baseline: policy gradients = [+10, -2, +5, -3, -1] (noisy, high variance)
With baseline (subtract mean=2): [+8, -4, +3, -5, -3] (same direction, but cleaner relative to baseline)
Result: Gradient points in same direction (increase high G, decrease low G) but with MUCH lower variance.
This reduces sample complexity significantly.
Advantage Estimation
The core of actor-critic is accurate advantage estimation:
A(s,a) = Q(s,a) - V(s)
= E[r + γV(s')] - V(s)
= E[r + γV(s') - V(s)]
Key insight: Advantage = "by how much does taking action a in state s beat the average for this state?"
Three ways to estimate advantage:
1. Monte Carlo (full return):
G_t = r_t + γr_{t+1} + γ²r_{t+2} + ... (full rollout)
A(s,a) = G_t - V(s)
- Unbiased but high variance
- Requires complete episodes or long horizons
2. TD(0) (one-step bootstrap):
A(s,a) = r + γV(s') - V(s)
- Low variance but biased (depends on critic accuracy)
- One-step lookahead only
- If V(s') is wrong, advantage is wrong
3. GAE - Generalized Advantage Estimation (best practice):
A_t = δ_t + (γλ)δ_{t+1} + (γλ)²δ_{t+2} + ...
δ_t = r_t + γV(s_{t+1}) - V(s_t) [TD error]
λ ∈ [0,1] trades off bias-variance:
- λ=0: pure TD(0) (low variance, high bias)
- λ=1: pure MC (high variance, low bias)
- λ=0.95: sweet spot (good tradeoff)
Why GAE: Exponentially decaying trace over multiple steps. Reduces variance without full MC, reduces bias without pure TD.
Actor-Critic Pitfall #1: Critic Not Learning Properly
Scenario: User trains actor-critic but critic loss doesn't decrease. Actor improves, but value function plateaus. Agent can't use accurate advantage estimates.
Problem:
# WRONG - critic loss computed incorrectly
critic_loss = mean((V(s) - G_t)^2) # Wrong target!
critic_loss.backward()
The bug: Critic should learn Bellman equation:
V(s) = E[r + γV(s')]
If you compute target as G_t directly, you're using Monte Carlo returns (too noisy). If you use r + γV(s'), you're bootstrapping properly.
Correct approach:
# RIGHT - Bellman bootstrap target
V_target = r + gamma * V(s').detach() # Detach next state value!
critic_loss = mean((V(s) - V_target)^2)
Why detach() matters: If you don't detach V(s'), gradient flows backward through value function, creating a moving target problem.
Red Flag: If critic loss doesn't decrease while actor loss decreases, critic isn't learning Bellman equation. Check:
- Target computation (should be r + γV(s'), not G_t alone)
- Detach next state value
- Critic network is separate from actor
- Different learning rates (critic typically higher than actor)
Critic as Baseline vs Critic as Q-Function
Important distinction:
A2C uses critic as baseline:
V(s) = value of being in state s
A(s,a) = r + γV(s') - V(s) [TD advantage]
Policy loss = -∇log π(a|s) * A(s,a)
SAC/TD3 use critic as Q-function:
Q(s,a) = expected return from taking action a in s
A(s,a) = Q(s,a) - V(s)
Policy loss = ∇log π(a|s) * Q(s,a) [deterministic policy gradient]
Why the difference: A2C updates actor and critic together (on-policy). SAC/TD3 decouple them (off-policy):
- Actor never sees the replay buffer
- Critic learns Q from replay buffer
- Actor uses critic's Q estimate (always lagging slightly)
Part 2: A2C - Advantage Actor-Critic
A2C Architecture
A2C = on-policy advantage actor-critic. Actor and critic train simultaneously with synchronized rollouts:
┌─────────────────────────────────────────┐
│ Environment │
└────────────┬────────────────────────────┘
│ states, rewards
▼
┌─────────────────────────────────────────┐
│ Actor π(a|s) Critic V(s) │
│ Policy network Value network │
│ Outputs: action Outputs: value │
└────────┬──────────────────┬─────────────┘
│ │
└──────┬───────────┘
│
┌──────▼───────────┐
│ Advantage │
│ A(s,a) = r+γV(s')-V(s)
└────────┬─────────┘
│
┌────────▼────────────┐
│ Actor Loss: │
│ -log π(a|s) * A(s,a)
│ │
│ Critic Loss: │
│ (V(s) - target)² │
└─────────────────────┘
A2C Training Loop
for episode in range(num_episodes):
states, actions, rewards, values = [], [], [], []
state = env.reset()
for t in range(horizon):
# Actor samples action from policy
action = actor(state)
# Step environment
next_state, reward = env.step(action)
# Get value estimate (baseline)
value = critic(state)
# Store for advantage computation
states.append(state)
actions.append(action)
rewards.append(reward)
values.append(value)
state = next_state
# Advantage estimation (GAE)
advantages = compute_gae(rewards, values, next_value, gamma, lambda)
# Actor loss (policy gradient with baseline)
actor_loss = -log_prob(actions, actor(states)) * advantages
actor.update(actor_loss)
# Critic loss (value function learning)
critic_targets = rewards + gamma * values[1:] + gamma * critic(next_state)
critic_loss = (critic(states) - critic_targets)^2
critic.update(critic_loss)
A2C vs A3C
A2C: Synchronous - all parallel workers update at same time (cleaner, deterministic)
Worker 1 ────┐
Worker 2 ────┼──► Global Model Update ──► All workers receive updated weights
Worker 3 ────┤
Worker N ────┘
Wait for all workers before next update
A3C: Asynchronous - workers update whenever they finish (faster wall clock time, messier)
Worker 1 ──► Update (1) ──► Continue
Worker 2 ──────► Update (2) ──────► Continue
Worker 3 ──────────► Update (3) ──────► Continue
No synchronization barrier (race conditions possible)
In practice: A2C is preferred. A3C was important historically (enables multi-GPU training without synchronization) but A2C is cleaner.
Part 3: SAC - Soft Actor-Critic
SAC Overview
SAC = Soft Actor-Critic. The current SOTA (state-of-the-art) for continuous control. Three key innovations:
- Entropy regularization: Add H(π(·|s)) to objective (maximize entropy + reward)
- Auto-tuning entropy coefficient: Learn α automatically (no manual tuning!)
- Off-policy learning: Learn from replay buffer (sample efficient)
SAC's Objective Function
Standard policy gradient maximizes:
J(π) = E[G_t]
SAC maximizes:
J(π) = E[G_t + α H(π(·|s))]
= E[G_t] + α E[H(π(·|s))]
Where:
- G_t = cumulative reward
- H(π(·|s)) = policy entropy (randomness)
- α = entropy coefficient (how much we value exploration)
Why entropy: Exploratory policies (high entropy) discover better strategies. Adding entropy to objective = agent explores automatically.
SAC Components
┌─────────────────────────────────────┐
│ Replay Buffer (off-policy data) │
└────────────┬────────────────────────┘
│ sample batch
▼
┌────────────────────────┐
│ Actor Network │
│ π(a|s) = μ(s) + σ(s) │ (Gaussian policy)
│ Outputs: mean, std │
└────────────────────────┘
│
▼
┌────────────────────────┐
│ Two Critic Networks │
│ Q1(s,a), Q2(s,a) │
│ Learn Q-values │
└────────────────────────┘
│
▼
┌────────────────────────┐
│ Target Networks │
│ Q_target1, Q_target2 │
│ (updated every N) │
└────────────────────────┘
│
▼
┌────────────────────────┐
│ Entropy Coefficient │
│ α (learned!) │
└────────────────────────┘
SAC Training Algorithm
# Initialize
actor = ActorNetwork()
critic1, critic2 = CriticNetwork(), CriticNetwork()
target_critic1, target_critic2 = copy(critic1), copy(critic2)
entropy_alpha = 1.0 # Learned!
target_entropy = -action_dim # Target entropy (usually -action_dim)
for step in range(num_steps):
# 1. Collect data (could be online or from buffer)
state = env.reset() if step % 1000 == 0 else next_state
action = actor.sample(state) # π(a|s)
next_state, reward = env.step(action)
replay_buffer.add(state, action, reward, next_state, done)
# 2. Sample batch from replay buffer
batch = replay_buffer.sample(batch_size=256)
states, actions, rewards, next_states, dones = batch
# 3. Critic update (Q-function learning)
# Compute target Q value using entropy-regularized objective
next_actions = actor.sample(next_states)
next_log_probs = actor.log_prob(next_actions, next_states)
# Use BOTH target critics, take minimum (overestimation prevention)
Q_target1 = target_critic1(next_states, next_actions)
Q_target2 = target_critic2(next_states, next_actions)
Q_target = min(Q_target1, Q_target2)
# Entropy-regularized target
y = reward + γ(1 - done) * (Q_target - α * next_log_probs)
# Update both critics
critic1_loss = MSE(critic1(states, actions), y)
critic1.update(critic1_loss)
critic2_loss = MSE(critic2(states, actions), y)
critic2.update(critic2_loss)
# 4. Actor update (policy gradient with entropy)
# Reparameterization trick: sample actions, compute log probs
sampled_actions = actor.sample(states)
sampled_log_probs = actor.log_prob(sampled_actions, states)
# Actor maximizes Q - α*log_prob (entropy regularization)
Q1_sampled = critic1(states, sampled_actions)
Q2_sampled = critic2(states, sampled_actions)
Q_sampled = min(Q1_sampled, Q2_sampled)
actor_loss = -E[Q_sampled - α * sampled_log_probs]
actor.update(actor_loss)
# 5. Entropy coefficient auto-tuning (SAC's KEY INNOVATION)
# Learn α to maintain target entropy
entropy_loss = -α * (sampled_log_probs + target_entropy)
alpha.update(entropy_loss)
# 6. Soft update target networks (every N steps)
if step % update_frequency == 0:
target_critic1 = τ * critic1 + (1-τ) * target_critic1
target_critic2 = τ * critic2 + (1-τ) * target_critic2
SAC Pitfall #1: Manual Entropy Coefficient
Scenario: User implements SAC but manually sets α=0.2 and training diverges. Agent explores randomly and never improves.
Problem: SAC's entire design is that α is learned automatically. Setting it manually defeats the purpose.
# WRONG - treating α as fixed hyperparameter
alpha = 0.2 # Fixed!
loss = Q_target - 0.2 * log_prob # Same penalty regardless of entropy
# Result: If entropy naturally low, penalty still high → policy forced random
# If entropy naturally high, penalty too weak → insufficient exploration
Correct approach:
# RIGHT - α is learned via entropy constraint
target_entropy = -action_dim # For Gaussian: typically -action_dim
# Optimize α to maintain target entropy
entropy_loss = -α * (sampled_log_probs.detach() + target_entropy)
alpha_optimizer.zero_grad()
entropy_loss.backward()
alpha_optimizer.step()
# α adjusts automatically:
# - If entropy too high: α increases (more penalty) → policy becomes more deterministic
# - If entropy too low: α decreases (less penalty) → policy explores more
Red Flag: If SAC agent explores randomly without improving, check:
- Is α being optimized? (not fixed value)
- Is target entropy set correctly? (usually -action_dim)
- Is log_prob computed with squashed action (after tanh)?
SAC Pitfall #2: Tanh Squashing and Log Probability
Scenario: User implements SAC with Gaussian policy but uses policy directly. Log probabilities are computed wrong. Training is unstable.
Problem: SAC uses tanh squashing to bound actions:
Raw action from network: μ(s) + σ(s)*ε, ε~N(0,1) → unbounded
Tanh squashed: a = tanh(raw_action) → bounded in [-1,1]
But policy probability must account for this transformation:
π(a|s) ≠ N(μ(s), σ²(s)) [Wrong! Ignores tanh]
π(a|s) = |det(∂a/∂raw_action)|^(-1) * N(μ(s), σ²(s))
= (1 - a²)^2 * N(μ(s), σ²(s)) [Right! Jacobian correction]
log π(a|s) = log N(μ(s), σ²(s)) - 2*log(1 - a²)
The bug: Computing log_prob without Jacobian correction:
# WRONG
log_prob = normal.log_prob(raw_action) - log(1 + exp(-2*x))
# (standard normal log prob, ignores squashing)
# RIGHT
log_prob = normal.log_prob(raw_action) - log(1 + exp(-2*x))
log_prob = log_prob - 2 * (log(2) - x - softplus(-2*x)) # Add Jacobian term
Or simpler:
# PyTorch way
dist = Normal(mu, sigma)
raw_action = dist.rsample() # Reparameterized sample
action = torch.tanh(raw_action)
log_prob = dist.log_prob(raw_action) - torch.log(1 - action.pow(2) + 1e-6).sum(-1)
Red Flag: If SAC policy doesn't learn despite updates, check:
- Are actions being squashed (tanh)?
- Is log_prob computed with tanh Jacobian term?
- Is squashing adjustment in entropy coefficient update?
SAC Pitfall #3: Two Critics and Target Networks
Scenario: User implements SAC with one critic and gets unstable learning. "I thought SAC just needed entropy regularization?"
Problem: SAC uses TWO critics because of Q-function overestimation:
Single critic Q(s,a):
- Targets computed as: y = r + γQ_target(s', a')
- Q_target is function of Q (updated less frequently)
- In continuous space, selecting actions via max isn't feasible
- Next action sampled from π (deterministic max removed)
- But Q-values can still overestimate (stochastic environment noise)
Two critics (clipped double Q-learning):
- Use both Q1 and Q2, take minimum: Q_target = min(Q1_target, Q2_target)
- Prevents overestimation (conservative estimate)
- Both updated simultaneously
- Asymmetric: both learn, but target uses minimum
Correct implementation:
# WRONG - one critic
target = reward + gamma * critic_target(next_state, next_action)
# RIGHT - two critics with min
Q1_target = critic1_target(next_state, next_action)
Q2_target = critic2_target(next_state, next_action)
target = reward + gamma * min(Q1_target, Q2_target)
# Both critics learn
critic1_loss = MSE(critic1(state, action), target)
critic2_loss = MSE(critic2(state, action), target)
# But actor only uses critic1 (or min of both)
Q_current = min(critic1(state, sampled_action), critic2(state, sampled_action))
actor_loss = -(Q_current - alpha * log_prob)
Red Flag: If SAC diverges, check:
- Are there two Q-networks?
- Does target use min(Q1, Q2)?
- Are target networks updated (soft or hard)?
Part 4: TD3 - Twin Delayed DDPG
Why TD3 Exists
TD3 = Twin Delayed DDPG. It addresses SAC's cost (two networks, more computation) with deterministic policy gradient (simpler).
DDPG (older): Deterministic policy, single Q-network, no entropy. Fast but unstable.
TD3 (newer): Three tricks to stabilize DDPG:
- Twin critics: Two Q-networks (clipped double Q-learning)
- Delayed actor updates: Update actor every d steps (not every step)
- Target policy smoothing: Add noise to target action before Q evaluation
TD3 Architecture
┌──────────────────────────────────┐
│ Replay Buffer │
└────────────┬─────────────────────┘
│
▼
┌───────────────────────┐
│ Actor μ(s) │
│ Deterministic policy │
│ Outputs: action │
└───────────────────────┘
│
▼
┌─────────────────────────┐
│ Q1(s,a), Q2(s,a) │
│ Two Q-networks │
│ (Triple: original+2) │
└─────────────────────────┘
│
▼
┌─────────────────────────┐
│ Delayed Actor Update │
│ (every d steps) │
└─────────────────────────┘
TD3 Training Algorithm
for step in range(num_steps):
# 1. Collect data
action = actor(state) + exploration_noise
next_state, reward = env.step(action)
replay_buffer.add(state, action, reward, next_state, done)
if step < min_steps_before_training:
continue
batch = replay_buffer.sample(batch_size)
states, actions, rewards, next_states, dones = batch
# 2. Critic update (BOTH Q-networks)
# Trick #3: Target policy smoothing
next_actions = actor_target(next_states)
noise = torch.randn_like(next_actions) * target_noise
noise = torch.clamp(noise, -noise_clip, noise_clip)
next_actions = torch.clamp(next_actions + noise, -1, 1) # Add noise, clip
# Clipped double Q-learning: use minimum
Q1_target = critic1_target(next_states, next_actions)
Q2_target = critic2_target(next_states, next_actions)
Q_target = torch.min(Q1_target, Q2_target)
y = rewards + gamma * (1 - dones) * Q_target
# Update both critics
critic1_loss = MSE(critic1(states, actions), y)
critic1_optimizer.zero_grad()
critic1_loss.backward()
critic1_optimizer.step()
critic2_loss = MSE(critic2(states, actions), y)
critic2_optimizer.zero_grad()
critic2_loss.backward()
critic2_optimizer.step()
# 3. Delayed actor update (Trick #2)
if step % policy_delay == 0:
# Deterministic policy gradient
Q1_current = critic1(states, actor(states))
actor_loss = -Q1_current.mean()
actor_optimizer.zero_grad()
actor_loss.backward()
actor_optimizer.step()
# Soft update target networks
for param, target_param in zip(critic1.parameters(), critic1_target.parameters()):
target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
for param, target_param in zip(critic2.parameters(), critic2_target.parameters()):
target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
for param, target_param in zip(actor.parameters(), actor_target.parameters()):
target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
TD3 Pitfall #1: Missing Target Policy Smoothing
Scenario: User implements TD3 with twin critics and delayed updates but training still unstable. "I have two critics, why isn't it stable?"
Problem: Target policy smoothing is critical. Without it:
Next action = deterministic μ_target(s') [exact, no exploration noise]
If Q-networks overestimate for certain actions:
- Target policy always selects that exact action
- Q-target biased high for that action
- Feedback loop: overestimation → more value → policy selects it more → more overestimation
With smoothing:
Next action = μ_target(s') + ε_smoothing
- Adds small random noise to target action
- Prevents exploitation of Q-estimation errors
- Breaks feedback loop by adding randomness to target action
Important: Noise is added at TARGET action, not current action!
- Current: exploration_noise (for exploration during collection)
- Target: target_noise (for stability, noise clip small)
Correct implementation:
# Trick #3: Target policy smoothing
next_actions = actor_target(next_states)
noise = torch.randn_like(next_actions) * target_policy_noise
noise = torch.clamp(noise, -noise_clip, noise_clip)
next_actions = torch.clamp(next_actions + noise, -1, 1)
# Then use these noisy actions for Q-target
Q_target = min(Q1_target(next_states, next_actions),
Q2_target(next_states, next_actions))
Red Flag: If TD3 diverges despite two critics, check:
- Is noise added to target action (not just actor output)?
- Is noise clipped (noise_clip prevents too much noise)?
- Are critic targets using smoothed actions?
TD3 Pitfall #2: Delayed Actor Updates
Scenario: User implements TD3 with target policy smoothing and twin critics, but updates actor every step. "Do I really need delayed updates?"
Problem: Policy updates change actor, which changes actions chosen. If you update actor every step while critics are learning:
Step 1: Actor outputs a1, Q(s,a1) = 5, Actor updated
Step 2: Actor outputs a2, Q(s,a2) = 3, Actor wants to stay at a1
Step 3: Critics haven't converged, oscillate between a1 and a2
Result: Actor chases moving target, training unstable
With delayed updates:
Steps 1-4: Update critics only, let them converge
Step 5: Update actor (once per policy_delay=5)
Steps 6-9: Update critics only
Step 10: Update actor again
Result: Critic stabilizes before actor changes, smoother learning
Typical settings:
policy_delay = 2 # Update actor every 2 critic updates
# or
policy_delay = 5 # More conservative, every 5 critic updates
Correct implementation:
if step % policy_delay == 0: # Only sometimes!
actor_loss = -critic1(state, actor(state)).mean()
actor_optimizer.zero_grad()
actor_loss.backward()
actor_optimizer.step()
# Update targets on same schedule
soft_update(critic1_target, critic1)
soft_update(critic2_target, critic2)
soft_update(actor_target, actor)
Red Flag: If TD3 training unstable, check:
- Is actor updated only every policy_delay steps?
- Are target networks updated on same schedule (policy_delay)?
- Policy_delay typically 2-5
SAC vs TD3 Decision Framework
Both are SOTA for continuous control. How to choose?
| Aspect | SAC | TD3 |
|---|---|---|
| Policy Type | Stochastic (Gaussian) | Deterministic |
| Exploration | Entropy maximization (automatic) | Target policy smoothing |
| Sample Efficiency | High (two critics) | High (two critics) |
| Stability | Very stable (entropy helps) | Stable (three tricks) |
| Computation | Higher (entropy tuning) | Slightly lower |
| Manual Tuning | Minimal (α auto-tuned) | Moderate (policy_delay, noise) |
| When to Use | Default choice, off-policy | When deterministic better, simpler noise |
Decision tree:
Do you prefer stochastic or deterministic policy?
- Stochastic (multiple possible actions per state) → SAC
- Deterministic (one action per state) → TD3
Sample efficiency critical?
- Yes, limited data → Both good, slight edge SAC
- No, lots of data → Either works
How much tuning tolerance?
- Want minimal tuning → SAC (α auto-tuned)
- Don't mind tuning policy_delay, noise → TD3 (simpler conceptually)
Exploration challenges?
- Complex exploration (entropy helps) → SAC
- Simple exploration (policy smoothing enough) → TD3
Practical recommendation: Start with SAC. It's more robust (entropy auto-tuning). Switch to TD3 only if you:
- Know you want deterministic policy
- Have tuning expertise for policy_delay
- Need slightly faster computation
Part 5: Continuous Action Handling
Gaussian Policy Representation
Actor outputs mean and standard deviation:
raw_output = actor_network(state)
mu = raw_output[:, :action_dim]
log_std = raw_output[:, action_dim:]
log_std = torch.clamp(log_std, min=log_std_min, max=log_std_max)
std = log_std.exp()
dist = Normal(mu, std)
raw_action = dist.rsample() # Reparameterized sample
Why log(std)?: Parameterize log scale instead of scale directly.
- Numerical stability (log prevents underflow)
- Gradient flow smoother
- Prevents std from becoming negative
Why clamp log_std?: Prevents std from becoming too small or large.
- Too small: policy becomes deterministic, no exploration
- Too large: policy becomes random, no learning
Typical ranges:
log_std_min = -20 # std >= exp(-20) ≈ 2e-9 (small exploration)
log_std_max = 2 # std <= exp(2) ≈ 7.4 (max randomness)
Continuous Action Squashing (Tanh)
Raw network output unbounded. Use tanh to bound to [-1,1]:
# After sampling from policy
action = torch.tanh(raw_action)
# action now in [-1, 1]
# Scale to environment action range [low, high]
action_scaled = (high - low) / 2 * action + (high + low) / 2
Pitfall: Log probability must account for squashing (already covered in SAC section).
Exploration Noise in Continuous Control
Off-policy methods (SAC, TD3) need exploration during data collection:
Method 1: Action space noise (simpler):
action = actor(state) + noise
noise = torch.randn_like(action) * exploration_std
action = torch.clamp(action, -1, 1) # Ensure in bounds
Method 2: Parameter noise (more complex):
Add noise to actor network weights periodically
Action = actor_with_noisy_weights(state)
Results in correlated action noise across timesteps (more natural exploration)
Typical settings:
# For SAC: exploration_std = 0.1 * max_action
# For TD3: exploration_std starts high, decays over time
Part 6: Common Bugs and Debugging
Bug #1: Critic Divergence
Symptom: Critic loss explodes, V(s) becomes huge (1e6+), agent breaks.
Causes:
- Wrong target computation: Using wrong Bellman target
- No gradient clipping: Gradients unstable
- Learning rate too high: Critic overshoots
- Value targets too large: Reward scale not normalized
Diagnosis:
# Check target computation
print("Reward range:", rewards.min(), rewards.max())
print("V(s) range:", v_current.min(), v_current.max())
print("Target range:", v_target.min(), v_target.max())
# Plot value function over time
plt.plot(v_values_history) # Should slowly increase, not explode
# Check critic loss
print("Critic loss:", critic_loss.item()) # Should decrease, not diverge
Fix:
# 1. Reward normalization
rewards = (rewards - rewards.mean()) / (rewards.std() + 1e-8)
# 2. Gradient clipping
torch.nn.utils.clip_grad_norm_(critic.parameters(), max_norm=1.0)
# 3. Lower learning rate
critic_lr = 1e-4 # Instead of 1e-3
# 4. Value function target clipping (optional)
v_target = torch.clamp(v_target, -100, 100)
Bug #2: Actor Not Learning (Constant Policy)
Symptom: Actor loss decreases but policy doesn't change. Same action sampled repeatedly. No improvement in return.
Causes:
- Policy output not properly parameterized: Mean/std wrong
- Critic signal dead: Q-values all same, no gradient
- Learning rate too low: Actor updates too small
- Advantage always zero: Critic perfect (impossible) or wrong
Diagnosis:
# Check policy output distribution
actions = [actor.sample(state) for _ in range(1000)]
print("Action std:", np.std(actions)) # Should be >0.01
print("Action mean:", np.mean(actions))
# Check critic signal
q_values = critic(states, random_actions)
print("Q range:", q_values.min(), q_values.max())
print("Q std:", q_values.std()) # Should have variation
# Check advantage
advantages = q_values - v_baseline
print("Advantage std:", advantages.std()) # Should be >0
Fix:
# 1. Ensure policy outputs have variance
assert log_std.mean() < log_std_max - 0.5 # Not clamped to max
assert log_std.mean() > log_std_min + 0.5 # Not clamped to min
# 2. Check critic learns
critic_loss should decrease
# 3. Increase actor learning rate
actor_lr = 3e-4 # Instead of 1e-4
# 4. Debug advantage calculation
if advantage.std() < 0.01:
print("WARNING: Advantages have no variation, critic might be wrong")
Bug #3: Entropy Coefficient Divergence (SAC)
Symptom: SAC entropy coefficient α explodes (1e6+), policy becomes completely random, agent stops learning.
Cause: Entropy constraint optimization unstable.
# WRONG - entropy loss unbounded
entropy_loss = -alpha * (log_probs + target_entropy)
# If log_probs >> target_entropy, loss becomes huge positive, α explodes
Fix:
# RIGHT - use log(α) to avoid explosion
log_alpha = torch.log(alpha)
log_alpha_loss = -log_alpha * (log_probs.detach() + target_entropy)
alpha_optimizer.zero_grad()
log_alpha_loss.backward()
alpha_optimizer.step()
alpha = log_alpha.exp()
# Or clip α
alpha = torch.clamp(alpha, min=1e-4, max=10.0)
Bug #4: Target Network Never Updated
Symptom: Agent learns for a bit, then stops improving. Training plateaus.
Cause: Target networks not updated (or updated too rarely).
# WRONG - never update targets
target_critic = copy(critic) # Initialize once
for step in range(1000000):
# ... training loop ...
# But target_critic never updated!
Fix:
# RIGHT - soft update every step (or every N steps for delayed methods)
tau = 0.005 # Soft update parameter
for step in range(1000000):
# ... critic update ...
# Soft update targets
for param, target_param in zip(critic.parameters(), target_critic.parameters()):
target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
# Or hard update (copy all weights) every N steps
if step % update_frequency == 0:
target_critic = copy(critic)
Bug #5: Gradient Flow Through Detached Tensors
Symptom: Actor loss computation succeeds, but actor parameters don't update.
Cause: Critic detached but actor expects gradients.
# WRONG
for step in range(1000):
q_value = critic(state, action).detach() # Detached!
actor_loss = -q_value.mean()
actor.update(actor_loss) # Gradient won't flow through q_value!
# Result: actor_loss always 0 (constant from q_value.detach())
# Actor parameters updated but toward constant target (no signal)
Fix:
# RIGHT - don't detach when computing actor loss
q_value = critic(state, action) # No detach!
actor_loss = -q_value.mean()
actor.update(actor_loss) # Gradient flows through q_value
# Detach where appropriate:
# - Value targets: v_target = (r + gamma * v_next).detach()
# - Stop gradient in critic: q_target = (r + gamma * q_next.detach()).detach()
# But NOT when computing actor loss
Part 7: When to Use Actor-Critic vs Alternatives
Actor-Critic vs Policy Gradient (REINFORCE)
| Factor | Actor-Critic | Policy Gradient |
|---|---|---|
| Variance | Low (baseline reduces) | High (full return) |
| Sample Efficiency | High | Low |
| Convergence Speed | Fast | Slow |
| Complexity | Two networks | One network |
| Stability | Better | Worse (high noise) |
Use Actor-Critic when: Continuous actions, sample efficiency matters, training instability
Use Policy Gradient when: Simple problem, don't need value function, prefer simpler code
Actor-Critic vs Q-Learning (DQN)
| Factor | Actor-Critic | Q-Learning |
|---|---|---|
| Action Space | Continuous (natural) | Discrete (requires all Q values) |
| Sample Efficiency | High | Very high |
| Stability | Good | Can diverge (overestimation) |
| Complexity | Two networks | One network (but needs tricks) |
Use Actor-Critic for: Continuous actions, robotics, control
Use Q-Learning for: Discrete actions, games, navigation
Actor-Critic (On-Policy A2C) vs Off-Policy (SAC, TD3)
| Factor | A2C (On-Policy) | SAC/TD3 (Off-Policy) |
|---|---|---|
| Sample Efficiency | Moderate | High (replay buffer) |
| Stability | Good | Excellent |
| Complexity | Simpler | More complex |
| Data Reuse | Limited (one pass) | High (replay buffer) |
| Parallel Training | Excellent (A3C) | Limited (off-policy break) |
Use A2C when: Want simplicity, have parallel workers, on-policy is okay
Use SAC/TD3 when: Need sample efficiency, offline data possible, maximum stability
Part 8: Implementation Checklist
Pre-Training Checklist
- Actor outputs mean and log_std separately
- Log_std clamped:
log_std_min <= log_std <= log_std_max - Action squashing with tanh (bounded to [-1,1])
- Log probability computation includes tanh Jacobian (SAC/A2C)
- Critic network separate from actor
- Critic loss is value bootstrap (r + γV(s'), not G_t)
- Two critics for SAC/TD3 (or one for A2C)
- Target networks initialized as copies of main networks
- Replay buffer created (for off-policy methods)
- Advantage estimation (GAE preferred, MC acceptable)
Training Loop Checklist
- Data collection uses current actor (not target)
- Critic updated with Bellman target:
r + γV(s').detach() - Actor updated with advantage signal:
-log_prob(a) * A(s,a)or-Q(s,a) - Target networks soft updated:
τ * main + (1-τ) * target - For SAC: entropy coefficient α being optimized
- For TD3: delayed actor updates (every policy_delay)
- For TD3: target policy smoothing (noise + clip)
- Gradient clipping applied if losses explode
- Learning rates appropriate (critic_lr typically >= actor_lr)
- Reward normalization or clipping applied
Debugging Checklist
- Critic loss decreasing over time?
- V(s) and Q(s,a) values in reasonable range?
- Policy entropy decreasing (exploration → exploitation)?
- Actor loss decreasing?
- Return increasing over episodes?
- No NaN or Inf in losses?
- Advantage estimates have variation?
- Policy output std not stuck at min/max?
Part 9: Comprehensive Pitfall Reference
1. Critic Loss Not Decreasing
- Wrong Bellman target (should be r + γV(s'))
- Critic weights not updating (zero gradients)
- Learning rate too low
- Target network staleness (not updated)
2. Actor Not Improving
- Critic broken (no signal)
- Advantage estimates all zero
- Actor learning rate too low
- Policy parameterization wrong (no variance)
3. Training Unstable (Divergence)
- Missing target networks
- Critic loss exploding (wrong target, high learning rate)
- Entropy coefficient exploding (SAC: should be log(α))
- Actor updates every step (should delay, especially TD3)
4. Policy Stuck at Random Actions (SAC)
- Manual α fixed (should be auto-tuned)
- Target entropy wrong (should be -action_dim)
- Entropy loss gradient wrong direction
5. Policy Output Clamped to Min/Max Std
- Log_std range too tight (check log_std_min/max)
- Network initialization pushing to extreme values
- No gradient clipping preventing adjustment
6. Tanh Squashing Ignored
- Log probability not adjusted for squashing
- Missing Jacobian term in SAC/policy gradient
- Action scaling inconsistent
7. Target Networks Never Updated
- Forgot to create target networks
- Update function called but not applied
- Update frequency too high (no learning)
8. Off-Policy Break (Experience Replay)
- Actor training on old data (should use current replay buffer)
- Data distribution shift (actions from old policy)
- Batch importance weights missing (PER)
9. Advantage Estimates Biased
- GAE parameter λ wrong (should be 0.95-0.99)
- Bootstrap incorrect (wrong value target)
- Critic too inaccurate (overcorrection)
10. Entropy Coefficient Issues (SAC)
- Manual tuning instead of auto-tuning
- Entropy target not set correctly
- Log(α) optimization not used (causes explosion)
Part 10: Real-World Examples
Example 1: SAC for Robotic Arm Control
Problem: Robotic arm needs to reach target position. Continuous joint angles.
Setup:
state_dim = 18 # 6 joint angles + velocities
action_dim = 6 # Joint torques
action_range = [-1, 1] # Normalized
actor = ActorNetwork(state_dim, action_dim) # Outputs μ, log_std
critic1 = CriticNetwork(state_dim, action_dim)
critic2 = CriticNetwork(state_dim, action_dim)
target_entropy = -action_dim # -6
alpha = 1.0
Training:
for step in range(1000000):
# Collect experience
state = env.reset() if done else next_state
action = actor.sample(state)
next_state, reward, done = env.step(action)
replay_buffer.add(state, action, reward, next_state, done)
if len(replay_buffer) < min_buffer_size:
continue
batch = replay_buffer.sample(256)
# Critic update
next_actions = actor.sample(batch.next_states)
next_log_probs = actor.log_prob(next_actions, batch.next_states)
q1_target = target_critic1(batch.next_states, next_actions)
q2_target = target_critic2(batch.next_states, next_actions)
target = batch.rewards + gamma * (1-batch.dones) * (
torch.min(q1_target, q2_tar
…(truncated)