# Actor Critic Methods

> Master A2C, A3C, SAC, TD3 - actor-critic methods for continuous control

- Skill: `diegosouzapw/actor-critic-methods` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add diegosouzapw/actor-critic-methods`
- Raw SKILL.md: https://api.skillmd.com/api/skills/diegosouzapw/actor-critic-methods/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: diegosouzapw (https://skillmd.com/u/diegosouzapw)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/diegosouzapw/actor-critic-methods

---


# Actor-Critic Methods

## When to Use This Skill

Invoke this skill when you encounter:

- **Algorithm Selection**: "Should I use actor-critic for my continuous control problem?"
- **SAC Implementation**: User implementing SAC and needs guidance on entropy coefficient tuning
- **TD3 Confusion**: "Why does TD3 have twin critics and delayed updates?"
- **Training Instability**: "My actor-critic diverges. How do I stabilize it?"
- **A2C/A3C Questions**: "What's the difference between A2C and A3C?"
- **Continuous Control**: User has continuous action space and needs appropriate algorithm
- **Critic Issues**: "My critic loss isn't decreasing" or "Advantage estimates are wrong"
- **SAC vs TD3**: "Which algorithm should I use for my problem?"
- **Entropy Tuning**: "How do I set the entropy coefficient α in SAC?"
- **Policy Gradient Variance**: "My policy gradients are too noisy. How do I reduce variance?"
- **Implementation Bugs**: Critic divergence, actor-critic synchronization, target network staleness
- **Continuous Action Handling**: Tanh squashing, log determinant Jacobian, action scaling

**This skill provides practical guidance for continuous action space RL using actor-critic methods.**

Do NOT use this skill for:

- Discrete action spaces (route to value-based-methods for Q-learning/DQN)
- Pure policy gradient without value baseline (route to policy-gradient-methods)
- Model-based RL (route to model-based-rl)
- Offline RL (route to offline-rl-methods)
- Theory foundations (route to rl-foundations)

---

## Core Principle

**Actor-critic methods achieve the best of both worlds: a policy (actor) for action selection guided by a value function (critic) for stable learning. They dominate continuous control because they're designed for infinite action spaces and provide sample-efficient learning through variance reduction.**

Key insight: Continuous control has infinite actions to explore. Value-based methods (compare all action values) are infeasible. Policy gradient methods (directly optimize policy) have high variance. **Actor-critic solves this: policy directly outputs action distribution (actor), value function provides stable baseline (critic) to reduce variance.**

Use them for:

- Continuous control (robot arms, locomotion, vehicle control)
- High-dimensional action spaces (continuous angles, forces, velocities)
- Sample-efficient learning from sparse experiences
- Problems requiring exploration via stochastic policies
- Continuous state/action MDPs (deterministic or stochastic environments)

**Do not use for**:

- Discrete small action spaces (too slow compared to DQN)
- Imitation learning focused on behavior cloning (use behavior cloning directly)
- Very high-dimensional continuous spaces without careful design (curse of dimensionality)
- Planning-focused problems (route to model-based methods)

---

## Part 1: Actor-Critic Foundations

### From Policy Gradient to Actor-Critic

You understand policy gradient from policy-gradient-methods. Actor-critic extends it with **a value baseline to reduce variance**.

**Pure Policy Gradient (REINFORCE)**:

```
∇J = E_τ[∇log π(a|s) * G_t]
```

**Problem**: G_t (cumulative future reward) has high variance. All rollouts pulled toward average. Noisy gradients = slow learning.

**Actor-Critic Solution**:

```
∇J = E_τ[∇log π(a|s) * (G_t - V(s))]
       = E_τ[∇log π(a|s) * A(s,a)]

where:
- Actor: π(a|s) = policy (action distribution)
- Critic: V(s) = value function (baseline)
- Advantage: A(s,a) = G_t - V(s) = "how much better than average"
```

**Why baseline helps**:

```
Without baseline: policy gradients = [+10, -2, +5, -3, -1] (noisy, high variance)
With baseline (subtract mean=2): [+8, -4, +3, -5, -3] (same direction, but cleaner relative to baseline)

Result: Gradient points in same direction (increase high G, decrease low G) but with MUCH lower variance.
This reduces sample complexity significantly.
```

### Advantage Estimation

The core of actor-critic is **accurate advantage estimation**:

```
A(s,a) = Q(s,a) - V(s)
       = E[r + γV(s')] - V(s)
       = E[r + γV(s') - V(s)]
```

**Key insight**: Advantage = "by how much does taking action a in state s beat the average for this state?"

**Three ways to estimate advantage**:

**1. Monte Carlo (full return)**:

```python
G_t = r_t + γr_{t+1} + γ²r_{t+2} + ... (full rollout)
A(s,a) = G_t - V(s)
```

- Unbiased but high variance
- Requires complete episodes or long horizons

**2. TD(0) (one-step bootstrap)**:

```python
A(s,a) = r + γV(s') - V(s)
```

- Low variance but biased (depends on critic accuracy)
- One-step lookahead only
- If V(s') is wrong, advantage is wrong

**3. GAE - Generalized Advantage Estimation** (best practice):

```python
A_t = δ_t + (γλ)δ_{t+1} + (γλ)²δ_{t+2} + ...
δ_t = r_t + γV(s_{t+1}) - V(s_t)  [TD error]

λ ∈ [0,1] trades off bias-variance:
- λ=0: pure TD(0) (low variance, high bias)
- λ=1: pure MC (high variance, low bias)
- λ=0.95: sweet spot (good tradeoff)
```

**Why GAE**: Exponentially decaying trace over multiple steps. Reduces variance without full MC, reduces bias without pure TD.

---

### Actor-Critic Pitfall #1: Critic Not Learning Properly

**Scenario**: User trains actor-critic but critic loss doesn't decrease. Actor improves, but value function plateaus. Agent can't use accurate advantage estimates.

**Problem**:

```python
# WRONG - critic loss computed incorrectly
critic_loss = mean((V(s) - G_t)^2)  # Wrong target!
critic_loss.backward()
```

**The bug**: Critic should learn Bellman equation:

```
V(s) = E[r + γV(s')]
```

If you compute target as G_t directly, you're using Monte Carlo returns (too noisy). If you use r + γV(s'), you're bootstrapping properly.

**Correct approach**:

```python
# RIGHT - Bellman bootstrap target
V_target = r + gamma * V(s').detach()  # Detach next state value!
critic_loss = mean((V(s) - V_target)^2)
```

**Why detach() matters**: If you don't detach V(s'), gradient flows backward through value function, creating a moving target problem.

**Red Flag**: If critic loss doesn't decrease while actor loss decreases, critic isn't learning Bellman equation. Check:

1. Target computation (should be r + γV(s'), not G_t alone)
2. Detach next state value
3. Critic network is separate from actor
4. Different learning rates (critic typically higher than actor)

---

### Critic as Baseline vs Critic as Q-Function

**Important distinction**:

**A2C uses critic as baseline**:

```
V(s) = value of being in state s
A(s,a) = r + γV(s') - V(s)  [TD advantage]
Policy loss = -∇log π(a|s) * A(s,a)
```

**SAC/TD3 use critic as Q-function**:

```
Q(s,a) = expected return from taking action a in s
A(s,a) = Q(s,a) - V(s)
Policy loss = ∇log π(a|s) * Q(s,a) [deterministic policy gradient]
```

**Why the difference**: A2C updates actor and critic together (on-policy). SAC/TD3 decouple them (off-policy):

- Actor never sees the replay buffer
- Critic learns Q from replay buffer
- Actor uses critic's Q estimate (always lagging slightly)

---

## Part 2: A2C - Advantage Actor-Critic

### A2C Architecture

A2C = on-policy advantage actor-critic. Actor and critic train simultaneously with synchronized rollouts:

```
┌─────────────────────────────────────────┐
│  Environment                            │
└────────────┬────────────────────────────┘
             │ states, rewards
             ▼
┌─────────────────────────────────────────┐
│  Actor π(a|s)     Critic V(s)          │
│  Policy network   Value network         │
│  Outputs: action  Outputs: value        │
└────────┬──────────────────┬─────────────┘
         │                  │
         └──────┬───────────┘
                │
         ┌──────▼───────────┐
         │ Advantage        │
         │ A(s,a) = r+γV(s')-V(s)
         └────────┬─────────┘
                  │
         ┌────────▼────────────┐
         │ Actor Loss:         │
         │ -log π(a|s) * A(s,a)
         │                     │
         │ Critic Loss:        │
         │ (V(s) - target)²    │
         └─────────────────────┘
```

### A2C Training Loop

```python
for episode in range(num_episodes):
    states, actions, rewards, values = [], [], [], []

    state = env.reset()
    for t in range(horizon):
        # Actor samples action from policy
        action = actor(state)

        # Step environment
        next_state, reward = env.step(action)

        # Get value estimate (baseline)
        value = critic(state)

        # Store for advantage computation
        states.append(state)
        actions.append(action)
        rewards.append(reward)
        values.append(value)

        state = next_state

    # Advantage estimation (GAE)
    advantages = compute_gae(rewards, values, next_value, gamma, lambda)

    # Actor loss (policy gradient with baseline)
    actor_loss = -log_prob(actions, actor(states)) * advantages
    actor.update(actor_loss)

    # Critic loss (value function learning)
    critic_targets = rewards + gamma * values[1:] + gamma * critic(next_state)
    critic_loss = (critic(states) - critic_targets)^2
    critic.update(critic_loss)
```

### A2C vs A3C

**A2C**: Synchronous - all parallel workers update at same time (cleaner, deterministic)

```
Worker 1  ────┐
Worker 2  ────┼──► Global Model Update ──► All workers receive updated weights
Worker 3  ────┤
Worker N  ────┘
Wait for all workers before next update
```

**A3C**: Asynchronous - workers update whenever they finish (faster wall clock time, messier)

```
Worker 1  ──► Update (1) ──► Continue
Worker 2  ──────► Update (2) ──────► Continue
Worker 3  ──────────► Update (3) ──────► Continue
No synchronization barrier (race conditions possible)
```

**In practice**: A2C is preferred. A3C was important historically (enables multi-GPU training without synchronization) but A2C is cleaner.

---

## Part 3: SAC - Soft Actor-Critic

### SAC Overview

SAC = Soft Actor-Critic. The current SOTA (state-of-the-art) for continuous control. Three key innovations:

1. **Entropy regularization**: Add H(π(·|s)) to objective (maximize entropy + reward)
2. **Auto-tuning entropy coefficient**: Learn α automatically (no manual tuning!)
3. **Off-policy learning**: Learn from replay buffer (sample efficient)

### SAC's Objective Function

Standard policy gradient maximizes:

```
J(π) = E[G_t]
```

SAC maximizes:

```
J(π) = E[G_t + α H(π(·|s))]
     = E[G_t] + α E[H(π(·|s))]
```

**Where**:

- G_t = cumulative reward
- H(π(·|s)) = policy entropy (randomness)
- α = entropy coefficient (how much we value exploration)

**Why entropy**: Exploratory policies (high entropy) discover better strategies. Adding entropy to objective = agent explores automatically.

### SAC Components

```
┌─────────────────────────────────────┐
│  Replay Buffer (off-policy data)    │
└────────────┬────────────────────────┘
             │ sample batch
             ▼
    ┌────────────────────────┐
    │  Actor Network         │
    │  π(a|s) = μ(s) + σ(s) │  (Gaussian policy)
    │  Outputs: mean, std    │
    └────────────────────────┘
             │
             ▼
    ┌────────────────────────┐
    │  Two Critic Networks   │
    │  Q1(s,a), Q2(s,a)     │
    │  Learn Q-values        │
    └────────────────────────┘
             │
             ▼
    ┌────────────────────────┐
    │  Target Networks       │
    │  Q_target1, Q_target2  │
    │  (updated every N)     │
    └────────────────────────┘
             │
             ▼
    ┌────────────────────────┐
    │  Entropy Coefficient   │
    │  α (learned!)          │
    └────────────────────────┘
```

### SAC Training Algorithm

```python
# Initialize
actor = ActorNetwork()
critic1, critic2 = CriticNetwork(), CriticNetwork()
target_critic1, target_critic2 = copy(critic1), copy(critic2)
entropy_alpha = 1.0  # Learned!
target_entropy = -action_dim  # Target entropy (usually -action_dim)

for step in range(num_steps):
    # 1. Collect data (could be online or from buffer)
    state = env.reset() if step % 1000 == 0 else next_state
    action = actor.sample(state)  # π(a|s)
    next_state, reward = env.step(action)
    replay_buffer.add(state, action, reward, next_state, done)

    # 2. Sample batch from replay buffer
    batch = replay_buffer.sample(batch_size=256)
    states, actions, rewards, next_states, dones = batch

    # 3. Critic update (Q-function learning)
    # Compute target Q value using entropy-regularized objective
    next_actions = actor.sample(next_states)
    next_log_probs = actor.log_prob(next_actions, next_states)

    # Use BOTH target critics, take minimum (overestimation prevention)
    Q_target1 = target_critic1(next_states, next_actions)
    Q_target2 = target_critic2(next_states, next_actions)
    Q_target = min(Q_target1, Q_target2)

    # Entropy-regularized target
    y = reward + γ(1 - done) * (Q_target - α * next_log_probs)

    # Update both critics
    critic1_loss = MSE(critic1(states, actions), y)
    critic1.update(critic1_loss)

    critic2_loss = MSE(critic2(states, actions), y)
    critic2.update(critic2_loss)

    # 4. Actor update (policy gradient with entropy)
    # Reparameterization trick: sample actions, compute log probs
    sampled_actions = actor.sample(states)
    sampled_log_probs = actor.log_prob(sampled_actions, states)

    # Actor maximizes Q - α*log_prob (entropy regularization)
    Q1_sampled = critic1(states, sampled_actions)
    Q2_sampled = critic2(states, sampled_actions)
    Q_sampled = min(Q1_sampled, Q2_sampled)

    actor_loss = -E[Q_sampled - α * sampled_log_probs]
    actor.update(actor_loss)

    # 5. Entropy coefficient auto-tuning (SAC's KEY INNOVATION)
    # Learn α to maintain target entropy
    entropy_loss = -α * (sampled_log_probs + target_entropy)
    alpha.update(entropy_loss)

    # 6. Soft update target networks (every N steps)
    if step % update_frequency == 0:
        target_critic1 = τ * critic1 + (1-τ) * target_critic1
        target_critic2 = τ * critic2 + (1-τ) * target_critic2
```

### SAC Pitfall #1: Manual Entropy Coefficient

**Scenario**: User implements SAC but manually sets α=0.2 and training diverges. Agent explores randomly and never improves.

**Problem**: SAC's entire design is that α is **learned automatically**. Setting it manually defeats the purpose.

```python
# WRONG - treating α as fixed hyperparameter
alpha = 0.2  # Fixed!
loss = Q_target - 0.2 * log_prob  # Same penalty regardless of entropy

# Result: If entropy naturally low, penalty still high → policy forced random
#         If entropy naturally high, penalty too weak → insufficient exploration
```

**Correct approach**:

```python
# RIGHT - α is learned via entropy constraint
target_entropy = -action_dim  # For Gaussian: typically -action_dim

# Optimize α to maintain target entropy
entropy_loss = -α * (sampled_log_probs.detach() + target_entropy)
alpha_optimizer.zero_grad()
entropy_loss.backward()
alpha_optimizer.step()

# α adjusts automatically:
# - If entropy too high: α increases (more penalty) → policy becomes more deterministic
# - If entropy too low: α decreases (less penalty) → policy explores more
```

**Red Flag**: If SAC agent explores randomly without improving, check:

1. Is α being optimized? (not fixed value)
2. Is target entropy set correctly? (usually -action_dim)
3. Is log_prob computed with squashed action (after tanh)?

---

### SAC Pitfall #2: Tanh Squashing and Log Probability

**Scenario**: User implements SAC with Gaussian policy but uses policy directly. Log probabilities are computed wrong. Training is unstable.

**Problem**: SAC uses tanh squashing to bound actions:

```
Raw action from network: μ(s) + σ(s)*ε, ε~N(0,1)  → unbounded
Tanh squashed: a = tanh(raw_action)  → bounded in [-1,1]
```

But policy probability must account for this transformation:

```
π(a|s) ≠ N(μ(s), σ²(s))  [Wrong! Ignores tanh]
π(a|s) = |det(∂a/∂raw_action)|^(-1) * N(μ(s), σ²(s))
       = (1 - a²)^2 * N(μ(s), σ²(s))  [Right! Jacobian correction]

log π(a|s) = log N(μ(s), σ²(s)) - 2*log(1 - a²)
```

**The bug**: Computing log_prob without Jacobian correction:

```python
# WRONG
log_prob = normal.log_prob(raw_action) - log(1 + exp(-2*x))
# (standard normal log prob, ignores squashing)

# RIGHT
log_prob = normal.log_prob(raw_action) - log(1 + exp(-2*x))
log_prob = log_prob - 2 * (log(2) - x - softplus(-2*x))  # Add Jacobian term
```

Or simpler:

```python
# PyTorch way
dist = Normal(mu, sigma)
raw_action = dist.rsample()  # Reparameterized sample
action = torch.tanh(raw_action)
log_prob = dist.log_prob(raw_action) - torch.log(1 - action.pow(2) + 1e-6).sum(-1)
```

**Red Flag**: If SAC policy doesn't learn despite updates, check:

1. Are actions being squashed (tanh)?
2. Is log_prob computed with tanh Jacobian term?
3. Is squashing adjustment in entropy coefficient update?

---

### SAC Pitfall #3: Two Critics and Target Networks

**Scenario**: User implements SAC with one critic and gets unstable learning. "I thought SAC just needed entropy regularization?"

**Problem**: SAC uses TWO critics because of Q-function overestimation:

```
Single critic Q(s,a):
- Targets computed as: y = r + γQ_target(s', a')
- Q_target is function of Q (updated less frequently)
- In continuous space, selecting actions via max isn't feasible
- Next action sampled from π (deterministic max removed)
- But Q-values can still overestimate (stochastic environment noise)

Two critics (clipped double Q-learning):
- Use both Q1 and Q2, take minimum: Q_target = min(Q1_target, Q2_target)
- Prevents overestimation (conservative estimate)
- Both updated simultaneously
- Asymmetric: both learn, but target uses minimum
```

**Correct implementation**:

```python
# WRONG - one critic
target = reward + gamma * critic_target(next_state, next_action)

# RIGHT - two critics with min
Q1_target = critic1_target(next_state, next_action)
Q2_target = critic2_target(next_state, next_action)
target = reward + gamma * min(Q1_target, Q2_target)

# Both critics learn
critic1_loss = MSE(critic1(state, action), target)
critic2_loss = MSE(critic2(state, action), target)

# But actor only uses critic1 (or min of both)
Q_current = min(critic1(state, sampled_action), critic2(state, sampled_action))
actor_loss = -(Q_current - alpha * log_prob)
```

**Red Flag**: If SAC diverges, check:

1. Are there two Q-networks?
2. Does target use min(Q1, Q2)?
3. Are target networks updated (soft or hard)?

---

## Part 4: TD3 - Twin Delayed DDPG

### Why TD3 Exists

TD3 = Twin Delayed DDPG. It addresses SAC's cost (two networks, more computation) with deterministic policy gradient (simpler).

**DDPG** (older): Deterministic policy, single Q-network, no entropy. Fast but unstable.

**TD3** (newer): Three tricks to stabilize DDPG:

1. **Twin critics**: Two Q-networks (clipped double Q-learning)
2. **Delayed actor updates**: Update actor every d steps (not every step)
3. **Target policy smoothing**: Add noise to target action before Q evaluation

### TD3 Architecture

```
┌──────────────────────────────────┐
│  Replay Buffer                   │
└────────────┬─────────────────────┘
             │
             ▼
    ┌───────────────────────┐
    │  Actor μ(s)           │
    │  Deterministic policy │
    │  Outputs: action      │
    └───────────────────────┘
             │
             ▼
    ┌─────────────────────────┐
    │  Q1(s,a), Q2(s,a)      │
    │  Two Q-networks        │
    │  (Triple: original+2)  │
    └─────────────────────────┘
             │
             ▼
    ┌─────────────────────────┐
    │  Delayed Actor Update   │
    │  (every d steps)        │
    └─────────────────────────┘
```

### TD3 Training Algorithm

```python
for step in range(num_steps):
    # 1. Collect data
    action = actor(state) + exploration_noise
    next_state, reward = env.step(action)
    replay_buffer.add(state, action, reward, next_state, done)

    if step < min_steps_before_training:
        continue

    batch = replay_buffer.sample(batch_size)
    states, actions, rewards, next_states, dones = batch

    # 2. Critic update (BOTH Q-networks)
    # Trick #3: Target policy smoothing
    next_actions = actor_target(next_states)
    noise = torch.randn_like(next_actions) * target_noise
    noise = torch.clamp(noise, -noise_clip, noise_clip)
    next_actions = torch.clamp(next_actions + noise, -1, 1)  # Add noise, clip

    # Clipped double Q-learning: use minimum
    Q1_target = critic1_target(next_states, next_actions)
    Q2_target = critic2_target(next_states, next_actions)
    Q_target = torch.min(Q1_target, Q2_target)

    y = rewards + gamma * (1 - dones) * Q_target

    # Update both critics
    critic1_loss = MSE(critic1(states, actions), y)
    critic1_optimizer.zero_grad()
    critic1_loss.backward()
    critic1_optimizer.step()

    critic2_loss = MSE(critic2(states, actions), y)
    critic2_optimizer.zero_grad()
    critic2_loss.backward()
    critic2_optimizer.step()

    # 3. Delayed actor update (Trick #2)
    if step % policy_delay == 0:
        # Deterministic policy gradient
        Q1_current = critic1(states, actor(states))
        actor_loss = -Q1_current.mean()

        actor_optimizer.zero_grad()
        actor_loss.backward()
        actor_optimizer.step()

        # Soft update target networks
        for param, target_param in zip(critic1.parameters(), critic1_target.parameters()):
            target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
        for param, target_param in zip(critic2.parameters(), critic2_target.parameters()):
            target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
        for param, target_param in zip(actor.parameters(), actor_target.parameters()):
            target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)
```

### TD3 Pitfall #1: Missing Target Policy Smoothing

**Scenario**: User implements TD3 with twin critics and delayed updates but training still unstable. "I have two critics, why isn't it stable?"

**Problem**: Target policy smoothing is critical. Without it:

```
Next action = deterministic μ_target(s')  [exact, no exploration noise]

If Q-networks overestimate for certain actions:
- Target policy always selects that exact action
- Q-target biased high for that action
- Feedback loop: overestimation → more value → policy selects it more → more overestimation
```

With smoothing:

```
Next action = μ_target(s') + ε_smoothing
- Adds small random noise to target action
- Prevents exploitation of Q-estimation errors
- Breaks feedback loop by adding randomness to target action

Important: Noise is added at TARGET action, not current action!
- Current: exploration_noise (for exploration during collection)
- Target: target_noise (for stability, noise clip small)
```

**Correct implementation**:

```python
# Trick #3: Target policy smoothing
next_actions = actor_target(next_states)
noise = torch.randn_like(next_actions) * target_policy_noise
noise = torch.clamp(noise, -noise_clip, noise_clip)
next_actions = torch.clamp(next_actions + noise, -1, 1)

# Then use these noisy actions for Q-target
Q_target = min(Q1_target(next_states, next_actions),
               Q2_target(next_states, next_actions))
```

**Red Flag**: If TD3 diverges despite two critics, check:

1. Is noise added to target action (not just actor output)?
2. Is noise clipped (noise_clip prevents too much noise)?
3. Are critic targets using smoothed actions?

---

### TD3 Pitfall #2: Delayed Actor Updates

**Scenario**: User implements TD3 with target policy smoothing and twin critics, but updates actor every step. "Do I really need delayed updates?"

**Problem**: Policy updates change actor, which changes actions chosen. If you update actor every step while critics are learning:

```
Step 1: Actor outputs a1, Q(s,a1) = 5, Actor updated
Step 2: Actor outputs a2, Q(s,a2) = 3, Actor wants to stay at a1
Step 3: Critics haven't converged, oscillate between a1 and a2
Result: Actor chases moving target, training unstable
```

With delayed updates:

```
Steps 1-4: Update critics only, let them converge
Step 5: Update actor (once per policy_delay=5)
Steps 6-9: Update critics only
Step 10: Update actor again
Result: Critic stabilizes before actor changes, smoother learning
```

**Typical settings**:

```python
policy_delay = 2  # Update actor every 2 critic updates
# or
policy_delay = 5  # More conservative, every 5 critic updates
```

**Correct implementation**:

```python
if step % policy_delay == 0:  # Only sometimes!
    actor_loss = -critic1(state, actor(state)).mean()
    actor_optimizer.zero_grad()
    actor_loss.backward()
    actor_optimizer.step()

    # Update targets on same schedule
    soft_update(critic1_target, critic1)
    soft_update(critic2_target, critic2)
    soft_update(actor_target, actor)
```

**Red Flag**: If TD3 training unstable, check:

1. Is actor updated only every policy_delay steps?
2. Are target networks updated on same schedule (policy_delay)?
3. Policy_delay typically 2-5

---

### SAC vs TD3 Decision Framework

**Both are SOTA for continuous control. How to choose?**

| Aspect | SAC | TD3 |
|--------|-----|-----|
| **Policy Type** | Stochastic (Gaussian) | Deterministic |
| **Exploration** | Entropy maximization (automatic) | Target policy smoothing |
| **Sample Efficiency** | High (two critics) | High (two critics) |
| **Stability** | Very stable (entropy helps) | Stable (three tricks) |
| **Computation** | Higher (entropy tuning) | Slightly lower |
| **Manual Tuning** | Minimal (α auto-tuned) | Moderate (policy_delay, noise) |
| **When to Use** | Default choice, off-policy | When deterministic better, simpler noise |

**Decision tree**:

1. **Do you prefer stochastic or deterministic policy?**
   - Stochastic (multiple possible actions per state) → SAC
   - Deterministic (one action per state) → TD3

2. **Sample efficiency critical?**
   - Yes, limited data → Both good, slight edge SAC
   - No, lots of data → Either works

3. **How much tuning tolerance?**
   - Want minimal tuning → SAC (α auto-tuned)
   - Don't mind tuning policy_delay, noise → TD3 (simpler conceptually)

4. **Exploration challenges?**
   - Complex exploration (entropy helps) → SAC
   - Simple exploration (policy smoothing enough) → TD3

**Practical recommendation**: Start with SAC. It's more robust (entropy auto-tuning). Switch to TD3 only if you:

- Know you want deterministic policy
- Have tuning expertise for policy_delay
- Need slightly faster computation

---

## Part 5: Continuous Action Handling

### Gaussian Policy Representation

Actor outputs **mean and standard deviation**:

```python
raw_output = actor_network(state)
mu = raw_output[:, :action_dim]
log_std = raw_output[:, action_dim:]
log_std = torch.clamp(log_std, min=log_std_min, max=log_std_max)
std = log_std.exp()

dist = Normal(mu, std)
raw_action = dist.rsample()  # Reparameterized sample
```

**Why log(std)?**: Parameterize log scale instead of scale directly.

- Numerical stability (log prevents underflow)
- Gradient flow smoother
- Prevents std from becoming negative

**Why clamp log_std?**: Prevents std from becoming too small or large.

- Too small: policy becomes deterministic, no exploration
- Too large: policy becomes random, no learning

Typical ranges:

```python
log_std_min = -20  # std >= exp(-20) ≈ 2e-9 (small exploration)
log_std_max = 2    # std <= exp(2) ≈ 7.4 (max randomness)
```

### Continuous Action Squashing (Tanh)

Raw network output unbounded. Use tanh to bound to [-1,1]:

```python
# After sampling from policy
action = torch.tanh(raw_action)
# action now in [-1, 1]

# Scale to environment action range [low, high]
action_scaled = (high - low) / 2 * action + (high + low) / 2
```

**Pitfall**: Log probability must account for squashing (already covered in SAC section).

### Exploration Noise in Continuous Control

**Off-policy methods** (SAC, TD3) need exploration during data collection:

**Method 1: Action space noise** (simpler):

```python
action = actor(state) + noise
noise = torch.randn_like(action) * exploration_std
action = torch.clamp(action, -1, 1)  # Ensure in bounds
```

**Method 2: Parameter noise** (more complex):

```
Add noise to actor network weights periodically
Action = actor_with_noisy_weights(state)
Results in correlated action noise across timesteps (more natural exploration)
```

**Typical settings**:

```python
# For SAC: exploration_std = 0.1 * max_action
# For TD3: exploration_std starts high, decays over time
```

---

## Part 6: Common Bugs and Debugging

### Bug #1: Critic Divergence

**Symptom**: Critic loss explodes, V(s) becomes huge (1e6+), agent breaks.

**Causes**:

1. **Wrong target computation**: Using wrong Bellman target
2. **No gradient clipping**: Gradients unstable
3. **Learning rate too high**: Critic overshoots
4. **Value targets too large**: Reward scale not normalized

**Diagnosis**:

```python
# Check target computation
print("Reward range:", rewards.min(), rewards.max())
print("V(s) range:", v_current.min(), v_current.max())
print("Target range:", v_target.min(), v_target.max())

# Plot value function over time
plt.plot(v_values_history)  # Should slowly increase, not explode

# Check critic loss
print("Critic loss:", critic_loss.item())  # Should decrease, not diverge
```

**Fix**:

```python
# 1. Reward normalization
rewards = (rewards - rewards.mean()) / (rewards.std() + 1e-8)

# 2. Gradient clipping
torch.nn.utils.clip_grad_norm_(critic.parameters(), max_norm=1.0)

# 3. Lower learning rate
critic_lr = 1e-4  # Instead of 1e-3

# 4. Value function target clipping (optional)
v_target = torch.clamp(v_target, -100, 100)
```

---

### Bug #2: Actor Not Learning (Constant Policy)

**Symptom**: Actor loss decreases but policy doesn't change. Same action sampled repeatedly. No improvement in return.

**Causes**:

1. **Policy output not properly parameterized**: Mean/std wrong
2. **Critic signal dead**: Q-values all same, no gradient
3. **Learning rate too low**: Actor updates too small
4. **Advantage always zero**: Critic perfect (impossible) or wrong

**Diagnosis**:

```python
# Check policy output distribution
actions = [actor.sample(state) for _ in range(1000)]
print("Action std:", np.std(actions))  # Should be >0.01
print("Action mean:", np.mean(actions))

# Check critic signal
q_values = critic(states, random_actions)
print("Q range:", q_values.min(), q_values.max())
print("Q std:", q_values.std())  # Should have variation

# Check advantage
advantages = q_values - v_baseline
print("Advantage std:", advantages.std())  # Should be >0
```

**Fix**:

```python
# 1. Ensure policy outputs have variance
assert log_std.mean() < log_std_max - 0.5  # Not clamped to max
assert log_std.mean() > log_std_min + 0.5  # Not clamped to min

# 2. Check critic learns
critic_loss should decrease

# 3. Increase actor learning rate
actor_lr = 3e-4  # Instead of 1e-4

# 4. Debug advantage calculation
if advantage.std() < 0.01:
    print("WARNING: Advantages have no variation, critic might be wrong")
```

---

### Bug #3: Entropy Coefficient Divergence (SAC)

**Symptom**: SAC entropy coefficient α explodes (1e6+), policy becomes completely random, agent stops learning.

**Cause**: Entropy constraint optimization unstable.

```python
# WRONG - entropy loss unbounded
entropy_loss = -alpha * (log_probs + target_entropy)
# If log_probs >> target_entropy, loss becomes huge positive, α explodes
```

**Fix**:

```python
# RIGHT - use log(α) to avoid explosion
log_alpha = torch.log(alpha)
log_alpha_loss = -log_alpha * (log_probs.detach() + target_entropy)
alpha_optimizer.zero_grad()
log_alpha_loss.backward()
alpha_optimizer.step()
alpha = log_alpha.exp()

# Or clip α
alpha = torch.clamp(alpha, min=1e-4, max=10.0)
```

---

### Bug #4: Target Network Never Updated

**Symptom**: Agent learns for a bit, then stops improving. Training plateaus.

**Cause**: Target networks not updated (or updated too rarely).

```python
# WRONG - never update targets
target_critic = copy(critic)  # Initialize once
for step in range(1000000):
    # ... training loop ...
    # But target_critic never updated!
```

**Fix**:

```python
# RIGHT - soft update every step (or every N steps for delayed methods)
tau = 0.005  # Soft update parameter
for step in range(1000000):
    # ... critic update ...
    # Soft update targets
    for param, target_param in zip(critic.parameters(), target_critic.parameters()):
        target_param.data.copy_(tau * param.data + (1-tau) * target_param.data)

# Or hard update (copy all weights) every N steps
if step % update_frequency == 0:
    target_critic = copy(critic)
```

---

### Bug #5: Gradient Flow Through Detached Tensors

**Symptom**: Actor loss computation succeeds, but actor parameters don't update.

**Cause**: Critic detached but actor expects gradients.

```python
# WRONG
for step in range(1000):
    q_value = critic(state, action).detach()  # Detached!
    actor_loss = -q_value.mean()
    actor.update(actor_loss)  # Gradient won't flow through q_value!

# Result: actor_loss always 0 (constant from q_value.detach())
# Actor parameters updated but toward constant target (no signal)
```

**Fix**:

```python
# RIGHT - don't detach when computing actor loss
q_value = critic(state, action)  # No detach!
actor_loss = -q_value.mean()
actor.update(actor_loss)  # Gradient flows through q_value

# Detach where appropriate:
# - Value targets: v_target = (r + gamma * v_next).detach()
# - Stop gradient in critic: q_target = (r + gamma * q_next.detach()).detach()
# But NOT when computing actor loss
```

---

## Part 7: When to Use Actor-Critic vs Alternatives

### Actor-Critic vs Policy Gradient (REINFORCE)

| Factor | Actor-Critic | Policy Gradient |
|--------|--------------|-----------------|
| **Variance** | Low (baseline reduces) | High (full return) |
| **Sample Efficiency** | High | Low |
| **Convergence Speed** | Fast | Slow |
| **Complexity** | Two networks | One network |
| **Stability** | Better | Worse (high noise) |

**Use Actor-Critic when**: Continuous actions, sample efficiency matters, training instability

**Use Policy Gradient when**: Simple problem, don't need value function, prefer simpler code

---

### Actor-Critic vs Q-Learning (DQN)

| Factor | Actor-Critic | Q-Learning |
|--------|--------------|-----------|
| **Action Space** | Continuous (natural) | Discrete (requires all Q values) |
| **Sample Efficiency** | High | Very high |
| **Stability** | Good | Can diverge (overestimation) |
| **Complexity** | Two networks | One network (but needs tricks) |

**Use Actor-Critic for**: Continuous actions, robotics, control

**Use Q-Learning for**: Discrete actions, games, navigation

---

### Actor-Critic (On-Policy A2C) vs Off-Policy (SAC, TD3)

| Factor | A2C (On-Policy) | SAC/TD3 (Off-Policy) |
|--------|-----------------|---------------------|
| **Sample Efficiency** | Moderate | High (replay buffer) |
| **Stability** | Good | Excellent |
| **Complexity** | Simpler | More complex |
| **Data Reuse** | Limited (one pass) | High (replay buffer) |
| **Parallel Training** | Excellent (A3C) | Limited (off-policy break) |

**Use A2C when**: Want simplicity, have parallel workers, on-policy is okay

**Use SAC/TD3 when**: Need sample efficiency, offline data possible, maximum stability

---

## Part 8: Implementation Checklist

### Pre-Training Checklist

- [ ] Actor outputs mean and log_std separately
- [ ] Log_std clamped: `log_std_min <= log_std <= log_std_max`
- [ ] Action squashing with tanh (bounded to [-1,1])
- [ ] Log probability computation includes tanh Jacobian (SAC/A2C)
- [ ] Critic network separate from actor
- [ ] Critic loss is value bootstrap (r + γV(s'), not G_t)
- [ ] Two critics for SAC/TD3 (or one for A2C)
- [ ] Target networks initialized as copies of main networks
- [ ] Replay buffer created (for off-policy methods)
- [ ] Advantage estimation (GAE preferred, MC acceptable)

### Training Loop Checklist

- [ ] Data collection uses current actor (not target)
- [ ] Critic updated with Bellman target: `r + γV(s').detach()`
- [ ] Actor updated with advantage signal: `-log_prob(a) * A(s,a)` or `-Q(s,a)`
- [ ] Target networks soft updated: `τ * main + (1-τ) * target`
- [ ] For SAC: entropy coefficient α being optimized
- [ ] For TD3: delayed actor updates (every policy_delay)
- [ ] For TD3: target policy smoothing (noise + clip)
- [ ] Gradient clipping applied if losses explode
- [ ] Learning rates appropriate (critic_lr typically >= actor_lr)
- [ ] Reward normalization or clipping applied

### Debugging Checklist

- [ ] Critic loss decreasing over time?
- [ ] V(s) and Q(s,a) values in reasonable range?
- [ ] Policy entropy decreasing (exploration → exploitation)?
- [ ] Actor loss decreasing?
- [ ] Return increasing over episodes?
- [ ] No NaN or Inf in losses?
- [ ] Advantage estimates have variation?
- [ ] Policy output std not stuck at min/max?

---

## Part 9: Comprehensive Pitfall Reference

### 1. Critic Loss Not Decreasing

- Wrong Bellman target (should be r + γV(s'))
- Critic weights not updating (zero gradients)
- Learning rate too low
- Target network staleness (not updated)

### 2. Actor Not Improving

- Critic broken (no signal)
- Advantage estimates all zero
- Actor learning rate too low
- Policy parameterization wrong (no variance)

### 3. Training Unstable (Divergence)

- Missing target networks
- Critic loss exploding (wrong target, high learning rate)
- Entropy coefficient exploding (SAC: should be log(α))
- Actor updates every step (should delay, especially TD3)

### 4. Policy Stuck at Random Actions (SAC)

- Manual α fixed (should be auto-tuned)
- Target entropy wrong (should be -action_dim)
- Entropy loss gradient wrong direction

### 5. Policy Output Clamped to Min/Max Std

- Log_std range too tight (check log_std_min/max)
- Network initialization pushing to extreme values
- No gradient clipping preventing adjustment

### 6. Tanh Squashing Ignored

- Log probability not adjusted for squashing
- Missing Jacobian term in SAC/policy gradient
- Action scaling inconsistent

### 7. Target Networks Never Updated

- Forgot to create target networks
- Update function called but not applied
- Update frequency too high (no learning)

### 8. Off-Policy Break (Experience Replay)

- Actor training on old data (should use current replay buffer)
- Data distribution shift (actions from old policy)
- Batch importance weights missing (PER)

### 9. Advantage Estimates Biased

- GAE parameter λ wrong (should be 0.95-0.99)
- Bootstrap incorrect (wrong value target)
- Critic too inaccurate (overcorrection)

### 10. Entropy Coefficient Issues (SAC)

- Manual tuning instead of auto-tuning
- Entropy target not set correctly
- Log(α) optimization not used (causes explosion)

---

## Part 10: Real-World Examples

### Example 1: SAC for Robotic Arm Control

**Problem**: Robotic arm needs to reach target position. Continuous joint angles.

**Setup**:

```python
state_dim = 18  # 6 joint angles + velocities
action_dim = 6  # Joint torques
action_range = [-1, 1]  # Normalized

actor = ActorNetwork(state_dim, action_dim)  # Outputs μ, log_std
critic1 = CriticNetwork(state_dim, action_dim)
critic2 = CriticNetwork(state_dim, action_dim)

target_entropy = -action_dim  # -6
alpha = 1.0
```

**Training**:

```python
for step in range(1000000):
    # Collect experience
    state = env.reset() if done else next_state
    action = actor.sample(state)
    next_state, reward, done = env.step(action)
    replay_buffer.add(state, action, reward, next_state, done)

    if len(replay_buffer) < min_buffer_size:
        continue

    batch = replay_buffer.sample(256)

    # Critic update
    next_actions = actor.sample(batch.next_states)
    next_log_probs = actor.log_prob(next_actions, batch.next_states)
    q1_target = target_critic1(batch.next_states, next_actions)
    q2_target = target_critic2(batch.next_states, next_actions)
    target = batch.rewards + gamma * (1-batch.dones) * (
        torch.min(q1_target, q2_tar

…(truncated)
