TD Recovery
How the Bridge Works (v2)
The bridge runs a background reconciler thread that continuously manages connectivity:
- Config polling (every 1s): Watches
.embody/envoy.json for mtime changes. Automatically switches to the new active instance when the config is updated.
- Heartbeat (every 10s, fixed --
HEARTBEAT_TICK_S): Pings the backend to detect connect/disconnect transitions, and reads how long TD's main thread has been away from Envoy's request loop.
- Process discovery: Detects new and exited TD processes via
find_all_td_pids(). Forces a config re-read when new TDs appear.
- Tool cache: Persists the tool list to disk so new sessions start with full tools immediately, without waiting for a backend round-trip.
- Single-attempt forwarding: Failed requests return an error immediately -- no per-request retry loop. The reconciler handles recovery in the background.
- Streaming forwards (A-46): Tool responses stream incrementally -- server-pushed notifications reach the client per-frame, an idle window bounds a stalled stream, and an absolute cap bounds the whole forward. If a suspected STREAMING defect is wedging every forward (and only then), the env var
EMBODY_BRIDGE_NO_STREAM=1 (set in .mcp.json's env block, then reopen the session) falls back to read-to-EOF with the same parser: pushed frames then arrive batched at the end, and the read is bounded ONLY by the per-recv socket timeout. It is a lever for a broken parser, not for a broken peer -- and note Envoy regenerates .mcp.json on config deploys, so re-add the var if it disappears.
The Envoy Liveness Watchdog (TD-side)
Connectivity self-heals at two independent layers, so a dropped connection almost never needs manual recovery:
- The bridge reconciler (client-side, described above) reconnects the STDIO bridge to a live Envoy.
- The Envoy liveness watchdog (TD-side, in EnvoyExt) revives the Envoy MCP server itself. It is a pure
run()-loop that probes the server socket every ~4s and revives whenever Envoy is enabled-but-down -- a dead socket while TD keeps running, OR a project.save() / extension reinit that took the server down -- force-freeing port 9870 if it is still held and rebinding in seconds. No operator, no timer. It is armed from EnvoyExt's __init__ and tied to the instance lifetime (one loop per instance, dying only when a reinit replaces the instance, whose __init__ then arms a fresh one), so a save/reinit whose post-reinit auto-start never completes still leaves a live watchdog to recover Envoy.
This is verified, not aspirational: killing the live listener socket self-heals in ~6s, and two project.save() cycles each had the watchdog fire (running=False), force-free the still-held port, and rebind in ~1s -- all with no restart, after which the bridge returns to connected:true on its own.
Recovery -- Manual Intervention
Most connectivity issues self-heal (see the two layers above). Before any manual action, read get_td_status and distinguish the cause:
connected:false while td_process_alive:true (Envoy unreachable but TD still running) is the dropped-socket zombie. The TD-side watchdog self-heals it in ~6-8s. WAIT ~10s and re-check get_td_status (or probe the port directly: python3 -c "import socket; socket.create_connection(('127.0.0.1',9870),0.4)"). Do NOT restart_td, relaunch TD, or toggle Envoy for this -- it defeats the watchdog and is almost never necessary. Only escalate if it genuinely has not recovered after ~15s.
- Editing an extension
.py (EnvoyExt / EmbodyExt / TDXNExt) does NOT need a restart either -- the source DATs have syncfile=True and reinit on change, so edits go live on their own.
envoy_unresponsive:true (TD alive, port accepting, Envoy silent 60s+ since unresponsive_since) means TD looks frozen, or stuck in one very long call; main_thread_stalled:true (Envoy answers, but TD's main thread has not come back to Envoy's request loop for 60s+ since stalled_since) means a blocking dialog or a call that never returned. The TD-side watchdog cannot help either -- it runs on that thread. Call list_dialogs first; with no dialog up, report the timestamp to the user and let them decide. Never kill TD yourself. A call pinned on the frozen TD is answered once envoy_unresponsive has held another 120s (180s of silence), saying it was abandoned.
Manual recovery below is only for when TD is actually down or the bridge process itself is broken:
- Call
get_td_status: This is always available (even when TD is down). It shows connection state, process liveness, instance registry, and any unregistered TD processes.
- If TD is not running: Call
launch_td. The bridge will launch TD with the configured .toe file and wait for Envoy to become reachable.
- If the wrong instance is active: Call
switch_instance to list or switch instances. The reconciler also auto-switches when .embody/envoy.json is edited.
- If the bridge process is stuck: Tell the user to reopen this session/conversation -- this is always the first recovery step. Only if that fails, suggest restarting the MCP server as a fallback.
Common failure: stale active instance
The most frequent cause of connectivity issues is .embody/envoy.json having active set to an instance whose TD process is no longer running. The bridge's reconciler detects this automatically via heartbeat failures and reports it in get_td_status. Use launch_td or switch_instance to recover.
Common failure: broken venv
Symptoms: Bridge process doesn't start at all, or starts and immediately exits. The envoy-bridge.log file inside Embody's logs directory (see the Logfolder parameter on the Embody COMP) is empty or shows a Python traceback about a missing interpreter. .mcp.json command points to a .venv/ Python that doesn't work.
Cause: The venv was created from a TD Python installation that has since been upgraded or removed. The home key in .venv/pyvenv.cfg points to a dead path (common on Windows with versioned TD directories like TouchDesigner.2025.32460/).
Diagnosis:
- Read
.mcp.json -- find the command path for the envoy server.
- Test it: run
<command> -c "print(1)" via Bash. If it fails with "No Python at ..." or similar, the venv is broken.
- Confirm by reading
.venv/pyvenv.cfg -- check if the home path points to an existing directory.
Fix:
- Delete the broken venv:
rm -rf <project_dir>/.venv
- Envoy will recreate it on next startup. Tell the user to toggle Envoy off and on in TD, or restart TD.
- After recreation, reopen the Claude Code session so the bridge reconnects with the new venv Python.
Prevention: Envoy validates the venv Python on startup and logs a warning if broken. Check TD textport for "failed to execute" warnings after TD upgrades.
A "Frozen" TD May Be a Masked Crash
A worker-thread-rich project (MCP servers, download threads, any threading.Thread) that touched TD objects from a worker -- including worker-side run()/td.run(), which silently corrupts TD state (see rules/td-python.md) -- can die as a native access violation inside libTD.dll far from any worker thread. TD's crash handler can then hang inside its own CrashAutoSave pickling (/sys/pickleUtils), so the process looks FROZEN: no crash dump, no CrashAutoSave written, no exit (Derivative-confirmed failure signature, 2026-08-17).
- Distinguish it from a real hang:
py-spy dump --pid <td_pid> --native works against TD's embedded Python 3.11, including hung processes. A masked crash shows the main thread stuck in crash-handler / pickling frames rather than in a cook or a Python loop.
- Do not kill the process yourself -- report the diagnosis and let the user decide.
- Root-cause hunt: audit every worker thread in the project for TD-object touches, especially
run() calls -- the corruption site is usually nowhere near the crash site.
- Operational mitigation (UNVERIFIED against official docs): Derivative support has reported that
TOUCH_QUICK_CRASH=1, set in the LAUNCHER environment (TD reads it at boot only), skips the slow crash handler so a supervisor can observe the death and relaunch. The variable does not appear in official documentation as of 2026-08-17 -- verify with Derivative before relying on it.
1---2name: td-recovery3description: MUST READ when Envoy/TD connectivity is broken and has not self-healed: bridge internals, manual recovery, stale instances, broken venv.4---5<!-- Generated by Embody/Envoy - Do not remove this comment -->67# TD Recovery89## How the Bridge Works (v2)1011The bridge runs a background reconciler thread that continuously manages connectivity:1213- **Config polling** (every 1s): Watches `.embody/envoy.json` for mtime changes. Automatically switches to the new active instance when the config is updated.14- **Heartbeat** (every 10s, fixed -- `HEARTBEAT_TICK_S`): Pings the backend to detect connect/disconnect transitions, and reads how long TD's main thread has been away from Envoy's request loop.15- **Process discovery**: Detects new and exited TD processes via `find_all_td_pids()`. Forces a config re-read when new TDs appear.16- **Tool cache**: Persists the tool list to disk so new sessions start with full tools immediately, without waiting for a backend round-trip.17- **Single-attempt forwarding**: Failed requests return an error immediately -- no per-request retry loop. The reconciler handles recovery in the background.18- **Streaming forwards (A-46)**: Tool responses stream incrementally -- server-pushed notifications reach the client per-frame, an idle window bounds a stalled stream, and an absolute cap bounds the whole forward. If a suspected STREAMING defect is wedging every forward (and only then), the env var `EMBODY_BRIDGE_NO_STREAM=1` (set in `.mcp.json`'s `env` block, then reopen the session) falls back to read-to-EOF with the same parser: pushed frames then arrive batched at the end, and the read is bounded ONLY by the per-recv socket timeout. It is a lever for a broken parser, not for a broken peer -- and note Envoy regenerates `.mcp.json` on config deploys, so re-add the var if it disappears.1920## The Envoy Liveness Watchdog (TD-side)2122Connectivity self-heals at two independent layers, so a dropped connection almost never needs manual recovery:2324- **The bridge reconciler** (client-side, described above) reconnects the STDIO bridge to a live Envoy.25- **The Envoy liveness watchdog** (TD-side, in EnvoyExt) revives the Envoy MCP server itself. It is a pure `run()`-loop that probes the server socket every ~4s and revives whenever Envoy is enabled-but-down -- a dead socket while TD keeps running, OR a `project.save()` / extension reinit that took the server down -- force-freeing port 9870 if it is still held and rebinding in seconds. No operator, no timer. It is armed from EnvoyExt's `__init__` and tied to the instance lifetime (one loop per instance, dying only when a reinit replaces the instance, whose `__init__` then arms a fresh one), so a save/reinit whose post-reinit auto-start never completes still leaves a live watchdog to recover Envoy.2627This is verified, not aspirational: killing the live listener socket self-heals in ~6s, and two `project.save()` cycles each had the watchdog fire (`running=False`), force-free the still-held port, and rebind in ~1s -- all with no restart, after which the bridge returns to `connected:true` on its own.2829## Recovery -- Manual Intervention3031Most connectivity issues self-heal (see the two layers above). Before any manual action, read `get_td_status` and distinguish the cause:3233- **`connected:false` while `td_process_alive:true`** (Envoy unreachable but TD still running) is the dropped-socket zombie. The TD-side watchdog self-heals it in ~6-8s. **WAIT ~10s and re-check `get_td_status`** (or probe the port directly: `python3 -c "import socket; socket.create_connection(('127.0.0.1',9870),0.4)"`). **Do NOT `restart_td`, relaunch TD, or toggle Envoy for this** -- it defeats the watchdog and is almost never necessary. Only escalate if it genuinely has not recovered after ~15s.34- Editing an extension `.py` (EnvoyExt / EmbodyExt / TDXNExt) does NOT need a restart either -- the source DATs have `syncfile=True` and reinit on change, so edits go live on their own.35- **`envoy_unresponsive:true`** (TD alive, port accepting, Envoy silent 60s+ since `unresponsive_since`) means TD looks frozen, or stuck in one very long call; **`main_thread_stalled:true`** (Envoy answers, but TD's main thread has not come back to Envoy's request loop for 60s+ since `stalled_since`) means a blocking dialog or a call that never returned. The TD-side watchdog cannot help either -- it runs on that thread. Call `list_dialogs` first; with no dialog up, report the timestamp to the user and let them decide. Never kill TD yourself. A call pinned on the frozen TD is answered once `envoy_unresponsive` has held another 120s (180s of silence), saying it was abandoned.3637Manual recovery below is only for when TD is actually down or the bridge process itself is broken:38391. **Call `get_td_status`**: This is always available (even when TD is down). It shows connection state, process liveness, instance registry, and any unregistered TD processes.402. **If TD is not running**: Call `launch_td`. The bridge will launch TD with the configured `.toe` file and wait for Envoy to become reachable.413. **If the wrong instance is active**: Call `switch_instance` to list or switch instances. The reconciler also auto-switches when `.embody/envoy.json` is edited.424. **If the bridge process is stuck**: Tell the user to **reopen this session/conversation** -- this is always the first recovery step. Only if that fails, suggest restarting the MCP server as a fallback.4344### Common failure: stale active instance4546The most frequent cause of connectivity issues is `.embody/envoy.json` having `active` set to an instance whose TD process is no longer running. The bridge's reconciler detects this automatically via heartbeat failures and reports it in `get_td_status`. Use `launch_td` or `switch_instance` to recover.4748### Common failure: broken venv4950**Symptoms**: Bridge process doesn't start at all, or starts and immediately exits. The `envoy-bridge.log` file inside Embody's logs directory (see the `Logfolder` parameter on the Embody COMP) is empty or shows a Python traceback about a missing interpreter. `.mcp.json` command points to a `.venv/` Python that doesn't work.5152**Cause**: The venv was created from a TD Python installation that has since been upgraded or removed. The `home` key in `.venv/pyvenv.cfg` points to a dead path (common on Windows with versioned TD directories like `TouchDesigner.2025.32460/`).5354**Diagnosis**:551. Read `.mcp.json` -- find the `command` path for the envoy server.562. Test it: run `<command> -c "print(1)"` via Bash. If it fails with "No Python at ..." or similar, the venv is broken.573. Confirm by reading `.venv/pyvenv.cfg` -- check if the `home` path points to an existing directory.5859**Fix**:601. Delete the broken venv: `rm -rf <project_dir>/.venv`612. Envoy will recreate it on next startup. Tell the user to toggle Envoy off and on in TD, or restart TD.623. After recreation, reopen the Claude Code session so the bridge reconnects with the new venv Python.6364**Prevention**: Envoy validates the venv Python on startup and logs a warning if broken. Check TD textport for "failed to execute" warnings after TD upgrades.6566## A "Frozen" TD May Be a Masked Crash6768A worker-thread-rich project (MCP servers, download threads, any `threading.Thread`) that touched TD objects from a worker -- including worker-side `run()`/`td.run()`, which silently corrupts TD state (see rules/td-python.md) -- can die as a native access violation inside libTD.dll far from any worker thread. TD's crash handler can then hang inside its own CrashAutoSave pickling (`/sys/pickleUtils`), so the process looks FROZEN: no crash dump, no CrashAutoSave written, no exit (Derivative-confirmed failure signature, 2026-08-17).6970- **Distinguish it from a real hang**: `py-spy dump --pid <td_pid> --native` works against TD's embedded Python 3.11, including hung processes. A masked crash shows the main thread stuck in crash-handler / pickling frames rather than in a cook or a Python loop.71- **Do not kill the process yourself** -- report the diagnosis and let the user decide.72- **Root-cause hunt**: audit every worker thread in the project for TD-object touches, especially `run()` calls -- the corruption site is usually nowhere near the crash site.73- **Operational mitigation (UNVERIFIED against official docs)**: Derivative support has reported that `TOUCH_QUICK_CRASH=1`, set in the LAUNCHER environment (TD reads it at boot only), skips the slow crash handler so a supervisor can observe the death and relaunch. The variable does not appear in official documentation as of 2026-08-17 -- verify with Derivative before relying on it.