Run Notebook on Colab
Executes a local Jupyter notebook on a Colab GPU runtime via SSH + papermill so the agent can see errors and fix them directly without manual copy-paste. File transfer is scp only — no git round-trip.
Inputs
| Input | Description | Example |
|---|---|---|
| Notebook path | Repo-relative .ipynb |
notebooks/ad_hoc/experiment_foo.ipynb |
| Colab hostname | User provides after manual bootstrap (Phase 2) | loud-turkey-abc.trycloudflare.com |
If the user gives a short name, resolve it by searching the repo's notebooks/ directory.
References (load when relevant)
kb/context/colab/colab-compatibility.md— papermill compatibility checklist (drive mount guards, magics, git-sync conflicts)kb/context/colab/colab-sync-protocol.md— scp-only file sync rule and "no git inside notebook" policykb/context/colab/gpu-review-patterns.md— GPU anti-patterns to flag during preflight (Critical only by default)
Prerequisites (per-repo)
cloudflaredinstalled locally (brew install cloudflared)- SSH key at
~/.ssh/colab_keymatching the public key installed by the bootstrap notebook - Repo contains
notebooks/ad_hoc/colab_ssh_bootstrap.ipynbandscripts/connect_colab.sh(or equivalent)
Steps
Phase 1 — Prepare the notebook (delegate to subagent)
Dispatch a general-purpose subagent (or experiment-code-change when within an iterative experiment) to:
- Read the notebook and check it against
kb/context/colab/colab-compatibility.md. - Apply execution-safe fixes (drive mount guard, replace magics with
os.makedirs/subprocess.run, gate anygit reset --hardcells behindSKIP_GIT_REPO_SYNC). - Return a one-line summary of what changed.
Phase 2 — Connect to Colab
Resolve the repo path via fetch-repo-path if not already known.
Automated (preferred): run the repo's scripts/trigger_colab_bootstrap.py to launch Chrome with a saved session, connect a GPU runtime, and relay the hostname back via ntfy.
Manual fallback: open the bootstrap notebook in Colab, Run All, and have the user paste the *.trycloudflare.com hostname.
Verify connectivity:
ssh <SSH_OPTS> root@<HOSTNAME> \
"echo OK && nvidia-smi --query-gpu=name --format=csv,noheader"
Phase 3 — Execute (delegate to subagent)
Dispatch a general-purpose subagent (or experiment-runtime when within an iterative experiment) to:
scpthe notebook to<REMOTE_REPO_ROOT>/<NOTEBOOK_PATH>on Colab.Launch detached — never run papermill in the foreground of an ssh session (a dropped tunnel or local tool timeout would SIGHUP a long training run):
ssh <SSH_OPTS> root@<HOSTNAME> \ "cd <REMOTE_REPO_ROOT> && nohup papermill <NOTEBOOK_PATH> <OUT_NB> --log-output \ > run_r<round>_a<attempt>.log 2>&1 & echo \$! > run_r<round>_a<attempt>.pid"Monitor with a per-job time budget (see "Headless-safe execution" below): poll the run log with short-lived ssh calls at interval
max(1 min, budget/10); the run is done when the pid is gone or the papermill exit marker appears in the log.On non-zero exit: read traceback, dispatch a code-change subagent to fix the failing cell, re-scp, re-run. Up to 3 retries; otherwise escalate.
On success: extract metrics / training curves from the papermill output log.
The lead never runs scp/ssh/papermill commands directly.
Headless-safe execution (mandatory)
Rules that keep unattended runs from stalling on permission prompts or killing long trainings:
- Job lifetime is decoupled from connection lifetime. Training always launches via
nohup ... &with a pid file and log file (Phase 3 step 2). Local ssh timeouts bound only the connection, never the remote job. - Per-job time budget, not fixed timeouts. Derive
time_budget_minutesfrom the planner'sestimated_runtime_class/ critic'sruntime_budgetwhen inside an experiment loop, or from the user/notebook for standalone long trainings. Poll interval scales with budget:max(1 min, budget/10). - Progress-based hang detection — never kill on elapsed time alone. Hung = run log size/mtime unchanged for
max(10 min, budget/3). A progressing job is NEVER killed, even over budget: log anover_budgetflag, continue, and mark the result's comparability. Remotepkill -f papermillis allowed ONLY on confirmed no-progress hang or explicit user abort, and ONLY as a remote command (ssh <SSH_OPTS> root@<HOSTNAME> 'pkill -f papermill') — never a local kill. - No confirmation-gated local commands (rm/rmdir/unlink/trash, mv, kill/killall/pkill, sudo, git push/reset/clean/restore): they block forever in unattended runs. Fresh unique filenames per round+attempt (
<nb>_out_r<round>_a<attempt>.*); never rm/mv prior outputs; place results withcpor by scp-ing directly to the final path. - Bootstrap cleanup: the Chrome instance from
trigger_colab_bootstrap.pyis left running or must self-terminate; neverpkillit from the agent.
SSH/SCP options (subagent applies)
-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null \
-o IdentityFile=~/.ssh/colab_key \
-o ConnectTimeout=5 -o ServerAliveInterval=5 -o ServerAliveCountMax=2 \
-o ProxyCommand="cloudflared access ssh --hostname <HOSTNAME>"
Connection-level timeouts only — they bound the ssh session, never the remote job (which is detached via nohup).
Paths
| Variable | Default value |
|---|---|
<REMOTE_REPO_ROOT> |
/content/drive/MyDrive/colab/<project_name>/<project_name> |
| Dataset cache | /content/drive/MyDrive/colab/data/<project_name> |
| SSH key | ~/.ssh/colab_key (key auth only) |
<project_name> defaults to the repo directory name; resolve with fetch-repo-path and override only if the user specifies one.
Output
- On success: notebook output path, parsed metrics, runtime in minutes.
- On failure after 3 retries: classified failure (per
kb/context/experiments/experiment-contracts.mdfailure classes if within an experiment loop), error trace, suggested next step.