THE LOOP
A screenshot shows a moment; it cannot show a flow. THE LOOP records the actual running target, watches the recording, and judges it against pass criteria you state in plain language. The loop only observes — YOU apply the fixes, then iterate. On pass it renders a before/after MP4 + GIF as proof.
UI / web page / desktop window
watch-skill loop start "<url | screen: | window:<title> | file>" "<pass criteria>" [--script '<json steps>']
# ...apply the suggested fixes to the code...
watch-skill loop iterate <loop_id>
- Pass criteria are ordinary sentences: "the checkout completes, no error toast ever appears, the total is never NaN". Negative claims ("never X") and exemplars ("a real price (like $29.00)") are both understood and enforced deterministically.
--scriptreplays the same clicks/fills every iteration, so the comparison is honest.- Only call
iterateafter you actually changed something. It diffs against the previous iteration: fixed / unchanged / new.
Generated video (Manim, Remotion, ffmpeg, AI-gen)
watch-skill loop video-gen "<what the video must show>" "<render command>" --output <file>
Re-runs the generator each iteration and judges the fresh render against the spec.
Game / simulation
watch-skill loop game "<canvas-url | window:<title> | screen:>" "<pass criteria>" [--run "<launch cmd>"]
Catches the failures screenshots miss: a NaN score counter, black flicker frames, sprites that vanish mid-motion.
Watch for a condition (monitoring)
watch-skill loop monitor "<folder | url | screen: | window:<title>>" "<condition>" [--interval 10] [--max-checks 10]
Bounded — it always terminates. Events land in events.jsonl under the
loop directory as structured records.
Record without judging
watch-skill capture "<target>" [--duration 10]
Capture alone never critiques; it just records, analyzes, and indexes.
Use loop start when there are pass criteria.