Product + Technical Spec: Target User Agentic Harness (v0)
Status: PRD draft (not implemented).
Objective
Build a repeatable harness that simulates how real target users interact with Brood:
- import realistic photo sets,
- provide project context,
- use Mother + Abilities over multiple turns,
- log explicit "open pondering" questions each turn,
- produce scored, reproducible run artifacts.
Feasibility Snapshot
Feasibility is high because the required primitives already exist:
brood chat --out ... --events ...supports multi-turn interaction.- Abilities already exist as slash commands (
/diagnose,/argue,/bridge,/swap_dna,/triforce, etc.). events.jsonlis append-only and machine-readable.brood_engine/harnessalready handles deterministic multi-step experiment runs and telemetry patterns.
The missing piece is orchestration for persona-driven user behavior, not core model capability.
v0 Scope
- Adapter: CLI-first (
brood chatPTY) for speed and determinism. - Persona + scenario pack input (JSON).
- Agentic turn loop with explicit reflection logging.
- Ability coverage tracking and stop conditions.
- Scorecard per scenario and aggregate report.
Non-Goals (v0)
- Pixel-level desktop UI automation.
- Perfect simulation of human behavior.
- Using reflection text as hidden reasoning. Reflection is explicit, user-visible artifact data.
Architecture
Components:
- Scenario pack loader: validates persona/scenario config.
- Session runner: executes one scenario against one persona.
- Chat adapter: sends utterances/commands to
brood chat, tailsevents.jsonl. - Policy layer: chooses next action from scenario goals + recent events.
- Reflection writer: writes explicit turn reflections to
reflections.jsonl. - Evaluator: computes success and rubric scores from artifacts/events/reflections.
Suggested code placement:
brood_engine/harness/user_sim.py(core loop + data models)scripts/target_user_harness.py(CLI entrypoint)docs/target_user_harness.schema.json(config contract)
Run Artifacts
For each scenario run:
session.json: resolved persona/scenario/policy.events.jsonl: raw engine events (from Brood run dir).transcript.jsonl: adapter-level IO (user_input,assistant_output,command).reflections.jsonl: explicit per-turn pondering.scorecard.json: rubric + pass/fail result.summary.json: compact outcome for dashboarding.
State Machine
States:
initstart_chatimport_inputsplan_turnactobservereflectscore_checkpointdoneerror
Transitions:
init -> start_chatwhen scenario validates.start_chat -> import_inputswhen PTY and events are live.import_inputs -> plan_turnafter/useand required multi-image context setup.plan_turn -> actwhen next action is chosen.act -> observeafter command/utterance dispatch.observe -> reflectafter event delta window closes.reflect -> score_checkpointevery turn.score_checkpoint -> plan_turnwhile stop condition is unmet.score_checkpoint -> doneon success, budget exhaustion, or max turns.- Any state ->
erroron unrecoverable adapter/engine failure.
Stop conditions:
- Success criteria satisfied.
max_turnsreached.max_cost_usdreached.max_runtime_sreached.- Consecutive failure threshold reached.
Action Model
Two action forms are supported:
utterance: natural language sent to Mother.command: explicit slash command to an Ability.
Supported command.name in v0:
usediagnosedescribecanvas_contextblendbridgeswap_dnaargueextract_ruleodd_one_outtriforcerecastqualityfastcheaperbetter
Interpretation:
- Use
utterancewhen simulating ambiguous user intent. - Use
commandwhen simulating power-user behavior and explicit Ability invocation.
Open Pondering Contract
Each turn must write one reflection record:
turn: integeraction_taken: normalized action recordobserved_signals: key events/artifacts in that turnwhat_worked: short textwhat_failed: short textopen_questions: array of unresolved user-style questionsnext_hypothesis: what the simulated user will try nextconfidence: float0.0..1.0
This keeps reasoning explicit and auditable.
Scoring (v0)
Per scenario score is weighted:
goal_progress(35%)artifact_quality_proxy(20%)ability_usage_quality(15%)efficiency_cost_time(15%)resilience(10%)reflection_quality(5%)
Pass gate:
goal_progress >= 0.7- at least one required ability used
- no fatal adapter errors
Three Benchmark Scenarios (v0)
drop_lookbook_rescue
- Persona: indie streetwear founder.
- Inputs: 3 rough iPhone product/lifestyle photos.
- Goal: produce 2 hero images for a launch post in <12 turns.
- Required abilities:
diagnose,bridgeorblend,argue.
tattoo_flash_direction_find
- Persona: tattoo artist exploring a motif system.
- Inputs: 2 references + 1 sketch.
- Goal: produce one coherent "centroid" concept + rule extraction.
- Required abilities:
triforce,extract_rule,odd_one_out.
creator_thumbnail_pack
- Persona: solo creator shipping a video by tonight.
- Inputs: selfie + screenshot + mood reference.
- Goal: produce 3 viable thumbnail candidates and select one direction.
- Required abilities:
swap_dnaorbridge,diagnose,argue.
Concrete sample config for these scenarios is in docs/target_user_harness.scenarios.sample.json.
Rollout Plan
Phase 1 (engine-level):
- Implement CLI adapter and deterministic loop.
- Run scenario pack using dryrun and one live provider profile.
Phase 2 (evaluation hardening):
- Add score calibration and failure taxonomy.
- Add aggregate report for multi-run comparisons.
Phase 3 (desktop fidelity):
- Add optional desktop adapter for true upload/canvas flow parity.
Risks and Mitigations
- Simulation drift from real users. Mitigation: calibrate scenario pack from real anonymized usage motifs.
- Agent over-optimizes rubric. Mitigation: separate actor policy from evaluator policy/model.
- Flaky provider/network responses. Mitigation: retries, bounded backoff, and deterministic seeds where possible.
Acceptance Criteria
- Run all three benchmark scenarios from one JSON pack.
- Produce complete artifact bundle per scenario.
- Emit reflection record every turn.
- Generate deterministic summary for repeated seeded runs.
- Complete without manual intervention in CLI adapter mode.