Product + Technical Spec: Target User Usability Harness (v0)
Status: PRD draft (not implemented).
Objective
Build a repeatable harness that simulates a target-user usability panel to surface pain points before human testing:
- import realistic photo sets,
- provide project context and constraints,
- drive Mother + Abilities over multiple turns,
- capture explicit "open pondering" and friction observations each turn,
- take UI screenshot evidence at key junctures so simulated users can react to visible state,
- produce scored, reproducible run artifacts with clustered pain-point evidence.
This harness is not a replacement for live user interviews; it is a pre-screening signal generator for likely UX friction in common scenarios.
Screenshot Requirement (Required)
- Every turn is driven by a screenshot captured from the app state.
- The simulator chooses actions from the latest image + events, not from hidden internal state alone.
- Screenshot paths are immutable and versioned in run artifacts.
- Reflection records must include which screenshot drove the decision and what changed since last capture.
Usability Framing
The harness should answer:
- Where do users stall or issue repeated fallback instructions?
- Which abilities are difficult to discover or sequence correctly?
- Which steps cause avoidable retries, confusion, or explicit user uncertainty?
- What error or ambiguity patterns recur across personas and scenarios?
- Which interface actions are too slow or feel fragile for target workflows?
Feasibility Snapshot
Feasibility is high because the required primitives already exist:
brood-rs chat --out ... --events ... supports multi-turn interaction.
- Abilities already exist as slash commands (
/diagnose, /argue, /bridge, /swap_dna, /triforce, etc.).
events.jsonl is append-only and machine-readable.
- Native CLI +
events.jsonl contracts already support deterministic multi-step experiment runs and telemetry capture.
desktop/src/canvas_app.js already has snapshot capture helpers that write image files under runDir for intent/inference flows.
The missing piece is orchestration for persona-driven user behavior and pain extraction, not core model capability.
v0 Scope
- Adapter: CLI-first (
brood-rs chat PTY) for speed and determinism.
- Persona + scenario pack input (JSON) with explicit usability hypotheses and expected friction points.
- Agentic turn loop with explicit reflection logging.
- Pain-point capture and severity scoring as first-class output.
- Ability coverage tracking and stop conditions.
- Scorecard per scenario and aggregate usability report.
- Key-juncture screenshot capture and reaction coupling.
Proposed Pain Output Contract
Every scenario run should produce:
- A per-turn friction stream with structured tags.
- A scenario-level pain summary grouped by category.
- Actionable mitigations where confidence allows, as suggestions for backlog triage.
Non-Goals (v0)
- Pixel-perfect desktop UI automation (v0 supports coarse scripted UI gestures only).
- Perfect simulation of human behavior.
- Using reflection text as hidden reasoning. Reflection is explicit, user-visible artifact data.
Pain Taxonomy (v0)
discoverability: user does not know which ability or flow to use next.
intent_ambiguity: unclear how to translate goal into a command.
setup_friction: import/context/setup actions are too cumbersome.
error_recovery: recoveries from provider or ability failures are slow/confusing.
control_confidence: user is uncertain whether output matches goal.
speed_timing: waiting or latency creates stress.
sequence_break: user cannot compose a correct multi-step sequence.
output_quality_gap: produced results are off-spec and require multiple retries.
handoff_confusion: user is unsure how to continue after an intermediate output.
Architecture
Components:
- Scenario pack loader: validates persona/scenario config.
- Session runner: executes one scenario against one persona.
- Chat adapter: sends utterances/commands to
brood-rs chat, tails events.jsonl.
- Screenshot adapter: captures, versions, and records app screenshots for each turn.
- Policy layer: chooses next action from scenario goals + recent events + latest screenshot.
- Reflection writer: writes explicit turn reflections to
reflections.jsonl with pain tags.
- Pain extractor: normalizes and validates friction observations and aggregates them by scenario.
- Evaluator: computes success and rubric scores from artifacts/events/reflections.
Suggested code placement:
rust_engine/crates/brood-cli/src/main.rs (CLI-side orchestration hooks)
docs/target_user_harness.schema.json (config contract)
docs/target_user_harness_pain.schema.json (if formalized later for shared validator usage)
Run Artifacts
For each scenario run:
session.json: resolved persona/scenario/policy.
events.jsonl: raw engine events (from Brood run dir).
transcript.jsonl: adapter-level IO (user_input, assistant_output, command).
screenshots.jsonl: ordered screenshot events (phase, path, turn, pre_or_post).
reflections.jsonl: explicit per-turn pondering.
pain_points.jsonl: structured friction annotations with taxonomy tags + severity.
scorecard.json: rubric + pass/fail result.
summary.json: compact outcome for dashboarding.
usability_report.md: concise, human-readable pain-point review and priority list.
Across the full harness invocation:
ux_improvements.md: prioritized cross-scenario UX improvements with evidence counts and suggested success metrics.
ten_x_competitive_edge.md: candidate feature bets and problem-solving angles that could create a 10x competitive advantage.
State Machine
States:
init
capture_snapshot
start_chat
import_inputs
plan_turn
act
observe
reflect
score_checkpoint
done
error
Transitions:
init -> capture_snapshot for initial baseline screenshot.
capture_snapshot -> start_chat when scenario initializes and screenshot is saved.
start_chat -> import_inputs when PTY and events are live.
import_inputs -> capture_snapshot before turn planning begins.
capture_snapshot -> plan_turn after bootstrap imports are reflected in UI.
plan_turn -> act when next action is chosen.
act -> capture_snapshot immediately after dispatch.
capture_snapshot -> observe after screenshot write.
observe -> reflect after event delta window closes.
reflect -> capture_snapshot before next action (except terminal states).
capture_snapshot -> score_checkpoint when scenario is done or failed.
score_checkpoint -> plan_turn while stop condition is unmet.
score_checkpoint -> done on success, budget exhaustion, or max turns.
- Any state ->
error on unrecoverable adapter/engine failure.
Screenshot policy:
- Capture
pre_action and post_action in each turn where available.
- Capture mandatory junctions: boot, import completion, generation completion, failure branch, and completion.
- In CLI mode, default capture strategy is
auto: source-dir image, then live macOS window capture, then synthetic fallback.
Stop conditions:
- Success criteria satisfied.
max_turns reached.
max_cost_usd reached.
max_runtime_s reached.
- Consecutive failure threshold reached.
Action Model
Three action forms are supported:
utterance: natural language sent to Mother.
command: explicit slash command to an Ability.
ui: coarse desktop interaction (for desktop_ui adapter), such as focus, click, drag, keystroke, and wait.
- Includes Mother controls:
mother_next_proposal, mother_confirm_suggestion, mother_reject_suggestion.
Supported command.name in v0:
use
diagnose
describe
canvas_context
blend
bridge
swap_dna
argue
extract_rule
odd_one_out
triforce
recast
quality
fast
cheaper
better
Interpretation:
- Use
utterance when simulating ambiguous user intent.
- Use
command when simulating power-user behavior and explicit Ability invocation.
- Use
ui when you need visible interaction traces (e.g., clicking canvas controls, moving images) before screenshot capture.
ui actions are treated as setup/interaction steps and do not consume max_turns; max_turns gates chat/command turns.
Open Pondering Contract
Each turn must write one reflection record:
turn: integer
action_taken: normalized action record
observed_signals: key events/artifacts in that turn
what_worked: short text
what_failed: short text
open_questions: array of unresolved user-style questions
next_hypothesis: what the simulated user will try next
confidence: float 0.0..1.0
screenshot_path: screenshot used for reasoning.
screenshot_phase: pre_action or post_action
screenshot_delta: list of notable visual changes.
pain_points: array of objects
category: one of the taxonomy values
severity: 0.0..1.0
symptom: short sentence describing user pain
evidence: event/event-id references or transcript snippets
likely_cause: optional diagnosis hypothesis
mitigation_hint: suggested improvement to reduce this pain in one sentence
This keeps usability issues explicit, auditable, and easy to triage across scenarios.
Scoring (v0)
Per scenario score is weighted:
goal_progress (35%)
artifact_quality_proxy (20%)
ability_usage_quality (12%)
efficiency_cost_time (12%)
resilience (8%)
reflection_quality (5%)
friction_profile (18%)
friction_profile is derived from:
- weighted pain-point severity per scenario,
- pain recurrence on unresolved questions,
- count of distinct taxonomy categories triggered.
Pass gate:
goal_progress >= 0.7
- at least one required ability used
- no fatal adapter errors
friction_profile >= 0.6 (higher is better; pain-adjusted score)
Benchmark Scenarios (v0)
image_feature_regression_triage
- Persona: product engineer shipping in-app image features.
- Inputs: source image + failing release output + target reference.
- Goal: isolate likely failure causes and produce one improved direction.
- Required abilities:
diagnose, canvas_context, argue.
- Likely pain targets:
error_recovery, intent_ambiguity, sequence_break.
creative_direction_iteration_lane
- Persona: creative technologist / design-infra operator.
- Inputs: primary subject + two style references.
- Goal: create two distinct directions, then select one production-safe path.
- Required abilities:
swap_dna, bridge, argue.
- Likely pain targets:
discoverability, control_confidence, output_quality_gap.
multi_provider_cost_reliability_sweep
- Persona: AI agency/studio ops lead.
- Inputs: client source + campaign reference + known baseline output.
- Goal: produce a client-ready output and a cheaper fallback path.
- Required abilities:
cheaper, blend, argue.
- Likely pain targets:
speed_timing, output_quality_gap, handoff_confusion.
agent_intake_discoverability_check (secondary but intentional)
- Persona: founder/devrel owner optimizing agent discoverability.
- Inputs: screenshots of
llms.txt entrypoints, intake contract, visibility probe.
- Goal: validate what an external agent can infer and identify one doc improvement.
- Required abilities:
describe, canvas_context, extract_rule.
- Likely pain targets:
discoverability, intent_ambiguity, setup_friction.
Rollout Plan
Phase 1 (engine-level):
- Implement CLI adapter and deterministic loop.
- Add screenshot-aware transcript/transition model and dryrun policy.
- Run scenario pack using dryrun and one live provider profile.
Phase 2 (evaluation hardening):
- Add score calibration and failure taxonomy.
- Add aggregate report for multi-run comparisons and recurring pain clusters.
- Add screenshot integrity checks and screenshot-derived reaction evidence.
Phase 3 (desktop fidelity):
- Ship desktop adapter that captures live app screenshots at each transition state.
- Add deterministic desktop playback harness and comparison snapshots.
Risks and Mitigations
- Simulation drift from real users.
Mitigation: calibrate scenario pack from real anonymized usage motifs.
- Agent over-optimizes rubric.
Mitigation: separate actor policy from evaluator policy/model.
- Flaky provider/network responses.
Mitigation: retries, bounded backoff, and deterministic seeds where possible.
Acceptance Criteria
- Run all benchmark scenarios from one JSON pack.
- Produce complete artifact bundle per scenario.
- Emit structured pain record each turn, including screenshot evidence fields.
- Generate deterministic summary for repeated seeded runs.
- Aggregate recurring pain themes into a usability summary with severity ranking.
- Capture all required screenshots with stable filenames and references.
- Complete without manual intervention in desktop adapter mode.
1---2name: product-plus-technical-spec-target-user-usability-harness-23description: This harness is not a replacement for live user interviews; it is a pre-screening signal generator for likely UX friction in common scenarios.4---5# Product + Technical Spec: Target User Usability Harness (v0)67Status: PRD draft (not implemented).89## Objective10Build a repeatable harness that simulates a target-user usability panel to surface pain points before human testing:11- import realistic photo sets,12- provide project context and constraints,13- drive Mother + Abilities over multiple turns,14- capture explicit "open pondering" and friction observations each turn,15- take UI screenshot evidence at key junctures so simulated users can react to visible state,16- produce scored, reproducible run artifacts with clustered pain-point evidence.1718This harness is not a replacement for live user interviews; it is a pre-screening signal generator for likely UX friction in common scenarios.1920## Screenshot Requirement (Required)21- Every turn is driven by a screenshot captured from the app state.22- The simulator chooses actions from the latest image + events, not from hidden internal state alone.23- Screenshot paths are immutable and versioned in run artifacts.24- Reflection records must include which screenshot drove the decision and what changed since last capture.2526## Usability Framing27The harness should answer:28- Where do users stall or issue repeated fallback instructions?29- Which abilities are difficult to discover or sequence correctly?30- Which steps cause avoidable retries, confusion, or explicit user uncertainty?31- What error or ambiguity patterns recur across personas and scenarios?32- Which interface actions are too slow or feel fragile for target workflows?3334## Feasibility Snapshot35Feasibility is high because the required primitives already exist:36- `brood-rs chat --out ... --events ...` supports multi-turn interaction.37- Abilities already exist as slash commands (`/diagnose`, `/argue`, `/bridge`, `/swap_dna`, `/triforce`, etc.).38- `events.jsonl` is append-only and machine-readable.39- Native CLI + `events.jsonl` contracts already support deterministic multi-step experiment runs and telemetry capture.40- `desktop/src/canvas_app.js` already has snapshot capture helpers that write image files under `runDir` for intent/inference flows.4142The missing piece is orchestration for persona-driven user behavior and pain extraction, not core model capability.4344## v0 Scope45- Adapter: CLI-first (`brood-rs chat` PTY) for speed and determinism.46- Persona + scenario pack input (JSON) with explicit usability hypotheses and expected friction points.47- Agentic turn loop with explicit reflection logging.48- Pain-point capture and severity scoring as first-class output.49- Ability coverage tracking and stop conditions.50- Scorecard per scenario and aggregate usability report.51- Key-juncture screenshot capture and reaction coupling.5253## Proposed Pain Output Contract54Every scenario run should produce:55- A per-turn friction stream with structured tags.56- A scenario-level pain summary grouped by category.57- Actionable mitigations where confidence allows, as suggestions for backlog triage.5859## Non-Goals (v0)60- Pixel-perfect desktop UI automation (v0 supports coarse scripted UI gestures only).61- Perfect simulation of human behavior.62- Using reflection text as hidden reasoning. Reflection is explicit, user-visible artifact data.6364## Pain Taxonomy (v0)65- `discoverability`: user does not know which ability or flow to use next.66- `intent_ambiguity`: unclear how to translate goal into a command.67- `setup_friction`: import/context/setup actions are too cumbersome.68- `error_recovery`: recoveries from provider or ability failures are slow/confusing.69- `control_confidence`: user is uncertain whether output matches goal.70- `speed_timing`: waiting or latency creates stress.71- `sequence_break`: user cannot compose a correct multi-step sequence.72- `output_quality_gap`: produced results are off-spec and require multiple retries.73- `handoff_confusion`: user is unsure how to continue after an intermediate output.7475## Architecture76Components:77- Scenario pack loader: validates persona/scenario config.78- Session runner: executes one scenario against one persona.79- Chat adapter: sends utterances/commands to `brood-rs chat`, tails `events.jsonl`.80- Screenshot adapter: captures, versions, and records app screenshots for each turn.81- Policy layer: chooses next action from scenario goals + recent events + latest screenshot.82- Reflection writer: writes explicit turn reflections to `reflections.jsonl` with pain tags.83- Pain extractor: normalizes and validates friction observations and aggregates them by scenario.84- Evaluator: computes success and rubric scores from artifacts/events/reflections.8586Suggested code placement:87- `rust_engine/crates/brood-cli/src/main.rs` (CLI-side orchestration hooks)88- `docs/target_user_harness.schema.json` (config contract)89- `docs/target_user_harness_pain.schema.json` (if formalized later for shared validator usage)9091## Run Artifacts92For each scenario run:93- `session.json`: resolved persona/scenario/policy.94- `events.jsonl`: raw engine events (from Brood run dir).95- `transcript.jsonl`: adapter-level IO (`user_input`, `assistant_output`, `command`).96- `screenshots.jsonl`: ordered screenshot events (`phase`, `path`, `turn`, `pre_or_post`).97- `reflections.jsonl`: explicit per-turn pondering.98- `pain_points.jsonl`: structured friction annotations with taxonomy tags + severity.99- `scorecard.json`: rubric + pass/fail result.100- `summary.json`: compact outcome for dashboarding.101- `usability_report.md`: concise, human-readable pain-point review and priority list.102103Across the full harness invocation:104- `ux_improvements.md`: prioritized cross-scenario UX improvements with evidence counts and suggested success metrics.105- `ten_x_competitive_edge.md`: candidate feature bets and problem-solving angles that could create a 10x competitive advantage.106107## State Machine108States:1091. `init`1102. `capture_snapshot`1113. `start_chat`1124. `import_inputs`1135. `plan_turn`1146. `act`1157. `observe`1168. `reflect`1179. `score_checkpoint`11810. `done`11911. `error`120121Transitions:122- `init -> capture_snapshot` for initial baseline screenshot.123- `capture_snapshot -> start_chat` when scenario initializes and screenshot is saved.124- `start_chat -> import_inputs` when PTY and events are live.125- `import_inputs -> capture_snapshot` before turn planning begins.126- `capture_snapshot -> plan_turn` after bootstrap imports are reflected in UI.127- `plan_turn -> act` when next action is chosen.128- `act -> capture_snapshot` immediately after dispatch.129- `capture_snapshot -> observe` after screenshot write.130- `observe -> reflect` after event delta window closes.131- `reflect -> capture_snapshot` before next action (except terminal states).132- `capture_snapshot -> score_checkpoint` when scenario is done or failed.133- `score_checkpoint -> plan_turn` while stop condition is unmet.134- `score_checkpoint -> done` on success, budget exhaustion, or max turns.135- Any state -> `error` on unrecoverable adapter/engine failure.136137Screenshot policy:138- Capture `pre_action` and `post_action` in each turn where available.139- Capture mandatory junctions: boot, import completion, generation completion, failure branch, and completion.140- In CLI mode, default capture strategy is `auto`: source-dir image, then live macOS window capture, then synthetic fallback.141142Stop conditions:143- Success criteria satisfied.144- `max_turns` reached.145- `max_cost_usd` reached.146- `max_runtime_s` reached.147- Consecutive failure threshold reached.148149## Action Model150Three action forms are supported:151- `utterance`: natural language sent to Mother.152- `command`: explicit slash command to an Ability.153- `ui`: coarse desktop interaction (for `desktop_ui` adapter), such as focus, click, drag, keystroke, and wait.154 - Includes Mother controls: `mother_next_proposal`, `mother_confirm_suggestion`, `mother_reject_suggestion`.155156Supported `command.name` in v0:157- `use`158- `diagnose`159- `describe`160- `canvas_context`161- `blend`162- `bridge`163- `swap_dna`164- `argue`165- `extract_rule`166- `odd_one_out`167- `triforce`168- `recast`169- `quality`170- `fast`171- `cheaper`172- `better`173174Interpretation:175- Use `utterance` when simulating ambiguous user intent.176- Use `command` when simulating power-user behavior and explicit Ability invocation.177- Use `ui` when you need visible interaction traces (e.g., clicking canvas controls, moving images) before screenshot capture.178- `ui` actions are treated as setup/interaction steps and do not consume `max_turns`; `max_turns` gates chat/command turns.179180## Open Pondering Contract181Each turn must write one reflection record:182- `turn`: integer183- `action_taken`: normalized action record184- `observed_signals`: key events/artifacts in that turn185- `what_worked`: short text186- `what_failed`: short text187- `open_questions`: array of unresolved user-style questions188- `next_hypothesis`: what the simulated user will try next189- `confidence`: float `0.0..1.0`190- `screenshot_path`: screenshot used for reasoning.191- `screenshot_phase`: `pre_action` or `post_action`192- `screenshot_delta`: list of notable visual changes.193- `pain_points`: array of objects194 - `category`: one of the taxonomy values195 - `severity`: `0.0..1.0`196 - `symptom`: short sentence describing user pain197 - `evidence`: event/event-id references or transcript snippets198 - `likely_cause`: optional diagnosis hypothesis199- `mitigation_hint`: suggested improvement to reduce this pain in one sentence200201This keeps usability issues explicit, auditable, and easy to triage across scenarios.202203## Scoring (v0)204Per scenario score is weighted:205- `goal_progress` (35%)206- `artifact_quality_proxy` (20%)207- `ability_usage_quality` (12%)208- `efficiency_cost_time` (12%)209- `resilience` (8%)210- `reflection_quality` (5%)211- `friction_profile` (18%)212213`friction_profile` is derived from:214- weighted pain-point severity per scenario,215- pain recurrence on unresolved questions,216- count of distinct taxonomy categories triggered.217218Pass gate:219- `goal_progress >= 0.7`220- at least one required ability used221- no fatal adapter errors222- `friction_profile >= 0.6` (higher is better; pain-adjusted score)223224## Benchmark Scenarios (v0)2251. `image_feature_regression_triage`226- Persona: product engineer shipping in-app image features.227- Inputs: source image + failing release output + target reference.228- Goal: isolate likely failure causes and produce one improved direction.229- Required abilities: `diagnose`, `canvas_context`, `argue`.230- Likely pain targets: `error_recovery`, `intent_ambiguity`, `sequence_break`.2312322. `creative_direction_iteration_lane`233- Persona: creative technologist / design-infra operator.234- Inputs: primary subject + two style references.235- Goal: create two distinct directions, then select one production-safe path.236- Required abilities: `swap_dna`, `bridge`, `argue`.237- Likely pain targets: `discoverability`, `control_confidence`, `output_quality_gap`.2382393. `multi_provider_cost_reliability_sweep`240- Persona: AI agency/studio ops lead.241- Inputs: client source + campaign reference + known baseline output.242- Goal: produce a client-ready output and a cheaper fallback path.243- Required abilities: `cheaper`, `blend`, `argue`.244- Likely pain targets: `speed_timing`, `output_quality_gap`, `handoff_confusion`.2452464. `agent_intake_discoverability_check` (secondary but intentional)247- Persona: founder/devrel owner optimizing agent discoverability.248- Inputs: screenshots of `llms.txt` entrypoints, intake contract, visibility probe.249- Goal: validate what an external agent can infer and identify one doc improvement.250- Required abilities: `describe`, `canvas_context`, `extract_rule`.251- Likely pain targets: `discoverability`, `intent_ambiguity`, `setup_friction`.252253## Rollout Plan254Phase 1 (engine-level):255- Implement CLI adapter and deterministic loop.256- Add screenshot-aware transcript/transition model and dryrun policy.257- Run scenario pack using dryrun and one live provider profile.258259Phase 2 (evaluation hardening):260- Add score calibration and failure taxonomy.261- Add aggregate report for multi-run comparisons and recurring pain clusters.262- Add screenshot integrity checks and screenshot-derived reaction evidence.263264Phase 3 (desktop fidelity):265- Ship desktop adapter that captures live app screenshots at each transition state.266- Add deterministic desktop playback harness and comparison snapshots.267268## Risks and Mitigations269- Simulation drift from real users.270 Mitigation: calibrate scenario pack from real anonymized usage motifs.271- Agent over-optimizes rubric.272 Mitigation: separate actor policy from evaluator policy/model.273- Flaky provider/network responses.274 Mitigation: retries, bounded backoff, and deterministic seeds where possible.275276## Acceptance Criteria277- Run all benchmark scenarios from one JSON pack.278- Produce complete artifact bundle per scenario.279- Emit structured pain record each turn, including screenshot evidence fields.280- Generate deterministic summary for repeated seeded runs.281- Aggregate recurring pain themes into a usability summary with severity ranking.282- Capture all required screenshots with stable filenames and references.283- Complete without manual intervention in desktop adapter mode.