Desktop App (Tauri)
Supported platform: macOS only (Desktop app). There is no web app, and Windows/Linux builds are not supported yet.
Category claim:
- Promptless, reference-first AI image generation and editing desktop for developers (multi-provider + reproducible runs).
The desktop app is image-first: import images, run Abilities, and inspect results in the bottom HUD.
Core Concepts
- Run: a folder on disk (created under
~/brood_runs/) that stores inputs, artifacts, receipts, and events.jsonl.
- Unit: the currently selected image (shown on the canvas in single view).
- Views:
Single view: one image on the canvas, with a filmstrip to browse artifacts.
Multi view: tiled layout of all images in the run (used for 2-photo actions).
Basic Workflow
- Click New Run (creates a run directory and starts the engine).
- Click Import Photos or drag-drop onto the canvas (copies files into
run_dir/inputs/).
- Use Abilities (right panel) to generate edits/variants.
- Use Export to write
run_dir/export.html for a lightweight shareable viewer.
Abilities
Single-image actions (work in Single view):
Recast: reimagine the image in a different medium/context (image output).
Create Layers: split one image into compositional layers as separate artifacts.
Background: White / Background: Sweep: background replacement edits.
Crop: Square: local crop (no model call).
Variations: zero-prompt variations of the active image.
Extraction actions (work from selected source images, typically in Multi view):
Extract DNA: collapse each selected source into a draggable DNA glyph.
Soul Leech: collapse each selected source into a draggable Soul glyph.
Two-image actions (require Multi view and exactly 2 photos loaded):
Combine: blend the two images into one (/blend).
Swap DNA: structure from one + surface qualities from the other (/swap_dna). Shift-click to invert.
Bridge: synthesize the aesthetic midpoint between two references (/bridge).
Notes:
- Some actions auto-switch the Image Model (e.g. 2-photo actions prefer
gemini-3-pro-image-preview). The agent portraits update to match.
- After a 2-photo action completes, Brood switches back to
Single view showing the output-only image. Use Multi view to return to the tiled layout.
Effect Tokens (DNA / Soul)
- Extraction visuals run on a dedicated Pixi overlay (
#effects-canvas) and are clipped exactly to the source tile bounds.
- When extraction completes, the source tile is tokenized: the source image box is removed from normal canvas interaction and replaced by a floating draggable glyph.
- The token lifecycle is explicit:
extracting -> ready -> dragging -> drop_preview -> applying -> consumed.
- Drag/drop rules:
- Valid drop target must be a different image than the source.
- Valid targets get a strong hover highlight.
- Drop plays a sink/absorb animation, then dispatches apply exactly once.
- Invalid drop cancels without dispatching apply.
- On successful apply:
- Target is edited in place (DNA/Soul transfer).
- The token is consumed.
- The extracted source image is removed from the canvas.
- On failed apply:
- Token recovers to a draggable
ready state (no stuck applying lock).
Canvas Context + Mother Reference Counts
- Tokenized source images are excluded from visible-canvas counts and selection logic.
- Effect glyphs are rendered only on the Pixi overlay (not the base work canvas), so realtime/intent snapshots do not include DNA/Soul glyphs.
- Practical result: after one extraction from
n images, Mother context and reference counts use n - 1 visible images until the effect is applied.
HUD + Tools
- The HUD prints
UNIT / DESC / SEL / GEN for the active image.
- The HUD keybar (buttons
1-9) activates canvas tools/actions. Common hotkeys:
L lasso
F fit-to-view
Esc clear selection / close panels
Mother Proposal + Gemini Context (v2)
Brood now sends two compact context packets that preserve user-selected proposal flow while making model behavior more aware of what happened on canvas.
brood.mother.proposal_context.v1 (during intent/proposal inference):
- Added to
mother_intent_infer-*.json as proposal_context.
- Encodes soft priors only: interaction focus, geometry hints, and compact spatial relations.
- Does not override explicit proposal lock semantics (
active_id, selected_ids, chosen proposal mode).
brood.gemini.context_packet.v2 (during image generation):
- Added to
mother_generate-*.json as gemini_context_packet.
- Includes
proposal_lock, ranked image slots, compact relations, and a capped must_not list.
- Includes a tiny
geometry_trace per image: cx, cy, relative_scale, iou_to_primary.
Scoring Math (high level)
Per-image interaction and geometry are normalized and combined into a soft weighting prior.
- Saturating interaction transform:
sat(c, k) = min(1, ln(1 + c) / ln(1 + k))
- Interaction base:
E = 0.35*sat(move,8) + 0.35*sat(resize,4) + 0.25*sat(selection,8) + 0.05*sat(action_grid,4)
- Recency + staleness:
- decay
exp(-age_ms / 90000)
- hard stale cutoff on transform activity:
age_transform_ms > 600000 (10 min) => interaction contribution 0
- Geometry score:
- size term uses
sqrt(area_ratio) normalization
- centrality term uses distance to
(0.5, 0.5)
- combined as
0.8*size + 0.2*centrality, then normalized
- Combined score (soft prior):
- intent proposal context uses:
(1 + 0.8*focus_score) * (1 + 0.5*geometry_score) * (1 + selected_bonus + active_bonus)
- Gemini generation context uses role priors and single-target guardrails.
Guardrails and compactness
- Single-target clamp logic is applied only when exactly one target exists.
must_not is deduped and capped to exactly 6 constraints.
- Relations are compact and confidence-gated (
OVERLAP / directional ADJACENT) to reduce prompt noise.
Debugging and verification
- Enable Gemini wire debug:
BROOD_DEBUG_GEMINI_WIRE=1 npm --prefix desktop run tauri dev
- Inspect the latest run:
~/brood_runs/run-*/mother_intent_infer-*.json -> proposal_context
~/brood_runs/run-*/_raw_provider_outputs/gemini-send-message-*.json -> exact Gemini chat.send_message payload parts
~/brood_runs/run-*/_raw_provider_outputs/gemini-receipt-*.json -> provider receipt copy
~/brood_runs/run-*/mother_generate-*.json -> generation payload containing gemini_context_packet
Files Written To The Run
run_dir/inputs/: imported photos
run_dir/receipt-*.json: generation/edit receipts
run_dir/events.jsonl: event stream consumed by the desktop UI
run_dir/visual_prompt.json: serialized canvas marks/layout (see docs/visual_prompting_v0.md)
1---2name: desktop-app-tauri-23description: Brood now sends two compact context packets that preserve user-selected proposal flow while making model behavior more aware of what happened on canvas.4---5# Desktop App (Tauri)67Supported platform: **macOS only** (Desktop app). There is no web app, and Windows/Linux builds are not supported yet.89Category claim:10- Promptless, reference-first AI image generation and editing desktop for developers (multi-provider + reproducible runs).1112The desktop app is image-first: import images, run Abilities, and inspect results in the bottom HUD.1314## Core Concepts15- **Run**: a folder on disk (created under `~/brood_runs/`) that stores inputs, artifacts, receipts, and `events.jsonl`.16- **Unit**: the currently selected image (shown on the canvas in single view).17- **Views**:18 - `Single view`: one image on the canvas, with a filmstrip to browse artifacts.19 - `Multi view`: tiled layout of all images in the run (used for 2-photo actions).2021## Basic Workflow221. Click **New Run** (creates a run directory and starts the engine).232. Click **Import Photos** or drag-drop onto the canvas (copies files into `run_dir/inputs/`).243. Use **Abilities** (right panel) to generate edits/variants.254. Use **Export** to write `run_dir/export.html` for a lightweight shareable viewer.2627## Abilities2829Single-image actions (work in `Single view`):30- `Recast`: reimagine the image in a different medium/context (image output).31- `Create Layers`: split one image into compositional layers as separate artifacts.32- `Background: White` / `Background: Sweep`: background replacement edits.33- `Crop: Square`: local crop (no model call).34- `Variations`: zero-prompt variations of the active image.3536Extraction actions (work from selected source images, typically in `Multi view`):37- `Extract DNA`: collapse each selected source into a draggable DNA glyph.38- `Soul Leech`: collapse each selected source into a draggable Soul glyph.3940Two-image actions (require `Multi view` and **exactly 2** photos loaded):41- `Combine`: blend the two images into one (`/blend`).42- `Swap DNA`: structure from one + surface qualities from the other (`/swap_dna`). Shift-click to invert.43- `Bridge`: synthesize the aesthetic midpoint between two references (`/bridge`).4445Notes:46- Some actions auto-switch the **Image Model** (e.g. 2-photo actions prefer `gemini-3-pro-image-preview`). The agent portraits update to match.47- After a 2-photo action completes, Brood switches back to `Single view` showing the output-only image. Use `Multi view` to return to the tiled layout.4849## Effect Tokens (DNA / Soul)50- Extraction visuals run on a dedicated Pixi overlay (`#effects-canvas`) and are clipped exactly to the source tile bounds.51- When extraction completes, the source tile is tokenized: the source image box is removed from normal canvas interaction and replaced by a floating draggable glyph.52- The token lifecycle is explicit: `extracting -> ready -> dragging -> drop_preview -> applying -> consumed`.53- Drag/drop rules:54 - Valid drop target must be a different image than the source.55 - Valid targets get a strong hover highlight.56 - Drop plays a sink/absorb animation, then dispatches apply exactly once.57 - Invalid drop cancels without dispatching apply.58- On successful apply:59 - Target is edited in place (DNA/Soul transfer).60 - The token is consumed.61 - The extracted source image is removed from the canvas.62- On failed apply:63 - Token recovers to a draggable `ready` state (no stuck `applying` lock).6465### Canvas Context + Mother Reference Counts66- Tokenized source images are excluded from visible-canvas counts and selection logic.67- Effect glyphs are rendered only on the Pixi overlay (not the base work canvas), so realtime/intent snapshots do not include DNA/Soul glyphs.68- Practical result: after one extraction from `n` images, Mother context and reference counts use `n - 1` visible images until the effect is applied.6970## HUD + Tools71- The HUD prints `UNIT / DESC / SEL / GEN` for the active image.72- The HUD keybar (buttons `1`-`9`) activates canvas tools/actions. Common hotkeys:73 - `L` lasso74 - `F` fit-to-view75 - `Esc` clear selection / close panels7677## Mother Proposal + Gemini Context (v2)78Brood now sends two compact context packets that preserve user-selected proposal flow while making model behavior more aware of what happened on canvas.7980- `brood.mother.proposal_context.v1` (during intent/proposal inference):81 - Added to `mother_intent_infer-*.json` as `proposal_context`.82 - Encodes soft priors only: interaction focus, geometry hints, and compact spatial relations.83 - Does not override explicit proposal lock semantics (`active_id`, `selected_ids`, chosen proposal mode).84- `brood.gemini.context_packet.v2` (during image generation):85 - Added to `mother_generate-*.json` as `gemini_context_packet`.86 - Includes `proposal_lock`, ranked image slots, compact relations, and a capped `must_not` list.87 - Includes a tiny `geometry_trace` per image: `cx`, `cy`, `relative_scale`, `iou_to_primary`.8889### Scoring Math (high level)90Per-image interaction and geometry are normalized and combined into a soft weighting prior.9192- Saturating interaction transform:93 - `sat(c, k) = min(1, ln(1 + c) / ln(1 + k))`94- Interaction base:95 - `E = 0.35*sat(move,8) + 0.35*sat(resize,4) + 0.25*sat(selection,8) + 0.05*sat(action_grid,4)`96- Recency + staleness:97 - decay `exp(-age_ms / 90000)`98 - hard stale cutoff on transform activity: `age_transform_ms > 600000` (10 min) => interaction contribution `0`99- Geometry score:100 - size term uses `sqrt(area_ratio)` normalization101 - centrality term uses distance to `(0.5, 0.5)`102 - combined as `0.8*size + 0.2*centrality`, then normalized103- Combined score (soft prior):104 - intent proposal context uses:105 - `(1 + 0.8*focus_score) * (1 + 0.5*geometry_score) * (1 + selected_bonus + active_bonus)`106 - Gemini generation context uses role priors and single-target guardrails.107108### Guardrails and compactness109- Single-target clamp logic is applied only when exactly one target exists.110- `must_not` is deduped and capped to exactly 6 constraints.111- Relations are compact and confidence-gated (`OVERLAP` / directional `ADJACENT`) to reduce prompt noise.112113### Debugging and verification114- Enable Gemini wire debug:115 - `BROOD_DEBUG_GEMINI_WIRE=1 npm --prefix desktop run tauri dev`116- Inspect the latest run:117 - `~/brood_runs/run-*/mother_intent_infer-*.json` -> `proposal_context`118 - `~/brood_runs/run-*/_raw_provider_outputs/gemini-send-message-*.json` -> exact Gemini `chat.send_message` payload parts119 - `~/brood_runs/run-*/_raw_provider_outputs/gemini-receipt-*.json` -> provider receipt copy120 - `~/brood_runs/run-*/mother_generate-*.json` -> generation payload containing `gemini_context_packet`121122## Files Written To The Run123- `run_dir/inputs/`: imported photos124- `run_dir/receipt-*.json`: generation/edit receipts125- `run_dir/events.jsonl`: event stream consumed by the desktop UI126- `run_dir/visual_prompt.json`: serialized canvas marks/layout (see `docs/visual_prompting_v0.md`)