PPT Master Skill
AI-driven multi-format SVG content generation system. Converts source documents into high-quality SVG pages through multi-role collaboration and exports to PPTX.
Core Pipeline: Source Document → Create Project → [Template] → Strategist → [Image_Generator] → Executor Live Preview → Quality Check → Post-processing → Export
[!CAUTION]
🚨 Global Execution Discipline (MANDATORY)
This workflow is a strict serial pipeline. The following rules have the highest priority — violating any one of them constitutes execution failure:
- SERIAL EXECUTION — Steps MUST be executed in order; the output of each step is the input for the next. Non-BLOCKING adjacent steps may proceed continuously once prerequisites are met, without waiting for the user to say "continue"
- BLOCKING = HARD STOP — Steps marked ⛔ BLOCKING require a full stop; the AI MUST wait for an explicit user response before proceeding and MUST NOT make any decisions on behalf of the user
- NO CROSS-PHASE BUNDLING — Cross-phase bundling is FORBIDDEN. (Note: the Strategist confirmation stage in Step 4 is ⛔ BLOCKING — the AI MUST present recommendations and wait for explicit user confirmation before proceeding. Once the user confirms, all subsequent non-BLOCKING steps — design spec output, SVG generation, speaker notes, and post-processing — may proceed automatically without further user confirmation)
- GATE BEFORE ENTRY — Each Step has prerequisites (🚧 GATE) listed at the top; these MUST be verified before starting that Step
- NO SPECULATIVE EXECUTION — "Pre-preparing" content for subsequent Steps is FORBIDDEN (e.g., writing SVG code during the Strategist phase)
- NO SUB-AGENT SVG GENERATION — Executor Step 6 SVG generation is context-dependent and MUST be completed by the current main agent end-to-end. Delegating page SVG generation to sub-agents is FORBIDDEN
- SEQUENTIAL PAGE GENERATION ONLY — In Executor Step 6, after the global design context is confirmed, SVG pages MUST be generated sequentially page by page in one continuous pass. Grouped page batches (for example, 5 pages at a time) are FORBIDDEN
- SPEC_LOCK RE-READ PER PAGE — Before generating each SVG page, Executor MUST
read_file <project_path>/spec_lock.md. All colors / fonts / icons / images MUST come from this file — no values from memory or invented on the fly. Executor MUST also look up the current page'spage_rhythm(anchor/dense/breathing),page_layouts(which template SVG to inherit, if any), andpage_charts(which chart template to adapt, if any). Empty / absent entries are intentional Strategist signals — see executor-base.md §2.1. This rule exists to resist context-compression drift on long decks and to break the uniform "every page is a card grid" default- SVG MUST BE HAND-WRITTEN, NOT SCRIPT-GENERATED — Every SVG page is written by the main agent directly, one page at a time (see rules 6 and 7). Writing or running a Python / Node / shell script that produces the SVG files in batch — looping over pages, templating from data, or emitting them via a generator — is FORBIDDEN, including under "save tokens", "quick draft", or "user is in a hurry" pretexts. The script-generation path was tried on a feature branch and abandoned: cross-page visual consistency depends on per-page authoring with full upstream context, which a generator script cannot reproduce
- FOLLOW DETERMINISTIC ROUTING RULES — Do not add blocking routing questions when this skill defines a route. If the user request violates a route precondition, state the required prerequisite and stop that route instead of asking the user to choose around the rule. Ordinary finite options, stylistic preferences, and recoverable details are surfaced with a recommended value plus alternatives at the next existing confirmation gate.
[!IMPORTANT]
🌐 Language & Communication Rule
- Response language: match the user's input and source materials. Explicit user override (e.g., "请用英文回答") takes precedence.
- User-facing option labels: when presenting confirmations, brief proposals, choices, or finite option sets, use the user's language for labels and explanations. English enum IDs / file fields may appear in parentheses for precision, but never rely on English-only labels such as
deck,layout,mirror, orfidelitywithout a localized explanation.- Template format:
design_spec.mdMUST follow its original English template structure (section headings, field names) regardless of conversation language. Content values may be in the user's language.
[!IMPORTANT]
🔌 Compatibility With Generic Coding Skills
ppt-masteris a repository-specific workflow, not a general application scaffold- Do NOT create
.worktrees/,tests/, branch workflows, or generic engineering structure by default- On conflict with a generic coding skill, follow this skill unless the user explicitly says otherwise
Rule Strength Labels
| Label | Meaning |
|---|---|
MUST |
Required behavior; violation is workflow failure |
MUST NOT |
Forbidden behavior |
DEFAULT |
Used when the user has not specified otherwise |
OPTIONAL |
Run only when explicitly triggered or when the route says so |
FALLBACK |
Recovery path after the primary path fails |
GATE |
Required checkpoint before entering the next step |
Cross-Cutting Authorities
| Concern | Authority | Contract |
|---|---|---|
| Main pipeline sequencing | This SKILL.md |
Owns Step 1-7 order, gates, role switching, and mandatory commands |
| Route selection | workflows/routing.md |
Owns deterministic route choice before the main pipeline or a standalone workflow |
| Workflow registry | workflows/index.md |
Owns standalone workflow trigger/precondition/output inventory |
| Artifact ownership | references/artifact-ownership.md |
Owns fact channels, source/derived artifact boundaries, and regeneration rules |
| Failure recovery | workflows/failure-recovery.md |
Owns stop/continue decisions for common failures |
| Confirm UI details | scripts/docs/confirm_ui.md |
Owns schema, launcher behavior, port strategy, and chat fallback details |
Main Pipeline Scripts
| Script | Purpose |
|---|---|
${SKILL_DIR}/scripts/source_to_md.py |
Unified source-to-Markdown dispatcher — default Step 1 entry for explicit file(s) or URL(s) |
${SKILL_DIR}/scripts/pptx_intake.py |
Standard PPTX intake enrichment — canvas / identity / slide geometry / tables / native chart data |
${SKILL_DIR}/scripts/project_manager.py |
Project init / validate / manage |
${SKILL_DIR}/scripts/icon_sync.py |
Copy chosen library icons into <project>/icons/ at selection time; missing names reported + non-zero (re-pick gate) |
${SKILL_DIR}/scripts/analyze_images.py |
Image analysis |
${SKILL_DIR}/scripts/latex_render.py |
LaTeX formula rendering (manifest-driven PNG assets) |
${SKILL_DIR}/scripts/image_gen.py |
AI image generation (multi-provider) |
${SKILL_DIR}/scripts/slice_images.py |
Slice one AI illustration sheet into individual spot-illustration elements |
${SKILL_DIR}/scripts/svg_quality_checker.py |
SVG quality check |
${SKILL_DIR}/scripts/total_md_split.py |
Speaker notes splitting |
${SKILL_DIR}/scripts/finalize_svg.py |
SVG post-processing (unified entry) |
${SKILL_DIR}/scripts/svg_to_pptx.py |
Export to PPTX |
${SKILL_DIR}/scripts/native_enhance_pptx.py |
Existing PPTX enhancement project init / validation / direct OOXML patch export |
${SKILL_DIR}/scripts/native_narration_pptx.py |
Backward-compatible entrypoint for existing PPTX notes / narration enhancement |
${SKILL_DIR}/scripts/update_spec.py |
Propagate a spec_lock.md color / font_family change across all generated SVGs |
For complete tool documentation, see ${SKILL_DIR}/scripts/README.md.
Windows note: if a
python3 ...command fails (common on python.org installs, which providepython.exebut notpython3.exe), rerun the same command withpythoninstead.
Template Index
| Index | Path | Purpose |
|---|---|---|
| Layout templates | ${SKILL_DIR}/templates/layouts/layouts_index.json |
Query available page layout templates |
| Brand presets | ${SKILL_DIR}/templates/brands/brands_index.json |
Query available brand identity presets (color / typography / logo / voice) |
| Visualization templates | ${SKILL_DIR}/templates/charts/charts_index.json |
Query available visualization SVG templates (charts, infographics, diagrams, frameworks) |
| Icon library | ${SKILL_DIR}/templates/icons/ |
See ${SKILL_DIR}/templates/icons/README.md; search icons on demand with ls templates/icons/<library>/ | grep <keyword> |
Standalone Workflows
Route authority: Use workflows/routing.md before entering the main pipeline or any standalone workflow.
Registry: Use workflows/index.md for the complete workflow list, triggers, preconditions, exclusions, outputs, and blocking points.
PPTX Route Boundary
| User intent | Route |
|---|---|
| Raw PPTX template plus new material/topic, generate a PPTX | template-fill-pptx |
| Existing PPTX, preserve page count/order and slide wording 1:1, improve layout | beautify-pptx |
| Existing PPTX as source material, rethink outline or change page count/order | Main pipeline via source_to_md.py plus PPTX intake |
| Build a reusable template package from a PPTX/design reference | create-template, then return with the generated template directory path |
| Finished PPTX, keep content/layout stable and add notes/audio/timing/transitions | native-enhance-pptx |
MUST: Raw .pptx template plus "generate PPTX" routes to template-fill-pptx by default. The SVG generation route consumes only an explicit template directory path that already contains a valid template design_spec.md.
MUST: Beautify is strictly 1:1. Any split, merge, drop, reorder, or page-count change routes to the main pipeline.
FALLBACK: Ambiguous requests such as "make this PPT more professional" require exactly one discriminator question: preserve original page count/order and slide wording, or treat the deck as source material and restructure it?
Workflow
Step 1: Source Content Processing
🚧 GATE: User has provided source material (PDF / DOCX / EPUB / URL / Markdown file / text description / conversation content — any form is acceptable).
No source content? When the user supplies only a topic name or requirements without any file or substantive description, run the
topic-researchworkflow first, then return here with its products as input.
When the user provides non-Markdown content, convert immediately through the unified dispatcher. It preserves the backend converters' existing behavior, routes by source type, and writes the standard Markdown plus conversion profile.
| User Provides | Action |
|---|---|
| PDF / DOCX / Office document / XLSX / XLSM / PPTX / EPUB / HTML / LaTeX / RST / web URL | python3 ${SKILL_DIR}/scripts/source_to_md.py <file_or_URL_or_dir> [<file_or_URL_or_dir> ...] |
| CSV / TSV | Read directly as plain-text table source |
| Markdown | Read directly |
For PPTX sources, Step 1 converts the deck to Markdown content; after Step 2
import-sources, standard PPTX intake is also written to <project>/analysis/.
Use source_to_md.py -t <type> only when extension detection is ambiguous.
Default local conversion writes Markdown/profile outputs beside each source file.
Use -o only when a specific output file/directory is required; with multiple
inputs or directory inputs, -o is an output directory. Backend converter details are documented in
scripts/docs/conversion.md.
Office vector assets (EMF/WMF) from DOCX/PPTX sources: Source conversion extracts embedded Office vector images (.emf/.wmf) alongside bitmap images when the source format exposes them. After
import-sources, these land inimages/together withimage_manifest.jsonand are first-class assets in §VIII Image Resource List.Do NOT convert EMF/WMF to PNG. The PPT Master pipeline preserves them as external references (
finalize_svg.pyskips them) andsvg_to_pptx.pyembeds them as PPTX-native media viaimage/x-emf/image/x-wmfMIME — PowerPoint renders them at full vector fidelity. Converting via LibreOffice/Inkscape introduces CJK font substitution drift and rasterization loss; the original EMF/WMF is always higher fidelity than the converted PNG.Browser-based live preview cannot render EMF (will show blank) — this is expected; the PPTX output is the source of truth.
✅ Checkpoint — Confirm source content is ready, proceed to Step 2.
Step 2: Project Initialization
🚧 GATE: Step 1 complete; source content is ready (Markdown file, user-provided text, or requirements described in conversation are all valid).
python3 ${SKILL_DIR}/scripts/project_manager.py init <project_name> --format <format>
Format options must be named with concrete dimensions. Default: ppt169 = 1280x720, viewBox="0 0 1280 720". Other examples: ppt43 = 1024x768, story = 1080x1920, banner = 1920x1080. For the full format list, see references/canvas-formats.md.
Import source content (choose based on the situation):
| Situation | Action |
|---|---|
| Has source files (PDF/MD/etc.) | python3 ${SKILL_DIR}/scripts/project_manager.py import-sources <project_path> <source_files_or_dirs...> --move |
| User provided text directly in conversation | No import needed — content is already in conversation context; subsequent steps can reference it directly |
For PPTX sources, import-sources automatically runs the standard intake enrichment:
python3 ${SKILL_DIR}/scripts/pptx_intake.py <project_path>/sources/<source.pptx> -o <project_path>/analysis
For each PPTX it writes <stem>.identity.json (canvas, theme palette/fonts, observed usage) and <stem>.slide_library.json (text slots, geometry, native tables, native chart caches), and merges that deck's Strategist-facing digest into the single multi-deck index analysis/source_profile.json (decks[], one self-contained entry per source deck, with prefixed artifact pointers). In the main generation path these are source facts and recommendation candidates, not replica constraints; beautify and template-fill workflows decide separately which fields become locked constraints.
Multi-deck: several PPTX files may be imported into one main-pipeline project — each gets its own <stem>.* artifacts and a deck entry in source_profile.json. source_profile.json stays the single must-read index (one entry for a one-deck project, several for a combined-source project). Stems must be distinct; re-importing the same stem replaces that deck's entry. The beautify / template-fill workflows remain single-deck (1:1 to one chosen source deck) and read that deck's <stem>.* artifacts.
⚠️ MUST use
--move(not copy): all source files — Step 1's generated Markdown, original PDFs / MDs / images — go intosources/viaimport-sources --move. If Step 1 wrote Markdown beside the original sources, pass that source path/directory once. If Step 1 used-oto write Markdown elsewhere, pass both the original source path(s)/directory and the Markdown output path(s)/directory. After execution they no longer exist at the original location. Intermediate artifacts (e.g.,_files/) are handled automatically.
✅ Checkpoint — Confirm project structure created successfully, sources/ contains all source files, converted materials are ready. Proceed to Step 3.
Step 3: Template Option
🚧 GATE: Step 2 complete; project directory structure is ready.
Default — free design. Proceed directly to Step 4. Do NOT query any *_index.json unless triggered. Do NOT ask the user. Do NOT proactively suggest, hint at, or fuzzy-match any template based on content, slug-like words, or vague style descriptions.
Hard boundary — raw PPTX template references are not Step 3 templates. PPTX-as-source remains valid in Step 1 / Step 2, and raw PPTX template + generated PPTX routes to template-fill. But if the user wants the SVG/template-based generation route from that PPTX, stop before Step 3. The user must first run workflows/create-template.md, then return with the generated template directory path. Step 3 only consumes an explicit template directory that already contains design_spec.md with kind: brand / kind: layout / kind: deck.
Do not reinterpret this boundary as 1:1 redesign or free SVG generation. Use template-fill for raw PPTX template + generated PPTX requests; use beautify only when the source deck's page count, order, and wording are preserved.
Template flow triggers ONLY on explicit directory paths supplied by the user in their initial message. The trigger rule is mechanical, not interpretive:
| User input contains | Step 3 action |
|---|---|
One or more explicit template directory paths (each resolves to a directory containing design_spec.md with kind: brand / kind: layout / kind: deck in its YAML frontmatter) |
Read each spec's kind, dispatch per the kind matrix below, fuse if multiple |
| Anything else — bare template names ("用 academic_defense"), style descriptions ("麦肯锡风格"), brand mentions ("招商银行风格"), vague intent ("想用个模板"), or silence | Skip Step 3, free design |
There is no slug matching, no name lookup, no fuzzy resolution. A name without a path does not trigger — the user must give a path the AI can cd into.
Style descriptions ("麦肯锡风格" / "Keynote 风" / "极简风" / etc.) never trigger Step 3. They flow into the Strategist confirmation stage as a style brief (color / typography / tone in fields e–g).
Bare names ("academic_defense", "招商银行", "anthropic") do NOT trigger Step 3 even if a matching directory exists in the library. The user must give a path. AI must not "helpfully" resolve a name to a path.
"What templates exist?" is out-of-band Q&A — answer by listing entries from
brands_index.json/layouts_index.json/decks_index.jsontogether with their paths. Listing alone does not advance the pipeline; the user must send a path back to trigger Step 3.
To create a new layout or deck, read
workflows/create-template.md. To create a new brand, readworkflows/create-brand.md.
Three template kinds
The architecture has three independent reference bundles. Full schema in docs/zh/templates-architecture.md. Summary:
| Kind | Physical dir | Contains | Frontmatter |
|---|---|---|---|
| brand | templates/brands/<id>/ |
identity-only segment: color / typography / logo / voice / icon style | kind: brand |
| layout | templates/layouts/<id>/ |
structure-only segment: canvas / page structure / page types / SVG roster | kind: layout |
| deck | templates/decks/<id>/ |
full replica: identity + structure + middle (template overview) segments | kind: deck |
Segment ownership (governs fusion override priority):
| Segment | Sections | Owner kind on fusion |
|---|---|---|
| Identity | Color Scheme / Typography / Logo / Voice & Tone / Icon Style | brand |
| Structure | Canvas / Page Structure / Page Types / SVG Roster | layout |
| Middle | Template Overview (use cases / design intent) | deck (no other kind writes this) |
Single-path dispatch
User path's kind |
Step 3 action |
|---|---|
kind: brand |
design_spec.md + non-image assets → <project>/templates/; logo / illustration / icon bitmaps → <project>/images/. Strategist locks identity segment as truth; structure stays free. |
kind: layout |
design_spec.md + SVG roster → <project>/templates/; any bitmap assets → <project>/images/. Strategist locks structure; identity decided in Strategist confirmation stage e–g. |
kind: deck |
design_spec.md + template SVGs → <project>/templates/; logos / backgrounds / other bitmaps → <project>/images/. Strategist locks all segments; Strategist confirmation stage narrows to deck-content fields (audience / page count / outline / tone tweaks). |
TEMPLATE_DIR=<user-supplied path>
# Bitmaps join the project's single runtime image pool (images/, referenced as
# ../images/); the spec + template SVGs + other non-image assets stay in
# templates/ as design reference the Strategist/Executor read but never render.
cp -r ${TEMPLATE_DIR}/* <project_path>/templates/
find <project_path>/templates -type f \( -iname '*.png' -o -iname '*.jpg' -o -iname '*.jpeg' -o -iname '*.gif' -o -iname '*.webp' -o -iname '*.bmp' \) -exec mv {} <project_path>/images/ \;
The same split applies to all three kinds — bitmaps always land in images/, the rest in templates/. The spec's kind field tells Strategist how to read the templates/ side; downstream code doesn't distinguish. (Template SVGs in templates/ are reference material only — the rendered pages live in svg_output/ and reference images via ../images/.)
Multi-path fusion
When the user gives two or more paths of different kinds, Step 3 fuses them into a single <project>/templates/design_spec.md. Default granularity is segment-level integer replacement — entire identity / structure / middle segments are taken from the highest-priority source for that segment, no implicit field-level mixing.
Override priority by segment:
| Combination | Identity from | Structure from | Middle from |
|---|---|---|---|
| brand only | brand | (free design) | (none) |
| layout only | (free design) | layout | (none) |
| deck only | deck | deck | deck |
| brand + layout | brand | layout | (none) |
| brand + deck | brand (overrides deck) | deck | deck |
| layout + deck | deck | layout (overrides deck) | deck |
| brand + layout + deck | brand | layout | deck |
Field-level micro-adjustment (e.g. "use anthropic brand but primary changed to #FF0000") is not part of Step 3 fusion — it flows into Strategist confirmation stage e–g as a normal user request.
Same-kind multiple paths — conflict resolution
When the user gives two paths of the same kind (e.g. brands/anthropic + brands/google), Step 3 surfaces a conflict prompt before fusing — like resolving a git merge conflict:
AI: 你给了两个 brand,检测到段级冲突:
- Color Scheme(Anthropic 橙红 vs Google 多色)
- Typography(Styrene/AnthropicSans vs GoogleSans/Roboto)
- Logo(Anthropic 标 vs Google 标)
- Voice & Tone(restrained vs friendly)
- Icon Style(stroke vs filled)
要 (a) 全部按 Anthropic / (b) 全部按 Google / (c) 逐段挑?
Rules:
- Default: no implicit ordering — every cross-source segment difference is reported as a conflict
- Only when the user picks
(c)does AI walk through each segment one by one - Field-level conflicts are out of scope — segment-level only
- Three or more same-kind paths are not supported — ask the user to converge to at most two
Fused spec provenance
When fusion happens (any multi-path case), the resulting <project>/templates/design_spec.md carries a provenance block immediately under its H1:
> **Fused from:**
> - deck: `templates/decks/招商银行/` (base)
> - brand: `templates/brands/anthropic/` (identity override)
> - layout: `templates/layouts/academic_defense/` (structure override)
> - conflicts resolved: Color Scheme from anthropic(user picked a)
Single-path Step 3 does not add provenance (the source is self-evident from the copied files).
✅ Checkpoint — Default path proceeds to Step 4 without user interaction. If the user supplied one or more explicit template paths, those have been dispatched (or fused) into <project_path>/templates/ before advancing.
Step 4: Strategist Phase (MANDATORY — cannot be skipped)
🚧 GATE: Step 3 complete; default free-design path taken, or (if triggered) template files copied into the project.
First, read the role definition:
Read references/strategist.md
⚠️ Mandatory gate: before writing
design_spec.md, Strategist MUSTread_file templates/design_spec_reference.mdand follow its full I–XI section structure. Seestrategist.mdSection 1.
Artifact ownership: fact-channel and source/derived artifact boundaries are defined in references/artifact-ownership.md. This Step uses those ownership rules; it does not redefine them.
<project_path>/analysis/ is the project's intermediate-analysis folder: the canonical home for machine-extracted source/asset facts — the PPTX intake bundle (source_profile.json index + per-deck <stem>.identity.json / <stem>.slide_library.json) and image_analysis.csv. It holds facts, not design contracts — design_spec.md / spec_lock.md stay at the project root. The MUST-read contract covers only the compact structured data files (.json / .csv); other artifacts that may live under analysis/ (e.g. a beautify source_svg_import/ vector reference package) are NOT bulk-read — they are read selectively only when a specific workflow step calls for them. Before the Strategist confirmation stage, Strategist MUST read the auto-extracted fact files already in analysis/ — currently source_profile.json (PPTX intake), when present. This file is the multi-deck index: read it once for the decks[] digests (canvas / chart / table entries per source deck), then open a specific deck's <stem>.identity.json / <stem>.slide_library.json only if you need its full raw facts. Use these entries as factual source context (format default + content facts); when several decks are present, synthesize across all of them. The source's palette / typography / visual identity are a reference, not a constraint: the main pipeline may inherit them where they fit the content and the confirmed style, or design fresh where they don't — the Strategist's judgment, never an obligation to either keep or discard. (Template-fill preserves the native source design by editing cloned slides directly; beautify defaults to the source identity but still follows the confirmed values; the main pipeline treats source identity as reference only and defaults to fresh design.) (image_analysis.csv lands later, at the image-analysis step below, and is the authoritative regenerated image-fact view there — re-derived from the live images/ folder, not a durable store.)
Channel ownership — read each fact once from its owning channel. In the main pipeline the content contract is the content-type files in sources/ — primarily <stem>.md, but also any user-supplied content the import archived there: .md / .markdown / .txt / .csv / .tsv / .json / .jsonl / .yaml / .yml (a metrics.json or data.csv may carry core content — judge by what the file holds). Text, tables, and chart data values come from these (ppt_to_md now transcribes native chart data into Markdown tables). Do NOT read pipeline sidecars in sources/ as content: *.conversion_profile.json (conversion audit) and *_files/image_manifest.json (asset index) are process metadata — open them only to audit a conversion or resolve assets, never as slide content. Converted-source originals archived in sources/ (.pdf / .pptx / .docx / .xlsx / .html / .epub / .tex / .rst / .ipynb / .typ, etc.) are read via their converted <stem>.md, not scanned directly in the main pipeline. The analysis/ chart / table entries are a structural digest for outline decisions (which slides carried charts, type, series names) — not a second copy of the values; do NOT also pull chart values from <stem>.slide_library.json in the main pipeline. The <stem>.slide_library.json full structured data is owned by the direct-PPTX workflows: template-fill uses it as the native fill contract; beautify uses it for native chart / table data while keeping slide text from the Markdown.
Strategist confirmation stage (full template: templates/design_spec_reference.md):
⛔ BLOCKING: present the Strategist confirmation stage and wait for explicit user confirmation or modification before outputting Design Specification & Content Outline. This is the single core confirmation gate — once the final confirmation lands, all subsequent steps proceed automatically. The default Confirm UI delivers the gate in three stages (direction → design system → images / execution; see below); the chat fallback mirrors the same staged order.
- Canvas format
- Page count range
- Target audience
- Style objective
- Color scheme
- Icon usage approach
- Typography plan, including formula rendering policy
- Image usage approach
Confirm UI Auto-Launch (Mandatory — default visual confirmation surface): by default the Strategist confirmation stage is presented through an interactive local page in three stages within one browser session — Stage 1 confirms the direction anchors; the AI then re-derives the design-system layer from the user's actual anchors; Stage 2 confirms that layer; the AI then re-derives image and execution choices from the confirmed direction + design system; Stage 3 confirms the final operational layer. Color swatches, live font previews, icon samples, image-style reference previews, and candidate picks appear where they help judgment; the chat path is the always-valid fallback. scripts/docs/confirm_ui.md owns the schema, server lifecycle, port strategy, and fallback details; this section keeps the orchestration contract. The split:
| Stage | Confirms | Driven by |
|---|---|---|
| 1 — direction anchors | canvas · audience + core message + content_divergence + delivery_purpose (PPT only — omitted on non-PPT canvases) (all §c key info) · mode + visual_style |
the source + user intent |
| 2 — design system (re-derived from Stage 1) | page count · color · typography (font + size) · icons · formula policy | the confirmed Stage 1 |
| 3 — images / execution (re-derived from Stage 1 + Stage 2) | image usage · generated-image style · AI-image generation path · generation mode · refine-spec toggle | the confirmed direction + design system |
Why three stages. Design-system fields are anchored by the same few choices (
visual_styleanchors color / icon / typography;delivery_purposesets the body size, page density, and the page-count recommendation). Image strategy depends on both the confirmed visual direction and the confirmed color system — its palette is color behavior only, while final HEX values follow Stage 2. Confirming direction first, then design system, then image / execution choices means each downstream stage fits the user's real choices instead of the AI's original assumptions. Page count is a derived field (content volume ×delivery_purpose), which is why it lives in Stage 2, not up front.
Steps:
⛔ Steps 2 → 3 → 4 are ONE uninterrupted run — do NOT yield to the user mid-flow. When an intermediate
--waitreturns, the AI immediately and autonomously re-derives and writes the next stage in the same turn: do not summarize, ask a question, report progress, or end the turn in between. The browser is sitting on a "deriving…" spinner polling for the next stage you must write — stopping here strands the page and the user must prod you in chat to finish (a bug, not the intended flow). Stage-1 and Stage-2 confirmations are intermediate machine handoffs, not stopping points. The single ⛔ BLOCKING wait is the final confirmation at the end of step 4. (Chat-fallback path — only when the page never opened — is the exception: there you do present each stage in chat and wait for a reply.)
- Write Stage 1 to
<project_path>/confirm_ui/recommendations.jsonwith"stage": "stage1"and only the anchor fields. New recommendations MUST use the canonicalstageselector. Enumerable anchors (canvas/mode/visual_style/delivery_purpose) name a recommended canonicalidin arecommendblock (the page lists common options fromconfirm_ui/static/catalogs.json);visual_stylealso carries the ≥3-stylevisual_style_spectrum(safe / shifted / bold — same hard rule as h.5).audienceandcontent_divergenceare plain{ "value": "<free text>" }.content_divergenceis the free-text field shown under audience in §c — how closely to follow the source vs how freely to reshape it (blank = balanced; facts stay sourced at every level); it is consumed by Strategist when authoring§IX, recorded indesign_spec.md §I, carries no page-count coupling, and is not written tospec_lock.md. Setlangto the page language (zh/en/ja); visible text matcheslang, or provide multilingualname_zh/name_en/name_ja+note_zh/note_en/note_ja— when the user's language is Japanese, setlang: "ja"and always include the_javariants (labels resolve in the page language first — ajapage falls back ja → en → zh, so missing_jalabels silently render in English; zh/en pages keep their zh↔en fallback and only try_jalast). - Launch + wait for Stage 1. Background launch; the parent returns when the page writes the stage-1
result.json. Long tool timeout — 600000 ms (the--wait≈590 s budget):
Page opens atpython3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --daemon --waithttp://localhost:5050— the same port as the Step 6 live preview (they never run at once: this page shuts down at the end of Step 4). If 5050 is held, the launcher auto-advances (5051, …) — read the actual URL from the launch log and report it. The page does not close after Stage 1: it shows a "deriving…" state and polls for Stage 2. Launch or wait failure is non-fatal: if it fails or times out (flask missing, port blocked, no GUI / remote / web host), do NOT troubleshoot — on any non-zero exit, re-checkresult.jsononce for a freshstatus: stage1-confirmedbefore dropping to the chat fallback. On success (exit 0 with a stage-1 result), do not pause or report — go straight to step 3 in the same turn. - Re-derive Stage 2 from the confirmed anchors, write it, then wait for the design-system handoff — immediately, same turn (the page is polling for it). Read the stage-1
result.json(status: stage1-confirmed). Using the user's actual confirmed anchors (not your originals), author the design-system candidates and overwriterecommendations.jsonwith"stage": "stage2": page count (content volume ×delivery_purpose); color and typography as generative ≥3-candidate fields (creative recommendations always offer real choice; fewer than 3 only on the honest-shortfall exception, with a stated reason; color: corepalettewith background/secondary_bg/primary/accent/secondary_accent/body_text; typography: CJK + Latin forheadingandbodywithcsspreview stacks +body_sizeas the body baseline in px (every canvas) — one fixed value per confirmeddelivery_purpose(text20 /balanced24 /presentation32), not a range; each typography candidate must include topic-matchedsample_heading/sample_heading_latin/sample_body/sample_body_latinpreview text, never a fixed unrelated industry sample); enumerableicons/formula_policy(recommendedid). The still-open page polls, renders Stage 2, and preserves the user's Stage 1 picks. Then attach to the already-running page, do not relaunch:
This returns when the page writes the stage-2python3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --wait-only --wait-stage stage2result.json(status: stage2-confirmed). On a non-zero exit, re-checkresult.jsononce before falling back to chat. - Re-derive Stage 3 from the confirmed anchors + design system, then wait for the final confirmation. Read the stage-2
result.json. Author the image and execution recommendations and overwriterecommendations.jsonwith"stage": "stage3":image_usageas one or more source ids (["ai"],["ai","provided"],["web","placeholder"], or["none"];noneis exclusive);image_strategy.candidatesas exactly three non-custom rendering × palette recommendations from h.5 whenimage_usageincludesai(the page adds the fourth Custom card itself); enumerableimage_ai_path/generation_modeandrefine_spec(recommendedid/ boolean). If the recommendation involves several image sources, keep the source list structured inrecommend.image_usageand write the usage rationale / page-role guidance intoimage_notes(for example, "封面和章节页用 AI 主视觉,产品页优先用户素材,行业背景页可用网络参考"). Writeimage_ai_pathonly whenimage_usageincludesai. Spot-illustration lean is not a candidate field here: it derives from the lockedvisual_style's illustration propensity and is expressed only in the recommendation rationale /image_notes, never as a new confirmation field. Generated-image style palettes are color behavior only; final image colors follow the confirmed Stage-2color. Custom image-strategy dimensions are handled by the built-in Custom card, are prose-only, and should not promise a gallery reference image. Then attach to the already-running page, do not relaunch (same 600000 ms budget):
This is the ⛔ BLOCKING completion: returns when the page writes the finalpython3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --wait-onlyresult.json(status: confirmed,stage: final, carrying Stage 1 + Stage 2 + Stage 3 fields). On a non-zero exit, re-checkresult.jsononce. Confirmed sizes are already px (the system is px-only — no pt anywhere, no conversion): writeresult.jsontypography.body_size/sizesintodesign_spec.md/spec_lock.md/ SVG verbatim.generation_mode: "split"/refine_spec: trueare explicit user choices. - Close the confirm page (Mandatory cleanup — every path). Shut the server down before leaving Step 4 so it cannot keep holding port 5050 (which Step 6 live preview reuses):
Idempotent and required regardless of whether Confirm was clicked: clicking the final Confirm already shuts the page down (then a no-op); the chat-fallback path leaves it running. Run it after reading the confirmation, before Step 5.python3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --shutdown
Always also print each stage's recommendations + URL in chat as the always-valid fallback. The chat fallback is staged too: if the page never opens or a wait times out with no fresh result, present Stage 1 in chat → get confirmation → re-derive → present Stage 2 → get confirmation → re-derive → present Stage 3 → get confirmation → take those values. Either path converges.
Honoring the confirmation (result.json is authoritative — Mandatory): the confirmed values override your own recommendations when you write design_spec.md / spec_lock.md. A user who changed any field changed it on purpose. In particular, map image_usage to §VIII Acquire Via (its value names differ from §h options — translate). image_usage may be either a legacy single string or a Confirm UI multi-select array; for arrays, apply every selected source. image_notes, when present, is a user-authored image inte
…(truncated)