PPT Master Skill
AI-driven multi-format SVG content generation system. Converts source documents into high-quality SVG pages through multi-role collaboration and exports to PPTX.
Core Pipeline: Source Document → Create Project → [Template] → Strategist → [Image_Generator] → Executor Live Preview → Quality Check → Post-processing → Export
[!CAUTION]
🚨 Global Execution Discipline (MANDATORY)
This workflow is a strict serial pipeline. The following rules have the highest priority — violating any one of them constitutes execution failure:
- SERIAL EXECUTION — Steps MUST be executed in order; the output of each step is the input for the next. Non-BLOCKING adjacent steps may proceed continuously once prerequisites are met, without waiting for the user to say "continue"
- BLOCKING = HARD STOP — Steps marked ⛔ BLOCKING require a full stop; the AI MUST wait for an explicit user response before proceeding and MUST NOT make any decisions on behalf of the user
- NO CROSS-PHASE BUNDLING — Cross-phase bundling is FORBIDDEN. (Note: the Eight Confirmations in Step 4 are ⛔ BLOCKING — the AI MUST present recommendations and wait for explicit user confirmation before proceeding. Once the user confirms, all subsequent non-BLOCKING steps — design spec output, SVG generation, speaker notes, and post-processing — may proceed automatically without further user confirmation)
- GATE BEFORE ENTRY — Each Step has prerequisites (🚧 GATE) listed at the top; these MUST be verified before starting that Step
- NO SPECULATIVE EXECUTION — "Pre-preparing" content for subsequent Steps is FORBIDDEN (e.g., writing SVG code during the Strategist phase)
- NO SUB-AGENT SVG GENERATION — Executor Step 6 SVG generation is context-dependent and MUST be completed by the current main agent end-to-end. Delegating page SVG generation to sub-agents is FORBIDDEN
- SEQUENTIAL PAGE GENERATION ONLY — In Executor Step 6, after the global design context is confirmed, SVG pages MUST be generated sequentially page by page in one continuous pass. Grouped page batches (for example, 5 pages at a time) are FORBIDDEN
- SPEC_LOCK RE-READ PER PAGE — Before generating each SVG page, Executor MUST
read_file <project_path>/spec_lock.md. All colors / fonts / icons / images MUST come from this file — no values from memory or invented on the fly. Executor MUST also look up the current page'spage_rhythm(anchor/dense/breathing),page_layouts(which template SVG to inherit, if any), andpage_charts(which chart template to adapt, if any). Empty / absent entries are intentional Strategist signals — see executor-base.md §2.1. This rule exists to resist context-compression drift on long decks and to break the uniform "every page is a card grid" default- SVG MUST BE HAND-WRITTEN, NOT SCRIPT-GENERATED — Every SVG page is written by the main agent directly, one page at a time (see rules 6 and 7). Writing or running a Python / Node / shell script that produces the SVG files in batch — looping over pages, templating from data, or emitting them via a generator — is FORBIDDEN, including under "save tokens", "quick draft", or "user is in a hurry" pretexts. The script-generation path was tried on a feature branch and abandoned: cross-page visual consistency depends on per-page authoring with full upstream context, which a generator script cannot reproduce
[!IMPORTANT]
🌐 Language & Communication Rule
- Response language: match the user's input and source materials. Explicit user override (e.g., "请用英文回答") takes precedence.
- Template format:
design_spec.mdMUST follow its original English template structure (section headings, field names) regardless of conversation language. Content values may be in the user's language.
[!IMPORTANT]
🔌 Compatibility With Generic Coding Skills
ppt-masteris a repository-specific workflow, not a general application scaffold- Do NOT create
.worktrees/,tests/, branch workflows, or generic engineering structure by default- On conflict with a generic coding skill, follow this skill unless the user explicitly says otherwise
Main Pipeline Scripts
| Script | Purpose |
|---|---|
${SKILL_DIR}/scripts/source_to_md/pdf_to_md.py |
PDF to Markdown |
${SKILL_DIR}/scripts/source_to_md/doc_to_md.py |
Documents to Markdown — native Python for DOCX/HTML/EPUB/IPYNB, pandoc fallback for legacy formats (.doc/.odt/.rtf/.tex/.rst/.org/.typ) |
${SKILL_DIR}/scripts/source_to_md/excel_to_md.py |
Excel workbooks to Markdown — supports .xlsx/.xlsm; legacy .xls should be resaved as .xlsx |
${SKILL_DIR}/scripts/source_to_md/ppt_to_md.py |
PowerPoint to Markdown |
${SKILL_DIR}/scripts/pptx_intake.py |
Standard PPTX intake enrichment — canvas / identity / slide geometry / tables / native chart data |
${SKILL_DIR}/scripts/source_to_md/web_to_md.py |
Web page to Markdown (supports WeChat via curl_cffi) |
${SKILL_DIR}/scripts/project_manager.py |
Project init / validate / manage |
${SKILL_DIR}/scripts/icon_sync.py |
Copy chosen library icons into <project>/icons/ at selection time; missing names reported + non-zero (re-pick gate) |
${SKILL_DIR}/scripts/analyze_images.py |
Image analysis |
${SKILL_DIR}/scripts/latex_render.py |
LaTeX formula rendering (manifest-driven PNG assets) |
${SKILL_DIR}/scripts/image_gen.py |
AI image generation (multi-provider) |
${SKILL_DIR}/scripts/svg_quality_checker.py |
SVG quality check |
${SKILL_DIR}/scripts/total_md_split.py |
Speaker notes splitting |
${SKILL_DIR}/scripts/finalize_svg.py |
SVG post-processing (unified entry) |
${SKILL_DIR}/scripts/svg_to_pptx.py |
Export to PPTX |
${SKILL_DIR}/scripts/native_enhance_pptx.py |
Existing PPTX enhancement project init / validation / direct OOXML patch export |
${SKILL_DIR}/scripts/native_narration_pptx.py |
Backward-compatible entrypoint for existing PPTX notes / narration enhancement |
${SKILL_DIR}/scripts/update_spec.py |
Propagate a spec_lock.md color / font_family change across all generated SVGs |
For complete tool documentation, see ${SKILL_DIR}/scripts/README.md.
Windows note: if a
python3 ...command fails (common on python.org installs, which providepython.exebut notpython3.exe), rerun the same command withpythoninstead.
Template Index
| Index | Path | Purpose |
|---|---|---|
| Layout templates | ${SKILL_DIR}/templates/layouts/layouts_index.json |
Query available page layout templates |
| Brand presets | ${SKILL_DIR}/templates/brands/brands_index.json |
Query available brand identity presets (color / typography / logo / voice) |
| Visualization templates | ${SKILL_DIR}/templates/charts/charts_index.json |
Query available visualization SVG templates (charts, infographics, diagrams, frameworks) |
| Icon library | ${SKILL_DIR}/templates/icons/ |
See ${SKILL_DIR}/templates/icons/README.md; search icons on demand with ls templates/icons/<library>/ | grep <keyword> |
Standalone Workflows
| Workflow | Path | Purpose |
|---|---|---|
topic-research |
workflows/topic-research.md |
Pre-pipeline — gather web sources when the user supplies only a topic with no source files |
template-fill |
workflows/template-fill-pptx.md |
Give a native PPTX template deck plus source material; select fitting pages (a page may be reused for several output slides) and fill text back without SVG conversion |
beautify |
workflows/beautify-pptx.md |
Re-layout an existing PPTX through the SVG pipeline — preserve its text verbatim, inherit its palette/fonts as truth, redo only layout; mirror of template-fill |
create-template |
workflows/create-template.md |
Standalone layout template creation workflow |
create-brand |
workflows/create-brand.md |
Standalone brand-only template creation (identity preset; no SVG page roster) |
resume-execute |
workflows/resume-execute.md |
Phase B entry — resume execution in a fresh chat after Phase A (Step 1–5) completed in another session (split mode) |
verify-charts |
workflows/verify-charts.md |
Chart coordinate calibration — run after SVG generation if the deck contains data charts |
customize-animations |
workflows/customize-animations.md |
Object-level PPTX animation customization — run only when the user explicitly asks to tune animation order/effects/timing |
native-enhance-pptx |
workflows/native-enhance-pptx.md |
Existing PPTX native enhancement — optimize a finished deck by appending notes / audio / auto-advance / page transitions without changing existing content or layout |
native-narration-pptx |
workflows/native-narration-pptx.md |
Compatibility reference for the notes / narration subset of native-enhance-pptx |
live-preview |
workflows/live-preview.md |
Browser-based live preview — auto-started during generation and re-enterable any time the user mentions "live preview", "preview", "看效果", or wants to click/select a slide element |
visual-review |
workflows/visual-review.md |
Per-page rubric-based visual self-check — run only when the user explicitly asks for a visual re-pass on the generated SVGs (between Executor and post-processing). Opt-in only; never invoked by the main pipeline. |
PPTX Route Boundary
When the user provides an existing .pptx, route by the role of the source deck:
| User intent | Route | Contract |
|---|---|---|
| Preserve the deck's page split, page order, and per-slide wording; improve layout / hierarchy / whitespace | beautify |
Source page count and order are 1:1; text and data values are frozen; visual identity is inherited after confirmation |
| Treat the deck as source material; rethink the story, merge / split / drop / reorder pages, or change page count | Main pipeline | ppt_to_md + PPTX intake provide content facts and candidates; Strategist may re-architect freely |
| Reuse the deck's native design with new material | template-fill |
Clone selected source slides and replace text / table / chart data directly in OOXML; no SVG generation |
| Harvest the deck as a reusable future template | create-template |
Build a template package, not a one-off generated deck |
| Keep the finished deck visually stable and append native optimizations such as notes / narration audio / automatic playback | native-enhance-pptx |
Archive the source PPTX into the project (projects/ sources move; external sources copy) and patch enhancement metadata/media directly in OOXML; no SVG generation |
Deciding axis (beautify vs main pipeline) — one question, one discriminator: is the source's page split a finished artifact to preserve, or a draft structure to overturn? The concrete discriminator is page count / order: if it changes at all — any split, merge, drop, or reorder — it is the main pipeline, never beautify. Beautify is strictly 1:1: same page count, same order, text verbatim, only layout / hierarchy / whitespace redone. Edge case made explicit: "keep all the content but split a crowded page so it reads better" still changes page count, so it is the main pipeline (re-pagination is re-architecture), not beautify.
Ambiguous requests such as "make this PPT more professional" or "optimize this deck" MUST be clarified with one question before routing: "Should the original page count/order and each slide's wording be preserved, or should the deck be treated as source material and restructured into a new story?" Preserve → beautify; restructure → main pipeline.
Workflow
Step 1: Source Content Processing
🚧 GATE: User has provided source material (PDF / DOCX / EPUB / URL / Markdown file / text description / conversation content — any form is acceptable).
No source content? When the user supplies only a topic name or requirements without any file or substantive description, run the
topic-researchworkflow first, then return here with its products as input.
When the user provides non-Markdown content, convert immediately:
| User Provides | Command |
|---|---|
| PDF file | python3 ${SKILL_DIR}/scripts/source_to_md/pdf_to_md.py <file> |
| DOCX / Word / Office document | python3 ${SKILL_DIR}/scripts/source_to_md/doc_to_md.py <file> |
| XLSX / XLSM / Excel workbook | python3 ${SKILL_DIR}/scripts/source_to_md/excel_to_md.py <file> |
| CSV / TSV | Read directly as plain-text table source |
| PPTX / PowerPoint deck | python3 ${SKILL_DIR}/scripts/source_to_md/ppt_to_md.py <file> for Markdown content; after Step 2 import-sources, standard PPTX intake is also written to <project>/analysis/ |
| EPUB / HTML / LaTeX / RST / other | python3 ${SKILL_DIR}/scripts/source_to_md/doc_to_md.py <file> |
| Web link | python3 ${SKILL_DIR}/scripts/source_to_md/web_to_md.py <URL> |
| WeChat / high-security site | python3 ${SKILL_DIR}/scripts/source_to_md/web_to_md.py <URL> (requires curl_cffi, included in requirements.txt) |
| Markdown | Read directly |
Office vector assets (EMF/WMF) from DOCX/PPTX sources:
doc_to_md.py/ppt_to_md.pyextract embedded Office vector images (.emf/.wmf) alongside bitmap images. Afterimport-sources, these land inimages/together withimage_manifest.jsonand are first-class assets in §VIII Image Resource List.Do NOT convert EMF/WMF to PNG. The PPT Master pipeline preserves them as external references (
finalize_svg.pyskips them) andsvg_to_pptx.pyembeds them as PPTX-native media viaimage/x-emf/image/x-wmfMIME — PowerPoint renders them at full vector fidelity. Converting via LibreOffice/Inkscape introduces CJK font substitution drift and rasterization loss; the original EMF/WMF is always higher fidelity than the converted PNG.Browser-based live preview cannot render EMF (will show blank) — this is expected; the PPTX output is the source of truth.
✅ Checkpoint — Confirm source content is ready, proceed to Step 2.
Step 2: Project Initialization
🚧 GATE: Step 1 complete; source content is ready (Markdown file, user-provided text, or requirements described in conversation are all valid).
python3 ${SKILL_DIR}/scripts/project_manager.py init <project_name> --format <format>
Format options: ppt169 (default), ppt43, xhs, story, etc. For the full format list, see references/canvas-formats.md.
Import source content (choose based on the situation):
| Situation | Action |
|---|---|
| Has source files (PDF/MD/etc.) | python3 ${SKILL_DIR}/scripts/project_manager.py import-sources <project_path> <source_files...> --move |
| User provided text directly in conversation | No import needed — content is already in conversation context; subsequent steps can reference it directly |
For PPTX sources, import-sources automatically runs the standard intake enrichment:
python3 ${SKILL_DIR}/scripts/pptx_intake.py <project_path>/sources/<source.pptx> -o <project_path>/analysis
For each PPTX it writes <stem>.identity.json (canvas, theme palette/fonts, observed usage) and <stem>.slide_library.json (text slots, geometry, native tables, native chart caches), and merges that deck's Strategist-facing digest into the single multi-deck index analysis/source_profile.json (decks[], one self-contained entry per source deck, with prefixed artifact pointers). In the main generation path these are source facts and recommendation candidates, not replica constraints; beautify and template-fill workflows decide separately which fields become locked constraints.
Multi-deck: several PPTX files may be imported into one main-pipeline project — each gets its own <stem>.* artifacts and a deck entry in source_profile.json. source_profile.json stays the single must-read index (one entry for a one-deck project, several for a combined-source project). Stems must be distinct; re-importing the same stem replaces that deck's entry. The beautify / template-fill workflows remain single-deck (1:1 to one chosen source deck) and read that deck's <stem>.* artifacts.
⚠️ MUST use
--move(not copy): all source files — Step 1's generated Markdown, original PDFs / MDs / images — go intosources/viaimport-sources --move. After execution they no longer exist at the original location. Intermediate artifacts (e.g.,_files/) are handled automatically.
✅ Checkpoint — Confirm project structure created successfully, sources/ contains all source files, converted materials are ready. Proceed to Step 3.
Step 3: Template Option
🚧 GATE: Step 2 complete; project directory structure is ready.
Default — free design. Proceed directly to Step 4. Do NOT query any *_index.json unless triggered. Do NOT ask the user. Do NOT proactively suggest, hint at, or fuzzy-match any template based on content, slug-like words, or vague style descriptions.
Template flow triggers ONLY on explicit directory paths supplied by the user in their initial message. The trigger rule is mechanical, not interpretive:
| User input contains | Step 3 action |
|---|---|
One or more explicit template directory paths (each resolves to a directory containing design_spec.md with kind: brand / kind: layout / kind: deck in its YAML frontmatter) |
Read each spec's kind, dispatch per the kind matrix below, fuse if multiple |
| Anything else — bare template names ("用 academic_defense"), style descriptions ("麦肯锡风格"), brand mentions ("招商银行风格"), vague intent ("想用个模板"), or silence | Skip Step 3, free design |
There is no slug matching, no name lookup, no fuzzy resolution. A name without a path does not trigger — the user must give a path the AI can cd into.
Style descriptions ("麦肯锡风格" / "Keynote 风" / "极简风" / etc.) never trigger Step 3. They flow into Strategist's Eight Confirmations as a style brief (color / typography / tone in confirmations e–g).
Bare names ("academic_defense", "招商银行", "anthropic") do NOT trigger Step 3 even if a matching directory exists in the library. The user must give a path. AI must not "helpfully" resolve a name to a path.
"What templates exist?" is out-of-band Q&A — answer by listing entries from
brands_index.json/layouts_index.json/decks_index.jsontogether with their paths. Listing alone does not advance the pipeline; the user must send a path back to trigger Step 3.
To create a new layout or deck, read
workflows/create-template.md. To create a new brand, readworkflows/create-brand.md.
Three template kinds
The architecture has three independent reference bundles. Full schema in docs/zh/templates-architecture.md. Summary:
| Kind | Physical dir | Contains | Frontmatter |
|---|---|---|---|
| brand | templates/brands/<id>/ |
identity-only segment: color / typography / logo / voice / icon style | kind: brand |
| layout | templates/layouts/<id>/ |
structure-only segment: canvas / page structure / page types / SVG roster | kind: layout |
| deck | templates/decks/<id>/ |
full replica: identity + structure + middle (template overview) segments | kind: deck |
Segment ownership (governs fusion override priority):
| Segment | Sections | Owner kind on fusion |
|---|---|---|
| Identity | Color Scheme / Typography / Logo / Voice & Tone / Icon Style | brand |
| Structure | Canvas / Page Structure / Page Types / SVG Roster | layout |
| Middle | Template Overview (use cases / design intent) | deck (no other kind writes this) |
Single-path dispatch
User path's kind |
Step 3 action |
|---|---|
kind: brand |
design_spec.md + non-image assets → <project>/templates/; logo / illustration / icon bitmaps → <project>/images/. Strategist locks identity segment as truth; structure stays free. |
kind: layout |
design_spec.md + SVG roster → <project>/templates/; any bitmap assets → <project>/images/. Strategist locks structure; identity decided in Eight Confirmations e–g. |
kind: deck |
design_spec.md + template SVGs → <project>/templates/; logos / backgrounds / other bitmaps → <project>/images/. Strategist locks all segments; Eight Confirmations narrows to deck-content fields (audience / page count / outline / tone tweaks). |
TEMPLATE_DIR=<user-supplied path>
# Bitmaps join the project's single runtime image pool (images/, referenced as
# ../images/); the spec + template SVGs + other non-image assets stay in
# templates/ as design reference the Strategist/Executor read but never render.
cp -r ${TEMPLATE_DIR}/* <project_path>/templates/
find <project_path>/templates -type f \( -iname '*.png' -o -iname '*.jpg' -o -iname '*.jpeg' -o -iname '*.gif' -o -iname '*.webp' -o -iname '*.bmp' \) -exec mv {} <project_path>/images/ \;
The same split applies to all three kinds — bitmaps always land in images/, the rest in templates/. The spec's kind field tells Strategist how to read the templates/ side; downstream code doesn't distinguish. (Template SVGs in templates/ are reference material only — the rendered pages live in svg_output/ and reference images via ../images/.)
Multi-path fusion
When the user gives two or more paths of different kinds, Step 3 fuses them into a single <project>/templates/design_spec.md. Default granularity is segment-level integer replacement — entire identity / structure / middle segments are taken from the highest-priority source for that segment, no implicit field-level mixing.
Override priority by segment:
| Combination | Identity from | Structure from | Middle from |
|---|---|---|---|
| brand only | brand | (free design) | (none) |
| layout only | (free design) | layout | (none) |
| deck only | deck | deck | deck |
| brand + layout | brand | layout | (none) |
| brand + deck | brand (overrides deck) | deck | deck |
| layout + deck | deck | layout (overrides deck) | deck |
| brand + layout + deck | brand | layout | deck |
Field-level micro-adjustment (e.g. "use anthropic brand but primary changed to #FF0000") is not part of Step 3 fusion — it flows into Strategist Eight Confirmations e–g as a normal user request.
Same-kind multiple paths — conflict resolution
When the user gives two paths of the same kind (e.g. brands/anthropic + brands/google), Step 3 surfaces a conflict prompt before fusing — like resolving a git merge conflict:
AI: 你给了两个 brand,检测到段级冲突:
- Color Scheme(Anthropic 橙红 vs Google 多色)
- Typography(Styrene/AnthropicSans vs GoogleSans/Roboto)
- Logo(Anthropic 标 vs Google 标)
- Voice & Tone(restrained vs friendly)
- Icon Style(stroke vs filled)
要 (a) 全部按 Anthropic / (b) 全部按 Google / (c) 逐段挑?
Rules:
- Default: no implicit ordering — every cross-source segment difference is reported as a conflict
- Only when the user picks
(c)does AI walk through each segment one by one - Field-level conflicts are out of scope — segment-level only
- Three or more same-kind paths are not supported — ask the user to converge to at most two
Fused spec provenance
When fusion happens (any multi-path case), the resulting <project>/templates/design_spec.md carries a provenance block immediately under its H1:
> **Fused from:**
> - deck: `templates/decks/招商银行/` (base)
> - brand: `templates/brands/anthropic/` (identity override)
> - layout: `templates/layouts/academic_defense/` (structure override)
> - conflicts resolved: Color Scheme from anthropic(user picked a)
Single-path Step 3 does not add provenance (the source is self-evident from the copied files).
✅ Checkpoint — Default path proceeds to Step 4 without user interaction. If the user supplied one or more explicit template paths, those have been dispatched (or fused) into <project_path>/templates/ before advancing.
Step 4: Strategist Phase (MANDATORY — cannot be skipped)
🚧 GATE: Step 3 complete; default free-design path taken, or (if triggered) template files copied into the project.
First, read the role definition:
Read references/strategist.md
⚠️ Mandatory gate: before writing
design_spec.md, Strategist MUSTread_file templates/design_spec_reference.mdand follow its full I–XI section structure. Seestrategist.mdSection 1.
<project_path>/analysis/ is the project's intermediate-analysis folder: the canonical home for machine-extracted source/asset facts — the PPTX intake bundle (source_profile.json index + per-deck <stem>.identity.json / <stem>.slide_library.json) and image_analysis.csv. It holds facts, not design contracts — design_spec.md / spec_lock.md stay at the project root. The MUST-read contract covers only the compact structured data files (.json / .csv); other artifacts that may live under analysis/ (e.g. a beautify source_svg_import/ vector reference package) are NOT bulk-read — they are read selectively only when a specific workflow step calls for them. Before the Eight Confirmations, Strategist MUST read the auto-extracted fact files already in analysis/ — currently source_profile.json (PPTX intake), when present. This file is the multi-deck index: read it once for the decks[] digests (canvas / chart / table entries per source deck), then open a specific deck's <stem>.identity.json / <stem>.slide_library.json only if you need its full raw facts. Use these entries as factual source context (format default + content facts); when several decks are present, synthesize across all of them. The source's palette / typography / visual identity are a reference, not a constraint: the main pipeline may inherit them where they fit the content and the confirmed style, or design fresh where they don't — the Strategist's judgment, never an obligation to either keep or discard. (Template-fill preserves the native source design by editing cloned slides directly; beautify defaults to the source identity but still follows the confirmed values; the main pipeline treats source identity as reference only and defaults to fresh design.) (image_analysis.csv lands later, at the image-analysis step below, and is the authoritative regenerated image-fact view there — re-derived from the live images/ folder, not a durable store.)
Channel ownership — read each fact once from its owning channel. In the main pipeline the content contract is the Markdown (sources/<stem>.md): text, tables, and chart data values all come from there (ppt_to_md now transcribes native chart data into Markdown tables). The analysis/ chart / table entries are a structural digest for outline decisions (which slides carried charts, type, series names) — not a second copy of the values; do NOT also pull chart values from <stem>.slide_library.json in the main pipeline. The <stem>.slide_library.json full structured data is owned by the direct-PPTX workflows: template-fill uses it as the native fill contract; beautify uses it for native chart / table data while keeping slide text from the Markdown.
Eight Confirmations (full template: templates/design_spec_reference.md):
⛔ BLOCKING: present the Eight Confirmations and wait for explicit user confirmation or modification before outputting Design Specification & Content Outline. This is the single core confirmation gate — once the final confirmation lands, all subsequent steps proceed automatically. The default Confirm UI delivers the gate in two tiers (anchors → re-derive → realization; see below); the chat fallback mirrors the same two steps.
- Canvas format
- Page count range
- Target audience
- Style objective
- Color scheme
- Icon usage approach
- Typography plan, including formula rendering policy
- Image usage approach
Confirm UI Auto-Launch (Mandatory — default visual confirmation surface): by default the Eight Confirmations are presented through an interactive local page in two tiers within one browser session — Tier 1 confirms the anchors; the AI then re-derives the realization layer from the user's actual anchors; Tier 2 confirms that layer. Color swatches, live font previews, candidate picks; the chat path is the always-valid fallback. The split (full field rules: scripts/docs/confirm_ui.md):
| Tier | Confirms | Driven by |
|---|---|---|
| 1 — anchors | canvas · audience + core message + content_divergence + delivery_purpose (PPT only — omitted on non-PPT canvases) (all §c key info) · mode + visual_style |
the source + user intent |
| 2 — realization (re-derived from Tier 1) | page count · color · typography (font + size) · icons · formula policy · image usage + strategy · generation mode · refine-spec toggle | the confirmed Tier 1 |
Why two tiers. Every realization field is anchored by the same few choices (
visual_styleanchors color / icon / typography / image;delivery_purposesets the body size, page density, and the page-count recommendation). Confirming anchors first, then re-deriving, means Tier 2's candidates fit the user's real anchors instead of your originals — the coherence reconciliation below is done by construction on this path. Page count is a derived field (content volume ×delivery_purpose), which is why it lives in Tier 2, not up front.
Steps:
⛔ Steps 2 → 3 → 4 are ONE uninterrupted run — do NOT yield to the user mid-flow. When the tier-1
--wait(step 2) returns, the AI immediately and autonomously continues to step 3 (re-derive + write Tier 2) and step 4 (--wait-only) in the same turn: do not summarize, ask a question, report progress, or end the turn in between. The browser is sitting on a "deriving…" spinner polling for the Tier 2 you must write next — stopping here strands the page and the user must prod you in chat to finish (a bug, not the intended flow). The tier-1 confirmation is an intermediate machine handoff, not a stopping point. The single ⛔ BLOCKING wait is the final confirmation at the end of step 4. (Chat-fallback path — only when the page never opened — is the exception: there you do present each tier in chat and wait for a reply.)
- Write Tier 1 to
<project_path>/confirm_ui/recommendations.jsonwith"tier": 1and only the anchor fields. Enumerable anchors (canvas/mode/visual_style/delivery_purpose) name a recommended canonicalidin arecommendblock (the page lists common options fromconfirm_ui/static/catalogs.json);visual_stylealso carries the ≥3-stylevisual_style_spectrum(safe / shifted / bold — same hard rule as h.5).audienceandcontent_divergenceare plain{ "value": "<free text>" }.content_divergenceis the free-text field shown under audience in §c — how closely to follow the source vs how freely to reshape it (blank = balanced; facts stay sourced at every level); it is consumed by Strategist when authoring§IX, recorded indesign_spec.md §I, carries no page-count coupling, and is not written tospec_lock.md. Setlangto the page language; visible text matcheslang, or provide bilingualname_zh/name_en+note_zh/note_en. - Launch + wait for Tier 1. Background launch; the parent returns when the page writes the tier-1
result.json. Long tool timeout — 600000 ms (the--wait≈590 s budget):
Page opens atpython3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --daemon --waithttp://localhost:5050— the same port as the Step 6 live preview (they never run at once: this page shuts down at the end of Step 4). If 5050 is held, the launcher auto-advances (5051, …) — read the actual URL from the launch log and report it. The page does not close after Tier 1: it shows a "deriving…" state and polls for Tier 2. Launch or wait failure is non-fatal: if it fails or times out (flask missing, port blocked, no GUI / remote / web host), do NOT troubleshoot — on any non-zero exit, re-checkresult.jsononce (a freshstatus: tier1-confirmed) before dropping to the chat fallback. On success (exit 0 with a tier-1 result), do not pause or report — go straight to step 3 in the same turn. - Re-derive Tier 2 from the confirmed anchors, then write it — immediately, same turn (the page is polling for it). Read the tier-1
result.json(status: tier1-confirmed). Using the user's actual confirmed anchors (not your originals), author the realization candidates and overwriterecommendations.jsonwith"tier": 2: page count (content volume ×delivery_purpose); color, typography, and generated-image style as generative ≥3-candidate fields (creative recommendations always offer real choice — same rule as h.5; fewer than 3 only on the honest-shortfall exception, with a stated reason; color: corepalettewith background/secondary_bg/primary/accent/secondary_accent/body_text; typography: CJK + Latin forheadingandbodywithcsspreview stacks +body_sizeas the body baseline in px (every canvas) — one fixed value per confirmeddelivery_purpose(text20 /balanced24 /presentation32), not a range; images:image_strategy.candidatesrendering × palette from h.5); enumerableicons/formula_policy/generation_mode(recommendedid);image_usage(ai/web/provided/placeholder/none, or a custom prose plan when several sources mix — never bare"custom"; writeimage_ai_pathonly when the plan includes AI). The still-open page polls, renders Tier 2, and preserves the user's Tier 1 picks. Closed fields (image_ai_path,formula_policy,generation_mode,refine_spec) stay finite; open fields (icons,image_usage, typography custom text) show a Custom box. - Wait for the final confirmation — attach to the already-running page, do not relaunch (same 600000 ms budget):
This is the ⛔ BLOCKING completion: returns when the page writes the finalpython3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --wait-onlyresult.json(status: confirmed,stage: final, carrying all Tier 1 + Tier 2 fields). On a non-zero exit, re-checkresult.jsononce. Confirmed sizes are already px (the system is px-only — no pt anywhere, no conversion): writeresult.jsontypography.body_size/sizesintodesign_spec.md/spec_lock.md/ SVG verbatim.generation_mode: "split"/refine_spec: trueare explicit user choices. - Close the confirm page (Mandatory cleanup — every path). Shut the server down before leaving Step 4 so it cannot keep holding port 5050 (which Step 6 live preview reuses):
Idempotent and required regardless of whether Confirm was clicked: clicking the final Confirm already shuts the page down (then a no-op); the chat-fallback path leaves it running. Run it after reading the confirmation, before Step 5.python3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --shutdown
Always also print each tier's recommendations + URL in chat as the always-valid fallback. The chat fallback is two-step too: if the page never opens or a wait times out with no fresh result, present Tier 1 in chat → get confirmation → re-derive → present Tier 2 → get confirmation → take those values. Either path converges.
Honoring the confirmation (result.json is authoritative — Mandatory): the confirmed values override your own recommendations when you write design_spec.md / spec_lock.md. A user who changed any field changed it on purpose. In particular, map image_usage to §VIII Acquire Via (its value names differ from §h options — translate):
result.json.image_usage |
§VIII Acquire Via |
h.5 + Step 5 generation |
|---|---|---|
ai (or a custom plan that includes AI) |
ai rows |
Run h.5 (lock rendering + palette); Step 5 generates |
web |
web rows |
None |
provided |
user rows |
None — never generate |
placeholder |
placeholder rows |
None |
none |
no image rows (§h option A) | None |
When the confirmed image_usage is not ai (and the plan has no AI part), do NOT run h.5, do NOT write ai rows, and do NOT generate images in Step 5 — regardless of what you recommended. The same "confirmed value wins" rule applies to every field (color → §III, typography → §IV, etc.).
Upstream override → re-derive untouched downstream (Mandatory — chat-fallback / single-pass path). On the two-tier page path this is already handled (Step 3 re-derives Tier 2 from the user's actual anchors). It still applies whenever anchors and realization are confirmed together — the two-step chat fallback collapsed into one bundle, or a legacy single-pass result.json. "Confirmed value wins" governs each field's own value — never recompute a value the user set (a size, canvas, or palette they edited stays verbatim). But a single-pass result.json can carry a changed anchor beside downstream fields still holding your original — now incoherent — recommendation (e.g. switched to dark-tech while the light palette you proposed is untouched). Before writing the spec, reconcile: when the user changed an anchor, re-derive the downstream fields the user did not themselves edit so they realize the new anchor; fields the user pinned stay as confirmed.
| Anchor the user changed | Re-derive (only the downstream fields the user left at your recommendation) |
|---|---|
visual_style (§d Layer 2 — anchors e–h) |
color neutral tiers (§e), icon library / stroke (§f), typography character (§g), image rendering (§h.5) |
mode (§d Layer 1) |
outline structure + register (§IX) |
delivery_purpose (§g) |
body baseline + per-page density / rhythm (§6.1) |
audience / core message (§c) |
tone across e–h, outline emphasis (§IX) |
color HEX (§e) |
h.5 palette (re-filter for the new HEX) |
Reconcile without a new blocking wait — fold the coherent values into design_spec.md / spec_lock.md and state the adjustment in the §8 next-step handoff (e.g. "you switched to dark-tech; the light palette you had left no longer fit, so background / accent were re-derived — tell me if you wanted the original"). Canvas is the explicit exception: font sizes are deliberately not rescaled on a canvas change (see strategist §g).
Opt-out: if the user has said they don't want the page (e.g. "不要网页" / "just confirm in chat" / "纯聊天确认"), skip the launch entirely (step 2) and present the Eight Confirmations in chat as before — steps 1, 3, 4 still apply (recommendations summary in chat; wait; take chat values).
The page is a confirmation surface only — Strategist still authors every recommendation; the page never generates content.
Mandatory — split-mode note (not a ninth confirmation): after listing the eight confirmation details, you MUST append exactly one short line (rendered in the user's language, prefixed with 💡) about generation mode. Pick the variant by qualitative read of Phase A signals — recommended page count, source-material bulk, whether topic-research ran with substantial web-fetch accumulation:
| Signal read | Line content |
|---|---|
| Heavy (long page count / bulky sources / heavy web-fetch accumulation) | State estimated page count and large source size; recommend switching to split mode after Step 5 — stop this chat, open a fresh window and input 继续生成 projects/<project_name> to enter Phase B (SVG generation + export); no response or "continue" = default continuous mode. |
| Normal (default) | State scale is moderate, default continuous mode generates in one go; if mid-way window switch is desired, input 继续生成 projects/<project_name> after Step 5 to switch to [split |
…(truncated)