PPT Master Skill
AI-driven multi-format SVG content generation system. Converts source documents into high-quality SVG pages through multi-role collaboration and exports to PPTX.
Core Pipeline: Source Document → Create Project → [Template] → Strategist → [Image_Generator] → Executor Live Preview → Quality Check → Post-processing → Export
[!CAUTION]
🚨 Global Execution Discipline (MANDATORY)
This workflow is a strict serial pipeline. The following rules have the highest priority — violating any one of them constitutes execution failure:
- SERIAL EXECUTION — Steps MUST be executed in order; the output of each step is the input for the next. Non-BLOCKING adjacent steps may proceed continuously once prerequisites are met, without waiting for the user to say "continue"
- BLOCKING = HARD STOP — Steps marked ⛔ BLOCKING require a full stop; the AI MUST wait for an explicit user response before proceeding and MUST NOT make any decisions on behalf of the user
- NO CROSS-PHASE BUNDLING — Cross-phase bundling is FORBIDDEN. (Note: the Eight Confirmations in Step 4 are ⛔ BLOCKING — the AI MUST present recommendations and wait for explicit user confirmation before proceeding. Once the user confirms, all subsequent non-BLOCKING steps — design spec output, SVG generation, speaker notes, and post-processing — may proceed automatically without further user confirmation)
- GATE BEFORE ENTRY — Each Step has prerequisites (🚧 GATE) listed at the top; these MUST be verified before starting that Step
- NO SPECULATIVE EXECUTION — "Pre-preparing" content for subsequent Steps is FORBIDDEN (e.g., writing SVG code during the Strategist phase)
- NO SUB-AGENT SVG GENERATION — Executor Step 6 SVG generation is context-dependent and MUST be completed by the current main agent end-to-end. Delegating page SVG generation to sub-agents is FORBIDDEN
- SEQUENTIAL PAGE GENERATION ONLY — In Executor Step 6, after the global design context is confirmed, SVG pages MUST be generated sequentially page by page in one continuous pass. Grouped page batches (for example, 5 pages at a time) are FORBIDDEN
- SPEC_LOCK RE-READ PER PAGE — Before generating each SVG page, Executor MUST
read_file <project_path>/spec_lock.md. All colors / fonts / icons / images MUST come from this file — no values from memory or invented on the fly. Executor MUST also look up the current page'spage_rhythm(anchor/dense/breathing),page_layouts(which template SVG to inherit, if any), andpage_charts(which chart template to adapt, if any). Empty / absent entries are intentional Strategist signals — see executor-base.md §2.1. This rule exists to resist context-compression drift on long decks and to break the uniform "every page is a card grid" default- SVG MUST BE HAND-WRITTEN, NOT SCRIPT-GENERATED — Every SVG page is written by the main agent directly, one page at a time (see rules 6 and 7). Writing or running a Python / Node / shell script that produces the SVG files in batch — looping over pages, templating from data, or emitting them via a generator — is FORBIDDEN, including under "save tokens", "quick draft", or "user is in a hurry" pretexts. The script-generation path was tried on a feature branch and abandoned: cross-page visual consistency depends on per-page authoring with full upstream context, which a generator script cannot reproduce
[!IMPORTANT]
🌐 Language & Communication Rule
- Response language: match the user's input and source materials. Explicit user override (e.g., "请用英文回答") takes precedence.
- Template format:
design_spec.mdMUST follow its original English template structure (section headings, field names) regardless of conversation language. Content values may be in the user's language.
[!IMPORTANT]
🔌 Compatibility With Generic Coding Skills
ppt-masteris a repository-specific workflow, not a general application scaffold- Do NOT create
.worktrees/,tests/, branch workflows, or generic engineering structure by default- On conflict with a generic coding skill, follow this skill unless the user explicitly says otherwise
Main Pipeline Scripts
| Script | Purpose |
|---|---|
${SKILL_DIR}/scripts/source_to_md/pdf_to_md.py |
PDF to Markdown |
${SKILL_DIR}/scripts/source_to_md/doc_to_md.py |
Documents to Markdown — native Python for DOCX/HTML/EPUB/IPYNB, pandoc fallback for legacy formats (.doc/.odt/.rtf/.tex/.rst/.org/.typ) |
${SKILL_DIR}/scripts/source_to_md/excel_to_md.py |
Excel workbooks to Markdown — supports .xlsx/.xlsm; legacy .xls should be resaved as .xlsx |
${SKILL_DIR}/scripts/source_to_md/ppt_to_md.py |
PowerPoint to Markdown |
${SKILL_DIR}/scripts/pptx_intake.py |
Standard PPTX intake enrichment — canvas / identity / slide geometry / tables / native chart data |
${SKILL_DIR}/scripts/source_to_md/web_to_md.py |
Web page to Markdown (supports WeChat via curl_cffi) |
${SKILL_DIR}/scripts/project_manager.py |
Project init / validate / manage |
${SKILL_DIR}/scripts/icon_sync.py |
Copy chosen library icons into <project>/icons/ at selection time; missing names reported + non-zero (re-pick gate) |
${SKILL_DIR}/scripts/analyze_images.py |
Image analysis |
${SKILL_DIR}/scripts/latex_render.py |
LaTeX formula rendering (manifest-driven PNG assets) |
${SKILL_DIR}/scripts/image_gen.py |
AI image generation (multi-provider) |
${SKILL_DIR}/scripts/svg_quality_checker.py |
SVG quality check |
${SKILL_DIR}/scripts/total_md_split.py |
Speaker notes splitting |
${SKILL_DIR}/scripts/finalize_svg.py |
SVG post-processing (unified entry) |
${SKILL_DIR}/scripts/svg_to_pptx.py |
Export to PPTX |
${SKILL_DIR}/scripts/update_spec.py |
Propagate a spec_lock.md color / font_family change across all generated SVGs |
For complete tool documentation, see ${SKILL_DIR}/scripts/README.md.
Windows note: if a
python3 ...command fails (common on python.org installs, which providepython.exebut notpython3.exe), rerun the same command withpythoninstead.
Template Index
| Index | Path | Purpose |
|---|---|---|
| Layout templates | ${SKILL_DIR}/templates/layouts/layouts_index.json |
Query available page layout templates |
| Brand presets | ${SKILL_DIR}/templates/brands/brands_index.json |
Query available brand identity presets (color / typography / logo / voice) |
| Visualization templates | ${SKILL_DIR}/templates/charts/charts_index.json |
Query available visualization SVG templates (charts, infographics, diagrams, frameworks) |
| Icon library | ${SKILL_DIR}/templates/icons/ |
See ${SKILL_DIR}/templates/icons/README.md; search icons on demand with ls templates/icons/<library>/ | grep <keyword> |
Standalone Workflows
| Workflow | Path | Purpose |
|---|---|---|
topic-research |
workflows/topic-research.md |
Pre-pipeline — gather web sources when the user supplies only a topic with no source files |
template-fill |
workflows/template-fill-pptx.md |
Give a native PPTX template deck plus source material; select fitting pages (a page may be reused for several output slides) and fill text back without SVG conversion |
beautify |
workflows/beautify-pptx.md |
Re-layout an existing PPTX through the SVG pipeline — preserve its text verbatim, inherit its palette/fonts as truth, redo only layout; mirror of template-fill |
create-template |
workflows/create-template.md |
Standalone layout template creation workflow |
create-brand |
workflows/create-brand.md |
Standalone brand-only template creation (identity preset; no SVG page roster) |
resume-execute |
workflows/resume-execute.md |
Phase B entry — resume execution in a fresh chat after Phase A (Step 1–5) completed in another session (split mode) |
verify-charts |
workflows/verify-charts.md |
Chart coordinate calibration — run after SVG generation if the deck contains data charts |
customize-animations |
workflows/customize-animations.md |
Object-level PPTX animation customization — run only when the user explicitly asks to tune animation order/effects/timing |
live-preview |
workflows/live-preview.md |
Browser-based live preview — auto-started during generation and re-enterable any time the user mentions "live preview", "preview", "看效果", or wants to click/select a slide element |
visual-review |
workflows/visual-review.md |
Per-page rubric-based visual self-check — run only when the user explicitly asks for a visual re-pass on the generated SVGs (between Executor and post-processing). Opt-in only; never invoked by the main pipeline. |
PPTX Route Boundary
When the user provides an existing .pptx, route by the role of the source deck:
| User intent | Route | Contract |
|---|---|---|
| Preserve the deck's page split, page order, and per-slide wording; improve layout / hierarchy / whitespace | beautify |
Source page count and order are 1:1; text and data values are frozen; visual identity is inherited after confirmation |
| Treat the deck as source material; rethink the story, merge / split / drop / reorder pages, or change page count | Main pipeline | ppt_to_md + PPTX intake provide content facts and candidates; Strategist may re-architect freely |
| Reuse the deck's native design with new material | template-fill |
Clone selected source slides and replace text / table / chart data directly in OOXML; no SVG generation |
| Harvest the deck as a reusable future template | create-template |
Build a template package, not a one-off generated deck |
Deciding axis (beautify vs main pipeline) — one question, one discriminator: is the source's page split a finished artifact to preserve, or a draft structure to overturn? The concrete discriminator is page count / order: if it changes at all — any split, merge, drop, or reorder — it is the main pipeline, never beautify. Beautify is strictly 1:1: same page count, same order, text verbatim, only layout / hierarchy / whitespace redone. Edge case made explicit: "keep all the content but split a crowded page so it reads better" still changes page count, so it is the main pipeline (re-pagination is re-architecture), not beautify.
Ambiguous requests such as "make this PPT more professional" or "optimize this deck" MUST be clarified with one question before routing: "Should the original page count/order and each slide's wording be preserved, or should the deck be treated as source material and restructured into a new story?" Preserve → beautify; restructure → main pipeline.
Workflow
Step 1: Source Content Processing
🚧 GATE: User has provided source material (PDF / DOCX / EPUB / URL / Markdown file / text description / conversation content — any form is acceptable).
No source content? When the user supplies only a topic name or requirements without any file or substantive description, run the
topic-researchworkflow first, then return here with its products as input.
When the user provides non-Markdown content, convert immediately:
| User Provides | Command |
|---|---|
| PDF file | python3 ${SKILL_DIR}/scripts/source_to_md/pdf_to_md.py <file> |
| DOCX / Word / Office document | python3 ${SKILL_DIR}/scripts/source_to_md/doc_to_md.py <file> |
| XLSX / XLSM / Excel workbook | python3 ${SKILL_DIR}/scripts/source_to_md/excel_to_md.py <file> |
| CSV / TSV | Read directly as plain-text table source |
| PPTX / PowerPoint deck | python3 ${SKILL_DIR}/scripts/source_to_md/ppt_to_md.py <file> for Markdown content; after Step 2 import-sources, standard PPTX intake is also written to <project>/analysis/ |
| EPUB / HTML / LaTeX / RST / other | python3 ${SKILL_DIR}/scripts/source_to_md/doc_to_md.py <file> |
| Web link | python3 ${SKILL_DIR}/scripts/source_to_md/web_to_md.py <URL> |
| WeChat / high-security site | python3 ${SKILL_DIR}/scripts/source_to_md/web_to_md.py <URL> (requires curl_cffi, included in requirements.txt) |
| Markdown | Read directly |
Office vector assets (EMF/WMF) from DOCX/PPTX sources:
doc_to_md.py/ppt_to_md.pyextract embedded Office vector images (.emf/.wmf) alongside bitmap images. Afterimport-sources, these land inimages/together withimage_manifest.jsonand are first-class assets in §VIII Image Resource List.Do NOT convert EMF/WMF to PNG. The PPT Master pipeline preserves them as external references (
finalize_svg.pyskips them) andsvg_to_pptx.pyembeds them as PPTX-native media viaimage/x-emf/image/x-wmfMIME — PowerPoint renders them at full vector fidelity. Converting via LibreOffice/Inkscape introduces CJK font substitution drift and rasterization loss; the original EMF/WMF is always higher fidelity than the converted PNG.Browser-based live preview cannot render EMF (will show blank) — this is expected; the PPTX output is the source of truth.
✅ Checkpoint — Confirm source content is ready, proceed to Step 2.
Step 2: Project Initialization
🚧 GATE: Step 1 complete; source content is ready (Markdown file, user-provided text, or requirements described in conversation are all valid).
python3 ${SKILL_DIR}/scripts/project_manager.py init <project_name> --format <format>
Format options: ppt169 (default), ppt43, xhs, story, etc. For the full format list, see references/canvas-formats.md.
Import source content (choose based on the situation):
| Situation | Action |
|---|---|
| Has source files (PDF/MD/etc.) | python3 ${SKILL_DIR}/scripts/project_manager.py import-sources <project_path> <source_files...> --move |
| User provided text directly in conversation | No import needed — content is already in conversation context; subsequent steps can reference it directly |
For PPTX sources, import-sources automatically runs the standard intake enrichment:
python3 ${SKILL_DIR}/scripts/pptx_intake.py <project_path>/sources/<source.pptx> -o <project_path>/analysis
For each PPTX it writes <stem>.identity.json (canvas, theme palette/fonts, observed usage) and <stem>.slide_library.json (text slots, geometry, native tables, native chart caches), and merges that deck's Strategist-facing digest into the single multi-deck index analysis/source_profile.json (decks[], one self-contained entry per source deck, with prefixed artifact pointers). In the main generation path these are source facts and recommendation candidates, not replica constraints; beautify and template-fill workflows decide separately which fields become locked constraints.
Multi-deck: several PPTX files may be imported into one main-pipeline project — each gets its own <stem>.* artifacts and a deck entry in source_profile.json. source_profile.json stays the single must-read index (one entry for a one-deck project, several for a combined-source project). Stems must be distinct; re-importing the same stem replaces that deck's entry. The beautify / template-fill workflows remain single-deck (1:1 to one chosen source deck) and read that deck's <stem>.* artifacts.
⚠️ MUST use
--move(not copy): all source files — Step 1's generated Markdown, original PDFs / MDs / images — go intosources/viaimport-sources --move. After execution they no longer exist at the original location. Intermediate artifacts (e.g.,_files/) are handled automatically.
✅ Checkpoint — Confirm project structure created successfully, sources/ contains all source files, converted materials are ready. Proceed to Step 3.
Step 3: Template Option
🚧 GATE: Step 2 complete; project directory structure is ready.
Default — free design. Proceed directly to Step 4. Do NOT query any *_index.json unless triggered. Do NOT ask the user. Do NOT proactively suggest, hint at, or fuzzy-match any template based on content, slug-like words, or vague style descriptions.
Template flow triggers ONLY on explicit directory paths supplied by the user in their initial message. The trigger rule is mechanical, not interpretive:
| User input contains | Step 3 action |
|---|---|
One or more explicit template directory paths (each resolves to a directory containing design_spec.md with kind: brand / kind: layout / kind: deck in its YAML frontmatter) |
Read each spec's kind, dispatch per the kind matrix below, fuse if multiple |
| Anything else — bare template names ("用 academic_defense"), style descriptions ("麦肯锡风格"), brand mentions ("招商银行风格"), vague intent ("想用个模板"), or silence | Skip Step 3, free design |
There is no slug matching, no name lookup, no fuzzy resolution. A name without a path does not trigger — the user must give a path the AI can cd into.
Style descriptions ("麦肯锡风格" / "Keynote 风" / "极简风" / etc.) never trigger Step 3. They flow into Strategist's Eight Confirmations as a style brief (color / typography / tone in confirmations e–g).
Bare names ("academic_defense", "招商银行", "anthropic") do NOT trigger Step 3 even if a matching directory exists in the library. The user must give a path. AI must not "helpfully" resolve a name to a path.
"What templates exist?" is out-of-band Q&A — answer by listing entries from
brands_index.json/layouts_index.json/decks_index.jsontogether with their paths. Listing alone does not advance the pipeline; the user must send a path back to trigger Step 3.
To create a new layout or deck, read
workflows/create-template.md. To create a new brand, readworkflows/create-brand.md.
Three template kinds
The architecture has three independent reference bundles. Full schema in docs/zh/templates-architecture.md. Summary:
| Kind | Physical dir | Contains | Frontmatter |
|---|---|---|---|
| brand | templates/brands/<id>/ |
identity-only segment: color / typography / logo / voice / icon style | kind: brand |
| layout | templates/layouts/<id>/ |
structure-only segment: canvas / page structure / page types / SVG roster | kind: layout |
| deck | templates/decks/<id>/ |
full replica: identity + structure + middle (template overview) segments | kind: deck |
Segment ownership (governs fusion override priority):
| Segment | Sections | Owner kind on fusion |
|---|---|---|
| Identity | Color Scheme / Typography / Logo / Voice & Tone / Icon Style | brand |
| Structure | Canvas / Page Structure / Page Types / SVG Roster | layout |
| Middle | Template Overview (use cases / design intent) | deck (no other kind writes this) |
Single-path dispatch
User path's kind |
Step 3 action |
|---|---|
kind: brand |
design_spec.md + non-image assets → <project>/templates/; logo / illustration / icon bitmaps → <project>/images/. Strategist locks identity segment as truth; structure stays free. |
kind: layout |
design_spec.md + SVG roster → <project>/templates/; any bitmap assets → <project>/images/. Strategist locks structure; identity decided in Eight Confirmations e–g. |
kind: deck |
design_spec.md + template SVGs → <project>/templates/; logos / backgrounds / other bitmaps → <project>/images/. Strategist locks all segments; Eight Confirmations narrows to deck-content fields (audience / page count / outline / tone tweaks). |
TEMPLATE_DIR=<user-supplied path>
# Bitmaps join the project's single runtime image pool (images/, referenced as
# ../images/); the spec + template SVGs + other non-image assets stay in
# templates/ as design reference the Strategist/Executor read but never render.
cp -r ${TEMPLATE_DIR}/* <project_path>/templates/
find <project_path>/templates -type f \( -iname '*.png' -o -iname '*.jpg' -o -iname '*.jpeg' -o -iname '*.gif' -o -iname '*.webp' -o -iname '*.bmp' \) -exec mv {} <project_path>/images/ \;
The same split applies to all three kinds — bitmaps always land in images/, the rest in templates/. The spec's kind field tells Strategist how to read the templates/ side; downstream code doesn't distinguish. (Template SVGs in templates/ are reference material only — the rendered pages live in svg_output/ and reference images via ../images/.)
Multi-path fusion
When the user gives two or more paths of different kinds, Step 3 fuses them into a single <project>/templates/design_spec.md. Default granularity is segment-level integer replacement — entire identity / structure / middle segments are taken from the highest-priority source for that segment, no implicit field-level mixing.
Override priority by segment:
| Combination | Identity from | Structure from | Middle from |
|---|---|---|---|
| brand only | brand | (free design) | (none) |
| layout only | (free design) | layout | (none) |
| deck only | deck | deck | deck |
| brand + layout | brand | layout | (none) |
| brand + deck | brand (overrides deck) | deck | deck |
| layout + deck | deck | layout (overrides deck) | deck |
| brand + layout + deck | brand | layout | deck |
Field-level micro-adjustment (e.g. "use anthropic brand but primary changed to #FF0000") is not part of Step 3 fusion — it flows into Strategist Eight Confirmations e–g as a normal user request.
Same-kind multiple paths — conflict resolution
When the user gives two paths of the same kind (e.g. brands/anthropic + brands/google), Step 3 surfaces a conflict prompt before fusing — like resolving a git merge conflict:
AI: 你给了两个 brand,检测到段级冲突:
- Color Scheme(Anthropic 橙红 vs Google 多色)
- Typography(Styrene/AnthropicSans vs GoogleSans/Roboto)
- Logo(Anthropic 标 vs Google 标)
- Voice & Tone(restrained vs friendly)
- Icon Style(stroke vs filled)
要 (a) 全部按 Anthropic / (b) 全部按 Google / (c) 逐段挑?
Rules:
- Default: no implicit ordering — every cross-source segment difference is reported as a conflict
- Only when the user picks
(c)does AI walk through each segment one by one - Field-level conflicts are out of scope — segment-level only
- Three or more same-kind paths are not supported — ask the user to converge to at most two
Fused spec provenance
When fusion happens (any multi-path case), the resulting <project>/templates/design_spec.md carries a provenance block immediately under its H1:
> **Fused from:**
> - deck: `templates/decks/招商银行/` (base)
> - brand: `templates/brands/anthropic/` (identity override)
> - layout: `templates/layouts/academic_defense/` (structure override)
> - conflicts resolved: Color Scheme from anthropic(user picked a)
Single-path Step 3 does not add provenance (the source is self-evident from the copied files).
✅ Checkpoint — Default path proceeds to Step 4 without user interaction. If the user supplied one or more explicit template paths, those have been dispatched (or fused) into <project_path>/templates/ before advancing.
Step 4: Strategist Phase (MANDATORY — cannot be skipped)
🚧 GATE: Step 3 complete; default free-design path taken, or (if triggered) template files copied into the project.
First, read the role definition:
Read references/strategist.md
⚠️ Mandatory gate: before writing
design_spec.md, Strategist MUSTread_file templates/design_spec_reference.mdand follow its full I–XI section structure. Seestrategist.mdSection 1.
<project_path>/analysis/ is the project's intermediate-analysis folder: the canonical home for machine-extracted source/asset facts — the PPTX intake bundle (source_profile.json index + per-deck <stem>.identity.json / <stem>.slide_library.json) and image_analysis.csv. It holds facts, not design contracts — design_spec.md / spec_lock.md stay at the project root. The MUST-read contract covers only the compact structured data files (.json / .csv); other artifacts that may live under analysis/ (e.g. a beautify source_svg_import/ vector reference package) are NOT bulk-read — they are read selectively only when a specific workflow step calls for them. Before the Eight Confirmations, Strategist MUST read the auto-extracted fact files already in analysis/ — currently source_profile.json (PPTX intake), when present. This file is the multi-deck index: read it once for the decks[] digests (canvas / chart / table entries per source deck), then open a specific deck's <stem>.identity.json / <stem>.slide_library.json only if you need its full raw facts. Use these entries as factual source context (format default + content facts); when several decks are present, synthesize across all of them. The source's palette / typography / visual identity are a reference, not a constraint: the main pipeline may inherit them where they fit the content and the confirmed style, or design fresh where they don't — the Strategist's judgment, never an obligation to either keep or discard. (Template-fill preserves the native source design by editing cloned slides directly; beautify defaults to the source identity but still follows the confirmed values; the main pipeline treats source identity as reference only and defaults to fresh design.) (image_analysis.csv lands later, at the image-analysis step below, and is the authoritative regenerated image-fact view there — re-derived from the live images/ folder, not a durable store.)
Channel ownership — read each fact once from its owning channel. In the main pipeline the content contract is the Markdown (sources/<stem>.md): text, tables, and chart data values all come from there (ppt_to_md now transcribes native chart data into Markdown tables). The analysis/ chart / table entries are a structural digest for outline decisions (which slides carried charts, type, series names) — not a second copy of the values; do NOT also pull chart values from <stem>.slide_library.json in the main pipeline. The <stem>.slide_library.json full structured data is owned by the direct-PPTX workflows: template-fill uses it as the native fill contract; beautify uses it for native chart / table data while keeping slide text from the Markdown.
Eight Confirmations (full template: templates/design_spec_reference.md):
⛔ BLOCKING: present the Eight Confirmations as a single bundled recommendation set and wait for explicit user confirmation or modification before outputting Design Specification & Content Outline. This is the single core confirmation point — once confirmed, all subsequent steps proceed automatically.
- Canvas format
- Page count range
- Target audience
- Style objective
- Color scheme
- Icon usage approach
- Typography plan, including formula rendering policy
- Image usage approach
Confirm UI Auto-Launch (Mandatory — default visual confirmation surface): by default the Eight Confirmations are presented through an interactive local page (color swatches, live font previews, candidate picks); the chat path is the always-valid fallback. Steps:
- Write the recommendations to
<project_path>/confirm_ui/recommendations.json(full schema + field mapping:scripts/docs/confirm_ui.md). Two kinds of field: enumerable (canvas / mode / visual_style / icons / formula policy / generation mode; plus image usage with a Custom path; plus AI source only when image usage may includeai) — the page lists common options fromconfirm_ui/static/catalogs.json, so you only name the recommended canonicalidin arecommendblock (canvas may be a catalog id likeppt169or a custom size/prose; style =mode+visual_style, two independent picks; icon ids are real libraries such astabler-outline, oremojifor system emoji; image usage usesai/web/provided/placeholder/none, or a custom prose plan when several sources must be combined; never write bare"custom"for image usage — write the actual mixed plan, e.g. "AI cover + user product assets + web industry images"; writeimage_ai_pathonly when recommendingimage_usage: "ai"or a custom plan that includes AI); generative (color, typography, generated-image style) — author ≥3 candidates each (creative recommendations always offer real choice, never a single silent option — same rule as strategist h.5; fewer than 3 only on the honest-shortfall exception, with a stated reason) (color: user-facing corepalettewith background/secondary_bg/primary/accent/secondary_accent/body_text; typography: CJK + Latin forheadingandbodywithcsspreview stacks, plusbody_sizeas the body baseline px; when recommending generated images,image_strategy.candidateswith rendering × palette combinations from strategist h.5).page_count/audience/content_divergenceare plain values (free text). Only open fields show a Custom box:canvas,mode,visual_style,icons,image_usage, and typography custom text. Closed fields (image_ai_path,formula_policy,generation_mode,refine_spec) stay finite.content_divergenceis a free-text field shown under audience in §c — the user states in their own words how closely to follow the source vs how freely to reshape it (blank = balanced; facts stay sourced at every level). Write it ascontent_divergence: { "value": "<prose or empty>" }. It is consumed by Strategist when authoring§IX, recorded indesign_spec.md §I, carries no page-count coupling, and is not written tospec_lock.md. Setlangto the page language; visible candidate text should matchlang, or provide bilingualname_zh/name_enandnote_zh/note_enfields. Reuse the same candidate thinking as strategist h.5. - Launch the page in the background and wait for the browser confirmation (the child server runs detached; the parent command returns after
result.jsonis freshly written). Run this command with a long tool timeout — 600000 ms — so the--wait(≈590 s budget) can complete:
Page opens atpython3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --daemon --waithttp://localhost:5050— the same port as the Step 6 live preview (they never run at once: this page shuts down at the end of Step 4, freeing the port). If another project already holds 5050, the launcher auto-advances to the next free port (5051, …) and serves this project there — read the actual URL from the launch log and report that. When the user clicks Confirm, the command exits 0 and Step 4 readsresult.jsonimmediately; do not require a second chat confirmation. Launch or wait failure is non-fatal: if it fails or times out (flask missing, port blocked, no GUI / remote / web host, browser never confirms in time), do NOT troubleshoot. The detached page stays open, so a slow user may confirm after the wait returns — therefore on any non-zero exit, re-check<project_path>/confirm_ui/result.jsononce (a freshstatus: confirmed) before dropping to the chat-summary fallback below. - Always also print the eight recommendations as a short summary in chat, with the URL. This keeps the chat fallback valid whether or not the browser opened. If the page never appears, the user simply confirms or edits in chat as before.
- This is the ⛔ BLOCKING wait. Preferred page path: the
--waitcommand returns after the page writes a fresh<project_path>/confirm_ui/result.json; immediately read that file and use its values. On a non-zero exit, re-checkresult.jsononce (per step 2) — a freshstatus: confirmedstill wins. Chat fallback path: only if no fresh result exists (page didn't open, wait timed out with no confirmation, or the user replies in chat with edits) take the chat values directly. Either path converges. A confirmedresult.jsonis an explicit user choice:generation_mode: "split"means split mode was chosen;refine_spec: truemeans the refine-spec workflow was chosen. - Close the confirm page (Mandatory cleanup — every path). Once you have the confirmed values (page or chat), shut the confirm server down before leaving Step 4 so it cannot keep holding port 5050 (which Step 6 live preview reuses):
This is idempotent and required regardless of whether Confirm was clicked: clicking Confirm already shuts the page down (this is then a no-op), but the chat-fallback path leaves the page running — without this cleanup it would block the live preview launch. Run it after reading the confirmation and before proceeding to Step 5.python3 ${SKILL_DIR}/scripts/confirm_ui/server.py <project_path> --shutdown
Honoring the confirmation (result.json is authoritative — Mandatory): the confirmed values override your own recommendations when you write design_spec.md / spec_lock.md. A user who changed any field changed it on purpose. In particular, map image_usage to §VIII Acquire Via (its value names differ from §h options — translate):
result.json.image_usage |
§VIII Acquire Via |
h.5 + Step 5 generation |
|---|---|---|
ai (or a custom plan that includes AI) |
ai rows |
Run h.5 (lock rendering + palette); Step 5 generates |
web |
web rows |
None |
provided |
user rows |
None — never generate |
placeholder |
placeholder rows |
None |
none |
no image rows (§h option A) | None |
When the confirmed image_usage is not ai (and the plan has no AI part), do NOT run h.5, do NOT write ai rows, and do NOT generate images in Step 5 — regardless of what you recommended. The same "confirmed value wins" rule applies to every field (color → §III, typography → §IV, etc.).
Opt-out: if the user has said they don't want the page (e.g. "不要网页" / "just confirm in chat" / "纯聊天确认"), skip the launch entirely (step 2) and present the Eight Confirmations in chat as before — steps 1, 3, 4 still apply (recommendations summary in chat; wait; take chat values).
The page is a confirmation surface only — Strategist still authors every recommendation; the page never generates content.
Mandatory — split-mode note (not a ninth confirmation): after listing the eight confirmation details, you MUST append exactly one short line (rendered in the user's language, prefixed with 💡) about generation mode. Pick the variant by qualitative read of Phase A signals — recommended page count, source-material bulk, whether topic-research ran with substantial web-fetch accumulation:
| Signal read | Line content |
|---|---|
| Heavy (long page count / bulky sources / heavy web-fetch accumulation) | State estimated page count and large source size; recommend switching to split mode after Step 5 — stop this chat, open a fresh window and input 继续生成 projects/<project_name> to enter Phase B (SVG generation + export); no response or "continue" = default continuous mode. |
| Normal (default) | State scale is moderate, default continuous mode generates in one go; if mid-way window switch is desired, input 继续生成 projects/<project_name> after Step 5 to switch to split mode. |
This line is required output every run — the user must always see the mode choice exists. Whether to act on it is the user's call. When the Confirm UI is used, this choice also appears as the in-page generation-mode toggle and is captured in result.json (generation_mode); the chat-summary fallback still prints this line.
Mandatory — spec-refinement note (not a ninth confirmation): after the split-mode line, you MUST append one short opt-in line (rendered in the user's language, prefixed with 💡) telling the user they may refine the spec first — Strategist will produce the full design spec, then stop for review/revision of any part of it before any generation, via the refine-spec workflow. Default is OFF: no request → the spec is written in one go and the pipeline auto-proceeds as usual. Only when the user explicitly asks in chat (e.g. "refine the spec first") or confirms refine_spec: true through Confirm UI does the refine-spec workflow take over after the Eight Confirmations. This line, like the split-mode line, is required output every run — the user must see the choice exists; whether to act on it is theirs. When the Confirm UI is used, this choice also appears as the in-page refine-spec toggle and is captured in result.json (refine_spec); the chat-summary fallback still prints this line.
Formula rendering policy lives inside item 7 (Typography plan):
| Policy | Behavior |
|---|---|
mixed (default) |
Strategist renders complex formula-worthy expressions as PNG assets; simple inline expressions remain editable text / Unicode |
render-all |
Strategist renders every formula-worthy expression as PNG assets |
text-only |
No formula rendering; formulas remain editable text / Unicode |
After the Eight Confirmations are approved and before outputting design_spec.md / spec_lock.md, if the confirmed formula policy is mixed or render-all and the content contains formula-worthy expressions, Strategist MUST:
- Identify explicit LaTeX and any source expressions that should be faithfully structured as formulas.
- Write
<project_path>/images/formula_manifest.jsonwith only the formulas selected for rendering. - Run:
python3 ${SKILL_DIR}/scripts/latex_render.py <project_path> - Include the rendered formula PNGs as
Acquire Via: formula,Status: Rendered,Type: Latex Formularows indesign_spec.md §VIII Image Resource List; also list them inspec_lock.md imageswith| no-crop.
The formula renderer uses a provider fallback chain by default: codecogs,quicklatex,mathpad,wikimedia. The first three are color-aware; Wikimedia is an availability fallback. Formula PNGs are transparent by default: manifest background is the temporary render matte and transparency-removal reference, not a retained final background unless transparent: false is set for that item. Do not scan spec_lock.md for $...$ or $$...$$. Dollar-delimited math in source material is only a signal for Strategist; the renderer consumes the explicit manifest.
If the user provided images or formula PNGs were rendered, run analysis before outputting the design spec. It writes analysis/image_analysis.csv — the authoritative regenerated image-fact view in the analysis/ folder, which MUST be read before authoring §VIII:
python3 ${SKILL_DIR}/scripts/analyze_images.py <project_path>/images
🔁 Image facts are regenerated on demand, never a durable store.
images/is a live working folder — pictures are extracted from the source at import, the user may drop or replace files at any time, and Step 5 writes web/AI images into it. The single source of truth is therefore the current contents ofimages/, andanalysis/image_analysis.csvis a regenerated view of it, not a fact to keep in sync. Re-runanalyze_images.py <project_path>/imagesimmediately before any step that reads image facts so the view reflects the live folder: before the §h image-usage recommendation (see strategist.md §h), here before authoring §VIII, after Step 5 acquisition (so web/AI files join the view), and again any time the user says they added or replaced images. This is the staleness strategy — re-derive on use, no cache to invalidate.
⚠️ Image handling: NEVER directly read / open / view image files (
.jpg,.png, etc.). All image info comes fromanalyze_images.pyoutput (analysis/image_analysis.csv) or the Design Spec's Image Resource List.
Output:
<project_path>/design_spec.md— human-readable design narrative<project_path>/spec_lock.md— machine-readable execution contract (skeleton: `templates/spec_lock_reference.
…(truncated)