rp-import-codegen
Generate source readers, transforms, setup/import entrypoints, and thin Wix write specs from approved migration artifacts.
Purpose
This skill turns the schema, mapping, and setup decisions into implementation files under the active migration project.
Required inputs
migrations/<project>/source-schema.jsonmigrations/<project>/mapping/mapping-plan.jsonmigrations/<project>/setup/setup-plan.jsonmigrations/<project>/setup/setup-requirements.jsonwhen setup affects write paths
Prefer the machine-readable artifacts above. Markdown review files are secondary renderings for humans, not the primary codegen contract.
Source read contract
Generating a correct reader requires platform-specific knowledge — auth model,
pagination, rate limits, and REST quirks. That knowledge lives in the matching source
adapter skill, not here, so this skill never names a platform. Resolve the adapter from
the platform field in source-schema.json via the naming convention rp-source-<platform>
(e.g. platform: "wordpress" → rp-source-wordpress) and read its "Read contract"
section. The operational facts should already be recorded in source-profile.md; use the
adapter to fill any gaps rather than guessing. rp-execute-import runs the reader you
generate and stays platform-agnostic — so the platform specifics must be baked into this
generated code, not deferred to execution.
Adding a new source platform therefore requires no change to this skill: a new
rp-source-<platform> adapter is enough.
Reading plugin-provided entities
Entities that came from a source plugin carry extra sourceMeta the generated reader must
respect. These are read mechanics, not mapping decisions:
origin: "embedded"— the records live inside a property of a parent record (propertyPath) on an already-fetched route (embeddedIn). Extract them from the parent fetch. Do not issue a second request per parent; some of these plugins have no route of their own at all.requiresParent— a sub-collection route parameterized by the parent id (.../{id}/notes). Read the parent first and iterate; order this after the parent entity in the dependency graph.channel: "plugin-rest-child"— a sub-resource on a CORE collection this plugin does not itself profile (e.g. WooCommerce orders), so it cannot userequiresParent(which points at another entity declared in the same profile).routecarries a{parentId}template (/wc/v3/orders/{parentId}/notes);parentRoutenames the collection route to iterate for real parent ids. Unlike discovery'ssampleChildEntities(which checks only a representative few parents), the generated reader must iterate every parent record fromparentRouteand fetch its child route — that full iteration is what actually migrates the data; the discovery-time sample only confirmed the shape is real.context— when the adapter says view and edit contexts return different data, request the one the entity declares. Requesting only one context can silently halve the result set.- Per-plugin credentials — some plugin APIs use their own key system rather than the platform credential. Load them from project-local config like any other secret, and fail loudly with the missing key name rather than falling back to unauthenticated reads.
recognized: false(derived entity) — no plugin-specific read contract exists; read it as an ordinarywp/v2-style collection using the shared transport. Do not invent plugin-specific handling for it.requestMethod/requestBody— the entity's real read path is not a plain GET collection; it is a declared non-GET method with a JSON body template (spec 0044). The generated reader must issue that exact method/body, not fall back to a GET.- A body value of the form
"$SAMPLED_IDS:<route>"names another route whose record count is UNKNOWN and possibly unbounded (e.g. an entire product catalog) — never resolve it against one fixed id list fetched in a single request, at discovery time or at execution time, because<route>may hold far more ids than fit in one page or one request body. The generated reader must instead: (1) paginate<route>to exhaustion (every page, not a capped sample); (2) split the collected ids into bounded batches (mirrorlib/sampled-ids-batch.js's default of 50 ids per batch — do not send an unbounded id list in one request); (3) issue one request per batch against this entity's own route, with the placeholder resolved to just that batch's ids; (4) normalize each batch's RAW, unmodified response (envelope resolution, thenresponseFragmentGroupSizereassembly below — never pre-coerce a non-array raw response to an empty list before normalizing it, since an enveloped response is legitimately an object, not an array, until unwrapped) and merge the batches, deduplicating byrecordKeyField(defaultid) — the same record can otherwise appear once per batch it happened to match; (5) stop and record the run as an explicit failure/deferred outcome on the first page or batch that cannot be read OR whose normalized result is not a valid array shape (a batch that doesn't match the declared envelope/fragment contract must defer the whole entity, not be counted as a legitimate empty batch — that would produce an exact-looking but false zero) — never emit a partial/undercounted or falsely-exact result as if it were complete.rp-source-wordpress/lib/sampled-ids-batch.jsimplements this exact algorithm (collectAllIds,queryInBatches) forwp-discovery.jsitself; a generated reader for another platform must reproduce the same five steps, not invent a different policy.
- A body value of the form
responseFragmentGroupSize— the response arrives as N separate flat single-key array entries per logical record instead of one object (a plugin bug some source APIs have). The generated reader must chunk the raw array into groups of N and merge each group withObject.assigninto one record before any mapping/transform runs or any count is emitted — per batch when combined with a$SAMPLED_IDSrequestBody above, mirroringwp-discovery.js'sinspectEntity/normalizeResponseRecordshandling of the same field.
CSV sources (platform: "csv")
When source-schema.json.platform === "csv", generate a file reader instead of an HTTP
reader and read rp-source-csv → "Read contract". The differences that matter:
- Vendor
lib/csv-parse.jsintosrc/lib/and import it, exactly as a WordPress reader vendorswp-http.js. Do not re-emit parsing: the sampler and the reader must agree on what the file contains. There is no auth, no pagination, and no rate limiting to generate. - Grouping is the reader's core job. Replay
sourceMeta.sourceFiles[].layout: withcontinuation: "blank-key"a blank key extends the current group; with asectionedpattern the discriminator column routes parent vs child rows andparentRefColumnresolves a child to its parent (it may holdid:123or a SKU, and the child is not guaranteed to follow its parent). - Iterate the file set by role, treating
sourceFiles[].partOfentries as continuations of the same logical stream and honoring each part's own column order. - Materialize
column-valuesentities: collect the distinct values of the source column and their ancestors into their own entity file, emitted depth-ascending so a parent category exists before its child, plus the linking relation. Do not invent Wix ids at extract time — theImportCrosswalkresolves them at import time. - Apply
sourceMeta.dialect.emptyPolicythrough the sharedcoerceEmptyhelper so a required Wix field is never fed an empty string the source did not have. - CSV values are plain text; write them as-is with no entity-decoding step.
Target write contract
Symmetrically, do not re-derive the Wix write surface here. It is identical for
every migration and is pre-verified in the rp-target-wix internal resource (see
CONVENTIONS.md), which ships shared Wix runtime code: verified request builders,
executors, and reusable execution logic for retries, throttling, checkpoints, and
reporting. Vendor copies of the shared runtime modules from rp-target-wix into the
project (like the source transport) and generate thin project-specific write specs and
transforms that call those shared primitives/runtime functions. Never hand-emit Wix
endpoints/bodies inline —
that path repeatedly shipped wrong shapes (lowercase Ricos plugin enums, {tag:{…}}
tag bodies, media.wixMedia.image.id featured images) that only failed at execution.
Validate by real call, not by doc example. MCP doc checks confirm an endpoint
exists; they do not confirm the request shape works — public examples have been
wrong (e.g. Ricos plugin enums shown lowercase that 400). Trust a Wix request shape only once a real call (or
tests/target-wix/contract-test.js in live mode, run from the repo root) has succeeded. Treat the
adapter's // VERIFIED: shapes as the source of truth over any docs example.
Adding a new source platform requires no change here, and the Wix surface change for
all migrations is a one-place edit in rp-target-wix (caught by its contract test),
not a per-project regeneration.
Workflow
- Read the discovery, mapping, and setup artifacts, plus the source adapter's read
contract and the
rp-target-wixwrite contract. If a mapping entity includestargetRef, validate its selected writer and reliability againstrp-target-wix/scripts/domain-knowledge.js summarize-entities; do not remap source entities independently. - Generate machine-readable execution artifacts:
execution/execution-manifest.jsonexecution/llm-handoff.json- when
SAFE_MODE=trueorDRY_RUN=true,execution/review/code-safety-review.md - optional supporting execution subplans only when the runtime truly needs them
- Generate setup code that can verify/provision Wix prerequisites by executing the machine setup artifacts through the shared setup runtime.
- Generate source reader code that can enumerate and fetch source entities. If the adapter
ships a shared transport module (auth, pagination, throttling, retries), vendor a copy of
it into the project (e.g.
migrations/<project>/src/lib/) and import from it instead of re-emitting that plumbing — the reader should hold only per-project orchestration. The reader must extract source records to durable files on disk; it must not require the whole source dataset to live in memory before import begins. - Generate transform code that maps source records into Wix-shaped objects — thin glue over
wix-build.js, which already owns slug sanitizing and the money/price/variant rules; see "Never hand-write slug sanitizing" below. Wix Stores products: Catalog V3 only. Catalog V1 is not supported by this workflow — there is nothing to generate for it. The only destination a migration ever writes to is aV3_CATALOGsite (guaranteed at provisioning — see0079-catalog-v3-guaranteed-retire-v1-gate.md), so generated code may assume V3 unconditionally: no version detection, no branch, no V1 shape, no V1 fallback for a failing V3 write. A V1 site (only reachable on a pre-existing site this run did not create) is a blocker the run halts on, never a case codegen handles. Concretely: never emit Catalog V1-style top-levelprice,sku, or variant inventory fields. For simple products, emit one variant undervariantsInfo.variants[]with variant-levelprice,sku, and physical properties. When the source product is subscription-based and the source payload exposes explicit recurring cadence in structured product data, emit nativeproduct.subscriptionDetailswith at leastallowOneTimePurchasesand onesubscriptions[]entry carryingtitle,description,frequency,interval, andautoRenewal. Do not emit placeholder-only metadata or skip those products on create when that cadence can be inferred deterministically; live create coverage for this shape was verified on July 26, 2026. Product HTML descriptions land inplainDescription, which Wix converts to rich content server-side — it is HTML, not a plain-text flattening, so there is no fidelity loss to accept. Do NOT call the Ricos conversion endpoint on the product path: it costs one HTTP round-trip per product ahead of a bulk create, and that burst is what the endpoint throttles with a 403. Two traps:plainDescriptionis silently ignored whendescriptionis also set (set exactly one), and it is capped at 16,000 characters — longer bodies need truncation or an info section, recorded inmapping-gaps.json. Ricos conversion remains correct for blog posts, whererichContentreally is a Ricos document. Media scope should be relationship-driven by default: generate imports only for media referenced by mapped entities, not for the entire unattached source library. When the target adapter says an entity can ingest external URLs directly (for example Wix Stores product media), generate that entity-native path instead of routing those files through the slower generic Media Manager import flow. - Generate thin project-specific write specs and import orchestration code that pass
those Wix-shaped objects into the shared
rp-target-wixruntime. Generated code should describe what to write and in what order, not how to implement retries, throttling, checkpointing, or audit logging. Prefer the selected entity'spreferredWrite; if it is notverified-live, surface an execution warning before consent. Native mappings must have a knownwriterId, a direct REST plan that callsnotifyMissingWriter, or an explicit unsupported/gap fallback. - Generate or wire the deterministic execution-state preparation step before any import
write. The generated runner should call the shared
execution-statepreparation contract, or the execution instructions should runskills/wix-replatform/scripts/execution-state-prepare.jsbefore the generated import entrypoint. - Generate the runnable setup/extraction/import entrypoints (see below) — the artifacts
rp-execute-setupandrp-execute-importactually run. This is required, not optional. Then run the sample-preview gate (see below) before the execution-plan approval. - Generate
execution/review/import-plan.mdplusexecution/review/import-plan.freshness.jsonusingskills/wix-replatform/scripts/artifact-freshness.js write .... The freshness metadata must include hashes forsource-schema.json,mapping/mapping-plan.json,setup/setup-verification.json, generated import code, and the target contract ledger revision. If codegen follows a newly promoted write contract, regenerated code and this metadata must reflect that promotion. - Document any manual code follow-up still required.
- Generate any post-import remediation helpers that are required to reach the accepted mapping fidelity when the source capture cannot express the relationship inline during the first create pass.
Post-codegen code-safety review checkpoint
When either SAFE_MODE=true or DRY_RUN=true for the active project, codegen must
produce a mandatory review artifact before the execution approval gate:
migrations/<project>/execution/review/code-safety-review.md
This review is performed by the agent, not delegated to the user. It verifies the generated code itself. It must check:
- every generated write path that can carry email/phone data passes
safeModeOptionsinto the shared Wix runtime or direct REST wrapper - every shared/native writer reached by generated code actually consumes those
safeModeOptionsand routes the request body through the shared sanitizer rather than silently ignoring the parameter - no generated writer path silently bypasses safe mode because of a missing function parameter or a direct call that skips the shared sanitizer
- dry-run uses the same generated write code path as live import, with Wix calls skipped
only at the shared
wix.sendboundary - dry-run reporting does not label would-send placeholders as live-created/imported site objects
- mapping-declared
safeModeReplacements[]are reflected in the generated write specs and runner wiring, and any entity that still carries outbound contact data without matching replacement-path coverage is treated as a review failure rather than left for the user to reason about manually
If the review finds a gap, execution approval must remain pending until the generated code is corrected and the review artifact is regenerated with a passing verdict. The user's role at this checkpoint is final go/no-go approval after the agent has already completed the review and surfaced the findings.
Codegen boundary
This skill owns the project-specific layer only.
It should generate:
- setup plan renderings and setup runner wiring
- source readers
- transforms
- per-entity write specs
- import ordering and dependency wiring
- project-local config loading
- execution review artifacts
- post-codegen code-safety review artifacts when safe mode or dry-run is enabled
It must not regenerate for each migration:
- raw Wix auth/client plumbing
- generic retry loops
- throttling behavior
- audit log shape
- compact execution report shape
- checkpoint store mechanics
- generic bulk-write orchestration
Those behaviors belong in rp-target-wix.
Final handoff expectations
The generated plan and downstream execution path must make the final state legible to the user. Treat these as required handoff details, not optional niceties:
- Surface the dashboard URL for the destination site.
- Surface the editor URL only when editor work was actually performed or the next required step is explicitly in the editor.
- Surface the preview URL only when public route/site verification is relevant to the completed work.
- Distinguish clearly between:
- catalog/data imported
- website/homepage built
- Do not imply that a successful Stores import means the site's homepage or full website experience exists. A migration can finish with valid catalog/product/cart/checkout routes while the homepage is still blank or unbuilt.
- If Shopify quick mode relies on public collection feeds such as
/collections/{handle}/products.jsonto recover category membership, bake that into the generated extractor/importer or emit an explicit remediation helper and call it out inexecution/review/import-plan.md.
Runnable setup/extraction/import entrypoints — the artifacts are the execution path (required)
rp-execute-setup and rp-execute-import run the migration by executing these
artifacts, never by the agent hand-issuing Wix MCP calls. So codegen must emit real,
runnable entrypoints:
src/setup/run-setup.jsor equivalent setup entrypoint — reads the machine setup artifacts and executes them through the shared setup runtime.src/extract/run-extract.jsor equivalent reader entrypoint — reads the source APIs and writes durable extracted files under the project. Extraction is a separate step from destination writes, so large migrations can be resumed or re-imported without re-reading the whole source.src/import/run-import.jsor equivalent import entrypoint — reads the extracted files, applies transforms, and writes to Wix in dependency order through the vendored shared write runtime. It must:- load project-local config files, then process env, for all expected env-style values (never hardcoded; fail fast if absent — do not fall back to any agent/MCP auth),
- run the notification-mute preflight before the first entity write whenever mute is in effect (see "Notification-mute preflight" below),
- consume extracted source files from disk rather than materializing the entire source in memory,
- apply idempotent dedupe keyed by source ID, using either a client-controlled source-id
field on the target or the authoritative local
state/crosswalk/crosswalk.ndjsondurablesourceId -> targetIdcrosswalk for native Wix entities with server-assigned IDs, reading existing target state under the rules in "Reading existing target state" below — an empty read is not an empty site, - honor
DRY_RUN=trueand--dry-runfor the safe-validation pass;--dry-runtakes precedence over config and enables the shared Wix runtime's dry-run mode. --samplemay be supported as a narrower dry-run-style validation mode, but it must not replace--dry-runas the primary safe-validation control.- honor deterministic selective resume flags:
--entity <entity>,--source-type <subtype>,--missing-only,--failed-only, and--deferred-only.--missing-only,--failed-only, and--deferred-onlyare mutually exclusive. These flags must drive the same import path as a full run after selecting a stable record set; do not generate one-off recovery drivers. - emit a stable
runIdat process start and include it in every audit event, dry-run request capture, placeholder crosswalk row, and summary artifact. - emit
execution/live-import-summary.jsonandexecution/completion-report.jsonthrough the shared completion-report runtime, with entity/subtype completeness counters for extracted, in-scope, attempted, imported, already present by crosswalk, deferred, failed, and skipped-out-of-scope records. - stop before dependent phases when required upstream entities fail. For example, do not create products after product-category failures unless the execution plan explicitly marks category assignments as non-blocking.
- pass
safeModeOptionsto every relevant writer path, including direct REST fallback paths and unverified native writers that can carry contact values.
- The dry-run is the same import code path with Wix calls skipped at the shared Wix
boundary — not a separate driver. Generated code must construct the same Wix client,
call the same writer helpers, run
SAFE_MODEsanitization, and pass request metadata (phase,operation,entity,sourceId,verification, and when neededresponseShape) intowix.sendso the runtime can capture would-send requests and return live-compatible dry-run placeholders.
Extraction format requirements:
- write extracted data to project-local files, not process memory
- chunk by entity and page/batch so a large source does not become one giant file
- write a manifest that lets the import step discover which entity files exist and in what order to consume them
- make the extracted files deterministic enough for resume, replay, and debugging
Notification-mute preflight (required)
When mute is in effect — always for WIX_SITE_STRATEGY=new (unconditional,
regardless of WIX_MUTE_NOTIFICATIONS), and for existing sites only when
WIX_MUTE_NOTIFICATIONS=on — the generated import script must contain a mandatory
preflight step, before the first entity write, that asserts the site is muted:
- Primary:
getSiteMuteState(wix)(rp-target-wix, VERIFIED 2026-08-04) returnsmuted: true. - Fallback (only if the status read is unavailable): an idempotent
muteSiteNotifications(wix, { reason })call — passing the same project-identifying reason as setup (RePlatform migration — <project>), because a re-mute overwrites the recorded reason (last caller wins).
If the preflight call fails or reads muted: false and the re-mute fails, the script
aborts before any write with a clear error routing back to setup — exit non-zero, no
degraded mode, no --skip-mute-preflight-style flag, no warning-and-continue. For
new-site projects the preflight is emitted unconditionally and never appears as a
skippable option; for existing-site projects with WIX_MUTE_NOTIFICATIONS=off (the
default), no preflight is emitted. The preflight logs its re-verification and outcome
into the run's execution/progress log (keyed by runId) — terminal reports read this
recorded state, never strategy/config inference. In dry-run mode the preflight follows
the shared dry-run contract like any other Wix call: the intent is captured and the
state read is skipped (stateKnown: false), not asserted as muted. The generated script
must never call unmuteSiteNotifications — unmute is an explicit owner request outside
the import path. A preflight mute-verification failure is recorded through the standard
rp-telemetry recorder as an error event (with error_code) like any other run
event — no new telemetry surface.
Use the target's BULK write path (required)
A generated importer must write through the target's bulk endpoint whenever one exists. Per-record creates are acceptable only when the target has no bulk equivalent, or for a deliberate single-record contract probe. This is not an optimization to add later: at 1000 products, per-record creates are 1000 round trips where bulk is ~11, and the latency difference is the difference between a minute and half an hour.
Derive the batch shape from the endpoint's own limits, all of them at once. Bulk
endpoints routinely cap several dimensions simultaneously and exceeding any one rejects
the entire request — Wix bulk product create caps products (100), variants (1000), options
(100), modifiers (100) and infoSections (100) per request, so with 2 options per product the
options cap binds at 50 products, not 100. Batch with ndjson.readBatchesByLimits and the
limits/cost helpers the target adapter exports (BULK_PRODUCT_LIMITS,
storesProductBulkCost); never batch on record count alone.
Three properties of bulk responses that generated code must handle explicitly, because each one silently corrupts a report if ignored:
- Bulk is not atomic. A
200can contain per-item failures. Walk everyresults[]entry; never infer success from the HTTP status. - Correlate by the response's own index field (
itemMetadata.originalIndexfor Wix), not by response position, and verify that every input is accounted for. A mis-correlated result crosswalks the wrong target id onto a source record. - Count the "undetailed failures" bucket. Servers drop failure detail past a threshold; those are still failures and must appear in the completion report.
Dedupe before building the batch — a skipped record must never reach the API — and record the crosswalk per successful item, not per batch, so an interrupted run resumes correctly.
Record streams are NDJSON, single documents are JSON (required)
Every file that holds a stream of records must be newline-delimited JSON (.ndjson), one
record per line — never a { "records": [ … ] } array. Vendor
rp-target-wix/lib/ndjson.js into the project (like wix-writers.js) and use it; do not
hand-roll line splitting, which gets chunk boundaries and CRLF wrong.
This applies to:
data/source-extract/<entity>.ndjson— the extractor's output- the
sourceId -> targetIdcrosswalk - audit logs (already NDJSON)
It exists because every downstream stage does the same three things with these files, and an array is the wrong shape for all of them:
- scan —
countRecordscounts lines; nothing is parsed and nothing is held in memory. A JSON array must be fully parsed to be counted. - batch — a bulk endpoint's page is
readBatches(file, 100). UsereadBatchesBywhen the target caps more than one dimension: Wix bulk product create allows ≤100 products AND ≤1000 variants per request, which is{ maxCount: 100, maxCost: 1000, cost: p => p.variantsInfo.variants.length }. - cursor / resume —
readSlice(file, { offset, limit })skips what is already done without rebuilding it into objects.
Two more properties that matter in practice: a producer can append records as it finds them instead of buffering the whole entity, and an interrupted write leaves a valid readable prefix — a truncated JSON array is unparseable, so a crash mid-extract loses everything.
Do not line-delimit single documents. A manifest, mapping-plan.json,
decisions.json, execution-manifest.json, preview-result.json or
completion-report.json is one object; it stays .json. NDJSON buys nothing there and makes
it unreadable.
Generated importers must stream these files (for await (const batch of readBatches(...))), not readAllRecords them. readAllRecords is an escape hatch for
genuinely small streams (a six-record category list) and is named to make its misuse on a
large stream obvious.
Projects generated before this rule can be moved forward with
convertLegacyJsonFile(jsonPath, ndjsonPath) rather than re-extracting.
Reading existing target state (required)
An idempotent importer has to know what is already on the site before it writes. Every bug in this section is the same bug: a read that returns nothing looks exactly like a site that contains nothing, and nothing-on-the-site is the branch that writes. None of them throw, so none of them show up in a dry-run.
The adapter's query* executors already return the array. queryStoresCategories,
queryStoresProducts, queryContacts, queryCoupons and queryOrders unwrap the response
before returning it, so the value is categories / products / etc. Generated code must
use it directly:
const existing = await W.queryStoresCategories(wix); // an array
const existing = (await W.queryStoresCategories(wix)).categories; // WRONG → undefined → []
The second form is what shipped, and reading .categories off an array yields undefined,
which the usual || [] turns into an empty array. It silently disabled a category dedupe
index, silently disabled a product name-match safety net, and made a setup verification
report 0 categories on a site that had 25 — all without one error line.
Never cursor-page through those executors. Unwrapping discards pagingMetadata, so the
cursor a loop needs is already gone; a loop built on them cannot advance past page one, and
reading .pagingMetadata off the returned array is the same undefined as above. Prefer the
adapter's sweep primitives, which own the loop and the failure semantics:
queryAllStoresCategories(wix),queryAllStoresProducts(wix),queryAllDataItems(...)
When a sweep is needed for an entity that has no queryAll* primitive yet, generate the loop
against the raw response and add the primitive to rp-target-wix rather than leaving the
loop in project code:
let cursor = null;
do {
const body = cursor ? { cursorPaging: { limit: 100, cursor } } : { cursorPaging: { limit: 100 } };
const response = await wix.send(W.buildQueryStoresProductsRequest(body)); // raw, not the executor
for (const p of response.products || []) { /* index it */ }
cursor = (response.pagingMetadata && response.pagingMetadata.cursors && response.pagingMetadata.cursors.next) || null;
} while (cursor);
An incomplete sweep must throw, not fall through. If any page fails, or the loop hits its page ceiling with a cursor still outstanding, the generated code must abort the import with a message naming the sweep. It must not continue with the partial index, and must not treat "the sweep failed" as "nothing exists" — that is precisely the state in which a re-run re-creates the entire catalog it already imported. The single most expensive failure in this whole pipeline is a duplicate import, and it arrives through an empty net.
A match must be ADOPTED into the crosswalk, not skipped. When a safety net (name match,
slug match, source-id field) finds that the target entity already exists, record it in the
crosswalk with its target id and revision and count it as reused. A bare continue that
skips the write without recording the id looks correct — nothing is duplicated — but every
later phase that resolves ids from the crosswalk then silently drops the record. Concretely:
the product was already on the site, so it was skipped, so it had no crosswalk row, so the
category-link phase could not resolve its id and it ended up in no category at all. Index the
net as name -> { id, revision }, not as a Set of names, so the id is available to adopt.
An ambiguous match is the one case that must not be adopted: if the source key is not unique (e.g. four source products share a title), the net cannot tell which existing entity corresponds to which source record. Import rather than guess, and report the ambiguity in the completion report.
Never hand-write slug sanitizing (required)
A source handle is not already a valid Wix slug. Wix rejects anything outside [a-z0-9-],
and because slug validation happens before the batch is applied, one bad slug fails the
entire bulk request — 100 products lost for one character. Shopify mints underscores from
decimal titles ("pH 5.5" → ph-5_5), so this is routine input, not an edge case.
It is already solved: wix-build.js exports toWixSlug and applies it automatically through the
coerce: 'slug' rule on product.slug in wix-target-spec.js. Generated transforms that call
the build layer (see the src/import/transforms/ note under "File targets") get it for free and
must not re-derive it — a per-project copy is how the underscore bug reached a live site in the
first place.
Two properties to preserve when a generated transform sets a slug explicitly:
- Sanitize in the build layer, not the writer, and keep both values. URL preservation needs
the original
sourceSlugalongside theplannedTargetSlugactually derived from it (see the URL preservation rules under "Codegen rules"), which a silent rewrite inside the writer would falsify.normalizeStoresProductV3therefore passes a slug through untouched. toWixSlugthrows when a value sanitizes to empty (an all-non-latin title, for example). That is a signal to supply a deterministic fallback — the record's source id — not to omit the slug and let Wix derive one, which breaks URL preservation with no trace.
Sample-preview gate
Between the extractor being generated and the execution-plan approval, show the user what their data actually became. The mapping review checkpoint validates intent; this validates structure — how source rows turned into entities — before any full run.
Required whenever the source cannot be read back from a live API — in particular every
platform: "csv" run, where a misread layout silently produces the wrong entity split.
- Generate the extractor first (
src/extract/run-extract.js). - Run it in sample mode (
--sample) to materialize a smalldata/source-extract/slice. - Write two artifacts under
migrations/<project>/preview/:preview-summary.md— a short human-readable structure preview: how rows became grouped entities (e.g. one product with its variants and images), per-entity record counts, and which columns landed in which Wix fields.preview-result.json—{ "status": "pending", "decidedAt": null, "decidedBy": null, "entityCounts": {...}, "warnings": [] }.
- Pause and ask the user to validate the structure. On accept, set
status: "accepted"withdecidedAt/decidedBy; on reject, setstatus: "rejected"and setapprovals.mapping.statusback topendingso the router returns torp-mapper. In explicit user-requested1-click mode(automationMode=one_click,source=user), validate the structure automatically from the preview artifacts, set the result toacceptedwithdecidedBy: "agent"when it passes, and continue without a user pause. Otherwise every CSV run pauses here, including a run whoseone_clickvalue was inferred or copied with a non-user decision source. - Record the artifacts on the codegen checkpoint (
checkpoints.codegen.artifactRefs+lastCompletedStep: "codegen.sample-preview").
This is a codegen sub-gate, not an orchestration phase: it adds no state to
orchestration-state.js. While preview-result.json says pending, the router keeps routing
back to this skill instead of advancing to the execution-plan gate — that routing behavior is
the enforcement, so the artifact must be written honestly. The gate rides on the existing
extract→import split and does not change how import writes to Wix.
Project-local config files
Generated code should treat migrations/<project>/config/ as the canonical home for all
values that are otherwise expected as environment variables. Use simple .env syntax and
load these files before reading config:
config/wix.envalways exists. DefaultDRY_RUN=false, except when the user has explicitly asked to start, create, prepare, or run the migration in dry-run mode; in that case scaffold or preserveDRY_RUN=true:WIX_SITE_STRATEGY= WIX_SITE_ID= WIX_AUTH_TOKEN= WIX_MUTE_NOTIFICATIONS= DRY_RUN=false SAFE_MODE=true SAFE_MODE_PHONE_NUMBER=+972 50 0000000config/source.<platform>.envexists after the source platform is known. For WordPress:WP_BASE_URL= WP_USERNAME= WP_APPLICATION_PASSWORD= WP_MEDIA_URL_REWRITE_FROM= WP_MEDIA_URL_REWRITE_TO= WC_CONSUMER_KEY= WC_CONSUMER_SECRET=For CSV sources this is
config/source.csv.env, which is not secret-bearing — all keys are optional hints (CSV_INPUT_ROOT,CSV_DELIMITER,CSV_ENCODING,CSV_VENDOR,CSV_MEDIA_URL_REWRITE_FROM,CSV_MEDIA_URL_REWRITE_TO). Generated readers must resolve input paths againstCSV_INPUT_ROOT(falling back to the project directory) rather than baking absolute paths into generated code.
Codegen rules:
- Generate a small dependency-free config loader in the runnable entrypoint or
src/lib/. - Load
config/wix.envand the selected source config before constructing source/Wix clients. - Default
DRY_RUNto disabled. Treattrue,1,yes, andonas enabled andfalse,0,no, andoffas disabled.--dry-runmust override config and enable dry-run.--no-dry-runmay be supported to overrideDRY_RUN=true. - When scaffolding a project that is in dry-run mode, generated review artifacts must say that leaving dry-run later requires explicit user approval for any step other than new-site creation, and that such overrides should be avoided when a dry-run or report is sufficient.
- Preserve an explicit
DRY_RUN=truefrom project config when regenerating code or config. Do not reset it tofalseduring later codegen passes. - Default safe mode to enabled when
SAFE_MODEis missing or blank. HonorSAFE_MODE=falsewhen the user set it before mapping: do not requiresafeModeReplacements[], do not replace contact values, and do not write safe-mode email recovery rows. - When safe mode is enabled, require
SAFE_MODE_PHONE_NUMBER, default missing/blank values to+972 50 0000000, and make the generated config explicit. - Real process environment variables may override file values.
- Blank values in config files must not overwrite non-empty process env values.
- If a required key is still missing after loading file + env, fail fast with the key name, not a downstream 401.
- In dry-run, missing or blank
WIX_AUTH_TOKENandWIX_SITE_IDarewould_block_livefindings, not blockers, unless a generated local artifact requires the site ID as a stable namespace. Do not mint a Wix CLI token solely for dry-run. WIX_MUTE_NOTIFICATIONSresolves by strategy when blank:new→on,existing→off; record the resolved value explicitly (mirrored from
…(truncated)