Minitest CLI
minitest is a command-line tool for automated mobile and web app testing. For
native apps it runs on virtual devices (simulators & emulators); for web apps it
runs browser-based targets configured on the app. An AI agent analyses the app
screen and verifies acceptance criteria you define. You manage everything through
the CLI: user stories, native builds, web lanes, runs, batches, and results.
Command shape
Every invocation follows one shape. The global options come before the subcommand, because they are parsed by the root callback:
minitest [--json] [--app <APP_ID>] <subcommand> [args…] [subcommand flags…]
minitest --json --app 713c6550-7d36-41aa-b1cd-3240b3a0dda1 user-story list
minitest --json --app $APP run verdicts <batch_id> --actionable
Agents must always pass --json. Without it the CLI prints Rich tables and
panels meant for humans — box-drawing characters, truncated columns, colour
escapes — which are lossy and unparseable. --json writes camelCase JSON to
stdout and keeps every diagnostic on stderr, so | jq is always safe.
--app may be replaced by export MINITEST_APP_ID=<uuid>; the flag wins when
both are set. Commands that are tenant-scoped rather than app-scoped (apps,
auth, flow-types) do not require it, though flow-types needs it as a tenant
hint when your account spans several tenants.
Three exceptions to the shape
| Command group | Deviation |
|---|---|
app-knowledge |
Declares its own --app, so it goes after the subcommand: minitest --json app-knowledge get --app <id>. The global pre-subcommand --app exits 2. |
auth api-key |
Declares its own --json, so it goes after the subcommand: minitest auth api-key list --tenant <id> --json. The global --json is silently ignored and a table is printed. |
flow-types |
--app must stay in the global position; flow-types list --app <id> exits 2. |
Commands that do not emit JSON
--json is still harmless on these, but their payload is markdown or a raw
value by design:
minitest init --agent,minitest maintenance --agent,minitest skill— print playbook markdown for the agent to read.minitest --app $APP env get <KEY>— prints the raw value verbatim only without--json; with--jsonit returns{"KEY": "value"}. Use the non-JSON form when capturing into a shell variable.
Prerequisites
- Install:
curl -fsSL https://raw.githubusercontent.com/minitap-ai/minitest-cli/main/install.sh | bash - Authenticate:
minitest auth login(opens browser for OAuth), check withminitest --json auth status - Set target app:
export MINITEST_APP_ID=<uuid>or pass--app <uuid>before any subcommand - Upgrade:
minitest upgrade(check the installed version withminitest --version)
Onboarding a new app (minitest init)
If you are onboarding an app from scratch, run minitest init first, in the
app's repository. It prints a short playbook covering only what is specific to
the CLI entry point — authenticating, resolving the app that was already created
for you, and where onboarding stops (the Minitest web app collects any needed
native build and launches the suite). For the suite itself the playbook sends
you to Designing a full test suite below;
follow that methodology rather than improvising one.
minitest init # prints the onboarding playbook (raw markdown when piped/non-interactive)
minitest init --agent # force raw markdown output regardless of context
The playbook is served by Minitap so it stays in step with the methodology; the
CLI falls back to an embedded copy when it cannot reach the API (which is the
normal case before minitest auth login).
Designing a full test suite
When the user wants to design a complete test suite from an app codebase —
not just add a story or two — follow the disciplined multi-wave workflow in
reference/test-suite-design.md. It takes an
app repository to a reviewed suite through ordered waves of read-only codebase
analysis: recon, surface mapping, gating/persona discovery, state modelling,
suite design, and adversarial verification.
Two companion references support that workflow:
reference/suite-schemas.md— the localsuite.yamlschema and the step-writing rules the workflow's artifacts follow.reference/minitest-target.md— the executor envelope: what the Minitest tester agent (Mini) can and cannot run or observe, so you never design a scenario it can't execute.
The workflow ends by applying the suite through the ordinary minitest CLI
commands already documented in this SKILL.md — test-profile create,
user-story create (with --criteria, --profile, --depends-on),
app-knowledge update, and so on. There is no special apply command: the
reviewed suite is replayed as regular CLI calls in dependency order.
If a scenario needs a file available in the test environment, minitest test-file upload it and bind it with minitest user-story-binding set-files
while applying the suite.
Maintaining existing tests (minitest maintenance)
Use minitest maintenance when the app UI/code changed and the customer wants
their Minitest user stories updated without sharing GitHub/code with Minitap.
Run it from the app repository. The customer's own coding agent reads local code;
the CLI sends only proposed test-flow edits and an opaque local HEAD SHA.
# Fetch the server-composed maintenance brain (same rules as cloud maintenance)
minitest maintenance --agent
# First agent step: opens a maintenance run and returns context as JSON
minitest --json maintenance context
The context output includes:
mode:auditon the first run,incrementalafter a watermark existsfromSha: null in audit mode; diff baseline in incremental modeheadSha: local HEAD SHA reported by the CLIcontext: current stories, stable criterion ids, dependencies, app memoryguardrail: pending Release Queue count; warn the user before continuing when non-zero
In incremental mode, inspect changes with git diff <fromSha>..HEAD locally.
In audit mode, inspect the full current codebase against every active story.
Post findings back through these mechanics-only commands:
minitest --json maintenance affected --file affected.json
minitest --json maintenance change --file change.json
minitest --json maintenance divergence --file divergence.json
minitest --json maintenance status --phase triage --message "Mapped affected stories" --progress 40
minitest --json maintenance complete --changed # or --no-changed when nothing changed
maintenance status,complete, andapplyprint nothing on stdout under--json— judge them by their exit code (0 = accepted), or drop--jsonto read the human confirmation.
The exact JSON fields for affected.json, change.json, and divergence.json
are specified by the maintenance brain fetched via minitest maintenance --agent;
follow that contract rather than inventing field names — the endpoint currently
returns a 500 for some tenants, and without it the payload shapes cannot be guessed
(the API rejects invented fields). complete --changed is
what advances the watermark to the current HEAD, so the next run resolves to
incremental mode; use --no-changed when nothing was edited.
At the end, always show the user the Release Queue review link so they can inspect the proposed edits in the webapp, and offer to apply them right away — it is one flow, not two modes:
minitest maintenance apply --review # prints the Release Queue link (no changes made)
minitest maintenance apply # applies all pending edits now
Leave --json off these two: they only ever print human text, and the review
link is the whole point of --review.
Present it as a single choice: surface the review link, then ask whether the user wants you to apply the edits now or review them manually through the link first. Auto-applying is simply the option they can pick, not a separate mode.
Never ask the user to paste code or diffs into Minitap. The privacy contract is: local code stays local; proposed criteria/dependency edits are the only payload.
Authentication
Three credential sources, in priority order:
MINITEST_TOKEN— raw bearer override (legacy; usually unset).MINITEST_API_KEY— tenant-scopedmtk_…key, recommended for CI/scripts.minitest auth login— interactive OAuth.
If both MINITEST_TOKEN and MINITEST_API_KEY are set, MINITEST_TOKEN wins and a stderr warning is emitted once per process.
Key rotation
mtk_ keys are mintable and revocable but do not expire. To rotate: mint a new key, update the secret in your CI/orchestrator, then revoke the old key:
minitest auth api-key mint --tenant <tenant-id> --name new-ci --json
minitest auth api-key list --tenant <tenant-id> --json
minitest auth api-key revoke --tenant <tenant-id> --key <old-key-id> --json
Note the trailing --json: this group owns its own flag, and the global
pre-subcommand --json is ignored here.
CI usage
env:
MINITEST_API_KEY: ${{ secrets.MINITEST_API_KEY }}
steps:
- run: minitest --json apps list
Treat MINITEST_API_KEY as a credential. Never commit it; rotate on suspected leak.
Global Flags
| Flag / environment | Effect |
|---|---|
--json |
camelCase JSON to stdout, diagnostics to stderr. Must appear before the subcommand. Always pass it when driving the CLI from an agent |
--app <id> |
Target app (overrides MINITEST_APP_ID). Must appear before the subcommand |
MINITEST_CHANNEL |
Overrides the X-Minitest-Channel header value (default cli) for provenance tagging of automated editors |
Exit Codes
| Code | Meaning |
|---|---|
| 0 | Success |
| 1 | General error |
| 2 | Authentication required (auth commands) |
| 3 | Network / API error |
| 4 | Resource not found |
| 5 | Build rejected as invalid |
| 6 | Conflict — re-read, rebuild, retry once |
Credentials rejected on a normal command exit 1, not 2 — only the auth
commands themselves use 2.
Only 3 is worth a blind retry. 6 comes from df alone and is explained in
commands/draft-features.md.
Core Workflow
1. Identify the app
minitest --json apps list # bare JSON array of {id, name, tenantId, …}
Dependency graph
Visualise the user-story dependency DAG as a Mermaid flowchart — useful for understanding the execution order before creating or modifying stories:
# Raw graph JSON (nodes + edges)
minitest --json apps dependencies <app_id>
# Same, taking the app from the global flag or MINITEST_APP_ID
minitest --json --app <app_id> apps dependencies
# Without --json: a Mermaid `flowchart TD` for a human to paste into a doc
minitest apps dependencies <app_id>
The Mermaid form labels each node "Story Name\n(type)" with directed edges
showing dependency relationships (parent → child).
Before changing dependencies, agents must simulate the delta through the CLI instead of computing graph effects themselves:
minitest --json --app <app_id> apps dependencies <app_id> --simulate \
--add <story_id>:<depends_on_id> \
--remove <story_id>:<depends_on_id>
Simulation is read-only and returns valid, cycle (including the cycle path
when invalid), addedEdges, removedEdges, affectedStories, runOrder
(parallel waves), and the resulting edges. An edge A:B means “A depends on
B”, so B runs first. --add and --remove require --simulate, and
--simulate requires at least one of them.
Creating apps
If the user does not yet have an app for the project, create one. The app
lives under a tenant; when the authenticated user belongs to a single
tenant the CLI auto-resolves it, otherwise pass --tenant <id> explicitly
(apps list exposes existing tenant IDs in JSON mode).
# Auto-resolve tenant (single-tenant users), native mobile app
minitest --json apps create --name "My Mobile App" --platform ios --platform android
# Full record as JSON, suitable for piping
minitest --json apps create --tenant <tenant_id> --name "My App" \
--platform web --web-url https://example.com \
--description "Customer portal" --slug "my-app" --icon ./icon.png
In a multi-tenant non-interactive context (CI, piped invocation), --tenant
is required: the command exits 1 with a clear error otherwise.
An app created with
--platform webstill has no web execution target until one is configured in the Minitest web app, andrun start --webrefuses to launch until then. Creating a web app end-to-end is not a CLI-only flow. There is also noapps delete, so avoid creating throwaway apps.
2. Create user stories
A user story describes a user journey to test. It has a name, a type, an optional description, and a list of acceptance criteria — plain-text assertions the AI agent will verify visually on the target screen (mobile device or browser).
--profile <profile_id> is optional. If omitted, Minitest auto-assigns the app's default profile when one is configured.
minitest --json --app <app_id> user-story create \
--name "User Login" \
--type login \
--profile <profile_id> \
--description "Email/password login from welcome screen" \
--criteria "The login screen shows email and password fields" \
--criteria "After submitting valid credentials, a loading indicator appears" \
--criteria "The home screen is displayed after successful login"
Use --depends-on to declare that this story must be run after another story
completes successfully (repeatable for multiple parents):
minitest --json --app <app_id> user-story create \
--name "View Order History" \
--type navigation \
--depends-on <login_story_id> \
--criteria "The order history screen is displayed"
User story types: login, registration, checkout, onboarding,
search, settings, navigation, form, profile, other, custom.
Ask first: Do not create
checkout, billing, or payment user stories until you know how this app expects payment to be exercised. The hard limit is narrow — the tester never enters a real card and never makes a real-money purchase — but everything short of that is testable: sandbox and test cards, in-app purchases and subscriptions through the RevenueCat Test Store or an Apple/Google platform sandbox, and whatever custom payment flow the app uses. What breaks a run is guessing, not the payment step itself.So when you find a paid flow during codebase analysis, ask the user which applies: a test card and its number, a build wired to the RevenueCat Test Store with its Test Store API key, an Apple or Google platform sandbox account, a bypass or promo code, a staging payment provider, or none of these — in which case the story stops before the charge and says so in its last criterion. The two RevenueCat environments are not interchangeable, so record which one this build uses: the Test Store mocks billing entirely and needs no store setup, while a platform sandbox drives the real App Store or Play purchase flow.
Record the answer where the tester will read it at run time: the flow type's
--usage-prompt(e.g."Paid plans are sandboxed: use card 4242 4242 4242 4242."), the app knowledge, or the persona's--about.Once you have that answer, author these stories like any other. Do not refuse a whole feature over the payment step it ends on, and never invent a restriction of your own or bake one into a flow type.
Test account requirement: Before creating user stories that require login or account-specific state, ensure the user provides test credentials via the Minitest web app's test configuration. User stories should only cover journeys the test account can actually perform.
Test Profiles
When generating user stories, create a test profile for every distinct role or subscription tier the app requires. Each profile represents a unique starting state the agent needs to run a story (e.g. "Free User", "Pro User", "Admin", "Driver").
Default to email-OTP personas. Give each profile a <prefix>@qa.minitap.ai username and NO password. Every @qa.minitap.ai address delivers into a shared inbox the tester agent reads at run time, so it can sign in (or sign up) by pulling the login/verification code itself — no real credentials to manage:
minitest --json --app $APP test-profile create \
--name "Pro User" --username "pro@qa.minitap.ai" \
--about "Pro subscription active, has saved items, payment method on file"
A non-@qa.minitap.ai username with no password is rejected unless the profile also carries --phone-number or --static-otp-code — those personas have their own way to sign in, so keep the address as given. Otherwise keep the @qa.minitap.ai domain (or leave the username blank to auto-generate one).
Bring-your-own account. Only when the user supplies a real account they own (and the app needs a password) pass both, via stdin:
printf "%s" "$PASSWORD" | `minitest --json --app $APP test-profile create \
--name "Pro User" --username "real-user@example.com" --password-stdin --about "..."
Specific account state (e.g. premium). To exercise a flow that needs a particular state, proactively create a <something>@qa.minitap.ai persona WITH an explicit password, then ask the user to link that exact email+password combo to a pro/specific-state account in their backend. The @qa.minitap.ai address keeps the inbox readable for any OTP while the password lets them pre-provision the account state.
No persona bound. If a story has no profile, the agent defaults to anonymous (skip login). If a flow forces authentication, it self-generates a <random>@qa.minitap.ai with a runtime password, signs up, and reads the inbox for the confirmation/OTP code — so unbound scenarios still work without you provisioning anything.
Phone-OTP personas. If the app authenticates via SMS one-time codes instead of email, set --phone-number (E.164, e.g. +14155551234) and --static-otp-code (the fixed code the customer's backend accepts for that number) instead of --username/--password:
printf "%s" "$OTP_CODE" | minitest --json --app $APP test-profile create \
--name "Driver" --phone-number "+14155551234" --static-otp-code-stdin \
--about "Verified driver, active delivery in progress"
Ask the user for the whitelisted phone number and its fixed OTP code — these must be pre-provisioned in the customer's backend, the same way a bring-your-own-account password is. A phone or static-OTP persona keeps whatever --username you give it (a real customer address, or none at all) — it is never forced onto @qa.minitap.ai and never has one invented for it, since it already has its own way to sign in.
Fill the about field with what makes each profile distinct (e.g. "Pro subscription active, has saved items, payment method on file"). This context is injected into the tester agent's prompt at run time.
If the app uses a third-party auth provider (e.g. Google OAuth), a shared Minitap account covers that flow — bind it to the relevant story instead of creating a new profile. Those shared pool addresses are also @qa.minitap.ai, so their inboxes are readable the same way.
Bind every story that requires authentication to its profile at creation time:
- Use
user-story create --profile <profile_id>when you need a specific profile. - If you omit
--profile, ensure the app default profile is already configured so story creation auto-binds it. - If needed, use
user-story-binding set-profileimmediately after creation.
Acceptance criteria rules:
- Must be visually verifiable (the agent only sees the screen)
- Must be specific and unambiguous
- One assertion per criterion
- Order them chronologically as they appear in the journey
Other user story commands:
minitest --json --app <app_id> user-story list
minitest --json --app <app_id> user-story get <user_story_id>
minitest --json --app <app_id> user-story update <user_story_id> --name "New Name"
minitest --json --app <app_id> user-story update <user_story_id> --add-criteria "New check"
minitest --json --app <app_id> user-story update <user_story_id> \
--set-criterion <criterion-id-or-index>="New text"
minitest --json --app <app_id> user-story update <user_story_id> \
--revert-criterion <criterion-id-or-index>=<version_id>
minitest --json --app <app_id> user-story update <user_story_id> \
--criteria "First check" --criteria "Second check" # full replace
minitest --json --app <app_id> user-story delete <user_story_id> --force
Acceptance criteria are versioned. For rewording, always prefer repeatable
--set-criterion <criterion-id-or-index>="new text": it changes one criterion while preserving its identity and version history.--revert-criterion <criterion-id-or-index>=<version_id>restores that version's exact text and records the change as a revert. Both flags conflict with--criteriaand--add-criteria.--criteriafully replaces the set and severs history for reworded criteria;--add-criteriaonly appends.
Use
criterionId, notid. Each entry ofacceptanceCriteria[]inuser-story get --jsoncarries both:idis the version id andcriterionIdis the stable identity.--set-criterionacceptscriterionIdor a 1-based index; passingidfails with "No criterion matches id".--revert-criterion <criterionId>=<id-of-the-target-version>takes one of each.
Story dependencies
Use --depends-on / --remove-dependency on user-story update to manage
which stories gate this one:
# Replace the full dependency set (all parents at once)
minitest --json --app <app_id> user-story update <story_id> \
--depends-on <parent_id_1> --depends-on <parent_id_2>
# Remove a single dependency without touching the others
minitest --json --app <app_id> user-story update <story_id> \
--remove-dependency <parent_id>
# Clear all dependencies (pass empty --depends-on list)
minitest --json --app <app_id> user-story update <story_id> --depends-on ""
--depends-onis a full replace: omitting a previously set parent removes it. Use--remove-dependencyfor a surgical delta when you only want to drop one parent. The two flags are mutually exclusive on the same invocation —--remove-dependencyis ignored when--depends-onis also provided.
Device count
A story's device count is how many virtual devices a single run provisions.
By default it is auto: one device per bound persona (minimum 1). Set an
explicit override with --device-count on user-story create / update. The
value is capped server-side at min(3, tenant device quota).
# Create a story that always runs on 2 devices
minitest --json --app <app_id> user-story create \
--name "Two-player match" --type other --device-count 2 \
--criteria "Both players see the shared game state"
# Override an existing story to 3 devices
minitest --json --app <app_id> user-story update <story_id> --device-count 3
# Reset back to auto (one device per bound persona)
minitest --json --app <app_id> user-story update <story_id> --device-count auto
Use more than one device only for flows that genuinely need concurrent devices:
- Real-time cross-account flows — two accounts interacting live (chat, multiplayer, sharing, calls) where one device acts and the other must observe.
- Session-conflict same-account tests — the same account signed in on two devices at once to exercise session invalidation, presence, or sync conflicts.
Single-user journeys (login, checkout, navigation) should stay on auto. list
and get show the effective device count when it is greater than 1; omitting
--device-count on create leaves the story on auto.
Camera media (web runs)
Use --camera-media on user-story create or user-story update to feed a
specific image or video into the virtual webcam during web runs (e.g. a story
that uploads an ID photo or scans a QR code). The value is either a local
file path to upload or an existing test-file ID to reuse:
# Upload a local image/video and attach it to a new story
minitest --app <app_id> user-story create \
--name "Scan QR to check in" --type navigation \
--camera-media ./fixtures/checkin-qr.png
# Reuse an already-uploaded test file by its ID
minitest --app <app_id> user-story update <story_id> \
--camera-media <test_file_id>
# Reset the story back to the built-in default camera feed
minitest --app <app_id> user-story update <story_id> --clear-camera-media
A UUID value is treated as an existing test-file ID and reused as-is; any other value is a local path that gets uploaded as a test file first. The file must be an image or video within the size caps — video ≤ 50 MB, image ≤ 25 MB.
--camera-mediaand--clear-camera-mediaare mutually exclusive.
3. Reading and managing flow types, and app knowledge
Needs minitest-cli ≥ 0.22.0. On an older CLI the
create/update/deletesubcommands below do not exist, and--json flow-types liststill returns a flat array of names. Check withminitest --versionand runminitest upgradebefore assuming a command is broken.
When generating user stories programmatically (e.g. from an exploration
subagent), validate every --type value against the live list of flow types
before calling user-story create — invalid values exit non-zero.
# Bare JSON array, with each custom type's id and presentation fields
minitest --json flow-types list
# [{"name": "login", "custom": false, "id": null, "icon": null, …},
# {"name": "Billing", "custom": true, "id": "3fa85…", "icon": "credit-card", …}]
# Just the names, for a membership check
minitest --json flow-types list | jq -r '.[].name'
The list is the built-in types plus your tenant's custom ones. Flow types are
tenant-scoped, so --app is only a tenant hint. It is nonetheless
required whenever your account spans more than one tenant — otherwise the
command exits non-zero asking which tenant to act on. Keep it in the global
position (minitest --json --app <id> flow-types list); a trailing --app
exits 2.
When no built-in type fits a journey, create your own instead of forcing it into
other:
minitest --json flow-types create --name "Subscription" \
--usage-prompt "Paid plans are sandboxed: use card 4242 4242 4242 4242."
# Optional presentation flags, defaults are tag/gray
minitest --json flow-types create --name "Subscription" --icon credit-card --color green
--usage-prompt is extra context handed to the testing agent whenever it runs a
story of that type — use it for domain rules that apply to the whole family of
flows. Creating a name that already exists (or that collides with a built-in)
fails with the server's conflict message.
Rename or restyle an existing custom type with update, addressing it by name
or by id. The type keeps its identity, so stories already on it stay on it:
minitest --json flow-types update "Subscription" --name "Billing"
minitest --json flow-types update "Billing" --usage-prompt "Card 4242… declines above 500."
delete removes a custom type and resets every user story on it to other,
so it refuses to run without --yes:
minitest --json flow-types delete "Billing" --yes
Pass the type name to any user-story command; the CLI resolves it for you (matching is case-insensitive):
minitest --json --app <app_id> user-story create --name "Upgrade to Pro" --type "Billing" \
--criteria "The plan badge reads Pro after checkout"
minitest --json --app <app_id> user-story list --type "Billing"
flow-types list wraps GET /api/v1/user-story-types (built-ins) plus
GET /api/v1/apps/<app_id>/custom-user-story-types; create, update and
delete wrap POST / PATCH / DELETE on the latter. The app in that path is
only addressing: the CLI picks one of your apps automatically when you do not
pass --app and your account has a single tenant.
For app-level prompt context (the markdown blob that conditions the AI agent
during runs), use app-knowledge:
Note the --app position: this group declares its own, so it goes after the
subcommand and the global pre-subcommand form exits 2.
# {"appId": …, "content": "<markdown>"} — no version metadata on read,
# and exit 4 when the app has no knowledge set yet
minitest --json app-knowledge get --app <app_id>
# Push a new version inline
minitest --json app-knowledge update --app <app_id> --content "# App Knowledge\n..."
# Or load it from a file (preferred for non-trivial markdown)
minitest --json app-knowledge update --app <app_id> --content-file ./app-knowledge.md
app-knowledge update calls PUT /api/v1/apps/{app_id}/app-knowledge and
prints the new versionNumber to stdout (full record with --json). Each
update creates a new prompt version — there is no rollback shortcut.
3b. See what exploration actually mapped (minitest screens)
Before you invent scenarios for an app, read the screen map: it is the list of screens the exploration crawl genuinely stood on, so it tells you what the app really contains instead of what its name or flow types imply.
# Every mapped screen, shallowest first
minitest --app <app_id> screens list
# One platform only
minitest --app <app_id> screens list --platform ios
# The navigation shape — the fastest way to see where a crawl stopped
minitest --app <app_id> screens list --tree
# Only the screens the crawl could not get past
minitest --app <app_id> screens list --blocked
# One screen: how to reach it, and where it leads
minitest --app <app_id> screens get "Onboarding age"
screens list wraps GET /api/v1/apps/{app_id}/screens, which returns the
whole map in one call (screenshots are signed in a single batch). Filtering by
--platform is server-side; --area and --blocked are applied client-side
over that one response.
With --json the command returns the map itself, and screenCount always
matches the screens array beside it (so a filtered call reports the filtered
count, not the server total). --tree is a human rendering only — under
--json it is ignored and you get the same map object, whose outgoing edges
carry the graph. Note the casing seam inherited from the API: the envelope is
camelCase (screenKey, displayName) while outgoing and context are
snake_case (to_screen_key, parked_reason, requires_auth).
Read the frontier line, printed under the table. It reports three things that each mean something different:
| Signal | Meaning |
|---|---|
| parked edge | The crawl saw a way onward and deliberately did not follow it (a duplicate branch, a login wall, a non-navigating toggle). Parked edges are the unexplored frontier. |
| blocked screen | The crawl reached the screen but could not get past it. screens get shows the reason, and the ask gating it. |
| edge leading to a screen with no row | The crawl recorded a step to a destination that was never written as a screen — usually the destination was named slightly differently than the screen later called itself. The map understates what was reached. |
Two shapes worth recognising in --tree output:
- A long unbranching chain — exploration never escaped a funnel (typically onboarding). Scenarios written from this map will all be onboarding scenarios.
- A wide, shallow tree — exploration never got past the lobby, usually
because of auth. Check
screens getforRequires auth.
An empty result is not an error. It means no crawl has run against a build for this app yet — the map is written by exploration as it walks, so it stays empty until then.
4. Upload native builds
For iOS/Android apps, upload your .apk (Android) or .ipa (iOS) build
artifacts. The platform is auto-detected from the file extension. Only .apk
and .ipa files are supported. Web apps do not upload builds; create the app
with --platform web --web-url ... and run the web lane with --web.
Important — virtual-device builds required:
Tests run on simulators/emulators, not physical devices. You must upload builds that are compatible with virtual devices:
- iOS: a Simulator build (
.ipabuilt for the iOS Simulator destination, not a physical device). In Xcode: build for "Any iOS Simulator Device" or a specific Simulator target.- Android: an x86_64-compatible
.apk. Ensure your Gradle build includes thex86_64ABI.Uploading a device-only build (e.g. an arm64 iOS archive or an arm-only Android APK) will cause test runs to fail.
minitest --json --app <app_id> build upload ./app-release.apk
minitest --json --app <app_id> build upload ./MyApp.ipa
minitest --json --app <app_id> build list --platform ios --page-size 1
Pass --platform ios|android to build upload when the filename does not end
in .ipa / .apk and auto-detection cannot fire.
For a repo-connected app, queue a build directly from GitHub without starting tests:
minitest --json --app <app_id> build from-commit [<full_sha>] \
[--platform ios|android|web] [--platform ...] [--force-full]
The SHA is optional. When omitted, the platform resolves the default branch's
HEAD. When supplied, it must be the full 40-character lowercase hexadecimal
SHA. --force-full bypasses incremental build caches. Inspect failures with
build list --status failed; use --kind web for web builds because web rows
may have no platform, and therefore are not selected by --platform web.
build list returns completed builds only unless you pass --status, so always
use --status failed to see failures at all. --status is repeatable. Valid
values: pending | completed | failed | cancelled — there is no
running status; an in-progress build is tracked via a separate heartbeat,
not a status value, so poll build list without --status running.
Every item in the build list JSON carries a guidance object next to the raw
envelope fields. It is null when the build did not fail, so the key is always
present:
{ "source": "fix_prompt", "text": "..." }
source tells you how much to trust text. Read it before acting:
source |
Meaning |
|---|---|
fix_prompt |
Machine-authored and actionable. Act on it directly. |
remediation |
Human-readable next step. Reliable, less specific. |
summary |
One-line classification only. Investigate before changing code. |
raw |
Unstructured builder stderr, no envelope was recorded. May be a stale heartbeat or a timeout rather than a real defect. Treat as a lead, not a diagnosis. |
withheld |
An internal failure we deliberately suppressed. Nothing for you to fix: retry the build, escalate if it persists. |
none |
No failure details were recorded at all. |
Guidance falls back down that ladder, so raw only appears when no fix prompt,
remediation, or summary exists. When source is withheld the CLI also blanks
errorSummary and nulls errorRemediation, errorFixPrompt, and errorRaw on
that item, so there is nothing to dig into. Otherwise the envelope fields are
left untouched alongside guidance.
5. Run tests
Execute a user story on either native lanes or the web lane. For native runs,
provide at least one of --ios-build or --android-build; single-platform apps
may omit the other. For web runs, pass --web by itself — do not combine it
with native build flags. Web runs use the app's configured web URL and default web
targets; there are no per-run --web-url, --browser, or --viewport overrides
in the CLI.
# Run a single user story (by name or UUID) and wait for results
minitest --json --app <app_id> run start "User Login" \
--ios-build <ios_build_id> \
--android-build <android_build_id>
# Android-only app
minitest --json --app <app_id> run start "User Login" \
--android-build <android_build_id>
# Web app (no build upload needed)
minitest --json --app <app_id> run start "User Login" --web
# Fire-and-forget (returns runId immediately — useful in CI)
minitest --json --app <app_id> run start "User Login" \
--ios-build <ios_build_id> \
--android-build <android_build_id> \
--no-watch
# Run ALL user stories at once (creates one batch, fire-and-forget)
minitest --json --app <app_id> run all \
--ios-build <ios_build_id> \
--android-build <android_build_id>
# Run ALL user stories on web targets only
minitest --json --app <app_id> run all --web
# Build a required commit SHA and run its suite, polling by default
minitest --json --app <app_id> run from-commit <full_sha> \
[--platform ios|android|web] [--platform ...] \
[--user-story <id-or-name>] [--no-watch] [--timeout <seconds>]
# Cancel a running or pending run
minitest --json --app <app_id> run cancel <run_id>
Under the hood, run start and run all create a batch. A single run is
just a batch with one user story. With --json --no-watch, run start emits
only {"runId": …, "status": …}.
--web requires a web execution target configured on the app in the Minitest
web app; without one the run is refused even though the app was created with
--platform web --web-url ….
run from-commit requires a full SHA and polls the batch until a verdict unless
--no-watch is passed. Its initial response can legitimately contain no story
runs while the commit build is being prepared. For a web-only app, explicitly
pass --platform web: omitting platforms defaults server-side to iOS and
Android. Failed web builds currently may have no error envelope or fix prompt.
6. Check results
# Check a specific run
minitest --json --app <app_id> run status <run_id>
# Poll until completion
minitest --json --app <app_id> run status <run_id> --watch
# List all runs for a user story
minitest --json --app <app_id> run list "User Login"
minitest --json --app <app_id> run list "User Login" --status failed
minitest --json --app <app_id> run list "User Login" --all
Run statuses: pending → running → completed | failed | cancelled
A completed run includes per-target results: pass/fail for each acceptance
criterion, fail reasons, recording URLs, and stable resultId, criterionId,
and criterionVersionId identifiers for each criterion result.
7. Work with batches
A batch groups runs triggered together (by run all, CI, or a single
run start). Use the batch group to inspect or cancel them.
minitest --json --app <app_id> batch list # recent batches
minitest --json --app <app_id> batch list --status running
minitest --json --app <app_id> batch list --commit-sha abc1234
minitest --json --app <app_id> batch list --user-story <id>
minitest --json --app <app_id> batch get <batch_id> # batch + its runs
minitest --json --app <app_id> batch cancel <batch_id> # cancels all pending/running runs
Batch statuses: pending | awaiting_build | `ru
…(truncated)