Repair Intake: $ARGUMENTS
Run one batch-repair cycle against the queue identified by $ARGUMENTS. Where lisa-intake
scans the ready role and moves work forward, repair-intake scans the stuck and
close-out roles and moves work unstuck or fully closed:
- Stalled in-progress — an item left in an in-progress role (build
claimed, PRDin_review) whose processing cycle died. It is technically "being worked" but nothing is happening, so it sits ignored forever. (The vendor PRD intakes explicitly leave an errored PRD inin_review"for the human to investigate from there" — that orphan is exactly what this skill recovers.) For a stalled build, repair-intake first diagnoses why it stalled by inspecting its PRs and deploys. A PR that already merged is recovered by applying the env transition build-intake never got to (its merge gate left the itemclaimedwhen the merge landed after its agent returned) — no re-dispatch. A PR that is merely behind its base (BEHIND, no conflict) is re-synced in place withgh pr update-branchso the already-enabled auto-merge can finally land — a clean rebase needs no human, and leaving it stranded is the exact gap that lets an auto-merge PR sit unmerged forever. A true merge conflict is first given one bounded in-place re-dispatch to the build agent — whosedrive-pr-to-mergefix-mode loop resolves conflicts — because a conflict, unlike a failing external check, is fixable by re-running the build; only a conflict that survives that single attempt (or that the agent says needs design input) becomes a fix ticket. The genuinely non-resolvable blockers — failing checks / unaddressed CodeRabbit orCHANGES_REQUESTEDreview / a failed deploy — get a build-ready leaf fix ticket with the item moved toblocked(blocked by that ticket) instead of blindly re-dispatching the agent, which would just churn against them. - Recoverable blocked — an item in
blockedwhose blocker may now be gone. The blocker is one of three classes, and repair re-checks all of them, not just dependencies: (a) anis blocked bydependency has since closed; (b) a validation / quality-gate self-block — the item was bounced toblockedby its own pre-flightverify/validategate (missing Validation Journey, Sign-in Required, Acceptance Criteria, etc.) with no dependency at all, and a human has since edited the item to add what the gate demanded; or (c) clarifying questions answered / an ambiguity research can now settle. A self-block (b) is the common one missed by dependency-only re-checks: nothing else is blocking it, so re-running the same gate against its current content is the only way to know it is now passable. - Terminal-open drift — an item already carrying its true terminal lifecycle role (for
example GitHub
status:done) but still open/active in the provider's native state. - Rollup drift — a parent/container item (Epic, Story, PRD, Linear Project, or equivalent)
whose own lifecycle state does not match the roll-up of its children's states per
leaf-only-lifecycle. This covers the completed case (all children terminal → close the parent out) and the intermediate-env case (all children shipped to an env likeOn Stg, but the parent never advanced — including a parent left stranded in a status it should never carry). - Stale-
readycontainer — a parent/container (open child work, or a childless Epic) wrongly carrying the build-ready role. This is a leaf-only-invariant violation the build-intake claim gate deliberately leaves for a human; repair-intake reconciles it by rolling the parent up from its children (with an audit note), so a container never sits inreadyindefinitely. - Missing official ready-label drift — a GitHub issue that is missing every configured Lisa
lifecycle label. repair-intake classifies it as a PRD or build ticket and adds the configured
readylabel (prd-readyfor a PRD, buildstatus:readyfor a ticket) so normal intake can see it; if the later intake/implement gate finds the item incomplete, it moves the item toblocked. - Missing native child link drift — a GitHub parent (a
ticketed/other open non-product-owned PRD, or a build Epic/Story container) whose children are discoverable — from the generated-work section/comment for a PRD, or from body parentage (Parent: #<n>/Parent Epic: #<n>) for a build container resolved via the documented hierarchy fallback — but whose native sub-issue list is missing one or more of those children. This is the common shape when children were created by an external generator (e.g. Codex) or an older write path that recorded parentage only in prose and never calledaddSubIssue. repair-intake replays theprd-backlink/github-write-issuenative-linking contract and attaches the missing same-repo children idempotently, so rollup and the GitHub UI can rely on the native graph again.
This skill is the symmetric counterpart to lisa-intake. It reuses the same queue-detection,
the same agent-team orchestration, the same "don't ask, just run" confirmation policy, and the
same per-item surfaces the vendor intakes use (lisa:<source>-to-tracker dry-run for PRDs;
lisa:<tracker>-agent + the scanner's lifecycle transitions for build) — it differs in which
roles it scans and, for stalled/blocked work, that it skips the claim step (the item is already
claimed/blocked). Close-out candidates do not dispatch agents; they only reconcile terminal
lifecycle state with provider-native closure and rollup state.
Public contract
/lisa:repair-intake <queue> [intake_mode=prd|build|both] [stale_after=2h] [max_candidates=100] [force=true]
| Token | Meaning | Default |
|---|---|---|
<queue> |
Same queue identifier lisa-intake accepts (see Source dispatch). Required. |
— |
intake_mode |
prd | build | both. Only meaningful for a GitHub org/repo (or bare github) that hosts both PRD and build label namespaces. both is unique to repair — a repair sweep usefully covers both lifecycles in one schedule. Absent → both when both namespaces exist, else whichever lifecycle exists. |
both for dual GitHub queues; otherwise infer |
stale_after |
How long since the last state-changing transition into the in-progress role, or since the last human / PR-side forward-progress activity, before an in-progress item counts as stalled. Automation self-comments do not reset this clock. Accepts 24h, 90m, 2d, or 0 (treat any in-progress item as stalled — manual recovery, also the only way to resume work on a provider that exposes no reliable timestamp). Overrides config. |
2h |
max_candidates |
Cap on how many stuck/close-out candidates to enumerate and evaluate. Repair every materially actionable candidate within this bounded set, then stop. Overrides config. | 100 |
force |
true bypasses the loop-prevention backoff window (so a manual re-run re-attempts items even if their fingerprint is unchanged). It does not change the staleness rule — use stale_after=0 for that. |
false |
Confirmation policy
Do NOT ask the caller whether to proceed. Once invoked with a queue, run the cycle to completion. The caller (a human at the CLI or a scheduled cron) has already authorized the run by invoking the skill; re-prompting defeats the purpose of a background repair sweep.
Specifically forbidden:
- Previewing projected scope (number of stuck items, projected re-dispatch count, write counts) and asking whether to continue.
- Offering A/B/C-style choices like "repair / skip / report-only" — the documented behavior IS the default.
- Pausing because many items are stuck, an item looks complex, or a repair is likely to land
the item back in
blocked. Returning an item toblockedwith a current, accurate note is a valid outcome of the repair lifecycle, not a failure. - Pausing because a re-dispatch looks expensive. The cost of one cycle is bounded by
max_candidatesand the actionable subset inside that cap; the cost of stalling a scheduled cron waiting on a human is unbounded.
The only legitimate reasons to stop early:
- Missing required input (no queue argument, missing project configuration). Surface the missing value and exit.
- The queue itself is misconfigured (Status property missing expected values, JIRA workflow can't reach required transitions). Surface and exit.
- No stuck/close-out candidates, or none actionable this cycle. Exit cleanly with the idle-case summary.
Orchestration: agent team
If you are NOT already operating inside an agent team (no prior successful team-creation or subagent-delegation tool call in this session, not spawned into a team context), the very first thing you do is establish team orchestration.
Use the team tool for the current runtime:
- Claude Code >= 2.1.178: there is no
TeamCreatetool; the team forms automatically when you spawn the first teammate withAgent. That first spawn should be the bounded specialist needed to start this flow. On older Claude Code that still exposesTeamCreate, the explicit team-create path is also acceptable. - Codex: do not call
TeamCreate; Codex does not expose that Claude tool. Usetool_searchwith a query likemulti-agent toolsto loadmulti_agent_v1, then usemulti_agent_v1.spawn_agentfor teammate delegation. Treat the first successfulspawn_agentcall as establishing team orchestration. - Other runtimes: use the current runtime's tool-discovery mechanism to discover and call the appropriate multi-agent/team tool.
If no team creation or subagent delegation tool is available, explicitly state that team orchestration is unavailable in this runtime, continue as the lead agent, and preserve the workflow's review, verification, and task-tracking obligations locally.
Until the team is established, the first Codex teammate has been spawned, or the no-team
fallback has been declared, do NOT call any of: TaskCreate, Skill, MCP tools
(Atlassian / Linear / GitHub / Notion), Read, Write, Edit, Bash, Grep, Glob.
The initial Claude Agent spawn described above is the only pre-team exception because it
establishes the team.
Scanning the queue, evaluating staleness, and dispatching per-item repairs — all of those are
tasks for the team you are about to create, not for the lead session before orchestration
exists.
If you ARE already inside an agent team (e.g., a teammate invoked this skill via the Skill tool), do NOT create a second team — many harnesses reject double-creates — and do NOT collapse the nested flow into a single inline worker. A nested team-first flow must still bring in the specialists it requires by adding them to the existing team, not by doing the work itself:
- Claude: teams are flat and only the lead can add named teammates, so do NOT call
Agentwith anamefrom a teammate (the harness rejects it: "Teammates cannot spawn other teammates — the team roster is flat"). Send the team lead a message naming the specialist teammate(s) this flow needs, their task assignments, and completion criteria, then coordinate through the shared task list until they finish. An anonymous subagent (Agentwithnameomitted) is permitted only for bounded one-shot work whose result returns directly to you — it is not a substitute for the required lifecycle specialists. - Codex: do NOT call
TeamCreate. If the lead/root agent is addressable (you were given its id/handle), send it a request tomulti_agent_v1.spawn_agentthe specialist agent(s), including each agent's prompt, ownership, and expected result. If no lead handle exists butspawn_agentis available to you, spawn only the bounded specialist agent(s) this flow needs,wait_agentfor their results, and relay those results upward to the parent/lead.
Treat the first successful lead-spawn request (or, on the Codex fallback, the first specialist spawn) as preserving team orchestration. Never satisfy a team-first lifecycle flow by doing all the work inline. The cycle's outer team is created by repair-intake. Each per-item repair it runs
(lisa:<source>-to-tracker for a PRD, lisa:<tracker>-agent for a build item) executes within
the same team — those skills' orchestration preambles detect the existing team and skip creating
a second one. One team per cron cycle.
Source dispatch
Detect the queue type from $ARGUMENTS using the exact same detection and disambiguation
rules as lisa-intake — read that skill's "Source dispatch" section for the authoritative
table; the detection is identical and only the per-item action changes (repair instead of
claim-and-advance). The essentials, inlined here so this skill is self-complete:
If $ARGUMENTS is... |
Queue / lifecycle | Source/tracker key | Candidates repaired |
|---|---|---|---|
| Notion database URL/ID | PRD (Notion) | source=notion | in_review, blocked, terminal/open PRDs, all-terminal generated-work rollups |
| Confluence space URL/key | PRD (Confluence) | source=confluence | in_review, blocked, terminal/open PRDs, all-terminal generated-work rollups |
| Confluence parent page URL/ID | PRD (Confluence, narrowed) | source=confluence | in_review, blocked, terminal/open PRDs, all-terminal generated-work rollups |
Linear workspace URL, team URL/key, or literal linear |
PRD (Linear) | source=linear | in_review, blocked, terminal/open PRDs, all-terminal generated-work rollups |
GitHub repo URL / org/repo (PRD namespace) |
PRD (GitHub) | source=github | in_review, blocked, terminal/open PRDs, missing PRD child links, all-terminal generated-work rollups |
GitHub repo URL / org/repo with tracker = github (build namespace) |
Build (GitHub) | tracker=github | claimed, blocked, terminal/open issues, parent rollups (intermediate-env + all-terminal), stale-ready containers |
GitHub repo URL / org/repo with an open issue missing configured lifecycle labels |
GitHub label normalization | per classified lifecycle | add configured prd.ready or build ready |
Literal github |
GitHub; route by intake_mode (prd / build / both) |
per lifecycle | per lifecycle above, plus GitHub ready-label normalization |
| JIRA project key or full JQL | Build (JIRA) | tracker=jira | claimed, blocked, terminal/closure verification, parent rollups (intermediate-env + all-terminal), stale-ready containers |
Disambiguation (same as lisa-intake): a notion.so/notion.site URL → Notion; an Atlassian
/wiki/spaces/<KEY> URL → Confluence (with /pages/<id> → parent-page narrowing); a
linear.app workspace/team URL or literal linear → Linear; a github.com URL / <org>/<repo>
token / literal github → GitHub; a bare token matching the JIRA project-key regex → JIRA
(else try Confluence space, then Linear team); a string with JQL operators → JQL. A single-item
URL is out of scope — this skill is batch-only; repair one item by hand via lisa-implement
(build) or by re-running lisa:<source>-to-tracker (PRD).
Role names for every vendor are resolved from .lisa.config.json per the config-resolution
rule — never hardcode status/label strings. The relevant repair roles:
| Lifecycle | Vendor | In-progress role key | Blocked role key | Terminal / rollup role key |
|---|---|---|---|---|
| Build | JIRA | jira.workflow.claimed (In Progress) |
jira.workflow.blocked (Blocked) |
env-resolved jira.workflow.done |
| Build | GitHub | github.labels.build.claimed (status:in-progress) |
github.labels.build.blocked (status:blocked) |
env-resolved github.labels.build.done (status:done) |
| Build | Linear | linear.labels.build.claimed (status:in-progress) |
linear.labels.build.blocked (status:blocked) |
env-resolved linear.labels.build.done (status:done) |
| PRD | Notion | notion.values.in_review (In Review) |
notion.values.blocked (Blocked) |
notion.values.shipped (Shipped) |
| PRD | GitHub | github.labels.prd.in_review (prd-in-review) |
github.labels.prd.blocked (prd-blocked) |
github.labels.prd.shipped (prd-shipped) |
| PRD | Linear | linear.labels.prd.in_review (prd-in-review) |
linear.labels.prd.blocked (prd-blocked) |
linear.labels.prd.shipped (prd-shipped) |
| PRD | Confluence | confluence.parents.in_review (page id) |
confluence.parents.blocked (page id) |
confluence.parents.shipped (page id) |
In addition to the lifecycle roles above, the build lifecycle defines the human_needed marker — an additive label (jira.labels.human_needed / github.labels.build.human_needed / linear.labels.build.human_needed, default Human Needed / human-needed) that rides alongside blocked when the block needs human-only input no agent or retry can supply (see config-resolution "Build markers"). repair-intake's interaction with the marker is asymmetric and is the whole point of the distinction below:
- The blocks repair-intake itself writes are the auto-recoverable kind — it files a build-ready fix ticket and moves the item
blockedblocked by that ticket, expecting the next cycle to self-heal. Those are nothuman_needed; if such an item arrives already carrying a stalehuman_neededmarker, repair-intake clears it (the block is no longer waiting on a human). - The blocks the vendor agent writes when repair-intake re-dispatches it (its pre-flight gate) carry
human_neededalready — the agent owns that marker. repair-intake leaves it in place.
Resolve with the standard role-read pattern (local overrides global, default fallback):
read_role() {
local path="$1" default="$2"
local local_v global_v
local_v=$(jq -r "${path} // empty" .lisa.config.local.json 2>/dev/null)
global_v=$(jq -r "${path} // empty" .lisa.config.json 2>/dev/null)
echo "${local_v:-${global_v:-$default}}"
}
# e.g. build/github:
CLAIMED=$(read_role '.github.labels.build.claimed' 'status:in-progress')
BLOCKED=$(read_role '.github.labels.build.blocked' 'status:blocked')
Access layer (which surface does each write)
repair-intake stays vendor-neutral; concrete reads/writes go through the same layers the vendor
intakes use. Never call Atlassian MCP or acli directly — go through lisa-atlassian-access.
| Vendor | Reads (scan / comments / links) | Writes (transition / comment / close-out) | Re-dispatch / re-validate |
|---|---|---|---|
| JIRA (build) | lisa-atlassian-access search-issues / lisa-jira-read-ticket |
lisa-atlassian-access transition / comment |
lisa-jira-agent |
| GitHub (build) | gh issue list / gh issue view --json / gh pr list / GraphQL sub-issues |
gh issue edit (labels) / gh issue comment / gh issue close --reason completed |
lisa-github-agent |
| Linear (build) | Linear MCP list_issues / get_issue / list_comments |
Linear MCP save_issue (labels) / save_comment |
lisa-linear-agent |
| Notion (PRD) | lisa-notion-access (query, page comments) |
lisa-notion-access write-page (status) / page comment |
lisa-notion-to-tracker (dry-run) |
| GitHub (PRD) | gh issue list/view (PRD labels) / GraphQL sub-issues / generated-work section |
gh issue edit / gh issue comment / gh issue close --reason completed |
lisa-github-to-tracker (dry-run) |
| Linear (PRD) | Linear MCP list_projects / get_project (+ sentinel feedback issue) |
Linear MCP save_project (labels) / save_comment |
lisa-linear-to-tracker (dry-run) |
| Confluence (PRD) | lisa-atlassian-access CQL |
lisa-atlassian-access page parentId update / comment |
lisa-confluence-to-tracker (dry-run) |
Staleness model
An in-progress item (build claimed, PRD in_review) is stalled when the last
state-changing transition into the in-progress role, or the last human / PR-side forward-progress
activity after that transition, is older than the stale_after threshold. blocked items are NOT
gated on staleness — their repairability is judged on current blocker/answer state, not elapsed
time.
Automation self-comments are not forward progress and must not reset the staleness clock. Status
comments like [claude-build-intake] PR remains open..., [codex-build-intake] Follow-up pushed...,
or [lisa-repair-intake] ... may be useful audit notes, but they cannot make a claimed item fresh
forever. If a provider exposes a changelog/history surface, prefer the timestamp of the last
transition into the claimed/In-Progress role over the item's generic updated timestamp. When the
history surface is unavailable, ignore comments whose author/marker clearly belongs to Lisa or its
automation agents, and use the newest human comment/edit or PR-side progress event instead.
A build claimed leaf whose linked PR has already merged (state == MERGED) is likewise NOT
gated on staleness. A merged PR is a settled terminal state, not in-flight work: the only thing
missing is the env transition build-intake never applied (its merge gate left the item claimed
because the merge landed after its agent returned). The recovery is judged on PR merge state, not
elapsed time — and crucially post-merge activity does not defer it. A freshly-merged PR keeps
producing activity that the signal below would otherwise read as keep-alive (a queued/in-progress
release or deploy check-run, a post-merge CodeRabbit summary comment), so gating merged-PR recovery
on staleness strands a completed leaf in claimed for as long as that activity keeps the clock
warm — exactly the failure that leaves a shipped Sub-task showing status:in-progress for a day
while its parents roll up against it. Recover it regardless of recent activity (Build claimed
decision tree step 0, and the dedicated high-confidence ordering bucket).
Threshold resolution
$ARGUMENTSstale_after=<dur>(one-off override) — always wins. ParseNh/Nm/Nd/0into hours..lisa.config.jsonintake.repair.staleAfterHours(durable project default).- Built-in default: 2 hours.
stale_after=0 means "treat any in-progress item as stalled" — a manual full-recovery lever,
and the only way to resume work on a provider that exposes no reliable activity timestamp.
Activity signal (state-change first, portable across vendors)
Compute the item's newest eligible activity timestamp from the highest-priority signal the vendor
exposes, and compare it to now - stale_after:
- Provider-native status/label transition time into the in-progress role, when the provider
exposes it cleanly (JIRA changelog transition to
claimed/ In Progress, GitHub label event, Linear state/label history, Notion/Confluence page move/status history). - Latest human lifecycle/progress comment or edit on the item (and, for Linear PRDs, the
sentinel feedback issue). Exclude automation self-comments and Lisa audit markers such as
[claude-build-intake],[codex-build-intake],[lisa-build-intake], and[lisa-repair-intake]. - For build items, latest PR-side forward-progress activity on the linked PR: newest commit, review, check-run, or PR comment.
- Provider-native item
updatedAt/last_edited_time/updatedonly when the provider cannot expose transition/comment authorship and the timestamp is not known to be driven by automation self-comments.
If ANY of these is newer than the threshold, the item is active → record it as active and
skip it (read-only). For build claimed, an open PR with recent commits/checks is active. For
PRD in_review, a recent comment or page edit is active.
Count only forward-progress signals as keep-alive: new commits, a review that was just
requested or posted, an in-progress/queued check run, a fresh progress comment. A settled
blocker state — a failing/errored check run, CONFLICTING mergeability, a CHANGES_REQUESTED
review, an unaddressed CodeRabbit/reviewer change request, or a failed deployment — is NOT
keep-alive activity: it does not reset the staleness clock. The clock runs from the last genuine
progress event, so a PR that has been sitting failed/conflicted/awaiting-changes for longer than
stale_after counts as stalled and is diagnosed below.
A merged linked PR is the same kind of non-keep-alive signal, in the other direction: the work
is settled and complete, so its post-merge check-runs and summary comments must NOT count as
keep-alive either. A build claimed leaf with a merged PR is recovered regardless of the staleness
clock (see the Staleness model note above and the dedicated ordering bucket); it is never recorded
active and skipped on the strength of post-merge activity.
If a provider cannot expose any reliable timestamp, do not auto-resume its in-progress
items unless the caller passed stale_after=0. (Dependency-cleared blocked repair still
proceeds — it is judged on blocker state, not time.)
Repair decision tree
Apply per candidate. Continue through the ordered list until every candidate inside the
max_candidates cap has been evaluated. Each candidate may trigger a write (lifecycle transition,
native close/archive/complete, re-dispatch, or refreshed note), be recorded read-only, or be
recorded under Errors. Do not stop after the first write; the cap is the batch boundary.
Build claimed (stalled in-progress) → diagnose blocker, else resume in place
First check for an already-merged PR — this check is NOT gated on staleness. Read the item's
linked PR state before applying the staleness gate (see "Stuck-cause diagnosis" step 1–2 for
discovery). If state == MERGED, recover it immediately via step 0's merged-PR arm regardless of
elapsed time or recent post-merge activity (per the Staleness model's merged-PR exemption): a merged
PR is a completed leaf, and deferring it behind the staleness clock is what strands shipped work in
claimed.
Only if the PR is not merged does the staleness gate apply. Once it passes, diagnose why it stalled by inspecting the item's PRs and deploys (see "Stuck-cause diagnosis" below). A stalled build usually stalled for a concrete external reason, and re-dispatching the agent at it will not fix a PR that cannot merge or a deploy that failed — it just churns.
- Diagnose PR & deploy state. Run "Stuck-cause diagnosis" below. It resolves, in order:
- PR already merged (checked first, staleness-exempt) → the build effectively completed; the
vendor build-intake's merge gate left the item
claimedbecause the merge landed after its agent returned. Do not re-dispatch or file anything — apply the scanner's post-agent env-resolvedclaimed → donetransition directly (step 2 below, env-resolved), and record it. This is the recovery arm for build-intake leaving merged-but-unadvanced items inclaimed. - PR only behind its base (a needed rebase) → mechanically resolvable, not a human blocker.
Re-sync the branch in place so the already-enabled auto-merge can land (see diagnosis step 3).
Keep the item
claimed; a later cycle confirms the merge and transitions. Do not file a fix ticket for a clean rebase. - A true merge conflict → not an immediate fix ticket. A conflict is fixable by re-running
the build, so attempt one in-place re-dispatch first (the resume sequence below; the vendor
agent re-enters
drive-pr-to-mergefix mode, which resolves conflicts). Only a conflict that survives that single attempt — the same conflicting head stillCONFLICTINGon a later cycle — or that the agent reports needs design input, falls through to the fix-ticket path (diagnosis step 5). - A real external blocker re-running the build cannot fix (failing checks /
CHANGES_REQUESTED/ unaddressed CodeRabbit; or a failed deploy) → do not dispatch the agent. File a build-ready leaf fix ticket for the blocker, move this itemclaimed → blockedwith anis blocked bylink to that ticket, and record it. The existing "Buildblocked→ unblock if cleared" path resumes this item on a later cycle once the fix ticket is terminal — a self-healing loop. Skip the resume steps below.
- PR already merged (checked first, staleness-exempt) → the build effectively completed; the
vendor build-intake's merge gate left the item
If the PR is healthy in-flight and no blocker is found, the work simply died mid-flight — run the same per-item sequence
the vendor build-intake runs, skipping the claim transition (the item is already claimed):
- Dispatch the item to the vendor agent —
lisa-jira-agent/lisa-github-agent/lisa-linear-agent(matching the queue's tracker) — with the item ref. If repair-intake is running as a teammate rather than the lead/root agent, return a structureddelegation-requestto the lead instead of spawning that named peer yourself; only the lead can add named teammates in Claude's flat roster. This resumes the work in place, preserving its existing branch/PR and prior comments. - On agent success, apply the scanner's post-agent transition yourself:
claimed → done, wheredoneis env-resolved exactly aslisa:<tracker>-build-intakeresolves it (perconfig-resolutionenv-keyeddone: explicittarget_envarg wins; else reverse-lookup the env from the resulting PR's base branch viadeploy.branches; ifdoneis a map and env is unresolvable, fail loudly — never guess). repair-intake owns this transition because it is standing in for the scanner that never got to finish it. - On a surfaced blocker (agent reports it cannot proceed), leave/move the item to
blockedwith a[lisa-repair-intake]note (see Loop prevention). When the surfaced blocker is something only a human can supply (credentials, access/permissions, a product or scoping decision), the item also carries thehuman_neededmarker — the vendor agent's pre-flight gate applies it; if repair-intake makes the block transition itself for such a reason, it adds the marker too. (A block that another tracked ticket or retry will clear is not human-needed — that is the auto-recoverable fix-ticket path above.)
Do not reset stalled in-progress items to
ready. Reset throws away state, makes a partially-built item look freshly human-approved to the nextlisa-intakeclaim, and forces a two-cycle recovery. Resume in place.
Stuck-cause diagnosis: PR & deploy blockers
Run this for every stalled claimed build item before considering an agent re-dispatch. The
goal is to distinguish "work died mid-flight, just resume it" from "work is blocked on a concrete
external state that resuming the agent cannot fix."
1. Find the associated PR(s) and deploy(s). From the item's linked PRs (GitHub: prefer the
native dev-link surface — gh issue view <n> --json closedByPullRequestsReferences — which lists
merged PRs that closed the issue, then gh pr list --search <issue-ref> --state all; JIRA:
dev-status / remote links; Linear: attachments and git-branch links) and the deploy(s) for the
resulting merge (the env-keyed deploy.branches mapping from config-resolution). The --state all
is load-bearing: gh pr list --search defaults to --state open, so a merged (closed) PR is
invisible on that surface — the exact state this recovery path exists to catch. A merged PR linked
only via search (no Closes # / native dev link) would otherwise never be discovered, and the leaf
would never recover. Read each PR with the vendor's native state, e.g. GitHub
gh pr view <n> --json state,mergedAt,mergeable,mergeStateStatus,reviewDecision,statusCheckRollup,comments,reviews.
2. PR already merged → recover, don't re-dispatch. If state == MERGED, the build is effectively
complete and the only thing missing is the env transition the build-intake never applied (its merge
gate left the item claimed because the merge landed after its agent returned). Do not re-dispatch
or file anything: apply the scanner's post-agent env-resolved claimed → done transition (the
resume-sequence step 2, env-resolved from the merged PR's base branch) and record it as a repair
write. Where the env deploy is observable, confirm it did not fail first; a failed post-merge deploy
falls through to the blocker path (step 4).
3. PR only behind its base → re-sync in place (mechanical, not a blocker). If the PR is clean but
behind its base — mergeStateStatus == BEHIND while mergeable != CONFLICTING and no required check
is failing — it does not need a human. This is exactly the case that strands a PR forever: GitHub
auto-merge will not advance a BEHIND branch on its own, so a PR opened with --auto sits unmerged
until something rebases it. Delegate this mechanical nudge + classification to the
drive-pr-to-merge skill in report mode — the single source of truth for the
"ensure auto-merge + re-sync a clean BEHIND branch" primitive — so this scanner does not
re-implement it:
drive-pr-to-merge pr=<n>
In report mode it ensures auto-merge is enabled and runs gh pr update-branch <n> only when the
PR is BEHIND-but-clean and the base branch's ruleset or classic branch protection requires
strict up-to-date status checks (strict_required_status_checks_policy / required_status_checks.strict).
If strict checks are off, it does not update the branch solely because the base moved; that avoids
CI cancellation storms in repos where updating the PR head restarts and cancels in-flight runs. It
never edits code, resolves threads, or dismisses reviews. It returns a classification (merged /
will-merge-after-resync / blocked:<reason>). On a merged / will-merge-after-resync result,
record this as a repair write (resynced), keep the item claimed, and move on — a later cycle sees
the now-CLEAN (or merged) PR and either lets auto-merge finish or applies the merged-PR recovery in
step 2. Only if gh pr update-branch itself reports a conflict it cannot apply does the PR become a
true conflict (step 4). Honor the backoff window so repeated cycles don't re-issue update-branch on
an unchanged head (Loop prevention). For JIRA/Linear items the PR is still the GitHub PR backing the
branch — operate on it the same way.
4. Classify as a blocker. Treat any of these as a real external blocker:
- True merge conflict —
mergeable = CONFLICTINGormergeStateStatus = DIRTY(overlapping changes a plain rebase cannot resolve), orgh pr update-branch(step 3) reported a conflict. A merelyBEHINDbranch is not here — it was re-synced in step 3. Unlike the other classes below, a conflict is resolvable by re-running the build, so step 5 gives it one in-place re-dispatch before filing — see its conflict-first rule. - Failing required checks —
statusCheckRolluphas aFAILURE/ERROR/TIMED_OUTconclusion, ormergeStateStatus = UNSTABLE/BLOCKEDdue to checks. - Change requests outstanding —
reviewDecision = CHANGES_REQUESTED, or unresolved CodeRabbit (or other reviewer) comments that request changes and have not been addressed by a newer commit. - Branch-protection / approvals blocked —
mergeStateStatus = BLOCKEDfor a reason other than a transient check still running. - Failed deploy — the deployment for the item's merge/branch reports a failed/errored status (failed deploy workflow run, failed deployment status, or the project's deploy check is red).
A check that is still queued/in progress, or a CLEAN/HAS_HOOKS mergeable PR with no
outstanding change request, is not a blocker — that is normal in-flight state. (Such a PR with
recent check/commit activity would already have been caught as active by the staleness gate.)
5. On a blocker found → file a leaf fix ticket + block the item.
Conflict-first exception (try to resolve before filing). A true merge conflict — and only a
conflict, not failing checks, change requests, or a failed deploy — is fixable by re-running the build:
the vendor agent's drive-pr-to-merge fix-mode loop resolves conflicts. So before filing a fix ticket
for a conflict, give the item one in-place re-dispatch: run the resume-in-place sequence (steps 1–3
of the parent path above), which re-enters drive-pr-to-merge in fix mode against the existing PR.
Bound it to a single attempt per conflicting head — when you re-dispatch, post a [lisa-repair-intake] conflict-resolve-attempt: <item-ref>@<head-sha> marker keyed on the PR head SHA. On a later cycle, if
that marker already exists for the same head SHA and the PR is still CONFLICTING, the attempt
failed: stop retrying and file the fix ticket below. File immediately (skip the attempt) if the agent
reports the conflict needs design input. Every other blocker class files the fix ticket with no
re-dispatch. Honor the backoff window / state fingerprint (Loop prevention) so the re-dispatch is never
re-issued against an unchanged conflicting head.
- File one build-ready leaf fix ticket per distinct blocker via
lisa-tracker-write(the vendor-neutral leaf writer + validation gate; never a vendor*-write-*skill directly),issue_type: Bugfor a failing-check/conflict/failed-deploy,Taskfor review-feedback follow-up,build_ready: trueso it auto-builds. The ticket MUST name: the blocked item + its PR/deploy URL, the exact blocker (conflict / which checks failed with their logs link / which change requests / which deploy run), three-audience description, and Gherkin acceptance criteria for "PR is mergeable / deploy is green." - Transition the stalled item
claimed → blockedand add anis blocked bylink to the new fix ticket (vendor-native: JIRA issue linkis blocked by; GitHub/LinearBlocked by:line- label). Post a
[lisa-repair-intake]note naming what it is blocked by and why. This block is auto-recoverable — the fix ticket will build and close on its own — so do not add thehuman_neededmarker, and if the item already carries a stalehuman_neededlabel, remove it here (it is no longer waiting on a human). Thehuman_neededmarker is reserved for blocks a human must clear; this one a later cycle clears automatically.
- label). Post a
- Record it as a repair write. Do not dispatch the vendor agent for this item this cycle.
The item now sits in blocked; once the fix ticket reaches a terminal state, the Build
blocked → unblock if cleared path (next section) detects the cleared is blocked by
dependency and resumes the original in place — a self-healing loop.
Idempotency. Before filing, check for an open fix ticket already carrying the marker
[lisa-repair-intake] blocker:<item-ref>/<blocker-key> (blocker-key is a stable slug of the
blocker, e.g. pr-1234/merge-conflict or pr-1234/checks-failing). If one exists, reference it
and ensure the is blocked by link is present rather than creating a duplicate. Honor the backoff
window and state fingerprint (Loop prevention) so re-runs over the same unchanged blocker are no-ops.
Build blocked → re-evaluate, unblock if cleared
- Read the block reason and classify the blocker (see Blocker classification & clearing). An item
may be held by a dependency, by a validation / quality-gate self-block, by a
deployed / runtime verification failure, by an ambiguity, or by more than one at once.
Re-check every class present — do not stop at "no
is blocked bylinks, therefore nothing to do." A self-block has zero dependencie
…(truncated)