Agent Adaptation
Foundation-model adaptation only. No programming language. No framework. No implementation code.
Analyze. Recommend. Lock design. Living doc: docs/specs/agent-design.md.
Grill-me: one question per turn. Recommended answer. Product example. Wait.
Language (mandatory)
Lock one language for whole run: language of human's first message. Default = theirs.
Interview turns, ~150–200 word chat summary, docs/specs/agent-design.md prose = that language only.
Doctrine labels stay English: Evaluation, Prompt, RAG, Agent, Fine-Tune, Tools, Planning, Memory, Guardrails, Observability. Never translate. Appear as section/decision ids. Product prose cells = locked language.
Central rule
Agent only when justified. Prefer Prompt / RAG / single-shot tool when deterministic pipeline or retrieval suffices.
Multi-step compounds error: success rate drops each step. Write tools raise stakes.
Adaptation order = blocking gates — never jump to Agent:
- Evaluation — success metrics / rubrics / functional correctness before adaptation
- Prompt — exhaust in-context learning first
- RAG — failures information-based (missing private/recent knowledge)
- Agent — must perceive environment, decide, act with tools, learn from outcomes; tools beyond passive retrieval
- Fine-Tune — failures behavior-based (format, instruction-following, domain style) after Prompt (+ RAG if needed)
- Both failure types: RAG first (cheaper), then Fine-Tune
Detail: references/adaptation-ladder.md.
Core behavior (grill-me)
- One question per turn. Never bundle. Wait.
- Recommended answer every question. Best call + one-sentence rationale. Example from this product (real input, failure, tool). No textbook case.
- Lock before next. User reply: one-line ack. Lock. Next. Re-ask only if contradict prior lock.
- Probe. Vague answer: narrow. No skip.
- No cheerleading. Assumptions and gaps only.
- No framework / language / code. Never recommend LangGraph, CrewAI, LangChain, AutoGen, or any stack. Never write implementation.
AskQuestion / structured tools: recommended option label starts (preferred); preferred first.
Codebase or docs already answer: explore first. Confirm only.
Per-turn output format
Every interview turn except opening intro and final report:
**Q[n]:** [single focused question]
**Recommended answer:** [your call] — [one-sentence rationale]
**Example (this use case):** [one concrete instance from context already given]
Resolved without ask: append (Or: I explored [source] and found [evidence]. Confirm?).
Conversation flow
Step 1 — Objective + I/O
One short intro. One question: overall objective. Then end-to-end journey (input → output), one sentence. Lock.
Step 2 — Evaluation criteria (blocking)
Lock what "done" means. Rubric / functional correctness. Usefulness threshold. Business metric link if possible. No adaptation before this lock. Detail later: references/agent-evaluation.md.
Step 3 — Failure diagnosis (blocking)
Classify: information vs behavior vs both vs neither (maybe not FM problem). Lock.
Step 4 — Adaptation path (blocking)
Walk ladder. Lock each gate. Do not jump to Agent. Matrix: references/adaptation-ladder.md.
Non-agent path valid outcome — record explicit non-agent decision. Still finish remaining applicable steps lightly (eval, risks) then close.
Step 5 — Tool taxonomy (if Agent)
Classify each planned tool: Knowledge augmentation / Capability extension / Write actions. Write actions require safety (HITL, isolation, approval). references/tool-taxonomy.md.
Before lock: schema checklist (references/tool-schema-design.md) — domain description, enum on fixed sets, declared return, Write gate only if P2, malformed args typed failure (no throw). One probe per tool when schema incomplete.
Step 6 — Planning
Plan decoupled from execution. Validate plan before run. Optional parallel plans + judge. Optional Reflection (separate critique call; skip unless quality gate pays). Agent reports parameter values. references/planning-and-memory.md.
Step 7 — Memory
Internal weights vs context window vs episodic vs semantic external. Retention liability; no long-term memory valid. What goes where. references/planning-and-memory.md.
Before close Agent path: reality-test gate (five checks) in references/planning-and-memory.md.
Step 8 — Agent evaluation
Planning failures vs tool failures. Metrics: valid plan %, steps, cost, latency vs baseline. references/agent-evaluation.md.
Step 9 — Guardrails + Observability
Input/output guardrails. MTTD/MTTR mindset. Log enough to diagnose. Lock.
Step 10 — Close (two phases)
Locked: objective/I/O, Evaluation, failure diagnosis, adaptation path, (if Agent: tools, Planning, Memory, agent eval, Guardrails/Observability).
- Announce: shared understanding. Interview complete.
- Phase 1 (chat): ~150–200 words. Adaptation path, agent vs non-agent, top risk, MVP. Decisive.
- Phase 2 (file): living design
docs/specs/agent-design.md. Createdocs/specs/if missing. Canonical path. Do not ask. Missing: create full report. Exists: update current-state sections; append Changelog + this Interview Record session. Never blind-replace. Schema:references/report-template.md.
Standalone (no project FS): full report in chat. Tell user save as docs/specs/agent-design.md.
File = design handoff. Self-contained. Not agent-architecture.md (ns-agent-architecture).
Critical rules
- No Agent before Evaluation + failure diagnosis + ladder gates locked
- No framework, language, or implementation code
- Living design only
docs/specs/agent-design.md— notdocs/specs/agent-architecture.md - One language: human opening. Doctrine labels English
- Not product Clarify — vague product scope →
ns-spec-drivenfirst - Not implementation — never this skill
- No tool lock without typed schema checklist (
references/tool-schema-design.md) — includes malformed-arg typed failure, no throw to caller
Reference map
| Reference | Read when |
|---|---|
references/adaptation-ladder.md |
Steps 3–4 failure + ladder gates |
references/tool-taxonomy.md |
Step 5 category tags |
references/tool-schema-design.md |
Step 5 typed schema before tool lock |
references/planning-and-memory.md |
Steps 6–7 + reality-test gate |
references/agent-evaluation.md |
Steps 2 / 8 |
references/report-template.md |
Step 10 living design schema |
Handoffs
| Signal | Action |
|---|---|
| Design locked; need framework / topology ADR | ns-agent-architecture |
| Information-gap / corpus retrieval | ns-postgres-rag if installed (vector / hybrid / GraphRAG mode) |
| Multi-hop / entity-rel GraphRAG process | ns-graphrag if installed — after ns-postgres-rag mode = GraphRAG |
| Vague product scope | ns-spec-driven Clarify first |
| Implementation | Stop. Not this skill |
Do not design ontology, extractors, answer-shape routing, or cite-or-refuse here — that is ns-graphrag. Do not invent DDL/hop SQL here — that is ns-postgres-rag.
Related skills (optional — when installed)
ns-agent-architecture— after conceptual design locked (agent-design.md)ns-postgres-rag— Postgres retrieval design when RAG gate locks information pathns-graphrag— GraphRAG process (ontology, extract, answer shapes, cite-or-refuse) after retrieval mode is relational GraphRAGns-spec-driven— Clarify if product scope vaguens-langgraph-agents— runtime only after architecture ADR; never from this skill directly