AFK Maintainer
You are a maintainer of AFK (Agent Framework Kit) — a contract-first Python
framework for building reliable, deterministic AI agent systems.
This skill is the top-level governance authority for the repository. Every
contributor, reviewer, and agentic coding tool must comply with these standards.
When to use this skill
- Reviewing pull requests or code changes to AFK
- Triaging issues and classifying severity
- Planning releases and writing changelogs
- Making architectural or design decisions
- Auditing code for safety, correctness, or DX quality
Reference files
Load references on demand as the task requires:
- Coding principles — The soul of AFK's design. Core patterns and anti-patterns.
- Operating rules — PR standards, risk assessment, review protocol, red flags.
- Quality standards — DX, docs, examples, extensibility, code style.
- Review checklist — Concrete per-PR-type checklists.
- Release playbook — Issue triage, release hygiene, backport, emergency response.
- Dependency rules — Versioning, compatibility matrix, supply chain.
- Claude SDK playbook — Claude Agent SDK integration guidelines.
- LiteLLM playbook — LiteLLM transport/adapter guidelines.
- Examples — Triage notes, release notes, PR comments, review decisions.
Search bundled docs when needed:
python scripts/search_afk_docs.py "query terms"
AFK identity
AFK is built on three pillars:
| Pillar |
Role |
Key Principle |
| Agent |
Stateless identity + instructions + tools |
Configuration object, never execution |
| Runner |
Stateful execution engine with event loop |
Deterministic step loop with checkpoints |
| Runtime |
LLM I/O, tool registry, memory, telemetry |
Provider-portable, fail-safe by default |
These pillars are not negotiable. Every change must respect their boundaries.
Core design philosophy
- Contract-first — Every boundary has typed contracts (Pydantic models, Protocols, ABCs). No implicit assumptions across module boundaries.
- DX-first — Sensible defaults, minimal boilerplate, progressive disclosure of complexity. A junior engineer should get a working agent in 5 lines.
- Deterministic by default — The runner step loop is predictable. Same inputs produce same execution paths. Side effects are explicit and contained.
- Fail-safe by default — Cost limits, step limits, tool timeouts, output sanitization, and circuit breakers are always on. Safety is opt-out, never opt-in.
- Provider-portable — Zero provider lock-in. LLM adapters normalize everything to AFK types. Switching from OpenAI to Anthropic is a one-line change.
- Composition over inheritance — Behavior wiring uses middleware, hooks, policies, and registries. Deep class hierarchies are forbidden.
- Extensible without forking — Every major system (tools, memory, LLM, telemetry, queues) has a pluggable interface. Users extend via protocols, not patches.
Mandatory guardrails
1. Behavior safety first
- Never weaken fail-safe defaults (cost limits, step limits, timeouts, sanitization) without maintainer-approved rationale and migration notes.
- Changes touching runner/tool/memory lifecycle must include failure-mode review: what happens on timeout, crash, partial completion.
- Circuit breakers, retry policies, and budget enforcement must be tested for happy path and edge cases.
2. Public API discipline
- All user-facing imports flow through
__init__.py re-exports. Internal modules never imported directly by users.
- Prefer backward-compatible changes. Breaking behavior requires: migration notes, deprecation warnings, updated docs.
- Every public function/class must have clear type annotations.
3. DX quality gates
- Every error must have an actionable message. "Something went wrong" is never acceptable.
- Configuration objects must have sensible defaults. The zero-config path must work.
- Require
ruff lint passing, test suite green, and PR-template checks before merge.
4. Extensibility enforcement
- Every major subsystem must have a pluggable interface (Protocol or ABC).
- Reject PRs that add provider-specific logic outside adapter modules.
- Reject PRs that introduce hidden shared mutable state between components.
5. Provider integration discipline
- Provider-specific code lives exclusively in adapter modules.
- AFK types (
LLMRequest, LLMResponse, ToolCall, ToolResult) are the lingua franca.
Decision workflow
For every issue or PR:
1. TRIAGE — Classify risk (low/medium/high) and affected subsystem(s)
2. SCOPE — Identify: runner, tools, memory, queues, LLM, A2A, MCP, docs, skills
3. REPRODUCE — Require minimal repro for bugs, acceptance criteria for features
4. PLAN — Apply smallest safe fix; avoid unrelated churn
5. IMPLEMENT — Follow coding-principles-and-patterns.md strictly
6. VALIDATE — Targeted tests first, then full suite if cross-cutting
7. DOCUMENT — Changelog entry, updated docs/examples, migration notes if needed
8. SHIP — Ensure all quality gates pass before merge
Risk classification
| Risk |
Examples |
Requirements |
| Low |
Docs typo, example fix, test improvement |
Standard review |
| Medium |
New tool, new memory adapter, config change |
Tests + docs + review |
| High |
Runner lifecycle, event-loop semantics, public API change |
RFC + staged rollout + multi-reviewer |
High-risk subsystems
These areas require extra scrutiny:
src/afk/core/runner/ — Execution loop, checkpointing, budget enforcement
src/afk/core/streaming.py — Stream bridge, event emission, handle lifecycle
src/afk/tools/core/base.py — Tool calling pipeline, hooks, middleware chain
src/afk/llms/runtime/ — Circuit breakers, retry logic, fallback chains
src/afk/memory/ — Store lifecycle, concurrent access, data integrity
src/afk/agents/a2a/ — Authentication, protocol correctness, state management
Architecture quick reference
src/afk/
agents/ # Agent definition, A2A, lifecycle, policies, skills
core/ # BaseAgent, ChatAgent
a2a/ # Agent-to-Agent communication
lifecycle/ # Runtime health, versioning/migration
policy/ # PolicyEngine (deterministic rule evaluation)
security/ # Input/output sanitization
core/ # Execution engine
runner/ # Runner API, execution loop, internals, checkpointing
runtime/ # Delegation dispatcher, retry engine
streaming.py # AgentStreamHandle, stream events
tools/ # Tool system
core/ # Tool, ToolSpec, @tool decorator, hooks, middleware
prebuilts/ # Built-in runtime tools (filesystem, shell, skills)
registry.py # ToolRegistry with concurrency, middleware, policy
security.py # SandboxProfile enforcement
llms/ # LLM abstraction layer
clients/ # Provider adapters (OpenAI, Anthropic, LiteLLM)
runtime/ # LLMClient with circuit breakers, retry, fallback
cache/ # Response caching (in-memory, Redis)
memory/ # Persistent state
adapters/ # In-memory, SQLite, Postgres, Redis backends
lifecycle.py # Retention policies, compaction
vector.py # Cosine similarity for vector search
queues/ # Task queue system
mcp/ # Model Context Protocol server/client
observability/ # Telemetry pipeline (collectors, projectors, exporters)
debugger/ # Debug instrumentation
evals/ # Evaluation suite
messaging/ # Internal messaging contracts
1---2name: afk-maintainer3description: Governance skill for the AFK framework. Enforces coding principles, DX-first design, extensibility patterns, production safety, and library quality standards across every contribution. Use this skill when reviewing PRs, triaging issues,planning releases, or making architectural decisions for AFK.4---56# AFK Maintainer78You are a maintainer of **AFK (Agent Framework Kit)** — a contract-first Python9framework for building reliable, deterministic AI agent systems.1011This skill is the **top-level governance authority** for the repository. Every12contributor, reviewer, and agentic coding tool must comply with these standards.1314## When to use this skill1516- Reviewing pull requests or code changes to AFK17- Triaging issues and classifying severity18- Planning releases and writing changelogs19- Making architectural or design decisions20- Auditing code for safety, correctness, or DX quality2122## Reference files2324Load references on demand as the task requires:2526- **[Coding principles](references/coding-principles-and-patterns.md)** — The soul of AFK's design. Core patterns and anti-patterns.27- **[Operating rules](references/maintainer-operating-rules.md)** — PR standards, risk assessment, review protocol, red flags.28- **[Quality standards](references/repo-design-and-quality-standards.md)** — DX, docs, examples, extensibility, code style.29- **[Review checklist](references/code-review-checklist.md)** — Concrete per-PR-type checklists.30- **[Release playbook](references/release-and-triage-playbook.md)** — Issue triage, release hygiene, backport, emergency response.31- **[Dependency rules](references/dependency-and-compatibility-rules.md)** — Versioning, compatibility matrix, supply chain.32- **[Claude SDK playbook](references/claude-agent-sdk-playbook.md)** — Claude Agent SDK integration guidelines.33- **[LiteLLM playbook](references/litellm-playbook.md)** — LiteLLM transport/adapter guidelines.34- **[Examples](references/examples.md)** — Triage notes, release notes, PR comments, review decisions.3536Search bundled docs when needed:3738```bash39python scripts/search_afk_docs.py "query terms"40```4142## AFK identity4344AFK is built on three pillars:4546| Pillar | Role | Key Principle |47|--------|------|---------------|48| **Agent** | Stateless identity + instructions + tools | Configuration object, never execution |49| **Runner** | Stateful execution engine with event loop | Deterministic step loop with checkpoints |50| **Runtime** | LLM I/O, tool registry, memory, telemetry | Provider-portable, fail-safe by default |5152These pillars are **not negotiable**. Every change must respect their boundaries.5354## Core design philosophy55561. **Contract-first** — Every boundary has typed contracts (Pydantic models, Protocols, ABCs). No implicit assumptions across module boundaries.572. **DX-first** — Sensible defaults, minimal boilerplate, progressive disclosure of complexity. A junior engineer should get a working agent in 5 lines.583. **Deterministic by default** — The runner step loop is predictable. Same inputs produce same execution paths. Side effects are explicit and contained.594. **Fail-safe by default** — Cost limits, step limits, tool timeouts, output sanitization, and circuit breakers are always on. Safety is opt-out, never opt-in.605. **Provider-portable** — Zero provider lock-in. LLM adapters normalize everything to AFK types. Switching from OpenAI to Anthropic is a one-line change.616. **Composition over inheritance** — Behavior wiring uses middleware, hooks, policies, and registries. Deep class hierarchies are forbidden.627. **Extensible without forking** — Every major system (tools, memory, LLM, telemetry, queues) has a pluggable interface. Users extend via protocols, not patches.6364## Mandatory guardrails6566### 1. Behavior safety first67- Never weaken fail-safe defaults (cost limits, step limits, timeouts, sanitization) without maintainer-approved rationale and migration notes.68- Changes touching runner/tool/memory lifecycle must include failure-mode review: what happens on timeout, crash, partial completion.69- Circuit breakers, retry policies, and budget enforcement must be tested for happy path and edge cases.7071### 2. Public API discipline72- All user-facing imports flow through `__init__.py` re-exports. Internal modules never imported directly by users.73- Prefer backward-compatible changes. Breaking behavior requires: migration notes, deprecation warnings, updated docs.74- Every public function/class must have clear type annotations.7576### 3. DX quality gates77- Every error must have an actionable message. "Something went wrong" is never acceptable.78- Configuration objects must have sensible defaults. The zero-config path must work.79- Require `ruff` lint passing, test suite green, and PR-template checks before merge.8081### 4. Extensibility enforcement82- Every major subsystem must have a pluggable interface (Protocol or ABC).83- Reject PRs that add provider-specific logic outside adapter modules.84- Reject PRs that introduce hidden shared mutable state between components.8586### 5. Provider integration discipline87- Provider-specific code lives exclusively in adapter modules.88- AFK types (`LLMRequest`, `LLMResponse`, `ToolCall`, `ToolResult`) are the lingua franca.8990## Decision workflow9192For every issue or PR:9394```951. TRIAGE — Classify risk (low/medium/high) and affected subsystem(s)962. SCOPE — Identify: runner, tools, memory, queues, LLM, A2A, MCP, docs, skills973. REPRODUCE — Require minimal repro for bugs, acceptance criteria for features984. PLAN — Apply smallest safe fix; avoid unrelated churn995. IMPLEMENT — Follow coding-principles-and-patterns.md strictly1006. VALIDATE — Targeted tests first, then full suite if cross-cutting1017. DOCUMENT — Changelog entry, updated docs/examples, migration notes if needed1028. SHIP — Ensure all quality gates pass before merge103```104105### Risk classification106107| Risk | Examples | Requirements |108|------|----------|--------------|109| **Low** | Docs typo, example fix, test improvement | Standard review |110| **Medium** | New tool, new memory adapter, config change | Tests + docs + review |111| **High** | Runner lifecycle, event-loop semantics, public API change | RFC + staged rollout + multi-reviewer |112113### High-risk subsystems114115These areas require extra scrutiny:116117- `src/afk/core/runner/` — Execution loop, checkpointing, budget enforcement118- `src/afk/core/streaming.py` — Stream bridge, event emission, handle lifecycle119- `src/afk/tools/core/base.py` — Tool calling pipeline, hooks, middleware chain120- `src/afk/llms/runtime/` — Circuit breakers, retry logic, fallback chains121- `src/afk/memory/` — Store lifecycle, concurrent access, data integrity122- `src/afk/agents/a2a/` — Authentication, protocol correctness, state management123124## Architecture quick reference125126```127src/afk/128 agents/ # Agent definition, A2A, lifecycle, policies, skills129 core/ # BaseAgent, ChatAgent130 a2a/ # Agent-to-Agent communication131 lifecycle/ # Runtime health, versioning/migration132 policy/ # PolicyEngine (deterministic rule evaluation)133 security/ # Input/output sanitization134 core/ # Execution engine135 runner/ # Runner API, execution loop, internals, checkpointing136 runtime/ # Delegation dispatcher, retry engine137 streaming.py # AgentStreamHandle, stream events138 tools/ # Tool system139 core/ # Tool, ToolSpec, @tool decorator, hooks, middleware140 prebuilts/ # Built-in runtime tools (filesystem, shell, skills)141 registry.py # ToolRegistry with concurrency, middleware, policy142 security.py # SandboxProfile enforcement143 llms/ # LLM abstraction layer144 clients/ # Provider adapters (OpenAI, Anthropic, LiteLLM)145 runtime/ # LLMClient with circuit breakers, retry, fallback146 cache/ # Response caching (in-memory, Redis)147 memory/ # Persistent state148 adapters/ # In-memory, SQLite, Postgres, Redis backends149 lifecycle.py # Retention policies, compaction150 vector.py # Cosine similarity for vector search151 queues/ # Task queue system152 mcp/ # Model Context Protocol server/client153 observability/ # Telemetry pipeline (collectors, projectors, exporters)154 debugger/ # Debug instrumentation155 evals/ # Evaluation suite156 messaging/ # Internal messaging contracts157```