# Agent Capability

> Add or change what an agent can do — a new tool the model can call, a new capability (knowledge, web search, charts, a guardrail, a compaction strategy), a tool rename, or a per-tool approval gate. Use whenever the ask is "give the agent a new tool/function/action", "add a capability", "let the agent do X". There is no @agent.tool and no assistant module in this project; every tool arrives through the capability registry.

- Skill: `vstorm-co/agent-capability` (Agent Skill, multi-file: 4 files)
- Install (CLI): `npx skillmds@latest add vstorm-co/agent-capability`
- Raw SKILL.md: https://api.skillmd.com/api/skills/vstorm-co/agent-capability/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: vstorm-co (https://skillmd.com/u/vstorm-co)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/vstorm-co/agent-capability

---


# Capabilities — the only way a tool reaches a model

**Read `docs/howto/add-capability.md` first.** It is the walkthrough and it is
current. This file is the decision layer around it: what shape the work should
take, and the traps that are silent.

An agent here is *data*. There is no module with an `Agent` object to decorate —
`@agent.tool`, `RunContext[Deps]` and `app/agents/assistant.py` do not exist. An
agent is assembled per run from the capabilities its spec names, so a new tool is
always part of a capability.

## Pick the shape first

| The ask | Do this |
|---|---|
| A tool that belongs to a decision somebody already makes | Add it to that capability's `_toolset.py` **and** its `tools=` tuple |
| A genuinely new switch an agent author would want | New capability package |
| Behaviour change, not a tool (reasoning, a guard, injected context) | New capability with `tools=()` — see `clock/`, `thinking/` |
| A third-party SaaS that already publishes an MCP server | **No code.** See the `mcp-connections` skill and `docs/mcp.md` |
| Same tool, different wording for one agent | Not code either — `tool_overrides` in that agent's spec |

That last row catches a lot of requests. A tool's name and description *are*
prompt, and an agent needing different behaviour from the same tool usually needs
them reworded per agent, not a second tool written.

## Layout, which a test enforces

```
backend/app/agents/capabilities/<name>/
  __init__.py       @register(...) — and nowhere else
  _capability.py    the AbstractCapability subclass
  _toolset.py       the tools, and the text the model reads
  README.md         why this exists and what it deliberately does not do
```

`tests/test_capability_layout.py` enforces all four. `@register` in a submodule
only fires if something imports that module, which is how a capability vanishes
from the Builder with every test still green.

A capability whose tools come from a library has no `_toolset.py` — `skills` and
`sandbox` are both in the test's `EXTERNAL_TOOLSET` set. The tool *text* is still
this repository's: declare it once in `_capability.py` and hand the same
descriptions to the library, so the catalog and the model read the same sentence
rather than two copies drifting in two repositories.

Read `clock/` for the smallest complete example, `knowledge/` for one with a config
schema, resources and a scope, `web_research/` for a conditional secret
requirement, `sandbox/` for per-tool approval and a resource the runner resolves.

## When one flag cannot describe the whole capability

`side_effecting` on `@register` is the capability's answer, and
`CapabilityToolInfo.side_effecting` overrides it per tool. Use the per-tool form
when a capability genuinely both reads and writes: marking the whole of `sandbox`
side-effecting makes an agent ask permission to list a directory, and not marking
it lets a write run unattended. An author would then hand-write a `tool_approval`
override per tool in every spec, and the one they forget is the dangerous one.

`None` — the default — defers to the capability, so every capability that
declares nothing behaves exactly as it did. A binding's `tool_approval` still
beats both: that is the operator's decision, and it wins over the code's.

## The three things that fail silently

**1. A tool missing from `tools=`.** It still runs. It just cannot be approved,
cannot be renamed per agent, and does not appear in the Builder. The dangerous
half is that an author adds a *side-effecting* tool, forgets to declare it, and it
runs unattended forever. `tests/test_capability_registry.py` is the drift test that
catches this — it builds every registered capability and compares the declared list
against the tools the model is actually offered.

**2. A module missing from `load_builtins()`.** The capability does not exist as far
as the Builder is concerned. Registration is an import, not a scan.

**3. A changed `id`.** Ids are in every published spec and in clients' git
repositories. Rename the class freely; never re-id.

## What the tool says, and what it says when the call was wrong

Both are prompt, and both have a house shape - `references/tool-text-and-failures.md`
has it, with the worked examples. The two things that keep being missed:

- **A docstring with no `Returns:`.** The model cannot infer the shape of the
  answer, that a list stopped at 100 entries, or that a failure still carries
  output. Every tool here now says so; a new one that does not is the odd one out.
- **A bare `raise ModelRetry`.** Use `steer(ctx, ...)` from
  `capabilities/_failures.py`. A retry past a tool's budget - one attempt, by
  default - does not fail the call, it ends the run with
  `UnexpectedModelBehavior`, so the same malformed chart sent twice takes the
  conversation with it. `steer` returns the message on the last attempt instead.
  A tool that needs it takes `ctx`; `RunContext` never reaches the schema.

And what stays a returned string: a result that is bad news, and a refusal. A
retry prompt on a refusal invites the model to look for a way around it.

## Then

- Add the module to `load_builtins()` in `_registry.py`.
- Write the `README.md`. Reasoning goes there, not in the commit message.
- Test it — `app/agents/**` is at **100% and CI fails below it**. See the
  `backend-tests` skill.
- If the capability needs a credential, declare it as a `SecretRequirement` *kind*,
  never an instance. See the `vault-secrets` skill.
- Update `docs/reference/capabilities.md` — it is the human-readable catalog and it
  is a snapshot of the registry.

## Depth

- `references/registry-contract.md` — every `@register` argument, `ctx.resources`,
  returning `None`, scopes, conditional secrets, returning a capability we did not
  write.
- `references/tool-text-and-failures.md` — the four parts of a tool's docstring,
  and which failures steer the model rather than answering it.
- `references/approval-and-overrides.md` — how `side_effecting`, `approval`,
  `tool_approval` and `tool_overrides` resolve, and why everything keys on the
  stable tool id.
- `docs/reference/capabilities.md` — what ships today.

