AI Harness
Use This For
Use this skill for work involving @purista/harness, @purista/harness-openai, or addon packages named @purista/harness-*.
Core Model
@purista/harness is a standalone, ESM-only agent runtime. It composes typed model aliases, tools, skills, agents, workflows, state, memory, sandboxing, logging, telemetry, and streaming behind one session API.
Keep these layers separate:
- configuration:
defineHarness()registers adapters, defaults, models, tools, skills, agents, and workflows - execution:
harness.getSession(id)returns typedsession.agents.*andsession.workflows.* - adapter code: provider, state, memory, sandbox, MCP, durable runtime, logger, and telemetry ports
- application integration: HTTP/SSE, queues, persistence, auth, and business state stay outside the harness unless represented by a port or tool
- optional governance: policy-as-code for tool decisions, approvals, audit, and policy-pack adapters is configured only when needed; ordinary agents do not require policy setup
Hard Rules
- Use
defineHarness()as the sole construction path. Do not invent standalonedefineAgent,defineWorkflow,defineTool,defineSkill, ordefineModelhelpers. - Use
defineHarnessModule<Required>()('module.id', { register })only for local static composition. Modules contribute normal definitions to the caller's builder; they are not remote plugins, manifests, loaders, or lifecycle owners. Module callbacks cannot build or recursively use a harness. - Use
@purista/harness-agent-pluginsonly for data-only Agent Plugins v1 packages. Require application-owned source/digest/trust approval, inspect diagnostics, then bind selected skills and MCP tools explicitly. A selected stdio server also needs an existing caller-owned data directory and an isolating sandbox implementing bothspawnandmountReadOnly; the local host-directory sandbox does not qualify. Never load package code, auto-install dependencies, auto-expose tools, or accept plugin-provided credentials. - MCP is a clean v2 integration pinned to
2026-07-28: use@modelcontextprotocol/client, modern stateless Streamable HTTP, and a spawn-capable sandbox for stdio. Do not add legacy MCP, HTTP+SSE, one-shot exec, or compatibility fallbacks. - Module definition ids compose additively. Treat duplicate module/definition ids as configuration errors; inspect only
harness.inspect().modulesfor content-free provenance. - Preserve builder inference by declaring models before agents and agents before workflows.
- Use inline helper callbacks for agents and workflows:
.agents(({ agent }) => ({ ... }))and.workflows(({ workflow }) => ({ ... })). - Child-agent delegation is disabled by default. Any workflow that calls
ctx.agents.<id>(input)must declareworkflow.delegation; preferdelegation.agentsallowlists and document budget/model overrides there. - Use
ctx.fanOut(...)for ordered, bounded workflow batches. Usectx.childTasks.start(...)only for workflow-owned isolated background work; task turns queue under the delegation parallel ceiling and never inherit parent history or widen agent permissions. mode: 'continuable'keeps an isolated in-process task conversation open for explicitsend(...)turns andclose(). Do not use it for durable workflow execution or claim cross-process recovery; use an application queue/worker adapter when work must survive a restart.- Configure
defaults.historyRetentionfor durable conversations that need a storage bound. It retains complete newest turns only and requires an atomicStateStore.replaceMessages;maxBytesis serialized UTF-8 storage size, never a token estimate. Use the model's context window/token tooling separately when selecting request context. - For at-least-once direct-agent delivery, pass the transport's stable message or delivery id as
InvokeOptions.idempotencyKey. Replaying the same successful invocation returns its recorded output without a second provider call or transcript; never derive this key from prompt content. - Use
session.release()at the end of an idle request to close live sandbox/MCP resources while preserving StateStore-backed history and runs.session.close()is destructive: it deletes the session record, history, runs, and persisted events. - Declare model capabilities truthfully. Capability arrays gate both TypeScript handles and runtime behavior.
- Prefer
object/object_streamfor structured generation. Do not use legacyjsoncapability names. - Keep RAG orchestration in application/workflow code. The harness provides embeddings and rerank operations, not vector storage.
- Keep HTTP/SSE protocol mapping outside the harness. Harness streams are typed
RunEventvalues. - Do not import PURISTA framework packages from harness or harness addon packages.
- Do not leak prompts, documents, tool inputs, or secrets through logs or telemetry.
telemetry({ contentCaptureMode: 'NO_CONTENT' })is the production default. - Skills are mounted files, not prompt text. Register directories with
.skills(...), allowlist skill ids per agent, keepreadavailable for skill-backed agents, and verifySKILL.mdbodies are not inlined into prompts, logs, traces, or persisted events. - Prefer
ctx.metricsfor application-owned counters, histograms, and operation durations inside workflow handlers, custom agent handlers, and TypeScript tool handlers. Do not call the low-levelTelemetryShimdirectly for app metrics. - Governance policy is optional and late-bound through
.governance(...)after agents/workflows are declared. Keep simple use cases on per-agent permissions; use governance only for composable/audited policy, approval, or external policy-pack interoperability.
Default Workflow
- Inspect implementation first when behavior matters:
packages/harness/src/harness/defineHarness.ts,models/registry.ts,agents/index.ts,skills/index.ts,ports/*, and provider package source. - Decide whether the task is one agent loop, a custom handler agent, or an orchestrating workflow.
- Define Zod schemas at every agent, workflow, and tool boundary.
- Configure model aliases with model-specific provider options, defaults, and the minimal required capabilities.
- Attach tools, skill directories, permissions, sandbox, memory, state, runtime requirements, logger, and telemetry explicitly.
- Decide which state is durable: session history/runs use
StateStore, session memory usesMemoryAdapter, and provider context is transient. Bound durable history with whole-turn retention and bound request context with the model's context/token limits. - Invoke through
harness.getSession(id), release idle sessions, and shut down the shared harness during process shutdown. Use destructive session close only for explicit conversation deletion. - Test with
@purista/harness/testingfakes/contracts before live-provider smoke tests. - For provider-loop regression tests, use the explicit sanitizer recorder and offline replay provider; do not capture production interaction content. Use diagnostic invariants only as explicitly invoked test checks.
Quick Pattern
import { z } from 'zod'
import { defineHarness, JsonLogger, inMemorySandbox } from '@purista/harness'
import { openai } from '@purista/harness-openai'
const harness = defineHarness({ name: 'support-ai' })
.logger(new JsonLogger({ level: 'info' }))
.telemetry({ contentCaptureMode: 'NO_CONTENT' })
.sandbox(inMemorySandbox())
.defaults({
historyRetention: { maxTurns: 50, maxBytes: 256_000 }
})
.models({
assistant: {
provider: openai({ apiKey: process.env.OPENAI_API_KEY! }),
model: process.env.OPENAI_MODEL ?? 'gpt-5-mini',
capabilities: ['object', 'tool_use']
}
})
.tools({
lookup_ticket: {
description: 'Look up one support ticket by id.',
input: z.object({ id: z.string() }),
output: z.object({ status: z.string(), summary: z.string() }),
handler: async (_ctx, input) => ({ status: 'open', summary: `Ticket ${input.id}` })
}
})
.agents(({ agent }) => ({
triage: agent({
model: 'assistant',
input: z.object({ ticketId: z.string() }),
output: z.object({ priority: z.enum(['low', 'normal', 'high']), reason: z.string() }),
builtinTools: false,
tools: ['lookup_ticket'],
instructions: 'Use lookup_ticket, then return a validated triage object.'
})
}))
.workflows(({ workflow }) => ({
triage_ticket: workflow({
input: z.object({ ticketId: z.string() }),
output: z.object({ priority: z.string(), reason: z.string() }),
delegation: { agents: ['triage'] },
handler: (ctx) => {
ctx.metrics.counter('support.triage.started', 1)
return ctx.metrics.duration('support.triage.duration', undefined, () => ctx.agents.triage(ctx.input))
}
})
}))
.build()
const session = await harness.getSession('tenant-a:user-42')
const result = await session.workflows.triage_ticket.prompt({ ticketId: 'T-123' })
await session.release()
await harness.shutdown()
The example's in-memory state store supports atomic history replacement for
local use. Production history retention needs a durable StateStore adapter that
implements replaceMessages atomically.
Read If Needed
references/configuration.mdfor package setup, builder order, sessions, state, sandbox, runtime capabilities, streaming, and shutdown.references/model-setup.mdfor provider aliases, OpenAI setup, defaults, capability-gated model handles, multimodal content, embeddings, and rerank.references/agents-workflows-tools.mdfor deciding between agents/workflows and wiring typed tools, permissions, MCP, and skill-mounted agents.references/agents-workflows-tools.mdalso covers optional governance policy and when to prefer it over simple permissions.references/skills.mdfor creating harness skill folders and registering/mounting them correctly.references/sandbox.mdfor in-memory/bash sandboxes, filesystem/exec APIs, snapshots, built-in tool risk, and custom sandbox adapters.references/state-sessions-streaming-errors.mdforStateStore, session lifecycle, memory/history, run events, error mapping, and replay.references/durable-feedback-operations.mdfor durable runtime checkpoints, adapter capabilities, feedback records, readiness, and operational runbooks.references/telemetry-observability.mdfor OpenTelemetry setup,TelemetryShim, span/metric names, logs, privacy, and adapter context propagation.references/adapters.mdfor creating and using provider, state store, memory, sandbox, durable runtime, logger, telemetry, tool/MCP, and addon adapter packages.references/testing.mdfor fake providers, type checks, contract tests, and live-provider boundaries.references/package-surface.mdfor exports, package boundaries, source files, public docs, and known source-vs-doc checks.
Mirror Maintenance
This directory is the canonical source for the AI Harness agent skill. Sync a runtime mirror explicitly after changing it, then verify byte-for-byte:
npm run skills:sync -- /path/to/installed/ai-harness
npm run skills:sync -- --check /path/to/installed/ai-harness