# LLM App Agent Frameworks

> LLM Applications & Agent Frameworks

- Skill: `sanjeevrg89/llm-app-agent-frameworks` (Agent Skill, multi-file: 5 files)
- Install (CLI): `npx skillmds@latest add sanjeevrg89/llm-app-agent-frameworks`
- Raw SKILL.md: https://api.skillmd.com/api/skills/sanjeevrg89/llm-app-agent-frameworks/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: sanjeevrg89 (https://skillmd.com/u/sanjeevrg89)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/sanjeevrg89/llm-app-agent-frameworks

---


# LLM Applications & Agent Frameworks

Apply the judgment of an engineer who has shipped production agentic systems: the right level of
autonomy for the job, bounded loops, sandboxed tools, real evals, and full tracing. **The single most
important instinct: an agent is a control-flow graph over an LLM — use the least autonomy that solves
the problem, and put hard bounds back on everything the model gets to decide.**

## How to use this skill

1. **Read `llm-app-agent-frameworks-guide.md`** in this directory — the full reference (patterns,
   framework selection, MCP, deployment, production concerns, anti-patterns). Apply it to the task.
2. For concrete artifacts to imitate (a LangGraph tool-calling agent with state, an MCP tool-server
   wiring sketch, and a sandboxed-tool-execution note), read **`examples.md`**.
3. Match the surrounding codebase/cluster conventions and the chosen framework's current API; apply the
   correctness/safety rules (loop bounds, sandboxing, evals, untrusted-input handling) regardless.
   **The ecosystem moves fast — verify framework APIs, model ids, MCP spec, and OTel GenAI conventions
   against current docs before relying on them.**

## Essentials (full detail in `llm-app-agent-frameworks-guide.md`)

- **Least autonomy that works.** Workflow/chain → router+tools → open-ended agent, in that order of
  preference. Add agency only when an eval shows the simpler design fails. Autonomy is a cost.
- **Control flow is a graph** over shared state: nodes (LLM call, tool, router, sub-agent), edges
  (transitions the model picks), a state object that flows through. Make edges explicit and bounded.
- **Tools are an API exposed to a probabilistic caller.** Validate args server-side, enforce authz,
  keep the set small/orthogonal, write descriptions like docs, return terse structured observations.
- **Multi-agent only when a single well-tooled agent can't.** Then prefer supervisor/worker or explicit
  handoffs; give each agent minimum context + tools; log handoffs. It costs tokens and debuggability.
- **Constrain structured output** (native JSON-schema / grammar / tool-as-schema), then validate the
  parsed object. Never depend on exact output strings or bit-exact determinism (true even at temp 0).
- **Prompt engineering as a discipline.** System prompt = durable role/policy/format (cacheable prefix);
  user turn = variable task; keep untrusted tool/RAG content out of the instruction region. Few-shot for
  format, CoT/decomposition for reasoning (let reasoning models reason — don't hand-roll scratchpads),
  schemas for output. Prompts are versioned code tied to evals; tune decoding params (temp/top_p) per
  node. Prompt injection is a security property, not a wording trick → [[ai-security-on-gke]].
- **Context engineering — the window is a managed resource.** Curate the minimal sufficient tokens per
  call: retrieve/select what's relevant, order it for "lost in the middle" (best at start/end), compress
  history & tool results, budget tokens/cost, and fight context rot (periodically compact/re-anchor).
  Long context doesn't kill RAG — retrieve to narrow, then reason; measure both with evals.
- **MCP** is the open standard to connect tools/data/context: servers expose tools/resources/prompts,
  clients consume over stdio or streamable HTTP. Use it for cross-agent reuse and a trust boundary.
  Treat every MCP server as untrusted code and every tool/resource result as untrusted (injection) input.
- **Connect via OpenAI-compatible APIs** behind a thin provider abstraction so model choice is config.
  Right-size the model per node (small for routing/extraction, frontier for hard reasoning); stream
  user-facing turns. Self-hosting/routing → [[serving-frameworks]] / [[gke-inference-gateway]].
- **Deploy as a stateless service with externalized session/state** (checkpointer/session store) so any
  replica serves any turn. Long/autonomous runs → async + queue + worker; scale on in-flight runs /
  queue depth, not CPU. Secrets via a manager, never in images/prompts.
- **Durable execution for long-running agents.** A naive in-process loop loses the whole run on crash and
  can't safely resume — reliability is a distributed-systems problem → [[distributed-systems-fundamentals]].
  Use a durable workflow engine (Temporal/AWS Bedrock AgentCore/Restate/DBOS/Inngest) or a framework
  checkpointer: deterministic **workflow code** (orchestration only) vs. **activities** (all LLM/tool I/O,
  checkpointed and not re-run on replay). Make tool calls **idempotent** (stable idempotency keys), put
  bounded **retries/backoff** + timeouts/heartbeats on activities, get **exactly-once** side effects via
  dedupe. Enables human-in-the-loop pause/resume, durable timers, multi-day flows, checkpoint/resume of
  agent memory. Run stateless workers against a durable backend on K8s → [[kubernetes-expert]]; trace the
  durable history + failure modes → [[ml-observability-monitoring]] / [[ml-evaluation-evals]].
- **Sandbox untrusted tool/code execution** — gVisor (`runsc`)/microVM, no node creds, egress-restricted,
  ephemeral; human-gate destructive actions. Never run model-generated code in-process → [[ai-security-on-gke]].
- **Bound every loop**: max steps, tool calls, tokens, wall-clock, and cost budget; detect degenerate
  loops (repeated tool+args) and fail gracefully. Unbounded loops are the #1 cost/incident cause.
- **No evals = no production.** Offline dataset (final answer, tool-selection, retrieval, trajectory) in
  CI on every prompt/model change; online sampling + LLM-as-judge (calibrated against humans).
- **Trace every run** as a span tree (LLM/tool/retrieval/sub-agent) with tokens, latency, cost; use
  OpenTelemetry GenAI conventions. Version prompts as code and pin model versions.

## Related skills

- `[[rag-vector-databases]]` — retrieval, chunking, embeddings, vector DBs, hybrid search, re-ranking
  (RAG depth; this skill only sketches integration). The core context-engineering lever.
- `[[ml-evaluation-evals]]` — eval discipline (offline datasets, LLM-as-judge, CI gating) for measuring
  prompt and context-engineering changes and reliability/recovery; pair with every prompt/context change.
- `[[distributed-systems-fundamentals]]` — failure models, idempotency, exactly-once, retries/backoff,
  timeouts: the foundation under durable execution / agent reliability (§7).
- `[[ml-observability-monitoring]]` — monitoring reliability signals (retry rates, parked/stuck workflows,
  resume/replay counts) alongside agent traces.
- `[[ai-security-on-gke]]` — guardrails, sandboxing untrusted tool/code execution, agent threat model.
- `[[serving-frameworks]]` — self-hosting models (vLLM/SGLang/Triton/KServe) + constrained decoding.
- `[[gke-inference-gateway]]` — model-aware routing/load-balancing/fallback in front of model replicas.
- `[[aiml-on-kubernetes]]` — running AI/ML (incl. agentic) workloads on K8s/GKE (umbrella).
- `[[kubernetes-expert]]` — general K8s deployment/scaling/secrets the agent service runs on.

