# Agentforce Cost Optimization

> Use when Agentforce run costs are climbing, you need to forecast scale, or you want to reduce tokens per conversation without hurting quality. Covers topic (now subagent) design impact on cost, prompt/template reuse, grounding size discipline, caching, and model-tier selection. Triggers: 'agentforce cost', 'tokens per conversation too high', 'reduce agentforce runs spend', 'forecast agentforce scale cost', 'einstein trust layer tokens'. NOT for capping spend per user with a budget gate and fallback — use agentforce/agent-rate-limit-strategy. NOT for org-wide model-tier and BYOLLM platform strategy — use architect/ai-platform-architecture.

- Skill: `pranavnagrecha/agentforce-cost-optimization` (Agent Skill, multi-file: 7 files)
- Install (CLI): `npx skillmds add pranavnagrecha/agentforce-cost-optimization`
- Raw SKILL.md: https://api.skillmd.com/api/skills/pranavnagrecha/agentforce-cost-optimization/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: PranavNagrecha (https://skillmd.com/u/pranavnagrecha)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/pranavnagrecha/agentforce-cost-optimization

---


# Agentforce Cost Optimization

Agentforce cost looks like "we'll just pay per run" right up until volume meets reality. A customer-service agent handling 200,000 conversations/month can consume 10× the tokens of a well-tuned version of the same agent — same quality, same subagents, different token discipline. The cost drivers are predictable: subagent instruction length, prompt template verbosity, grounding payload size, tool-call round-trips, and model tier. None of these are free to change, but they all respond to focused work.

The job is to measure first, then optimize the top three contributors. Most orgs find that subagent instructions and grounding dominate — often 60-80% of tokens per conversation. Once those are disciplined, the remaining optimizations (template reuse, model-tier selection) become viable.

> **Terminology.** This skill leads with *subagent* because that is the current
> product term — Salesforce renamed agent *topics* to *subagents* in April 2026,
> with no change to functionality. It deliberately keeps *topic* in metadata and
> API names and in search keywords, because those did not change, and readers
> arriving with the older vocabulary still need to find this skill.

---

## Before Starting

- Pull 7 days of Agentforce runs; compute average and p95 token counts per conversation.
- Inventory subagents, prompt templates, and grounding sources.
- Confirm model tier currently in use and any rate-limit headroom.
- Confirm business tolerance for quality-vs-cost tradeoffs.

## Core Concepts

### What Tokens Are You Paying For?

Every conversation pays for:

1. **System prompt** — the framework-level Agentforce prompt.
2. **Subagent instructions** — active subagent's instructions injected verbatim.
3. **Prompt template** — any custom template rendered per turn.
4. **Grounding** — retrieved content from Data Cloud, Knowledge, or explicit variables.
5. **Conversation history** — full turn history on each call.
6. **Tool output** — action results returned into context.

### The 80/20 Rule

For most agents, subagent instructions + grounding = 60-80% of token spend. Conversation history grows linearly in long sessions. Tool output is lumpy but occasionally large (SOQL result sets dumped raw into context).

### Reducing Subagent Instruction Tokens

- Delete department-name preamble ("As a customer service agent working for Acme Insurance...").
- Collapse redundant examples; 2 good examples outperform 10 mediocre ones.
- Externalize static policy ("always use formal English") into the system prompt instead of per-subagent.

### Reducing Grounding Tokens

- Retrieve k=3, not k=10, unless evaluation shows quality improves.
- Chunk sizes: 300-500 tokens usually beats 1000-2000.
- Reranker before final injection when using Data Cloud retrievers.
- Strip boilerplate (legal footers, headers) from Knowledge articles before indexing.

### Conversation History Discipline

Long sessions inflate every turn's token count. Patterns:
- Summarize older turns ("Summary of first 5 turns: …") rather than sending verbatim.
- Archive turns beyond a threshold; keep only the last N in active context.

### Model Tier Selection

Not every action needs the most capable model. Use tiered routing:
- Classification / intent detection → smaller model.
- Reasoning / final response → larger model.
- Tool-calling / structured output → mid-tier is often enough.

### Caching Opportunities

- Subagent instructions are stable across conversations — framework should cache; you don't need to change anything unless your template is dynamic.
- Grounding retrieval can cache per query; watch freshness needs.

---

## Common Patterns

### Pattern 1: Subagent Instruction Audit And Trim

Per subagent, measure instruction token count. Target 150-300 tokens per subagent instruction. Trim anything above 500 without a compelling reason.

### Pattern 2: k-3 Retriever With Reranker

Retrieve 10 candidates; rerank; inject top 3. Cuts grounding tokens 70% vs retrieve-10-inject-10.

### Pattern 3: Conversation Summarization Trigger

After N turns or M tokens of history, replace older turns with a one-line summary.

### Pattern 4: Tiered Model Routing

Route classification / intent steps to a smaller model; reasoning/response to the capable model.

### Pattern 5: Tool Output Projection

When a tool returns a large payload (e.g. SOQL result), project the fields the agent actually needs instead of dumping the full response.

---

## Decision Guidance

| Situation | Recommended Approach | Reason |
|---|---|---|
| Token usage high, unknown contributor | Instrument and measure first | Avoid guessing |
| Subagent instructions > 500 tokens | Trim (Pattern 1) | Biggest win |
| Grounding k ≥ 5 without evaluation | Reduce k + rerank (Pattern 2) | Second biggest win |
| Long conversations | Summarize (Pattern 3) | Linear savings per turn |
| Classification step using largest model | Switch to smaller tier (Pattern 4) | Cheap wins |
| Tool returns wide records | Project fields (Pattern 5) | Eliminates silent waste |

## Review Checklist

- [ ] Per-conversation token metrics collected and dashboarded.
- [ ] Top 3 token contributors identified per agent.
- [ ] Subagent instruction length audited.
- [ ] Grounding k and chunk size justified.
- [ ] Long-conversation strategy exists.
- [ ] Model tier routing considered.
- [ ] Tool output projection in place.

## Recommended Workflow

1. Measure — 7 days of run data broken down by token source.
2. Identify top 3 contributors.
3. Optimize subagent instructions first.
4. Optimize grounding second.
5. Add conversation summarization if sessions are long.
6. Apply tier routing where quality allows.
7. Re-measure; document cost savings.

---

## Salesforce-Specific Gotchas

1. Trust Layer adds tokens — masking, citation, guardrails all add context weight.
2. Grounding sources can include large boilerplate (Knowledge article footers); index selectively.
3. Tool output is counted even if the agent ignores it.
4. Managed subagents may have opaque instruction length; audit via runtime logs.
5. Switching model tier changes quality — do not do this without A/B evaluation.

## Proactive Triggers

- Subagent instruction > 500 tokens → Flag High.
- Retriever k ≥ 10 without reranker → Flag High.
- Average conversation > 20 turns with no summarization → Flag Medium.
- Classification step on flagship model → Flag Medium.
- Token growth > 15%/month without volume growth → Flag High.

## Output Artifacts

| Artifact | Description |
|---|---|
| Cost model | Tokens per conversation by contributor |
| Optimization plan | Prioritized trim list with expected savings |
| Tier routing design | Step → model mapping |

## Related Skills

- `agentforce/agent-topic-design` — subagent structure quality.
- `agentforce/prompt-builder-templates` — prompt template hygiene.
- `agentforce/data-cloud-grounding-for-agentforce` — grounding retrieval.
- `agentforce/agentforce-observability` — measurement infrastructure.

