prompt-compression
Sometimes you genuinely need a big external blob (a 5k-line log, a long transcript, a spec dump) in context. Pasting it whole is expensive and dilutes attention. Compress it to the parts that carry signal first.
Why this exists (evidence)
- Prompt compression (LLMLingua, Jiang et al., arXiv:2310.05736 and follow-ups) shows large prompts can be compressed multiple-x with little task-performance loss, cutting cost and latency, because most tokens in a big blob are low-information.
- It also fights context rot: fewer irrelevant tokens means the model attends to what matters.
When to use
- A SINGLE large input you cannot avoid including: long logs, stack traces, transcripts, large docs, data dumps, search results.
- Before pasting that blob into context or a sub-agent prompt.
- NOT a substitute for context-warden: that governs the whole SESSION (what to keep/drop over time); this compresses ONE oversized input on the way in.
The method
- Identify the signal the task needs from the blob: the error + its frames, the relevant section, the rows that matter, the decisions, the numbers.
- Extract, don't summarize loosely: keep exact identifiers (names, signatures, error strings, IDs, line refs), drop boilerplate, repetition, timestamps, banners, passing/no-op lines.
- Structure the residue: a short ordered extract or a small table, with a pointer back to the source (file:line / log range) so detail is recoverable on demand.
- State the compression: note what was dropped (e.g. "kept the 3 ERROR frames, dropped 4.8k INFO lines") so nothing looks hidden.
How to run it
- Logs: grep the error/levels you need, keep those frames + surrounding context, drop the rest.
- Transcripts/docs: extract the decisions/claims/sections relevant to the task; reference the rest by location.
- Heavy/automated: if an LLMLingua-style compressor or a summarizer tool is available, run it on the blob, then verify identifiers survived.
Composes with
context-warden: warden keeps the SESSION lean; prompt-compression shrinks a specific INPUT before it enters. Use together.retrieval-router: prefer retrieving only the needed slice over compressing the whole; compress when you truly must include a lot.run-cost: big blobs are top cost drivers; compress before paying for them repeatedly.
Honest limits
- Lossy by nature: aggressive compression can drop a detail that mattered. Keep exact identifiers and a pointer to the source so you can re-expand.
- For precise work, retrieving the exact slice (retrieval-router) beats compressing the whole. Compression is for when you cannot avoid the bulk.
- Cited ratios are from the papers' setups; measure your own loss.