llmkit — LLM Extraction Framework
Shared package at ~/research/packages/llmkit/ (pip install -e). Provides deterministic, auditable LLM extraction with Pydantic validation and file-backed caching. Designed for academic research where referees need to understand exactly what happened.
Core API
from llmkit import LLMCache, ExtractionSchema, extract, audit_sample
from llmkit.cache import text_hash, content_hash
from llmkit.extract import ExtractionResult
LLMCache(directory: Path)
File-backed cache. Each entry is a JSON file with three sections: _cache_meta, input_text, extraction.
cache = LLMCache(Path("cache_dir"))
key = cache.key(doc_id, text_hash, model) # composite key (back-compat)
key = cache.key(doc_id, text_hash, model, schema_name="my_task") # schema-aware key (recommended for new tasks)
hit = cache.get(key) # by composite key
hit = cache.get_by_doc(doc_id) # legacy fallback (doc_id.json)
cache.put(key, extraction, doc_id=..., text_hash=..., ...)
entries = cache.iter_entries()
schema_name is opt-in so existing caches (procure, EJ) stay valid. When you pass it, schema_name is mixed into the key — preventing collisions when two extraction tasks run against the same (doc_id, text) on the same model. New tasks should pass it; established tasks can migrate by renaming cache files (the new key is computable from each file's _cache_meta).
ExtractionSchema(BaseModel)
Base class for Pydantic schemas. Subclass and set schema_name/schema_version as ClassVar[str]:
from typing import ClassVar
from llmkit import ExtractionSchema
class MySchema(ExtractionSchema):
schema_name: ClassVar[str] = "my_task"
schema_version: ClassVar[str] = "v1"
field_a: str = ""
extract(...) -> ExtractionResult
Main entry point. Checks cache → calls LLM → validates → caches.
result = extract(
doc_id="ABC", text="...",
system_prompt="...", user_prompt="...",
schema=MySchema, model="gpt-4o-mini",
prompt_file="my_prompt.txt",
cache=cache, client=openai_client,
reextract=False, # True to redo stale entries
use_structured_outputs=True, # recommended — server-side schema enforcement
schema_in_cache_key=True, # recommended for new tasks (see above)
)
result.valid # bool — did Pydantic validation pass
result.parsed # MySchema instance (or None if invalid)
result.raw # dict — raw LLM JSON
result.cached # bool — loaded from cache
result.stale # bool — cached but prompt changed since
result.usage # dict — token counts
audit_sample(results, n=50, ...)
Stratified sampling for human review. See llmkit/audit.py.
build_report(cache, *, prices=None) / render_markdown(report)
Rolls a cache directory up into the reporting items a journal data editor needs for a paper using LLM-generated variables, per the checklist in Coqueret, Llull, Oswald, Pérignon, Scheuch & Vilhuber (2026), Randomness in large language models. Groups entries by (model, schema, prompt_hash, source_commit) and emits, per configuration: a model card (name, served version, dates queried), the verbatim system prompt + one example user message, the generation parameters actually used, and a cost statement (total input/output tokens, estimated USD, elapsed wall-clock). The rendered Markdown is framed around the two reproducibility checks (their §5.2): free/exact reproduction from deposited outputs vs. paid regeneration from the model (a fresh draw, not a byte copy).
python -m llmkit.report <cache_dir> -o REPLICATION_LLM.md
llmkit-report <cache_dir> --price gpt-4o-mini=0.15/0.60 # override $/1M in/out
Token counts always come from the cache's per-call usage; the USD estimate
uses a small built-in OpenAI price table (DEFAULT_PRICES, clearly caveated —
verify before quoting) or --price overrides. Models with no price show n/a
and are excluded from the USD total but still contribute token counts. The
cache holds one draw per unique input, so the totals describe a
single-draw regeneration. Where temperature == 0, the report adds the
paper's caveat that this is not deterministic. Legacy entries (no
_cache_meta) are counted but carry no tokens. See llmkit/report.py.
Cache design
Cache key = hash(doc_id, text_hash, model[, schema_name])
- Prompt changes do NOT invalidate cache — intentional, to avoid re-extraction during iterative development
- Text changes DO invalidate — different text = different key = cache miss
- Model changes DO invalidate — different model = different key
- schema_name is opt-in — pass it for new tasks (prevents cross-task collisions on the same document); omit it to stay compatible with caches written before the schema-aware option existed
Cache metadata (audit trail, not part of key)
Each cached JSON file contains:
{
"_cache_meta": {
"doc_id": "2Q6JZRW4NZ243MR",
"text_hash": "f3a1b2c4d5e6f7a8",
"prompt_hash": "9791af5d5643f4dc",
"model": "gpt-4o-mini",
"model_version": "gpt-4o-mini-2024-07-18",
"temperature": 0,
"max_tokens": 4000,
"schema_name": "irregularity_extraction",
"schema_version": "v1",
"source_commit": "ccfb959",
"validation_status": "valid",
"finish_reason": "stop",
"timestamp": "2026-03-23T10:01:06+00:00",
"usage": {"prompt_tokens": 7838, "completion_tokens": 1737},
"api_params": {"response_format": "json_object", "top_p": 1}
},
"messages": [
{"role": "system", "content": "You are an expert..."},
{"role": "user", "content": "Extract all irregularities..."}
],
"extraction": { ... LLM output ... }
}
Three top-level sections: _cache_meta (audit trail), messages (exact API input), extraction (LLM output). No redundancy — document text lives inside the user message.
Staleness and --reextract
entry.is_stale(current_prompt_hash="...") # True if prompt changed
- Day-to-day: run normally, cache accumulates, prompt edits don't trigger re-extraction
- Before submission: pass
--reextractto bring all entries in line with final prompt/data - Legacy entries (no metadata) are always considered stale
Backward compatibility
Old-style caches ({doc_id}.json with raw extraction, no metadata wrapper) are readable via get_by_doc(). The _load() method detects both formats.
LLM call modes
extract() supports two response modes. Both write the same cache shape; the _cache_meta.api_params.response_format field records which was used.
use_structured_outputs=True (recommended)
Calls client.chat.completions.parse(response_format=schema). OpenAI enforces the Pydantic schema server-side via constrained decoding. The model literally cannot produce out-of-schema field names or types.
- No JSON-keyword requirement on the prompt.
- No need to spell out nested field names in the system prompt — the schema does it.
- Requires a Structured-Outputs-capable model (gpt-4o, gpt-4o-mini, gpt-4.1, o-series).
_cache_meta.api_params.response_format == "structured_outputs";schema_nameis recorded.
use_structured_outputs=False (legacy / default for back-compat)
Calls client.chat.completions.create(response_format={"type": "json_object"}).
- Guarantees the output parses as JSON, but does not enforce the schema — field names are at the model's discretion. Always spell the schema out in the system prompt under this mode.
- Requires the literal word "json" to appear somewhere in the messages.
- Default
Falseso existing caches/scripts don't change behavior; new tasks should setTrue.
__bool__ and the c = cache or DEFAULT trap
LLMCache is always truthy (__bool__ returns True). Before this was set explicitly, LLMCache.__len__ made an empty cache falsy, so the common idiom c = passed_cache or DEFAULT_CACHE silently substituted the default whenever a caller passed a brand-new empty cache. If you maintain wrapper code that predates this fix, prefer c = DEFAULT_CACHE if cache is None else cache regardless — defensive against any future override that might want to make an LLMCache falsy.
Per-project setup
Each project defines its own schemas and prompts; llmkit provides the machinery.
Directory structure
project/source/llm/
__init__.py
schemas.py # Pydantic models (inherit ExtractionSchema)
irregularity.py # Wrapper wiring llmkit to project config
prompts/
irregularity_system.txt
irregularity_user.txt
Prompt versioning
Prompts are plain text files tracked by git. No version suffixes in filenames — git history IS the version control. The prompt's content hash is stored in cache metadata, and source_commit records which repo state produced the extraction.
To reconstruct the exact prompt for a cached entry:
git show <source_commit>:source/llm/prompts/irregularity_system.txt
Project wrapper pattern
The project wrapper handles legacy cache fallback and project-specific config:
# source/llm/my_task.py
from llmkit import LLMCache, extract
from llmkit.cache import content_hash, text_hash
from source.llm.schemas import MySchema
CACHE = LLMCache(DATA_DIR / "my_cache")
def extract_my_task(*, doc_id, text, client, reextract=False):
# 1. Try new composite key
# 2. Try legacy get_by_doc() fallback
# 3. Call LLM via extract()
...
See procure/source/llm/irregularity.py for the full pattern including legacy cache handling.
Existing implementation: procure
The procure project has a working extraction pipeline:
- Schema:
source/llm/schemas.py→IrregularityExtraction(document-level + per-irregularity fields) - Prompts:
source/llm/prompts/irregularity_system.txt,irregularity_user.txt - Wrapper:
source/llm/irregularity.py→extract_irregularities() - Script:
source/analysis/llm_extract_irregularities.py - Cache:
DATA_DIR/tce_sp/consulta/llm_cache/(~6,141 legacy entries + new-format entries) - Downstream:
source/analysis/build_consulta_llm.pyreadsllm_extract_irregularities.json
CLI:
python3 -m source.analysis.llm_extract_irregularities # use cache
python3 -m source.analysis.llm_extract_irregularities --reextract # redo stale
python3 -m source.analysis.llm_extract_irregularities --force # redo all
python3 -m source.analysis.llm_extract_irregularities --dry-run # preview
When this skill triggers automatically
Consult this skill (even without explicit /llmkit invocation) whenever you're about to:
- Set up LLM-based structured extraction in any project
- Work with cached LLM outputs or cache metadata
- Write Pydantic schemas for LLM validation
- Design prompts for extraction tasks
- Implement audit/review workflows for LLM outputs
- Debug cache staleness or re-extraction issues
Gotchas
- Cache key excludes prompt — by design. Don't add prompt to the key.
modelin key vsmodel_versionin metadata —modelis what you requested (e.g. "gpt-4o-mini"),model_versionis what the API actually used (e.g. "gpt-4o-mini-2024-07-18"). Both are saved.- Legacy cache
_usage— old entries store_usageinside the extraction dict. The wrapper strips it out before validation. Don't rely on_usagebeing inextractionfor new entries. ExtractionSchemausesClassVar— not underscore-prefixed attrs (Pydantic v2 treats_-prefixed as private attrs, which aren't JSON-serializable).messagesstores the full API input — each file is a self-contained audit record. Expect cache files to be large for long documents. Access document text viaentry.messages[1]["content"].diariosmoved — the package is now at~/research/packages/diarios/(was~/research/diarios/). Imports unchanged.- Cache-key changes need a rename, not a re-extract — switching a task to
schema_in_cache_key=Trueinvalidates existing cache lookups. Both the old and new key are deterministic functions of the file's_cache_meta(doc_id,text_hash,model, optionalschema_name), so a one-shot script can rename every cache file from the old key to the new one without re-calling the LLM. Do that, not--reextract, when migrating an established task.