Agent Scorecard Output Quality Framework
Configurable quality evaluation for AI agent outputs. Define criteria, run evaluations, track quality over time. No LLM-as-judge, no API calls, pattern-based automated checks.
Configurable quality evaluation for AI agent outputs. Define criteria, run evaluations, track quality over time.
Agent Scorecard gives you a structured, repeatable way to measure whether your AI agent is producing good output — and whether it's getting better or worse over time. No LLM-as-judge, no API calls, no external dependencies. Everything runs locally with pattern-based automated checks and optional human scoring.
The Problem
You changed your agent's system prompt. Is the output better now? You don't know. You added a new tool. Did response quality degrade? You have a feeling, but no data. Quality management for AI agents is mostly vibes.
Agent Scorecard replaces vibes with numbers.
What It Does
1. Define Quality Dimensions (config_example.json)
- Configure what "quality" means for your use case
- Set dimensions: accuracy, completeness, tone, format compliance, consistency — or your own
- Define rubrics (what does a 1 vs a 5 look like for each dimension?)
- Set weights (accuracy matters more than tone? Give it 2× weight)
- Set pass/fail thresholds per dimension
2. Evaluate (scorecard.py)
- Automated mode: Pattern-based checks run instantly with zero API calls
- Response length analysis (too short? too long?)
- Format compliance (expected headers, lists, code blocks present?)
- Sycophancy detection ("Great question!" markers)
- Filler/hedge word density ("basically", "perhaps", "I think")
- Required section verification
- Style consistency (sentence length variation)
- Manual mode: Interactive rubric-guided human scoring
- Blended mode: Combine auto scores with human judgment (averaged)
- Aggregate scoring with configurable method (weighted average, minimum, geometric mean)
3. Track (scorecard_track.py)
- Append every evaluation to a JSONL history file
- Filter by agent, task type, time period
- Compute trends per dimension (improving, degrading, stable)
- Linear regression slope for quantified direction
- Sparkline visualisations in terminal
4. Compare (scorecard_track.py)
- Before/after comparison (last N evals vs previous N)
- Per-dimension delta with direction indicators
- Perfect for measuring the impact of config changes
5. Report (scorecard_report.py)
- Single evaluation reports (markdown or JSON)
- History summary reports with tables and sparklines
- Per-dimension breakdowns with rubric reference
- Export to files or stdout
Quick Start
# 1. Configure
cp config_example.json scorecard_config.json
# Edit dimensions, thresholds, and weights for your use case
# 2. Evaluate a response
python3 scorecard.py --config scorecard_config.json --input response.txt
# 3. Evaluate and save to history
python3 scorecard.py --config scorecard_config.json --input response.txt --save history.jsonl
# 4. Manual scoring mode
python3 scorecard.py --config scorecard_config.json --input response.txt --manual --save history.jsonl
# 5. View trends
python3 scorecard_track.py --history history.jsonl --summary
# 6. Compare before/after (last 10 vs previous 10)
python3 scorecard_track.py --history history.jsonl --compare 10
# 7. Generate a report
python3 scorecard_report.py --config scorecard_config.json --history history.jsonl
Programmatic Usage
from scorecard import Scorecard, _load_config
cfg = _load_config("scorecard_config.json")
sc = Scorecard(cfg)
text = open("agent_response.txt").read()
result = sc.evaluate(text, agent="my-agent", task_type="code-review")
print(result.summary())
# Overall: 3.85/5 (PASS)
# ✓ Accuracy: 4.0/5 (threshold 3, weight 2.0) [auto]
# ✓ Completeness: 3.5/5 (threshold 3, weight 1.5) [auto]
# ...
# Save for tracking
import json
with open("history.jsonl", "a") as f:
f.write(json.dumps(result.to_dict()) + "\n")
Use Cases
- Prompt engineering: Measure whether prompt changes improve output quality
- Model comparison: Same task, different models — which scores higher?
- Agent regression testing: Catch quality degradation before it ships
- Team quality standards: Define shared rubrics for consistent evaluation
- Continuous monitoring: Track quality trends over days/weeks/months
- A/B testing: Quantified before/after comparisons
What's Included
| File |
Purpose |
scorecard.py |
Main evaluation engine — define, evaluate, score |
scorecard_track.py |
Historical tracking and trend analysis |
scorecard_report.py |
Report generation (markdown, JSON) |
config_example.json |
Full configuration template with all tunables |
LIMITATIONS.md |
What this tool doesn't do |
LICENSE |
MIT License |
Requirements
- Python 3.8+
- No external dependencies (stdlib only)
- Works on any OS
- Platform-agnostic (works with any AI agent framework)
Configuration
See config_example.json for the complete reference. Key areas:
DIMENSIONS — Quality dimensions with rubrics, weights, thresholds, and auto-checks
AUTO_CHECKS — Tuning for each pattern-based check (markers, thresholds, penalties)
AGGREGATE_METHOD — How to combine dimension scores ("weighted_average", "minimum", "geometric_mean")
HISTORY_FILE — Where to store evaluation history
REPORT_OUTPUT_DIR — Where reports are saved
quality-verified
License
MIT — See LICENSE file.
⚠️ Security Note — Config File
Configuration is loaded from a JSON file. This is safe to share — no code execution.
- Config path is validated for existence and size (1MB cap) before loading
- Must be a
.json file — raises ValueError if given a non-JSON path
- Keep your config under version control; it defines your quality rubrics and scoring weights
⚠️ Disclaimer
This software is provided "AS IS", without warranty of any kind, express or implied.
USE AT YOUR OWN RISK.
- The author(s) are NOT liable for any damages, losses, or consequences arising from
the use or misuse of this software — including but not limited to financial loss,
data loss, security breaches, business interruption, or any indirect/consequential damages.
- This software does NOT constitute financial, legal, trading, or professional advice.
- Users are solely responsible for evaluating whether this software is suitable for
their use case, environment, and risk tolerance.
- No guarantee is made regarding accuracy, reliability, completeness, or fitness
for any particular purpose.
- The author(s) are not responsible for how third parties use, modify, or distribute
this software after purchase.
By downloading, installing, or using this software, you acknowledge that you have read
this disclaimer and agree to use the software entirely at your own risk.
DATA DISCLAIMER: This software processes and stores data locally on your system.
The author(s) are not responsible for data loss, corruption, or unauthorized access
resulting from software bugs, system failures, or user error. Always maintain
independent backups of important data. This software does not transmit data externally
unless explicitly configured by the user.
Support & Links
Built with OpenClaw — thank you for making this possible.
🛠️ Need something custom? Custom OpenClaw agents & skills starting at $500. If you can describe it, I can build it. → Hire me on Fiverr
1---2name: agent-scorecard-output-quality-framework3description: Configurable quality evaluation for AI agent outputs. Define criteria, run evaluations, track quality over time. No LLM-as-judge, no API calls, pattern-based automated checks.4license: MIT5---67# Agent Scorecard Output Quality Framework89Configurable quality evaluation for AI agent outputs. Define criteria, run evaluations, track quality over time. No LLM-as-judge, no API calls, pattern-based automated checks.1011---1213**Configurable quality evaluation for AI agent outputs. Define criteria, run evaluations, track quality over time.**1415Agent Scorecard gives you a structured, repeatable way to measure whether your AI agent is producing good output — and whether it's getting better or worse over time. No LLM-as-judge, no API calls, no external dependencies. Everything runs locally with pattern-based automated checks and optional human scoring.1617---1819## The Problem2021You changed your agent's system prompt. Is the output better now? You don't know. You added a new tool. Did response quality degrade? You have a feeling, but no data. Quality management for AI agents is mostly vibes.2223Agent Scorecard replaces vibes with numbers.2425## What It Does2627### 1. **Define Quality Dimensions** (`config_example.json`)28- Configure what "quality" means for your use case29- Set dimensions: accuracy, completeness, tone, format compliance, consistency — or your own30- Define rubrics (what does a 1 vs a 5 look like for each dimension?)31- Set weights (accuracy matters more than tone? Give it 2× weight)32- Set pass/fail thresholds per dimension3334### 2. **Evaluate** (`scorecard.py`)35- **Automated mode:** Pattern-based checks run instantly with zero API calls36 - Response length analysis (too short? too long?)37 - Format compliance (expected headers, lists, code blocks present?)38 - Sycophancy detection ("Great question!" markers)39 - Filler/hedge word density ("basically", "perhaps", "I think")40 - Required section verification41 - Style consistency (sentence length variation)42- **Manual mode:** Interactive rubric-guided human scoring43- **Blended mode:** Combine auto scores with human judgment (averaged)44- Aggregate scoring with configurable method (weighted average, minimum, geometric mean)4546### 3. **Track** (`scorecard_track.py`)47- Append every evaluation to a JSONL history file48- Filter by agent, task type, time period49- Compute trends per dimension (improving, degrading, stable)50- Linear regression slope for quantified direction51- Sparkline visualisations in terminal5253### 4. **Compare** (`scorecard_track.py`)54- Before/after comparison (last N evals vs previous N)55- Per-dimension delta with direction indicators56- Perfect for measuring the impact of config changes5758### 5. **Report** (`scorecard_report.py`)59- Single evaluation reports (markdown or JSON)60- History summary reports with tables and sparklines61- Per-dimension breakdowns with rubric reference62- Export to files or stdout6364---6566## Quick Start6768```bash69# 1. Configure70cp config_example.json scorecard_config.json71# Edit dimensions, thresholds, and weights for your use case7273# 2. Evaluate a response74python3 scorecard.py --config scorecard_config.json --input response.txt7576# 3. Evaluate and save to history77python3 scorecard.py --config scorecard_config.json --input response.txt --save history.jsonl7879# 4. Manual scoring mode80python3 scorecard.py --config scorecard_config.json --input response.txt --manual --save history.jsonl8182# 5. View trends83python3 scorecard_track.py --history history.jsonl --summary8485# 6. Compare before/after (last 10 vs previous 10)86python3 scorecard_track.py --history history.jsonl --compare 108788# 7. Generate a report89python3 scorecard_report.py --config scorecard_config.json --history history.jsonl90```9192## Programmatic Usage9394```python95from scorecard import Scorecard, _load_config9697cfg = _load_config("scorecard_config.json")98sc = Scorecard(cfg)99100text = open("agent_response.txt").read()101result = sc.evaluate(text, agent="my-agent", task_type="code-review")102103print(result.summary())104# Overall: 3.85/5 (PASS)105# ✓ Accuracy: 4.0/5 (threshold 3, weight 2.0) [auto]106# ✓ Completeness: 3.5/5 (threshold 3, weight 1.5) [auto]107# ...108109# Save for tracking110import json111with open("history.jsonl", "a") as f:112 f.write(json.dumps(result.to_dict()) + "\n")113```114115---116117## Use Cases118119- **Prompt engineering:** Measure whether prompt changes improve output quality120- **Model comparison:** Same task, different models — which scores higher?121- **Agent regression testing:** Catch quality degradation before it ships122- **Team quality standards:** Define shared rubrics for consistent evaluation123- **Continuous monitoring:** Track quality trends over days/weeks/months124- **A/B testing:** Quantified before/after comparisons125126## What's Included127128| File | Purpose |129|------|---------|130| `scorecard.py` | Main evaluation engine — define, evaluate, score |131| `scorecard_track.py` | Historical tracking and trend analysis |132| `scorecard_report.py` | Report generation (markdown, JSON) |133| `config_example.json` | Full configuration template with all tunables |134| `LIMITATIONS.md` | What this tool doesn't do |135| `LICENSE` | MIT License |136137## Requirements138139- Python 3.8+140- No external dependencies (stdlib only)141- Works on any OS142- Platform-agnostic (works with any AI agent framework)143144## Configuration145146See `config_example.json` for the complete reference. Key areas:147148- **`DIMENSIONS`** — Quality dimensions with rubrics, weights, thresholds, and auto-checks149- **`AUTO_CHECKS`** — Tuning for each pattern-based check (markers, thresholds, penalties)150- **`AGGREGATE_METHOD`** — How to combine dimension scores ("weighted_average", "minimum", "geometric_mean")151- **`HISTORY_FILE`** — Where to store evaluation history152- **`REPORT_OUTPUT_DIR`** — Where reports are saved153154---155156## quality-verified157158159## License160161MIT — See `LICENSE` file.162163164---165166167## ⚠️ Security Note — Config File168169Configuration is loaded from a JSON file. This is safe to share — no code execution.170171- Config path is validated for existence and size (1MB cap) before loading172- Must be a `.json` file — raises `ValueError` if given a non-JSON path173- Keep your config under version control; it defines your quality rubrics and scoring weights174175## ⚠️ Disclaimer176177This software is provided "AS IS", without warranty of any kind, express or implied.178179**USE AT YOUR OWN RISK.**180181- The author(s) are NOT liable for any damages, losses, or consequences arising from 182 the use or misuse of this software — including but not limited to financial loss, 183 data loss, security breaches, business interruption, or any indirect/consequential damages.184- This software does NOT constitute financial, legal, trading, or professional advice.185- Users are solely responsible for evaluating whether this software is suitable for 186 their use case, environment, and risk tolerance.187- No guarantee is made regarding accuracy, reliability, completeness, or fitness 188 for any particular purpose.189- The author(s) are not responsible for how third parties use, modify, or distribute 190 this software after purchase.191192By downloading, installing, or using this software, you acknowledge that you have read 193this disclaimer and agree to use the software entirely at your own risk.194195196**DATA DISCLAIMER:** This software processes and stores data locally on your system. 197The author(s) are not responsible for data loss, corruption, or unauthorized access 198resulting from software bugs, system failures, or user error. Always maintain 199independent backups of important data. This software does not transmit data externally 200unless explicitly configured by the user.201202---203204## Support & Links205206| | |207|---|---|208| 🐛 **Bug Reports** | TheShadowyRose@proton.me |209| ☕ **Ko-fi** | [ko-fi.com/theshadowrose](https://ko-fi.com/theshadowrose) |210| 🛒 **Gumroad** | [shadowyrose.gumroad.com](https://shadowyrose.gumroad.com) |211| 🐦 **Twitter** | [@TheShadowyRose](https://twitter.com/TheShadowyRose) |212| 🐙 **GitHub** | [github.com/TheShadowRose](https://github.com/TheShadowRose) |213| 🧠 **PromptBase** | [promptbase.com/profile/shadowrose](https://promptbase.com/profile/shadowrose) |214215*Built with [OpenClaw](https://github.com/openclaw/openclaw) — thank you for making this possible.*216217---218219🛠️ **Need something custom?** Custom OpenClaw agents & skills starting at $500. If you can describe it, I can build it. → [Hire me on Fiverr](https://www.fiverr.com/s/jjmlZ0v)