Rag Evaluator
AI-powered RAG (Retrieval-Augmented Generation) evaluation toolkit. Configure, benchmark, compare, and optimize your RAG pipelines from the command line. Track prompts, evaluations, fine-tuning experiments, costs, and usage — all with persistent local logging and full export capabilities.
Commands
Run rag-evaluator <command> [args] to use.
| Command |
Description |
configure |
Configure RAG evaluation settings and parameters |
benchmark |
Run benchmarks against your RAG pipeline |
compare |
Compare results across different RAG configurations |
prompt |
Log and manage prompt templates and variations |
evaluate |
Evaluate RAG output quality and relevance |
fine-tune |
Track fine-tuning experiments and parameters |
analyze |
Analyze evaluation results and identify patterns |
cost |
Track and log API/inference costs |
usage |
Monitor token usage and API call volumes |
optimize |
Log optimization strategies and results |
test |
Run test cases against RAG configurations |
report |
Generate evaluation reports |
stats |
Show summary statistics across all categories |
export <fmt> |
Export data in json, csv, or txt format |
search <term> |
Search across all logged entries |
recent |
Show recent activity from history log |
status |
Health check — version, data dir, disk usage |
help |
Show help and available commands |
version |
Show version (v2.0.0) |
Each domain command (configure, benchmark, compare, etc.) works in two modes:
- Without arguments: displays the most recent 20 entries from that category
- With arguments: logs the input with a timestamp and saves to the category log file
Data Storage
All data is stored locally in ~/.local/share/rag-evaluator/:
- Each command creates its own log file (e.g.,
configure.log, benchmark.log)
- A unified
history.log tracks all activity across commands
- Entries are stored in
timestamp|value pipe-delimited format
- Export supports JSON, CSV, and plain text formats
Requirements
- Bash 4+ with
set -euo pipefail strict mode
- Standard Unix utilities:
date, wc, du, tail, grep, sed, cat
- No external dependencies or API keys required
When to Use
- Evaluating RAG pipeline quality — log evaluation scores, compare retrieval strategies, and track improvements over time
- Benchmarking different configurations — run benchmarks across embedding models, chunk sizes, or retrieval methods and compare results side by side
- Tracking costs and usage — monitor API costs and token usage across experiments to stay within budget
- Managing prompt engineering — log prompt variations, test them against your pipeline, and analyze which templates perform best
- Generating reports for stakeholders — export evaluation data as JSON/CSV for dashboards, or generate text reports summarizing RAG performance
Examples
# Configure a new evaluation run
rag-evaluator configure "model=gpt-4 chunks=512 overlap=50 top_k=5"
# Run a benchmark and log results
rag-evaluator benchmark "latency=230ms recall@5=0.82 precision@5=0.71"
# Compare two retrieval strategies
rag-evaluator compare "bm25 vs dense: bm25 recall=0.78, dense recall=0.85"
# Track evaluation scores
rag-evaluator evaluate "faithfulness=0.91 relevance=0.87 coherence=0.93"
# Log API cost for a run
rag-evaluator cost "run-042: $0.23 (1.2k tokens input, 800 tokens output)"
# View summary statistics
rag-evaluator stats
# Export all data as CSV
rag-evaluator export csv
# Search for specific entries
rag-evaluator search "gpt-4"
# Check recent activity
rag-evaluator recent
# Health check
rag-evaluator status
Output
All commands output to stdout. Redirect to a file if needed:
rag-evaluator report "weekly summary" > report.txt
rag-evaluator export json # saves to ~/.local/share/rag-evaluator/export.json
Configuration
Set DATA_DIR by modifying the script, or use the default: ~/.local/share/rag-evaluator/
Powered by BytesAgain | bytesagain.com | hello@bytesagain.com
1---2name: ragaai-catalyst3description: Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like a ragaai catalyst, python, agentic-ai.4---5
6# Rag Evaluator
7
8AI-powered RAG (Retrieval-Augmented Generation) evaluation toolkit. Configure, benchmark, compare, and optimize your RAG pipelines from the command line. Track prompts, evaluations, fine-tuning experiments, costs, and usage — all with persistent local logging and full export capabilities.
9
10## Commands
11
12Run `rag-evaluator <command> [args]` to use.
13
14| Command | Description |
15|---------|-------------|
16| `configure` | Configure RAG evaluation settings and parameters |
17| `benchmark` | Run benchmarks against your RAG pipeline |
18| `compare` | Compare results across different RAG configurations |
19| `prompt` | Log and manage prompt templates and variations |
20| `evaluate` | Evaluate RAG output quality and relevance |
21| `fine-tune` | Track fine-tuning experiments and parameters |
22| `analyze` | Analyze evaluation results and identify patterns |
23| `cost` | Track and log API/inference costs |
24| `usage` | Monitor token usage and API call volumes |
25| `optimize` | Log optimization strategies and results |
26| `test` | Run test cases against RAG configurations |
27| `report` | Generate evaluation reports |
28| `stats` | Show summary statistics across all categories |
29| `export <fmt>` | Export data in json, csv, or txt format |
30| `search <term>` | Search across all logged entries |
31| `recent` | Show recent activity from history log |
32| `status` | Health check — version, data dir, disk usage |
33| `help` | Show help and available commands |
34| `version` | Show version (v2.0.0) |
35
36Each domain command (configure, benchmark, compare, etc.) works in two modes:
37- **Without arguments**: displays the most recent 20 entries from that category
38- **With arguments**: logs the input with a timestamp and saves to the category log file
39
40## Data Storage
41
42All data is stored locally in `~/.local/share/rag-evaluator/`:
43
44- Each command creates its own log file (e.g., `configure.log`, `benchmark.log`)
45- A unified `history.log` tracks all activity across commands
46- Entries are stored in `timestamp|value` pipe-delimited format
47- Export supports JSON, CSV, and plain text formats
48
49## Requirements
50
51- Bash 4+ with `set -euo pipefail` strict mode
52- Standard Unix utilities: `date`, `wc`, `du`, `tail`, `grep`, `sed`, `cat`
53- No external dependencies or API keys required
54
55## When to Use
56
571. **Evaluating RAG pipeline quality** — log evaluation scores, compare retrieval strategies, and track improvements over time
582. **Benchmarking different configurations** — run benchmarks across embedding models, chunk sizes, or retrieval methods and compare results side by side
593. **Tracking costs and usage** — monitor API costs and token usage across experiments to stay within budget
604. **Managing prompt engineering** — log prompt variations, test them against your pipeline, and analyze which templates perform best
615. **Generating reports for stakeholders** — export evaluation data as JSON/CSV for dashboards, or generate text reports summarizing RAG performance
62
63## Examples
64
65```bash
66# Configure a new evaluation run
67rag-evaluator configure "model=gpt-4 chunks=512 overlap=50 top_k=5"
68
69# Run a benchmark and log results
70rag-evaluator benchmark "latency=230ms recall@5=0.82 precision@5=0.71"
71
72# Compare two retrieval strategies
73rag-evaluator compare "bm25 vs dense: bm25 recall=0.78, dense recall=0.85"
74
75# Track evaluation scores
76rag-evaluator evaluate "faithfulness=0.91 relevance=0.87 coherence=0.93"
77
78# Log API cost for a run
79rag-evaluator cost "run-042: $0.23 (1.2k tokens input, 800 tokens output)"
80
81# View summary statistics
82rag-evaluator stats
83
84# Export all data as CSV
85rag-evaluator export csv
86
87# Search for specific entries
88rag-evaluator search "gpt-4"
89
90# Check recent activity
91rag-evaluator recent
92
93# Health check
94rag-evaluator status
95```
96
97## Output
98
99All commands output to stdout. Redirect to a file if needed:
100
101```bash
102rag-evaluator report "weekly summary" > report.txt
103rag-evaluator export json # saves to ~/.local/share/rag-evaluator/export.json
104```
105
106## Configuration
107
108Set `DATA_DIR` by modifying the script, or use the default: `~/.local/share/rag-evaluator/`
109
110---
111Powered by BytesAgain | bytesagain.com | hello@bytesagain.com