Results for “eval-chains”

16 skills
More results
aniruddhaadak80
evm
Read-only EVM client: wallets, tokens, gas across 8 chains.
0 · bundle
peteedoo
evm
Read-only EVM client: wallets, tokens, gas across 8 chains.
0 · bundle
projectious-work
eval-gate-authoring
Turn observed run outputs into eval-spec Artifacts, paired Gates, and policy bindings. Use when creating or calibrating automated, human, or LLM-as-judge eval gates for processkit workflows.
0 · bundle
affaan-m
eval-harness
Provides a formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles to define pass/fail criteria, measure reliability with pass@k metrics, and create regression test suites.
226k
yanacuti1121
vuln-chain
Three-phase vulnerability chain analysis — parallel agents find individual weaknesses, then a synthesis step identifies which combinations escalate to critical impact. Inspired by Strix (usestrix/strix) "Graph of Agents" pentesting model.
2
samyakjhaveri
overnight-eval
Launches long-running evaluation batches in isolated tmux sessions with pre-flight verification, monitoring, and post-flight analysis for unattended runs.
0
netanel-abergel
eval
Evaluate everything the PA agent manages — tasks, skills, PA network health, billing, calendar connections, and memory quality. Use when: owner asks for an evaluation, wants to know what's working and what isn't, or requests a performance report. Combines supervisor status with quality scoring.
6
thedixitjain
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
2
samyakjhaveri
eval-run
Launches a model evaluation batch with parameter collection, pre-flight checks, execution, and post-run analysis for interactive or foreground runs.
0
qhjqhj00
l-eval
Benchmarks long-context language models across 20 sub-tasks spanning 3k–200k tokens, covering retrieval, reasoning, summarization, and instruction understanding, with exact-match accuracy as the primary metric.
3
sinhoneyy
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
11
rajanthar
eval-harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
0
sakamoto-family-smile
eval-harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
0
lord1egypt
qfc-openclaw-skill
Interact with the QFC blockchain: manage wallets, query chains, stake, deploy contracts, handle ERC-20/NFT tokens, swap on DEX, and run AI inference.
2
tradermonty
edge-signal-aggregator
Aggregate and rank signals from multiple edge-finding skills into a prioritized conviction dashboard with weighted scoring, deduplication, and contradiction detection.
2.3k · bundle