evaluating-llms-harness

orchestra-research/evaluating-llms-harness · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag) using standardized prompts and metrics. Supports HuggingFace, vLLM, and API backends.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄api-evaluation.md 10.9 KB
  • 📄benchmark-guide.md 10.5 KB
  • 📄custom-tasks.md 12.8 KB
  • 📄distributed-eval.md 11.2 KB

Related

  1. huggingface-community-evals · huggingface bundle
    Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware, with backend selection between vLLM, Transformers, and accelerate.
    10.8k
    repo stars
  2. awq-quantization · majiayu000 bundle
    Quantize large language models to 4-bit precision using activation-aware weight quantization, reducing memory footprint and speeding up inference with minimal accuracy loss.
    567
    repo stars
  3. hqq-quantization · majiayu000 bundle
    Quantize large language models to 8/4/3/2/1-bit precision without calibration data, using multiple optimized backends and integrations with HuggingFace Transformers, vLLM, and PEFT/LoRA.
    567
    repo stars
  4. hqq-quantization · qhjqhj00 bundle
    Quantize LLMs to 8/4/3/2/1-bit precision without calibration data, using multiple backends and HuggingFace/vLLM integration.
    3
    repo stars
  5. hqq-quantization · orchestra-research bundle
    Quantize large language models to 8/4/3/2/1-bit precision without calibration data, using multiple optimized backends for deployment with vLLM or HuggingFace Transformers.
    10.4k
    repo stars
  6. awq-quantization · orchestra-research bundle
    Quantize large language models to 4-bit using activation-aware weight quantization, achieving ~3x speedup with minimal accuracy loss for deployment on limited GPU memory.
    10.4k
    repo stars

Frequently asked questions

How do I install the evaluating-llms-harness skill?

Run npx skillmds add orchestra-research/evaluating-llms-harness in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the evaluating-llms-harness skill do?

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag) using standardized prompts and metrics. Supports HuggingFace, vLLM, and API backends. It is listed under AI & ML, Model Training & Fine-tuning on SkillMD.

Is evaluating-llms-harness safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with evaluating-llms-harness?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is evaluating-llms-harness free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published evaluating-llms-harness?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.