# Results

> Documents benchmark results for a multi-agent system and provides commands to run benchmark suites.

- Skill: `jorcan/results` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add jorcan/results`
- Raw SKILL.md: https://api.skillmd.com/api/skills/jorcan/results/raw
- Safety review: PASS (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools, Testing & QA
- Tags: Benchmark, Humaneval, Loki Mode, Multi Agent, Swebench
- Author: jorcan (https://skillmd.com/u/jorcan)
- Updated: 2026-08-22
- Page: https://skillmd.com/skills/jorcan/results

---

# Loki Mode Benchmark Results

**Generated:** 2026-01-05 09:31:14

## Overview

This directory contains benchmark results for Loki Mode multi-agent system.

## Methodology

Loki Mode uses its multi-agent architecture to solve each problem:
1. **Architect Agent** analyzes the problem
2. **Engineer Agent** implements the solution
3. **QA Agent** validates with test cases
4. **Review Agent** checks code quality

This mirrors real-world software development more accurately than single-agent approaches.

## Running Benchmarks

```bash
# Setup only (download datasets)
./benchmarks/run-benchmarks.sh all

# Execute with Claude
./benchmarks/run-benchmarks.sh humaneval --execute
./benchmarks/run-benchmarks.sh humaneval --execute --limit 10  # First 10 only
./benchmarks/run-benchmarks.sh swebench --execute --limit 5    # First 5 only

# Use different model
./benchmarks/run-benchmarks.sh humaneval --execute --model opus
```

