# Embrace 3k Eval

> embrace-3k-eval

- Skill: `qhjqhj00/embrace-3k-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/embrace-3k-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/embrace-3k-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/embrace-3k-eval

---


# embrace-3k-eval

> EmbRACE-3K: Embodied Reasoning and Action in Complex Environments — Lin et al. (2025) (arXiv:2507.10548, 2025)

## What this evaluates

Evaluates embodied vision-language models on step-wise, instruction-driven navigation and interaction in complex, partially observable environments. It probes three core capabilities: active exploration, dynamic spatial-semantic reasoning, and multi-stage goal execution under closed-loop perception-action constraints.

## Datasets

- **EmbRACE-3K** — total 3000; splits: test (-1)

## Metrics

- `success` **(primary)** — range: percent
  - Percentage of tasks where the agent successfully completes the instruction within the allowed step limit.

## Input / output format

**Input**: Egocentric RGB frames, 6-DoF agent pose, and natural language task instruction.

**Output**: Discrete action sequence and step-wise natural language reasoning/justification for each action.

## Scoring recipe

```python
def compute_success(predictions, gold):
    correct = 0
    for pred, task in zip(predictions, gold):
        if pred.completed and pred.steps <= 32:
            correct += 1
    return (correct / len(gold)) * 100
```

## Common pitfalls

- Trajectories are strictly filtered to a maximum of 32 steps; evaluating agents that exceed this limit will artificially penalize valid long-horizon strategies.
- Task types vary significantly in difficulty (Basic vs. Multi-stage/Interaction); reporting only aggregate success masks capability gaps in complex reasoning categories.
- Egocentric observations and 6-DoF poses must be temporally aligned; mismatched frame-pose pairs break the closed-loop perception-action evaluation.

## Evidence (verbatim from paper)

> The dataset enables evaluation of VLMs on three core embodied capabilities—Exploration, Dynamic Spatial-Semantic Reasoning, and Multi-stage Goal Execution—where state-of-the-art models achieve <20% success in zero-shot settings, revealing a critical gap in embodied reasoning.

## Citation

```bibtex
@misc{lin2025embrace3k,
  title={EmbRACE-3K: Embodied Reasoning and Action in Complex Environments},
  author={Lin et al. (2025)},
  year={2025},
  note={arXiv:2507.10548}
}
```

- arXiv: 2507.10548

