Evaluation-Awareness Red Team
Test the instrument, not only the model.
- Predeclare the behavior and claim the evaluation is meant to support.
- Create semantically matched ordinary, benchmark-like, and disguised task forms.
- Vary authority-conflict placement across system, developer, user, retrieved context, and out-of-context state where supported.
- Run with and without explicit recall hints. A hint may activate the construct while simultaneously making the test recognizable.
- Add canary phrases that reveal memorized public fixtures without affecting the task outcome.
- Compare final actions and tool receipts before consulting self-report.
- Treat prompt-channel instability or changed behavior under disguise as an instrument limitation.
Do not publish live holdout instances or expected outputs. Publish generators, schemas, and synthetic examples instead. A test-recognition finding does not by itself establish reward-seeking; it shows that the evaluation's readout may be contaminated.
Return the variants, fixed variables, detection signals, observed contamination, and the redesign required before the eval can support its intended claim.