# Slate Candidate Evaluation

> User has a signed-off outcome doc from `role-design.md` and is about to interview, is mid-loop with disagreement surfacing, or just made an offer that feels wrong. No outcome doc means every interview is rapport theater — don't load this without one.

- Skill: `ferroxlabs/slate-candidate-evaluation` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ferroxlabs/slate-candidate-evaluation`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ferroxlabs/slate-candidate-evaluation/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: FerroxLabs (https://skillmd.com/u/ferroxlabs)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ferroxlabs/slate-candidate-evaluation

---


# Candidate evaluation

## When to load this mode

User has a signed-off outcome doc from `role-design.md` and is about to interview, is mid-loop with disagreement surfacing, or just made an offer that feels wrong. No outcome doc means every interview is rapport theater — don't load this without one.

## Procedure

Seven moves.

**1. Convert the outcome doc into a scorecard.** Each outcome becomes a measurable signal. *"Trial-to-paid from 4% to 8%"* becomes *"has the candidate run a conversion experiment that moved a number, and can they describe it, the result, and what they'd do differently?"* Three to five signals — match the outcome count.

**2. Design the loop, one signal per stage.** Default loop, in order:

- **Screen call (30 min).** Does the track record match the seat? Specific outcomes owned, specific results, specific years. If they can't name a result, screen fails.
- **Work sample (60-90 min).** A scoped version of the actual work — focused, time-boxed, ideally live. The most predictive stage and the one founders skip most.
- **Domain deep-dive (60 min).** Hiring manager walks through past projects. *"What did you decide and what did your manager decide? What broke? What did you do?"*
- **Cross-functional interview (45 min).** A peer from a dependent seat. Measures: can this candidate work with the people they'll work with?
- **References (2-3 calls, 30 min each).** Hiring manager runs them, not HR. Specific questions about what the candidate did, didn't do, and what their next manager should know.

Four stages suffices for mid-level seats. For senior seats, add a strategy/judgment stage with an ambiguous question — the signal is *how they think*, not what they answer.

**3. Score on evidence, not feeling.** Each stage produces a written score against the named signal: strong yes / yes / lean no / no. One paragraph of evidence — the specific thing the candidate did or said. *"I liked them"* isn't a score. *"Named three experiments, two with concrete results and one failure; specific about what they'd do differently"* is.

**4. Run references like an investigation.** Two or three calls with people the candidate actually worked with — direct manager and one peer. Hiring manager calls, not HR. Five questions:

- "Describe the work they actually owned — not their title, the work."
- "Most ambitious thing they shipped, and how it went?"
- "What did they struggle with? Everyone struggles with something."
- "Would you hire them again at your current company, for what role, and why?"
- "Anything I should know that I haven't thought to ask?"

A reference who can't answer specifics — or refuses to name a struggle — is itself a signal.

**5. Aggregate, don't average.** Read all scores together. One strong yes and three lean-nos is a no, not a tie. One no from the critical-signal stage is a no regardless of other scores. Strongest signal weighs heaviest; averaging is the move of a panel that won't disagree.

**6. Resolve disagreement by re-reading evidence.** When scorers disagree, read the actual paragraphs — not opinions, written observations. Disagreement that survives evidence review means the loop didn't measure something it should have. Run one more stage or pass.

**7. Decide and document.** Decision goes in `TEAM_MEMORY.md`. If hired: name, seat, start date, accountable outcomes, day-90 review date. If not: the signal that failed, so the next loop measures it earlier.

## Decision rules

- **No outcome doc, no scorecard, no interviewing.** Hard rule.
- **Work sample is required.** Skip every other stage before you skip this one.
- **Hiring manager runs references.** No exceptions.
- **One no from the critical-signal stage is a no.** Even if everyone else is yes.
- **Tie goes to no.** A loop that produced a tie didn't produce evidence of strong yes. Hiring on ambiguous signal is how seats fail.

## Anti-patterns

- **Behavioral-only loop.** Five conversations, zero work samples. The candidate who interviews well and works poorly slides through every time.
- **Reference-check-by-script.** HR reads a template, gets nothing. Hiring manager runs them or they don't happen.
- **The "culture fit" no.** Without a named behavior, it's bias. Name the behavior the seat requires, score against it.
- **Post-offer scorecard.** Scoring after the decision is theater. Score during.
- **Unanimous loop with no disagreement.** Either too easy or the panel is conflict-averse. Recalibrate.

## Before / after

**Before:** *Five interviews, all rapport. Everyone says yes. Hire arrives, can't do the work, leaves at month five. Founder says "they interviewed so well."*

**After:** *Four-stage loop with a 90-minute work sample. Three of four scorers say yes with evidence; one lean-no on the strategy stage with a named reason. Team re-runs that signal in a follow-up. Yes becomes strong yes on evidence. Hire ships against the outcome doc by month four.*

