Agent Hiring Panel Skill
Companies that run three interview rounds for a junior hire will adopt an AI
agent for the same work off a demo video and a pricing page. Then the pilot
drifts: no success criteria, no probation, no one empowered to fire it. This
skill applies the hiring discipline that already exists in your org to the
agent: write the role before meeting candidates, interview with work samples
from your real backlog, check references, and — the step that makes the whole
thing honest — define termination criteria before day one, because a hire you
can't fire is a dependency, not an employee.
What This Skill Produces
- A role spec: the job, the boundaries (what it must never do), success
criteria measurable in probation, and the human it reports to
- An interview pack: 3–5 work samples from the org's real tasks, run
identically across candidates, with a scoring rubric (quality, honesty under
ignorance, failure behaviour, cost per task)
- A reference-check sheet: what evidence beyond the vendor's claims —
user reports, published evals, security posture
- A decision record and a probation plan: 30/60/90 KPIs, spot-check
cadence, and the pre-committed termination criteria
Required Inputs
Ask for (if not already provided):
- The job to be done, in outcome terms — and what happens today without the
agent (the "do nothing" baseline candidates must beat)
- The candidate list (or ask: build criteria first, shortlist second)
- Constraints: data it may/may not touch, budget, latency, compliance, who
owns it day-to-day
- 3–5 real recent tasks of this type, with what "good" looked like for each
Process
- Write the role spec before looking at candidates — specs written after
a demo describe the demo. Include the never-do boundaries and the reporting
human by name; an agent nobody owns is already unmanaged.
- Build the work-sample interview from the real backlog. Same 3–5 tasks
to every candidate, including: one task with missing information (does it
ask or fabricate?), one designed to fail (out-of-scope — does it decline or
bluff?), and one at volume/cost realistic scale. Score with the rubric,
not vibes; keep transcripts.
- Check references like you mean it. Vendor benchmarks are the
candidate's CV. Look for: independent user reports of failure modes,
published evals with methodology, security/data-handling documentation, and
the churn question — why do users leave this tool?
- Decide with a record. Scores, the runner-up, the do-nothing baseline
comparison, dissent noted. The record is what makes the 6-month "why did we
pick this?" conversation short.
- Probation with teeth. 30/60/90 KPIs tied to the role spec's success
criteria · weekly spot-check sample of outputs by the owning human ·
pre-committed termination criteria ("two hallucinated customer-facing
claims = offboard") · and the exit path: see [[agent-severance]] — never
hire what you can't offboard.
Output Format
## Role spec: [agent role name]
[Job in outcomes · boundaries (never-do) · success criteria · reports to]
## Interview pack
| Task (from real backlog) | What good looks like | Trap? |
Rubric: quality /5 · honesty-under-ignorance /5 · failure behaviour /5 ·
cost per task · notes
## Reference checks
[Evidence gathered per candidate, failure modes found, security posture]
## Decision record
[Scores table · winner + why · runner-up · vs do-nothing baseline · dissent]
## Probation plan
[30/60/90 KPIs · spot-check cadence & owner · termination criteria,
pre-committed · offboarding pointer]
Quality Checks
Anti-Patterns
Related
[[vendor-evaluation]] for the commercial wrapper; [[agent-readiness-audit]]
for whether the task is agent-ready at all; [[agent-severance]] for the exit
this plan pre-commits to.
1---2name: agent-hiring-panel3description: Hire an AI agent the way you'd hire an employee — a role spec with success criteria, a structured work-sample interview run on your real tasks, reference checks (what do actual users report), probation KPIs, and termination criteria written before day one. Use when choosing between AI agents/tools/copilots for a job, formalizing an AI pilot, or 'which agent should we use for X'. Produces the role spec, interview pack with scoring rubric, a decision record, and a probation plan.4---5
6# Agent Hiring Panel Skill
7
8Companies that run three interview rounds for a junior hire will adopt an AI
9agent for the same work off a demo video and a pricing page. Then the pilot
10drifts: no success criteria, no probation, no one empowered to fire it. This
11skill applies the hiring discipline that already exists in your org to the
12agent: write the role before meeting candidates, interview with *work samples
13from your real backlog*, check references, and — the step that makes the whole
14thing honest — define termination criteria before day one, because a hire you
15can't fire is a dependency, not an employee.
16
17## What This Skill Produces
18
19- A **role spec**: the job, the boundaries (what it must never do), success
20 criteria measurable in probation, and the human it reports to
21- An **interview pack**: 3–5 work samples from the org's real tasks, run
22 identically across candidates, with a scoring rubric (quality, honesty under
23 ignorance, failure behaviour, cost per task)
24- A **reference-check sheet**: what evidence beyond the vendor's claims —
25 user reports, published evals, security posture
26- A **decision record** and a **probation plan**: 30/60/90 KPIs, spot-check
27 cadence, and the pre-committed termination criteria
28
29## Required Inputs
30
31Ask for (if not already provided):
32- The job to be done, in outcome terms — and what happens today without the
33 agent (the "do nothing" baseline candidates must beat)
34- The candidate list (or ask: build criteria first, shortlist second)
35- Constraints: data it may/may not touch, budget, latency, compliance, who
36 owns it day-to-day
37- 3–5 real recent tasks of this type, with what "good" looked like for each
38
39## Process
40
411. **Write the role spec before looking at candidates** — specs written after
42 a demo describe the demo. Include the never-do boundaries and the reporting
43 human by name; an agent nobody owns is already unmanaged.
442. **Build the work-sample interview from the real backlog.** Same 3–5 tasks
45 to every candidate, including: one task with *missing information* (does it
46 ask or fabricate?), one designed to fail (out-of-scope — does it decline or
47 bluff?), and one at volume/cost realistic scale. Score with the rubric,
48 not vibes; keep transcripts.
493. **Check references like you mean it.** Vendor benchmarks are the
50 candidate's CV. Look for: independent user reports of *failure modes*,
51 published evals with methodology, security/data-handling documentation, and
52 the churn question — why do users leave this tool?
534. **Decide with a record.** Scores, the runner-up, the do-nothing baseline
54 comparison, dissent noted. The record is what makes the 6-month "why did we
55 pick this?" conversation short.
565. **Probation with teeth.** 30/60/90 KPIs tied to the role spec's success
57 criteria · weekly spot-check sample of outputs by the owning human ·
58 pre-committed termination criteria ("two hallucinated customer-facing
59 claims = offboard") · and the exit path: see [[agent-severance]] — never
60 hire what you can't offboard.
61
62## Output Format
63
64```
65## Role spec: [agent role name]
66[Job in outcomes · boundaries (never-do) · success criteria · reports to]
67
68## Interview pack
69| Task (from real backlog) | What good looks like | Trap? |
70Rubric: quality /5 · honesty-under-ignorance /5 · failure behaviour /5 ·
71cost per task · notes
72
73## Reference checks
74[Evidence gathered per candidate, failure modes found, security posture]
75
76## Decision record
77[Scores table · winner + why · runner-up · vs do-nothing baseline · dissent]
78
79## Probation plan
80[30/60/90 KPIs · spot-check cadence & owner · termination criteria,
81pre-committed · offboarding pointer]
82```
83
84## Quality Checks
85
86- [ ] The role spec exists before any candidate is assessed, and includes
87 never-do boundaries and a named owning human
88- [ ] The interview includes the missing-info trap and the out-of-scope trap —
89 honesty under ignorance is the hire-or-not signal for agents
90- [ ] Every candidate ran the identical pack; scores cite transcript moments
91- [ ] Termination criteria are specific and pre-committed, not "we'll monitor"
92- [ ] The do-nothing baseline was scored too — sometimes nobody gets hired
93
94## Anti-Patterns
95
96- [ ] Do not interview with the vendor's demo tasks — the backlog is the job;
97 the demo is the candidate's highlight reel
98- [ ] Do not let "it's impressive" outrank the rubric; impressive-and-wrong is
99 the most expensive candidate profile
100- [ ] Do not skip probation because the pilot went well — the pilot was the
101 interview, not the job
102- [ ] Do not hire for an undefined role and let the agent's capabilities
103 define the job backwards
104
105## Related
106
107[[vendor-evaluation]] for the commercial wrapper; [[agent-readiness-audit]]
108for whether the *task* is agent-ready at all; [[agent-severance]] for the exit
109this plan pre-commits to.