# Interview Question Kit

> Generates a structured candidate-interview question set where every behavioral (STAR) and role-specific item maps to a defined competency and ships with strong/borderline/red-flag answer anchors plus delivery notes. Use when someone asks "write interview questions for this role", "build an interview loop", "turn this rubric into questions per interviewer", or is planning a hiring interview, building a question bank, or standardizing how a panel assesses candidates. Do NOT use to consolidate completed scorecards into a hire decision - use interview-debrief-synthesizer instead; do NOT use for user-research or customer discovery interviews - use interview-guide-builder instead; do NOT use to build the competency rubric or scorecard itself - use hiring-scorecard or screening-rubric-builder instead.

- Skill: `skillmedev/interview-question-kit` (Agent Skill)
- Install (CLI): `npx skillmds@latest add skillmedev/interview-question-kit`
- Raw SKILL.md: https://api.skillmd.com/api/skills/skillmedev/interview-question-kit/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Product & Planning
- Author: SkillMedev (https://skillmd.com/u/skillmedev)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/skillmedev/interview-question-kit

---


# Interview Question Kit

Turn a role's competencies into a fair, comparable question set with per-question scoring anchors, so interviewers assess evidence instead of rapport. The costly mistake this prevents is the unstructured loop: five interviewers improvising different questions, scoring on likability, producing scorecards that cannot be compared - which is how teams hire the best storyteller instead of the best candidate, and how legally risky questions slip in.

## Operating procedure

Order matters: questions written before the competency map exist to sound clever, not to gather evidence.

### Step 1: Gather the rubric

Collect the role's defined, job-relevant competencies and their weights, plus the job description, loop size (how many interviewers, how long each session), and seniority level. If no rubric exists, derive 4-6 competencies from the job description, label them as derived, and confirm them before writing a single question. Never write a question that does not map to a competency on this list.

### Step 2: Allocate competencies across the loop

Assign each competency to one or two interviewers who probe it deeply, rather than every interviewer skimming all of them - overlap wastes candidate time and produces shallow, redundant evidence. Write 2-3 items per competency.

### Step 3: Write behavioral items in STAR form

Ask for specific past situations ("Tell me about a time you shipped under a hard deadline with incomplete requirements"), not hypotheticals - past behavior is evidence; hypotheticals are performance. For each item, list follow-up probes that separately pull out Situation, Task, Action, and Result, because candidates default to narrating the team's work instead of their own. The probe "What did *you* do, specifically?" earns its place in every kit.

### Step 4: Add a role-specific work sample

Pair behavioral items with a short, realistic job task: a design discussion, a code reading, a mock customer email, a prioritization exercise. Keep it role-true, time-boxed, and identical across candidates. At least one work-sample item per loop.

### Step 5: Write strong/borderline/red-flag anchors for every item

Spell out what a strong answer demonstrates, what a borderline answer looks like, and the red flags - tied to the competency, not to likability. Anchors are what make two interviewers' scores comparable.

### Step 6: Standardize delivery

Specify question order, time budget per item, and which follow-ups are allowed, so the same role gets the same loop every time. A 45-60 minute session realistically fits 3-4 behavioral items with probes, or 2 items plus a work sample - more than that produces surface answers.

### Step 7: Screen for compliance

Remove or rewrite any item a candidate could reasonably read as probing a protected characteristic: age, family or marital status, pregnancy, religion, disability, health, national origin - and salary history where restricted. Do this as a final pass on the full set, because risky questions often emerge from innocent-sounding "culture" items.

## Worked artifact: one kit entry (copy this structure per item)

```
COMPETENCY: Ownership under ambiguity (weight: 25%)
INTERVIEWER: [FILL: name] - Session 2 of 4 (45 min; this item: 12 min)

QUESTION (behavioral, STAR):
"Tell me about a time you were handed a project with unclear requirements
and a fixed deadline. Walk me through what happened."

PROBES (use as needed, in order):
- Situation: "What was the context - team size, what was at stake?"
- Task: "What specifically were you responsible for, versus the team?"
- Action: "What did YOU do first? What did you decide not to do?"
- Result: "How did it land? What would you do differently?"

ANCHORS:
- Strong: names the ambiguity explicitly, describes how they scoped or
  de-risked it themselves (stakeholder alignment, cutting scope, spike),
  owns a concrete personal action, states a measurable result AND a lesson.
- Borderline: real story but narrates the team's work; result stated
  without their causal role; needed every probe to surface specifics.
- Red flag: hypothetical answer despite redirection; blames others for the
  ambiguity; cannot name a single decision they personally made.

SCORE: 1 (red flag) / 2 (borderline) / 3 (strong) - record evidence quotes, not impressions.
```

A full kit repeats this block for every item, prefaced by a one-page loop plan: competency-to-interviewer matrix, session order, and time budgets.

## Deliverable

Produce the complete kit: a competency-to-interviewer allocation matrix, 2-3 items per competency in the block format above (each with STAR probes and all three anchors), at least one work-sample exercise with its own anchors, and delivery notes (order, time budgets, allowed follow-ups). The kit is the input to scorecards - not a hire recommendation.

## Do NOT

- Do NOT generate questions touching age, family or marital status, pregnancy, religion, disability, health, national origin, or salary history where restricted - these create legal exposure regardless of intent.
- Do NOT write hypothetical "what would you do" questions where a past-behavior or work-sample item would yield real evidence.
- Do NOT allow free-form, candidate-specific tangents that break comparability across candidates.
- Do NOT assess anything outside the defined, job-relevant competencies - "culture fit" without a defined competency is bias with a nicer name.
- Do NOT issue a verdict, ranking, or hire/no-hire call; the panel decides, and the kit only informs that decision.
- Do NOT overload a session - more than 4 behavioral items in 45 minutes guarantees shallow answers on all of them.

## Quality bar

- Every question traces to exactly one competency on the rubric; no orphan questions.
- Every question has all three anchors (strong, borderline, red flag), each tied to the competency.
- At least one work-sample item per loop, identical across candidates.
- The loop is reproducible: another interviewer could run it identically from the kit alone.
- The compliance pass is done on the final set, and derived competencies are labeled as derived.

## Escalation and neighbors

This kit is not legal advice; for roles in regulated industries, unionized environments, or jurisdictions with strict interview law, have HR or employment counsel review the final set. Upstream, build the rubric with hiring-scorecard or screening-rubric-builder and the posting with job-description-writer (bias-check it with jd-bias-scrubber). Downstream, consolidate completed scorecards with interview-debrief-synthesizer. For user-research interviews - a different craft with different rules - use interview-guide-builder.

