Evaluating Developer Proficiency
Evaluate demonstrated capability against a fixed blueprint. Observe evidence first, apply deterministic gates second, and report uncertainty instead of producing an unconstrained model score.
Require an evaluation contract
Before evaluating, identify:
- capability and target level;
- blueprint or explicit criterion version;
- required evidence dimensions;
- challenge types;
- observable rubric criteria;
- blocking gates and minimum evidence;
- retest policy when applicable.
If the criteria are missing or ambiguous, do not silently invent them.
Assemble evidence-producing challenges
Choose challenges that can reveal the required dimensions: concept knowledge, code reading, debugging, implementation, or technical reasoning. Avoid multiple-choice-only evaluation when performance is required.
Keep challenge scope proportional to the criterion being tested. Do not add irrelevant difficulty to make an assessment feel rigorous.
Observe before judging
Record what the candidate actually demonstrated for each rubric criterion. Preserve failed gates, partial demonstrations, contradictions, and provenance. Separate evaluator interpretation from the raw response or artifact.
Apply gates and derive the result
Determine whether required criteria and evidence classes are satisfied. A strong average cannot compensate for a failed blocking dimension. Keep proficiency and confidence separate, and cap confidence when provenance is external or unverified.
Produce a structured report
Return:
- blueprint/version;
- target and demonstrated level;
- confidence;
- criterion observations;
- evidence/provenance;
- failed or uncertain gates;
- actionable next gaps;
- retest recommendation when useful.
Use language such as Skill Assessment or Proficiency Report, not accredited certification.
Ownership boundaries
This method owns assessment assembly, rubric-based observation, gate application, structured results, and feedback. It does not redefine level criteria ad hoc, sequence a career roadmap, or use free-form model scoring as the final authority.