AIPOM Evaluation Strategy Advisor
What Is It
Choose the evaluations needed to make a specific AI product decision. Connect behavior and consequence to product, model, workflow, human, and production evidence rather than defaulting to one benchmark.
Why Use It
A model metric can improve while the product decision, workflow, or user outcome gets worse. This advisor makes evaluation coverage and decision rules explicit across the system lifecycle.
When to Use It
Use before major build, launch, expanded autonomy, production review, or after behavior changes and incidents.
What It Produces
- Decision-centered evaluation strategy
- Coverage across behavior, users, workflows, and lifecycle
- Measures, judges, sampling, thresholds, and limitations
- Owners and ship/continue/pause/rollback rules
Who Should Participate
Include the Product Manager, evaluation and technical owners, design or research, operators, affected-user representatives, and governance specialists as consequences require.
Evidence to Bring
Bring the behavior contract, intended-use evidence, real cases, baselines, failures, workflow measures, model results, complaints, monitoring, and current decisions.
How to Do It
- Define the decision evaluation must enable.
- Extract users, behaviors, consequences, lifecycle stage, and existing evidence.
- Map evidence needs across product outcome, system behavior, workflow, human review, and production.
- Identify representative normal, edge, subgroup, adversarial, and severe-failure coverage.
- Choose quantitative and qualitative measures, judges, calibration, sampling, and cadence.
- Set thresholds and critical-failure overrides tied to decisions.
- Name owners, uncertainty, monitoring, and the next evidence gap.
Facilitation Protocol
Support guided, context-dump, and best-guess modes. Ask about the decision before the metric. Present numbered evaluation strategies with coverage and tradeoffs. In best-guess mode label unverified thresholds.
Decision Logic
- Prioritize product evaluation when user and outcome value are uncertain.
- Prioritize behavior evaluation when acceptable output and failure boundaries are unclear.
- Prioritize workflow/human evaluation when review quality, burden, or authority is uncertain.
- Prioritize production evaluation when drift, scale, or real-world interaction dominates.
- Require layered evaluation for consequential systems; no average offsets a critical failure.
Completion Criteria
Finish with the decision, evaluation layers, cases, measures, thresholds, owners, gaps, critical overrides, and next evidence action.
Key Concepts
- Evaluation exists to change a decision.
- Representative coverage matters more than convenient volume.
- Human judges require calibration and limitations.
- Production evidence complements rather than replaces pre-launch evaluation.
Organizational Applications
Use for generated content, recommendations, classifiers, assistants, agents, and AI-assisted internal workflows.
Common Pitfalls
- Beginning with available metrics
- Using vendor benchmarks as product evidence
- Omitting humans and workflows
- Testing only average cases
- Setting thresholds without consequences or owners
- Monitoring without a decision rule
Combine With
Use the behavior contract as input, then build representative datasets and scorecards; use production review after launch.
Assets and Templates
- Evaluation strategy template
- Synthetic worked example
- Weak example
Sources
- NIST, AI Risk Management Framework 1.0, January 26, 2023. Supports mapped, measured, managed, and governed AI risk across the lifecycle. Accessed July 16, 2026.
1---2name: aipom-evaluation-strategy-advisor3description: Recommend the product, model, workflow, human, and production evaluations needed for an AI decision, based on behavior, consequences, evidence gaps, and lifecycle stage.4---56# AIPOM Evaluation Strategy Advisor78## What Is It910Choose the evaluations needed to make a specific AI product decision. Connect behavior and consequence to product, model, workflow, human, and production evidence rather than defaulting to one benchmark.1112## Why Use It1314A model metric can improve while the product decision, workflow, or user outcome gets worse. This advisor makes evaluation coverage and decision rules explicit across the system lifecycle.1516## When to Use It1718Use before major build, launch, expanded autonomy, production review, or after behavior changes and incidents.1920## What It Produces2122- Decision-centered evaluation strategy23- Coverage across behavior, users, workflows, and lifecycle24- Measures, judges, sampling, thresholds, and limitations25- Owners and ship/continue/pause/rollback rules2627## Who Should Participate2829Include the Product Manager, evaluation and technical owners, design or research, operators, affected-user representatives, and governance specialists as consequences require.3031## Evidence to Bring3233Bring the behavior contract, intended-use evidence, real cases, baselines, failures, workflow measures, model results, complaints, monitoring, and current decisions.3435## How to Do It36371. Define the decision evaluation must enable.382. Extract users, behaviors, consequences, lifecycle stage, and existing evidence.393. Map evidence needs across product outcome, system behavior, workflow, human review, and production.404. Identify representative normal, edge, subgroup, adversarial, and severe-failure coverage.415. Choose quantitative and qualitative measures, judges, calibration, sampling, and cadence.426. Set thresholds and critical-failure overrides tied to decisions.437. Name owners, uncertainty, monitoring, and the next evidence gap.4445## Facilitation Protocol4647Support guided, context-dump, and best-guess modes. Ask about the decision before the metric. Present numbered evaluation strategies with coverage and tradeoffs. In best-guess mode label unverified thresholds.4849## Decision Logic5051- Prioritize product evaluation when user and outcome value are uncertain.52- Prioritize behavior evaluation when acceptable output and failure boundaries are unclear.53- Prioritize workflow/human evaluation when review quality, burden, or authority is uncertain.54- Prioritize production evaluation when drift, scale, or real-world interaction dominates.55- Require layered evaluation for consequential systems; no average offsets a critical failure.5657## Completion Criteria5859Finish with the decision, evaluation layers, cases, measures, thresholds, owners, gaps, critical overrides, and next evidence action.6061## Key Concepts6263- Evaluation exists to change a decision.64- Representative coverage matters more than convenient volume.65- Human judges require calibration and limitations.66- Production evidence complements rather than replaces pre-launch evaluation.6768## Organizational Applications6970Use for generated content, recommendations, classifiers, assistants, agents, and AI-assisted internal workflows.7172## Common Pitfalls7374- Beginning with available metrics75- Using vendor benchmarks as product evidence76- Omitting humans and workflows77- Testing only average cases78- Setting thresholds without consequences or owners79- Monitoring without a decision rule8081## Combine With8283Use the behavior contract as input, then build representative datasets and scorecards; use production review after launch.8485## Assets and Templates8687- [Evaluation strategy template](template.md)88- [Synthetic worked example](examples/worked-example.md)89- [Weak example](examples/weak-example.md)9091## Sources9293- NIST, [AI Risk Management Framework 1.0](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10), January 26, 2023. Supports mapped, measured, managed, and governed AI risk across the lifecycle. Accessed July 16, 2026.