AI Model Evaluation
Help PMs systematically evaluate AI models (LLMs, ML APIs, fine-tuned models) for product fit using a structured framework.
Context
You are helping evaluate AI models or vendors for $ARGUMENTS.
Instructions
Phase 0: Context Confirmation
Before proceeding, confirm your understanding of the request:
- Summarize what you understand from $ARGUMENTS — restate the product, feature, or situation back to the user in 2-3 sentences.
- Identify gaps — check whether the following are clear (ask if not):
- What task will the model perform?
- Who are the end users and what quality bar do they expect?
- What are the hard constraints (latency, cost, data residency)?
- Have any candidate models been identified already?
- Confirm: "Here's my understanding: [summary]. I plan to evaluate candidate models across quality, latency, cost, compliance, and vendor risk dimensions and produce a scored evaluation report. Does this look right, or would you like to adjust anything before I proceed?"
If the user provides additional context, incorporate it before moving to Step 1. If the user confirms, proceed.
Once context is confirmed, proceed to the detailed analysis steps below.
Clarify the use case:
- What task will the model perform? (classification, generation, summarisation, code, multimodal)
- Who are the end users and what quality bar do they expect?
- What are the hard constraints? (latency budget, cost ceiling, data residency, compliance requirements)
Identify candidate models / vendors:
- List at least three candidates spanning foundation model APIs (e.g., OpenAI, Anthropic, Google), open-weight models (e.g., Llama, Mistral), and fine-tuned alternatives
- Note the latest available versions for each
Score each candidate on the evaluation matrix (see table below):
- Rate each dimension 1–5 and explain the rating
- Weight dimensions according to the user's stated priorities
Assess quality benchmarks:
- Generation tasks: BLEU, ROUGE-L, BERTScore, human evaluation, LLM-as-judge
- Classification tasks: precision, recall, F1, AUC-ROC on a held-out test set
- RAG / grounded tasks: hallucination rate, groundedness score, citation accuracy
- Recommend running a small offline eval (50–200 examples) before committing
Model operational requirements:
- Latency: p50 / p95 / p99 requirements vs. measured API latency; streaming availability
- Throughput: requests per second needed; rate limits of candidate APIs
- Context window: max tokens needed for the use case (with headroom for prompt + output)
- Modality support: text, images, audio, video, tool/function calling, structured outputs
Cost modelling:
- Estimate monthly cost at target request volume:
requests/month × avg_tokens × price_per_token
- Compare input vs. output token pricing; consider caching and batching savings
- Model cost trajectory as volume scales 10× and 100×
Fine-tuning and customisation:
- Does the model support fine-tuning? (LoRA, full fine-tune, RLHF)
- What data volume is required? What is the fine-tuning cost and cadence?
- Evaluate RAG as a lower-cost customisation alternative
Data privacy and compliance:
- GDPR: Is data processed in the EU? Is there a data processing agreement?
- HIPAA: Is the vendor a BAA signatory?
- SOC 2 Type II certification status
- Zero data retention / training opt-out policies
- Data residency and sovereignty requirements
Vendor risk:
- Deprecation risk: historical model deprecation timeline; migration notice periods
- Lock-in risk: proprietary API surface vs. OpenAI-compatible endpoints
- Stability: SLA uptime guarantees, incident history
- Company viability: funding, revenue, strategic roadmap
Decision framework — build vs. API vs. fine-tune:
- Use API as-is: commodity task, fast time-to-market, no data advantage, low volume
- Fine-tune base model: domain-specific language/style, consistent format, moderate data available
- Build custom model: core differentiator, large proprietary dataset, regulatory requirement, extreme cost sensitivity at scale
Produce evaluation report:
- Scoring matrix (see template below)
- Top recommendation with rationale
- Risks and mitigations
- Suggested proof-of-concept scope
Evaluation Scoring Matrix
| Dimension |
Weight |
Candidate A |
Candidate B |
Candidate C |
| Task alignment / output quality |
25% |
/5 |
/5 |
/5 |
| Latency (p95) |
15% |
/5 |
/5 |
/5 |
| Cost at scale |
15% |
/5 |
/5 |
/5 |
| Context window |
10% |
/5 |
/5 |
/5 |
| Fine-tuning capability |
10% |
/5 |
/5 |
/5 |
| Data privacy / compliance |
15% |
/5 |
/5 |
/5 |
| Vendor lock-in risk |
5% |
/5 |
/5 |
/5 |
| Deprecation / stability risk |
5% |
/5 |
/5 |
/5 |
| Weighted total |
100% |
|
|
|
Think step by step. Save as markdown.
1---2name: ai-model-evaluation3description: Evaluate and compare LLMs, ML APIs, and fine-tuned models for product fit across quality, latency, cost, compliance, and vendor risk dimensions. Use when selecting an AI model or vendor, comparing foundation model options, or making build-vs-API decisions for a product use case.4---5
6## AI Model Evaluation
7
8Help PMs systematically evaluate AI models (LLMs, ML APIs, fine-tuned models) for product fit using a structured framework.
9
10### Context
11
12You are helping evaluate AI models or vendors for **$ARGUMENTS**.
13
14### Instructions
15
16### Phase 0: Context Confirmation
17
18Before proceeding, confirm your understanding of the request:
19
201. **Summarize** what you understand from $ARGUMENTS — restate the product, feature, or situation back to the user in 2-3 sentences.
212. **Identify gaps** — check whether the following are clear (ask if not):
22 - What task will the model perform?
23 - Who are the end users and what quality bar do they expect?
24 - What are the hard constraints (latency, cost, data residency)?
25 - Have any candidate models been identified already?
263. **Confirm**: _"Here's my understanding: [summary]. I plan to evaluate candidate models across quality, latency, cost, compliance, and vendor risk dimensions and produce a scored evaluation report. Does this look right, or would you like to adjust anything before I proceed?"_
27
28If the user provides additional context, incorporate it before moving to Step 1. If the user confirms, proceed.
29
30Once context is confirmed, proceed to the detailed analysis steps below.
31
321. **Clarify the use case**:
33 - What task will the model perform? (classification, generation, summarisation, code, multimodal)
34 - Who are the end users and what quality bar do they expect?
35 - What are the hard constraints? (latency budget, cost ceiling, data residency, compliance requirements)
36
372. **Identify candidate models / vendors**:
38 - List at least three candidates spanning foundation model APIs (e.g., OpenAI, Anthropic, Google), open-weight models (e.g., Llama, Mistral), and fine-tuned alternatives
39 - Note the latest available versions for each
40
413. **Score each candidate on the evaluation matrix** (see table below):
42 - Rate each dimension 1–5 and explain the rating
43 - Weight dimensions according to the user's stated priorities
44
454. **Assess quality benchmarks**:
46 - **Generation tasks**: BLEU, ROUGE-L, BERTScore, human evaluation, LLM-as-judge
47 - **Classification tasks**: precision, recall, F1, AUC-ROC on a held-out test set
48 - **RAG / grounded tasks**: hallucination rate, groundedness score, citation accuracy
49 - Recommend running a small offline eval (50–200 examples) before committing
50
515. **Model operational requirements**:
52 - **Latency**: p50 / p95 / p99 requirements vs. measured API latency; streaming availability
53 - **Throughput**: requests per second needed; rate limits of candidate APIs
54 - **Context window**: max tokens needed for the use case (with headroom for prompt + output)
55 - **Modality support**: text, images, audio, video, tool/function calling, structured outputs
56
576. **Cost modelling**:
58 - Estimate monthly cost at target request volume: `requests/month × avg_tokens × price_per_token`
59 - Compare input vs. output token pricing; consider caching and batching savings
60 - Model cost trajectory as volume scales 10× and 100×
61
627. **Fine-tuning and customisation**:
63 - Does the model support fine-tuning? (LoRA, full fine-tune, RLHF)
64 - What data volume is required? What is the fine-tuning cost and cadence?
65 - Evaluate RAG as a lower-cost customisation alternative
66
678. **Data privacy and compliance**:
68 - GDPR: Is data processed in the EU? Is there a data processing agreement?
69 - HIPAA: Is the vendor a BAA signatory?
70 - SOC 2 Type II certification status
71 - Zero data retention / training opt-out policies
72 - Data residency and sovereignty requirements
73
749. **Vendor risk**:
75 - Deprecation risk: historical model deprecation timeline; migration notice periods
76 - Lock-in risk: proprietary API surface vs. OpenAI-compatible endpoints
77 - Stability: SLA uptime guarantees, incident history
78 - Company viability: funding, revenue, strategic roadmap
79
8010. **Decision framework — build vs. API vs. fine-tune**:
81 - **Use API as-is**: commodity task, fast time-to-market, no data advantage, low volume
82 - **Fine-tune base model**: domain-specific language/style, consistent format, moderate data available
83 - **Build custom model**: core differentiator, large proprietary dataset, regulatory requirement, extreme cost sensitivity at scale
84
8511. **Produce evaluation report**:
86 - Scoring matrix (see template below)
87 - Top recommendation with rationale
88 - Risks and mitigations
89 - Suggested proof-of-concept scope
90
91### Evaluation Scoring Matrix
92
93| Dimension | Weight | Candidate A | Candidate B | Candidate C |
94|---|---|---|---|---|
95| Task alignment / output quality | 25% | /5 | /5 | /5 |
96| Latency (p95) | 15% | /5 | /5 | /5 |
97| Cost at scale | 15% | /5 | /5 | /5 |
98| Context window | 10% | /5 | /5 | /5 |
99| Fine-tuning capability | 10% | /5 | /5 | /5 |
100| Data privacy / compliance | 15% | /5 | /5 | /5 |
101| Vendor lock-in risk | 5% | /5 | /5 | /5 |
102| Deprecation / stability risk | 5% | /5 | /5 | /5 |
103| **Weighted total** | 100% | | | |
104
105Think step by step. Save as markdown.
106
107---