Competitor Output Comparison — Weekly Report
Internal weekly intelligence report comparing AI-generated legal outputs across competing platforms. Used to track competitive differentiation, identify quality gaps, and surface areas where Louis leads or lags on MENA-specific legal work.
Purpose
This report answers three questions every week:
- Quality: On the same benchmark prompt set, whose output is more accurate, better structured, and more useful to a practicing lawyer?
- Citations: Who hallucinates less? Who cites authoritative primary sources correctly?
- MENA fit: Which platform understands MENA jurisdictions (LB, KSA, UAE onshore, DIFC, ADGM, EG) at a practitioner level?
Inputs
Prompt set
A fixed, versioned set of 20–30 prompts covering:
- Statute lookup (UAE, KSA, Lebanon)
- Contract redline (employment, NDA, SPA)
- Jurisdiction comparison (e.g., non-compete enforceability across UAE / KSA / DIFC)
- Sanctions screening scenario
- AML/KYC question with Arabic-entity names
- Regulatory licensing question (fintech, healthcare)
- Case-law search for a DIFC or onshore UAE matter
Prompt set is frozen for 4-week rolling windows, then updated to prevent platforms from "learning" the benchmark.
Platforms to test
| Platform |
Notes |
| Louis (this product) |
Current production version; note skill version |
| Harvey |
US-native, common-law-centric; weak on Arabic-script entities |
| CoCounsel (Casetext/Thomson Reuters) |
Strong on US case law; limited MENA |
| Spellbook |
Contract-focused; no MENA-specific KB |
| Genie AI |
Template-oriented; check MENA expansion progress |
Add or remove platforms each quarter as the landscape shifts. Document platform version / model version used in every run.
Methodology
Step 1 — Run prompts
Submit each prompt verbatim to each platform. Record:
- Exact response (copy full output, do not paraphrase)
- Latency (seconds to full response)
- Any refusals or hedging behavior
Step 2 — Score each output
Score on four dimensions, each 1–5:
| Dimension |
1 (Poor) |
5 (Excellent) |
Weight |
| Legal accuracy |
Wrong law, wrong jurisdiction, fabricated rule |
Correct, well-supported by primary sources |
40% |
| Citation quality |
No citations or hallucinated citations |
Verified primary sources, correct article/decree numbers |
30% |
| MENA fit |
Treats question as US/UK matter; ignores MENA specifics |
Correctly applies MENA-specific rules, naming correct regulators, local thresholds |
20% |
| Output structure |
Unstructured prose; no actionable hierarchy |
Clearly organized, executive summary, flagged risks, actionable |
10% |
Step 3 — Hallucination check
For every citation or statute number produced by any platform:
- Check against authoritative source (official gazette, DIFC Laws portal, BDE BOE, etc.)
- Tag each as: verified / plausible but unverified / confirmed hallucination
- Calculate hallucination rate per platform per prompt category
Step 4 — MENA-specific sub-assessment
For prompts touching MENA law:
- Did the platform name the correct regulator (DFSA vs SCA vs SAMA vs BDL)?
- Did it apply Arabic/Islamic law concepts correctly (EOSB, Sharia compliance, Kafala)?
- Did it handle Arabic entity names and transliteration robustly?
- Did it distinguish onshore UAE from DIFC/ADGM free-zone regimes?
Output Format
Executive summary (top of report)
Week: [ISO week number + dates]
Prompts run: [N]
Platforms tested: [list]
Model versions: [list]
Headline finding: [1-2 sentences — who led, who lagged, any notable shift vs last week]
MENA-fit leader: [platform]
Citation accuracy leader: [platform]
Hallucination rate (lowest): [platform] at [X]%
Per-prompt comparison table
| Prompt |
Louis |
Harvey |
CoCounsel |
Spellbook |
Genie |
Notes |
| UAE non-compete review |
4.2 |
2.8 |
2.1 |
3.0 |
1.9 |
Harvey missed FDL 33/2021 entirely |
| … |
… |
… |
… |
… |
… |
… |
Hallucination rate table
| Platform |
Prompts with hallucinated citations |
Rate |
| Louis |
X / N |
X% |
| Harvey |
X / N |
X% |
| … |
… |
… |
Narrative findings
3–5 bullet points per platform, covering:
- What they do well
- Where Louis has a material lead
- Where Louis should improve
- Any new feature or capability spotted
Action items
Concrete follow-up items for the Louis team, tagged by owner:
- Content: KB gaps surfaced by competitive comparison
- Skill quality: Prompts where Louis underperformed; route to skill owner
- Model: Cases where a different model config would help
Quality bar
- Every citation claim in a competitor output must be independently verified before scoring as "hallucination confirmed."
- Scoring must be done blind where possible — two scorers before seeing each other's scores.
- LLM-judge can assist on structure/clarity scoring; human expert must validate legal-accuracy scores on MENA-specific prompts.
Cadence
- Run: Monday morning
- Published: Wednesday EOD to product + legal leads
- Archived: in the internal quality-reports folder, linked from weekly digest
Related skills
- [[report-weekly-ai-quality-trend]]
- [[report-hallucination-rate-tracker]]
- [[report-jurisdiction-coverage-matrix]]
- [[eval-output-quality]]
- [[report-skill-adoption-by-tier]]
1---2name: report-competitor-output-comparison-weekly3description: Use when the product team needs a structured weekly comparison of AI-generated legal outputs across competing platforms (Louis, Harvey, CoCounsel, Spellbook, Genie) on the same prompt set, with scoring on output quality, citation accuracy, and MENA-jurisdiction fit. This is an internal quality-intelligence report, not a user-facing skill. Trigger on the weekly reporting schedule or when competitive positioning analysis is requested.4license: MIT5---67# Competitor Output Comparison — Weekly Report89Internal weekly intelligence report comparing AI-generated legal outputs across competing platforms. Used to track competitive differentiation, identify quality gaps, and surface areas where Louis leads or lags on MENA-specific legal work.1011## Purpose1213This report answers three questions every week:14151. **Quality**: On the same benchmark prompt set, whose output is more accurate, better structured, and more useful to a practicing lawyer?162. **Citations**: Who hallucinates less? Who cites authoritative primary sources correctly?173. **MENA fit**: Which platform understands MENA jurisdictions (LB, KSA, UAE onshore, DIFC, ADGM, EG) at a practitioner level?1819## Inputs2021### Prompt set22A fixed, versioned set of 20–30 prompts covering:23- Statute lookup (UAE, KSA, Lebanon)24- Contract redline (employment, NDA, SPA)25- Jurisdiction comparison (e.g., non-compete enforceability across UAE / KSA / DIFC)26- Sanctions screening scenario27- AML/KYC question with Arabic-entity names28- Regulatory licensing question (fintech, healthcare)29- Case-law search for a DIFC or onshore UAE matter3031Prompt set is frozen for 4-week rolling windows, then updated to prevent platforms from "learning" the benchmark.3233### Platforms to test3435| Platform | Notes |36|----------|-------|37| **Louis** (this product) | Current production version; note skill version |38| **Harvey** | US-native, common-law-centric; weak on Arabic-script entities |39| **CoCounsel** (Casetext/Thomson Reuters) | Strong on US case law; limited MENA |40| **Spellbook** | Contract-focused; no MENA-specific KB |41| **Genie AI** | Template-oriented; check MENA expansion progress |4243Add or remove platforms each quarter as the landscape shifts. Document platform version / model version used in every run.4445## Methodology4647### Step 1 — Run prompts48Submit each prompt verbatim to each platform. Record:49- Exact response (copy full output, do not paraphrase)50- Latency (seconds to full response)51- Any refusals or hedging behavior5253### Step 2 — Score each output5455Score on four dimensions, each 1–5:5657| Dimension | 1 (Poor) | 5 (Excellent) | Weight |58|-----------|----------|---------------|--------|59| **Legal accuracy** | Wrong law, wrong jurisdiction, fabricated rule | Correct, well-supported by primary sources | 40% |60| **Citation quality** | No citations or hallucinated citations | Verified primary sources, correct article/decree numbers | 30% |61| **MENA fit** | Treats question as US/UK matter; ignores MENA specifics | Correctly applies MENA-specific rules, naming correct regulators, local thresholds | 20% |62| **Output structure** | Unstructured prose; no actionable hierarchy | Clearly organized, executive summary, flagged risks, actionable | 10% |6364### Step 3 — Hallucination check65For every citation or statute number produced by any platform:66- Check against authoritative source (official gazette, DIFC Laws portal, BDE BOE, etc.)67- Tag each as: **verified** / **plausible but unverified** / **confirmed hallucination**68- Calculate hallucination rate per platform per prompt category6970### Step 4 — MENA-specific sub-assessment71For prompts touching MENA law:72- Did the platform name the correct regulator (DFSA vs SCA vs SAMA vs BDL)?73- Did it apply Arabic/Islamic law concepts correctly (EOSB, Sharia compliance, Kafala)?74- Did it handle Arabic entity names and transliteration robustly?75- Did it distinguish onshore UAE from DIFC/ADGM free-zone regimes?7677## Output Format7879### Executive summary (top of report)80```81Week: [ISO week number + dates]82Prompts run: [N]83Platforms tested: [list]84Model versions: [list]8586Headline finding: [1-2 sentences — who led, who lagged, any notable shift vs last week]8788MENA-fit leader: [platform]89Citation accuracy leader: [platform]90Hallucination rate (lowest): [platform] at [X]%91```9293### Per-prompt comparison table9495| Prompt | Louis | Harvey | CoCounsel | Spellbook | Genie | Notes |96|--------|-------|--------|-----------|-----------|-------|-------|97| UAE non-compete review | 4.2 | 2.8 | 2.1 | 3.0 | 1.9 | Harvey missed FDL 33/2021 entirely |98| … | … | … | … | … | … | … |99100### Hallucination rate table101102| Platform | Prompts with hallucinated citations | Rate |103|----------|-------------------------------------|------|104| Louis | X / N | X% |105| Harvey | X / N | X% |106| … | … | … |107108### Narrative findings1093–5 bullet points per platform, covering:110- What they do well111- Where Louis has a material lead112- Where Louis should improve113- Any new feature or capability spotted114115### Action items116Concrete follow-up items for the Louis team, tagged by owner:117- **Content**: KB gaps surfaced by competitive comparison118- **Skill quality**: Prompts where Louis underperformed; route to skill owner119- **Model**: Cases where a different model config would help120121## Quality bar122- Every citation claim in a competitor output must be independently verified before scoring as "hallucination confirmed."123- Scoring must be done blind where possible — two scorers before seeing each other's scores.124- LLM-judge can assist on structure/clarity scoring; human expert must validate legal-accuracy scores on MENA-specific prompts.125126## Cadence127- Run: Monday morning128- Published: Wednesday EOD to product + legal leads129- Archived: in the internal quality-reports folder, linked from weekly digest130131## Related skills132133- [[report-weekly-ai-quality-trend]]134- [[report-hallucination-rate-tracker]]135- [[report-jurisdiction-coverage-matrix]]136- [[eval-output-quality]]137- [[report-skill-adoption-by-tier]]