Synthetic conversation generation
Real transcripts are the richest source of test cases and the fastest way to leak
PII into repos, eval sets, and vendor sandboxes. Poor synthetic data is useless;
poorly redacted real data is a compliance incident.
Goal: representative coverage without identifiable customers, with explicit
labelling of what is synthetic vs derived.
Choose the source strategy
| Approach |
When to use |
Risk |
| Fully synthetic |
Greenfield scenarios, adversarial cases, rare edge types |
Unrealistic phrasing if not validated |
| Redacted production |
Match real phrasing and driver mix |
Re-identification if redaction weak |
| Structured templates + variation |
Scale coverage of known drivers |
Misses messy real-world turns |
| Human-authored gold cases |
High-stakes regulated scenarios |
Expensive; small set |
Usually combine: redacted real for phrasing realism, synthetic for edges and
attacks, human gold for policy boundaries.
Fully synthetic: make it realistic
- Seed from contact-driver taxonomy, not from article titles.
- Vary register, language, typos, anger, multi-turn confusion.
- Include negative cases: out-of-scope, missing info, customer wrong about facts.
- Validate a sample against real traffic distribution — if 30% of volume is
billing disputes, the set should not be 30% password resets.
Label every synthetic case synthetic: true in metadata. Never mix unlabelled synthetic
into calibration sets without documenting origin.
Redacted production: redaction that holds
Minimum removals or replacements:
- Names, emails, phones, addresses, account numbers, order ids, payment details
- URLs with tokens, internal agent names, ticket ids tied to real people
- Rare quasi-identifiers (specific amounts + dates + product combo)
Replace with consistent fictitious tokens (CUSTOMER_A, ORDER_001) across turns
so multi-turn coherence remains.
After redaction, run a re-identification review on a sample: could someone internal
recognise the customer from timing, product, or story? If yes, strip more or discard.
Do not use production exports in shared eval repos without legal/process sign-off.
Edge-case coverage checklist
Ensure explicit cases for:
- Multi-turn clarification and correction ("no, I meant the other card")
- Language mix and code-switching
- Anger, legal threats, vulnerability signals
- Prompt injection and social engineering (in sandbox only)
- Bot should defer: no KB, account-specific, authentication required
- Long silence gaps, channel switches (email → chat references)
- Attachment references (metadata only; no real files)
Edge cases are where production bots fail silently — overweight them in the set
relative to volume.
Leakage risks when generating from real transcripts
| Failure |
Consequence |
| LLM paraphrase of real chat |
Embeds memorised PII or rare facts |
| Few-shot examples from production |
Eval set becomes training leakage |
| Synthetic "based on ticket 12345" in prompt |
Identifiers slip into output |
| Shared vendor thread with raw paste |
Data leaves your boundary |
Rules:
- Do not prompt with full raw transcripts in external models without clearance.
- Generate structure first (driver, turns, expected outcome), then language.
- Scan outputs with the same PII detectors used in production logging.
- Keep generation prompts out of retrieval and out of agent KB.
Versioning and hygiene
- Store fixtures in version control with manifest: origin, date, language, driver.
- Freeze eval sets; changes are version bumps, not silent edits.
- Separate regression tier (incident-derived) from exploratory sets.
- Rotate redacted samples when retention policy expires source tickets.
Traps
- Synthetic only in English when bot serves six languages.
- Happy-path bias — synthetic customers are too polite and too literate.
- Duplicate near-copies — fifty variants of the same dispute inflate scores.
- Using synthetic for grader calibration without human review — circular quality.
Present results to the user
- Strategy — synthetic vs redacted vs hybrid; rationale.
- Manifest — count by driver, language, origin label, edge-case tags.
- Redaction protocol — fields removed, token scheme, approval path.
- Coverage gaps — drivers or languages underrepresented.
- Leakage controls — generation environment, scanning, storage boundaries.
- Sample cases — 2–3 examples with metadata (no real PII).
- Maintenance — refresh cadence, versioning, and retirement rules.
1---2name: cx-synthetic-conversation-generation3description: Use to build test conversations for QA and AI eval without putting production PII in fixtures — synthetic vs redacted real data, edge-case coverage, and leakage risk when generation is done poorly. Trigger for "synthetic test conversations", "eval data without PII", "generate test tickets", "redact transcripts for testing", "fixture conversations for the bot", or avoiding production data in test sets.4---56# Synthetic conversation generation78Real transcripts are the richest source of test cases and the fastest way to **leak9PII into repos, eval sets, and vendor sandboxes.** Poor synthetic data is useless;10poorly redacted real data is a compliance incident.1112Goal: **representative coverage without identifiable customers**, with explicit13labelling of what is synthetic vs derived.1415## Choose the source strategy1617| Approach | When to use | Risk |18| --- | --- | --- |19| **Fully synthetic** | Greenfield scenarios, adversarial cases, rare edge types | Unrealistic phrasing if not validated |20| **Redacted production** | Match real phrasing and driver mix | Re-identification if redaction weak |21| **Structured templates + variation** | Scale coverage of known drivers | Misses messy real-world turns |22| **Human-authored gold cases** | High-stakes regulated scenarios | Expensive; small set |2324Usually combine: **redacted real for phrasing realism**, **synthetic for edges and25attacks**, **human gold for policy boundaries**.2627## Fully synthetic: make it realistic2829- Seed from **contact-driver taxonomy**, not from article titles.30- Vary **register, language, typos, anger, multi-turn confusion**.31- Include **negative cases**: out-of-scope, missing info, customer wrong about facts.32- Validate a sample against **real traffic distribution** — if 30% of volume is33 billing disputes, the set should not be 30% password resets.3435Label every synthetic case `synthetic: true` in metadata. Never mix unlabelled synthetic36into calibration sets without documenting origin.3738## Redacted production: redaction that holds3940Minimum removals or replacements:4142- Names, emails, phones, addresses, account numbers, order ids, payment details43- URLs with tokens, internal agent names, ticket ids tied to real people44- Rare quasi-identifiers (specific amounts + dates + product combo)4546**Replace with consistent fictitious tokens** (`CUSTOMER_A`, `ORDER_001`) across turns47so multi-turn coherence remains.4849After redaction, run a **re-identification review** on a sample: could someone internal50recognise the customer from timing, product, or story? If yes, strip more or discard.5152Do not use production exports in **shared eval repos** without legal/process sign-off.5354## Edge-case coverage checklist5556Ensure explicit cases for:5758- Multi-turn clarification and correction ("no, I meant the other card")59- Language mix and code-switching60- Anger, legal threats, vulnerability signals61- Prompt injection and social engineering (in sandbox only)62- Bot should defer: no KB, account-specific, authentication required63- Long silence gaps, channel switches (email → chat references)64- Attachment references (metadata only; no real files)6566**Edge cases are where production bots fail silently** — overweight them in the set67relative to volume.6869## Leakage risks when generating from real transcripts7071| Failure | Consequence |72| --- | --- |73| LLM paraphrase of real chat | Embeds memorised PII or rare facts |74| Few-shot examples from production | Eval set becomes training leakage |75| Synthetic "based on ticket 12345" in prompt | Identifiers slip into output |76| Shared vendor thread with raw paste | Data leaves your boundary |7778Rules:7980- **Do not prompt with full raw transcripts** in external models without clearance.81- **Generate structure first** (driver, turns, expected outcome), then language.82- **Scan outputs** with the same PII detectors used in production logging.83- **Keep generation prompts out of retrieval** and out of agent KB.8485## Versioning and hygiene8687- Store fixtures in version control with **manifest**: origin, date, language, driver.88- **Freeze eval sets**; changes are version bumps, not silent edits.89- Separate **regression tier** (incident-derived) from **exploratory** sets.90- Rotate redacted samples when retention policy expires source tickets.9192## Traps9394- **Synthetic only in English** when bot serves six languages.95- **Happy-path bias** — synthetic customers are too polite and too literate.96- **Duplicate near-copies** — fifty variants of the same dispute inflate scores.97- **Using synthetic for grader calibration without human review** — circular quality.9899## Present results to the user1001011. **Strategy** — synthetic vs redacted vs hybrid; rationale.1022. **Manifest** — count by driver, language, origin label, edge-case tags.1033. **Redaction protocol** — fields removed, token scheme, approval path.1044. **Coverage gaps** — drivers or languages underrepresented.1055. **Leakage controls** — generation environment, scanning, storage boundaries.1066. **Sample cases** — 2–3 examples with metadata (no real PII).1077. **Maintenance** — refresh cadence, versioning, and retirement rules.