# Test Generation Benchmark: Code-based vs TDD

> For new PDD development, prefer TDD approach - generates more comprehensive tests focused on intent, not implementation accidents.

- Skill: `tools-only/test-generation-benchmark-code-based-vs-tdd` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add tools-only/test-generation-benchmark-code-based-vs-tdd`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tools-only/test-generation-benchmark-code-based-vs-tdd/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: tools-only (https://skillmd.com/u/tools-only)
- Updated: 2026-09-29
- Page: https://skillmd.com/skills/tools-only/test-generation-benchmark-code-based-vs-tdd

---

# Test Generation Benchmark: Code-based vs TDD

We benchmarked two ways to generate tests with `pdd test`:

**Code-based (original):** `pdd test prompt.prompt src/module.py`
- Analyzes actual implementation code

**TDD/Example-based (new):** `pdd test prompt.prompt examples/module_example.py`
- Analyzes usage examples showing intended behavior
- Triggered by `_example` suffix in filename

## Results

| Metric | Code-based | Example-based |
|--------|-----------|---------------|
| Test functions | 20 | 37 |
| Tests passed | 17 | 34 |
| Private method refs | 0 | 0 |

## Key Findings

1. **Example-based generated 85% more tests** - more granular, better failure isolation

2. **Neither referenced private methods** - both respected "don't test internals" guidance

3. **Example-based tests are more portable** - one assertion per function vs batched loops

## When to Use Each

| Use Example-based (TDD) | Use Code-based |
|------------------------|----------------|
| Writing tests first | Testing existing code |
| API contract matters | Need coverage analysis |
| Implementation will change | Verifying specific behaviors |

## TL;DR

For new PDD development, prefer TDD approach - generates more comprehensive tests focused on intent, not implementation accidents.

---

Try it yourself:
```bash
cd examples/test_generation_benchmark
make benchmark
```

**Feedback welcome!** Does this match your experience? Any edge cases we should test?

