Harness Test Data
Test factories, fixtures, database seeding, and test data isolation. Establishes patterns for creating realistic, composable test data without coupling tests to specific database states.
When to Use
- Setting up test data factories for a new domain model or entity
- Migrating from shared test fixtures to isolated factory-based test data
- Establishing database seeding for development, staging, or test environments
- NOT when writing the tests themselves (use harness-tdd or harness-e2e instead)
- NOT when designing the database schema (use harness-database instead)
- NOT when testing data pipeline transformations (use harness-data-validation instead)
Process
Phase 1: DETECT -- Identify Models and Existing Patterns
Catalog domain models. Scan for:
- ORM model definitions (Prisma schema, TypeORM entities, Django models, SQLAlchemy models)
- Database migration files that reveal table structures and relationships
- TypeScript/Python type definitions for domain objects
Map model relationships. For each model, identify:
- Required fields and their types
- Foreign key relationships and cardinality (one-to-one, one-to-many, many-to-many)
- Unique constraints, enums, and validation rules
- Default values and auto-generated fields (IDs, timestamps)
Inventory existing test data patterns. Search for:
- Factory files (fishery, factory-bot, factory_boy, rosie)
- Fixture files (JSON, YAML, SQL seed files)
- Inline test data (objects constructed directly in test files)
- Shared test setup files (beforeAll/beforeEach with data creation)
Identify test data problems. Flag:
- Tests that share mutable data (one test's setup affects another)
- Hardcoded IDs or magic values that break when database is reset
- Missing cleanup leading to test pollution
- Overly complex setup that obscures test intent
Report findings. Summarize: models found, existing patterns, and specific problems to address.
Phase 2: DESIGN -- Choose Patterns and Plan Structure
Select the factory pattern. Based on the project's language and conventions:
- TypeScript/JavaScript: fishery (type-safe factories with traits), or a custom builder pattern
- Python: factory_boy (Django/SQLAlchemy integration), or Faker-based builders
- Go: custom builder functions with functional options pattern
- Ruby: factory_bot with traits and transient attributes
Design the factory API. Each factory must support:
- Default creation:
UserFactory.build() returns a valid object with sensible defaults
- Override:
UserFactory.build({ name: 'Custom' }) overrides specific fields
- Traits:
UserFactory.build({ trait: 'admin' }) applies a named set of overrides
- Associations:
ProjectFactory.build() automatically creates a related User owner
- Batch creation:
UserFactory.buildList(5) returns an array
Plan data relationships. Define how factories handle foreign keys:
- Lazy association: create the related record only when needed
- Explicit association: pass an existing related record to avoid duplicates
- Transient attributes: factory parameters that control behavior but are not persisted
Design cleanup strategy. Choose based on test infrastructure:
- Transaction rollback: wrap each test in a transaction (fastest, requires framework support)
- Truncation: truncate tables between tests in dependency order
- Deletion: delete records created by the test using tracked IDs
- Database recreation: drop and recreate the test database per suite (slowest, most isolated)
Define seed data tiers. Separate:
- Reference data: enums, categories, roles -- loaded once, read-only
- Scenario data: realistic datasets for development and demos
- Test data: minimal data created per-test via factories
Phase 3: SCAFFOLD -- Generate Factories and Seed Scripts
Create the factory directory structure. Follow the project's conventions:
tests/factories/ or src/__tests__/factories/ for unit/integration test factories
seeds/ or prisma/seed.ts for database seeding scripts
tests/fixtures/ for static fixture data (JSON, YAML)
Generate a factory for each domain model. Each factory file contains:
- Default attribute definitions using realistic fake data (Faker for names, emails, dates)
- Traits for common variations (active/inactive, admin/member, draft/published)
- Association handling for required relationships
- Type safety: factory output matches the model type definition
Generate a factory index. Create a barrel file that exports all factories for easy importing:
import { UserFactory, ProjectFactory, TaskFactory } from '../factories';
Create seed scripts. Generate:
- Reference data seeder: loads enums, categories, and lookup tables
- Development seeder: creates a realistic dataset for local development
- Test seeder: minimal baseline data required by most tests
Create cleanup utilities. Generate:
- Database cleanup function that truncates or deletes in correct dependency order
- Test lifecycle hooks (beforeEach/afterEach) that integrate cleanup
- Transaction wrapper for test isolation (if supported by the ORM)
Verify factories produce valid data. Write a smoke test that builds one instance of every factory and validates it against the model schema.
Phase 4: VALIDATE -- Verify Isolation, Composability, and Correctness
Test factory defaults. For each factory, verify:
build() returns a valid object that passes model validation
- Required fields are populated with realistic values
- Unique fields generate unique values across multiple builds
- Associations are created when needed and reused when provided
Test factory composition. Verify:
- Traits compose correctly:
UserFactory.build({ traits: ['admin', 'verified'] }) applies both
- Overrides take precedence over defaults and traits
- Batch creation produces distinct records with unique identifiers
Test data isolation. Run the test suite with factory-generated data and verify:
- Tests pass in any execution order (run with randomized order flag)
- No test reads data created by another test
- Cleanup runs correctly between tests (no orphaned records)
Test seed scripts. Verify:
- Seed scripts are idempotent (running twice does not create duplicates)
- Reference data seeder can run against an empty database
- Development seeder creates a realistic, navigable dataset
Run harness validate. Confirm the project passes all harness checks with factory infrastructure in place.
Graph Refresh
If a knowledge graph exists at .harness/graph/, refresh it after code changes to keep graph queries accurate:
harness scan [path]
Harness Integration
harness validate -- Run in VALIDATE phase after all factories and seed scripts are created. Confirms project health.
harness check-deps -- Run after SCAFFOLD phase to ensure test factory dependencies (Faker, fishery) are in devDependencies, not dependencies.
emit_interaction -- Used at design checkpoints to present factory pattern options and cleanup strategy choices to the human.
- Grep -- Used in DETECT phase to find inline test data, hardcoded IDs, and existing factory patterns.
- Glob -- Used to catalog model definitions, migration files, and existing fixture files.
Success Criteria
- Every domain model has a corresponding factory with sensible defaults
- Factories produce valid objects that pass model validation without any overrides
- No test file contains inline object construction for domain models (all use factories)
- Tests pass in any execution order, confirming data isolation
- Seed scripts are idempotent and documented
- Cleanup runs between tests with no orphaned records
harness validate passes with factory infrastructure in place
Examples
Example: Fishery Factories for a TypeScript Project
SCAFFOLD -- User factory with traits:
// tests/factories/user.factory.ts
import { Factory } from 'fishery';
import { faker } from '@faker-js/faker';
import { User } from '../../src/types/user';
export const UserFactory = Factory.define<User>(({ sequence, params, transientParams }) => ({
id: `user-${sequence}`,
email: params.email ?? faker.internet.email(),
name: params.name ?? faker.person.fullName(),
role: params.role ?? 'member',
status: params.status ?? 'active',
createdAt: params.createdAt ?? faker.date.past(),
updatedAt: new Date(),
}));
// Traits
export const AdminUser = UserFactory.params({ role: 'admin' });
export const InactiveUser = UserFactory.params({ status: 'inactive' });
SCAFFOLD -- Project factory with association:
// tests/factories/project.factory.ts
import { Factory } from 'fishery';
import { faker } from '@faker-js/faker';
import { Project } from '../../src/types/project';
import { UserFactory } from './user.factory';
export const ProjectFactory = Factory.define<Project>(({ sequence, associations }) => ({
id: `project-${sequence}`,
name: faker.commerce.productName(),
description: faker.lorem.sentence(),
owner: associations.owner ?? UserFactory.build(),
ownerId: associations.owner?.id ?? `user-${sequence}`,
status: 'active',
createdAt: faker.date.past(),
}));
Example: factory_boy for a Django Project
SCAFFOLD -- Django model factories:
# tests/factories.py
import factory
from factory.django import DjangoModelFactory
from faker import Faker
from myapp.models import User, Organization, Project
fake = Faker()
class OrganizationFactory(DjangoModelFactory):
class Meta:
model = Organization
name = factory.LazyFunction(lambda: fake.company())
slug = factory.LazyAttribute(lambda o: o.name.lower().replace(' ', '-'))
class UserFactory(DjangoModelFactory):
class Meta:
model = User
email = factory.LazyFunction(lambda: fake.unique.email())
name = factory.LazyFunction(lambda: fake.name())
organization = factory.SubFactory(OrganizationFactory)
class Params:
admin = factory.Trait(is_staff=True, is_superuser=True)
class ProjectFactory(DjangoModelFactory):
class Meta:
model = Project
name = factory.LazyFunction(lambda: fake.catch_phrase())
owner = factory.SubFactory(UserFactory)
organization = factory.LazyAttribute(lambda p: p.owner.organization)
Usage in tests:
def test_project_belongs_to_owner_organization():
project = ProjectFactory.create()
assert project.organization == project.owner.organization
def test_admin_can_delete_any_project():
admin = UserFactory.create(admin=True)
project = ProjectFactory.create()
assert admin.has_perm('delete_project', project)
Rationalizations to Reject
| Rationalization |
Reality |
"The test only needs one user — I'll just hardcode userId: 1 rather than building a factory." |
Hardcoded IDs cause silent test failures when the database is reset, when tests run in parallel, or when another test creates a conflicting record. Factories with sequences or UUIDs exist precisely to avoid this class of failure. One hardcoded ID is how fragile test suites start. |
| "The factory produces objects that are close enough to valid — the test just needs to override two fields anyway." |
A factory that requires overrides to produce a valid object has wrong defaults. The factory's zero-override output must pass model validation. If it does not, the factory is documenting the wrong defaults and tests that rely on overrides will break when the model changes. |
| "Cleanup is handled by rolling back the test transaction — I don't need explicit teardown." |
Transaction rollback works until it does not: tests that span multiple connections, tests that call external APIs, or tests that write to a queue or file system all escape the transaction. Explicit cleanup is the only strategy that covers all cases, including tests that use features you have not built yet. |
| "We can share the same seeded dataset across all integration tests to avoid the overhead of per-test factories." |
Shared mutable data means test A's side effect becomes test B's precondition. When tests fail intermittently based on execution order, the root cause is always shared state. The overhead of per-test factory creation is small compared to the cost of debugging order-dependent failures. |
| "The model only has three fields — writing a factory is more overhead than just constructing the object inline." |
Today's three-field model becomes tomorrow's ten-field model with required foreign keys. Inline construction scales linearly with model complexity. A factory written once absorbs all future field additions in one place. The overhead argument inverts as the codebase grows. |
Gates
- No hardcoded IDs in factories. Factories must generate unique IDs per instance. Hardcoded IDs cause collision failures when tests run in parallel. Use sequences or UUIDs.
- No production data in test fixtures. Test data must be synthetic. If a fixture file contains real customer names, emails, or PII, it must be replaced with Faker-generated data before merging.
- Factories must produce valid objects. A factory
build() with zero overrides must return an object that passes model validation. If it requires manual overrides to be valid, the defaults are wrong.
- Cleanup must be explicit. Do not rely on test framework teardown happening "eventually." Every test or test suite that creates database records must have an explicit cleanup step that runs even when tests fail.
Escalation
- When models have circular dependencies (User has Projects, Project has Owner User): Use lazy evaluation or two-pass creation. Create the User first without Projects, create the Project with the User, then optionally update the User. Document the pattern in the factory file.
- When the database schema is too complex for factories (50+ models): Prioritize factories for the models that appear most frequently in tests. Use a tiered approach: core factories first, then add factories for secondary models as tests demand them.
- When seed data conflicts with migration state: Seed scripts must be updated whenever migrations change the schema. If seeds fail after a migration, fix the seeds immediately -- do not skip seeding.
- When test isolation requires database-level features (row-level security, multi-tenancy): Factory cleanup may need tenant-aware truncation. Escalate to ensure the cleanup strategy respects the application's multi-tenancy model.
1---2name: harness-test-data3description: Harness Test Data4---5# Harness Test Data67> Test factories, fixtures, database seeding, and test data isolation. Establishes patterns for creating realistic, composable test data without coupling tests to specific database states.89## When to Use1011- Setting up test data factories for a new domain model or entity12- Migrating from shared test fixtures to isolated factory-based test data13- Establishing database seeding for development, staging, or test environments14- NOT when writing the tests themselves (use harness-tdd or harness-e2e instead)15- NOT when designing the database schema (use harness-database instead)16- NOT when testing data pipeline transformations (use harness-data-validation instead)1718## Process1920### Phase 1: DETECT -- Identify Models and Existing Patterns21221. **Catalog domain models.** Scan for:23 - ORM model definitions (Prisma schema, TypeORM entities, Django models, SQLAlchemy models)24 - Database migration files that reveal table structures and relationships25 - TypeScript/Python type definitions for domain objects26272. **Map model relationships.** For each model, identify:28 - Required fields and their types29 - Foreign key relationships and cardinality (one-to-one, one-to-many, many-to-many)30 - Unique constraints, enums, and validation rules31 - Default values and auto-generated fields (IDs, timestamps)32333. **Inventory existing test data patterns.** Search for:34 - Factory files (fishery, factory-bot, factory_boy, rosie)35 - Fixture files (JSON, YAML, SQL seed files)36 - Inline test data (objects constructed directly in test files)37 - Shared test setup files (beforeAll/beforeEach with data creation)38394. **Identify test data problems.** Flag:40 - Tests that share mutable data (one test's setup affects another)41 - Hardcoded IDs or magic values that break when database is reset42 - Missing cleanup leading to test pollution43 - Overly complex setup that obscures test intent44455. **Report findings.** Summarize: models found, existing patterns, and specific problems to address.4647### Phase 2: DESIGN -- Choose Patterns and Plan Structure48491. **Select the factory pattern.** Based on the project's language and conventions:50 - **TypeScript/JavaScript:** fishery (type-safe factories with traits), or a custom builder pattern51 - **Python:** factory_boy (Django/SQLAlchemy integration), or Faker-based builders52 - **Go:** custom builder functions with functional options pattern53 - **Ruby:** factory_bot with traits and transient attributes54552. **Design the factory API.** Each factory must support:56 - Default creation: `UserFactory.build()` returns a valid object with sensible defaults57 - Override: `UserFactory.build({ name: 'Custom' })` overrides specific fields58 - Traits: `UserFactory.build({ trait: 'admin' })` applies a named set of overrides59 - Associations: `ProjectFactory.build()` automatically creates a related `User` owner60 - Batch creation: `UserFactory.buildList(5)` returns an array61623. **Plan data relationships.** Define how factories handle foreign keys:63 - Lazy association: create the related record only when needed64 - Explicit association: pass an existing related record to avoid duplicates65 - Transient attributes: factory parameters that control behavior but are not persisted66674. **Design cleanup strategy.** Choose based on test infrastructure:68 - **Transaction rollback:** wrap each test in a transaction (fastest, requires framework support)69 - **Truncation:** truncate tables between tests in dependency order70 - **Deletion:** delete records created by the test using tracked IDs71 - **Database recreation:** drop and recreate the test database per suite (slowest, most isolated)72735. **Define seed data tiers.** Separate:74 - Reference data: enums, categories, roles -- loaded once, read-only75 - Scenario data: realistic datasets for development and demos76 - Test data: minimal data created per-test via factories7778### Phase 3: SCAFFOLD -- Generate Factories and Seed Scripts79801. **Create the factory directory structure.** Follow the project's conventions:81 - `tests/factories/` or `src/__tests__/factories/` for unit/integration test factories82 - `seeds/` or `prisma/seed.ts` for database seeding scripts83 - `tests/fixtures/` for static fixture data (JSON, YAML)84852. **Generate a factory for each domain model.** Each factory file contains:86 - Default attribute definitions using realistic fake data (Faker for names, emails, dates)87 - Traits for common variations (active/inactive, admin/member, draft/published)88 - Association handling for required relationships89 - Type safety: factory output matches the model type definition90913. **Generate a factory index.** Create a barrel file that exports all factories for easy importing:9293 ```94 import { UserFactory, ProjectFactory, TaskFactory } from '../factories';95 ```96974. **Create seed scripts.** Generate:98 - Reference data seeder: loads enums, categories, and lookup tables99 - Development seeder: creates a realistic dataset for local development100 - Test seeder: minimal baseline data required by most tests1011025. **Create cleanup utilities.** Generate:103 - Database cleanup function that truncates or deletes in correct dependency order104 - Test lifecycle hooks (beforeEach/afterEach) that integrate cleanup105 - Transaction wrapper for test isolation (if supported by the ORM)1061076. **Verify factories produce valid data.** Write a smoke test that builds one instance of every factory and validates it against the model schema.108109### Phase 4: VALIDATE -- Verify Isolation, Composability, and Correctness1101111. **Test factory defaults.** For each factory, verify:112 - `build()` returns a valid object that passes model validation113 - Required fields are populated with realistic values114 - Unique fields generate unique values across multiple builds115 - Associations are created when needed and reused when provided1161172. **Test factory composition.** Verify:118 - Traits compose correctly: `UserFactory.build({ traits: ['admin', 'verified'] })` applies both119 - Overrides take precedence over defaults and traits120 - Batch creation produces distinct records with unique identifiers1211223. **Test data isolation.** Run the test suite with factory-generated data and verify:123 - Tests pass in any execution order (run with randomized order flag)124 - No test reads data created by another test125 - Cleanup runs correctly between tests (no orphaned records)1261274. **Test seed scripts.** Verify:128 - Seed scripts are idempotent (running twice does not create duplicates)129 - Reference data seeder can run against an empty database130 - Development seeder creates a realistic, navigable dataset1311325. **Run `harness validate`.** Confirm the project passes all harness checks with factory infrastructure in place.133134### Graph Refresh135136If a knowledge graph exists at `.harness/graph/`, refresh it after code changes to keep graph queries accurate:137138```139harness scan [path]140```141142## Harness Integration143144- **`harness validate`** -- Run in VALIDATE phase after all factories and seed scripts are created. Confirms project health.145- **`harness check-deps`** -- Run after SCAFFOLD phase to ensure test factory dependencies (Faker, fishery) are in devDependencies, not dependencies.146- **`emit_interaction`** -- Used at design checkpoints to present factory pattern options and cleanup strategy choices to the human.147- **Grep** -- Used in DETECT phase to find inline test data, hardcoded IDs, and existing factory patterns.148- **Glob** -- Used to catalog model definitions, migration files, and existing fixture files.149150## Success Criteria151152- Every domain model has a corresponding factory with sensible defaults153- Factories produce valid objects that pass model validation without any overrides154- No test file contains inline object construction for domain models (all use factories)155- Tests pass in any execution order, confirming data isolation156- Seed scripts are idempotent and documented157- Cleanup runs between tests with no orphaned records158- `harness validate` passes with factory infrastructure in place159160## Examples161162### Example: Fishery Factories for a TypeScript Project163164**SCAFFOLD -- User factory with traits:**165166```typescript167// tests/factories/user.factory.ts168import { Factory } from 'fishery';169import { faker } from '@faker-js/faker';170import { User } from '../../src/types/user';171172export const UserFactory = Factory.define<User>(({ sequence, params, transientParams }) => ({173 id: `user-${sequence}`,174 email: params.email ?? faker.internet.email(),175 name: params.name ?? faker.person.fullName(),176 role: params.role ?? 'member',177 status: params.status ?? 'active',178 createdAt: params.createdAt ?? faker.date.past(),179 updatedAt: new Date(),180}));181182// Traits183export const AdminUser = UserFactory.params({ role: 'admin' });184export const InactiveUser = UserFactory.params({ status: 'inactive' });185```186187**SCAFFOLD -- Project factory with association:**188189```typescript190// tests/factories/project.factory.ts191import { Factory } from 'fishery';192import { faker } from '@faker-js/faker';193import { Project } from '../../src/types/project';194import { UserFactory } from './user.factory';195196export const ProjectFactory = Factory.define<Project>(({ sequence, associations }) => ({197 id: `project-${sequence}`,198 name: faker.commerce.productName(),199 description: faker.lorem.sentence(),200 owner: associations.owner ?? UserFactory.build(),201 ownerId: associations.owner?.id ?? `user-${sequence}`,202 status: 'active',203 createdAt: faker.date.past(),204}));205```206207### Example: factory_boy for a Django Project208209**SCAFFOLD -- Django model factories:**210211```python212# tests/factories.py213import factory214from factory.django import DjangoModelFactory215from faker import Faker216from myapp.models import User, Organization, Project217218fake = Faker()219220class OrganizationFactory(DjangoModelFactory):221 class Meta:222 model = Organization223224 name = factory.LazyFunction(lambda: fake.company())225 slug = factory.LazyAttribute(lambda o: o.name.lower().replace(' ', '-'))226227class UserFactory(DjangoModelFactory):228 class Meta:229 model = User230231 email = factory.LazyFunction(lambda: fake.unique.email())232 name = factory.LazyFunction(lambda: fake.name())233 organization = factory.SubFactory(OrganizationFactory)234235 class Params:236 admin = factory.Trait(is_staff=True, is_superuser=True)237238class ProjectFactory(DjangoModelFactory):239 class Meta:240 model = Project241242 name = factory.LazyFunction(lambda: fake.catch_phrase())243 owner = factory.SubFactory(UserFactory)244 organization = factory.LazyAttribute(lambda p: p.owner.organization)245```246247**Usage in tests:**248249```python250def test_project_belongs_to_owner_organization():251 project = ProjectFactory.create()252 assert project.organization == project.owner.organization253254def test_admin_can_delete_any_project():255 admin = UserFactory.create(admin=True)256 project = ProjectFactory.create()257 assert admin.has_perm('delete_project', project)258```259260## Rationalizations to Reject261262| Rationalization | Reality |263| ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |264| "The test only needs one user — I'll just hardcode `userId: 1` rather than building a factory." | Hardcoded IDs cause silent test failures when the database is reset, when tests run in parallel, or when another test creates a conflicting record. Factories with sequences or UUIDs exist precisely to avoid this class of failure. One hardcoded ID is how fragile test suites start. |265| "The factory produces objects that are close enough to valid — the test just needs to override two fields anyway." | A factory that requires overrides to produce a valid object has wrong defaults. The factory's zero-override output must pass model validation. If it does not, the factory is documenting the wrong defaults and tests that rely on overrides will break when the model changes. |266| "Cleanup is handled by rolling back the test transaction — I don't need explicit teardown." | Transaction rollback works until it does not: tests that span multiple connections, tests that call external APIs, or tests that write to a queue or file system all escape the transaction. Explicit cleanup is the only strategy that covers all cases, including tests that use features you have not built yet. |267| "We can share the same seeded dataset across all integration tests to avoid the overhead of per-test factories." | Shared mutable data means test A's side effect becomes test B's precondition. When tests fail intermittently based on execution order, the root cause is always shared state. The overhead of per-test factory creation is small compared to the cost of debugging order-dependent failures. |268| "The model only has three fields — writing a factory is more overhead than just constructing the object inline." | Today's three-field model becomes tomorrow's ten-field model with required foreign keys. Inline construction scales linearly with model complexity. A factory written once absorbs all future field additions in one place. The overhead argument inverts as the codebase grows. |269270## Gates271272- **No hardcoded IDs in factories.** Factories must generate unique IDs per instance. Hardcoded IDs cause collision failures when tests run in parallel. Use sequences or UUIDs.273- **No production data in test fixtures.** Test data must be synthetic. If a fixture file contains real customer names, emails, or PII, it must be replaced with Faker-generated data before merging.274- **Factories must produce valid objects.** A factory `build()` with zero overrides must return an object that passes model validation. If it requires manual overrides to be valid, the defaults are wrong.275- **Cleanup must be explicit.** Do not rely on test framework teardown happening "eventually." Every test or test suite that creates database records must have an explicit cleanup step that runs even when tests fail.276277## Escalation278279- **When models have circular dependencies (User has Projects, Project has Owner User):** Use lazy evaluation or two-pass creation. Create the User first without Projects, create the Project with the User, then optionally update the User. Document the pattern in the factory file.280- **When the database schema is too complex for factories (50+ models):** Prioritize factories for the models that appear most frequently in tests. Use a tiered approach: core factories first, then add factories for secondary models as tests demand them.281- **When seed data conflicts with migration state:** Seed scripts must be updated whenever migrations change the schema. If seeds fail after a migration, fix the seeds immediately -- do not skip seeding.282- **When test isolation requires database-level features (row-level security, multi-tenancy):** Factory cleanup may need tenant-aware truncation. Escalate to ensure the cleanup strategy respects the application's multi-tenancy model.