Property-Based Testing
Overview
Property-based testing (PBT) generates random inputs and verifies that properties hold for all of them. Instead of testing specific examples, you test invariants.
When PBT beats example-based tests:
- Serialization pairs (encode/decode)
- Pure functions with clear contracts
- Validators and normalizers
- Data structure operations
Property Catalog
| Property |
Formula |
When to Use |
| Roundtrip |
decode(encode(x)) == x |
Serialization, conversion pairs |
| Idempotence |
f(f(x)) == f(x) |
Normalization, formatting, sorting |
| Invariant |
Property holds before/after |
Any transformation |
| Commutativity |
f(a, b) == f(b, a) |
Binary/set operations |
| Associativity |
f(f(a,b), c) == f(a, f(b,c)) |
Combining operations |
| Identity |
f(x, identity) == x |
Operations with neutral element |
| Inverse |
f(g(x)) == x |
encrypt/decrypt, compress/decompress |
| Oracle |
new_impl(x) == reference(x) |
Optimization, refactoring |
| Easy to Verify |
is_sorted(sort(x)) |
Complex algorithms |
| No Exception |
No crash on valid input |
Baseline (weakest) |
Strength hierarchy (weakest to strongest):
No Exception -> Type Preservation -> Invariant -> Idempotence -> Roundtrip
Always aim for the strongest property that applies.
Pattern Detection
Use PBT when you see:
| Pattern |
Property |
Priority |
encode/decode, serialize/deserialize |
Roundtrip |
HIGH |
toJSON/fromJSON, pack/unpack |
Roundtrip |
HIGH |
| Pure functions with clear contracts |
Multiple |
HIGH |
normalize, sanitize, canonicalize |
Idempotence |
MEDIUM |
is_valid, validate with normalizers |
Valid after normalize |
MEDIUM |
| Sorting, ordering, comparators |
Idempotence + ordering |
MEDIUM |
| Custom collections (add/remove/get) |
Invariants |
MEDIUM |
| Builder/factory patterns |
Output invariants |
LOW |
When NOT to Use
- Simple CRUD without transformation logic
- UI/presentation logic
- Integration tests requiring complex external setup
- Code with side effects that cannot be isolated
- Prototyping where requirements are fluid
- Tests where specific examples suffice and edge cases are understood
Library Quick Reference
| Language |
Library |
Import |
| Python |
Hypothesis |
from hypothesis import given, strategies as st |
| TypeScript/JS |
fast-check |
import fc from 'fast-check' |
| Rust |
proptest |
use proptest::prelude::* |
| Go |
rapid |
import "pgregory.net/rapid" |
| Java |
jqwik |
@Property annotations |
| Haskell |
QuickCheck |
import Test.QuickCheck |
For library-specific syntax and patterns: Use @ed3d-research-agents:internet-researcher to get current documentation.
Input Strategy Best Practices
Constrain early: Build constraints INTO the strategy, not via assume()
# GOOD
st.integers(min_value=1, max_value=100)
# BAD - high rejection rate
st.integers().filter(lambda x: 1 <= x <= 100)
Size limits: Prevent slow tests
st.lists(st.integers(), max_size=100)
st.text(max_size=1000)
Realistic data: Match real-world constraints
st.integers(min_value=0, max_value=150) # Real ages, not arbitrary ints
Reuse strategies: Define once, use across tests
valid_users = st.builds(User, ...)
@given(valid_users)
def test_one(user): ...
@given(valid_users)
def test_two(user): ...
Settings Guide
# Development (fast feedback)
@settings(max_examples=10)
# CI (thorough)
@settings(max_examples=200)
# Nightly/Release (exhaustive)
@settings(max_examples=1000, deadline=None)
Quality Checklist
Before committing PBT tests:
Red Flags
- Tautological:
assert sorted(xs) == sorted(xs) tests nothing
- Only "no crash": Always look for stronger properties
- Vacuous: Multiple
assume() calls filter out most inputs
- Reimplementation:
assert add(a, b) == a + b if that's how add is implemented
- Missing edge cases: No
@example([]), @example([1]) decorators
- Overly constrained: Many
assume() calls means redesign the strategy
Common Mistakes
| Mistake |
Fix |
| Testing mock behavior |
Test real behavior |
| Reimplementing function in test |
Use algebraic properties |
| Filtering with assume() |
Build constraints into strategy |
| No edge case examples |
Add @example decorators |
| One property only |
Add multiple properties (length, ordering, etc.) |
1---2name: property-based-testing3description: Use when writing tests for serialization, validation, normalization, or pure functions - provides property catalog, pattern detection, and library reference for property-based testing4---56# Property-Based Testing78## Overview910Property-based testing (PBT) generates random inputs and verifies that properties hold for all of them. Instead of testing specific examples, you test invariants.1112**When PBT beats example-based tests:**13- Serialization pairs (encode/decode)14- Pure functions with clear contracts15- Validators and normalizers16- Data structure operations1718## Property Catalog1920| Property | Formula | When to Use |21|----------|---------|-------------|22| **Roundtrip** | `decode(encode(x)) == x` | Serialization, conversion pairs |23| **Idempotence** | `f(f(x)) == f(x)` | Normalization, formatting, sorting |24| **Invariant** | Property holds before/after | Any transformation |25| **Commutativity** | `f(a, b) == f(b, a)` | Binary/set operations |26| **Associativity** | `f(f(a,b), c) == f(a, f(b,c))` | Combining operations |27| **Identity** | `f(x, identity) == x` | Operations with neutral element |28| **Inverse** | `f(g(x)) == x` | encrypt/decrypt, compress/decompress |29| **Oracle** | `new_impl(x) == reference(x)` | Optimization, refactoring |30| **Easy to Verify** | `is_sorted(sort(x))` | Complex algorithms |31| **No Exception** | No crash on valid input | Baseline (weakest) |3233**Strength hierarchy** (weakest to strongest):34```35No Exception -> Type Preservation -> Invariant -> Idempotence -> Roundtrip36```3738Always aim for the strongest property that applies.3940## Pattern Detection4142**Use PBT when you see:**4344| Pattern | Property | Priority |45|---------|----------|----------|46| `encode`/`decode`, `serialize`/`deserialize` | Roundtrip | HIGH |47| `toJSON`/`fromJSON`, `pack`/`unpack` | Roundtrip | HIGH |48| Pure functions with clear contracts | Multiple | HIGH |49| `normalize`, `sanitize`, `canonicalize` | Idempotence | MEDIUM |50| `is_valid`, `validate` with normalizers | Valid after normalize | MEDIUM |51| Sorting, ordering, comparators | Idempotence + ordering | MEDIUM |52| Custom collections (add/remove/get) | Invariants | MEDIUM |53| Builder/factory patterns | Output invariants | LOW |5455## When NOT to Use5657- Simple CRUD without transformation logic58- UI/presentation logic59- Integration tests requiring complex external setup60- Code with side effects that cannot be isolated61- Prototyping where requirements are fluid62- Tests where specific examples suffice and edge cases are understood6364## Library Quick Reference6566| Language | Library | Import |67|----------|---------|--------|68| Python | Hypothesis | `from hypothesis import given, strategies as st` |69| TypeScript/JS | fast-check | `import fc from 'fast-check'` |70| Rust | proptest | `use proptest::prelude::*` |71| Go | rapid | `import "pgregory.net/rapid"` |72| Java | jqwik | `@Property` annotations |73| Haskell | QuickCheck | `import Test.QuickCheck` |7475**For library-specific syntax and patterns:** Use `@ed3d-research-agents:internet-researcher` to get current documentation.7677## Input Strategy Best Practices78791. **Constrain early:** Build constraints INTO the strategy, not via `assume()`80 ```python81 # GOOD82 st.integers(min_value=1, max_value=100)8384 # BAD - high rejection rate85 st.integers().filter(lambda x: 1 <= x <= 100)86 ```87882. **Size limits:** Prevent slow tests89 ```python90 st.lists(st.integers(), max_size=100)91 st.text(max_size=1000)92 ```93943. **Realistic data:** Match real-world constraints95 ```python96 st.integers(min_value=0, max_value=150) # Real ages, not arbitrary ints97 ```98994. **Reuse strategies:** Define once, use across tests100 ```python101 valid_users = st.builds(User, ...)102103 @given(valid_users)104 def test_one(user): ...105106 @given(valid_users)107 def test_two(user): ...108 ```109110## Settings Guide111112```python113# Development (fast feedback)114@settings(max_examples=10)115116# CI (thorough)117@settings(max_examples=200)118119# Nightly/Release (exhaustive)120@settings(max_examples=1000, deadline=None)121```122123## Quality Checklist124125Before committing PBT tests:126127- [ ] Not tautological (assertion doesn't compare same expression)128- [ ] Strong assertion (not just "no crash")129- [ ] Not vacuous (inputs not over-filtered by `assume()`)130- [ ] Edge cases covered with explicit examples (`@example`)131- [ ] No reimplementation of function logic in assertion132- [ ] Strategy constraints are realistic133- [ ] Settings appropriate for context134135## Red Flags136137- **Tautological:** `assert sorted(xs) == sorted(xs)` tests nothing138- **Only "no crash":** Always look for stronger properties139- **Vacuous:** Multiple `assume()` calls filter out most inputs140- **Reimplementation:** `assert add(a, b) == a + b` if that's how add is implemented141- **Missing edge cases:** No `@example([])`, `@example([1])` decorators142- **Overly constrained:** Many `assume()` calls means redesign the strategy143144## Common Mistakes145146| Mistake | Fix |147|---------|-----|148| Testing mock behavior | Test real behavior |149| Reimplementing function in test | Use algebraic properties |150| Filtering with assume() | Build constraints into strategy |151| No edge case examples | Add @example decorators |152| One property only | Add multiple properties (length, ordering, etc.) |