Comparison Techniques
Overview
This document describes various techniques for comparing behavior between original and migrated repositories to detect differences, regressions, and semantic changes.
Test-Based Comparison
Unit Test Comparison
Approach: Run the same unit tests on both versions and compare results.
Advantages:
- Fast execution
- Isolated functionality testing
- Easy to identify specific failures
Limitations:
- Only tests what's explicitly tested
- May miss integration issues
- Requires good test coverage
Implementation:
# Run tests on both versions
pytest tests/unit/ --json-report --json-report-file=original_unit.json
pytest tests/unit/ --json-report --json-report-file=migrated_unit.json
# Compare results
python scripts/compare_test_results.py original_unit.json migrated_unit.json
Integration Test Comparison
Approach: Test interactions between components.
Advantages:
- Catches integration issues
- Tests realistic scenarios
- Validates component interactions
Limitations:
- Slower execution
- More complex setup
- Harder to isolate failures
End-to-End Test Comparison
Approach: Test complete user workflows.
Advantages:
- Tests real user scenarios
- Validates entire system
- Catches UI/UX differences
Limitations:
- Very slow
- Brittle tests
- Environment-dependent
Execution Trace Comparison
Function Call Tracing
Approach: Capture all function calls, arguments, and return values.
Advantages:
- Detailed execution visibility
- Catches subtle differences
- Language-agnostic concept
Limitations:
- Performance overhead
- Large trace files
- Noise from irrelevant calls
Implementation:
import sys
import functools
def trace_calls(func):
@functools.wraps(func)
def wrapper(*args, **kwargs):
print(f"Call: {func.__name__}({args}, {kwargs})")
result = func(*args, **kwargs)
print(f"Return: {result}")
return result
return wrapper
Control Flow Tracing
Approach: Track execution paths through code.
Advantages:
- Identifies logic differences
- Shows branching behavior
- Useful for debugging
Limitations:
- Requires instrumentation
- May miss data flow issues
Data Flow Tracing
Approach: Track how data flows through the system.
Advantages:
- Identifies data transformation issues
- Shows variable dependencies
- Catches state management problems
Limitations:
- Complex to implement
- High overhead
Output Comparison
Stdout/Stderr Comparison
Approach: Compare program output for identical inputs.
Advantages:
- Simple to implement
- No code modification needed
- Works for any program
Limitations:
- Sensitive to formatting
- May miss internal state differences
- Timing-dependent output
Implementation:
# Capture output from both versions
./original_program < input.txt > original_output.txt 2>&1
./migrated_program < input.txt > migrated_output.txt 2>&1
# Compare
diff original_output.txt migrated_output.txt
File Output Comparison
Approach: Compare files generated by both versions.
Advantages:
- Validates file generation
- Checks data persistence
- Easy to automate
Limitations:
- Only checks file output
- May miss transient state
- Timestamp differences
API Response Comparison
Approach: Compare HTTP responses for identical requests.
Advantages:
- Tests API contracts
- Validates serialization
- Checks status codes
Limitations:
- Requires running servers
- Network-dependent
- May have timing issues
Implementation:
import requests
def compare_api_responses(original_url, migrated_url, endpoint):
orig_resp = requests.get(f"{original_url}{endpoint}")
mig_resp = requests.get(f"{migrated_url}{endpoint}")
assert orig_resp.status_code == mig_resp.status_code
assert orig_resp.json() == mig_resp.json()
Property-Based Testing
Invariant Checking
Approach: Define properties that should always hold.
Advantages:
- Finds edge cases
- Tests many inputs automatically
- Language-agnostic properties
Limitations:
- Requires property definition
- May be slow
- False positives possible
Example Properties:
- Idempotence:
f(f(x)) == f(x) - Commutativity:
f(x, y) == f(y, x) - Associativity:
f(f(x, y), z) == f(x, f(y, z))
Implementation:
from hypothesis import given, strategies as st
@given(st.integers())
def test_absolute_value_property(x):
original_result = original_abs(x)
migrated_result = migrated_abs(x)
# Property: abs(x) >= 0
assert original_result >= 0
assert migrated_result >= 0
# Equivalence: both should return same value
assert original_result == migrated_result
Metamorphic Testing
Approach: Test relationships between inputs and outputs.
Advantages:
- No oracle needed
- Finds subtle bugs
- Works for complex systems
Example:
# Metamorphic relation: sorting twice = sorting once
def test_sort_metamorphic(input_list):
orig_once = original_sort(input_list)
orig_twice = original_sort(orig_once)
mig_once = migrated_sort(input_list)
mig_twice = migrated_sort(mig_once)
assert orig_once == orig_twice # Idempotence
assert mig_once == mig_twice # Idempotence
assert orig_once == mig_once # Equivalence
Differential Testing
Fuzzing-Based Comparison
Approach: Generate random inputs and compare outputs.
Advantages:
- Finds unexpected differences
- No manual test writing
- Good coverage
Limitations:
- May generate invalid inputs
- Hard to reproduce failures
- Requires input generation strategy
Implementation:
import random
def differential_fuzz(original_func, migrated_func, iterations=1000):
differences = []
for i in range(iterations):
# Generate random input
input_data = generate_random_input()
try:
orig_result = original_func(input_data)
mig_result = migrated_func(input_data)
if orig_result != mig_result:
differences.append({
'input': input_data,
'original': orig_result,
'migrated': mig_result
})
except Exception as e:
# Handle exceptions
pass
return differences
Mutation Testing
Approach: Mutate inputs and check if both versions respond similarly.
Advantages:
- Tests robustness
- Finds edge cases
- Validates error handling
Limitations:
- Requires mutation strategies
- May be slow
- Hard to interpret results
Performance Comparison
Benchmark Comparison
Approach: Measure execution time for identical workloads.
Advantages:
- Identifies performance regressions
- Quantifiable metrics
- Easy to automate
Limitations:
- Environment-dependent
- May have variance
- Doesn't catch functional issues
Implementation:
import time
def benchmark_comparison(original_func, migrated_func, input_data, iterations=100):
# Benchmark original
start = time.time()
for _ in range(iterations):
original_func(input_data)
original_time = time.time() - start
# Benchmark migrated
start = time.time()
for _ in range(iterations):
migrated_func(input_data)
migrated_time = time.time() - start
return {
'original_time': original_time,
'migrated_time': migrated_time,
'speedup': original_time / migrated_time
}
Memory Usage Comparison
Approach: Compare memory consumption.
Advantages:
- Identifies memory leaks
- Validates optimization
- Important for resource-constrained systems
Limitations:
- Platform-dependent
- Garbage collection affects results
- Hard to measure accurately
State Comparison
Database State Comparison
Approach: Compare database contents after operations.
Advantages:
- Validates data persistence
- Checks transactions
- Finds data corruption
Limitations:
- Requires database access
- Timing-dependent
- Schema differences
Memory State Comparison
Approach: Compare in-memory state at checkpoints.
Advantages:
- Catches state management issues
- Validates object state
- Useful for debugging
Limitations:
- Requires instrumentation
- Performance overhead
- Complex comparison logic
Best Practices
1. Use Multiple Techniques
Combine different comparison methods for comprehensive validation:
- Tests for functional correctness
- Traces for execution details
- Outputs for observable behavior
- Performance for efficiency
2. Set Appropriate Tolerances
Define acceptable differences:
# Exact match for critical functionality
assert original_result == migrated_result
# Tolerance for floating-point
assert abs(original_result - migrated_result) < 1e-6
# Structural equivalence for JSON
assert json.loads(original) == json.loads(migrated)
3. Prioritize Differences
Classify differences by severity:
- Critical: Core functionality broken
- High: Important features affected
- Medium: Edge cases or minor issues
- Low: Cosmetic or acceptable differences
4. Automate Comparison
Integrate into CI/CD pipeline:
# .github/workflows/behavior-check.yml
- name: Compare behavior
run: |
python scripts/behavior_checker.py \
--original original-repo \
--migrated . \
--fail-on critical
5. Document Intentional Changes
Mark expected differences:
{
"expected_differences": [
{
"type": "performance_improvement",
"description": "Optimized sorting algorithm",
"acceptable": true
}
]
}
Troubleshooting
High False Positive Rate
Causes:
- Timing-dependent behavior
- Non-deterministic code
- Environment differences
Solutions:
- Fix random seeds
- Use deterministic algorithms
- Normalize timestamps
- Retry flaky tests
Missing Real Differences
Causes:
- Insufficient test coverage
- Tolerances too loose
- Wrong comparison method
Solutions:
- Increase test coverage
- Tighten tolerances
- Use multiple comparison methods
- Add property-based tests
Performance Overhead
Causes:
- Excessive tracing
- Large test suites
- Complex comparisons
Solutions:
- Sample traces instead of full capture
- Parallelize test execution
- Use incremental comparison
- Profile and optimize comparison code