# Comparison Techniques

> This document describes various techniques for comparing behavior between original and migrated repositories to detect differences, regressions, and semantic changes.

- Skill: `tools-only/comparison-techniques` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add tools-only/comparison-techniques`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tools-only/comparison-techniques/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: tools-only (https://skillmd.com/u/tools-only)
- Updated: 2026-09-29
- Page: https://skillmd.com/skills/tools-only/comparison-techniques

---

# Comparison Techniques

## Overview

This document describes various techniques for comparing behavior between original and migrated repositories to detect differences, regressions, and semantic changes.

## Test-Based Comparison

### Unit Test Comparison

**Approach**: Run the same unit tests on both versions and compare results.

**Advantages**:
- Fast execution
- Isolated functionality testing
- Easy to identify specific failures

**Limitations**:
- Only tests what's explicitly tested
- May miss integration issues
- Requires good test coverage

**Implementation**:
```bash
# Run tests on both versions
pytest tests/unit/ --json-report --json-report-file=original_unit.json
pytest tests/unit/ --json-report --json-report-file=migrated_unit.json

# Compare results
python scripts/compare_test_results.py original_unit.json migrated_unit.json
```

### Integration Test Comparison

**Approach**: Test interactions between components.

**Advantages**:
- Catches integration issues
- Tests realistic scenarios
- Validates component interactions

**Limitations**:
- Slower execution
- More complex setup
- Harder to isolate failures

### End-to-End Test Comparison

**Approach**: Test complete user workflows.

**Advantages**:
- Tests real user scenarios
- Validates entire system
- Catches UI/UX differences

**Limitations**:
- Very slow
- Brittle tests
- Environment-dependent

## Execution Trace Comparison

### Function Call Tracing

**Approach**: Capture all function calls, arguments, and return values.

**Advantages**:
- Detailed execution visibility
- Catches subtle differences
- Language-agnostic concept

**Limitations**:
- Performance overhead
- Large trace files
- Noise from irrelevant calls

**Implementation**:
```python
import sys
import functools

def trace_calls(func):
    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        print(f"Call: {func.__name__}({args}, {kwargs})")
        result = func(*args, **kwargs)
        print(f"Return: {result}")
        return result
    return wrapper
```

### Control Flow Tracing

**Approach**: Track execution paths through code.

**Advantages**:
- Identifies logic differences
- Shows branching behavior
- Useful for debugging

**Limitations**:
- Requires instrumentation
- May miss data flow issues

### Data Flow Tracing

**Approach**: Track how data flows through the system.

**Advantages**:
- Identifies data transformation issues
- Shows variable dependencies
- Catches state management problems

**Limitations**:
- Complex to implement
- High overhead

## Output Comparison

### Stdout/Stderr Comparison

**Approach**: Compare program output for identical inputs.

**Advantages**:
- Simple to implement
- No code modification needed
- Works for any program

**Limitations**:
- Sensitive to formatting
- May miss internal state differences
- Timing-dependent output

**Implementation**:
```bash
# Capture output from both versions
./original_program < input.txt > original_output.txt 2>&1
./migrated_program < input.txt > migrated_output.txt 2>&1

# Compare
diff original_output.txt migrated_output.txt
```

### File Output Comparison

**Approach**: Compare files generated by both versions.

**Advantages**:
- Validates file generation
- Checks data persistence
- Easy to automate

**Limitations**:
- Only checks file output
- May miss transient state
- Timestamp differences

### API Response Comparison

**Approach**: Compare HTTP responses for identical requests.

**Advantages**:
- Tests API contracts
- Validates serialization
- Checks status codes

**Limitations**:
- Requires running servers
- Network-dependent
- May have timing issues

**Implementation**:
```python
import requests

def compare_api_responses(original_url, migrated_url, endpoint):
    orig_resp = requests.get(f"{original_url}{endpoint}")
    mig_resp = requests.get(f"{migrated_url}{endpoint}")

    assert orig_resp.status_code == mig_resp.status_code
    assert orig_resp.json() == mig_resp.json()
```

## Property-Based Testing

### Invariant Checking

**Approach**: Define properties that should always hold.

**Advantages**:
- Finds edge cases
- Tests many inputs automatically
- Language-agnostic properties

**Limitations**:
- Requires property definition
- May be slow
- False positives possible

**Example Properties**:
- Idempotence: `f(f(x)) == f(x)`
- Commutativity: `f(x, y) == f(y, x)`
- Associativity: `f(f(x, y), z) == f(x, f(y, z))`

**Implementation**:
```python
from hypothesis import given, strategies as st

@given(st.integers())
def test_absolute_value_property(x):
    original_result = original_abs(x)
    migrated_result = migrated_abs(x)

    # Property: abs(x) >= 0
    assert original_result >= 0
    assert migrated_result >= 0

    # Equivalence: both should return same value
    assert original_result == migrated_result
```

### Metamorphic Testing

**Approach**: Test relationships between inputs and outputs.

**Advantages**:
- No oracle needed
- Finds subtle bugs
- Works for complex systems

**Example**:
```python
# Metamorphic relation: sorting twice = sorting once
def test_sort_metamorphic(input_list):
    orig_once = original_sort(input_list)
    orig_twice = original_sort(orig_once)

    mig_once = migrated_sort(input_list)
    mig_twice = migrated_sort(mig_once)

    assert orig_once == orig_twice  # Idempotence
    assert mig_once == mig_twice    # Idempotence
    assert orig_once == mig_once    # Equivalence
```

## Differential Testing

### Fuzzing-Based Comparison

**Approach**: Generate random inputs and compare outputs.

**Advantages**:
- Finds unexpected differences
- No manual test writing
- Good coverage

**Limitations**:
- May generate invalid inputs
- Hard to reproduce failures
- Requires input generation strategy

**Implementation**:
```python
import random

def differential_fuzz(original_func, migrated_func, iterations=1000):
    differences = []

    for i in range(iterations):
        # Generate random input
        input_data = generate_random_input()

        try:
            orig_result = original_func(input_data)
            mig_result = migrated_func(input_data)

            if orig_result != mig_result:
                differences.append({
                    'input': input_data,
                    'original': orig_result,
                    'migrated': mig_result
                })
        except Exception as e:
            # Handle exceptions
            pass

    return differences
```

### Mutation Testing

**Approach**: Mutate inputs and check if both versions respond similarly.

**Advantages**:
- Tests robustness
- Finds edge cases
- Validates error handling

**Limitations**:
- Requires mutation strategies
- May be slow
- Hard to interpret results

## Performance Comparison

### Benchmark Comparison

**Approach**: Measure execution time for identical workloads.

**Advantages**:
- Identifies performance regressions
- Quantifiable metrics
- Easy to automate

**Limitations**:
- Environment-dependent
- May have variance
- Doesn't catch functional issues

**Implementation**:
```python
import time

def benchmark_comparison(original_func, migrated_func, input_data, iterations=100):
    # Benchmark original
    start = time.time()
    for _ in range(iterations):
        original_func(input_data)
    original_time = time.time() - start

    # Benchmark migrated
    start = time.time()
    for _ in range(iterations):
        migrated_func(input_data)
    migrated_time = time.time() - start

    return {
        'original_time': original_time,
        'migrated_time': migrated_time,
        'speedup': original_time / migrated_time
    }
```

### Memory Usage Comparison

**Approach**: Compare memory consumption.

**Advantages**:
- Identifies memory leaks
- Validates optimization
- Important for resource-constrained systems

**Limitations**:
- Platform-dependent
- Garbage collection affects results
- Hard to measure accurately

## State Comparison

### Database State Comparison

**Approach**: Compare database contents after operations.

**Advantages**:
- Validates data persistence
- Checks transactions
- Finds data corruption

**Limitations**:
- Requires database access
- Timing-dependent
- Schema differences

### Memory State Comparison

**Approach**: Compare in-memory state at checkpoints.

**Advantages**:
- Catches state management issues
- Validates object state
- Useful for debugging

**Limitations**:
- Requires instrumentation
- Performance overhead
- Complex comparison logic

## Best Practices

### 1. Use Multiple Techniques

Combine different comparison methods for comprehensive validation:
- Tests for functional correctness
- Traces for execution details
- Outputs for observable behavior
- Performance for efficiency

### 2. Set Appropriate Tolerances

Define acceptable differences:
```python
# Exact match for critical functionality
assert original_result == migrated_result

# Tolerance for floating-point
assert abs(original_result - migrated_result) < 1e-6

# Structural equivalence for JSON
assert json.loads(original) == json.loads(migrated)
```

### 3. Prioritize Differences

Classify differences by severity:
- **Critical**: Core functionality broken
- **High**: Important features affected
- **Medium**: Edge cases or minor issues
- **Low**: Cosmetic or acceptable differences

### 4. Automate Comparison

Integrate into CI/CD pipeline:
```yaml
# .github/workflows/behavior-check.yml
- name: Compare behavior
  run: |
    python scripts/behavior_checker.py \
      --original original-repo \
      --migrated . \
      --fail-on critical
```

### 5. Document Intentional Changes

Mark expected differences:
```json
{
  "expected_differences": [
    {
      "type": "performance_improvement",
      "description": "Optimized sorting algorithm",
      "acceptable": true
    }
  ]
}
```

## Troubleshooting

### High False Positive Rate

**Causes**:
- Timing-dependent behavior
- Non-deterministic code
- Environment differences

**Solutions**:
- Fix random seeds
- Use deterministic algorithms
- Normalize timestamps
- Retry flaky tests

### Missing Real Differences

**Causes**:
- Insufficient test coverage
- Tolerances too loose
- Wrong comparison method

**Solutions**:
- Increase test coverage
- Tighten tolerances
- Use multiple comparison methods
- Add property-based tests

### Performance Overhead

**Causes**:
- Excessive tracing
- Large test suites
- Complex comparisons

**Solutions**:
- Sample traces instead of full capture
- Parallelize test execution
- Use incremental comparison
- Profile and optimize comparison code

