# Testing

> How to write good tests. Use when writing tests, improving test coverage, or evaluating test quality. Also invoked by other skills — BDD at RED phase, tdd-review at GREEN gate, refactor at PROTECT phase, and debug. Core test quality knowledge across all workflows.

- Skill: `arcadeai/testing-2` (Agent Skill)
- Install (CLI): `npx skillmds@latest add arcadeai/testing-2`
- Raw SKILL.md: https://api.skillmd.com/api/skills/arcadeai/testing-2/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: ArcadeAI (https://skillmd.com/u/arcadeai)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/arcadeai/testing-2

---


# Writing Good Tests

Tests prove the system behaves correctly. Every test — unit, integration, E2E, eval — must verify **observable behavior**, not implementation details.

**Core Principle:** TEST BEHAVIOR, NOT IMPLEMENTATION

---

## Philosophy: Behavior-Biased Testing

**What this means:** At every test level, assert on what the system _does_ (outputs, side effects, user-visible outcomes) — never on _how_ it does it (internal state, mock call counts, private methods).

**Why:** Tests coupled to implementation break on every refactor. Behavioral tests survive refactoring because behavior doesn't change — only the internals do.

**Scope preference:** When multiple test types can verify a behavior, prefer the highest scope that covers the behavior with acceptable feedback speed. Higher scope = more confidence that the real system works.

```text
Prefer (highest confidence):
  E2E        → proves the user can do the thing
  Integration → proves components work together
  Unit        → proves the algorithm is correct
Fallback (lowest scope):
```

**When to drop to a lower scope:**

- Pure function with many edge cases (20+ combinations) → unit test
- Internal service boundary, no UI involved → integration test
- Algorithm with complex logic (parsing, math, state machines) → unit test
- Only one module's contract matters → integration test

**When to stay at higher scope:**

- User-facing feature or workflow → E2E
- Multiple modules must cooperate → integration or E2E
- "If this breaks, users notice immediately" → E2E

State the test-scope choice and why in one line, unless it's obvious from context.

Use the decision tree, bug detection matrix, and edge cases in this skill.

---

## Iron Laws

Non-negotiable at every test level. Violating these produces tests that pass but catch nothing.

### 1. Test Behavior, Not Implementation

```typescript
// WRONG — tests internal state
expect(component.state.count).toBe(1);
expect(mockFn).toHaveBeenCalledWith('internal-detail');

// RIGHT — tests observable behavior
expect(screen.getByText('Count: 1')).toBeVisible();
expect(result).toEqual({ total: 42 });
```

This applies at EVERY level:

- **Unit:** assert on return values, not on which helpers were called
- **Integration:** assert on API responses, not on which service methods fired
- **E2E:** assert on what the user sees, not on DOM structure
- **Eval:** grade the output quality, not the path the LLM took

### 2. Every Test Needs a Meaningful Assertion

If your assertion would pass for ANY input, it asserts nothing.

```typescript
// WRONG — asserts nothing useful
expect(() => processData(input)).not.toThrow();
expect(result).toBeTruthy();
expect(result).toBeDefined();

// RIGHT — asserts specific behavior
expect(processData(input)).toEqual({ status: 'ok', count: 3 });
expect(result.errors).toHaveLength(0);
```

### 3. Tests Must Fail First

A new test that passes immediately is testing nothing — or testing something that already works (no value added). For new behavior: RED then GREEN. For existing code: if a characterization test fails, you found a bug.

### 4. One Test, One Behavior

If a test name has "and" in it, split it. Each test verifies ONE observable outcome.

```typescript
// WRONG
it('validates input and saves to database', ...);

// RIGHT
it('rejects input missing required field', ...);
it('saves valid input to database', ...);
```

### 5. Tests Must Be Independent

No test depends on another test's side effects. Fresh state per test. Run in any order.

---

## Anti-Patterns

The most common ways AI-generated tests go wrong. Watch for all of them.

| Pattern                     | Problem                                          | Fix                                                                          |
| --------------------------- | ------------------------------------------------ | ---------------------------------------------------------------------------- |
| **Coverage theater**        | High line coverage, tests catch no bugs          | Every test should fail if you break the behavior it guards                   |
| **Mock everything**         | Tests only verify mock wiring, not real behavior | Use real dependencies where practical; mock only external services           |
| **Duplicate tests**         | 20 tests with same structure, different values   | Use parameterized/table-driven tests: `it.each(...)`                         |
| **Happy-path only**         | Misses edge cases where real bugs live           | Always include: empty input, boundary values, error paths                    |
| **Hardcoded magic values**  | Timestamps, IDs, paths break across environments | Use builders, relative values, or factories                                  |
| **Snapshot overuse**        | Large snapshots pass review without scrutiny     | Prefer targeted assertions; snapshots only for large stable structures       |
| **Testing private methods** | Couples tests to implementation                  | Test through the public API                                                  |
| **Exact UI text matching**  | Breaks on copy changes                           | Use regex `/submit/i` or data-testid attributes                              |
| **Bug-locking**             | Tests written against buggy code encode the bug  | Write tests BEFORE implementation (TDD), or verify behavior is correct first |
| **Scope defaulting**        | AI defaults to unit tests for everything         | Ask "what's the highest scope with acceptable feedback speed?" first         |

---

## Behavioral Testing by Type

### Unit Tests — Behavioral

Test the contract (inputs → outputs), not the internals.

```typescript
// Behavioral: asserts on output
it('applies 20% discount for VIP users', () => {
  expect(calculateDiscount(100, { tier: 'VIP' })).toBe(80);
});

// Non-behavioral: asserts on internal call
it('calls applyRate with 0.2', () => {
  calculateDiscount(100, { tier: 'VIP' });
  expect(applyRate).toHaveBeenCalledWith(0.2);
});
```

### Integration Tests — Behavioral

Test that components produce correct combined outcomes with real dependencies.

```typescript
// Behavioral: asserts on combined outcome
it('returns user profile with computed permissions', async () => {
  const response = await api.get('/users/1/profile');
  expect(response.data.permissions).toContain('edit_posts');
});

// Non-behavioral: asserts on which services were called
it('calls UserService then PermissionService', async () => { ... });
```

### E2E Tests — Behavioral

Test what the user can see and do. E2E tests are naturally behavioral — lean into this.

```typescript
// Behavioral: user-visible outcome
test('user creates account and sees dashboard', async ({ page }) => {
  await page.goto('/signup');
  await page.fill('[name="email"]', 'test@example.com');
  await page.fill('[name="password"]', 'secure123');
  await page.click('button:has-text("Sign Up")');
  await expect(page).toHaveURL('/dashboard');
  await expect(page.getByText('Welcome')).toBeVisible();
});
```

### LLM Evals — Behavioral

Grade what the output _achieves_, not the path the model took. Use deterministic assertions first, LLM-as-judge second.

```yaml
# Deterministic assertion (cheap, run every commit)
- type: javascript
  value: JSON.parse(output).intent === 'order_pizza'

# LLM-as-judge (for subjective quality, run on PR/schedule)
- type: llm-rubric
  value: |
    PASS: Correctly identifies pizza order, confirms size and type
    FAIL: Wrong intent, ignores key details, or generic response
```

Eval-specific principles:

- **Grade outcomes, not paths** — the LLM can take any route to the right answer
- **Binary PASS/FAIL over scales** — "3 vs 4" is meaningless; force clarity
- **One dimension per scorer** — don't bundle factuality + tone + completeness
- **Deterministic checks first** — regex, schema validation, required fields before LLM-as-judge

### Wiring Tests — Mock Only the Process Boundary

A suite that fakes _internal_ seams can be fully green while the real wiring is
broken. Name the **process boundary** you mock — the network, filesystem, clock,
or subprocess at the edge of your code — and mock _only_ that. Every entry point
or command that wires modules together gets **≥1 test built from real
collaborators**, faking nothing but that boundary.

```typescript
// ✅ Wiring test: real config → real corpus walk → real orchestrator,
//    mocking only the network boundary (the one external service).
it('files an issue from a real ticket corpus', async () => {
  writeTicket(tmp, 'T1', { status: 'done' });
  const sent = [];
  const result = await syncTracker(tmp, { http: stubHttp(sent) }); // boundary only
  expect(sent).toHaveLength(1);
  expect(result.projected).toBe(1);
});

// ❌ Internal-seam mock: hand-builds the corpus the orchestrator should compute,
//    so a broken config→corpus wiring (e.g. passing cwd where a dir is expected)
//    passes every assertion. This is how a real bug survived 21 scenarios + 61 tests.
it('projects tickets', () => {
  expect(orchestrate({ tickets: [fakeTicket], config: fakeConfig })).toEqual(...);
});
```

If the only way to exercise a code path is through an injected internal value
(`provider: 'none'`, a hand-passed `repoVisibility`) and never through the real
command, that path has **no wiring test** — add one.

---

## Writing Approach

### Match Existing Style

Before writing any test, find existing tests near the code under test. Match their imports, describe/it structure, helpers, and patterns. Don't introduce new conventions into an established test suite.

If no existing tests: use AAA pattern (Arrange-Act-Assert).

### Design Before Writing

List planned tests before coding. For each test, name:

- **What behavior** it verifies (not what code it calls)
- **What the key assertion is** (not "it doesn't throw")
- **Why this test matters** (what bug would slip through without it?)

Aim for: happy path + edge cases + error cases + at least one test the implementation could plausibly get wrong.

### One Test at a Time

Write one test → run it → verify it fails (or passes for characterization) → move to next. Never write all tests at once then run them.

---

## Patterns

### Test Data Builders

```typescript
function buildUser(overrides = {}) {
  return { id: 'test-1', name: 'Test User', role: 'member', ...overrides };
}

it('applies VIP discount', () => {
  const user = buildUser({ role: 'vip' });
  expect(calculateDiscount(user)).toBe(0.2);
});
```

### Async Testing — Never Use Arbitrary Timeouts

```typescript
// WRONG
await sleep(3000);
await page.waitForTimeout(500);

// RIGHT — wait for condition
await expect.poll(() => getStatus()).toBe('ready');
await waitFor(() => expect(element).toBeVisible());
```

### Descriptive Test Names

```typescript
// WRONG
it('works correctly');
it('should handle edge case');

// RIGHT — describes the behavior
it('returns 401 when API key is missing');
it('preserves user input after validation error');
```

---

## Working Around An Upstream Bug? Ship A Tripwire

When you pin a dependency to dodge someone else's bug, the workaround is
temporary by intent and permanent in practice — comments and tickets do not
reach whoever eventually holds the code.

Write a test that asserts the dependency is **still pinned at the last
known-bad version**, with a file header naming what to delete when it fails.
The next upgrade turns CI red in front of the person doing the upgrade.

Assert the pin, never the bug — a test that reproduces the bug goes green when
upstream fixes it, silently, leaving the dead workaround behind.

Warranted when removal depends on someone else's release _and_ the failure
mode is silent. Apply the upstream-workaround tripwire rules in this skill.

---

## Quick Reference

| Need                             | Action                                          |
| -------------------------------- | ----------------------------------------------- |
| Full test type selection guide   | This skill's test-selection sections            |
| Upstream-workaround tripwire     | This skill's tripwire section                   |
| Smoke/live/release lane guidance | The `$safeword:verify` workflow                 |
| LLM eval design guide            | The repository's configured evaluation guidance |
| Test definition template (BDD)   | The `$safeword:bdd` workflow                    |
| Test quality review              | `$safeword:audit`                               |
| Feature-level TDD with scenarios | `$safeword:bdd`                                 |
| Debugging failing tests          | `$safeword:debug`                               |

