Harness Visual Regression
Screenshot comparison, visual diff detection, and baseline management. Catches unintended CSS regressions, layout shifts, and rendering inconsistencies before they reach production.
When to Use
- Adding visual regression coverage for UI components or pages
- Reviewing visual changes in a pull request before merge
- Updating baselines after intentional design changes
- NOT when testing interactive user flows (use harness-e2e instead)
- NOT when testing component behavior or state (use unit tests or harness-tdd instead)
- NOT when auditing accessibility compliance (use harness-accessibility instead)
Process
Phase 1: DETECT -- Identify UI Components and Rendering Infrastructure
Scan for existing visual test infrastructure. Search for:
- Storybook configuration (
.storybook/, *.stories.tsx, *.stories.ts)
- Visual testing tools (Chromatic config,
percy.yml, Playwright screenshot tests)
- Existing baseline directories (
screenshots/, __image_snapshots__/, visual-tests/)
Catalog testable components. Identify UI surfaces that benefit from visual testing:
- Shared design system components (buttons, forms, modals, navigation)
- Page-level layouts (dashboard, settings, landing page)
- Responsive breakpoints (mobile, tablet, desktop)
- Theme variants (light mode, dark mode)
- States (loading, empty, error, populated)
Determine the rendering strategy. Choose how screenshots are captured:
- Storybook + Chromatic/Percy: best for component libraries with existing stories
- Playwright screenshots: best for full-page and integration-level visual tests
- Jest + jest-image-snapshot: best for lightweight component rendering with jsdom or happy-dom
- Cypress + Percy plugin: best when Cypress is already the E2E framework
Identify viewport and theme matrix. Define the combinations to test:
- Viewports: 375px (mobile), 768px (tablet), 1280px (desktop), 1920px (wide)
- Themes: light, dark (if supported)
- Locales: LTR, RTL (if internationalized)
Report findings. Summarize: components to cover, rendering strategy, viewport matrix, and estimated baseline count.
Phase 2: BASELINE -- Capture Reference Screenshots
Configure the visual testing tool. Set up:
- Screenshot output directory with
.gitkeep or add to .gitignore as appropriate
- Threshold for pixel-level diff tolerance (recommended: 0.1% for component tests, 0.5% for full-page tests)
- Anti-aliasing handling to avoid false positives across different rendering engines
- Font loading: wait for web fonts to load before capture, or use a system font fallback in test mode
Stabilize rendering for deterministic screenshots. Address common sources of non-determinism:
- Disable CSS animations and transitions in test mode
- Mock dates and times to prevent timestamp-based changes
- Replace dynamic content (avatars, ads, user-generated content) with stable placeholders
- Set a fixed random seed for any randomized UI elements
Capture baseline screenshots. For each component in the test matrix:
- Render the component in each viewport and theme combination
- Wait for fonts, images, and lazy-loaded content to fully render
- Capture and save the screenshot to the baseline directory
Review baselines manually. Before committing, visually inspect every baseline screenshot. Confirm:
- The component renders correctly at each viewport
- No rendering artifacts (clipped text, missing icons, broken layouts)
- The screenshot captures the full component without excessive whitespace
Commit baselines. Add baseline screenshots to version control with a descriptive commit message. These baselines become the source of truth for future comparisons.
Phase 3: COMPARE -- Run Visual Diffs Against Baselines
Execute visual comparison. Run the visual test suite, which:
- Renders each component in the same viewport/theme matrix
- Captures a new screenshot for each combination
- Compares the new screenshot against the stored baseline pixel-by-pixel
- Reports differences that exceed the configured threshold
Classify each diff. For every screenshot that exceeds the threshold:
- Intentional change: the diff corresponds to a deliberate design update in the current PR. Mark for baseline update.
- Regression: the diff is unintended and represents a visual bug. Flag for investigation.
- Environmental noise: the diff is caused by rendering differences (sub-pixel anti-aliasing, font hinting). Increase threshold or stabilize rendering.
Investigate regressions. For each regression:
- Identify the CSS or component change that caused the visual shift
- Determine if the change is localized (one component) or cascading (layout shift affecting multiple components)
- Check if the change is caused by a dependency update (CSS framework, icon library)
Update baselines for intentional changes. When a visual change is confirmed intentional:
- Re-capture the baseline for affected screenshots
- Review the updated baseline to confirm it matches the design intent
- Commit updated baselines alongside the code change
Generate a diff report. Produce a summary showing:
- Total screenshots compared
- Screenshots unchanged (passed)
- Screenshots with intentional changes (baselines updated)
- Screenshots with regressions (flagged for fix)
Phase 4: REPORT -- Generate Visual Diff Report and Approval Workflow
Create a visual diff summary for PR review. Include:
- Side-by-side comparison images for each changed screenshot
- Diff overlay highlighting the pixels that changed
- Percentage of pixels changed per screenshot
- Grouped by component and viewport
Integrate with CI pipeline. Configure the visual test suite to:
- Run automatically on every pull request
- Block merge when unapproved visual changes are detected
- Provide a link to the visual diff report in the PR status check
Define the approval workflow. Establish:
- Who can approve visual changes (design team, frontend lead)
- How approvals are recorded (PR comment, Chromatic approval, Percy review)
- What constitutes an "auto-approve" (changes below threshold, test-only files)
Run harness validate. Confirm the project passes all harness checks with visual testing infrastructure in place.
Document the visual testing workflow. Record:
- How to run visual tests locally
- How to update baselines after intentional changes
- How to add visual tests for new components
- Where to find the diff report in CI
Graph Refresh
If a knowledge graph exists at .harness/graph/, refresh it after code changes to keep graph queries accurate:
harness scan [path]
Harness Integration
harness validate -- Run in REPORT phase after visual testing infrastructure is complete. Confirms project health.
harness check-deps -- Run after BASELINE phase to verify visual testing dependencies are in devDependencies.
emit_interaction -- Used to present visual diff results and request human approval for baseline updates.
- Glob -- Used in DETECT phase to find Storybook stories, existing screenshots, and component files.
- Grep -- Used to search for CSS animation properties, dynamic content patterns, and non-deterministic rendering.
Success Criteria
- Every shared design system component has visual baselines for at least mobile and desktop viewports
- Visual diffs are deterministic: running the same code produces the same screenshots every time
- No false positives: environmental noise (font rendering, anti-aliasing) does not trigger diff failures
- Intentional changes are distinguished from regressions in the diff report
- Baselines are committed to version control and updated alongside code changes
- CI blocks merge when unapproved visual changes are detected
harness validate passes with visual testing infrastructure in place
Examples
Example: Playwright Visual Regression for a React App
BASELINE -- Capture component screenshots:
// visual-tests/components.spec.ts
import { test, expect } from '@playwright/test';
const viewports = [
{ name: 'mobile', width: 375, height: 812 },
{ name: 'desktop', width: 1280, height: 720 },
];
for (const viewport of viewports) {
test.describe(`${viewport.name} viewport`, () => {
test.use({ viewport: { width: viewport.width, height: viewport.height } });
test('dashboard renders correctly', async ({ page }) => {
await page.goto('/dashboard');
await page.waitForLoadState('networkidle');
// Disable animations for deterministic screenshots
await page.addStyleTag({
content:
'*, *::before, *::after { animation: none !important; transition: none !important; }',
});
await expect(page).toHaveScreenshot(`dashboard-${viewport.name}.png`, {
maxDiffPixelRatio: 0.005,
fullPage: true,
});
});
test('settings page renders correctly', async ({ page }) => {
await page.goto('/settings');
await page.waitForLoadState('networkidle');
await page.addStyleTag({
content:
'*, *::before, *::after { animation: none !important; transition: none !important; }',
});
await expect(page).toHaveScreenshot(`settings-${viewport.name}.png`, {
maxDiffPixelRatio: 0.005,
});
});
});
}
Example: Storybook with Chromatic
DETECT output:
Storybook: v7.6 detected (.storybook/main.ts)
Stories: 47 stories across 23 components
Chromatic: not configured
Existing baselines: none
Components without stories: Modal, Toast, DatePicker (3 gaps)
BASELINE -- Configure Chromatic and run first build:
// package.json (relevant scripts)
{
"scripts": {
"chromatic": "chromatic --project-token=${CHROMATIC_PROJECT_TOKEN}",
"chromatic:ci": "chromatic --project-token=${CHROMATIC_PROJECT_TOKEN} --exit-zero-on-changes --auto-accept-changes main"
}
}
// .storybook/preview.ts -- stabilize rendering
import { Preview } from '@storybook/react';
const preview: Preview = {
parameters: {
chromatic: {
pauseAnimationAtEnd: true,
viewports: [375, 768, 1280],
},
},
decorators: [
(Story) => (
<div style={{ fontFamily: 'Arial, sans-serif' }}>
<Story />
</div>
),
],
};
export default preview;
Rationalizations to Reject
| Rationalization |
Reality |
| "The baseline screenshots are committed locally but not pushed — CI can capture its own baselines on the first run." |
CI baselines captured without human review become the source of truth for a rendering state nobody verified. Every baseline must be manually inspected before committing. A CI-generated baseline for a broken layout will pass every future comparison until someone notices the visual bug manually. |
| "The diff is only 0.3% of pixels — it's probably just font rendering noise, not a real regression." |
"Probably" is not a classification. Investigate the diff to determine if it is environmental noise or a real change before dismissing it. If it is noise, fix the source (fonts, animations) and lower the threshold. If it is a real change, update the baseline with intent. Skipping investigation means regressions hide behind noise tolerance. |
| "We have 300 components — it's not practical to have baselines for all of them." |
Start with shared design system components (buttons, inputs, modals) and page-level layouts. Partial coverage is better than none, and it is easier to add coverage incrementally than to audit an entire codebase for visual regressions after the fact. Prioritize high-visibility, high-change-frequency surfaces first. |
| "The visual test suite takes 20 minutes — let's just run it manually before releases instead of in CI." |
Manual pre-release checks are skipped under deadline pressure. Visual regression is most valuable on every PR, where the author is still in context and the fix is immediate. A 20-minute suite is a signal to optimize (affected-story detection, parallelization), not to remove CI integration. |
| "The component changed intentionally — I'll just auto-accept the diff without reviewing the new screenshot." |
Auto-accepting without review permanently resets the baseline to whatever was rendered, correct or not. Every baseline update requires a human to look at the new screenshot and confirm it matches design intent. The review is not a formality — it is the entire point of the approval workflow. |
Gates
- No non-deterministic screenshots. If the same code produces different screenshots on consecutive runs, the rendering is not stabilized. Fix animations, dynamic content, and font loading before capturing baselines.
- No uncommitted baselines. Baseline screenshots must be in version control. If baselines exist only on a developer's machine, CI cannot compare against them. Commit baselines with the code that creates them.
- No threshold above 1%. A pixel diff threshold above 1% hides real regressions. If environmental noise requires a higher threshold, fix the noise source (fonts, animations) rather than raising the threshold.
- No visual tests without review workflow. Visual tests that run but whose results are never reviewed provide false confidence. Every visual diff must have a defined approval path.
Escalation
- When screenshots differ between local and CI environments: This is usually caused by different font rendering, display scaling, or browser versions. Standardize by running visual tests in Docker with a fixed browser version and system fonts. Do not try to match local and CI rendering -- pick one as the source of truth.
- When baseline updates flood a PR with hundreds of changed screenshots: Group changes by root cause. If a single CSS variable change cascades to 200 screenshots, approve the root cause and batch-update baselines. Consider whether the cascade indicates a design system architecture issue.
- When dynamic content (user avatars, timestamps, ads) causes false positives: Mock or replace dynamic content in the test environment. Use Storybook args or Playwright route interception to inject stable placeholder content.
- When the visual test suite takes too long (> 15 minutes): Prioritize components by change frequency. Run the full visual suite nightly, and only test changed components on each PR using affected-story detection.
1---2name: harness-visual-regression3description: Harness Visual Regression4---5# Harness Visual Regression67> Screenshot comparison, visual diff detection, and baseline management. Catches unintended CSS regressions, layout shifts, and rendering inconsistencies before they reach production.89## When to Use1011- Adding visual regression coverage for UI components or pages12- Reviewing visual changes in a pull request before merge13- Updating baselines after intentional design changes14- NOT when testing interactive user flows (use harness-e2e instead)15- NOT when testing component behavior or state (use unit tests or harness-tdd instead)16- NOT when auditing accessibility compliance (use harness-accessibility instead)1718## Process1920### Phase 1: DETECT -- Identify UI Components and Rendering Infrastructure21221. **Scan for existing visual test infrastructure.** Search for:23 - Storybook configuration (`.storybook/`, `*.stories.tsx`, `*.stories.ts`)24 - Visual testing tools (Chromatic config, `percy.yml`, Playwright screenshot tests)25 - Existing baseline directories (`screenshots/`, `__image_snapshots__/`, `visual-tests/`)26272. **Catalog testable components.** Identify UI surfaces that benefit from visual testing:28 - Shared design system components (buttons, forms, modals, navigation)29 - Page-level layouts (dashboard, settings, landing page)30 - Responsive breakpoints (mobile, tablet, desktop)31 - Theme variants (light mode, dark mode)32 - States (loading, empty, error, populated)33343. **Determine the rendering strategy.** Choose how screenshots are captured:35 - **Storybook + Chromatic/Percy:** best for component libraries with existing stories36 - **Playwright screenshots:** best for full-page and integration-level visual tests37 - **Jest + jest-image-snapshot:** best for lightweight component rendering with jsdom or happy-dom38 - **Cypress + Percy plugin:** best when Cypress is already the E2E framework39404. **Identify viewport and theme matrix.** Define the combinations to test:41 - Viewports: 375px (mobile), 768px (tablet), 1280px (desktop), 1920px (wide)42 - Themes: light, dark (if supported)43 - Locales: LTR, RTL (if internationalized)44455. **Report findings.** Summarize: components to cover, rendering strategy, viewport matrix, and estimated baseline count.4647### Phase 2: BASELINE -- Capture Reference Screenshots48491. **Configure the visual testing tool.** Set up:50 - Screenshot output directory with `.gitkeep` or add to `.gitignore` as appropriate51 - Threshold for pixel-level diff tolerance (recommended: 0.1% for component tests, 0.5% for full-page tests)52 - Anti-aliasing handling to avoid false positives across different rendering engines53 - Font loading: wait for web fonts to load before capture, or use a system font fallback in test mode54552. **Stabilize rendering for deterministic screenshots.** Address common sources of non-determinism:56 - Disable CSS animations and transitions in test mode57 - Mock dates and times to prevent timestamp-based changes58 - Replace dynamic content (avatars, ads, user-generated content) with stable placeholders59 - Set a fixed random seed for any randomized UI elements60613. **Capture baseline screenshots.** For each component in the test matrix:62 - Render the component in each viewport and theme combination63 - Wait for fonts, images, and lazy-loaded content to fully render64 - Capture and save the screenshot to the baseline directory65664. **Review baselines manually.** Before committing, visually inspect every baseline screenshot. Confirm:67 - The component renders correctly at each viewport68 - No rendering artifacts (clipped text, missing icons, broken layouts)69 - The screenshot captures the full component without excessive whitespace70715. **Commit baselines.** Add baseline screenshots to version control with a descriptive commit message. These baselines become the source of truth for future comparisons.7273### Phase 3: COMPARE -- Run Visual Diffs Against Baselines74751. **Execute visual comparison.** Run the visual test suite, which:76 - Renders each component in the same viewport/theme matrix77 - Captures a new screenshot for each combination78 - Compares the new screenshot against the stored baseline pixel-by-pixel79 - Reports differences that exceed the configured threshold80812. **Classify each diff.** For every screenshot that exceeds the threshold:82 - **Intentional change:** the diff corresponds to a deliberate design update in the current PR. Mark for baseline update.83 - **Regression:** the diff is unintended and represents a visual bug. Flag for investigation.84 - **Environmental noise:** the diff is caused by rendering differences (sub-pixel anti-aliasing, font hinting). Increase threshold or stabilize rendering.85863. **Investigate regressions.** For each regression:87 - Identify the CSS or component change that caused the visual shift88 - Determine if the change is localized (one component) or cascading (layout shift affecting multiple components)89 - Check if the change is caused by a dependency update (CSS framework, icon library)90914. **Update baselines for intentional changes.** When a visual change is confirmed intentional:92 - Re-capture the baseline for affected screenshots93 - Review the updated baseline to confirm it matches the design intent94 - Commit updated baselines alongside the code change95965. **Generate a diff report.** Produce a summary showing:97 - Total screenshots compared98 - Screenshots unchanged (passed)99 - Screenshots with intentional changes (baselines updated)100 - Screenshots with regressions (flagged for fix)101102### Phase 4: REPORT -- Generate Visual Diff Report and Approval Workflow1031041. **Create a visual diff summary for PR review.** Include:105 - Side-by-side comparison images for each changed screenshot106 - Diff overlay highlighting the pixels that changed107 - Percentage of pixels changed per screenshot108 - Grouped by component and viewport1091102. **Integrate with CI pipeline.** Configure the visual test suite to:111 - Run automatically on every pull request112 - Block merge when unapproved visual changes are detected113 - Provide a link to the visual diff report in the PR status check1141153. **Define the approval workflow.** Establish:116 - Who can approve visual changes (design team, frontend lead)117 - How approvals are recorded (PR comment, Chromatic approval, Percy review)118 - What constitutes an "auto-approve" (changes below threshold, test-only files)1191204. **Run `harness validate`.** Confirm the project passes all harness checks with visual testing infrastructure in place.1211225. **Document the visual testing workflow.** Record:123 - How to run visual tests locally124 - How to update baselines after intentional changes125 - How to add visual tests for new components126 - Where to find the diff report in CI127128### Graph Refresh129130If a knowledge graph exists at `.harness/graph/`, refresh it after code changes to keep graph queries accurate:131132```133harness scan [path]134```135136## Harness Integration137138- **`harness validate`** -- Run in REPORT phase after visual testing infrastructure is complete. Confirms project health.139- **`harness check-deps`** -- Run after BASELINE phase to verify visual testing dependencies are in devDependencies.140- **`emit_interaction`** -- Used to present visual diff results and request human approval for baseline updates.141- **Glob** -- Used in DETECT phase to find Storybook stories, existing screenshots, and component files.142- **Grep** -- Used to search for CSS animation properties, dynamic content patterns, and non-deterministic rendering.143144## Success Criteria145146- Every shared design system component has visual baselines for at least mobile and desktop viewports147- Visual diffs are deterministic: running the same code produces the same screenshots every time148- No false positives: environmental noise (font rendering, anti-aliasing) does not trigger diff failures149- Intentional changes are distinguished from regressions in the diff report150- Baselines are committed to version control and updated alongside code changes151- CI blocks merge when unapproved visual changes are detected152- `harness validate` passes with visual testing infrastructure in place153154## Examples155156### Example: Playwright Visual Regression for a React App157158**BASELINE -- Capture component screenshots:**159160```typescript161// visual-tests/components.spec.ts162import { test, expect } from '@playwright/test';163164const viewports = [165 { name: 'mobile', width: 375, height: 812 },166 { name: 'desktop', width: 1280, height: 720 },167];168169for (const viewport of viewports) {170 test.describe(`${viewport.name} viewport`, () => {171 test.use({ viewport: { width: viewport.width, height: viewport.height } });172173 test('dashboard renders correctly', async ({ page }) => {174 await page.goto('/dashboard');175 await page.waitForLoadState('networkidle');176 // Disable animations for deterministic screenshots177 await page.addStyleTag({178 content:179 '*, *::before, *::after { animation: none !important; transition: none !important; }',180 });181 await expect(page).toHaveScreenshot(`dashboard-${viewport.name}.png`, {182 maxDiffPixelRatio: 0.005,183 fullPage: true,184 });185 });186187 test('settings page renders correctly', async ({ page }) => {188 await page.goto('/settings');189 await page.waitForLoadState('networkidle');190 await page.addStyleTag({191 content:192 '*, *::before, *::after { animation: none !important; transition: none !important; }',193 });194 await expect(page).toHaveScreenshot(`settings-${viewport.name}.png`, {195 maxDiffPixelRatio: 0.005,196 });197 });198 });199}200```201202### Example: Storybook with Chromatic203204**DETECT output:**205206```207Storybook: v7.6 detected (.storybook/main.ts)208Stories: 47 stories across 23 components209Chromatic: not configured210Existing baselines: none211Components without stories: Modal, Toast, DatePicker (3 gaps)212```213214**BASELINE -- Configure Chromatic and run first build:**215216```json217// package.json (relevant scripts)218{219 "scripts": {220 "chromatic": "chromatic --project-token=${CHROMATIC_PROJECT_TOKEN}",221 "chromatic:ci": "chromatic --project-token=${CHROMATIC_PROJECT_TOKEN} --exit-zero-on-changes --auto-accept-changes main"222 }223}224```225226```typescript227// .storybook/preview.ts -- stabilize rendering228import { Preview } from '@storybook/react';229230const preview: Preview = {231 parameters: {232 chromatic: {233 pauseAnimationAtEnd: true,234 viewports: [375, 768, 1280],235 },236 },237 decorators: [238 (Story) => (239 <div style={{ fontFamily: 'Arial, sans-serif' }}>240 <Story />241 </div>242 ),243 ],244};245246export default preview;247```248249## Rationalizations to Reject250251| Rationalization | Reality |252| -------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |253| "The baseline screenshots are committed locally but not pushed — CI can capture its own baselines on the first run." | CI baselines captured without human review become the source of truth for a rendering state nobody verified. Every baseline must be manually inspected before committing. A CI-generated baseline for a broken layout will pass every future comparison until someone notices the visual bug manually. |254| "The diff is only 0.3% of pixels — it's probably just font rendering noise, not a real regression." | "Probably" is not a classification. Investigate the diff to determine if it is environmental noise or a real change before dismissing it. If it is noise, fix the source (fonts, animations) and lower the threshold. If it is a real change, update the baseline with intent. Skipping investigation means regressions hide behind noise tolerance. |255| "We have 300 components — it's not practical to have baselines for all of them." | Start with shared design system components (buttons, inputs, modals) and page-level layouts. Partial coverage is better than none, and it is easier to add coverage incrementally than to audit an entire codebase for visual regressions after the fact. Prioritize high-visibility, high-change-frequency surfaces first. |256| "The visual test suite takes 20 minutes — let's just run it manually before releases instead of in CI." | Manual pre-release checks are skipped under deadline pressure. Visual regression is most valuable on every PR, where the author is still in context and the fix is immediate. A 20-minute suite is a signal to optimize (affected-story detection, parallelization), not to remove CI integration. |257| "The component changed intentionally — I'll just auto-accept the diff without reviewing the new screenshot." | Auto-accepting without review permanently resets the baseline to whatever was rendered, correct or not. Every baseline update requires a human to look at the new screenshot and confirm it matches design intent. The review is not a formality — it is the entire point of the approval workflow. |258259## Gates260261- **No non-deterministic screenshots.** If the same code produces different screenshots on consecutive runs, the rendering is not stabilized. Fix animations, dynamic content, and font loading before capturing baselines.262- **No uncommitted baselines.** Baseline screenshots must be in version control. If baselines exist only on a developer's machine, CI cannot compare against them. Commit baselines with the code that creates them.263- **No threshold above 1%.** A pixel diff threshold above 1% hides real regressions. If environmental noise requires a higher threshold, fix the noise source (fonts, animations) rather than raising the threshold.264- **No visual tests without review workflow.** Visual tests that run but whose results are never reviewed provide false confidence. Every visual diff must have a defined approval path.265266## Escalation267268- **When screenshots differ between local and CI environments:** This is usually caused by different font rendering, display scaling, or browser versions. Standardize by running visual tests in Docker with a fixed browser version and system fonts. Do not try to match local and CI rendering -- pick one as the source of truth.269- **When baseline updates flood a PR with hundreds of changed screenshots:** Group changes by root cause. If a single CSS variable change cascades to 200 screenshots, approve the root cause and batch-update baselines. Consider whether the cascade indicates a design system architecture issue.270- **When dynamic content (user avatars, timestamps, ads) causes false positives:** Mock or replace dynamic content in the test environment. Use Storybook args or Playwright route interception to inject stable placeholder content.271- **When the visual test suite takes too long (> 15 minutes):** Prioritize components by change frequency. Run the full visual suite nightly, and only test changed components on each PR using affected-story detection.