Polished screenshots
A screenshot that will be looked at by a stranger is a piece of design, and it fails in
predictable ways. This skill is a renderer (scripts/shotkit.mjs) plus the judgment about
when to reach for which part of it.
The one rule
A frame cannot rescue a bad capture. Padding, shadow and a mac window around a shot whose text is cut mid-glyph makes the defect more obvious, not less — you have drawn a nice border around the mistake and lit it. Spend the effort in this order:
- Capture clean and deterministic
- Crop at a boundary the UI actually has
- Frame it
- Check it
Most bad screenshots are lost at step 1 or 2, and no amount of step 3 gets them back.
Quick start
# Installed into this project (npx skills add patlf/shotkit):
S="node .claude/skills/shotkit/scripts/shotkit.mjs"
# Installed globally (npx skills add -g patlf/shotkit):
# S="node $HOME/.claude/skills/shotkit/scripts/shotkit.mjs"
# Installed as a Claude Code plugin (/plugin install shotkit@shotkit) — the cache
# path moves on every update, so resolve it rather than hardcoding it:
# S="node ${CLAUDE_PLUGIN_ROOT}/scripts/shotkit.mjs"
# Installed as a dependency of the project:
# S="npx shotkit"
# Capture a live page and frame it in one pass
$S capture http://localhost:3000 --out shots/hero.png \
--selector ".app-shell" --viewport 1440x900 --preset mac --title "Acme"
# Frame a screenshot you already have (or one from ⌘⇧4 / Shottr / CleanShot)
$S frame raw.png --out shots/hero.png --preset clean
# Frame a whole set consistently
$S frame raw/*.png --out shots/ --preset clean --scale 2
# Dissolve into the page instead of sitting on a backdrop — one file, both themes
$S frame raw.png --out shots/hero.png --preset clean --bg transparent --fade bottom --radius 12
# Audit what you have
$S check shots/
# Shrink PNGs a pipeline you already have produced
$S optimize shots/ --max-kb 400
No dependencies. It uses Playwright if the project has it, otherwise a local Chrome/Chromium.
capture needs Playwright; frame and check work with either.
Step 0 — Agree the style before rendering anything
Do not render a set until the style has been agreed out loud. Not one sample shot first, not "defaults now, adjust later" — the defaults are a house style, and a house style nobody chose is what people mean when they say screenshots look generic. Rendering first also makes the correction expensive: by the time it comes there are forty images and a page pointing at them.
Read the project first, but read it to decide what to propose — not to decide whether to ask.
| Read | What it settles |
|---|---|
PRODUCT.md, DESIGN.md |
Register, which is the decision everything else follows from. A product that describes itself as a tool, an instrument, something you work in takes auto or a quiet backdrop and no mesh — a campaign surface behind a working UI reads as a product that is being sold rather than used. A product whose landing page is the product can carry solid or mesh. |
| The existing screenshots | If the repo already ships shots, match them. check a directory to read back the sizes, and look at what preset, backdrop and padding they used. A set half in the old style is worse than a set entirely in either. |
| The design tokens | The CSS custom properties, the theme file, the Tailwind config. A backdrop built from the product's own surface colours belongs to the page in a way a named backdrop never quite does — pass them as linear:<a>,<b> rather than reaching for the table below. The same file gives you the radius the UI actually uses. |
| Where the shot will sit | A README renders on both a light and a dark GitHub page, so docs, bare or a dissolve. A tile or card that already has its own shadow and radius takes bare. A hero sits on a surface the page already owns. |
| Whether the product has themes | If the site or the docs ship light and dark, so do the screenshots — two captures, <picture> + prefers-color-scheme. This one is settled by the repo, not by taste. |
Then ask, in a single round: one message carrying every question, each with the default you
derived, so "all defaults" is a valid one-word answer. In Claude Code that is one
AskUserQuestion call with four questions in it, not four turns of interrogation. If the request
has not already said where the shots will be seen — README, docs, landing page, changelog, store
listing — establish that first, because it decides the rest.
1. Backdrop. Never ask this open-endedly. "What background do you want?" hands the design back to the person who asked you to do it. Offer the derived default, two named candidates that suit the register you just read, and both escape hatches:
auto (default) |
derived from the shot's own edge pixels — it cannot clash, because every colour in it came from the shot |
blur |
the shot itself, scaled up and defocused behind itself; CleanShot's signature |
| a named backdrop | propose two by name out of the quiet / solid / mesh tables below — not the whole list |
| the project's own colours | linear:<a>,<b> straight out of the design tokens |
| an image they supply | a wallpaper, a brand texture, a photo — image:<path>, and say up front that it needs --bg-blur 40 --bg-dim .45 or the UI on top of it is unreadable |
Quote the cost while you are asking: a mesh or solid backdrop is ~200 KB per shot before
--optimize, where auto and the quiet family are ~60.
2. Finish — how the shot meets whatever is behind it.
| framed (default) | padding on four sides, layered shadow, hairline rim |
| dissolve | --bg transparent --fade bottom. The shot fades out to real alpha and blends into whatever the page already is — Linear's look. It carries no backdrop of its own, so it is correct in light and dark at once, and it is the honest answer when the panel is taller than the slot it has to live in. |
| bleed | --pad "30 30 0" — the shot runs off one edge, but the backdrop stays |
| tilt | --tilt -14 — a landing-page hero, once. It dates fast. |
3. Corners. --radius is measured in the source page's own CSS pixels, so it is directly
comparable to the radii inside the shot: 12px cards sitting in a 4px frame read as two products
photographed together. Match the outermost radius the UI itself uses — the window, the panel, the
card the crop runs out to — or go one step above it. The presets ship 8–14. With --chrome mac
the frame is impersonating a macOS window, and those are ~10–12; a 24 there stops reading as a
window and starts reading as a sticker. --radius 0 is a real choice for a technical register,
and never a good accident.
4. Chrome. No chrome (default), macOS window, or browser with a URL bar. Ask it as is "this is an app" part of the message? — chrome is a claim about what the reader is looking at, rather than a decoration. Light, dark or both rides along here if the table above did not already settle it.
Four questions is the ceiling. Do not ask back something the repo already answers — propose it as the default and move on — and if the request already named the style outright ("dark, no chrome, on our purple"), that is the answer: skip the round and go straight to stating it.
Either way, state the choice in one line before rendering anything — "clean preset, auto
backdrop, dissolve at the bottom, 12px corners, no chrome, light and dark" — so it gets corrected
once instead of forty times. Then record it where the project keeps its design decisions, next to
the command that reproduces the set. A style nobody wrote down lasts exactly as long as the
session that chose it.
Replacing an existing set is the same decision, plus one. Old shots were captured at some scale, cropped at some height and framed in some style, and the new ones have to agree with the layout that was built around them, not just with each other. Measure the slot they go into before you capture: the rendered width in the page, the breakpoint where that width changes, and whether anything crops or bleeds them. Then capture at a scale that covers the largest of those.
Step 1 — Capture
capture applies the hygiene that separates a repeatable shot from a lucky one. It waits for
document.fonts.ready (a shot taken mid-swap shows the fallback face, which is the difference
between our type and some type), freezes animations and transitions, hides the caret and
the scrollbars, forces prefers-reduced-motion, and parks the pointer in a corner so no stray
hover state leaks in.
Everything else is on you, and it is the part that matters:
- Real content, not lorem ipsum. Placeholder text in a marketing shot tells the reader the product has nothing to show.
- No relative timestamps. "3 minutes ago" dates the screenshot the moment it is published.
- Nothing personal. Real names, emails, avatars, API keys, internal URLs, "Downloads (847)".
Use
--hidefor banners andblur/redactannotations for anything that survives. - Match the reader's theme. If the page has light and dark modes, capture
--theme lightand--theme darkand serve them with a<picture>+prefers-color-schemesource. A light-mode screenshot on a dark-mode page is a jarring white rectangle. --scale 2. A 1× capture upscaled into a 2× layout goes soft exactly where the type is smallest, which is where the reader is looking.
Useful flags: --selector clips to an element (measured, not guessed), --bleed adds context
around it — but check the result, because bleed catches whatever is next to the element.
--click and --wait-for drive the UI into the state worth showing. --full-page for whole
pages.
Step 2 — Crop where the UI has a seam
This is the step people skip. A crop height picked because it "looked about right" lands wherever it lands, which is usually the middle of a row of text.
Cut at a divider, a section heading, a card edge, the end of a list — somewhere the UI already
has a horizontal line. If nothing is close, change the viewport or the scroll position instead
of accepting a cut through a word. check finds these after the fact, but the cheap fix is to
pick the boundary while you still have the page open.
Deliberately bleeding content off an edge is a legitimate device — it says "this continues" rather than "this is all there is." It only works when it is obviously intentional: the cut runs through a large region or a whole repeated row, never through a single line of type.
Step 3 — Frame
Use the preset chosen in Step 0 for the entire set. Mixing presets is what makes a gallery look assembled from whatever was lying around.
| Preset | Backdrop | Chrome | Radius | Use for |
|---|---|---|---|---|
clean |
derived from the shot | none | 10 | The default. A UI region, a panel, a component. |
mac |
derived | macOS title bar | 12 | A desktop app, or a whole window. |
browser |
slate | URL bar | 12 | Anything where "this is a website" is the point. |
hero |
derived | macOS | 14 | Above the fold. Generous padding, tall shadow. |
docs |
transparent | none | 8 | Inline in a README or docs page. |
flat |
derived | none | 8 | Dense grids, where shadows on every tile become noise. |
bare |
transparent | none | 10 | Rounded corners only, for placing on a page that supplies its own background. |
Every number in that table is a starting point rather than a setting — --bg, --pad, --shadow
and --radius each override the preset they came with. Radius is the one worth overriding most
often, for the reason in Step 0: it is in the same CSS pixels as the UI inside the shot, so a
frame whose corners disagree with the corners already visible in the screenshot reads as two
objects rather than one.
Three things the renderer does that are worth knowing about, because they are the things that usually get done wrong by hand:
- The shadow is stacked, not single. Six layers from a tight contact shadow out to a wide
ambient one. A single
0 20px 60px rgba(0,0,0,.3)produces a uniform grey halo with no contact, which is the clearest tell of a screenshot that was decorated rather than lit. - The source is never resampled. It is placed at exactly
naturalWidth / --scaleCSS px on a whole device pixel.--scaletells the renderer what DPR the source was captured at — it does not resize anything. Get it wrong and you lose the sharpness you captured at 2× for. - The hairline follows the shot, not the backdrop. The rim is drawn 1px inside the shot's own
edge, so what it has to stand out against is the shot: a light UI takes a dark rim and a dark UI
takes a light one, whatever is behind them. It is measured from the shot's pixels at render
time; override with
--hairline light|dark|off.
If the shot is going into a card or tile that already has its own shadow and radius, use bare
or docs. Two shadows on one object looks like a mistake, because it is one.
Backdrops
--bg is the single biggest lever on how a screenshot reads. Run
$S with no arguments for the full list.
auto (default) |
Samples the shot's own edge pixels and derives a backdrop from them in OKLCH, with chroma capped hard. The pair reads as one object under one light, and it can never clash — every colour in it came from the shot. |
blur |
The screenshot itself, painted again behind itself, scaled up and defocused. CleanShot's signature look. Same guarantee as auto, more depth, more bytes. |
quiet — paper slate sand mint blush arctic dusk ink graphite |
Low-chroma two-stop gradients. Use when the screenshot has to be read — a docs page, a feature walkthrough, a comparison table. |
solid — cobalt azure teal emerald lemon amber tangerine crimson fuchsia violet |
One saturated hue, lit from the top-left: the lightness swings ±0.055 and the hue rotates a few degrees warm toward the light. Loud, confident, and still a single colour rather than a gradient effect. The house style for changelogs, release posts and social cards, where a set reads as a set because only the hue changes. |
mesh — aurora ember tide orchid moss sunset noir porcelain |
A base colour with three offset radial bleeds, which is what makes a mesh read as light rather than as a gradient. Use at the top of a landing page, where the image is an object on a surface and the surface is allowed to be one. porcelain is the only light member. |
image:<path> |
A photo or wallpaper. Pair with --bg-blur 40 and --bg-dim .45 — an unblurred photo behind a UI screenshot makes the UI unreadable, every time. |
grid: / dots: |
grid:#dfe3e8,#f8f9fa,26 — line colour, ground, step. Reads as technical rather than promotional; good for developer products. |
transparent |
Real alpha, including through the shadow. For placing on a page that supplies its own ground; with --fade it becomes the dissolve. |
| custom | mesh:<base>,<a>,<b>, linear:<a>,<b>[,deg], radial:<a>,<b>, or any CSS background value. |
The solid family is generated from one OKLCH triple per hue rather than written as CSS.
Interpolating a gradient between two saturated sRGB colours dips in chroma through the middle and
goes muddy; in OKLCH the hue and saturation hold along the ramp, so it reads as one colour lit
unevenly. Each chroma is 95% of the most sRGB can hold at that lightness and hue, solved rather
than eyeballed — asking for more does not get more, because the browser gamut-maps the two stops
by different amounts and bends the ramp. Adding a hue is one line in SOLIDS.
That ceiling is also why the warm ones are light: yellow at full chroma simply is an L≈0.87
colour, and a dark saturated yellow is olive. A light backdrop needs a bigger shadow — a white
panel on lemon is only 1.5:1, so pair it with --shadow deep, which is what separates them.
Two knobs that finish a backdrop: --vignette .25 darkens the corners, which pushes the shot
forward without touching its own contrast. --grain .2 dithers away 8-bit banding on a long
gradient — leave it off unless you can actually see steps, because noise is the one thing PNG
cannot compress and it roughly tripled the file in testing.
Saturated backdrops are expensive. A mesh backdrop costs ~200 KB of PNG where a flat one
costs 60. --optimize quantises to a 256-entry palette with error diffusion, which is what
palette encoders are good at on flat-ish art like this — measured at 201 KB → 70 KB with no
banding that survives looking for it. It is opt-in rather than automatic because a screenshot
with a photograph in it would not fare as well. It uses pngquant or oxipng when they are
installed and falls back to Pillow, and it never quantises an image with real transparency,
which would turn a soft shadow edge into a hard one.
Bleeding a shot off its frame
--pad takes one to four values in CSS order, and a zero edge is how you bleed: --pad "30 30 0"
puts backdrop on three sides and none on the fourth, so the shot runs off the bottom. The corners
on a zero edge square themselves off, because content that continues does not have a rounded
corner.
Reach for it when the shot is much taller than the space it has to live in. Containing the whole
thing means scaling it down until the type stops being readable, which costs the reader more than
the crop does. Match the shadow to the padding: soft reaches 64px and a 30px margin will clip it
square.
Fading an edge, and the dissolve
--fade bottom dissolves an edge instead of slicing it. The mask covers the shadow too, so the
shot does not dissolve while still casting a hard shadow underneath, and the corners on that edge
go square because content that continues does not have a rounded corner. --fade-depth controls
how far in the fade starts (default 0.28).
It does two different jobs, depending on what is behind it:
- Into a backdrop. The honest fix when a panel is genuinely taller than the space it has to live in: a hard cut says "this is all there is" and is wrong, a fade says "this continues" and is true.
- Into nothing.
--bg transparent --fade bottomfades the shot out to real alpha, so it blends into whatever surface the page itself supplies — the look Linear uses on its marketing pages. The output is RGBA, it needs no backdrop decision at all, and it is right on a light page and a dark one at once, which makes it a strong default for a README. Pair it withclean, notbare:barepads to zero, a zero edge squares its corners by design, and that leaves--radiusnothing to round.
Neither is a licence to skip Step 2 — a fade through a single line of type still looks like an accident.
Other things worth reaching for
--tilt -14puts the shot in perspective. Cheap, and it dates fast — use it on a landing page hero and nowhere else. Clamped to ±25°, and the padding grows to hold the corner that swings forward.--tintmultiplies a colour over the shot. Useful for a background layer in a stack, bad for anything the reader is meant to read.--ratio 16:9pads out to an exact proportion for OpenGraph cards and store listings. It only ever adds space, so nothing is cropped.
Step 4 — Check
$S check shots/
Reports, and exits non-zero on:
- an edge that cuts through content — the band just inside each edge is scanned for the sharp light/dark transitions that half-glyphs produce
- images over the file-size budget (
--max-kb, default 400) - images too small for where they will be displayed
- a set whose members are different sizes — the thing that makes a grid jitter
check reads the pixels, so it catches what review misses: whoever chose the crop was looking
at the content, not at the boundary.
optimize <dir> applies the same quantisation as --optimize to files that came from somewhere
else, so an existing capture pipeline gets the size win without being rewritten. Both belong in a
verify script — check exits non-zero on a finding.
Annotations
Arrows, numbered badges, highlight boxes, spotlight dimming, blur, pixelation, solid redaction,
a drawn macOS cursor and a click indicator. Coordinates are in source-image CSS pixels — natural
pixels divided by --scale, the same space the shot is laid out in.
Prefer pixelate over blur for anything that must not be recoverable. A Gaussian blur is a
convolution, and with the font and layout known a short string like a six-digit code can be
brute-forced back out of it; quantising to blocks throws the information away instead.
$S frame raw.png --out out.png --preset clean --annotate anno.json --accent "#e5484d"
The full shape spec, and guidance on how many annotations one image can carry, is in
references/recipes.md. Read it when you are annotating; skip it otherwise.
Where this goes wrong
| Symptom | Cause |
|---|---|
| Soft or fuzzy type | --scale does not match the DPR the source was captured at |
| Shadow cut off square at the canvas edge | --pad is tighter than the shadow reaches; raise it or use a shorter --shadow |
| Shot floats with no separation on a dark backdrop | Dark shadow on dark ground does nothing — the hairline carries it; keep --hairline on |
| Backdrop fights the UI | --bg auto sampled a colourful banner at the edge; name a backdrop instead |
| Photo backdrop makes the UI unreadable | image: without --bg-blur and --bg-dim |
| File is 600 KB | A gradient backdrop, or --grain. Try --optimize, or --format jpg |
Text visible through a blur annotation |
Use pixelate or redact — blur is not redaction |
| Set looks messy in a grid | Mixed presets or mixed sizes — check a directory to confirm |
| Framed shot looks worse than the raw one | The shot is going somewhere that already frames it; use bare |
| Frame's corners disagree with the UI's | --radius left at whatever the preset shipped; match the shot's own outer radius |
--radius appears to do nothing |
A --pad 0 edge or a --fade edge squares its own corners by design — and bare pads to zero on all four |
Sources other than a browser
frame takes any PNG or JPEG, so a macOS ⌘⇧4, a Shottr or CleanShot export, a simulator
recording still, or a Figma export all go through the same pipeline. macOS window captures
already carry a rounded alpha corner and the system's own shadow — capture with
screencapture -o to suppress that shadow, then let frame supply a consistent one, or the two
shadows will stack.
Using this outside Claude Code
scripts/shotkit.mjs is a plain Node CLI with no dependencies and no knowledge of any agent
harness, so nothing here is tied to one tool. For another agent, either point it at this SKILL.md
or copy the ## The one rule and ## Quick start sections into that project's AGENTS.md.