# Dataflow Diagram

> Draw a dataflow or pipeline diagram for a program that already exists in code — a DAG, a processing chain, a stream topology, a service graph. Use when asked to "diagram the pipeline", "show the dataflow", "chart the DAG", "visualize how the data moves", "make a graph of the program", or when a design doc or PR needs a picture of a running system. Symptoms that this skill applies: a hand-maintained diagram that no longer matches the code, a chart whose legend disagrees with the chart, a glyph reused for two different kinds of thing, a diagram that silently omits a node nobody classified, a legend placed by eye that overlaps an edge, or a picture that is a screenshot rather than a source file.

- Skill: `iopsystems/dataflow-diagram-2` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add iopsystems/dataflow-diagram-2`
- Raw SKILL.md: https://api.skillmd.com/api/skills/iopsystems/dataflow-diagram-2/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- Author: iopsystems (https://skillmd.com/u/iopsystems)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/iopsystems/dataflow-diagram-2

---


# Dataflow diagram

A diagram of a running program is a claim about that program. Every claim it
makes must be one the code can be held to, and every claim the code makes
must appear. Both halves fail quietly, which is why this is a generator
problem before it is a drawing problem.

## What this is, and what it is not yet

The **principles** — derive rather than draw, fail loudly on the
unclassified, treat every channel as a claim, check placement rather than
eyeball it — come from failures that were observed, diagnosed, and fixed.
They should survive contact with a different domain.

The **conventions** — the specific palette, rounded-versus-square, the edge
styles, the frame grammar — are defaults distilled from one project's
diagrams. They are internally consistent and worth adopting wholesale
rather than assembling from scratch, but they have not been tried against a
dataflow shaped differently.

So **say which conventions you are adopting and ask where they fight the
domain**:

- Does compute-versus-data actually partition this system's nodes, or is
  there a third kind that fits neither?
- Does the palette survive the medium it will be read in?
- Is there a distinction the reader needs that no channel here carries?

When a default does not fit, **that override is the finding this skill is
missing.** Capture it with its reason and bring it back here: a convention
rejected for a stated reason is worth more than one nobody has tested, and
this skill has more of the second kind than it should.

Two such contacts have now happened. The `architecture-diagram` skill
carries these principles into system-architecture charts — a build-time
structure chart plus runtime thread/request charts — and records where the
defaults fought (arrows dropped from the structure chart, provenance
demoted from a channel, geometry emitted directly). Use it for that duo;
this skill remains the home for dataflow and pipeline charts of a single
running program. The second contact was a request-lifecycle chart of an
io_uring runtime library, whose reader found the defects the sections on
placement and on the perceptual gate below now name.

This skill ships in two forms with one structure: the skill itself is
complete working defaults — enough for a single-use chart in a repository
you cannot modify, or a figure for a talk, where the generator script
travels with the artifact and the chart is stamped with the commit it
describes, a dated snapshot. The `dataflow-diagram-skill` template adds a
charter: a *delta* recording the bindings no default can supply, plus any
deviation with its reason. A convention absent from the charter means the
default applies.

## Derive, never draw

The nodes and edges come from the program's own structures — its topic
registry, its step declarations, its wiring function — read at generation
time. Not from a list beside them.

A hand-maintained diagram is a second source of truth. It is correct on the
day it is drawn and wrong on the first refactor, and nothing reports the
divergence: the picture keeps rendering. Deriving it means a renamed topic
either appears renamed or breaks the generator, and both are better than a
drawing that quietly describes last quarter's design. It is also what makes
the chart cheap to keep: a generated chart is regenerated, a drawn one
renegotiated.

## Fail loudly on anything unclassified

A generator maps program elements onto visual properties through tables:
this operation is that role, this topic is that category. **An element no
table covers must stop the run, not be skipped.**

Silent omission is the worst failure available here, because the output
still looks complete: a reader cannot tell a pipeline with four stages from
a five-stage pipeline whose fifth stage nobody classified.

The same applies in reverse: a table entry naming an element the program no
longer has is a stale classification, and should fail just as loudly.

## Every visual property is a claim

Shape, fill, line style, decoration, position: each is a channel, and a
reader who sees two things drawn differently will look for the difference.
A channel spent on ornament teaches a distinction that does not exist.

Choose the channel from what the distinction *is*:

- **Shape** — what kind of thing it is. A container of history, a unit of
  computation, a fixed value. Shape is the strongest channel; spend it on
  the taxonomy the reader most needs.
- **Fill** — which family within a kind: the role an operation plays, the
  boundary a topic sits on.
- **Line style** — conditionality. Solid for always, dashed for sometimes,
  dotted for a dependency that is not dataflow.
- **Decoration** — a property layered on a kind, such as whether a run
  records it.

A glyph asserts its own structure. A six-celled queue asserts a history a
reader can look back over; using it for a parameter that is written once
asserts something untrue about the parameter. Match the glyph to the claim
or pick another glyph.

**Absence must be legible.** When a property is encoded, the elements
lacking it should read as lacking it rather than as unmarked. The one input
in a chart with no recording mark is a finding, and it is only a finding
because the marked ones are visibly marked.

**Never encode a value you had to guess.** An unknown that renders as a
plausible mark is worse than an unknown that renders as a gap.

## Draw the state, do not bury it

Before a shape can distinguish compute from data, both have to be nodes. The
failure comes first as a stateful component folded into the label of the thing
that reads or writes it: "recorded streams → rings", "writes to the queue",
"the cache".

A component that holds data across invocations — a ring, a queue, a cache, a
store, a cursor — is not an implementation detail of its reader. Draw it. The
properties that make it worth drawing belong to nothing else in that box: what
it holds, how much, what happens when it fills, who else writes it, and whether
a reader can fall behind and lose records.

Two facts surface the moment it becomes a node, and neither is visible while it
is a phrase:

- **Where paths converge.** Two producers meeting at one buffer is an
  architectural claim a label cannot make. In one case that fact moved the
  boundary of a chart's central claim: two paths drawn as separate pipelines
  turned out to meet at the ring, so what was "identical from the program down"
  was really identical from the buffer down, one node earlier.
- **Where records can be lost.** Loss happens at the thing with a capacity. A
  chart that draws only processes has nowhere to put the loss, so it does not
  show it, and a reader concludes the path is lossless.

The test: **if a box names both a process and something that outlives it, split
it.** A file being read and the buffer it is read into are two nodes, and only
one of them can lap.

## Compute and data must not share a silhouette

The first thing a reader resolves is *what kind of thing is this*, and they
resolve it from outline before they resolve fill or text. Give the two
halves of a dataflow different corners:

- **Compute — rounded.** Operations, steps, stages, transforms: anything
  that runs. `shape=box, style="rounded,filled"`.
- **Data — square.** Queues, topics, buffers, parameters, stores: anything
  that holds. Straight corners, whatever the shape otherwise is.

Rounded-versus-square is deliberately quiet. It should register as a
texture across the whole chart rather than as a label on each node, which
is what lets fill stay free for the distinction *within* each half.

Then split data by what it promises:

- **A history** — a segmented glyph, cells in a row, reading as a queue
  with depth. It asserts that a reader can look back over a window.
- **One value** — a solid glyph with no segmentation, such as a dog-eared
  page (`shape=note`) for a parameter fixed before the run.

Segmentation is the claim. A six-celled queue drawn on a write-once
parameter promises a history the parameter does not have.

## Vary a shape before adding one

The shape inventory is a vocabulary the reader has to learn, and it is the
one channel with no gradient — two shapes are either the same or they are
not. Every shape added costs every future reader a lookup, so **a new shape
asserts a new kind of thing, and nothing less than that earns one.**

Where the thing is the same kind carrying a property, annotate the shape it
already has:

- **Fill** for which family it belongs to — an input queue and a state
  queue are both queues.
- **A frame** for a property layered on: recorded, tapped, exported.
- **Line style** on that frame for whether the property always holds.
- **Segmentation, size, a corner** for structure within the kind.

These compose. One shape plus three annotation channels expresses more
distinctions than four unrelated shapes, and the reader decodes them
incrementally: *queue, so a history — framed, so recorded — dashed, so
only sometimes.*

The test for whether something has earned a new shape: **would a reader who
knows the base shape still recognize it?** A framed queue is a queue. A
tinted queue is a queue. A page with a folded corner is not a queue, and
should only appear because a parameter genuinely is not a history — a
different kind, not a queue with a property.

Adding a shape is also the change most likely to strand the key, since a
key that fakes shapes can approximate a fill but not a silhouette.

## A default palette

Adopt this or replace it wholesale, but do not mix: a palette works because
its members were chosen against each other.

**Operation fills** — ColorBrewer Accent (`d3.schemeAccent`), pastel rather
than saturated because these are large filled areas a reader looks at for
a long time:

| Role | Hex |
|---|---|
| sensing / ingest | `#BEAED4` |
| estimation | `#7FC97F` |
| detection / decision | `#FFFF99` |
| control / output | `#FDC086` |

**Data tints**, lighter than the operation fills so data reads as quieter
than compute:

| Kind | Hex |
|---|---|
| external input | `#E6DEF2` |
| derived | `#EDEDED` |
| state carried across ticks | `#FBE3CB` |
| parameter | `#DDE4CF` |

**Edge colors** — `#4D4D4D` for ordinary dataflow, then `#CC79A7` and
`#D55E00` for the two paths worth separating. Those two are Okabe–Ito
colors, chosen for colorblind-safe separation; keep them if you change
everything else, because edge color is the channel with no shape to fall
back on.

**Panel** — `#F2F2F2` fill with a `#9E9E9E` border, for the key.

## An edge vocabulary

| Style | Means |
|---|---|
| solid | dataflow that constrains evaluation order |
| dashed, `constraint=false` | a value crossing a cycle boundary — real flow, no ordering claim, free to point against the layout |
| dotted | a dependency that is not dataflow, such as a parameter reaching an operation |
| `penwidth=2` | the product: what the pipeline exists to emit |

Drawing deferred reads dashed *and* unconstrained is what keeps a cyclic
program legible as a DAG: the cycle is visible as data without the layout
engine trying to satisfy it.

Decoration layers onto a glyph without changing its kind, so give it the
same solid/dashed grammar the edges use: a solid frame for a property that
always holds, a dashed frame for one that holds conditionally, no frame for
absent — three states from one channel. Scale the frame with the
glyph — a padding that is a hairline at chart size is a slab on a key's
smaller copy.

## Position carries meaning too

Lay the graph out along the direction of flow (`rankdir=LR` for a
pipeline) and pin each topological level to one rank. Columns then *are*
levels, and a column's height is how much of that work is independent —
which the reader gets for free from a layout they were going to read
anyway.

Where a program has a boundary the reader must respect — asynchronous
arrival on one side, deterministic evaluation on the other — pin the
inputs into a single column so the boundary can be drawn as one rule.

Number the compute nodes with their position in the evaluation order, as a
circled digit (U+2460 onward). That is a real fact about the program, and
numbering is only honest when order carries information.

Where one chart would carry two materially different implementations of the
same flow, give each its own panel on an obvious comparison axis — top and
bottom — and repeat the stages they share. Interleaving two backends'
operation names inside one set of nodes makes the shared lifecycle look like
shared mechanism, which is the opposite of what a comparison is for.
Repetition costs a reader less than one node carrying two vocabularies, and
what the variants genuinely share belongs in the prose beside the chart.

## The key is a graph, not a picture of one

Draw the key with the same node shapes the chart draws, from the same
generator. A key assembled out of table cells or text can only approximate
a shape, and an approximated shape teaches the reader a glyph the chart
never uses.

That usually means the key is its own graph, merged with the chart at
render time so the artifact stays one file. Order its entries by channel —
shapes, then line styles, then symbols — so a reader looking up a mark
searches one band rather than the whole panel.

Show the glyph, never its color's name. "Lilac means input" asks the
reader to translate, fails a colorblind reader entirely, and is
inconsistent with every other entry that shows the thing itself.

A key earns its place once the chart carries more than about two
orthogonal channels — the point at which a reader stops inferring the
encodings from context and starts guessing. Below that it is furniture: a
three-node chain with one edge style needs no legend, and adding one
implies distinctions the chart does not draw.

Keep it subordinate. Smaller type than the chart, tight rows, and seated
in the chart's own whitespace rather than stacked beneath it if the layout
has room — a key that reads as a second diagram competes with the first.

## Placement is computed, not eyeballed

An inset key must be checked against the laid-out graph, not positioned by
looking at it. Check against **node boxes and edge splines both**: an edge
routed through otherwise-empty space is invisible to a box-only check,
which will pass a layout it never looked at.

Make the check fail the build rather than warn. The free region moves
whenever an element is added, and a warning about a diagram nobody is
currently looking at is a warning nobody reads.

A layout engine is doing more placement work than the inset key makes
visible, and a generator that emits geometry itself inherits all of it —
anchoring, centering a multiline label as one line group, distributing the
connectors on one shape, routing to a resolved border, keeping children inside
their parent. `architecture-diagram` carries that set, because emitting
geometry directly is its default. Bounds are the weakest check of the group and
the only one most generators run: every one of those defects sits comfortably
inside the viewport.

## Verify the rendering, not the source

Layout engines accept attributes they ignore. A setting on the wrong graph,
an attribute the shape does not honor, a size that excludes the border it
draws — each fails by doing nothing, and a generator that emitted the
attribute looks correct.

After any change meant to alter layout, confirm the rendered geometry
actually moved: the bounding box, the node positions, the measured gap. If
the number is identical, the change did not land, however plausible the
source looks.

Keep the output a text format whose source is the artifact — a `.dot` or
`.d2` that renders to SVG, not a pasted raster. A text source diffs and
greps; an image is opaque to every tool.

## The perceptual gate

Assertions catch drift, not confusion. A chart whose every claim is derived
and whose every element is in bounds can still label the wrong thing, name
one concept twice, or lay a group out so that the eye reads a difference the
program does not have. Nothing in the generator will report that, so a human
reads every new chart and every visual change, and approval of one revision
does not carry to the next.

When the raster preview that review depends on is unavailable, say so and
name what stood in for it — markup validity, deterministic regeneration
compared by hash, bounds and collision checks, a reviewer reading the
committed artifact. A gate skipped without that record reads afterwards as a
gate passed.

Then convert what the reader found into an assertion. A visual defect fixed
in coordinates comes back at the next layout change; the same defect fixed as
a check on the generator's output does not. This is the only mechanism that
turns one reader's afternoon into a property of the chart.

One thing readers catch that no check will: a label that spends one word on
two meanings. A stage named for `poll` in a chart comparing readiness polling
with task polling names neither, and the reader cannot tell which one the
node is about. Reach for the lifecycle word — `schedule` — and leave the
mechanism to the prose, unless the mechanism is the distinction the chart
exists to draw. This is `technical-prose`'s one-name-one-thing rule, and a
label, with no sentence around it to recover in, is where it hides best.

## Prose in the chart

Prose in and around a chart — node labels, the key, the caption, the charter —
follows `technical-prose`, including its default of American spelling. A
convention the user or the project states overrides that default; a mixture of
both overrides nothing and reads as several authors.

## Red flags

- A node list or edge list maintained by hand beside the code it describes.
- A generator that `continue`s past an element it cannot classify.
- The legend built by a different mechanism than the chart.
- A color named in words in the key.
- A glyph whose structure asserts more than the thing it marks possesses.
- An inset element positioned by adjusting a number until it looked right.
- A collision check that inspects nodes but not edges, or that silently
  matched fewer elements than the graph contains.
- A multiline label centered by its first baseline rather than by its line
  group.
- Connectors attached to one shape at points chosen one at a time.
- A child shape crossing the bounds of the cluster that contains it.
- Two implementations interleaved in one flow rather than drawn as panels.
- A label carrying a word that means two things in the chart it sits in.
- A reported visual defect fixed in coordinates rather than in an assertion.
- A chart shipped without a human reading it, or a review gate skipped for an
  unavailable preview with no record of what stood in for it.
- A layout attribute added without confirming the rendered size changed.
- A diagram checked in as an image with no source beside it.
- Two visual channels carrying the same distinction, or one channel
  carrying two.
- Presenting the conventions here as settled when applying them, rather
  than as defaults the domain may override.
- Finishing a diagram without asking which defaults fought the domain.
- Compute and data sharing a silhouette, so kind can only be read from the
  label.
- A segmented glyph on something with no history to segment.
- A stateful component named inside another node's label rather than drawn.
- A box naming both a process and something that outlives it.
- A new shape introduced for what is really a property of an existing one.
- More shapes in the chart than kinds of thing in the program.
- Palette members borrowed from two different palettes.
- A legend on a chart with one edge style and one node kind.
- Numbered nodes whose numbers encode no order.

