# Picking A Format

> Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

- Skill: `thedixitjain/picking-a-format` (Agent Skill)
- Install (CLI): `npx skillmds add thedixitjain/picking-a-format`
- Raw SKILL.md: https://api.skillmd.com/api/skills/thedixitjain/picking-a-format/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Web & Frontend
- Author: thedixitjain (https://skillmd.com/u/thedixitjain)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/thedixitjain/picking-a-format

---



# Picking a format

Kreuzberg has two orthogonal format knobs. Get them right up front and the
downstream code stays simple.

| Knob                | What it controls                                  | Values                                 | Default          |
| ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- |
| `--format`          | How the CLI prints the result                     | `text`, `json`                         | `text` (`extract`), `json` (`batch`) |
| `--content-format`  | How extracted content is rendered inside `result` | `plain`, `markdown`, `djot`, `html`    | `plain`          |
| `--token-reduction` | Strip whitespace / boilerplate for LLM contexts   | `off`, `light`, `moderate`, `aggressive` | `off`          |

`--format json` always returns the full `ExtractionResult` (content +
metadata + tables + images). `--format text` prints just `content`.
`--content-format` is what shows up inside that `content` field.

## Decision tree

```text
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│       --format text --content-format markdown
├── Vector store / RAG indexer
│       --format json --content-format markdown
│       (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│       --format json --content-format plain
│       (cleanest text + structured metadata)
├── Human review / archival
│       --format text --content-format markdown
├── HTML re-rendering / web display
│       --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│       --format json --content-format djot
└── Token-budget-constrained pipeline
        --format text --content-format plain
        (drops markup; add --token-reduction moderate for further savings)
```

## Examples

Feed a PDF directly into an LLM:

```bash
kreuzberg extract paper.pdf --content-format markdown
```

Index a corpus into a RAG store with tables and headings preserved:

```bash
kreuzberg batch docs/*.pdf --format json --content-format markdown \
  | jq -c '.[] | {path: .metadata.path, content: .content, tables: .tables}'
```

Strip a file to bare text for a token-tight summarizer:

```bash
kreuzberg extract long.pdf \
  --content-format plain \
  --token-reduction moderate
```

Pull metadata only, ignore content:

```bash
kreuzberg extract file.pdf --format json | jq '.metadata'
```

## When in doubt

- **Default to `markdown`** as the content format. It is the best
  compromise across LLMs, RAG, and human review, and Kreuzberg has the
  most faithful renderer for it.
- Reach for `plain` only when downstream cannot tolerate any markup.
- Reach for `djot` only if you're already in a djot/pandoc pipeline.
- Reach for `html` only when re-rendering for the web.

## Token-reduction (orthogonal)

`--token-reduction` collapses whitespace, strips repeated headers/footers,
and trims boilerplate. It composes with any `--content-format`:

- `off` (default), `light`, `moderate`, `aggressive`, `maximum`.

Use `moderate` as a safe starting point for LLM context windows. `maximum`
is lossy — verify before relying on it.

See `references/cli-reference.md` for the full flag set and
`references/configuration.md` for the equivalent `output_format` and
`token_reduction` keys in `kreuzberg.toml`.

---

**Source:** [`hashgraph-online/awesome-codex-plugins`](https://github.com/hashgraph-online/awesome-codex-plugins) → `plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/picking-a-format/SKILL.md`

