# Htr Transcription

> Transcribe handwritten historical documents with the HTRflow MCP tools — batch images into one htr_transcribe call and present the interactive viewer. Use when the user wants to transcribe or OCR handwriting, read old handwriting, or digitize a manuscript.

- Skill: `ai-riksarkivet/htr-transcription` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ai-riksarkivet/htr-transcription`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ai-riksarkivet/htr-transcription/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: ai-riksarkivet (https://skillmd.com/u/ai-riksarkivet)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ai-riksarkivet/htr-transcription

---


# HTR Transcription

Transcribe handwritten historical documents using the HTRflow MCP server.
Returns an interactive viewer, per-line transcription JSON, and archival exports.

## Tools

- `htr_transcribe` — Transcribe images and return result URLs

## Workflow

### 1. Determine image source

- **http/https URLs** (IIIF links, public image URLs): Use directly — skip to step 2.
- **Local files or attachments**: Must be uploaded first. Use the `/upload-files` skill, then continue to step 2.

### 2. Transcribe

Call `htr_transcribe` once with ALL image URLs in a single call.

**Batching rule**: Never call `htr_transcribe` multiple times for separate
images. Each call runs an expensive GPU pipeline — batch everything.

### 3. Present results

After transcription, present results as an **inline artifact** for the viewer
and **downloadable links** for data exports.

#### 4a. Inline viewer artifact

Download the viewer HTML, then inline all external dependencies (OpenSeadragon
JS and images) so the artifact is fully self-contained (the artifact sandbox
blocks external requests).

```bash
curl -sL "{viewer_url}" -o /home/claude/viewer.html
```

Then run this Python script to embed dependencies:

```python
import re, base64, urllib.request

with open("/home/claude/viewer.html", "r") as f:
    html = f.read()

# Inline OpenSeadragon JS (CDN script -> inline script)
osd_match = re.search(r'<script src="(https://cdn[^"]+openseadragon[^"]+)">\s*</script>', html)
if osd_match:
    with urllib.request.urlopen(osd_match.group(1)) as resp:
        osd_js = resp.read().decode()
    html = html.replace(osd_match.group(0), f"<script>{osd_js}</script>")

# Embed all Gradio image URLs as base64 data URIs
for url in set(re.findall(
    r'https://riksarkivet-htr-demo\.hf\.space/gradio_api/file=[^\s"]+\.(?:jpg|png)', html
)):
    with urllib.request.urlopen(url) as resp:
        img_data = resp.read()
    ext = "jpeg" if url.endswith(".jpg") else "png"
    data_uri = f"data:image/{ext};base64,{base64.b64encode(img_data).decode()}"
    html = html.replace(url, data_uri)

with open("/mnt/user-data/outputs/viewer.html", "w") as f:
    f.write(html)
```

Then call `present_files` with `/mnt/user-data/outputs/viewer.html` to render
the interactive viewer as an inline artifact.

#### 4b. Export links

Provide the remaining URLs as clickable download links:

> - **Transcription data**: [pages_url] (per-line JSON)
> - **Export**: [export_url] (archival export)

Do NOT reproduce document text as plain text in your response — present
the artifact and links instead.

## Options

### Language

| Value       | Use when                            |
|-------------|-------------------------------------|
| `swedish`   | Swedish handwriting (default)       |
| `norwegian` | Norwegian handwriting               |
| `english`   | English handwriting                 |
| `medieval`  | Medieval scripts                    |

### Layout

| Value        | Use when                                         |
|--------------|--------------------------------------------------|
| `single_page`| Single pages, snippets, cropped regions (default)|
| `spread`     | Two-page book openings (Swedish only)            |

### Export format

| Value      | Description                             |
|------------|-----------------------------------------|
| `alto_xml` | ALTO XML — standard archival (default)  |
| `page_xml` | PAGE XML — alternative archival format  |
| `json`     | JSON — structured data format           |

### Custom pipeline

`custom_yaml` accepts a raw HTRflow YAML config string. Overrides
`language` and `layout`. Use only when user explicitly provides one.

Example — English modern handwriting with a custom TrOCR model:

```yaml
steps:
- step: Segmentation
  settings:
    model: yolo
    model_settings:
      model: Riksarkivet/yolov9-lines-within-regions-1
- step: TextRecognition
  settings:
    model: TrOCR
    model_settings:
      model: microsoft/trocr-base-handwritten
    generation_settings:
       batch_size: 16
- step: OrderLines
```

