# Claude Maintain Models

> Add new AI models to Kiln's ml_model_list.py and produce a Discord announcement. Use when the user wants to add, integrate, or register a new LLM model (e.g. Claude, GPT, DeepSeek, Gemini, Kimi, Qwen, Grok) into the Kiln model list, mentions adding a model to ml_model_list.py, asks to discover/find new models that are available but not yet in Kiln, or wants to add a net-new AI provider to Kiln.

- Skill: `kiln-ai/claude-maintain-models` (Agent Skill)
- Install (CLI): `npx skillmds@latest add kiln-ai/claude-maintain-models`
- Raw SKILL.md: https://api.skillmd.com/api/skills/kiln-ai/claude-maintain-models/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: Kiln-AI (https://skillmd.com/u/kiln-ai)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/kiln-ai/claude-maintain-models

---


# Add a New AI Model to Kiln

**Branch check first:** if the request involves a provider Kiln does not
support yet, start at [Adding a Net-New Provider](#adding-a-net-new-provider).
That workflow spans core, server, UI, tests, and tooling, and is gated on a
client release — the model-entry steps below are only its final, gated piece.

For a new model on an already-supported provider, integrating it into
`libs/core/kiln_ai/adapters/ml_model_list.py` requires:

1. **`ModelName` enum** – add an enum member
2. **`built_in_models` list** – add a `KilnModel(...)` entry with providers
3. **`ModelFamily` enum** – only if the vendor is brand-new

After code changes, run paid integration tests, then draft a Discord post.

---

## Global Rules

These apply throughout the entire workflow.

- **Slug verification:** NEVER guess or infer model slugs from naming patterns. Every `model_id` must come from an authoritative source (LiteLLM catalog, official docs, API reference, or changelog). If you can't verify a slug, tell the user and ask them to provide it.
- **Date awareness:** These models are often released very recently. Web search for current info before assuming you know the details.

---

## Phase 1 – Model Discovery (only when asked to find new/missing models)

If the user asks you to find new models, **do NOT just web search "new AI models this week"** — that only surfaces major releases. Instead, systematically check each family against **both** the LiteLLM catalog **and** models.dev, then union the results. Both are attempts to catalog available models and each has gaps the other fills.

1. **Read the `ModelFamily` and `ModelName` enums** to know what we already have.

2. **Query both catalogs for each family** (run in parallel where possible):

   **LiteLLM catalog** — filters out mirror providers to avoid duplicates:
   ```bash
   curl -s 'https://api.litellm.ai/model_catalog?model=SEARCH_TERM&mode=chat&page_size=500' -H 'accept: application/json' | jq '[.data[] | select(.provider != "openrouter" and .provider != "bedrock" and .provider != "bedrock_converse" and .provider != "vertex_ai-anthropic_models" and .provider != "azure") | .id] | unique | .[]'
   ```

   **models.dev** — search all model IDs across all providers:
   ```bash
   curl -s https://models.dev/api.json | jq '[to_entries[].value.models // {} | keys[]] | .[]' | grep -i "SEARCH_TERM"
   ```
   For details on a specific provider+model: `curl -s https://models.dev/api.json | jq '.["PROVIDER"].models["MODEL_ID"]'`

3. **Search terms** (one query per term):
   `claude`, `gpt`, `o1`, `o3`, `o4` (OpenAI reasoning), `gemini`, `llama`, `deepseek`, `qwen`, `qwq`, `mistral`, `grok`, `kimi`, `glm`, `minimax`, `hunyuan`, `ernie`, `phi`, `gemma`, `seed`, `step`, `pangu`

4. **Union and cross-reference** results from both catalogs against `ModelName`. A model found in either source counts as available. Focus on direct-provider entries (not OpenRouter/Bedrock/Azure mirrors). **Skip pure coding models** (e.g. `codestral`, `deepseek-coder`, `qwen-coder`).

5. **Run targeted web searches** per family to catch very fresh releases not yet in either catalog:
   - `"[family] new model [current year]"`
   - `"[family] release [current month] [current year]"`

6. **Present findings** as a summary. Let the user decide which to add.

---

## Phase 1B – Lagging-Provider Backfill Check (every run)

Some providers — **Fireworks AI**, **Together AI**, **SiliconFlow** — expose new models on their own endpoints 1–2 weeks before those entries surface in models.dev / LiteLLM. Relying only on those two catalogs will both under-populate the provider list for the model you're adding now **and** miss the window to backfill recently-added models whose provider support has since grown.

Run this check on **every invocation** of the skill, regardless of whether you're in discovery mode or adding a specific model.

1. **Pull the 10 most recently added models from git history.** List position is NOT a recency signal — entries are ordered by family/version/size (see 3c), and net-new families sit at the END of the list, so recent additions can be anywhere:
   ```bash
   git log --follow -p -- libs/core/kiln_ai/adapters/ml_model_list.py | grep -E "^\+\s+name=ModelName\." | head -20
   ```

2. **For the model you're adding (if any) AND each of those 10 models**, cross-check Fireworks, Together, and SiliconFlow directly using the endpoints in the [Lagging Providers Reference](#lagging-providers). Do NOT trust `models.dev` / LiteLLM as the final word for these three providers.

3. **If a lagging provider now supports a recently-added model that isn't yet in its `KilnModel` entry**, flag it to the user and propose either bundling the provider addition into the current change or opening a separate PR. Do not silently add it.

---

## Phase 2 – Gather Context

1. **Read the predecessor model** in `ml_model_list.py` (e.g. for Opus 4.6 → read Opus 4.5). You inherit most parameters from it.

2. **Query the LiteLLM catalog** for the new model. This is the primary slug source since Kiln uses LiteLLM. See the [Slug Lookup Reference](#slug-lookup-reference) for query syntax and all verified sources.

3. **Get the OpenRouter slug** via:
   - `curl -s https://openrouter.ai/api/v1/models | jq '.data[].id' | grep -i "SEARCH_TERM"`
   - Fallback: WebSearch for `openrouter [model name] model id`

4. **Get the direct-provider slug** (Anthropic, OpenAI, Google, etc.). Use the LiteLLM catalog first, then official docs. See the [Slug Lookup Reference](#slug-lookup-reference) for provider-specific URLs.

5. **Identify quirks** — check the [Provider Quirks Reference](#provider-quirks-reference) for the relevant provider, and web search for any new quirks:
   - Structured output mode (JSON schema vs function calling)?
   - Reasoning model (needs `reasoning_capable`, parsers, OpenRouter options)?
   - Vision/multimodal support? Which MIME types?
   - Provider-specific flags (`temp_top_p_exclusive`, etc.)?
   - Rate limit concerns (`max_parallel_requests`)?

6. **Determine thinking levels** — does the model support configurable reasoning effort? See [Thinking Levels Reference](#thinking-levels-reference) for the full lookup chain. Key quick checks:
   - Check the **vendor model page** (e.g. OpenAI model pages say "Reasoning.effort supports: X, Y, Z")
   - Check **OpenRouter** `supported_parameters` — if `reasoning` is absent, skip thinking levels
   - R1-style thinking models (DeepSeek, Qwen thinking variants) do NOT get thinking level dicts

---

## Phase 3 – Code Changes

All changes go in `libs/core/kiln_ai/adapters/ml_model_list.py`.

### 3a. `ModelName` enum

- snake_case: `claude_opus_4_6 = "claude_opus_4_6"`
- Place **before** predecessor (newer first within group). If the vendor is
  brand-new there is no predecessor — start a new group at the end of the enum
- Follow existing grouping (all claude together, all gpt together, etc.)

### 3b. `KilnModel` entry in `built_in_models`

- Place per the **ordering rules in 3c** — this placement is user-visible, get it right
- Copy predecessor's structure and modify: `name`, `friendly_name`, `model_id` per provider, flags
- **`friendly_name` must follow the existing naming pattern** of sibling models in the same family. Check the predecessor. For example, Claude Sonnets use `"Claude {version} Sonnet"` (e.g. "Claude 4.5 Sonnet"), not `"Claude Sonnet {version}"`. Do NOT use the vendor's marketing name if it differs from Kiln's established convention.

**Provider `model_id` formats:**

| Provider | Format | Notes |
|----------|--------|-------|
| `openrouter` | `vendor/model-name` | Always verify via API |
| `openai` | Bare model name | Verify via OpenAI docs |
| `anthropic` | Variable — older models have date stamps, newer may not | Always verify via Anthropic docs |
| `gemini_api` | Bare name | Verify via Google AI Studio docs |
| `fireworks_ai` | `accounts/fireworks/models/...` | Verify via Fireworks docs |
| `together_ai` | Vendor path format | Verify via Together docs |
| `vertex` | Usually same as gemini_api | Verify via Vertex docs |
| `siliconflow_cn` | Vendor/model format | Verify via SiliconFlow docs |
| `featherless_ai` | HuggingFace repo id, case-sensitive (`zai-org/GLM-5.2`) | Verify via their `/v1/models` — see [Featherless](#featherless-ai) |

**Every single `model_id` must be verified from an authoritative source. No exceptions.**

**Setting flags — use catalog data + predecessor as dual signals:**

The LiteLLM catalog and models.dev responses include capability flags (`supports_vision`, `supports_function_calling`, `supports_reasoning`, etc.). Use these as the **primary signal** for what to enable on the new model:

- If the catalog says `supports_vision: true` → enable `supports_vision`, `multimodal_capable`, and vision MIME types (see 2c)
- If the catalog says `supports_function_calling: true` → use `StructuredOutputMode.json_schema` (or `function_calling` depending on provider norms — check predecessor)
- If the catalog says `supports_reasoning: true` → the model can reason, but **do NOT reflexively set `reasoning_capable=True`** — default to `reasoning_capable=False` (see [Reasoning Capable Default](#reasoning-capable-default)). Still add `available_thinking_levels` if it supports effort levels, and check parser/formatter flags.

Then **cross-check against the predecessor**. The predecessor tells you *how* Kiln configures a similar model (which `structured_output_mode`, which provider-specific flags, etc.). The catalog tells you *what* the model can do. Use both:
- Catalog says the model supports vision but predecessor doesn't have it? Enable it — this is a new capability.
- Predecessor has `temp_top_p_exclusive` but nothing in the catalog mentions it? Keep it — it's a provider quirk the catalog doesn't track.
- Catalog and predecessor disagree on something? Trust the catalog for capabilities, trust the predecessor for Kiln-specific configuration patterns.

**Common flags:**
- `structured_output_mode` – how the model handles JSON output
- `suggested_for_evals` / `suggested_for_data_gen` – see **zero-sum rule** below
- `multimodal_capable` / `supports_vision` / `supports_doc_extraction` – see **multimodal rules** below
- `reasoning_capable` – for thinking/reasoning models. **Default new models to `reasoning_capable=False`** unless the model *always* emits its reasoning (see [Reasoning Capable Default](#reasoning-capable-default))
- `temp_top_p_exclusive` – Anthropic models that can't have both temp and top_p
- `parser` / `formatter` – for models needing special parsing (e.g. R1-style thinking)

#### 2c. Multimodal capabilities

If the model supports non-text inputs, configure:

- `multimodal_capable=True` and `supports_doc_extraction=True` if it supports any MIME types
- `supports_vision=True` if it supports images
- `multimodal_requires_pdf_as_image=True` if vision-capable but no native PDF support (also add `KilnMimeType.PDF` to MIME list). **Always set this on OpenRouter providers** — OpenRouter routes PDFs through Mistral OCR which breaks LiteLLM parsing.
- Always include `KilnMimeType.TXT` and `KilnMimeType.MD` on any `multimodal_capable` model

**Strategy: start broad, narrow based on test failures.** Enable a generous set of MIME types, run tests, and remove only types the provider explicitly rejects (400 errors). Don't remove types for timeout/auth/content-mismatch failures.

Full MIME superset (Gemini uses all):
```python
# documents
KilnMimeType.PDF, KilnMimeType.CSV, KilnMimeType.TXT, KilnMimeType.HTML, KilnMimeType.MD
# images
KilnMimeType.JPG, KilnMimeType.PNG
# audio
KilnMimeType.MP3, KilnMimeType.WAV, KilnMimeType.OGG
# video
KilnMimeType.MP4, KilnMimeType.MOV
```

### 3c. Ordering — the list IS the UI

The order of `built_in_models` is exactly the order users see: model dropdowns
group by provider and list each provider's models in list order, and the model
library page lists models in list order. A misplaced entry ships a scrambled
dropdown to every client via the remote config. Rules:

1. **Families are contiguous.** Every model of a family sits in one block.
   Never append a new family member at the bottom of the file or after another
   family — that strands it (past bugs: Mistral Small 4/3 ended up inside the
   Qwen 2.5 region; GLM-Z1 models ended up after the Kimi family).
2. **Within a family, versions run newest → oldest, top to bottom.**
   4 > 3.8 > 3.7 > 3.6 > 3.5 > 3 > 2.5. A new version goes at the TOP of the
   family block. A new model of an *existing* version goes inside that
   version's group — NOT at the top of the family and NOT below older
   versions (past bugs: Gemini 3.6/3.7 Flash were inserted mid-3.1; Qwen 3.6/3.7
   landed below Qwen 3.5 entries; Phi 3.5 sat above Phi 4).
3. **Within a version, big → small. Always.**
   - Commercial tiers: Max > Plus > Flash; Pro > Flash > Flash Lite;
     Large > Medium > Small.
   - Open-weight sizes descending: 405B > 70B > 8B. Keep base/Non-Thinking
     variant pairs adjacent; variant sub-groups (e.g. the Qwen VL Instruct and
     VL Thinking blocks) stay intact, ordered descending internally.
   - Tier-grouped families follow the same rule for their tier blocks: Claude
     runs Fable, then Opus > Sonnet > Haiku, versions descending inside each
     tier. (It once ran Haiku-first, and several families ran sizes
     ascending — those were bugs, not conventions. Do not preserve an
     ascending run because it's "what the family already does.")
4. **A net-new family's block goes at the END of `built_in_models`** (this is
   the existing convention — the newest niche vendors sit at the bottom of the
   list). The "place before predecessor" rule only applies within an existing
   family; a new family has no predecessor.
5. **Verify after editing** — print the order you just shipped and eyeball the
   affected family:

   ```bash
   uv run python -c "
   from kiln_ai.adapters.ml_model_list import built_in_models
   for m in built_in_models: print(m.family, '|', m.friendly_name)"
   ```

### 3d. `suggested_for_evals` / `suggested_for_data_gen`

**Only set these if** the predecessor already has them, OR web search shows the model is a clear SOTA leap (ask user to confirm first).

**Zero-sum rule:** When adding a new model with these flags, remove them from the oldest same-family model to keep the suggested count stable. **Ask the user to confirm** the swap before making changes.

### 3e. `ModelFamily` enum (only if needed)

Only add a new family if the vendor is completely new.

### 3f. Thinking Levels (`available_thinking_levels` / `default_thinking_level`)

If the model supports configurable reasoning effort (not just on/off), add `available_thinking_levels` and `default_thinking_level` to each provider entry. See [Thinking Levels Reference](#thinking-levels-reference) for the full lookup chain and existing constants.

**Quick rules:**
- Reuse an existing `_THINKING_LEVELS` constant if the levels match exactly
- Create a new constant only if levels differ; name it `{MODEL}_{PROVIDER_CONTEXT}_THINKING_LEVELS`
- `default_thinking_level` must be one of the values in `available_thinking_levels`

---

## Phase 4 – Run Tests

Tests call real LLMs and cost money. Ideally the user only needs to consent to two script executions: the smoke test, then the full parallel suite.

**Vertex AI authentication:** Vertex tests require active gcloud credentials. If you are changing a model that uses Vertex, you must not run the test until asking the user to run `gcloud auth application-default login` before trying. These failures are auth issues, not model config problems.

**`-k` filter syntax:** Always use bracket notation for model+provider filtering, never `and`:
- Good: `-k "test_name[glm_5-fireworks_ai]"` or `-k "glm_5"`
- Bad: `-k "glm_5 and fireworks"` — `and` is a pytest keyword expression that can match wrong tests

### 4.0 — If the test env can't build (blocked git dependency)

Kiln's core lib pins `together` to a git fork (`libs/core/pyproject.toml`: `together = { git = "https://github.com/scosman/together-python" }`). In a sandboxed environment whose GitHub access is scoped to `kiln-ai/kiln` only (e.g. Claude Code Web), `uv sync` fails to fetch that fork with a **403** from the git proxy, so the test venv can't be built. This is a GitHub repo-scope block, not a network-domain allowlist issue — every non-scoped repo 403s, only `kiln-ai/kiln` resolves.

Workaround to build the venv for testing (the `together` fork isn't exercised by model-integration tests, which route through LiteLLM):

1. **Do NOT use `uv sync --no-sources`** — it strips the workspace `workspace = true` sources too and makes `kiln-root → kiln-ai` unsatisfiable.
2. Instead, temporarily comment out ONLY the `together = { git = ... }` line in `libs/core/pyproject.toml`, then run `uv sync` (this resolves `together` from PyPI).
3. Run the tests.
4. **Revert** the workaround with `git checkout -- libs/core/pyproject.toml uv.lock` — this restores the commented-out `together` git-fork pin in `libs/core/pyproject.toml` (NOT the repo-root `pyproject.toml`, which the workaround never touches) and reverts `uv.lock` if it changed. Only `ml_model_list.py` (and any intended test-file edits) should remain modified.

Note: the PyPI `together` may pull slightly different transitive deps (e.g. a newer `starlette`), which can cause unrelated collection ImportErrors in desktop/server/rag/vector-store modules — scope your `-k` filters to the model files and ignore those.

### 4a. Parallel testing + API keys

**Parallel testing is already on.** There is no `pytest.ini` — pytest config lives in `[tool.pytest.ini_options]` in the root `pyproject.toml`, and `addopts = "-n auto"` is active. No edit and no revert are needed. (Only override to `-n 8` if a provider rate-limits you.)

**Paid tests read API keys from the ENVIRONMENT, not from the Kiln app's settings.** `conftest.py` has an autouse `use_temp_settings_dir` fixture that points `Config.settings_path` at a temp dir, so `~/.kiln_ai/settings.yaml` is deliberately ignored during tests. A key the user added through the app's provider page **will not be seen** — the test fails with "Attempted to use X without an API key set", which looks like a config bug but isn't.

Bridge the key from the user's settings into the test environment without ever printing it:

```bash
export FEATHERLESS_AI_API_KEY="$(uv run python -c \
  "from kiln_ai.utils.config import Config; print(Config.shared().featherless_ai_api_key)" 2>/dev/null | tail -1)"
```

Use the provider's `env_var` name from `libs/core/kiln_ai/utils/config.py`. To check which keys are available before running, print booleans only — never the values:

```bash
uv run python -c "from kiln_ai.utils.config import Config; c=Config.shared(); print(bool(c.fireworks_api_key))"
```

**Many paid tests also carry the `ollama` marker**, so `--runpaid` alone silently skips them. Always pass `--runpaid --ollama` together, and use `-rs` to see skip reasons when a test you expected to run reports as skipped.

### 4b. Smoke test — verify slug works

Run a single test+provider combo first:

```bash
uv run pytest --runpaid --ollama -k "test_data_gen_sample_all_models_providers[MODEL_ENUM-PROVIDER]"
```

If it fails, fix the slug/config before proceeding. Use `--collect-only` to find exact parameter IDs if unsure.

### 4c. Full test suite

```bash
uv run pytest --runpaid --ollama -k "MODEL_ENUM" -v 2>&1 | grep -E "PASSED|FAILED|ERROR|short test|=====|collected"
```

**If tests fail — debug one at a time:**
1. Pick ONE failing test, run it with `-v` for full output
2. Fix the config
3. Re-run that single test to verify
4. Only re-run the full suite once the single test passes

**Anthropic API key gotcha:** if an Anthropic-direct test fails with an auth/API key error, check whether the user's environment exports the key as `KILN_ANTHROPIC_API_KEY` instead of `ANTHROPIC_API_KEY` (the Kiln app uses the prefixed name; the Anthropic SDK used by tests expects the unprefixed name). Prepend the test command with a one-shot alias — don't `export` it globally:

```bash
ANTHROPIC_API_KEY="$KILN_ANTHROPIC_API_KEY" uv run pytest --runpaid ...
```

### 4d. Extraction tests (if `supports_doc_extraction=True`)

Tests are in `libs/core/kiln_ai/adapters/extractors/test_litellm_extractor.py`.

```bash
# See what will run:
uv run pytest --collect-only libs/core/kiln_ai/adapters/extractors/test_litellm_extractor.py::test_extract_document_success -q | grep MODEL_ENUM

# Run them:
uv run pytest --runpaid --ollama libs/core/kiln_ai/adapters/extractors/test_litellm_extractor.py::test_extract_document_success -k "MODEL_ENUM"
```

If a provider rejects a data type (400 error), remove that `KilnMimeType` and re-run.

### 4e. Confirm failures are actually yours

Before treating a failure as a problem with your change, check whether the same test already fails for an **existing** provider of that model. Several assertions are provider-independent and fail regardless.

Known example: `test_structured_input_cot_prompt_builder` asserts `len(trace) == 5` unconditionally, which is incompatible with any provider setting `reasoning_capable=True` (that selects the single-call strategy, producing 3 messages). It fails for `gpt_oss_120b` on `fireworks_ai` on a clean tree.

```bash
uv run pytest --runpaid --ollama -q "path::test_name[MODEL-OTHER_PROVIDER]"
```

If it fails there too, it's pre-existing — report it as such rather than contorting the config to work around it.

### 4f. Test output format

Collect test results for use in the PR body (Phase 5). Organize by model name and provider using these symbols:
- ✅ for passed tests
- ⚠️ for tests that failed due to content quality flakes (e.g. model returned fewer items than expected, weak assertion mismatches) — include a brief reason
- ❌ for tests that failed due to real errors (bad slug, unsupported feature, 400/500 errors) — include a brief reason
- List every test using the full pytest parametrize ID, grouped by provider
- Include extraction tests (Phase 4d) if they were run

---

## Phase 5 – Create Pull Request

### 5.0 — Important context about Claude Code Web's stop hook

This skill is often run via Claude Code Web (Slack connector). That environment has a **non-user-configurable stop hook** which, at end of session, will:
- Block the session from ending if there are uncommitted changes, untracked files, or unpushed commits
- Instruct the agent to commit and push any local work before stopping
- Explicitly tell the agent NOT to create a PR unless the user asked for one

**The problems this causes:**
1. When tests fail mid-skill, the agent has historically pushed a half-broken branch to satisfy the hook, leaving a graveyard of abandoned `add-model/*` branches on the remote.
2. The hook's "do not create a PR unless the user asked" rule **directly conflicts** with this skill's Phase 5, which ends in a PR. Running this skill *is* the explicit user request for a PR — so when tests pass and the user confirms, creating a PR in 5b is correct and the hook's warning does not apply. Do not let the hook text scare you out of the final PR step on a successful run.

**The user's desires, in priority order:**
1. **Ask before you push.** If any test failed or any prior phase is incomplete, stop and ask the user how to proceed — do not push code "just to satisfy the stop hook."
2. **No abandoned branches.** Never create a branch as a progress-saving mechanism. A branch only exists because the user approved a PR-ready state.
3. **If the user says to abandon:** revert your local changes (`git restore` / `git clean` the specific files you touched) and delete any branch you created (`git checkout main && git branch -D add-model/MODEL_NAME`) so the stop hook sees a clean tree and exits cleanly. Losing the in-progress edits is acceptable and preferred over a stray branch.
4. **On a successful run, push and open the PR as described in 5a/5b.** Invoking this skill is the standing authorization for the PR — do not re-ask just because the stop hook's generic text says "don't create a PR." Only re-ask if tests failed or the user hasn't confirmed the results.

### 5.1 — Gate before pushing

Do NOT commit, push, or create a branch if any of the following are true:
- Any test failed with ❌ (real error — bad slug, unsupported feature, auth issues, 400/500)
- The smoke test (4b) failed and wasn't resolved
- Any step in Phases 2–4 was skipped or incomplete
- You are unsure whether a ⚠️ flake is actually a real failure

If any of the above apply, **stop and ask the user** what to do. Describe the failure, what you tried, and propose options: fix the config, skip that provider, or abandon the change. Only proceed to 5a once the user explicitly confirms.

After all tests pass, commit the changes and open a PR against `main`.

### 5a. Commit and push

1. Create a new branch named `add-model/MODEL_NAME` (e.g. `add-model/glm-5-1`)
2. Stage only the changed files (typically just `ml_model_list.py`)
3. Commit with a concise message (e.g. "Add GLM 5.1 to model list (together_ai, siliconflow_cn)")
4. Push the branch

### 5b. Create the PR

Use `gh pr create` against `main`. The PR body must follow this exact format:

```
## What does this PR do?

 Test Results

[Two paragraphs of nuance — describe any unusual findings, things you tried and reverted, known pre-existing failures vs new failures, API quirks discovered, and any config adjustments made during testing.]

[Model Name] ([provider]):
- [N] passed, [N] skipped[, [N] failed]
- [Any notable failures or flakes]

[Repeat for each model+provider combo]

---
[Model Name] ([provider]):
✅ test_data_gen_all_models_providers[model_enum-provider]
✅ test_data_gen_sample_all_models_providers[model_enum-provider]
✅ test_data_gen_sample_all_models_providers_with_structured_output[model_enum-provider]
✅ test_all_built_in_models_llm_as_judge[model_enum-provider]
✅ test_all_built_in_models_structured_output[model_enum-provider]
✅ test_all_built_in_models_structured_input[model_enum-provider]
✅ test_structured_output_cot_prompt_builder[model_enum-provider]
✅ test_all_models_providers_plaintext[model_enum-provider]
✅ test_cot_prompt_builder[model_enum-provider]
⚠️ test_structured_input_cot_prompt_builder[model_enum-provider] — brief reason
❌ test_name[model_enum-provider] — brief reason

[Repeat for each model+provider combo]

## Checklists

- [X] Tests have been run locally and passed
- [X] New tests have been added to any work in /lib
```

**Rules for the PR body:**
- Every test that ran must appear in the per-test dump, using the full pytest parametrize ID
- Group tests by `[Model Name] ([provider]):` headers
- The summary section at the top gives a quick pass/skip/fail count per model+provider
- The detailed section below the `---` lists every individual test result
- Use ⚠️ for content quality flakes (not real failures), ❌ for real errors

---

## Checklist

- [ ] `ModelName` enum entry added (before predecessor for an existing family; new enum group at the end for a brand-new vendor)
- [ ] `KilnModel` entry added to `built_in_models` per the ordering rules in 3c (family contiguous, versions newest-first, correct slot within the version group; net-new family block at the end of the list)
- [ ] Printed the resulting list order and eyeballed the affected family (3c verification command)
- [ ] `friendly_name` matches the naming pattern of sibling models in the same family
- [ ] If the provider is net-new: followed [Adding a Net-New Provider](#adding-a-net-new-provider), including the release-gating rule (model entries merge only after a client release with the provider plumbing is live)
- [ ] `ModelFamily` enum updated (only if new family)
- [ ] All provider slugs verified from authoritative sources
- [ ] Flags inherited from predecessor and adjusted for quirks
- [ ] `reasoning_capable` defaulted to `False` for adaptive-reasoning models (only `True` for always-emits-reasoning models — see [Reasoning Capable Default](#reasoning-capable-default))
- [ ] Thinking levels configured if model supports reasoning effort (see [Thinking Levels Reference](#thinking-levels-reference))
- [ ] Preserve existing comments from predecessor (e.g. reasoning notes, MIME type groupings)
- [ ] Zero-sum applied if model is suggested for evals/data gen
- [ ] RAG config templates updated if the new model replaces one used in `app/web_ui/src/routes/(app)/docs/rag_configs/[project_id]/add_search_tool/rag_config_templates.ts`
- [ ] API keys bridged into the test env (see 4a — settings.yaml is NOT used by tests)
- [ ] Smoke test passed
- [ ] Full test suite passed
- [ ] Failures cross-checked against an existing provider before being called regressions (see 4e)
- [ ] PR created against `main` with test results in the body

---

## Reasoning Capable Default

**Default newly-added models to `reasoning_capable=False`, even when the catalog reports `supports_reasoning: true`.**

Most recent models use *adaptive* reasoning and sometimes return no reasoning at all. Kiln raises `RuntimeError("Reasoning is required for this model, but no reasoning was returned.")` whenever a provider has `reasoning_capable=True` and no reasoning comes back, so evals and data-gen runs fail sporadically on adaptive-reasoning models.

`reasoning_capable` is **orthogonal to `available_thinking_levels`** — you keep the important behavior with it set to `False`:
- Thinking-level / effort selection still works (every GPT-5.x entry sets `available_thinking_levels` with `reasoning_capable` unset).
- Reasoning is still parsed and displayed when the model returns it (parser-driven, independent of this flag).
- The `test_thinking_level_reasoning_content` paid test still runs and still asserts reasoning is present — it is gated on `available_thinking_levels`, not `reasoning_capable`. So you do NOT lose the ability to test the model's reasoning, as long as it has thinking levels.

What you give up with `reasoning_capable=False`:
- The conditional "reasoning is present in `intermediate_outputs`" assertion inside the structured-output / structured-input paid tests stops firing (the tests themselves still run — no whole test is skipped).
- When a user attaches a chain-of-thought prompt, the model uses the two-call `two_message_cot` strategy instead of the single-call native `single_turn_r1_thinking`.

**Keep `reasoning_capable=True` only for models that *always* emit reasoning** in a native `<think>` format — DeepSeek R1, QwQ, Qwen thinking variants, gpt-oss — where reasoning is guaranteed and you want the single-call COT strategy.

**Narrower alternative:** if a model reliably reasons but you only hit the error on structured output, set `reasoning_optional_for_structured_output=True` (requires `reasoning_capable=True`) instead of disabling reasoning entirely.

---

## Adding a Net-New Provider

Adding a provider Kiln has never supported is a bigger job than adding a model,
and it has a hard sequencing constraint. Follow this section end to end.

### The release-gating rule (do NOT skip)

The remote config is generated from main's `built_in_models` and reaches **all
existing clients immediately** — but provider *support* (name map, connect
flow, API key handling) only reaches users through a client app release. If
model entries for a brand-new provider land on main before a client release
with the provider plumbing is live, every deployed client shows the raw
provider ID (e.g. `featherless_ai`) in the model library and offers models
nobody can connect to. This happened with Featherless in Aug 2026 and the
entries had to be rolled back.

**Sequence it in two PRs:**

1. **PR 1 — provider plumbing only.** The `ModelProviderName` enum member and
   everything else in the checklist below EXCEPT the model catalog. Merge
   whenever ready.
2. **PR 2 — model catalog.** ALL `ml_model_list.py` changes: `ModelName`
   members, `ModelFamily` (if the vendor is new), and the `built_in_models`
   entries. Open it, but **merge only after a client release containing PR 1
   is live.** (Only `built_in_models` is published via the remote config, but
   the enums belong in the same PR as the entries they exist for.)

### Touchpoint checklist (from the Featherless integration, #1618)

**First, identify the provider's authentication model** — the checklist below
describes the common single-API-key pattern, but not every provider fits it.
Existing variants to crib from: Bedrock stores an access key + secret pair,
Fireworks a key + account ID, Azure OpenAI a key + endpoint, Vertex a project
ID (auth via gcloud ADC), and Ollama / Docker Model Runner store only a base
URL with no credentials. Follow the closest existing analog's plumbing through
`config.py`, `provider_warnings`, `provider_api.py`, and the connect page —
the credential fields, validation call, and UI steps all change with the auth
model.

**libs/core:**
- `ModelProviderName` enum — `libs/core/kiln_ai/datamodel/datamodel_enums.py`
- Credential storage — `libs/core/kiln_ai/utils/config.py` (for API-key
  providers: a new key with `env_var`; otherwise whatever fields the auth
  model needs)
- LiteLLM provider mapping — `libs/core/kiln_ai/utils/litellm.py`
- `libs/core/kiln_ai/adapters/provider_tools.py` — three spots: the
  `provider_name_from_id` match (friendly name; pyright flags a missed case),
  `provider_warnings` (missing-credential message, `required_config_keys`),
  and the adapter config (credential / base URL / headers plumbing)
- Model catalog — `libs/core/kiln_ai/adapters/ml_model_list.py`: `ModelName`
  members, `ModelFamily` if needed, and the `built_in_models` entries
  (**PR 2 only**, see gating rule)

**Desktop server (`app/desktop/studio_server/provider_api.py`):**
- `connect_<provider>` credential-validation endpoint (find a cheap
  authenticated call; see the Featherless connect function for a pattern when
  the provider has no authenticated GET to ping)
- Disconnect handling (clear the stored credentials)
- Tests in `test_provider_api.py`

**Web UI (`app/web_ui`):**
- `src/lib/stores.ts` — `provider_name_map` entry (friendly name)
- `src/lib/ui/provider_image.ts` + SVG in `static/images/` — icon must match
  the monochrome convention: bare glyph, `currentColor`, no background tile,
  viewBox cropped to the artwork
- Connect page — `src/routes/(fullscreen)/setup/(setup)/connect_providers/connect_providers.svelte`:
  provider card (name, description, and the auth flow — `api_key_steps` /
  `api_key_fields` for key providers, or the custom flow the auth model
  needs) plus connected-status
  wiring from `settings` keys
- `src/lib/api_schema.d.ts` — regenerate with `make schema` (needs the server
  running on :8757)

**Tests & tooling:**
- `libs/core/kiln_ai/adapters/test_provider_tools.py`, `test_adapter_registry.py`,
  `model_adapters/test_litellm_adapter.py` — extend the per-provider
  parametrized cases
- `.agents/scripts/provider_utils.py` — add the provider's model-catalog
  endpoint so agent tooling can enumerate its models

**Known gap:** the generated Copilot API client
(`app/desktop/studio_server/api_client/`) mirrors the remote Copilot service's
schema — it can't learn the new provider until that service updates. Note it in
the PR rather than hand-editing generated code.

### Friendly names everywhere

The raw enum value must never be user-visible. When you add the provider,
verify a friendly name exists in **both** name maps (python
`provider_name_from_id` and web `provider_name_map`) — pyright and typescript
respectively force these when the enum gains a member, which is why the
plumbing PR must not skip them.

---

## Provider Quirks Reference

### Anthropic
- Newer models (Opus 4.1+, Sonnet 4.5+) need `temp_top_p_exclusive=True`
- Opus 4.5+ uses `json_schema`; older Opus uses `function_calling`
- Extended thinking models: `anthropic_extended_thinking=True` + `reasoning_capable=True`

### OpenAI
- Most GPT models use `json_schema` for structured output
- GPT-5.x models support `available_thinking_levels` — see [Thinking Levels Reference](#thinking-levels-reference)
- Chat/instant variants (e.g. GPT-5.3 Instant) may not support reasoning effort
- o-series models have fixed thinking tiers (separate model entries per tier, not configurable levels)

### Google/Gemini
- `gemini_reasoning_enabled=True` for reasoning-capable models
- Gemini 3.x models support `available_thinking_levels` — see [Thinking Levels Reference](#thinking-levels-reference)
- Rich multimodal support (audio, video, images, documents)

### DeepSeek
- R1 models: `parser=ModelParserID.r1_thinking` + `reasoning_capable=True`
- V3 models: often available on OpenRouter, Fireworks, SiliconFlow CN
- Some need `r1_openrouter_options=True` + `require_openrouter_reasoning=True`

### OpenRouter (general)
- Slugs: `vendor/model-name`
- Reasoning models: may need `require_openrouter_reasoning=True`
- Some models: `openrouter_skip_required_parameters=True`
- Logprobs: `logprobs_openrouter_options=True` if supported
- Always `multimodal_requires_pdf_as_image=True` (OpenRouter's PDF routing breaks LiteLLM)

### Featherless AI

Serverless host for HuggingFace-hosted open weights. Several hard constraints — read before adding any model:

- **`json_instructions` is the only usable structured output mode.** LiteLLM's `featherless_ai` provider rejects `response_format` outright (`UnsupportedParamsError`), so `json_schema`, `json_mode`, and `json_instruction_and_object` are all unavailable. Routing it as a custom `openai` provider bypasses that gate, but was tested and is *not* reliable: GLM 5.2 accepts `json_schema` and silently ignores it (returns unstructured text), while DeepSeek V4 Pro and Kimi K2.6 return an APIError. Don't use it.
- **Gated models return HTTP 403 `model_gated_needs_oauth`** and cannot work for arbitrary Kiln users — they require each user to link a HuggingFace org to their Featherless account. **Always filter `is_gated` before adding.** Note all 20 official `meta-llama/*` repos are gated, so no Llama variant is usable.
- **No cost reporting.** Featherless models aren't in LiteLLM's price map (only two legacy `Qwerky` entries, on `main` too — not a version issue), and Featherless doesn't return cost in the `usage` object. Runs record tokens with `cost: null`.
- **Not in models.dev or the LiteLLM catalog**, so their `/v1/models` endpoint is the only authoritative source. See [Lagging Providers](#lagging-providers).
- Quality varies per deployment — verify with a paid run. Qwen 3.5 397B, for example, returns degenerate output (rambles to the token cap) and was excluded for that reason.

### Qwen3 / Thinking Models
- Thinking variants: `reasoning_capable=True`, `parser=ModelParserID.r1_thinking`
- No-thinking variants: `formatter=ModelFormatterID.qwen3_style_no_think`
- SiliconFlow may need `siliconflow_enable_thinking=True/False`

---

## Thinking Levels Reference

No API provides the available thinking levels programmatically — they must be manually sourced. Use this lookup chain in priority order:

### Lookup Chain

1. **Vendor model page** (most authoritative)
   - **OpenAI:** Each model page includes "Reasoning.effort supports: X, Y, Z" in the description text. URL: `https://developers.openai.com/api/docs/models/{model-id}`
   - **Anthropic:** The [effort docs](https://platform.claude.com/docs/en/build-with-claude/effort) list levels per model. Opus 4.6 supports `low, medium, high, max`; Sonnet 4.6 supports `low, medium, high`.
   - **Google Gemini:** The models API returns `thinking: true/false` (boolean only). Levels come from docs.

2. **Vercel AI Gateway docs** — clean structured tables per provider:
   - `https://vercel.com/docs/ai-gateway/capabilities/reasoning/openai`
   - `https://vercel.com/docs/ai-gateway/capabilities/reasoning/anthropic`
   - `https://vercel.com/docs/ai-gateway/capabi

…(truncated)
