vLLM Tool Parsers — Navigation Map
This skill points to the right source file, template, or GH issue. The source code is authoritative — read it. Do not paraphrase from this skill when the actual file is available.
Where things live
Assume a local vllm-project/vllm checkout is accessible. Every reference below is relative to that repo root.
| Target | Read |
|---|---|
| All tool parsers | vllm/tool_parsers/ (one file per parser) |
Parser base class + ToolParserManager |
vllm/tool_parsers/abstract_tool_parser.py |
Shared helpers (partial_json_loads, find_common_prefix, make_valid_python, partial_tag_overlap, compute_tool_delta, handle_single_tool) |
vllm/tool_parsers/utils.py |
| Built-in parser registry | vllm/tool_parsers/__init__.py — _TOOL_PARSERS_TO_REGISTER maps CLI name → module → class |
| Unified parser engine (new) | vllm/parser/ — one class per model (qwen3.py, gemma4.py, deepseek_v4.py, deepseek_v32.py, seed_oss.py, …), abstract_parser.py, and engine/ (parser_engine.py, streaming_parser_engine.py, incremental_lexer.py, token_id_scanner.py) |
| Adapter construction | vllm/parser/engine/registered_adapters.py — make_adapters(XParser) returns (XParserReasoningAdapter, XParserToolAdapter); the tool side is then subclassed in vllm/tool_parsers/*_engine_tool_parser.py to attach structural_tag_model |
| CLI flag definitions | vllm/entrypoints/openai/cli_args.py — grep tool_call_parser, enable_auto_tool_choice, tool_parser_plugin |
| Non-streaming serving invocation | vllm/entrypoints/openai/chat_completion/serving.py — grep extract_tool_calls |
| Streaming serving loop + tail flush | same file — grep extract_tool_calls_streaming, prev_tool_call_arr |
| Plugin import wiring | vllm/entrypoints/openai/api_server.py — grep import_tool_parser |
| Responses API tool handling | vllm/entrypoints/openai/responses/serving.py + vllm/entrypoints/openai/parser/responses_parser.py |
| Per-parser Jinja chat templates | examples/tool_chat_template_<family>.jinja |
| Per-parser tests (executable spec) | tests/tool_parsers/test_<name>_tool_parser.py + tests/tool_parsers/common_tests.py |
| User-facing docs | docs/features/tool_calling.md |
If the operator's question is "what does parser X do" — read vllm/tool_parsers/X_tool_parser.py. Don't rely on this skill's paraphrase.
Except for the 13 names on the unified-parser path, where that file is a
stub of a few lines and the logic lives in vllm/parser/<model>.py:
| CLI name(s) | Registry class | Real implementation |
|---|---|---|
qwen3_coder, qwen3_xml, mimo |
Qwen3EngineToolParser |
vllm/parser/qwen3.py |
gemma4 |
Gemma4EngineToolParser |
vllm/parser/gemma4.py |
deepseek_v4 |
DeepSeekV4EngineToolParser |
vllm/parser/deepseek_v4.py |
deepseek_v32 |
DeepSeekV32EngineToolParser |
vllm/parser/deepseek_v32.py |
seed_oss |
SeedOssEngineToolParser |
vllm/parser/seed_oss.py |
glm45, glm47 |
Glm47MoeModelToolParser |
vllm/parser/glm47_moe.py |
kimi_k2 |
KimiK2ToolParser |
vllm/parser/kimi_k2.py |
minimax_m2 |
MinimaxM2ToolParser |
vllm/parser/minimax_m2.py |
mistral |
MistralToolParser |
vllm/parser/mistral.py — moved onto this path at v0.27.0 (PR #48947) |
inkling |
InklingEngineToolParser |
vllm/parser/inkling.py — new at v0.27.0 |
_engine_ in the filename is not the marker. glm47_moe_tool_parser.py,
kimi_k2_tool_parser.py, minimax_m2_tool_parser.py and mistral_tool_parser.py
have ordinary names and are still stubs. The test is whether the file imports
the adapter: grep -l "registered_adapters import" vllm/tool_parsers/*.py.
This is the same refactor described in vllm-reasoning-parsers — a single
per-model parser now backs both the tool and reasoning adapters (RFC
#32713, still formally OPEN
and stale-bot-marked while the code ships). Practical consequence: a grammar
change to vllm/parser/qwen3.py moves tool and reasoning behaviour at once —
they are no longer independent surfaces for those models.
The CLI contract
Two flags, both required together for auto tool choice:
vllm serve <model> --enable-auto-tool-choice --tool-call-parser <name> [--chat-template <path>]
--enable-auto-tool-choicealone →TypeError: --enable-auto-tool-choice requires --tool-call-parser(seecli_args.py).--tool-call-parseralone → legal. Parser still runs fortool_choice="required"and named, and on Responses API.- No
autosentinel. Name a concrete parser. --tool-parser-plugin <path.py>→ third-party file that calls@ToolParserManager.register_module("name").--reasoning-parseris independent but several tool parsers assume a</think>has closed — match them (see "Reasoning pairing" below).- Chat template often matters. Each parser has a reference Jinja at
examples/tool_chat_template_<family>.jinja. Wrong template → model never emits the sentinels the parser expects.
Parser → model family index
Use this to pick the CLI name. Then read the parser file and the matching Jinja for details — the wrapping tokens, streaming strategy, and quirks live there, not here.
--tool-call-parser |
Model families | Reference template |
|---|---|---|
hermes |
Hermes-2/3, Qwen2.5-Instruct, Qwen3-Instruct (text), QwQ | tool_chat_template_hermes.jinja |
longcat |
LongCat-Flash-Chat | (inherits hermes) |
mistral |
Mistral-Instruct (all), Mistral-Large-2506+ (v≥11 format auto-detected) | tool_chat_template_mistral.jinja (also _mistral3.jinja, _mistral_parallel.jinja) |
llama3_json / llama4_json |
Llama 3.1/3.2/3.3/4 (JSON flavor) | tool_chat_template_llama3.1_json.jinja, _llama3.2_json.jinja, _llama4_json.jinja |
pythonic |
Llama-3.2-{1B,3B}, ToolACE-8B | tool_chat_template_llama3.2_pythonic.jinja, tool_chat_template_toolace.jinja |
llama4_pythonic |
Llama-4 Scout/Maverick | tool_chat_template_llama4_pythonic.jinja |
olmo3 |
Olmo-3-7B/32B | (HF default) |
qwen3_coder / qwen3_xml / mimo |
Qwen3-Coder-480B/30B, Qwen3-XML family | tool_chat_template_qwen3coder.jinja — all three names are one class at v0.25.1 (Qwen3EngineToolParser); the separate coder/xml files were deleted |
deepseek_v3 / deepseek_v31 / deepseek_v32 / deepseek_v4 |
DeepSeek-V3/R1, V3.1, V3.2, V4 | tool_chat_template_deepseekv3.jinja, _deepseekv31.jinja, _deepseekr1.jinja |
cohere_command3 / cohere_command4 |
Command-A, Command-R7B (3); Command-A-Reasoning/Vision (4) | <|START_ACTION|> grammar (HF default) |
apertus |
Apertus | (HF default) |
lfm2 |
LFM2 | (HF default) |
minicpm5 |
MiniCPM-5 | (HF default — no tool_chat_template_minicpm5.jinja ships) |
poolside_v1 |
Poolside (GLM-4-style grammar) | (HF default) |
hy_v3 |
Hunyuan V3 (newer than hunyuan_a13b) |
(HF default) |
glm45 / glm47 |
GLM-4.5/4.6, GLM-4.7 | tool_chat_template_glm4.jinja |
granite / granite-20b-fc / granite4 |
Granite-3.0/3.1, Granite-20B-FC, Granite-4.0 | tool_chat_template_granite.jinja, _granite_20b_fc.jinja |
phi4_mini_json |
Phi-4-mini | tool_chat_template_phi4_mini.jinja |
jamba |
Jamba-1.5 | (HF default, sentinel must be in vocab) |
internlm |
InternLM-2.5 | tool_chat_template_internlm2_tool.jinja |
kimi_k2 |
Kimi-K2 Instruct / Thinking | (HF default) |
kimi_k3 |
Kimi-K3 (XTML <|open|>tools<|sep|> channels) — new at v0.27.0 |
(HF default) |
inkling |
Inkling — new at v0.27.0; typed <|content_text|>/<|content_thinking|>/<|content_invoke_tool_json|> blocks |
(HF default) |
minimax_m2 / minimax_m3 |
MiniMax-M2 / M3 | the bare minimax name was removed at v0.25.1 — --tool-call-parser minimax no longer resolves |
step3 / step3p5 |
Step-3 VL / Step-3.5-Flash | (HF default) |
seed_oss |
Seed-OSS | (HF default) |
hunyuan_a13b |
Hunyuan-A13B | (HF default) |
ernie45 |
ERNIE-4.5 thinking | (HF default) |
gemma4 / functiongemma |
Gemma-4-IT / FunctionGemma-270m | tool_chat_template_gemma4.jinja, _functiongemma.jinja |
gigachat3 |
GigaChat-3 | (HF default) |
xlam |
Salesforce xLAM Llama & Qwen | tool_chat_template_xlam_llama.jinja, _xlam_qwen.jinja |
openai |
gpt-oss-20b/120b (Harmony channels) | (no Jinja — built-in renderer) |
gpt-oss/Harmony changed at v0.27.0 (PR #45560). json_object/json_schema
response_format is now rewritten into a Harmony-aware structural_tag in
HarmonyParser.adjust_request, so constrained decoding governs the whole
generation instead of only the post-<|channel|>final<|message|> region. Without
builtin tools the grammar validates tool name + arguments; with builtin tools it
falls back to "some tool is called" only.
Don't trust this table to be complete — verify with:
grep -E "^\s+\"" vllm/tool_parsers/__init__.py # lists registered names
ls examples/tool_chat_template_*.jinja # lists shipped templates
ls vllm/tool_parsers/*_tool_parser.py # lists source files
Framework contract (mental model)
Worth carrying as mental model, because it's spread across multiple files and easy to miss:
ToolParsersubclass implementsextract_tool_calls(non-streaming, stateless) andextract_tool_calls_streaming(stateful, per-delta). Seevllm/tool_parsers/abstract_tool_parser.py.- Serving-layer invariants (guaranteed to the streaming method):
current_text == previous_text + delta_textcurrent_token_ids == previous_token_ids + delta_token_ids- Deltas may span multiple tokens.
- Four state fields the parser MUST maintain:
prev_tool_call_arr: list[dict]— serving reads[i]["arguments"]at stream end to flush the tail. If empty at end,finish_reasonbecomesstopnottool_calls.current_tool_id: int— starts-1, increments per call.current_tool_name_sent: bool— flip True once name flushed for current tool.streamed_args_for_tool: list[str]— cumulative args already emitted per tool index. Append on every flush or the tail double-streams.
- Optional:
adjust_request(request)— setskip_special_tokens=False, inject grammar, etc.supports_required_and_named: bool = True— flip False if the output shape breaks guided JSON. - Return-value contract for streaming:
None= "consumed, nothing to emit";DeltaMessage(content=...)= pass-through;DeltaMessage(tool_calls=[DeltaToolCall(...)])= tool progress.
For the parse_delta refactor see RFC #11522 (closed 2025-09-05) and its follow-on PRs #38755 (merged 2026-04-08), #39728 (merged 2026-04-13), #39446 (merged 2026-04-14). Align new parsers with the parse_delta shape rather than copying older HACKs.
Reasoning-parser pairing
Several tool parsers gate on a reasoning-end sentinel. Mismatched pair = tool parser sees reasoning text as content, misses sentinels, or emits from inside <think>. Table below is a pointer — verify by reading the tool-parser file for the adjust_request/is_reasoning_end interaction.
| Tool parser | Pair with | Why |
|---|---|---|
hermes (Qwen3 thinking) |
qwen3 |
Gates on </think> |
deepseek_v3 (R1) |
deepseek_r1 |
Gates on </think> |
seed_oss |
seed_oss |
Gates on </seed:think> |
hunyuan_a13b |
hunyuan_a13b |
Excludes <think>…</think> region |
minimax_m2, minimax_m3 |
same name | Interleaved thinking / exclusion zone |
kimi_k2 |
kimi_k2 |
Implicit end via <|tool_calls_section_begin|> |
ernie45 |
ernie45 |
Expects </think>\n\n\n<tool_call> |
mistral (reasoning variants) |
mistral |
Tokenized reasoning section |
Sibling skill: vllm-reasoning-parsers — defer there for reasoning-side questions.
Diagnostic playbook
When a user reports a broken tool call, work down this list. Each step names where to look.
- Both flags set? Check serve command for
--enable-auto-tool-choice --tool-call-parser <name>. - Right parser for the model? Cross-check with the table above AND with
vllm/tool_parsers/__init__.pyregistry. - Chat template matches? Inspect the actual template bytes — don't trust the filename. See step 7.
- Reasoning parser paired? If the model has
<think>/</seed:think>/ harmony channels,--reasoning-parsermust match. Checkvllm/reasoning/for the registry. - Streaming vs non-streaming? Grep the parser file: if
extract_tool_calls_streamingreturnsNoneunconditionally or raisesNotImplementedError, streaming isn't supported (e.g.phi4_mini_json,openai). finish_reason=stop? Thenprev_tool_call_arris empty at stream end — parser failed to populate it. Enable vLLM debug logs and trace.- Raw model output vs what parser sees. Bypass the parser: call
/v1/completions(no tool parser) with the same prompt. Dump raw bytes:
This distinguishes "model not emitting sentinels" (template/training problem) from "parser not matching sentinels" (parser bug).curl -sS $VLLM/v1/completions -H 'content-type: application/json' \ -d '{"model":"...","prompt":"...","max_tokens":200}' \ | python3 -c "import sys,json,unicodedata; t=json.load(sys.stdin)['choices'][0]['text']; \ [print(hex(ord(ch)), unicodedata.name(ch,'?'), repr(ch)) for ch in t if ord(ch)>0x7E][:40]" - vLLM version vs known bugs. See "Bug archaeology" below.
- Custom template? Diff the deployed Jinja against
examples/tool_chat_template_<family>.jinja. Character-level. Full-width vs ASCII Unicode gotchas bite here.
Bug archaeology (don't trust static lists — search)
Tool-parser bugs churn quickly. Teach yourself the pattern:
# Open bugs affecting a specific parser
gh search issues --repo vllm-project/vllm "<parser-name> tool" --state open \
--json title,url,state,updatedAt --limit 20
# Recently merged fix PRs
gh search prs --repo vllm-project/vllm "<parser-name>" --state merged \
--json title,url,mergedAt --limit 20
# Streaming-specific
gh search issues --repo vllm-project/vllm "tool parser streaming <parser-name>" \
--json title,url,state,updatedAt --limit 20
# Broad sweep if parser name unknown
gh search issues --repo vllm-project/vllm "extract_tool_calls_streaming" --limit 30
On finding a referenced issue/PR, read it directly (gh issue view N --repo vllm-project/vllm --comments). Never paraphrase from a cached memory — the fix may have landed since.
Umbrella RFC to know about: #11522 — "Refactor tool parsers to eliminate coding errors". Tracks the parse_delta refactor (PRs #39446, #39728, #38755). Any new parser work should align with it.
In-tree HACKs to recognize
Helpful to know what to look for when reading a parser. Grep the codebase for these:
grep -rn "HACK" vllm/tool_parsers/
grep -rn "TODO" vllm/tool_parsers/
grep -rn "prev_tool_call_arr = \[{\"arguments\": {}}\]" vllm/tool_parsers/
The prev_tool_call_arr = [{"arguments": {}}] plant is the classic "force finish_reason=tool_calls" workaround. At v0.27.0 exactly five parsers carry it: pythonic, llama4_pythonic, olmo3, lfm2, minicpm5xml — the pythonic family plus the two that reuse its flush shape. (mistral does not; it populates prev_tool_call_arr properly.) When writing a new parser, prefer the parse_delta shape (RFC #11522).
Writing a custom parser
Copy the existing parser closest in format. Don't reinvent.
# Match the target shape to an existing parser
grep -lE "<your_sentinel>" vllm/tool_parsers/ # sometimes the sentinel itself is in the code
ls vllm/tool_parsers/ # scan names by format
For the skeleton + detailed checklist see references/custom-parser-plugin.md. Key points:
- Register with
@ToolParserManager.register_module(["name"])(decorator path is lazy — plugin loader imports the file). - Launch:
--tool-parser-plugin /abs/path/file.py --tool-call-parser name. - Closest starting point for most models:
vllm/tool_parsers/pythonic_tool_parser.py(kwargs Python syntax),hermes_tool_parser.py(JSON-in-tags), orstep3p5_tool_parser.py(expat streaming XML —qwen3xml_tool_parser.pywas deleted in the unified-engine refactor). - Must honor the four state fields above. Read
vllm/tool_parsers/abstract_tool_parser.pyfor exact signatures. - Read
tests/tool_parsers/common_tests.py— reviewers expect this harness.
Before upstreaming, read AGENTS.md at the repo root for the duplicate-work + PR-description policy.
References in this skill
Compact supplementary maps. Each points back at the source rather than duplicating it.
references/parser-index.md— one-line-per-parser index with file path, wrapping tokens at a glance, and the things unique to that parser that aren't obvious from the filename (e.g. "full-width|U+FF5C sentinels", "non-streaming only", "schema-aware type coercion").references/streaming-pitfalls.md— the things that bite across parsers: full-width pipes, special-token stripping, empty-args stalls, HACKs, diagnostic flow. Points atchat_completion/serving.pyfor the flush contract.references/custom-parser-plugin.md— plugin scaffolding checklist with file:line anchors into the base class and the canonical examples.references/sources.md— verification log of the external GitHub issues/PRs/source files cited by this skill, each with aLast verifieddate. Consult before re-citing a claim; re-probe if the entry is stale (>90 days).
(Older references/json-family.md, pythonic-xml-family.md, known-bugs.md removed — they duplicated source and rotted. Use Grep/gh search instead.)
Last verified: 2026-08-11 against vLLM v0.27.0. See references/sources.md for per-reference probe details and timestamps.