vLLM reasoning parsers
Target: operators wiring up --reasoning-parser NAME on a chat-completion endpoint, or developers authoring a parser for a new thinking model. Source of truth: vllm/reasoning/ on main.
What a reasoning parser actually does
When a reasoning-trained model emits a single token stream like
<think>user asked X, let me check Y...</think>The answer is 42.
vLLM splits this into two fields on the chat-completion response: reasoning (the CoT) and content (the final answer). --reasoning-parser NAME selects the class that does the split. Without it, the whole stream lands in content.
Field-name note. On current
mainthe field isreasoning(seeChatMessage.reasoning/DeltaMessage.reasoninginvllm/entrypoints/openai/chat_completion/protocol.py). Pre-v0.19 code and many third-party docs / clients call itreasoning_content. If a client is readingreasoning_contentagainst a current-main server it will seenullevery time even when the parser ran correctly.
The parser is also the gating authority for xgrammar / structured output: by default, grammar enforcement is held off until is_reasoning_end(input_ids) flips true, so the model thinks freely before being constrained to JSON. Flip that default with --structured-outputs-config.enable_in_reasoning=true — then the grammar applies from token 0 regardless of reasoning state (useful for structured CoT).
The contract (ReasoningParser ABC)
vllm/reasoning/abs_reasoning_parsers.py. Every parser implements:
| Method | Called by | Purpose |
|---|---|---|
is_reasoning_end(input_ids) |
xgrammar, non-streaming serving, tool-call gate | Has </think>-equivalent been emitted yet? |
is_reasoning_end_streaming(input_ids, delta_ids) |
xgrammar per decode step | Same, but cheap — checks only the delta |
extract_content_ids(input_ids) -> list[int] |
structured output | Token IDs of post-reasoning content |
extract_reasoning(model_output, request) -> (reasoning, content) |
non-streaming chat completion | Full-string split; either field may be None |
extract_reasoning_streaming(previous_text, current_text, delta_text, previous_token_ids, current_token_ids, delta_token_ids) -> DeltaMessage | None |
streaming chat completion | Per-delta split; returns DeltaMessage(reasoning=..., content=...) or None to swallow |
Optional:
count_reasoning_tokens(token_ids)— for usage accounting. Default returns0.BaseThinkingReasoningParserimplements it via depth counter so nested<think>…<think>…</think>…</think>are handled.adjust_request(request)— mutate request (rare).prepare_structured_tag(original, tool_server)— emit astructural_tagJSON (GPT-OSS uses this for harmony).reasoning_start_str/reasoning_end_str— expose delimiter strings soReasoningConfig.initialize_token_idscan derive the multi-token ID sequences automatically.
Streaming vs. non-streaming are two independent code paths. A parser that's correct on extract_reasoning can be buggy on extract_reasoning_streaming and vice versa. Every parser must be tested on both.
The CLI path (what --reasoning-parser qwen3 actually triggers)
vllm/engine/arg_utils.py:985 (v0.27.1; grep for --reasoning-parser — line moves between releases) declares reasoning_parser: str = StructuredOutputsConfig.reasoning_parser. api_server.py validates it against ReasoningParserManager.list_registered() at startup (invalid name = fast-fail with the list of registered names).
On request, OpenAIServingChat instantiates a fresh parser per request via self.reasoning_parser_cls(tokenizer, chat_template_kwargs=chat_template_kwargs) (vllm/entrypoints/openai/chat_completion/serving.py:240). "Fresh per request" is load-bearing for stateful parsers — see Hunyuan in the matrix.
Registry: vllm/reasoning/__init__.py has a _REASONING_PARSERS_TO_REGISTER dict of name → (filename, ClassName) that feeds ReasoningParserManager.register_lazy_module. Lazy import means a broken parser file won't crash vLLM startup until somebody selects it.
Plugin path: --reasoning-parser-plugin /path/to/my_parser.py calls ReasoningParserManager.import_reasoning_parser(path), which importlib-loads the file. The file registers itself via @ReasoningParserManager.register_module(["my-name"]) at import time. Then --reasoning-parser my-name selects it.
The 15 things that go wrong
See references/pitfalls.md for each with repros and fixes. Quick index:
reasoning_contentisnullon DeepSeek-R1 — chat template injected<think>into the prompt, so the model never emitted a start token. Parser must tolerate missing start (baseBaseThinkingReasoningParserdoes, via the.partition(start_token)pattern).contentis empty, CoT inreasoning_contentwithenable_thinking=False— parser didn't branch onchat_template_kwargs. Qwen3 / DeepSeek-V3 / Kimi K2 route toIdentityReasoningParser(or internal flag) when thinking is off. DeepSeek-V3 has two thinking-default variants:deepseek_v3(default off) andholo2=DeepSeekV3ReasoningWithThinkingParser(default on).glm45used to be in that second group and no longer is — as of v0.25.1 it moved onto the new adapter path alongsideglm47, so don't assumeglm45andholo2still behave alike.Gibberish JSON when
guided_json+enable_thinking=False— the reasoning parser'sis_reasoning_endon the prompt must return True so xgrammar enforces from token 0. Serving layer caches this result asprompt_is_reasoning_end_arr[i]. If the parser only checks for</think>in input_ids and the thinking-disabled chat template emits<think>\n\n</think>\n\nin the prompt, this works; if it does something else, xgrammar silently stays gated.Truncated output = wrong field — when the model hits
max_tokensmid-reasoning, is the whole output "reasoning" or "content"? Qwen3 withenable_thinking=True(default) →(model_output, None)= all reasoning. Qwen3 withenable_thinking=False→(None, model_output)= all content. DeepSeek-R1 base →(model_output, None). Know the parser's policy before shipping.Nested
<think>tags — base parser'spartitionstops at the first</think>; if a CoT contains a literal</think>substring (sometimes seen in distillation artifacts), content gets split wrong.count_reasoning_tokensuses a depth counter so counts are right, but the split is not.Tool calls break with reasoning enabled — tool parser only sees the
contenthalf ofextract_reasoning's return. If the reasoning parser returns(everything, None), tool parser sees nothing. Kimi K2 handles this by treating<|tool_calls_section_begin|>as an implicit reasoning-end marker.Stateful parser reused across requests — Hunyuan A13B's
extract_reasoning_streamingis a token-ID state machine withself.current_state/self.token_buffer. Second concurrent request on the same instance = interleaved garbage. vLLM already instantiates per-request, but custom plugins must not hoist state to class-level.Multi-token delimiter —
<think>may encode as a single token (DeepSeek-R1 vocab) or as multiple (<think,>in some tokenizers). Single-token path:vocab.get("<think>"). Multi-token:tokenizer.encode("<think>")and do sequence match — the current in-tree worked example iskimi_k3_reasoning_parser.py, whose<|open|>think<|sep|>marker is a 3-token sequence, so its token-id helpers search for the subsequence rather than one id. (The old example, GPT-OSSreasoning_end_token_ids_prefix, was deleted at v0.27.0 — see pitfall 10.)BaseThinkingReasoningParserraises at init ifvocab.getreturns None — don't inherit from it unless the delimiter is a single vocab entry.Single-token delta spam — if delta is exactly one token that is
<think>or</think>, returnNonefrom streaming so the client doesn't see an empty delta. Almost every parser has this skip.Harmony / GPT-OSS is different — no
<think>tag; the channel markers are multi-token. At v0.27.0 (PR #45560)GptOssReasoningParserwas gutted:is_reasoning_end/is_reasoning_end_streamingnowreturn Trueunconditionally, andextract_reasoning,extract_reasoning_streamingandextract_content_idsall raiseNotImplementedErrorpointing atHarmonyParser. The old backward-scan machinery (reasoning_end_token_ids_prefix,reasoning_max_num_between_tokens = 20,eom_token_id) is deleted — don't look for it, and don't cite it as the multi-token-delimiter example. Boundary detection plus all parsing now live inHarmonyParser(vllm/parser/harmony.py). Don't copy DeepSeek-R1 for a harmony model.Mistral moved onto the unified engine at v0.27.0 (PR #48947).
--reasoning-parser mistralis unchanged, but the registry class is nowMistralParserReasoningAdapterandvllm/reasoning/mistral_reasoning_parser.pyis a three-line re-export — the logic is invllm/parser/mistral.py, shared with--tool-call-parser mistral. The hardValueError: The tokenizer must be an instance of MistralTokenizer.is gone (it existed at v0.25.1;git grepfinds it nowhere at v0.27.0):is_mistral_tokenizer()now selects a path, falling back to the genericadjust_requestwhen the tokenizer isn't Mistral or can't do grammar. Still prefer--tokenizer-mode mistral— grammar-enforcedtool_choiceneeds the mistral-common path.Granite is regex-on-text — "Here is my thought process:" / "Here is my response:" are phrases, not special tokens. Streaming parser has to buffer partial matches across deltas (
Here is my thou…) which makes it the most complex parser in the tree. Readgranite_reasoning_parser.py:140+before modifying.Implicit reasoning-end — Kimi K2 ends at
</think>or<|tool_calls_section_begin|>. MiniMax M2 never emits<think>, only</think>. Customis_reasoning_endmust encode these facts or xgrammar gates at the wrong moment.--enable-reasoningis gone — older docs / Stack Overflow answers still reference it. Since roughly v0.8 the only flag is--reasoning-parser NAME; the enable/disable is implicit in whether one is passed.reasoning_contentis always null — but parser worked fine. Current-main response field isreasoning, notreasoning_content(renamed inprotocol.py). Client-side name mismatch that looks exactly like a parser failure. Before debugging parsers,jq '.choices[0].message | keys'to see what fields actually exist — ifreasoningis there, it's just a client rename.
The per-model matrix
references/parser-matrix.md — one row per registered name (29 at v0.27.0: deepseek_r1, deepseek_v3, deepseek_v4, poolside_v1, cohere_command3, cohere_command4, ernie45, gemma4, glm45, glm47, openai_gptoss, granite, holo2, hunyuan_a13b, hy_v3, inkling, kimi_k2, kimi_k3, mimo, minimax_m2, minimax_m2_append_think, minimax_m3, mistral, nemotron_v3, olmo3, qwen3, seed_oss, step3, step3p5) with: delimiter style, start-token-in-prompt-or-output, thinking-disable mechanism, truncation policy, structured-output gating peculiarities.
Routing (which family each name belongs to — <think> two-token, delegating wrapper, stateful, harmony, phrase-regex, tokenizer-gated) lives in the matrix Family column. The non-obvious cases worth knowing before reading it: openai_gptoss is harmony (extract_reasoning raises NotImplementedError), mistral requires MistralTokenizer, hunyuan_a13b is a token-ID state machine, granite is phrase-regex on text, and nemotron_v3 swaps reasoning↔content on enable_thinking=False.
Writing a custom parser
references/writing-custom-parser.md for the step-by-step. This
ReasoningParser subclass shape remains the supported path for out-of-tree
parsers — it is what --reasoning-parser-plugin loads, and 17 of the 29
in-tree names still use it.
Be aware there is now a second, in-tree-only shape: vllm/parser/<model>.py
defines one parser class per model, and
make_adapters(XParser) in vllm/parser/engine/registered_adapters.py derives
both XParserReasoningAdapter and XParserToolAdapter from it — the unified
reasoning+tool design of RFC
#32713. 12 of the 29
names are on it at v0.27.0: deepseek_v4, gemma4, glm45, glm47,
inkling, kimi_k2, mimo, minimax_m2, mistral, nemotron_v3, qwen3,
seed_oss. If you open one of those as a worked example you will find a
three-line re-export and the logic elsewhere.
Do not infer membership from the filename. Only some are named
*_engine_reasoning_parser.py; kimi_k2_reasoning_parser.py,
minimax_m2_reasoning_parser.py, glm47_moe_reasoning_parser.py and
mistral_reasoning_parser.py are ordinary names that are nonetheless shims. The
test is the import:
grep -l "registered_adapters import" vllm/reasoning/*.py.
Copy a legacy-path parser when writing your own. Shape:
from vllm.reasoning import ReasoningParser, ReasoningParserManager
from vllm.entrypoints.openai.engine.protocol import DeltaMessage
@ReasoningParserManager.register_module(["my_model"])
class MyReasoningParser(ReasoningParser):
def __init__(self, tokenizer, *args, **kwargs):
super().__init__(tokenizer, *args, **kwargs)
# resolve delimiter token IDs via self.vocab or tokenizer.encode(...)
def is_reasoning_end(self, input_ids): ...
def extract_content_ids(self, input_ids): ...
def extract_reasoning(self, model_output, request): ...
def extract_reasoning_streaming(self, prev_text, cur_text, delta_text,
prev_ids, cur_ids, delta_ids): ...
Register decorator → lazy entry in ReasoningParserManager.lazy_parsers. Pass the file to vllm serve ... --reasoning-parser-plugin /path/to/my_parser.py --reasoning-parser my_model.
Shortcuts when delimiter is a single vocab token: inherit BaseThinkingReasoningParser (vllm/reasoning/basic_parsers.py), override start_token / end_token properties only. That covers 90% of the surface — see deepseek_r1_reasoning_parser.py (32 lines) for the minimal concrete subclass.
Testing a parser
Unit tests live in tests/reasoning/ (DeepSeek R1, Qwen3, Granite, Hunyuan, GPT-OSS, Kimi K2 all have coverage). Run a single file with:
.venv/bin/python -m pytest tests/reasoning/test_qwen3_reasoning_parser.py -v
Every new parser needs tests for:
- Non-streaming: start-in-prompt, start-in-output, truncated (no end), nested, empty.
- Streaming: delta that is only the start token, only the end token, end token spanning two deltas, content-before-end-token in same delta, content-after-end-token in same delta.
is_reasoning_endon both pre-reasoning, mid-reasoning, and post-reasoning token ID sequences.- Thinking-disabled path if applicable.
Companion skills
vllm-chat-templates— how the chat template injects<think>into the prompt (or not); whyenable_thinking=Falsecan change what the parser sees.vllm-tool-parsers— tool-call extraction runs on the post-reasoningcontentonly. A wrong reasoning split silently breaks tool calling.vllm-configuration—chat_template_kwargsflows throughChatCompletionRequest.chat_template_kwargsinto the parser__init__.vllm-performance-tuning— structured output tuning; reasoning parser choice affects xgrammar gating latency.vllm-input-modalities— reasoning parsers don't apply to embedding / rerank / ASR endpoints.
Upstream issue/PR anchors for each pitfall → references/sources.md.