vLLM MTP: Faster Inference That Quietly Broke Tool Calls
Multi-token prediction cut our latency, then quietly corrupted tool calls. What MTP does, how the failure presented, and the settings that made it safe again.
Filed under Local LLM Inference· see every report on this topic
Verdict. Ship the reliability lane before the throughput lane if your local LLM must make tool calls. On 18 August 2026, vLLM 0.24.0, Qwen3.8-27B BF16, a 262K-context lane, and MTP=3 delivered a real speed win. On 20 August, concurrent long-context tool work became repeated
!reasoning, empty assistant messages, andfinish_reason=lengthbefore any tool call. We removed MTP from the production launcher. Speed that burns an agent’s retry budget is operational debt.
This is one dated failure class on one Blackwell workstation, not a claim that speculative decoding is broken, MTP is always unsafe, or local inference is a bad buy. The workstation is an already-paid-for 96GB RTX PRO 6000 serving private-context Hermes, GBrain, and Kanban workers. The narrower decision: when does faster decode stop paying for its reliability work?
The operator and the leak: agents need tools, not just tokens
The operator is a founder running a small agent stack on one local lane. The model reads a long worker prompt, selects a tool, emits valid arguments, and continues from its result. If the first tool call does not happen, the workflow has stopped; it has not merely slowed down.
On 18 August, eight non-thinking clients completed 6,180 tokens in 19.553 seconds: 316.056 aggregate completion tokens/s. All eight ended normally. The scheduler peak was eight running, zero waiting; MTP accepted 3,595 of 7,758 draft tokens (46.34%); peak temperature was 72C; the report marked all_ok: true. (Source: /home/kit/projects/vllm-qwen38-27b-bf16/reports/heavy-load-8-20260818-185540.json: total_completion_tokens, wall_elapsed_s, aggregate_completion_tok_s, peak_running, peak_waiting, metrics_delta, peak_temperature_c.)
That result did not describe every workload. The same day, eight thinking-mode clients each exhausted max_tokens=900, returned no visible content, and ended with length. They still reached 212.126 tokens/s, 52.32% MTP acceptance, and 90C. Hidden reasoning consumed the whole 900-token allowance: an output-budget failure, not proof of model corruption. (Source: /home/kit/projects/vllm-qwen38-27b-bf16/reports/heavy-load-8-20260818-185758.json: results[*].completion_tokens, content_chars, finish_reason, aggregate_completion_tok_s, metrics_delta.mtp_acceptance_pct, peak_temperature_c.)
A local lane keeps private context on-box and makes failures inspectable. It also gives the operator the pager: output caps, retry policy, queue policy, thermal policy, rollback time, and the obligation not to call a fast benchmark a production pass.
What MTP changes, and what it does not
Multi-token prediction (MTP) is speculative decoding here. A draft path proposes several future tokens; vLLM verifies them against the target model. Accepted drafts move decoding faster, so acceptance is meaningful only for the workload that will run.
The 18 August direct API baseline made MTP look worthwhile:
| Case | Observed result | MTP detail | Source |
|---|---|---|---|
| Draft acceptance | 2,304 draft; 1,152 accepted | Exactly 50.0% | metrics_delta |
| Short non-thinking | 493 tokens in 8.962s | 55.012 end-to-end tok/s | short_nonthinking |
| Medium reasoning | 339 tokens in 4.661s | 72.732 tok/s | reasoning_medium |
| Cold long prefill | Correct needle from 60,040-token prompt in 16.161s | Cold result | long_prefill_first |
| Warm prefix repeat | Identical cached prompt in 0.564s | Prefix reuse | long_prefill_prefix_cache_repeat |
(Source: /home/kit/projects/vllm-qwen38-27b-bf16/reports/raw-vllm-benchmark-20260818-184336.json, metrics_delta, and cases named short_nonthinking, reasoning_medium, long_prefill_first, and long_prefill_prefix_cache_repeat.) Short decode, cold prefill, and warmed prefix reuse are separate measurements, not one speed number.
Capacity had a boundary before the later incident. With max_num_seqs=1, two nominally simultaneous 536-token clients serialized at 10.404 and 20.822 seconds. After the 18 August change to max_num_seqs=8, telemetry still put maximum KV-cache concurrency at 1.69 at 262K context. Eight admits short or normal batching; it does not make eight full-context sessions physically real. (Source: raw report notes, concurrent_2/concurrent_3; /Users/kit/.hermes/kit-log.md, 18 August 2026 lines 71–75.)
For the engine, model, and capacity trade-offs, see the Blackwell field guide. This report covers the reversal: a lane can win tokens/s and fail the job.
Choose the lane by the work that can fail
| Approach | Good for | Cost shape | Catch |
|---|---|---|---|
| MTP=3 throughput lane | Short non-thinking batches | GPU time; 316.056 aggregate tok/s at c8 | Later degenerated on concurrent long-context tools |
| MTP-off reliability lane | Always-on agent tool loops | Owned GPU; lower speed unmeasured after removal | Production lane after 20 Aug |
| No speculative decoding | Baseline and fault isolation | GPU time; measure it | Do not call it slower before testing |
| Hosted API | Buying operations away | Requests/tokens, context, retries, uptime | Get dated pricing for your traffic |
| Do nothing first | Output-cap failures | $0 config and policy work | Fix token budget before engines |
The eight thinking-client run supports the $0 route. Before changing models or paying an API bill, raise or bound client output policy, select thinking deliberately, and log the finish reason.
A hosted comparison needs its own receipt: (input tokens × input rate) + (output tokens × output rate) + retry/continuation tokens + any fixed service charges. Include the prompt sizes and tool-call count your agents use. This article has no dated equivalent API-pricing receipt and makes no dollar break-even claim. What the Lab costs explains the decision framing: do the comparison; do not invent a monthly total.
The failure: ! consumed the retry budget
This was not the 900-token output-budget trap. On 20 August, six Kanban workers sent concurrent long-context tool requests. All six reached finish_reason=length before their first tool call. Stored assistant text was empty; recorded reasoning was repeated !, token ID 0. One task made four attempts and produced 229,376 reasoning characters. Five peers made three attempts and produced 163,840 reasoning characters each. (Source: /Users/kit/.hermes/kit-log.md, 20 August 2026, 12:40 IST, lines 46–50.)
A hard task can reason for a long time. A response almost entirely composed of one punctuation token cannot make tool progress, and a continuation prompt does not repair it. Stop the request, retain the record, and quarantine the lane or workload. Do not turn a broken stream into a more expensive broken stream.
Hermes amplified the blast radius. Its finish_reason=length path appended a continuation prompt and raised requested output from 16,384 to 32,768, eventually as high as 114,688 tokens per failed turn. A stale Discord gateway session separately auto-continued for about 38 minutes at 100% GPU, roughly 300W, and 89C. These are incident observations, not energy-cost math. (Source: /Users/kit/.hermes/kit-log.md, 20 August 2026, 12:18 and 12:25 IST, lines 36–39 and 52–56.)
After restart, the exact 44,334-character worker system prompt plus task made correct kanban_context calls with thinking enabled and disabled. Prompt wording was not a sufficient explanation. This does not isolate a CUDA kernel or prove a universal root cause. The supported label is inference degeneration under vLLM 0.24.0 / Qwen3.8-27B / MTP=3 / concurrent long-context-tool workload. (Source: /Users/kit/.hermes/kit-log.md, 20 August 2026, 12:40 IST, lines 46–49.)
vLLM issue #35800 gives outside corroboration, not a diagnosis: it reports MTP plus Qwen tool-call degradation on Blackwell across fresh sessions until restart. That supports suspicion and rollback; it does not show that every local ! loop has the same cause.
The rollback that restored tool calls
At 13:32 IST, with explicit owner approval, we:
- Removed
--speculative-configfrom the launcher. - Ran
bash -n. - Kept remote backup
serve-qwen38-27b.sh.bak-20260820-132128-mtp-enabled. - Verified the deployed SHA-256 matched.
- Restarted only the
qwen38-vllmsupervisor program.
The verified runtime command had no speculative configuration. A bounded tool smoke returned finish_reason=tool_calls, 71 reasoning characters, and valid get_status {}. The settled lane had zero running/queued, 16W, and 62C. (Source: /Users/kit/.hermes/kit-log.md, 20 August 2026, 13:32 IST, lines 27–30.)
One recovery detail matters: supervisorctl clear qwen38-vllm recreated an unlinked supervisor output log after restart. It did not restart the service. Keep repair actions distinct in the incident record. (Source: /Users/kit/.hermes/kit-log.md, 20 August 2026, 13:32 IST, line 30.)
When not to use speculative decoding, or local inference
Skip MTP for the tool lane until it survives long system prompts, concurrent tool requests, retries, cancellations, and real client parsing. A clean curl response is not enough.
Skip a local always-on lane when nobody owns logs, output limits, a kill switch, and rollback. A hosted API may be the better purchase when an availability contract matters more than private context and diagnostic control. That is a cost-shape decision, not a provider-price claim.
Fix these $0 controls before changing the serving engine:
- Set a per-attempt output ceiling.
- Do not auto-continue a one-token repetition signature.
- Bound agent dispatch to proven concurrency, not configured admission count.
- Record tool-call success and first-tool-call time; aggregate tokens/s cannot replace either.
Bottom line: tool success is the meter
MTP=3 earned a throughput claim here: 316.056 aggregate completion tokens/s at eight non-thinking clients, with 46.34% MTP acceptance. It did not earn a place in the always-on tool lane after the 20 August repeated-! degeneration and retry amplification.
Run local when private context and operational control justify owning the lane. Start with reliability, fix output-budget and retry-policy traps for $0, and bring speculative decoding back only after it survives real tool work. Keep MTP on a tool lane only if an observation period shows zero degeneration events, first-tool-call success at least as good as the MTP-off baseline, throughput within a tolerance you declare before the run starts, and no auto-continuation loop past a cap you set ahead of time. Otherwise keep it in a bounded throughput lane. A model that cannot reach its first tool call is not fast.
FAQs
Does speculative decoding hurt tool-call reliability?
It can in a specific lane. On 20 August 2026, vLLM 0.24.0 with Qwen3.8-27B BF16, MTP=3, and concurrent long-context tool work degenerated into repeated ! reasoning before any tool call. The same prompt worked after a restart with MTP removed. That is evidence about this configuration, not a general verdict on speculative decoding.
How do I benchmark a vLLM lane honestly?
Separate short decode, cold long-context prefill, warm-prefix reuse, real-concurrency traffic, and tool-call success. Record output caps, finish reasons, acceptance, queue state, and thermals. A throughput number without those controls cannot tell you whether the lane is safe for agents.
When is a local 96GB lane cheaper than API calls?
Only after you price the actual alternative: requests or tokens, prompt and output context, retries, and always-on agent volume. This workstation was already paid for and kept private context on-box; this article does not claim a universal API-versus-GPU break-even price.
What should an agent do after finish_reason=length?
Do not blindly auto-continue a response whose reasoning is almost entirely one repeated punctuation token. Stop it, preserve the request record, and treat it as an incident. In this case, length retries enlarged the requested output cap from 16,384 to 32,768 and eventually as high as 114,688 tokens per failed turn.
More on this decision, three ways to look at it:
Sources
/home/kit/projects/vllm-qwen38-27b-bf16/reports/heavy-load-8-20260818-185540.json, observed eight-client non-thinking load result, 18 August 2026./home/kit/projects/vllm-qwen38-27b-bf16/reports/heavy-load-8-20260818-185758.json, observed eight-client thinking output-budget result, 18 August 2026./home/kit/projects/vllm-qwen38-27b-bf16/reports/raw-vllm-benchmark-20260818-184336.json, observed direct-API baseline, prefill, cache, and serialization results, 18 August 2026./Users/kit/.hermes/kit-log.md, dated incident, remediation, and verification entries cited inline (18 and 20 August 2026).- vLLM issue #35800, upstream Blackwell/Qwen/MTP corroboration; not proof of this incident’s exact root cause.
Get the next verdict before it's everywhere.
One email when a new lab post or cost table ships. No spam, no confirmation step — unsubscribe anytime.