We run a share of our work on our own : a single desktop-class GPU box with 128 GB of , serving through and to coding agents that read, edit and review real repositories for an hour at a time.

The numbers that are easy to see in that setup are per second, the word "lossless" on a speedup, and a capability probe that passes. On this site we call those extensions: measurable stand-ins for the thing we want. The thing we want, the , is an agent finishing real work. Local inference is full of knobs whose failure doesn't raise an error. The stand-in stays green while the work quietly gets worse.

Most of what first looked to us like a weak model turned out to be one of those knobs. The figures below come from one box, often with two or three runs per arm, so read them as evidence rather than . The lessons apply on cloud GPUs too, because they live in the serving configuration.

The context you set isn't the context you get

We considered halving our main server's to get a second concurrent . Before doing it, we measured. Of about 1,600 real agent requests, 18.6% were longer than 131,072 tokens (128K). The 90th percentile was 151K and the longest was 209K. Agent transcripts only grow, so long sessions spend most of their time past the point where a 128K window would have cut them.

At our 256K window there were zero truncations. At 128K, roughly one request in five would have been cut, and in our setup that raises no error. The agent doesn't fail. It forgets the start of its own session and carries on, and that reads as a model getting worse partway through a task.

We have direct receipts on a small 7B model served at an 8,192-token context. Its tool-loop prompt alone was about 5.4K tokens on the first turn, and the server log showed truncation. An evaluation of that model scored 2/9 at 8,192 and 4/9 at 16,384. The first number was false. It measured our window, not the model.

There is a second layer. The agent client keeps its own context limit and output reserve, and the server can't see them. In one case the server was fine at 32,768 tokens, but the client was set to a 16,384 context with 8,192 reserved for output. After the system prompt and tool definitions, about 8K was left for actual work, and the model looked unusable. Whichever limit is smaller wins, and nothing tells you which one that is.

The chat template is part of the model

A is the small program that turns messages and tool definitions into the text format the model was trained on. It ships inside the model file, and it can be wrong.

A newly released coding model shipped with a stale template. The template expected <thought> tags and one tool-call format. The weights emitted <think> and a different one. Served as published, the reasoning field came back empty, tool calls leaked into the reply as raw text, and the agent ended its run after one turn having done nothing. Serving the vendor's own template file fixed reasoning and tool parsing together, and the same weights then did about 40 minutes of real multi-turn work.

On vLLM the equivalent piece is the tool and reasoning parser. With the right ones, a model completed a 19-page review walk with 38 tool calls. Without them, the identical run printed its first tool call as text and exited in 32 seconds. With no tool parser configured at all, vLLM refused every request that carried tools with HTTP 400, and our probe scored 0/8.

A strict template can fail in the other direction. One embedded template raised an error whenever a system message appeared anywhere but first. One version of an agent client appends a system message after the user turn, so every request it sent got HTTP 500, while an older version of the same client worked against the same server. A minimal override that renders the system message inline fixed it with no measured cost: the probe went to 8/8 and was 85.78 against a same-day band of 84.18–88.16. Whether the model actually honors a system message in the middle of a conversation is something we have not tested.

No client setting can fix a template. When a model "can't tool-call" or "isn't supported by the engine," we now check the template first and the engine last.

Thinking budgets

Reasoning models spend tokens thinking before they answer, and llama.cpp lets you cap that with a . The cap only works if the server is also splitting thinking from the answer with a reasoning parser. Without one, the budget does nothing and says nothing. We once concluded the budget was inert because we searched the log for a message the current build no longer writes. The guard was firing; our check was wrong. A guard you can't observe firing may not be firing, and the check itself can lie.

When the budget does fire, where you set it matters. Our client sent max_tokens of 16,384, and we had set the reasoning budget to 16,384 too. The server cut the thinking and hit the output cap on the same token: zero answer tokens, and the agent run died on a length stop. Healthy thinking on that task was under about 2K tokens, so we reset the budget to 6,144. The budget has to sit well below whatever output cap the client sends.

The budget also doesn't remove thinking so much as move it. On a 35B model at 0, we cut the budget from 256 to 128, 64 and 32 tokens. Thinking shrank from 784 to 140 characters, but the answer grew from 1,339 at the 128 budget to 2,786 at 32, as the model did its reasoning in the answer instead. Accuracy held at 12/12 all the way down to a 32-token budget. What broke it was capping total output: 1/3 correct at 512 tokens and 0/3 at 256 or less, each ending in a visible length stop.

Turning thinking off is not a safe shortcut either. With thinking disabled, a small mixture-of-experts model passed our 8/8 synthetic probe, then issued ten identical ls -F calls on the first turn of a real task and had to be killed.

The opposite failure is having no budget at all. vLLM parses thinking but, in the version we ran, does not cap it. A reasoning model there fell into an unbounded loop of empty thinking blocks mid-task, generating at 66 t/s until we aborted it. Its reasoning parser had logged that it failed to initialize. Our unattended and overnight agent runs now go to llama.cpp only. Engine choice is a safety choice too.

Speculative decoding: lossless, not free

has a guess the next few tokens and the full model check them all in one pass. Accepted guesses are kept and the first wrong one is replaced, so the output is exactly what the full model would have written. Only the speed changes. It works because single-stream decode is limited by memory bandwidth, which leaves compute to spare for checking several tokens at once.

Multi-token prediction () uses extra prediction heads trained into the model itself as the drafter. uses a small, separately trained drafter model that proposes a block of up to 16 tokens per step. It can guess further ahead, but only if it matches the target model.

On a 35B mixture-of-experts coding model, MTP took decode from 67 to 95.6 t/s, about 1.43×, for identical output. That is the largest free speedup we have found, and the knob with the most ways to fail silently.

The right didn't transfer even between two generations of the same model, with the same architecture and the same built-in heads:

Draft depthOlder generation (t/s)Newer generation (t/s)
None66.2168.46
180.4181.16
287.0571.57
389.3061.87

At depth 3 the newer model was slower than with no speculation at all. went from 0.83 to 0.65 across the depths on one and collapsed from 0.82 to 0.37 on the other. Copying the old setting across would have cost 26%.

The heads can also be present and dead. One model file shipped only part of the MTP block, and the engine discarded it with four unremarkable log lines and no change in speed. Loading the standalone head file took a dense 27B model from 12.55 to 28.63 t/s on a short prompt and from 9.93 to 20.62 at 63K tokens, 2.1 to 2.3×. A model file's name tells you nothing about whether its heads work. The engine's own log does.

DFlash showed the matching problem plainly. A drafter released for its target gave 9.57 to 22.80 t/s, about 2.4×. A community drafter paired with a different 35B model lost: 9.1% slower on a short prompt and 27.0% slower at 65K, with 15 tokens drafted for 3 to 5 accepted. We think the cause is the mismatch between drafter and target, but we haven't proven it.

When comparing drafters, the number to read is acceptance length, not acceptance rate. Single-token MTP accepted 84.9–87.7% of its guesses for a mean of 1.85–1.88 tokens per step. A block drafter accepted only 67–75%, but for 3.02–3.26 tokens per step, and won by 34–58%.

Finally, "lossless" describes the output, not the cost. On a slot-based llama.cpp server with a 30B mixture-of-experts model, speculation spends compute that concurrent requests need:

Concurrent requestsSpeculation off (t/s)MTP on (t/s)Effect
164.469.1+7%
499.582.4−17%
16140.896.0−32%

The same table hides a second lesson. Going from 1 to 16 concurrent requests bought about 2.2× the aggregate tokens. On a slot-based server, more slots mostly means admission, not throughput. Several agent sessions sharing one slot also defeat the : in one test five sessions on one slot missed the cache 40 times out of 40 and never finished, while one session alone ran 99% cached.

Benchmarks at the wrong depth

The two generations in the draft-depth table benched within 1.2% of each other on a short prompt: 87.17 against 89.30 t/s. On a real 18-page agent review, the older one finished in 14m50s and the newer in 25m45s. The older model was 42% faster while taking more turns, 205 against 128. That review lives around 148K tokens of context, and that is where the two models' drafting heads diverge.

Decode speed itself decays with depth. The same server ran 87.17 t/s on a short prompt, 55.99 at 65K and 35.6 at about 230K in live use. Draft acceptance falls with depth too, from 0.838 to 0.776 for the first drafted token between a short prompt and 63K. A short benchmark measures a model at the depth agents spend almost no time at. We now bench at the depth we serve at.

The box itself has to be settled before a benchmark is worth reading. Right after unloading a 31 GB model, we measured 53.23 t/s against 61.24 for the identical configuration 90 seconds later, 13% apart. The two early runs agreed with each other to two decimal places, which made them convincing, and they produced two confident wrong theories before we found the cause. Agreement between runs tells you the measurement is repeatable. It doesn't tell you it's right.

Vendor recipes are hypotheses

We treat a model card's recommended settings as a starting point to test. Several didn't survive agent work.

One card's headline temperature was 1.0. On a review task that produced three repeated headings, 7 duplicate findings out of 75, and zero file-and-line citations. Temperature 0.5, which the same card listed further down as its tool-calling setting, produced no repetition and 19 citations. Speed was not the trade-off: the higher temperature cost nothing in t/s.

A community fix for repetition, a of 1.08 over the last 4,096 tokens, made things worse in a way we didn't expect. Agent output repeats legitimately: paths, headings, a report rewritten in place. The penalty punished the model for re-emitting its own file, so its writes shrank. It also cut draft acceptance from 0.73 to 0.42–0.49, because the drafter doesn't apply the penalty and now disagreed with the model it was drafting for.

The largest factor wasn't a model setting at all. The same local model on the same task under four agent took 46 minutes and committed, 68 minutes and committed, 93 minutes without a commit, and in the fourth case made four attempts and delivered nothing. That is about 2× on the same model before counting the harness that failed. Within one harness, re-injecting the project's conventions on every page instead of only the first moved review coverage from 11 to 18 of 20 pages and citations from 19 to 103. No sampler setting we tried moved results that much.

What we check now

What a quiet failure looks like

None of these failures raised an error we would have seen in normal use. A truncated context, a stale template, an inert budget, a dead set of heads and a benchmark run on an unsettled box all return a healthy-looking number. Each one looked, from the outside, like a model that wasn't very good.

That is the trap at the scale of a serving stack. Tokens per second, a passing probe and "lossless" are worth measuring, but they are not the goal. The goal is an agent that finishes the work, so that is what we measure last, and what we trust first.