Closing the Prose Gap With a Mixture of Speculators
date: 19 Aug, 2026
I saw this tweet from Tom Lynch (of Actual Computer) getting Kimi K3 set up a few weeks ago.
And what really struck me was the massive gap between code and prose for decode. ~240 tok/s for code vs ~115 tok/s is a massive delta. I have a few theories for why there's such a big gap, the first one is just training, the DSPARK draft model was probably trained on a mix that favored code, because that's what people use these models for. The second explanation is that it's just a lot easier to predict code, there are tons of situations in code where the next token or several tokens are either fixed as part of a larger piece of syntax, or just trivial to predict, but that's not really a feature of prose.
// given this state
fn do_something_and_return_this_struct()// predict starting here
// this prediction is pretty much trivial
fn do_something_and_return_this_struct() -> this_struct {
// and that was 8 tokens(ish)! <space><arrow><space><this><_><struct><space><{>
// Tom's setup is only configured to predict 7 tokens per step,
// so this was basically free lunch for the DSPARK model!
There are way fewer situations like the above that present themselves in prose.
My Experiments
For economic reasons, for all of my experiments I'm using deepseek-v4-flash, with DSPARK, on a single B300 (Sidenote: dsv4-flash on a Modal B300 is an awesome way to free yourself from the clutches of closed models).
| domain | decode tok/s | accepted tokens |
|---|---|---|
| code | 399.9 | 6.29 |
| prose | 262.8 | 2.69 |
So, very similar situation, likely for the same reasons. We're also measuring accepted tokens per step, which is a much better measurement here for comparison between runs, especially across different machines.
Mixture of speculators
Alright, I'm going to assume you are familiar with the mixture of experts idea and what that generally entails. For this use case, we are going to finetune the DSPARK draft model on more prose, then we just need a routing system to send requests to either the normal draft model or the prose model. For this experiment I used a super simple heuristic based method that runs on the CPU in about ~50µs, we can err on the side of just sending things to the original drafter, as there won't be any negative effects from that, it'll just be the same as it was originally. However it's easy to imagine using an embedding model like Qwen-0.6b-embedding and then training a classifier if you wanted to handle several types of prompts with more fidelity. Really the best environment I can think of for this method is internal to a company, you can train on prompts from different parts of the company, and then routing is trivial based on who is sending it in.
Anyway, I finetuned the last layer of the DSPARK model on about 25 million tokens, which is uh, not very many, but GPU time is again, a bit expensive for me. Generating the training data here is pretty trivial, things like the answer to the prompt being correct don't particularly matter, the only thing that matters is predicting the output from the model, so you can just generate rollouts and save all the logprobs. You then train the DSPARK model on predicting those logprobs.
The actual implementation was extremely simple thanks to the MoE/routing primitives that are already present in vLLM, only about ~350 LoC to implement this. So how did it do? Here's a comparison on five prompts, all prose, of varying difficulty:
| prompt | base tok/s | finetune tok/s | base accepted | finetune accepted |
|---|---|---|---|---|
| p0 | 435.5 | 456.8 | 5.79 | 6.11 |
| p1 | 228.9 | 259.4 | 2.46 | 2.63 |
| p2 | 283.3 | 292.7 | 2.89 | 2.96 |
| p3 | 171.8 | 213.7 | 2.57 | 2.77 |
| p4 | 194.2 | 230.5 | 2.28 | 2.41 |
| avg | 262.8 | 290.6 | 3.20 | 3.38 |
So, small improvements, about ~.2 extra tokens accepted per step and ~30 tok/s. Nothing crazy, but I attribute this to lack of sufficient finetuning/training. 25M Tokens is basically nothing, lots of people use a billion or more every day, which is why this technique could actually be incredibly powerful. All you have to do is save your prompts, and you could train or finetune a speculator on your codebase, your workflows, and your business logic.
Free Lunch?
There's some free lunch here, or well, there was when I ran these experiments two weeks ago, vLLM just released the feature I'm about to talk about, theirs is better I'm sure. This idea being simply, when we know the speculator is much worse in this domain, can we just lower the speculation step size? Of course we can, without training a new model or anything I was able to lower it down to five tokens predicted per step. This pretty much just lets us cut down on the amount of wasted work. My version of this got again about ~30 tok/s improvement over not doing it on prose, but the vLLM implementation is much cooler/better overall.
Notes
If you have an inference setup running for your team/business, or even just mixed personal use, you should start recording all your prompts, not just for this technique. There are lots of ways in which that could be useful, finetune the actual base model on what you actually do! Or get free speedups from your speculator with basically zero downside, if it's just you, just finetune it on your prompts and don't even bother with the multiple dispatch system.
If you have this data already and want someone to come set this up for you or would otherwise be open to supporting further experimentation by me, let me know! Send me an email: [email protected]
Also I would just like to note that this is not an original idea, I came up with it on my own without prior knowledge of it, but researched current solutions before implementing something. Online Speculative Decoding [1] continuously finetunes draft models on the live query distribution, DistillSpec [2] is the "train the drafter on the target's logprobs" recipe I used above, and SpecDec++ [3] formalizes the adaptive speculation length trick from the Free Lunch section. And of course the original speculative decoding paper [4]. Though none of the above were implemented in vLLM or SGLang when I did this, and SpecDec++ is implemented in vLLM now.
Sources
[1] Online Speculative Decoding (Liu et al., 2023)
[2] DistillSpec: Improving Speculative Decoding via Knowledge Distillation (Zhou et al., 2023)
[3] SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths (Huang et al., 2024)
[4] Fast Inference from Transformers via Speculative Decoding (Leviathan et al., 2022)