Drafters and Speculative Decoding: Useful Tokens, Correct State
DFlash, DFlash 2 and DSpark: block proposals, target verification, retained state and the cost of a complete speculative cycle.
The target was 200 tokens per second from a 27B model on one Mac, on 5K–10K-token coding prompts. Faster kernels were one approach; getting more useful tokens from each expensive model pass was another. Five rounds of drafter training cost roughly $400 and produced no serving gain.
The promising numbers were real, but they described different parts of the problem. A candidate pool contained much better continuations than the selector could reliably choose. Extra draft rows were cheap to produce, but expensive to verify.
Several tokens from one target pass
Ordinary autoregressive decoding chooses the next token before it can use that token to calculate the following one. A token is a piece of text; the full model is the target.
A smaller model, the drafter, proposes a continuation in advance. The target can then evaluate those known inputs together. A round includes drafting, verification and the state updates needed to continue.
For the prefix “I saw,” the draft might be “the small blue bird.” The target checks “the” after “I saw,” then “small” after “I saw the,” and so on. Each row sees only its preceding prefix. Later proposed words stay masked out.
Draft. The drafter proposes four tokens after the committed prefix. These are candidates, not valid history.
Conceptual sequence. Playback timing is for explanation, not a benchmark. Use Step to read at your own pace.
In this example, “the small” survives and the target supplies “green” instead of “blue.” The rejected suffix disappears. Several useful tokens have come from one target pass, even though the entire proposal was not accepted.
The GPU can reuse weight tiles across those verification rows. For the affine Q4 format used here, calculated weight-only arithmetic intensity rises from 3.56 FLOPs/byte at one row to 28.44 at eight rows. More rows give the projections more work for each weight byte read.
The useful output still depends on how far the proposal survives. Efficiently computing rows that will be discarded does not make generation fast.
Better proposals need some dependence
Original DFlash drafts a masked block in parallel, using target hidden features as context. Positions can interact within that proposed block; the target's verification remains causal.
DFlash 2 adds candidate-pair scores and a lightweight sequential selector. A candidate at one position can then depend on the token actually chosen before it. Short dynamic convolutions also mix neighboring representations.
DSpark uses a parallel backbone followed by a small autoregressive Markov correction head. Both designs keep the expensive representation work parallel and spend a little sequential work on coherence.
Parallel computation with different dependencies between proposals.
Original DFlash
Positions interact within the draft block. Target verification remains causal.
DFlash 2
and candidate-pair scores
The candidate walk carries the dependence from one chosen token to the next.
DSpark
Markov correction head
Later corrections depend on earlier choices; the parallel backbone supplies context.
The selector did not capture the headroom
The candidate pool looked promising. A target-informed oracle could improve useful tokens per round by about 52.1%. It could see which candidates the target preferred, information a cheap serving selector did not have.
The selector-only pilot improved offline estimated progress by 5.68% on its training-side evaluation and just 0.07% on held-out prompts. It missed the 3% gate. The broader five-round training effort produced no serving gain for its roughly $400 cost.
A pool can contain the right answer without providing enough information to select it. The engineering problem was to recover that choice cheaply, at the prefixes the running engine actually visits. The oracle gave a useful target for investigation, but no deployable selector.
Cheap drafting, expensive verification
Increasing the native draft from eight to sixteen rows moved complete GPU cost from 3.007 to 3.442 ms. The paired increase was 0.438 ms.
The separate width experiment priced the target side. Sixteen-row verification added 6.63 ms at 2K after rest, 7.93 ms after soak, and 8.96 ms at 32K. Gate/up projections grew about 5%, while down/output projections grew about 31%. Attention and recurrent-state work added more.
The smaller cost on the draft side could not answer whether a wider round would pay off. More accepted tokens had to cover the complete verification and commit cost too.
Shortening blocks had the opposite problem. Proposals beyond the first three contributed about 0.845 useful tokens per cycle in the saved cohort. Removing them would lose roughly 23.6% of progress. The smaller graph would need comparable time savings merely to break even.
The rejected state matters as much as the rejected text
After the example round, committed history ends at “I saw the small.” The selected correction, “green,” can remain a pending anchor whose own state row is calculated next.
Attention must hide rejected key/value rows. The drafter must retain only accepted target features. Gated DeltaNet, the recurrent mixer, needs the recurrent matrix and convolution history after exactly the accepted prefix.
A length counter is not enough to repair a recurrent matrix already changed by rejected tokens. The engine can replay accepted updates from a saved state or retain intermediate checkpoints. A deferred commit can postpone the work, provided the next reader receives the correct state.
Sampling also needs the probability of the proposal that was actually made. Classical speculative sampling accepts with and uses a correction distribution after rejection. A selector that changes the proposal must change the reported accordingly. Named Block Verification uses its own joint acceptance and correction procedure.
Progress divided by complete round time
The rate is useful tokens divided by the time required to produce them. One native snapshot averaged 3.597 tokens per cycle at 40.631 ms, or 88.538 tok/s pooled.
At the same cycle duration, 150 tok/s would require about 6.095 useful tokens per round. Reaching the 200 tok/s goal would require 8.126. A conventional seven-proposal round can supply at most eight new tokens, so that target requires a faster cycle, a different route or both. Wide prompt lookup is one such route for repeated text.
Fusion can reduce work inside the drafter. Draft-ahead can overlap ready GPU work with host handling. Neither removes the need to verify proposals and commit the right state. The Pulsar results show which of those changes survived the complete serving loop, following the kernel experiments.
Methods and measurements
Experiment summary
The measurements exposed three gaps: between drafting and verification cost, between candidate coverage and usable selection, and between shorter blocks and useful output. Here means useful tokens per cycle. Each row below retains its own measurement boundary.
| Experiment | Result | Scope |
|---|---|---|
| Complete native draft, 8 versus 16 rows | 3.007 versus 3.442 ms; paired increase 0.438 ms [0.431, 0.445] | Fixed anchor after a real 5,120-token prefill; five draft layers, final norm, restricted 98,304-row head, selector and selection. No accepted-length or serving-throughput measurement. |
| Separate 16-versus-8 verification experiment | +6.63 ms at 2K rested; +7.93 ms at 2K after soak; +8.96 ms at 32K warm | Added cycle cost through the wide-lookup chain-verification route, not the native draft fixture or an implemented tree verifier. |
| Candidate-pool oracle versus trained selection | Target-informed oracle: about +52.1% G. Selector pilot: +5.68% offline in-sample, +0.07% held-out | The oracle has target information unavailable to a cheap selector. The pilot failed its 3% gate; neither result is candidate live throughput. |
| Shortening after three accepted proposals | Recorded survival tail: 0.845 useful tokens/cycle out of about 3.588 G; calculated loss: 23.6% | Unconditional shortening of this cohort. A selective policy needs a separate evaluation. |
The draft delta is the paired estimate, rather than subtraction of the rounded standalone times. Measurement supplement and source receipts.
From a proposed block to a causal target pass
With ordinary autoregressive decoding, the target chooses the next token before it can receive that token as input for the following step.
A cheaper drafter can supply a proposed continuation in advance. The target then evaluates that known continuation with causal masking and checks whether to retain it.
For example:
Real prefix: I saw
Draft: the small blue bird
The target evaluates the next-token distribution after the real prefix, after “I saw the,” after “I saw the small,” and after “I saw the small blue.” Those input tokens have already been proposed, so their target computations can be organized into one multi-position forward pass.
Each verification row remains causal: the distribution checking “small” sees “I saw the”; the later proposals “blue bird” are masked out.
The word fragments here are illustrative. Actual tokenization may split them differently.
The foundational algorithms show how to use a cheaper proposal while preserving the target distribution. A weak proposal can still be correct; it just wastes more verification work. Leviathan, Kalman and Matias, Chen and colleagues
Weight reuse across verification rows
An eight-row target pass contains one anchor and seven proposed positions. Its projections can reuse the same weight tile across those rows. Attention still needs correct masks, and recurrent layers still need valid state transitions, but the engine can avoid seven separate complete target passes.
For the affine Q4 group64 format in these experiments, 64 four-bit codes occupy 32 bytes, and the BF16 scale and bias add 4 bytes. That is 36 bytes per group, or byte per parameter. A projection with token rows performs roughly floating-point operations against weight bytes:
| Token rows | Calculated weight-only intensity |
|---|---|
| 1 | 3.56 FLOPs/byte |
| 8 | 28.44 FLOPs/byte |
| 16 | 56.89 FLOPs/byte |
| 32 | 113.78 FLOPs/byte |
These calculations describe the weight-reuse opportunity. They omit activation, state and KV traffic, register pressure, tile tails, dequantization instructions and recurrent dependencies. They are not hardware-counter readings. Format and calculation.
The gain depends on how many proposals survive. If only the first one is useful, work on the later rows may be discarded.
A wider target pass can therefore have excellent matrix efficiency and poor generation efficiency at the same time.
Greedy verification
In greedy decoding, each proposal is compared with the target's argmax at the corresponding prefix. The verifier retains the matching prefix, then uses the target's choice at the first mismatch and discards the remaining proposals. If every proposal matches, the target has a distribution for one additional token after the proposed block.
The greedy guarantee assumes the target computation and its state reproduce the intended target path. A batch-size-dependent numerical change can flip an argmax even when the surrounding acceptance logic is correct.
Sampling needs a correction rule
At nonzero temperature, matching the target's argmax would replace sampling with a different policy.
Let be the target's next-token probability and the proposal probability at the same prefix. Classical speculative sampling accepts a proposed token , drawn from q, with probability:
When the proposal is rejected, the replacement distribution is:
The target's full probability distribution matters. It is not sufficient to keep only the probability of the rejected token.
A three-token vocabulary makes the correction explicit:
| Token | Target p | Draft q | Acceptance probability | Probability mass accepted directly |
|---|---|---|---|---|
| A | 0.50 | 0.20 | 1 | 0.20 |
| B | 0.30 | 0.20 | 1 | 0.20 |
| C | 0.20 | 0.60 | 1/3 | 0.20 |
The total rejection probability is 0.40. The positive part of p−q is 0.30 for A, 0.10 for B, and zero for C. Normalizing gives replacement probabilities 0.75, 0.25 and zero.
The final probability of A is:
For B it is . For C it remains 0.20. The final distribution is p.
This also explains the overlap statistic:
Better overlap means fewer corrections for this one conditional distribution. It does not by itself predict the accepted length of a whole proposed sequence.
If p=q, rejection has probability zero, so the zero denominator of the residual expression is never needed. If q gives a token zero probability but p does not, the correction distribution can still supply that token.
Distribution, seeded output and quality
Different valid algorithms can consume random numbers in different orders. They can preserve the same output distribution while generating different text for a particular seed.
The validation needed depends on the promise:
| Promise | What must be checked |
|---|---|
| Same mathematical target distribution | Correct proposal, acceptance and correction procedure |
| Same greedy output | Same target argmax path and state |
| Same seeded sampled output | Matching random-number use as well as arithmetic and policy |
| Similar task quality | Representative task evaluation |
The probabilities p and q describe the policies actually served. Temperature, truncation, masks, penalties and selector transformations must be reflected in the relevant probabilities. A raw backbone softmax may no longer describe the drafter's proposal after a candidate selector changes it.
Joint Block Verification
Computing target rows together does not determine the acceptance algorithm.
The classical procedure above can use a multi-row target pass. The named Block Verification algorithm instead makes a joint decision about prefix acceptance and uses its own correction procedure. Its guarantee depends on that complete algorithm.
Pulsar, formerly fastkernel, includes a block-verification route for eligible sampled batches. The introductory p/q example explains classical speculative sampling; it is not pseudocode for that joint verifier. Mixing one method's acceptance decision with another method's residual draw is a correctness bug. Block Verification
DFlash, DFlash 2, and DSpark
DFlash, DFlash 2 and DSpark move most proposal computation into a parallel block. They differ in how they represent context and carry dependence between adjacent proposals.
DFlash: parallel masked-block drafting
Instead of asking a small autoregressive model to produce each proposal through a separate forward pass, DFlash uses a block diffusion drafter.
The input contains an anchor and masked future positions. Selected hidden representations from the target provide context. A small draft network predicts the block in one forward pass, with bidirectional interactions inside the draft block.
Bidirectional interactions improve the proposed block. The autoregressive target then checks that proposal with causal masking.
Original DFlash projects selected target features into a representation injected through draft-layer K/V paths. The exact layer taps, block length and attention window belong to a particular checkpoint configuration. DFlash paper, official implementation
DFlash 2: candidate-pair selection
Independent-looking choices at adjacent positions can form an incoherent continuation. A token that looks plausible at position three may be less plausible after the actual token chosen at position two.
DFlash 2 retains candidate tokens at each position, scores adjacent candidate pairs, and uses a lightweight sequential walk to choose or sample successors. The expensive block computation and pair scoring can remain parallel, while the final path selection carries the predecessor dependence.
It also adds short dynamic convolutions that mix neighboring draft representations.
This local walk differs from a global Viterbi search. Its conditional probability must match the actual predecessor chosen during the walk. DFlash 2 announcement
For the Qwen3.8-27B DFlash 2 checkpoint used in these experiments, the configuration has five draft layers, target taps [5, 19, 33, 47, 61], an eight-position block, a top-16 candidate pool, a rank-256 selector, two-tap convolutions and a 2,048-position sliding attention window. Those are checkpoint facts, not universal requirements of the method. Checkpoint configuration
DSpark: a parallel backbone with a small sequential head
DSpark also addresses dependence between neighboring proposals. Its expensive block backbone runs in parallel, then a comparatively small autoregressive Markov head conditions each proposed token on its predecessor.
The resulting division of work is:
expensive representation work across the block: parallel
inexpensive predecessor-dependent token correction: sequential
A confidence head can support verification-length decisions. That does not mean every implementation exposes the same adaptive scheduler. A fixed threshold in a reference evaluator and a production scheduler that accounts for serving capacity are different systems. DSpark paper, Markov-head implementation
DFlash 2 and DSpark both spend a little sequential work to make a parallel draft more coherent. Their heads, training choices and implementations differ. A result for one is not evidence that swapping its name or checkpoint into the other runtime will preserve speed or acceptance.
Candidate coverage and selector performance
A candidate pool can contain the target's preferred token without giving the serving selector enough information to choose it cheaply. The native headroom study's target-informed oracle improved estimated by about 52.1%, while the tested realizable eight-row selectors did not gain.
The subsequent selector-only training pilot scored saved stock-serving anchors. Its 5.68% in-sample improvement fell to 0.07% on held-out prompts, below the predeclared 3% gate. This was offline estimated useful progress, not a closed-loop candidate-throughput measurement. The failure belongs to that representation, objective and data experiment; it does not rule out other drafter-training methods.
Pool coverage, teacher-path estimates, marginal per-position overlap and emitted tokens in a live loop measure different things. The oracle identified headroom, but did not supply a selector capable of realizing it. Headroom and selector evidence.
Fusion within the drafter
A parallel drafter still contains projections, attention, normalization, a vocabulary head, selection and state updates. Each contributes to draft latency.
Original DFlash moves expensive proposal generation into a parallel masked-block forward pass. DFlash 2 and DSpark then add relatively cheap predecessor-dependent selection. That sequential selection is deliberate: it repairs a weakness of treating neighboring proposal positions independently.
The obvious fusion regions are therefore inside a draft layer: query/key/value (QKV) projection, attention preparation, attention and its reduction, output projection, residual addition and normalization. Keeping some intermediate values local can save memory traffic. A vocabulary head or a sequential Markov walk has a different shape and may deserve a separate kernel.
A useful draft fusion preserves three contracts:
- Proposal probability: the verifier needs the actual distribution that generated each candidate after the selector and sampling transforms. A faster kernel cannot silently keep reporting an older q.
- Context: only committed target features become reusable draft context. Rejected target rows must not leak into the next draft.
- Ordering: the next draft's anchor and retained context depend on the current verification result. Putting it on another queue or on the Neural Engine does not erase that dependency.
The draft-ahead path moved the next draft behind the current GPU submission, allowing it to run while the host handled the result. This reduced host waiting while preserving the dependency on the current verification.
Selection, context commit and handoff can absorb an isolated draft-layer saving. Lithos's DSpark refinement study illustrates this: its final isolated draft improvement did not establish a corresponding generation-throughput improvement on the small measured prompt set. DSpark refinement study
Committing state after rejection
In the running example, verification produces:
Existing prefix: I saw
Proposed block: the small blue bird
Accepted prefix: the small
Correction: green
Only “the small” survives. The verifier selects “green” at the next position, and the proposed suffix “blue bird” is discarded. The state commit retains the rows through “I saw the small”; “green” can remain a pending anchor whose own state row is computed next.
The target may already have executed calculations for the rejected suffix. The engine must distinguish computed scratch from committed history.
Attention exposes only the retained prefix as valid key/value (KV) history, with matching causal masks and logical positions. Rejected rows remain invalid for future cache reuse. The drafter likewise commits only retained target features and corresponding context; proposals from a different anchor, head mapping or random-number sequence cannot be reused.
A recurrent mixer such as Gated DeltaNet (GDN) needs the state after exactly the retained prefix. That includes convolution history as well as the recurrent matrix. Replaying the prefix requires a stable pre-update state.
A recurrent state cannot generally be repaired by decrementing a length counter. The rejected input has already changed the matrix.
One valid implementation stores enough information to replay the retained updates from the prior state. Another may keep intermediate checkpoints. A deferred implementation can postpone materializing the final recurrent state, provided every later reader first receives the logically committed state.
Cancellation, lane reuse, preemption, prefix caching and disk offload all become part of this contract. A request that ends may discard its pending state; a snapshot that will be reused must preserve it.
There is also a counting detail: an emitted correction or bonus token can become the next round's pending anchor. Its text may have been selected before the model has computed that token's own KV/state row. The emitted-token count can therefore differ from the number of materialized state rows.
Useful progress per complete cycle
The rate depends on two quantities:
- : useful output tokens produced per cycle, with the counting convention stated.
- : the complete cycle duration, including the work needed to make those tokens usable.
For a run, the basic rate is:
This is not generally the same as averaging the individual ratios . Likewise, a mean of per-prompt rates is not interchangeable with a pooled tokens-over-time rate.
A measured cycle and its conditional limits
One matched native status snapshot recorded 21,876 output tokens, 6,081 B1 cycles and 247,080.0927 ms of cycle time. Dividing those totals gives:
| Quantity | Calculated from the recorded totals |
|---|---|
| Useful output per cycle, | 3.597435 tokens |
| Cycle duration, | 40.631490 ms |
| Pooled native throughput | 88.538 tok/s |
The snapshot includes 24 benchmark requests and one warm request. The available totals do not allow exact subtraction of the warm request. The separately published 99.2 tok/s was a mean of per-prompt rates, so combining it with this pooled cycle duration would invent a retained length. Snapshot scope.
Holding cycle duration fixed at 40.631490 ms produces the following planning calculation:
| Desired throughput | Required useful tokens per cycle |
|---|---|
| 100 tok/s | 4.063 |
| 130 tok/s | 5.282 |
| 150 tok/s | 6.095 |
| 200 tok/s | 8.126 |
| 300 tok/s | 12.189 |
These are required-progress calculations, not achievable-rate forecasts. A conventional seven-proposal round can produce at most eight new tokens when every proposal survives and the extra target token is available, subject to anchor, stop and output-budget handling. With both that width and the measured cycle duration unchanged, the higher rows in the table require more progress than the route supplies.
Wider prompt lookup uses another route, and a faster cycle changes the calculation. An eight-row limit therefore says nothing universal about the maximum throughput of the Mac.
Tokens per step and time per step in a second runtime
Lithos's logs expose completion counts, generation steps and GPU decode time. The clean comparison's second block shows substantial variation in progress per step across workloads:
| Prompt | Completion tokens per logged step | Logged GPU ms per step | Client-observed tok/s |
|---|---|---|---|
| Code, magicoder-0 | 6.71 | 47.92 | 143.92 |
| Tip calculator | 5.72 | 53.00 | 106.21 |
| Math, numina-0 | 5.65 | 47.94 | 116.01 |
| Code file, r3-0 | 4.38 | 51.48 | 85.02 |
| Five chat prompts | 3.01–3.27 | 48.04–52.10 | 57.34–66.94 |
| Multilingual, aya-0 | 2.71 | 47.83 | 55.20 |
The code prompt produces substantially more tokens per logged step than chat, while GPU step times remain in a similar range. This makes proposal quality and round cost useful separate diagnostics. The columns have different boundaries: for the code prompt, dividing the unrounded token and GPU-time figures gives 139.95 tok/s, whereas the client recorded 143.92 tok/s. An exact throughput decomposition requires counts and time from the same interval.
The earlier Pulsar snapshot and these Lithos logs also come from different cohorts. They cannot establish that one runtime's round time caused the matched client-throughput difference. Per-request log calculations
The separate Lithos round-latency fixture measured nine samples at each context size:
| Context | Median GPU time | Median wall time | Proposals accepted / tokens committed |
|---|---|---|---|
| 128 | 48.04 ms | 48.85 ms | 7 / 8 |
| 32,000 | 64.53 ms | 65.24 ms | 6 / 7 |
The fixture includes acceptance, commitment and the next draft. Its retained counts differ between contexts, so it is not a fixed-work context-only comparison.
Wider drafting and wider verification
The opening draft fixture covered the complete native proposal path after a real 5,120-token prefill. Its extra rows were cheap compared with the added cost in the separate wide-lookup verification experiment. Neither fixture measured a deployed adaptive-drafting speedup; a tree verifier would also require correct parent-state and mask handling.
In the verification experiment, gate/up projections grew only about 5%, while down/output projections grew about 31%. GDN scan/commit and attention added further cost. Weight reuse helped some operations much more than others.
The complete width cost is:
with overlap accounted for from the actual timeline rather than by summing overlapping spans.
The cost of dropping the tail
In the saved survival cohort, the probability mass beyond three accepted proposals contributed approximately 0.304 + 0.230 + 0.175 + 0.136 = 0.845 useful tokens per cycle. Against a baseline around 3.588 G, removing that tail loses about 23.6% of useful progress. These are a different cohort from the native status snapshot above.
Unconditionally shortening the block would need a comparable reduction in cycle time just to break even. The saving depends on the complete smaller graph, including verification and commit. Survival-curve evidence.
A causal adaptive policy may still help some requests. The policy must use information available at the decision point and preserve the sampling algorithm's assumptions. Selecting a width after inspecting the sampled suffix that would be discarded can bias the proposal process.