Fast inference on Mac

Drafters and Speculative Decoding: Useful Tokens, Correct State

DFlash, DFlash 2 and DSpark: block proposals, target verification, retained state and the cost of a complete speculative cycle.

The target was 200 tokens per second from a 27B model on one Mac, on 5K–10K-token coding prompts. Faster kernels were one approach; getting more useful tokens from each expensive model pass was another. Five rounds of drafter training cost roughly $400 and produced no serving gain.

The promising numbers were real, but they described different parts of the problem. A candidate pool contained much better continuations than the selector could reliably choose. Extra draft rows were cheap to produce, but expensive to verify.

Tokens per round against round time: Lithos workloads vary from about 2.7 to 6.7 tokens per logged step, with a separate Pulsar pooled snapshot near 3.6 tokens and 40.6 milliseconds. Tokens per round against round time: Lithos workloads vary from about 2.7 to 6.7 tokens per logged step, with a separate Pulsar pooled snapshot near 3.6 tokens and 40.6 milliseconds.
Lithos per-request diagnostics and a separate Pulsar pooled snapshot. Different cohorts and timing boundaries; the chart shows the two levers, not a matched causal comparison.

Several tokens from one target pass

Ordinary autoregressive decoding chooses the next token before it can use that token to calculate the following one. A token is a piece of text; the full model is the target.

A smaller model, the drafter, proposes a continuation in advance. The target can then evaluate those known inputs together. A round includes drafting, verification and the state updates needed to continue.

For the prefix “I saw,” the draft might be “the small blue bird.” The target checks “the” after “I saw,” then “small” after “I saw the,” and so on. Each row sees only its preceding prefix. Later proposed words stay masked out.

EXPLORE THE MECHANISMProposed tokens are not committed history.
1Draft2Verify3Accept prefix4Discard suffix5Commit state
Candidate block4 proposed · illustrative rejection at token 3
01theproposal
02smallproposal
03blueproposal
04birdproposal
One target blockCausal rows to evaluate
p1the given “I saw”
p2small given “I saw the”
p3blue given “I saw the small”
p4bird given “I saw the small blue”
Each row sees only its preceding prefix. These are four rows in one verification pass.
Valid retained stateKV · draft context · GDN · convolution
I sawunchanged
Pending anchor—Correction will be selected after rejection.

Draft. The drafter proposes four tokens after the committed prefix. These are candidates, not valid history.

1 / 5

Conceptual sequence. Playback timing is for explanation, not a benchmark. Use Step to read at your own pace.

In this example, “the small” survives and the target supplies “green” instead of “blue.” The rejected suffix disappears. Several useful tokens have come from one target pass, even though the entire proposal was not accepted.

The GPU can reuse weight tiles across those verification rows. For the affine Q4 format used here, calculated weight-only arithmetic intensity rises from 3.56 FLOPs/byte at one row to 28.44 at eight rows. More rows give the projections more work for each weight byte read.

The useful output still depends on how far the proposal survives. Efficiently computing rows that will be discarded does not make generation fast.

Better proposals need some dependence

Original DFlash drafts a masked block in parallel, using target hidden features as context. Positions can interact within that proposed block; the target's verification remains causal.

DFlash 2 adds candidate-pair scores and a lightweight sequential selector. A candidate at one position can then depend on the token actually chosen before it. Short dynamic convolutions also mix neighboring representations.

DSpark uses a parallel backbone followed by a small autoregressive Markov correction head. Both designs keep the expensive representation work parallel and spend a little sequential work on coherence.

DRAFT ARCHITECTURESThree approaches to drafting a block

Parallel computation with different dependencies between proposals.

Original DFlash

Target hidden features
Parallel masked block
Proposals at each position

Positions interact within the draft block. Target verification remains causal.

DFlash 2

Target hidden features
Parallel candidate pool
and candidate-pair scores
Lightweight sequential selector

The candidate walk carries the dependence from one chosen token to the next.

DSpark

Parallel draft backbone
Autoregressive
Markov correction head
Corrected proposals

Later corrections depend on earlier choices; the parallel backbone supplies context.

High-level dataflow, not layer counts or measured timings. All proposals still require target verification.

The selector did not capture the headroom

The candidate pool looked promising. A target-informed oracle could improve useful tokens per round by about 52.1%. It could see which candidates the target preferred, information a cheap serving selector did not have.

The selector-only pilot improved offline estimated progress by 5.68% on its training-side evaluation and just 0.07% on held-out prompts. It missed the 3% gate. The broader five-round training effort produced no serving gain for its roughly $400 cost.

A pool can contain the right answer without providing enough information to select it. The engineering problem was to recover that choice cheaply, at the prefixes the running engine actually visits. The oracle gave a useful target for investigation, but no deployable selector.

Cheap drafting, expensive verification

Increasing the native draft from eight to sixteen rows moved complete GPU cost from 3.007 to 3.442 ms. The paired increase was 0.438 ms.

The separate width experiment priced the target side. Sixteen-row verification added 6.63 ms at 2K after rest, 7.93 ms after soak, and 8.96 ms at 32K. Gate/up projections grew about 5%, while down/output projections grew about 31%. Attention and recurrent-state work added more.

The smaller cost on the draft side could not answer whether a wider round would pay off. More accepted tokens had to cover the complete verification and commit cost too.

Shortening blocks had the opposite problem. Proposals beyond the first three contributed about 0.845 useful tokens per cycle in the saved cohort. Removing them would lose roughly 23.6% of progress. The smaller graph would need comparable time savings merely to break even.

The rejected state matters as much as the rejected text

After the example round, committed history ends at “I saw the small.” The selected correction, “green,” can remain a pending anchor whose own state row is calculated next.

Attention must hide rejected key/value rows. The drafter must retain only accepted target features. Gated DeltaNet, the recurrent mixer, needs the recurrent matrix and convolution history after exactly the accepted prefix.

A length counter is not enough to repair a recurrent matrix already changed by rejected tokens. The engine can replay accepted updates from a saved state or retain intermediate checkpoints. A deferred commit can postpone the work, provided the next reader receives the correct state.

Sampling also needs the probability of the proposal that was actually made. Classical speculative sampling accepts with min⁡(1,p/q)\min(1,p/q) and uses a correction distribution after rejection. A selector that changes the proposal must change the reported qq accordingly. Named Block Verification uses its own joint acceptance and correction procedure.

Progress divided by complete round time

The rate is useful tokens divided by the time required to produce them. One native snapshot averaged 3.597 tokens per cycle at 40.631 ms, or 88.538 tok/s pooled.

At the same cycle duration, 150 tok/s would require about 6.095 useful tokens per round. Reaching the 200 tok/s goal would require 8.126. A conventional seven-proposal round can supply at most eight new tokens, so that target requires a faster cycle, a different route or both. Wide prompt lookup is one such route for repeated text.

Fusion can reduce work inside the drafter. Draft-ahead can overlap ready GPU work with host handling. Neither removes the need to verify proposals and commit the right state. The Pulsar results show which of those changes survived the complete serving loop, following the kernel experiments.

Methods and measurements

Experiment summary

The measurements exposed three gaps: between drafting and verification cost, between candidate coverage and usable selection, and between shorter blocks and useful output. Here GG means useful tokens per cycle. Each row below retains its own measurement boundary.

ExperimentResultScope
Complete native draft, 8 versus 16 rows3.007 versus 3.442 ms; paired increase 0.438 ms [0.431, 0.445]Fixed anchor after a real 5,120-token prefill; five draft layers, final norm, restricted 98,304-row head, selector and selection. No accepted-length or serving-throughput measurement.
Separate 16-versus-8 verification experiment+6.63 ms at 2K rested; +7.93 ms at 2K after soak; +8.96 ms at 32K warmAdded cycle cost through the wide-lookup chain-verification route, not the native draft fixture or an implemented tree verifier.
Candidate-pool oracle versus trained selectionTarget-informed oracle: about +52.1% G. Selector pilot: +5.68% offline in-sample, +0.07% held-outThe oracle has target information unavailable to a cheap selector. The pilot failed its 3% gate; neither result is candidate live throughput.
Shortening after three accepted proposalsRecorded survival tail: 0.845 useful tokens/cycle out of about 3.588 G; calculated loss: 23.6%Unconditional shortening of this cohort. A selective policy needs a separate evaluation.

The draft delta is the paired estimate, rather than subtraction of the rounded standalone times. Measurement supplement and source receipts.

From a proposed block to a causal target pass

With ordinary autoregressive decoding, the target chooses the next token before it can receive that token as input for the following step.

A cheaper drafter can supply a proposed continuation in advance. The target then evaluates that known continuation with causal masking and checks whether to retain it.

For example:

Real prefix: I saw
Draft:       the small blue bird

The target evaluates the next-token distribution after the real prefix, after “I saw the,” after “I saw the small,” and after “I saw the small blue.” Those input tokens have already been proposed, so their target computations can be organized into one multi-position forward pass.

Each verification row remains causal: the distribution checking “small” sees “I saw the”; the later proposals “blue bird” are masked out.

The word fragments here are illustrative. Actual tokenization may split them differently.

The foundational algorithms show how to use a cheaper proposal while preserving the target distribution. A weak proposal can still be correct; it just wastes more verification work. Leviathan, Kalman and Matias, Chen and colleagues

Weight reuse across verification rows

An eight-row target pass contains one anchor and seven proposed positions. Its projections can reuse the same weight tile across those rows. Attention still needs correct masks, and recurrent layers still need valid state transitions, but the engine can avoid seven separate complete target passes.

For the affine Q4 group64 format in these experiments, 64 four-bit codes occupy 32 bytes, and the BF16 scale and bias add 4 bytes. That is 36 bytes per group, or 9/169/16 byte per parameter. A projection Y=XWTY=XW^T with MM token rows performs roughly 2MNK2MNK floating-point operations against NK(9/16)NK(9/16) weight bytes:

weight-only arithmetic intensity=32M9 FLOPs/byte.\text{weight-only arithmetic intensity}=\frac{32M}{9}\text{ FLOPs/byte}.
Token rowsCalculated weight-only intensity
13.56 FLOPs/byte
828.44 FLOPs/byte
1656.89 FLOPs/byte
32113.78 FLOPs/byte

These calculations describe the weight-reuse opportunity. They omit activation, state and KV traffic, register pressure, tile tails, dequantization instructions and recurrent dependencies. They are not hardware-counter readings. Format and calculation.

The gain depends on how many proposals survive. If only the first one is useful, work on the later rows may be discarded.

A wider target pass can therefore have excellent matrix efficiency and poor generation efficiency at the same time.

Greedy verification

In greedy decoding, each proposal is compared with the target's argmax at the corresponding prefix. The verifier retains the matching prefix, then uses the target's choice at the first mismatch and discards the remaining proposals. If every proposal matches, the target has a distribution for one additional token after the proposed block.

The greedy guarantee assumes the target computation and its state reproduce the intended target path. A batch-size-dependent numerical change can flip an argmax even when the surrounding acceptance logic is correct.

Sampling needs a correction rule

At nonzero temperature, matching the target's argmax would replace sampling with a different policy.

Let p(x)p(x) be the target's next-token probability and q(x)q(x) the proposal probability at the same prefix. Classical speculative sampling accepts a proposed token xx, drawn from q, with probability:

α(x)=min⁡(1,p(x)q(x)).\alpha(x)=\min\left(1,\frac{p(x)}{q(x)}\right).

When the proposal is rejected, the replacement distribution is:

r(x)=max⁡(p(x)−q(x),0)∑ymax⁡(p(y)−q(y),0).r(x)= \frac{\max(p(x)-q(x),0)} {\sum_y\max(p(y)-q(y),0)}.

The target's full probability distribution matters. It is not sufficient to keep only the probability of the rejected token.

A three-token vocabulary makes the correction explicit:

TokenTarget pDraft qAcceptance probabilityProbability mass accepted directly
A0.500.2010.20
B0.300.2010.20
C0.200.601/30.20

The total rejection probability is 0.40. The positive part of p−q is 0.30 for A, 0.10 for B, and zero for C. Normalizing gives replacement probabilities 0.75, 0.25 and zero.

The final probability of A is:

0.20+0.40(0.75)=0.50.0.20+0.40(0.75)=0.50.

For B it is 0.20+0.40(0.25)=0.300.20+0.40(0.25)=0.30. For C it remains 0.20. The final distribution is p.

This also explains the overlap statistic:

A=∑xmin⁡(p(x),q(x))=1−12∑x∣p(x)−q(x)∣.A=\sum_x\min(p(x),q(x)) =1-\frac{1}{2}\sum_x|p(x)-q(x)|.

Better overlap means fewer corrections for this one conditional distribution. It does not by itself predict the accepted length of a whole proposed sequence.

If p=q, rejection has probability zero, so the zero denominator of the residual expression is never needed. If q gives a token zero probability but p does not, the correction distribution can still supply that token.

Distribution, seeded output and quality

Different valid algorithms can consume random numbers in different orders. They can preserve the same output distribution while generating different text for a particular seed.

The validation needed depends on the promise:

PromiseWhat must be checked
Same mathematical target distributionCorrect proposal, acceptance and correction procedure
Same greedy outputSame target argmax path and state
Same seeded sampled outputMatching random-number use as well as arithmetic and policy
Similar task qualityRepresentative task evaluation

The probabilities p and q describe the policies actually served. Temperature, truncation, masks, penalties and selector transformations must be reflected in the relevant probabilities. A raw backbone softmax may no longer describe the drafter's proposal after a candidate selector changes it.

Joint Block Verification

Computing target rows together does not determine the acceptance algorithm.

The classical procedure above can use a multi-row target pass. The named Block Verification algorithm instead makes a joint decision about prefix acceptance and uses its own correction procedure. Its guarantee depends on that complete algorithm.

Pulsar, formerly fastkernel, includes a block-verification route for eligible sampled batches. The introductory p/q example explains classical speculative sampling; it is not pseudocode for that joint verifier. Mixing one method's acceptance decision with another method's residual draw is a correctness bug. Block Verification

DFlash, DFlash 2, and DSpark

DFlash, DFlash 2 and DSpark move most proposal computation into a parallel block. They differ in how they represent context and carry dependence between adjacent proposals.

DFlash: parallel masked-block drafting

Instead of asking a small autoregressive model to produce each proposal through a separate forward pass, DFlash uses a block diffusion drafter.

The input contains an anchor and masked future positions. Selected hidden representations from the target provide context. A small draft network predicts the block in one forward pass, with bidirectional interactions inside the draft block.

Bidirectional interactions improve the proposed block. The autoregressive target then checks that proposal with causal masking.

Original DFlash projects selected target features into a representation injected through draft-layer K/V paths. The exact layer taps, block length and attention window belong to a particular checkpoint configuration. DFlash paper, official implementation

DFlash 2: candidate-pair selection

Independent-looking choices at adjacent positions can form an incoherent continuation. A token that looks plausible at position three may be less plausible after the actual token chosen at position two.

DFlash 2 retains candidate tokens at each position, scores adjacent candidate pairs, and uses a lightweight sequential walk to choose or sample successors. The expensive block computation and pair scoring can remain parallel, while the final path selection carries the predecessor dependence.

It also adds short dynamic convolutions that mix neighboring draft representations.

This local walk differs from a global Viterbi search. Its conditional probability must match the actual predecessor chosen during the walk. DFlash 2 announcement

For the Qwen3.8-27B DFlash 2 checkpoint used in these experiments, the configuration has five draft layers, target taps [5, 19, 33, 47, 61], an eight-position block, a top-16 candidate pool, a rank-256 selector, two-tap convolutions and a 2,048-position sliding attention window. Those are checkpoint facts, not universal requirements of the method. Checkpoint configuration

DSpark: a parallel backbone with a small sequential head

DSpark also addresses dependence between neighboring proposals. Its expensive block backbone runs in parallel, then a comparatively small autoregressive Markov head conditions each proposed token on its predecessor.

The resulting division of work is:

expensive representation work across the block: parallel
inexpensive predecessor-dependent token correction: sequential

A confidence head can support verification-length decisions. That does not mean every implementation exposes the same adaptive scheduler. A fixed threshold in a reference evaluator and a production scheduler that accounts for serving capacity are different systems. DSpark paper, Markov-head implementation

DFlash 2 and DSpark both spend a little sequential work to make a parallel draft more coherent. Their heads, training choices and implementations differ. A result for one is not evidence that swapping its name or checkpoint into the other runtime will preserve speed or acceptance.

Candidate coverage and selector performance

A candidate pool can contain the target's preferred token without giving the serving selector enough information to choose it cheaply. The native headroom study's target-informed oracle improved estimated GG by about 52.1%, while the tested realizable eight-row selectors did not gain.

The subsequent selector-only training pilot scored saved stock-serving anchors. Its 5.68% in-sample improvement fell to 0.07% on held-out prompts, below the predeclared 3% gate. This was offline estimated useful progress, not a closed-loop candidate-throughput measurement. The failure belongs to that representation, objective and data experiment; it does not rule out other drafter-training methods.

Pool coverage, teacher-path estimates, marginal per-position overlap and emitted tokens in a live loop measure different things. The oracle identified headroom, but did not supply a selector capable of realizing it. Headroom and selector evidence.

Fusion within the drafter

A parallel drafter still contains projections, attention, normalization, a vocabulary head, selection and state updates. Each contributes to draft latency.

Original DFlash moves expensive proposal generation into a parallel masked-block forward pass. DFlash 2 and DSpark then add relatively cheap predecessor-dependent selection. That sequential selection is deliberate: it repairs a weakness of treating neighboring proposal positions independently.

The obvious fusion regions are therefore inside a draft layer: query/key/value (QKV) projection, attention preparation, attention and its reduction, output projection, residual addition and normalization. Keeping some intermediate values local can save memory traffic. A vocabulary head or a sequential Markov walk has a different shape and may deserve a separate kernel.

A useful draft fusion preserves three contracts:

  • Proposal probability: the verifier needs the actual distribution that generated each candidate after the selector and sampling transforms. A faster kernel cannot silently keep reporting an older q.
  • Context: only committed target features become reusable draft context. Rejected target rows must not leak into the next draft.
  • Ordering: the next draft's anchor and retained context depend on the current verification result. Putting it on another queue or on the Neural Engine does not erase that dependency.

The draft-ahead path moved the next draft behind the current GPU submission, allowing it to run while the host handled the result. This reduced host waiting while preserving the dependency on the current verification.

Selection, context commit and handoff can absorb an isolated draft-layer saving. Lithos's DSpark refinement study illustrates this: its final isolated draft improvement did not establish a corresponding generation-throughput improvement on the small measured prompt set. DSpark refinement study

Committing state after rejection

In the running example, verification produces:

Existing prefix: I saw
Proposed block:  the small blue bird
Accepted prefix: the small
Correction:      green

Only “the small” survives. The verifier selects “green” at the next position, and the proposed suffix “blue bird” is discarded. The state commit retains the rows through “I saw the small”; “green” can remain a pending anchor whose own state row is computed next.

The target may already have executed calculations for the rejected suffix. The engine must distinguish computed scratch from committed history.

Attention exposes only the retained prefix as valid key/value (KV) history, with matching causal masks and logical positions. Rejected rows remain invalid for future cache reuse. The drafter likewise commits only retained target features and corresponding context; proposals from a different anchor, head mapping or random-number sequence cannot be reused.

A recurrent mixer such as Gated DeltaNet (GDN) needs the state after exactly the retained prefix. That includes convolution history as well as the recurrent matrix. Replaying the prefix requires a stable pre-update state.

A recurrent state cannot generally be repaired by decrementing a length counter. The rejected input has already changed the matrix.

One valid implementation stores enough information to replay the retained updates from the prior state. Another may keep intermediate checkpoints. A deferred implementation can postpone materializing the final recurrent state, provided every later reader first receives the logically committed state.

Cancellation, lane reuse, preemption, prefix caching and disk offload all become part of this contract. A request that ends may discard its pending state; a snapshot that will be reused must preserve it.

There is also a counting detail: an emitted correction or bonus token can become the next round's pending anchor. Its text may have been selected before the model has computed that token's own KV/state row. The emitted-token count can therefore differ from the number of materialized state rows.

Useful progress per complete cycle

The rate depends on two quantities:

  • GG: useful output tokens produced per cycle, with the counting convention stated.
  • TT: the complete cycle duration, including the work needed to make those tokens usable.

For a run, the basic rate is:

throughput=∑iGi∑iTi.\text{throughput}= \frac{\sum_i G_i}{\sum_i T_i}.

This is not generally the same as averaging the individual ratios Gi/TiG_i/T_i. Likewise, a mean of per-prompt rates is not interchangeable with a pooled tokens-over-time rate.

A measured cycle and its conditional limits

One matched native status snapshot recorded 21,876 output tokens, 6,081 B1 cycles and 247,080.0927 ms of cycle time. Dividing those totals gives:

QuantityCalculated from the recorded totals
Useful output per cycle, GG3.597435 tokens
Cycle duration, TT40.631490 ms
Pooled native throughput88.538 tok/s

The snapshot includes 24 benchmark requests and one warm request. The available totals do not allow exact subtraction of the warm request. The separately published 99.2 tok/s was a mean of per-prompt rates, so combining it with this pooled cycle duration would invent a retained length. Snapshot scope.

Holding cycle duration fixed at 40.631490 ms produces the following planning calculation:

Desired throughputRequired useful tokens per cycle
100 tok/s4.063
130 tok/s5.282
150 tok/s6.095
200 tok/s8.126
300 tok/s12.189

These are required-progress calculations, not achievable-rate forecasts. A conventional seven-proposal round can produce at most eight new tokens when every proposal survives and the extra target token is available, subject to anchor, stop and output-budget handling. With both that width and the measured cycle duration unchanged, the higher rows in the table require more progress than the route supplies.

Wider prompt lookup uses another route, and a faster cycle changes the calculation. An eight-row limit therefore says nothing universal about the maximum throughput of the Mac.

Tokens per step and time per step in a second runtime

Lithos's logs expose completion counts, generation steps and GPU decode time. The clean comparison's second block shows substantial variation in progress per step across workloads:

PromptCompletion tokens per logged stepLogged GPU ms per stepClient-observed tok/s
Code, magicoder-06.7147.92143.92
Tip calculator5.7253.00106.21
Math, numina-05.6547.94116.01
Code file, r3-04.3851.4885.02
Five chat prompts3.01–3.2748.04–52.1057.34–66.94
Multilingual, aya-02.7147.8355.20

The code prompt produces substantially more tokens per logged step than chat, while GPU step times remain in a similar range. This makes proposal quality and round cost useful separate diagnostics. The columns have different boundaries: for the code prompt, dividing the unrounded token and GPU-time figures gives 139.95 tok/s, whereas the client recorded 143.92 tok/s. An exact throughput decomposition requires counts and time from the same interval.

The earlier Pulsar snapshot and these Lithos logs also come from different cohorts. They cannot establish that one runtime's round time caused the matched client-throughput difference. Per-request log calculations

The separate Lithos round-latency fixture measured nine samples at each context size:

ContextMedian GPU timeMedian wall timeProposals accepted / tokens committed
12848.04 ms48.85 ms7 / 8
32,00064.53 ms65.24 ms6 / 7

The fixture includes acceptance, commitment and the next draft. Its retained counts differ between contexts, so it is not a fixed-work context-only comparison.

Wider drafting and wider verification

The opening draft fixture covered the complete native proposal path after a real 5,120-token prefill. Its extra rows were cheap compared with the added cost in the separate wide-lookup verification experiment. Neither fixture measured a deployed adaptive-drafting speedup; a tree verifier would also require correct parent-state and mask handling.

In the verification experiment, gate/up projections grew only about 5%, while down/output projections grew about 31%. GDN scan/commit and attention added further cost. Weight reuse helped some operations much more than others.

The complete width cost is:

Tcycle=Tdraft+Tverify+Taccept/commit+Tunhidden host work,T_\text{cycle} =T_\text{draft} +T_\text{verify} +T_\text{accept/commit} +T_\text{unhidden host work},

with overlap accounted for from the actual timeline rather than by summing overlapping spans.

The cost of dropping the tail

In the saved survival cohort, the probability mass beyond three accepted proposals contributed approximately 0.304 + 0.230 + 0.175 + 0.136 = 0.845 useful tokens per cycle. Against a baseline around 3.588 G, removing that tail loses about 23.6% of useful progress. These are a different cohort from the native status snapshot above.

Unconditionally shortening the block would need a comparable reduction in cycle time just to break even. The saving depends on the complete smaller graph, including verification and commit. Survival-curve evidence.

A causal adaptive policy may still help some requests. The policy must use information available at the decision point and preserve the sampling algorithm's assumptions. Selecting a width after inspecting the sampled suffix that would be discarded can bias the proposal process.