# Original measurements: bandwidth, heat, cycle limits and drafter economics

Data supplement for the three-article series. Requested 2026-10-09.
Use this material prominently, not as a generic benchmarking footnote.

The contribution is our measured behavior, controls and failed hypotheses on
this implementation. “Original findings from our experiments” is supported.
“Nobody else measured this” or “world first” is not established by the evidence.

## Placement in the series

- **Megakernels article:** measured bandwidth reference, compulsory versus
  removable work, arithmetic intensity, dependency costs, negative controls.
- **Drafter article:** useful output per pass, complete draft cost versus verify
  cost, candidate-pool headroom versus realizable selection, accepted state.
- **Pulsar article:** hot versus warmed operating regimes, component gains
  versus serving gains, thermal/power observations, and actual positive results.

All local result paths below are relative to:

    /Users/abhishek/Desktop/meowkernels

## 1. The bandwidth we actually observed

The saved bw2 probe identifies an Apple M5 Max and a **4 GiB buffer**. It sweeps
three access patterns, four grid sizes and two threadgroup sizes. Each run
prints the best observed configuration:

| Probe run | Best reported rate | Winning pattern / geometry |
|---|---:|---|
| run 1 | 563.7 GB/s | chunk8, 5,242,880 threads, 256 threads/group |
| run 2 | 570.2 GB/s | stride4, 5,242,880 threads, 1,024 threads/group |
| run 3 | 572.5 GB/s | chunk8, 5,242,880 threads, 256 threads/group |

Sources: results/roofline/bw2/run1.txt, run2.txt, run3.txt.

Correct description: **roughly 570 GB/s in the best configurations of this
recorded streaming probe**. These are per-run maxima over a finite sweep,
not three random samples of one fixed configuration. Do not manufacture a
confidence interval from those three maxima.

The older results/roofline/alu.txt reports 512.9 GB/s with another probe over
8 GiB. The difference warns that probe/access geometry affects the result.
It does not establish that the hardware itself improved between runs.

### Critical warm-up qualification

The three bw2 text receipts do not themselves record a sustained warm-up
duration or a P-state/thermal plateau. Do not relabel them a fully validated
“warm maximum bandwidth” experiment without recovering the corresponding
harness and condition evidence.

Other experiments explicitly recorded representative GPU warm-up. For example,
results/megakernel-q4-rhs-pipeline/RESULTS.md reports 3.02329 accumulated GPU
seconds of warm-up, excluded from timing, followed by balanced comparisons and
A/A controls. That proves the stated preparation for that fixture, not that
three seconds is universally sufficient for every workload or thermal regime.

## 2. Warm, hot, cache-warm and memory-resident are different conditions

Use four distinct terms:

1. **Initialized:** pipelines, buffers and layouts have been prepared.
2. **GPU-warmed:** representative work has moved the device into a repeatable
   operating regime for the intended test.
3. **Thermally sustained:** enough work has run to expose the sustained power/
   cooling regime. A hot machine can run slower than a recently warmed one.
4. **Cache/residency state:** prompt-cache hits, file-backed pages, compressed
   memory and competing models can change the work or memory access cost.

No one label stands in for the others. In particular, a warm prompt cache can
remove prefill while the GPU is cold, and a long warm-up can eventually move a
machine into a lower sustained GPU clock regime.

### Our recorded thermal/power observations

| Observation | Evidence | What it establishes |
|---|---|---|
| Same stock code measured about 35.70 ms/cycle in a high-P-state sequential test; earlier readings were36.31/37.48/37.91 ms | results/serve-pstate/RESULTS.md | Operating conditions moved the result by an amount comparable to small optimizations; sequential order was not balanced. |
| A larger paired study associated one higher P-state with about 0.880 ms lower cycle time | Same file, 36 requests | Regression diagnostic, not an independently controlled frequency experiment. P-state IDs are not calibrated MHz. |
| External fan added during the October7 ab112 session; mean P-state subsequently climbed | results/speed-night/sampler.md, 20:2x entry | A recorded condition change. Do not call its full observed effect a code improvement. |
| Before-fan candidate/reference time ratio1.0006 over 36 pairs; with-fan0.9895 over 60 pairs | Same entry | Different subsets and times. This is not a randomized causal estimate of cooling. |
| pmset said nominal while IOReport said fair | results/speed-night/LOG.md and sampler.md | A coarse thermal label can miss a relevant operating-state difference. |

There is **no matched hot-versus-cool streaming-bandwidth curve established by
these specific receipts**. We measured cycle-time/P-state behavior, and we
measured a streaming bandwidth reference. Do not invent a numeric hot-memory
bandwidth value by combining them.

Lower GPU frequency can reduce effective kernel throughput without proving
that DRAM frequency or maximum memory-channel bandwidth changed by the same
percentage.

## 3. Our conditional memory arithmetic

For the October7 source pin c86ba286, the ordinary B1 projection-payload audit
reported:

| Payload | Logical bytes |
|---|---:|
| Target projections | 14.434467840 GB |
| Draft and context projections included in that route | 1.323417600 GB |
| Total | 15.757885440 GB |
| Target FFN gate/up/down subset | 9.625927680 GB |

Source: results/idea-scout/SPLASH130-1.6X.md, backed by that pinned model layout
and projection call graph. The mutable256-row head segment is already included
in the 98,304-row restricted draft head; do not add it twice.

Using570 GB/s as a **conditional reference rate**:

    15.757885440 GB / 570 GB/s = 27.6454 ms

This is the time to move that logical payload at that rate. It is not measured
DRAM traffic and is not an unconditional lower bound on every future design.
Actual traffic changes with cache reuse, selected graph, fallback, repeated
reads, discarded drafting, context length and model format.

A rigorous bandwidth lower bound needs a justified minimum number of actual
bytes crossing the limiting interface and a justified upper bound on its
sustained service rate. An achieved streaming measurement is not automatically
that universal upper bound.

Do not subtract27.65 ms from a measured 40 ms cycle and call the entire
remainder removable overhead. The remainder includes real computation,
other memory traffic, state work and dependencies; those contributions may
overlap rather than form a simple additive partition.

## 4. The arithmetic intensity of verification

For our affine Q4 group64 representation:

    64 four-bit codes = 32 bytes
    BF16 scale + BF16 bias = 4 bytes
    total = 36 bytes/group = 9/16 byte/parameter

For Y = X W-transpose, with M token rows, the projection performs roughly
2MNK floating-point operations and reads NK times9/16 weight bytes:

    weight-only arithmetic intensity = 32M/9 FLOPs per byte

| Rows | Weight-only intensity |
|---|---:|
| 1 | 3.56 FLOPs/byte |
| 8 | 28.44 FLOPs/byte |
| 16 | 56.89 FLOPs/byte |
| 32 | 113.78 FLOPs/byte |

These are calculations, not hardware-counter readings. They omit activation,
state and KV traffic and do not capture register pressure, tile tails,
dequantization instructions or recurrent dependencies.

The useful finding is that row reuse has a mathematical opportunity while the
complete graph can still lose. Our actual width-cost results below expose that
gap between a projection-level argument and a serving decision.

## 5. A tokens-per-cycle limit is conditional on the route

For a stable run:

    throughput = sum(useful emitted tokens) / sum(complete cycle time)

The matched native status snapshot for run-033951 contains 21,876 output tokens,
6,081 B1 cycles and 247,080.0927 ms cycle time:

- G =3.597435 tokens/cycle.
- T =40.631490 ms/cycle.
- Pooled native throughput =88.538 tok/s.

It includes 24 benchmark requests and one warm request; its warm time cannot
be subtracted exactly from the available status snapshot. The published99.2
tok/s was a mean of per-prompt rates, not this pooled native metric.
Do not multiply99.2 by 40.63 ms to invent a retained length.

At that fixed cycle time, the arithmetic required G would be:

| Desired rate | Required useful tokens/cycle |
|---|---:|
| 100 tok/s | 4.063 |
| 130 tok/s | 5.282 |
| 150 tok/s | 6.095 |
| 200 tok/s | 8.126 |
| 300 tok/s | 12.189 |

These are planning calculations, not achievable-rate forecasts.

A conventional seven-proposal round can produce at most eight new tokens
when all proposals survive and the extra target token is available, subject to
the runtime's anchor, stop and budget bookkeeping. Wider prompt lookup uses a
different route. It can therefore produce demo rates above a bound calculated
for the ordinary eight-row drafter path.

Never present “300 tok/s is impossible on this Mac” as a consequence of the
eight-row calculation. State the unchanged-cycle, unchanged-width assumptions.

## 6. What we measured about drafting that the headline misses

### Cheap extra draft rows can require expensive extra verification

The native draft fixture used a real prefilled 5,120-token context and a fixed
anchor, including five draft layers, final norm, the restricted 98,304-row head,
selector and selection:

| Draft | Complete GPU cost |
|---|---:|
| 8 rows | 3.007 ms |
| 16 rows | 3.442 ms |

Paired complete-draft delta: **+0.438 ms [0.431,0.445]**.
Source: results/native16/RESULTS.md.
No accepted length or serving throughput was measured in that fixture.

The separate October7 verification experiment found:

| Context / condition | Added16-versus 8 cycle cost |
|---|---:|
| 2K rested | 6.63 ms |
| 2K after soak | 7.93 ms |
| 32K warm | 8.96 ms |

Source: results/headroom/WIDTH.md. This prices chain verification through the
wide-lookup route; a real tree also needs correct parent-state/mask handling.

The gate/up projections grew only about 5%, but down/output projections grew
about 31%; GDN scan/commit and attention added more. Larger verification was
not uniformly free just because some matrix operations reused weights well.

### Pool coverage, selection and live progress are different quantities

The native headroom study reported a target-informed in-pool oracle improvement
of about 52.1% G, while tested realizable eight-row selectors did not gain.
The oracle uses target information unavailable to a cheap serving selector.

The selector-only training pilot then improved offline estimated G by 5.68%
in-sample and 0.07% on held-out prompts, failing its 3% gate. That is evidence
about transfer failure for this representation/objective/data experiment,
not proof that every future drafter training method will fail.

Sources: results/headroom/RESULTS.md and results/drafter-v5/RESULTS.md.

Do not turn a target-informed oracle, a teacher-path estimate, marginal
per-position overlap, and closed-loop emitted G into interchangeable numbers.

### The tail of the survival curve prices short blocks

The saved cohort's survival mass beyond three accepted proposals was
approximately 0.304+0.230+0.175+0.136 =0.845 useful tokens/cycle.
Removing that tail from a baseline around 3.588 loses about 23.6% G.

That makes unconditional shortening a demanding trade: the smaller graph must
save a comparable fraction of cycle time merely to break even. A selective
policy may differ; its decision needs causal inputs and independent evaluation.
Source: results/headroom/RESULTS.md and the October7 speed-profile notes.

## 7. A physics prediction that did transfer, at the right scale

Deferred GDN commit targeted one extra read of about 151 MB of recurrent state.
At 570 GB/s, that is a reference transfer time of approximately 0.265 ms.

Measured qualified block-mode changes:

- 2K: **-0.289 ms/cycle [-0.340,-0.209]**.
- 32K: **-0.230 ms/cycle [-0.243,-0.179]**.

The subsequent paired serving study, on top of no-store, measured raw
on/off decode-time ratio **0.9954 [0.9915,0.9993]** over 84 identical-text pairs.
This is about 0.46% lower decode time in that serving cohort, not the full
component saving pasted into a headline.

Sources: results/speed-night/LOG.md, blk-defer-2k-b,
blk-defer-32k-d and ab-defer.

This is a strong article example: name the redundant traffic, predict its
scale, preserve state correctness, measure the component, and measure serving.
It shows useful engineering without inventing a large gain.

## 8. What is genuinely distinctive about our evidence

Use these as original experiment narratives:

1. The measured gap between a speculative pool oracle and a realizable selector.
2. The separate native price of draft16 and target verification16.
3. The withdrawal of a plausible register explanation after a better control.
4. Identical-output component savings that shrink or disappear in real serving.
5. State-replay traffic predicted at hundreds of microseconds and confirmed at
   that scale, with a smaller full-serving gain.
6. A local competitor comparison whose more flattering initial results were
   withdrawn when residency conditions changed the conclusion.

The general principles are not claimed as inventions. The valuable material is
the implementation-specific data, controls, counterexamples and corrections.
No global novelty search establishes priority over all prior work.

## 9. Figures the writing agent can build from this data

- Three different limits on one diagram: advertised hardware rate, observed
  streaming probe, achieved complete-graph performance. Do not label them equal.
- Warm-up versus sustained heat versus prompt-cache warmth, with clearly
  separate axes/conditions.
- G/T tradeoff table: extra draft rows are cheap, extra verified/state rows are
  not necessarily cheap.
- A matched component-to-serving waterfall for GDN defer, with overlap and
  measurement scopes explicit.
- Pool oracle → realizable selector → offline held-out estimate → native
  serving, with no arrows implying guaranteed transfer.

Do not draw a numeric hot/cold bandwidth curve from missing data. If the author
wants that curve, it remains a separate controlled experiment, not a charting task.
