# Additional experiment receipts and interpretation

Reviewed 2026-10-09 against saved files. No GPU measurements were rerun.
Local source paths below are relative to the meowkernels results archive.

## Projection shapes and synchronization

- `results/roofline/rows-scaling.txt`: best eight-row payload rates across listed kernels: draft context 241, FFN down 247, mixer output 272, full input 422, GDN input 483, FFN gate/up 530, vocabulary head 554 GB/s. These are reported projection payload rates, not direct DRAM-counter readings or a universal rate for every shape.
- `results/down-geometry/RESULTS.md`: dependent-call FFN down 118.8→99.7us for split8, with about 1 BF16 ULP difference; short-K output 36.7→37.3us. Independent-call timings approximately equal. No measured blanket 6 ms exactness penalty is established.
- `results/mk-timeline/RESULTS.md`: separate 407.6 µs, persistent 428.0 µs, persistent with software clock 450.7 µs. Clock overhead 5.3%. 24 of 160 workers hold early down tickets forabout 205 µs; average worker idle about 85 µs. Idle-worker accounting is not recoverable critical-path time. Later calibration recorded phase-clock variation and confirmed that quarter readiness remained useful.
- `results/edge-probe/RESULTS.md`: waiters variant versus matched control +3.98,+3.17,+3.11us/transition; last8+4.31,+3.67,+4.10 µs. Three 24-round invocations, 16 layers, thermal state fair, Chrome open. Byte-identical checks passed. The separate exit-spread clock experiment was marked invalid and is not used.
- `results/mk-parity-gate/handoff/GAP-LOCATION-VERDICT-20260928-140648.md`: lean_plain versus stock +4.55 µs/transition across 672 rounds with Chrome closed;  +4.97 µs across 432 rounds with Chrome open. Distinct cohorts; do not infer Chrome causality.

## Lithos per-step diagnostics

Source: `results/lithos-h2h/clean-034203/b2-lithos-t1.log` and `rows.jsonl`.

| Prompt | completion tokens / steps | GPU ms / step | client tok/s |
|---|---:|---:|---:|
| magicoder-0 | 6.705882 | 47.918086 | 143.922 |
| lithos-demo | 5.720670 | 53.001885 | 106.205 |
| numina-0 | 5.647059 | 47.941134 | 116.006 |
| r3-0 | 4.384615 | 51.477637 | 85.023 |
| aya-0 | 2.708333 | 47.834280 | 55.200 |
| five chat prompts | 3.012048–3.271565 | 48.041980–52.104929 | 57.339–66.942 |

GPU timing and client throughput have different boundaries. For magicoder the first two columns imply 139.945 tok/s, not the observed client 143.922. October 7 Pulsar pooled counters are a different cohort and do not establish matched causal attribution.

`results/lithos-h2h/official-tool/round-128.json`: nine samples, GPU median 48.035500ms, wall median 48.848000ms; accepted 7, committed 8.
`round-32000.json`: nine samples, GPU median 64.527000ms, wall median 65.241375ms; accepted 6, committed 7; prefill 151438.824 ms.
The fixture includes acceptance, commit and next drafting. These are not identical-retention context comparisons. Published-paper comparisons and release checksum provenance were not independently revalidated in this audit.

## Resident-condition comparison

Sources: `results/lithos-h2h/greedy-031406` versus `clean-034203` blocks 1–2; `identity.txt`, `log.txt`, `rows.jsonl`.

The build identity files match. Geometric means over ten common prompts, later stopped / earlier resident: Pulsar 1.233951, Lithos 1.436704. Earlier rates are 18.9595% and 30.3962% below later rates.
The earlier resident server logged 0 bytes of growth and 0.1 CPU seconds in each block. This does not prove zero GPU activity. The later log identifies no server on port 8000. Sessions were sequential and conditions / warm-up histories varied. Memory pressure remains a hypothesis for the observed difference; the entire shift is not a controlled causal effect of residency.

## Recorded client-delivery peaks

Offline rerun of `results/lithos-h2h/peak.py` on all twelve JSON recordings in `results/demo-video-lithos/raw`.
The script retokenizes complete output using the first matching cached Qwen3.8 tokenizer; token timestamps are assigned to the chunk containing each token's final character. Peak = maximum arrivals in 1.0 seconds. Average = (retokenized count − 1) / (last − first token arrival). TTFT = first nonempty chunk arrival.

| Recording | peaks tok/s | retokenized averages tok/s | TTFT |
|---|---|---|---|
| Pulsar tip runs 1–3 |206 /206 /206|150.7 /147.8 /145.4|~0.11s|
| Lithos tip runs 1–3 |154 /152 /152|111.7 /111.9 /111.8|~0.21s|
| Pulsar edit run 1 |445|321.3|2.27s|
| Lithos edit run 1 |160|142.4|3.28s|

Sources: `{fastkernel-1.1.3,lithos-metal}-{tip,E1}-run*.json`.
Edit retokenized counts 1893/1883 differ from server counts 1894/1884. Native-summary averages 321.42/142.46 retain their original definition. One-second delivery peaks include chunk bursts; they are not instantaneous GPU throughput. This audit does not establish identical methodology to another product's video overlay.

## Excluded claims

No world-first priority claim, blanket occupancy causality, measured 6 ms bit-exactness penalty, or completed result from the still-running experiments is asserted. The proposed 37 ms shape decomposition remains an unverified estimate.
