Megakernels on Mac: What Fusion Saves and What It Costs
Persistent workers, memory reuse, and the dependencies that survive fusion. Why an improved megakernel still lost to optimized separate kernels.
Read the essayEngineering field notes · 3 parts
The goal was 200 tokens per second from a 27B model on one Mac. These are the mechanisms, the approaches that lost, and the measurements behind Pulsar.
Persistent workers, memory reuse, and the dependencies that survive fusion. Why an improved megakernel still lost to optimized separate kernels.
Read the essayDFlash, DFlash 2 and DSpark: block proposals, target verification, retained state and the cost of a complete speculative cycle.
Read the essayMaking a 27B model fast on one Mac: the speed gains, the moving baseline, and the kernel and drafter experiments that failed.
Read the essay