← All writing

Engineering field notes · 3 parts

Fast inference on Mac.

The goal was 200 tokens per second from a 27B model on one Mac. These are the mechanisms, the approaches that lost, and the measurements behind Pulsar.