The Speed Floor Moves: Wafer-Scale Inference and the Tokens-Per-Second Axis
GPT-5.6 Sol runs at 750 tokens per second on Cerebras wafer-scale silicon — roughly 15x a GPU. When frontier inference gets an order of magnitude faster, latency becomes the competitive axis.
26 min readMichael Eakins
Artificial IntelligenceInferenceHardware+2