State of Inference

Inference economics, read like a market — not a press release.

HYPERSCALER 2026 CAPEX $725B ▲ 77% YoY APPLE FY25 CAPEX $12.7B ▼ <2% of peers V4.1-FLASH KV CACHE 890B/tok ▼ 4x smaller SHARE OF GPU SPEND: INFERENCE ~80% HYPERSCALER 2026 CAPEX $725B ▲ 77% YoY APPLE FY25 CAPEX $12.7B ▼ <2% of peers V4.1-FLASH KV CACHE 890B/tok ▼ 4x smaller SHARE OF GPU SPEND: INFERENCE ~80%

ARCHIVE

DeepSeek's KV cache diet is the real story in V4.1-Flash

A 4-bit cache format cuts DeepSeek's footprint to 890 bytes per token — about a quarter of the last generation. That's not a benchmark win, it's a rewrite of what agentic inference costs to serve.

Same model, different harness

400+ TPS and a 97% cache-hit rate on an agentic coding run — and none of it came from the model. What harnesses actually control in inference performance.

The DeepSeek model nobody has confirmed yet

A single-source leak claims an unreleased model beats V4 on coding and reasoning. Why the rumor matters regardless of whether it's true.

Apple's quiet bet against the GPU boom

$12.7B in capex against $725B from the other four hyperscalers. Apple isn't behind on AI — it made a different bet on who pays for inference.

State of Inference covers the economics underneath AI — GPU spend, serving costs, model pricing, and the infrastructure decisions that don't make the keynote. Editorially independent, data-first, no vendor talking points.