DeepSeek's KV cache diet is the real story in V4.1-Flash
A 4-bit cache format cuts DeepSeek's footprint to 890 bytes per token — about a quarter of the last generation. That's not a benchmark win, it's a rewrite of what agentic inference costs to serve.