Whether Anthropic compresses the input:output ratio. A move from 5× to 4× would signal pricing pressure from below.
GLM-5.1 listing on Bedrock and Azure Foundry. Each addition tightens the procurement noose.
DeepSeek V5 timing. Mid-2026 release with another step-function is base case.
Whether the deflation rate at your workload mix holds or starts decelerating toward 3-5×.
The pricing-strategy tell
Anthropic prices output at 5× input. Chinese open models sit closer to 3-4×. Output token generation is more expensive in compute terms, but not by that much. The 5× ratio captures a brand premium, not just compute. The Chinese pricing is closer to actual serving cost.
Workload mix matters more than headline
A pure RAG workload sees the open-vs-closed economics narrow because input pricing converges across providers. An agentic or code-gen workload sees the gap widen because that is where the frontier premium is hidden.
Routing implication
For input-heavy workblocks, the open-vs-closed economics still favor open but the urgency is lower. For output-heavy generation and agent loops, open-weight routing is closer to a 25-30× cost saving than 21×. Build the routing layer with workload-aware pricing.
Watch over next 6 months
GLM-5.1's Opus parity claim verifying on independent SWE-Bench and stable LMArena vote counts.