The Inference Cost Curve: Why Token Prices Keep Falling
Speculative decoding, MoE pruning and purpose-built silicon are pushing effective token prices down an order of magnitude per year.
0 reactions ยท anonymous
The compounding tricks
It's not one breakthrough โ it's a stack of compounding optimisations: speculative decoding, activation sparsity, quantisation to 4-bit, and inference chips with huge on-chip memory.
Extrapolating the trend, the marginal cost of a GPT-4-class token in two years will be a rounding error. The economics of AI products will shift from paying for intelligence to paying for trust and workflow.
Discussion
0// No comments yet โ be the first to start the discussion.