The Inference Cost Curve: Why Token Prices Keep Falling

Speculative decoding, MoE pruning and purpose-built silicon are pushing effective token prices down an order of magnitude per year.

Demo Admin Author 10 min read 1.9k views
0 reactions ยท anonymous
The Inference Cost Curve: Why Token Prices Keep Falling

The compounding tricks

It's not one breakthrough โ€” it's a stack of compounding optimisations: speculative decoding, activation sparsity, quantisation to 4-bit, and inference chips with huge on-chip memory.

Extrapolating the trend, the marginal cost of a GPT-4-class token in two years will be a rounding error. The economics of AI products will shift from paying for intelligence to paying for trust and workflow.

Demo Admin

Building incoffeed โ€” a daily editorial on tech, AI, fintech and business. Co-founder & editor.

Discussion

0
Commenting as guest sign in
0/2000
// No comments yet โ€” be the first to start the discussion.