LLM output prices span 100× — and most SMBs don't notice
The public model catalog in July 2026 shows a 100×+ spread between the most expensive output (GPT-5.6 Sol at $30/1M tokens) and the cheapest competitive one (DeepSeek V4 Flash at $0.28/1M). Anyone running single-vendor is paying the worst price per task — and the quality gain doesn't always justify it.
For tasks like classification, structured extraction and simple RAG, models in the $0.25–$5/1M range deliver quality equivalent to frontier models in 80–90% of cases. The difference shows up in heavy multi-step reasoning, complex code and long contexts.
On a typical SMB workload (60% light, 30% medium, 10% heavy tasks), cost-optimal policy routing cuts monthly spend by 65–85% while maintaining perceived quality. The condition is a gateway with fallback, observability and audit — otherwise the savings vanish at the first provider hiccup.
The single-vendor risk isn't just financial: it's regulatory (your data policy is married to the vendor's policy) and operational (any pricing change or deprecation lands directly on your P&L).
Frontier model for heavy tasks, cheap model for the rest. Without a gateway that's a project. With a gateway it's configuration.
