AI inference costs are falling at an unprecedented rate, but this deflation hides a stark divide between high-margin frontier labs and low-margin open-weight competitors. While average token prices have dropped over 50%, the underlying market structure remains stable with distinct profit tiers.
Quantization, a technique used to make AI models more efficient by reducing the number of bits needed to represent information, has limitations that are becoming more apparent. A study by researchers from Harvard, Stanford, MIT, Databricks, and Carnegie Mellon found that quantized models perform worse if the original model was trained extensively on large datasets. This poses challenges for AI companies that rely on large models to improve answer quality and then quantize them to reduce costs. The study suggests that while lower precision can make models more robust, extremely low precision may degrade quality unless the model is very large. The findings highlight the need for careful data curation and the development of new architectures to support low precision training.