Challenging cost headwinds for frontier AI models ahead

The initial AI players in the US, Claude ,and ChatGPT, seems to find gold with their models.But as costs of frontier models balloon, and open models get closer in capability, the competition is heating up both on capability and cost.
The AI leaders continue spending more money to train the latest and bigger models (ChatGPT Sol is ~2 trillion parameters, Claude Fable at 6 trillion parameters , and China’s Kimi K3 is 2.8 trillion). Each of the latest frontier model is more than 10x the original ChatGPT model at 175 billion parameters.

In terms of cost, Claude Fable 5 API access costs double that of Claude Opus 4.8 with Fable 5 priced at $10 per million input tokens and $50 per million output tokens. Chat GPT Sol costs 5.00 per million input tokens and $30.00 per million output tokens for short-context requests (272k) token and higher prices for larger context windows. Kimi K3 model costs far less at $3 per million input tokens and $15 per million output tokes. For the frontier models in the US, the price points are at least double what the K3 model charges.

For companies that have seen token maxxing by employees and shocking API costs, the use of the right model will not come down to just the latest. Instead focussing on the ROI of the project will mean picking the most efficient model to minimize the cost. A key barrier for corporate adoption by K3 would be the comfort of sending data to Chinese based models. But a major hyperscaler, (Microsoft) is currently testing the K3 model for Copilot( https://cryptobriefing.com/microsoft-kimi-k3-ai-inference-costs/).

OpenAI and Microsoft had an early alliance with Microsoft’s CEO bringing Sam Altman back to OpenAi after his abrupt firing by the board. Since the event, 3 years ago, and marked by the removal of Azure exclusivity for Open AI, the partnership has fizzled out. Part of the problem is the high costs of closed model like Chat GPT. This allows Microsoft to test other models as the brains behind CoPilot.

Microsoft’s reach into the corporate world is dominated by its Office licenses, Windows OS, and Azure Cloud platform. With default status to be able to turn on AI models into Outlook, Excel, and Word finding the right balance of cost vs margin will drive whether their corporate base end up on Microsoft’s AI platform or seek other alternatives. Microsoft has not had the same headlines that ChatGPT, Claude, or even Google with the Gemini 3.5 launch garnered.

If Microsoft ends up with a Chinese K3 model, powering its CoPilot for tasks, the AI spend by corporations could see a big shift from USA based closed models to open models even if it is China based. The costs of tokens, compared to Fable and Sol would decrease 70% and 50% respectively. The moat then would the ideas a firm has.

With AI, the coding bottleneck is no longer the developer typing the code. Instead now with slack mentions, and agents, the develop spec is the bottleneck. If one can turn that into an always on, always running AI coding machine, then using a good enough model (K3) to generate code has a huge cost advantage while also giving firms an almost always updating product pipeline. The limits for firms then becomes how good of an idea do you have to build out.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top