LLM rivals bunch up on performance – putting pressure on price

Kimi K3 debuts - the closest a Chinese lab has come to the frontier lead... yet.

HFS set up an LLM tracker to monitor performance vs cost and context. We set it to refresh once a week. In the current rate of launches, its not enough. We have to had to run a refresh twice again already this week, the latest today as Kimi puts K3 among the front runners and reframes enterprise choices.


We are monitoring 56 models. And it’s getting increasingly tight at the top. While the cutting edge has not moved since June 6, the queue right behind @Anthropic’s Fable 5 has bunched right up.


So while Fable 5 (‘Artificial Analysis intelligence index -AA 60) holds #1 for a sixth week Moonshot AI‘s Kimi K3 just joined OpenAI (5.6) in creating a three-way race.


K3 debuted yesterday (July 16) at AA 57, the #3 model family behind Fable 5 (60) and GPT-5.6 Sol (59), and the closest a Chinese lab has come to the frontier lead… yet.

LLM rivals bunch up on performance - putting pressure on price


Kimi K3 is a 2.8T-parameter MoE, 1M context, native image+video input challenger. At AA 57 its a 13-point jump over Kimi K2.6, with open weights promised by July 27. 2.6 comes in at a 70c blended cost.


Inkling (Jul 15) has also entered the fray this week with Thinking Machines Lab‘s first-ever model, debuting as a 975B open-weights MoE (Apache 2.0) at AA 41, SWE-bench 77.6%, GPQA 87.2%. It comes in at #25 on our tracker.


Frontier velocity remains at ≈ +2.9 pts/month, but the peak has been flat at 60 since June 9 — the longest single pause at the bleeding edge of the whole of 2026. The leader index moved from 46 to 60 between February and July 2026. That was quite a leap. Enterprises now have a wide array of choices of very capable models – and the pressure is now building on price as an easy differentiator. HFS believes the harnesses (of cost, control, and context) should play just as great a part in deploying models for enterprise impact.


Even so, Kimi K3 resets frontier pricing. $0.30 cached / $3 input / $15 output ≈ $2.31/M blended — roughly 47% below GPT-5.6 Sol ($4.35) and 70% below Fable 5 ($7.70) at a 2–3 point AA discount. However, K3 used ~1.9× Sol’s output tokens on AA’s eval, eroding a good chunk of the list-price edge. DeepSeek V4 Flash remains the raw value king at AA 40 for $0.06/M.


GLM-5.2 (AA 51) is still the top downloadable open-weights model, 9 points off the closed frontier but that gap may collapse to 3 points if Moonshot ships K3 weights on July 27, as promised.


Meanwhile DeepReinforce’s MIT-licensed Ornith-1.0-397B is the new top open coding model (82.4% SWE-bench Verified), and Inkling gives the US its leading open-weights entry.


Benchmark movement is concentrating in coding and agents. SWE-bench Verified: Claude Mythos 5 (95.5%) and Fable 5 (95%) still lead by ~10 points over the best non-Anthropic model; K3’s coding scores are vendor-reported only so far.

The Bottom Line: Performance has flattened – for now. HFS expects further leaps in H2 as compute and energy bottle necks resolve. Focus on control for the most effective deployments.

More HFS Quick Takes

Sign In

Sign up for a free
research account

With the exception of our Horizons reports, most of our research is available for free on our website. Sign up for a free account and start realizing the power of insights now.

By registering you agree to our privacy policy.

I hereby consent that HFS Research can process my personal data.

Digests/Newsletters: Overviews of the latest news, insight, and research by HFS.

HFS Events: Exclusive invitations to HFS webinars, roundtables, and summits, bringing together key industry stakeholders focused on major innovations impacting business operations.

Premium Access

Our premium subscription gives enterprise clients access to our complete library of proprietary research, direct access to our industry analysts, and other benefits.

Contact us at [email protected] for more information on premium access.

    Contact Ask HFS AI Support