LLM Benchmarks
September 2026

Gemini enters the Top Five for the first time, Qwen3.8 27B stands out as a compact option for local deployments, and DeepSeek delivers strong results at particularly low cost. The key takeaway, however, remains the same as in previous months: it's not the newest model or the highest reasoning effort that guarantees the best performance — what matters is choosing the right configuration for the specific use case.

September brings a substantial expansion of our benchmark: Newly added models and additional reasoning configurations take the ranking from 187 to 225 entries. Gemini pushes into the leading group, Qwen offers a compelling option for local deployments, and DeepSeek shows how much capability is available at a moderate price.

But this month, too, delivers a clear reminder: a newer model, a larger parameter count, or more reasoning does not automatically produce better results.

LLM Benchmarks September 2026:
225 Models Compared

Gemini reaches the Top Five – but newer is not always better

Gemini 3.7 Flash scores 95 points in both its High and Medium reasoning configurations, placing fourth and seventh respectively in our ranking. The Medium configuration achieves the same rounded overall score at a lower estimated workload cost: €1.10 versus €1.68 for High.

Interestingly, Gemini 3.8 Flash does not improve on this result. Its best configurations score 94 points, with the highest-ranked entry appearing in eleventh place.

Document analysis also deserves a closer look. Gemini 3.7 Flash scores 90 points in the documents category with High reasoning, but 97 with Medium – matching GPT-5.6 Sol. Strong overall performance therefore does not make every Gemini configuration an equally strong choice for documents.

We are currently revising our recommendations for document analysis. In an upcoming dedicated benchmark, we will examine these limitations in greater depth and take a closer look at how Gemini compares with OpenAI. Stay tuned for the announcement.

 

Qwen3.8 27B: a model to watch for local deployment

For teams running models on their own infrastructure, Qwen3.8 27B is one of this month's most interesting additions.

With Extra High reasoning, it scores 92 points – the same rounded overall result as the best tested configurations of the much larger Qwen3.8 2.4T A95B. It also scores 100 points in both Code+Eng and CRM.

Quantized versions make deployment on a GPU with 24 GB of VRAM a realistic option. Unsloth lists a requirement of approximately 16–19 GB for its four-bit configurations, although actual memory needs depend on context length and runtime settings. Unsloth deployment guide.

Our benchmark result was obtained via a hosted API, however. The 92-point score should therefore not be interpreted as a verified result for any particular local quantization.

 

DeepSeek: strong results at a moderate cost

DeepSeek V4 Flash 0731 was already present in last month's ranking. This month, its additional reasoning configurations reveal a much stronger result: High and Max both reach 90 points, compared with 78 points for the configuration without reasoning.

The High configuration combines 90 points with an estimated workload cost of €0.12, making it a particularly attractive candidate for cost-sensitive applications.

DeepSeek has also introduced V4 Flash Vision Exp, an experimental model that processes images alongside text. DeepSeek announcement.

In our enterprise benchmark, its Low reasoning configuration scores 90 points at an estimated cost of €0.34. This is an encouraging result, though it is not a dedicated assessment of image understanding.

Once again, more reasoning does not guarantee a better outcome: Vision Exp's Max configuration scores 88 points while pushing the estimated cost up to €1.64.

 

Grok 4.6: an update without a higher headline score

Last month, Grok 4.5 entered the top ten with 94 points. Grok 4.6 does not improve on this headline result.

Its Medium and High configurations also score 94 points, while Low achieves 93 and Extra High drops to 92. The Medium configuration is faster than Grok 4.5 in our measurements, but its estimated workload cost is slightly higher.

For existing Grok users, this is reason to evaluate the specific configuration carefully rather than assume an upgrade automatically improves quality.

Fable 5.1 turns last month's disappointment into a stronger result

Our August report highlighted Fable 5's decline from 90 to 83 points when we retested it after access was restored.

Fable 5.1 now scores 94 points with Max reasoning – 11 points above that repeated test and four points above the original result.

This is a substantial improvement in our benchmark and a fitting continuation of last month's story: model capabilities can change considerably, which is why repeated evaluation is essential.

Meta's Muse shows strengths, but trails the leaders

Meta's new Muse Glimmer 30B scores between 82 and 84 points across the tested reasoning configurations – below Qwen3.8 27B's best result of 92.

The overall score conceals a mixed capability profile. Muse achieves 100 points in Code+Eng across all three configurations, while its reasoning scores range from 62 to 75.

It is therefore worth examining for specific workloads, but it does not yet match the strongest alternatives across our full set of enterprise tasks.

Other additions include NVIDIA Nemotron 3.5 Lightning, scoring 85 points, and GLM 5.3 Flash, reaching 84.

The takeaway: choose the configuration, not just the model

September offers more practical choices: Gemini now belongs to the leading group, Qwen3.8 27B deserves attention for local deployment, and DeepSeek delivers strong results at a moderate cost.

The broader lesson is equally important: model size, release order, and reasoning effort are not reliable shortcuts for choosing the best system. Test the specific configuration against your own workload – and compare quality, cost, and speed together.

All costs quoted above are the benchmark's estimated workload costs, not per-request or per-million-token API list prices.

Transform your digital projects with the best AI language models!

Discover the transformative power of the best language models and revolutionize your digital products with AI! Stay future-ready, boost efficiency, and secure a clear competitive edge. We help you take your business value to the next level.

* required

Wir verwenden die von Ihnen an uns gesendeten Angaben nur, um auf Ihren Wunsch hin mit Ihnen Kontakt im Zusammenhang mit Ihrer Anfrage aufzunehmen. Alle weiteren Informationen können Sie unseren Datenschutzhinweisen entnehmen.

Solve captcha, please!

captcha image
Martin Warnung
Sales Consultant TIMETOACT GROUP Österreich GmbH +43 664 881 788 80