LLM Benchmarks
August 2026

New models, new rankings: GPT-5.6 Sol, Claude Opus 5, and Grok 4.5 shake up the top of the TIMETOACT LLM Benchmarks in August 2026.

The TIMETOACT LLM Benchmarks for August 2026 show: GPT-5.6 Sol strengthens OpenAI's position, Claude Opus 5 delivers impressive results, Grok 4.5 moves up the rankings, while Fable 5 disappoints.

  • GPT-5.6 Sol: Scores 96 points to claim second place, while lower pricing makes the GPT-5.6 family even more attractive for enterprise use.

  • GPT-5.6 Terra: Delivers a strong performance with 91 points and becomes even more compelling following a 20% price reduction.

  • GPT-5.6 Luna: Scores 90 points and positions itself as a highly cost-efficient option thanks to an 80% price cut.

  • Claude Opus 5: Consistently scores 92 points across all reasoning levels and performs particularly well on complex enterprise tasks.

  • Claude Fable 5: Drops from 90 to 83 points in our repeated test, demonstrating that a model's performance can change significantly even when its name remains the same.

  • Grok 4.5: Enters the top 10 with 94 points, excelling especially in coding, CRM, and integration tasks.

  • Qwen3.7 Max: Despite maintaining its strong performance, it is pushed down to third place by GPT-5.6 Sol.

LLM Benchmarks August 2026:
187 Models Compared

GPT-5.6 joins the leading group

GPT-5.6 Sol scored 96 points to claim second place, pushing Qwen3.7 Max down to third. Terra achieved 91 points and Luna 90. At the same time, OpenAI has made the GPT-5.6 family even more attractive with significant price reductions: Terra is now 20% cheaper, while Luna's price has been reduced by 80%.

🔗 OpenAI Announcement

Fable 5 disappoints after its return

After Anthropic restored access to Fable 5, we decided to benchmark it again. Its score dropped from 90 to 83 points, demonstrating that a model's behavior can change significantly even when its name remains the same. Anthropic announcement

🔗Anthropic-Announcement

Claude Opus 5 catches up with the leaders

We evaluated the model across three reasoning levels—Low, Medium, and High—and all three achieved the same score of 92 points. Claude now performs significantly better on our enterprise benchmark tasks, while the nearly identical results suggest that increasing reasoning effort does not necessarily lead to better overall performance.

 

Grok 4.5 breaks into the top 10

With 94 points, Grok 4.5 enters the rankings in ninth place. It performed particularly well in coding, CRM, and integration tasks. xAI positions the model for software development and agentic workflows—a claim that is also supported by the overwhelmingly positive feedback from the developer community.

🔗 xAI Announcement

🔗 Developer discussion

 

 

Summary

August 2026 highlights how competitive the race at the top has become. While GPT-5.6 further strengthens OpenAI's leading position, Claude Opus 5 reaches a new level of quality, and Grok 4.5 establishes itself as a serious contender. At the same time, the significant performance drop of Fable 5 is a reminder that AI models are not static products. Continuous benchmarking is therefore essential for making informed architecture and model selection decisions.

Transformieren Sie Ihre digitalen Projekte mit den besten KI-Sprachmodellen!

Entdecken Sie die transformative Kraft der besten Sprachmodelle und revolutionieren Sie Ihre digitalen Produkte mit KI! Bleiben Sie zukunftsorientiert, steigern Sie die Effizienz und sichern Sie sich einen klaren Wettbewerbsvorteil. Wir unterstützen Sie dabei, Ihren Business Value auf das nächste Level zu heben.

* required

Wir verwenden die von Ihnen an uns gesendeten Angaben nur, um auf Ihren Wunsch hin mit Ihnen Kontakt im Zusammenhang mit Ihrer Anfrage aufzunehmen. Alle weiteren Informationen können Sie unseren Datenschutzhinweisen entnehmen.

Solve captcha, please!

captcha image
Martin Warnung
Sales Consultant TIMETOACT GROUP Österreich GmbH +43 664 881 788 80