LLM Benchmarks
October 2023

The TIMETOACT LLM Benchmarks provide an up-to-date comparison of various Large Language Models to assess their suitability for use in product development.

LLM Benchmarks October 2023 

Highlights & news from the October benchmarks

9 new benchmarks

We've added 9 new benchmarks to the suite. These benchmarks focus on the "Documents," "Integration," and "Reason" categories. This makes the evaluation of model capabilities more precise, and the total number of different evaluations rises from 85 to 134.

One example is situations where large language models create and process structured data.

In the Integration category, we now test the ability of large language models to understand and manipulate text in CSV, TSV, JSON, and YAML formats.

Another example concerns our work on business assistants and information retrieval systems for clients. In such cases, large language models need to identify, find, and evaluate relevant pieces of information. Our evaluations help measure various aspects of this capability. In addition to these new evaluations, we've improved the performance of some existing evaluations by introducing few-shot examples and better prompts. Most large language models respond very positively to this.

More guidance

Guidance is a process that helps large language models generate the desired text. It works by directing the model's attention to specific text elements (tokens).

As we gain more experience in getting better results from large language models, we're incorporating these insights into the benchmarks. Our October release already includes guidance in some of the evaluations, further improving the performance of some models.

In the coming months, we plan to provide even deeper guidance for models in task-specific areas.

Impressive newcomer: Mistral 7B

Mistral 7B is a new model from a French AI company of the same name. Although it is significantly smaller than the other models, it has outperformed the base configurations of Llama2 70B as well as all models in the 7B and 13B size classes.

That's genuinely impressive. It's worth paying more attention to this model over the coming months. Its cost and throughput characteristics make it even more attractive for local deployments.

Another highlight of this model is that it was released under the Apache license, which is clearer and less restrictive than the Llama 2 license. There are no "Google" clauses or potential ambiguity regarding using this model for non-English languages. Our model labeling reflects this change in the table.

Let’s turn benchmarks into your competitive advantage and build a custom AI solution tailored to your business.

Discover the transformative power of leading language models and revolutionize your digital products with AI. Stay ahead of the curve, boost efficiency, and gain a clear competitive advantage. We help you take your business value to the next level.

* required

Wir verwenden die von Ihnen an uns gesendeten Angaben nur, um auf Ihren Wunsch hin mit Ihnen Kontakt im Zusammenhang mit Ihrer Anfrage aufzunehmen. Alle weiteren Informationen können Sie unseren Datenschutzhinweisen entnehmen.

Solve captcha, please!

captcha image
Martin Warnung
Sales Consultant TIMETOACT GROUP Österreich GmbH +43 664 881 788 80