LLM Benchmarks
January 2024

The TIMETOACT LLM Benchmarks provide an up-to-date comparison of different Large Language Models, evaluating how well suited they are for use in product development.

LLM Benchmarks January 2024

Mistral 7B OpenChat v3 (0106) beats ChatGPT-3.5

Open-source models are getting better month by month. That's great news, since you can run these models locally on your own machines. Our clients really appreciate that option!

Progress in open-source models: more local variants becoming viable

Mistral 7B OpenChat-3.5: a top contender among open-source models

Among these open-source models, one of the best contenders is a fine-tuned version of Mistral 7B called Mistral 7B OpenChat-3.5, which offers a very good balance of quality, performance, and the compute resources required to run locally on your own hardware.

Third version's edge: Mistral 7B OpenChat-3.5 beats ChatGPT-3.5 on business tasks

So far we've mainly used the first version of Mistral 7B OpenChat-3.5 (fine-tune). However, our latest benchmarks show that the third version is significantly better. It even beats ChatGPT-3.5 on our business-oriented tasks!

That said, there's still a long way to go, since the version of ChatGPT-3.5 that Mistral 7B OpenChat-3.5 beat is the oldest and weakest one. Still, it's a start, and we're very curious to see what the coming weeks and months bring.

Why aren't there benchmarks for Mixtral 8x7B or Mistral-Medium?

You've probably already heard of another, significantly more capable version of Mistral: Mixtral 8x7B. This model uses a different architecture called Sparse Mixture of Experts (MoE).

This MoE model is more capable than Mistral 7B. It also powers MistralAI's hosted mistral-small model. We don't have either of these models in this LLM benchmark. The Mistral-Medium model, which is even more capable than Mixtral 8x7B, is also not included.

Why we haven't added them to the benchmarks yet

The second generation of Mistral models was trained to perform better than the first generation. These models give better answers, but sometimes don't stick precisely to the response format specified in the prompt and few-shot examples. They're more verbose than necessary.

This verbosity makes the model less useful for business-oriented tasks and business process automation.

Of course, one could add post-processing when using the model to strip out extra explanations. However, it would be unfair to the other models to add Mistral-specific post-processing to the TIMETOACT LLM Benchmarks, so we won't be doing that.

TIMETOACT's benchmarks helped the Mistral AI team better understand issues with the model

We reported the issue to the Mistral AI team and provided reproducible examples. The TIMETOACT LLM Benchmarks helped clarify the impact and narrow down the scope: the first model generation didn't have this problem.

The Mistral AI team is making rapid progress and is already working on the issue:

"Our models sometimes just tend toward verbosity... Our team is also working on this and will improve it."

Benchmark update after the issue is fixed

At this point, we've decided not to include the current benchmark results for mistral-tiny/mistral-small/mistral-medium, since this is just a temporary issue that will soon be resolved. Once that happens, we'll publish an update to the benchmarks.

Planning synthetic benchmarks: Wizards, Q&A, and RAGs

Even though pure LLM benchmarks are interesting and exciting, they can sometimes be too technical and not very relevant.

As we've learned from our clients, an LLM's performance is just one of many factors that contribute to the value of a complete product or service. We're currently considering creating another synthetic benchmark — one that evaluates the performance of complete AI systems on business-specific tasks. For example:

  • Finding the correct answer within a lengthy PDF
  • Synthesizing a correct answer that requires looking across multiple documents
  • Correctly evaluating and processing incoming documentation

We want your input: which business-specific tasks should we benchmark?

Are there other business-specific tasks you'd like to see in our benchmark? We'd love your feedback and would be happy to hear from you!

You're also warmly invited to contact us to get a sneak peek at such a benchmark before it's published.

Let’s turn benchmarks into your competitive advantage and build a custom AI solution tailored to your business.

Discover the transformative power of leading language models and revolutionize your digital products with AI. Stay ahead of the curve, boost efficiency, and gain a clear competitive advantage. We help you take your business value to the next level.

* required

Wir verwenden die von Ihnen an uns gesendeten Angaben nur, um auf Ihren Wunsch hin mit Ihnen Kontakt im Zusammenhang mit Ihrer Anfrage aufzunehmen. Alle weiteren Informationen können Sie unseren Datenschutzhinweisen entnehmen.

Solve captcha, please!

captcha image
Martin Warnung
Sales Consultant TIMETOACT GROUP Österreich GmbH +43 664 881 788 80