The TIMETOACT LLM Benchmarks provide an up-to-date comparison of various Large Language Models to assess their suitability for use in product development.
LLM Benchmarks
November 2023
LLM Benchmarks November 2023
Evaluating new ChatGPT models
We've evaluated the new models recently released by OpenAI (a summary of their presentation can be found here).
GPT-4 Turbo ("GPT-4 Turbo v3/1106-preview" in the table) is cheaper and less capable than the previous version. We observed a quality regression on complex coding tasks and few-shot CRM tasks.
The new GPT-3 Turbo ("GPT-3.5 v3/1106") is also cheaper and significantly faster than its predecessor. However, it too is less capable than the previous version. The degradation in model capability mirrors what we saw with the larger GPT-4 models.
On the positive side, both models offer the following:
- More recent training data
- Better capabilities in the reasoning category
Better language support in GPT-4 Turbo
GPT-4 Turbo has a hidden surprise that OpenAI hasn't mentioned anywhere.
From experience, we know that OpenAI's language models are generally really good at English. They can also be reasonably good at Chinese, German, Spanish, and Romance/Germanic languages. However, questions posed to the language models in other languages tend to produce much worse results compared to questions in English. The smaller the language's reach, the less written material is available online, and the worse the results.
This is where GPT-4 Turbo suddenly shines. Experts working with ML in niche languages report that ChatGPT is not only significantly better at communicating in small languages, but also shows a deeper understanding of the associated history and culture.
That's great news for language models! But it doesn't stop there.
Mistral 7B OpenChat catches up and matches ChatGPT 3
In the October LLM benchmarks, we introduced a new open-source model called Mistral 7B. It was sufficient for its size, but still lagged behind Llama2 70B Hermes and ChatGPT.
The November benchmarks introduce a new variant of Mistral 7B called Mistral 7B OpenChat. It outperforms Llama 70B Hermes and matches ChatGPT 3.5 (the first version).
The most impressive aspect of this model is its small size of just 7B, which means it can run without any trouble on standard hardware.
For example, on a task generating search terms for product catalogs, we were able to process 10-15 products per second. All of that on a server with an NVIDIA 3090 GPU using vLLM.
Mistral 7B OpenChat would also be particularly interesting for companies that want good LLM capabilities without depending on the uptime of OpenAI's services (they recently had outages).
Beam search improves language model accuracy
Looking closely at our previous benchmarks compared to the November benchmarks, you'd notice that smaller local models "suddenly" gained 4-5 accuracy points. We enabled beam search optimization. This algorithm automatically evaluates several answer options before selecting the best one.
This optimization will now be enabled for all local models going forward. Later, we plan to follow industry best practices and introduce even more variants of assistance into the benchmark.
Let’s turn benchmarks into your competitive advantage and build a custom AI solution tailored to your business.
Discover the transformative power of leading language models and revolutionize your digital products with AI. Stay ahead of the curve, boost efficiency, and gain a clear competitive advantage. We help you take your business value to the next level.