New and improved LLM leaderboard

Today, we celebrate the release of a brand new leaderboard that ranks the performance of large language models on Icelandic language and culture benchmarks. We look forward to hearing your feedback on the leaderboard, which is hosted on Hugging Face and can be accessed through the URL www.stigatafla.is.

Two years ago we released our first version of a leaderboard for the performance of large language models in Icelandic (An Icelandic leaderboard for large language models). While that leaderboard has proven to be an excellent guide for us and others in selecting the right models for various tasks, a lot has happened in the intervening years; the models have improved significantly, and keen observers have no doubt noticed that most of the old benchmarks have become saturated and are no longer useful for identifying progress [1].

It had, therefore, become evident that we needed new, better, and more difficult benchmarks to replace the old ones.

To address this issue, we brought on board a group of talented summer interns through the Student Innovation Fund to handcraft benchmark examples: Alfa Magdalena B Jórunnardóttir, Stefán Atli Árnason, and Thelma Karen Siggeirsdóttir. The result was a new set of test questions covering a wide range of topics: text generation skills; grammar; knowledge of Icelandic culture, such as music, television content, Icelandic wildlife, and handicrafts; idioms; and well-known phrases.

One major difference between the new benchmarks and the old ones is that neither the test questions nor the answers are made public. This way, we prevent the data from accidentally leaking into the models' training data and thereby contaminating the test results.

The new leaderboard contains other information, though, that might be of interest, such as the number of tokens a model uses to solve the problems, the evaluation date, a breakdown into subcategories where applicable, and the exact model settings used for the evaluation.

The first publication of our new leaderboard shows the performance of selected models that were recently released or have garnered special attention for their good performance in Icelandic. We will continue to update the leaderboard over the next few years with results for new models as they are released. We are also very open to adding new benchmarks to the collection, whether they are born out of Miðeind's own research or come from others interested in the status of Icelandic in AI models. As of this writing, the new model GPT-6 Astra from OpenAI is at the top of the leaderboard with an average score of about 70%. The leaderboard and more detailed information about the tests can be accessed on the Hugging Face website via a new link: www.stigatafla.is.

Benchmarks in the new leaderboard:

  • WikiQA. This test, which measures knowledge of Icelandic history and culture, remains unchanged from the old leaderboard, except that GPT-5.6 Luna now serves as the judge. Its agreement score on manually labeled examples is comparable to that of the previous evaluation model, GPT-4o.

  • Pop Culture. This is a two-part test that measures, on the one hand, knowledge of well-known Icelandic song lyrics and, on the other, well-known phrases from Icelandic media and public life. In the lyrics section, models are asked, for example, to complete the Icelandic national anthem, while in the phrases section, they are asked about lines from literature and television shows, such as answering who said, “My time will come!” It should be noted that this is currently by far the most difficult test, as it directly measures how much information the models have memorized.

  • Terminology. Here we test knowledge of specialized vocabulary. In the handicrafts section, models are asked to explain the meaning of various knitting and crochet terms, such as “fitja upp” (to cast on). In the other section, they are asked for Icelandic translations of English and Latin names for a vast number of plants and animals found in Iceland.

  • Generation Quality. The generation quality test is a new addition where models are required to compose text according to a specification, for example: “Write a short story in Icelandic about a reindeer that gets lost in downtown Egilsstaðir.” The texts are then run through the proofreading tool Málfríður to assess the error rate, and the words are looked up in BÍN (the Database of Modern Icelandic Inflection) to evaluate the proportion of non-words. Each of these methods is imperfect for assessing text quality on its own, but put together they correlate quite well with what a human would evaluate.

  • Grammar Correction. In this test, the model is given Icelandic sentences that contain either one grammatical error or none and is tasked with correcting them if necessary.

  • Case Government. Many Icelandic verbs take a subject in an oblique case (“mér leikur forvitni á” [I [dat] am curious], “hana vantar” [she [acc] needs]). Furthermore, the verb alone often determines the case of the object. In this benchmark, the model is given a noun phrase in the nominative case and must inflect it correctly given its role in the sentence structure, or construct a phrase given a verb and an object. The best models get close to a perfect score here, so the test is mainly useful for identifying weaker performers.

  • Idioms. This test is twofold: The first part presents well-known idioms with a blank to fill in, e.g., “Að vera þrándur í _ einhvers” (götu) [To be a thorn in someone's side]. In the second part, the model is asked to explain the meaning of idioms, e.g., “What does the idiom ‘að lepja dauðann úr skel’ mean?” [literally: to lap death from a shell]. Then, a language model is put in the role of a judge and given the correct meaning for reference. The answer “to eke out a living in great poverty” would get full marks, but “to eat a spoiled oyster” would probably get none.


[1] Some were also of low quality, such as the machine-translated benchmark ARC-Challenge, which was heavily geared towards American culture, and the machine translation introduced unfortunate errors, such as when it asked about the discoveries of the scientist Louis Guðmundsson regarding pasteurization (https://arxiv.org/pdf/2603.16406).

Sign up for our mailing list!

Post Tags:
Share this post: