CTO
25.8.2026
We are pleased to introduce Lísa, a tool that extracts text and images from PDF, PNG and JPG documents using optical character recognition, or OCR. Lísa handles large documents — up to 1,000 pages or 100 MB in size — and returns a .docx, .odt, .md or .txt text file that you can process further. It copes well with a wide range of documents and languages, and is particularly strong on Icelandic text and special Icelandic characters.
Once you have selected a document, it appears on the input side of the interface, where you can page through it and zoom in and out as needed. You can also choose to extract the images from the document during OCR: each image is then cropped out separately and placed in the appropriate location in the resulting text file (unless you have chosen .txt).
Lesa skjal (read document) starts the job, and a progress bar tracks how far along it is. You can close the tab and come back later - the job keeps running on the server.
When the OCR job has finished, the Sækja .docx (Retrieve .docx) button gives you the file in Word format, or you can pick other file types from the arrow button beside it. Download links are valid for one hour.
You can also click Þýðing (translation) to send the document straight on to document translation in Erlendur or Yfirlestur (proofreading) below to access document proofreading in Málfríður. PDF, PNG and JPG files can likewise be dragged straight into Málfríður and Erlendur, where OCR then runs as a pre-processing step ahead of translation and proofreading. From there you can move an OCR-processed document directly into the editor with the Færa í ritil (move to editor) button on the input side of the interface.
Below you can see an example of a PDF page on the left and the OCR-processed .docx output on the right, where an image on the page has been detected and cropped out, and then included with the text below it in the appropriate style.
All Málstaður users can now use this tool, with each page processed costing 10 ISK in credit / usage cost. This means that users with a fixed (föst) subscription can run OCR on up to 1,000 pages a month using their monthly credit.
In the near future we aim to improve the experience further with the following additions:
Google Drive integration, letting users pick documents straight from Google Drive and save results back there.
Job history, keeping track of previous OCR jobs, much like the transcription history in Hreimur.
An API for OCR at api.malstadur.is, available to everyone with a flexible (frjáls) subscription.
Greater accuracy in OCR, including by keeping text styling consistent from page to page.
We're excited to see how people put OCR to use, and we have plenty more in the pipeline. Keep an eye out for new features that make working with Icelandic text faster and easier. You can follow along on our website, Facebook, Instagram and LinkedIn.
As always, our aim is to bring the best available language technology and AI for Icelandic to Icelandic society — in an open and responsible manner.
If you want to follow future projects at Miðeind, we can let you know when there is something new to report.