At CHUV, small language models tested on French-language clinical data

CHUV’s Biomedical Data Science Center, together with teams in infectious diseases, oncology, internal medicine, clinical neurosciences and clinical informatics, evaluated six small open-source language models (8 to 24 billion parameters) on de-identified French-language clinical data from CHUV. Seven tasks were tested, from information extraction to clinical decision support.

Llama 3.1 and Mistral Small achieve near-perfect performance at extracting information. At identifying personal data, they perform markedly worse than RoBERTa, a specialised model, and one of the models introduced non-existent information in more than half of the translations of discharge letters. On complex tasks, evaluations range from neutral to unsatisfactory. The study is published in the Journal of Medical Internet Research.

https://www.chuv.ch/fr/actualites/actualite/news_actu_1170/les-petits-modeles-de-langage-a-lepreuve-du-terrain-clinique-au-chuv