[2605.16215] Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
Skip to main content
arXiv is now an independent nonprofit!<br>Learn more<br>×
Search arXiv
Press Enter to search · Advanced search
-->
Computer Science > Artificial Intelligence
arXiv:2605.16215 (cs)
[Submitted on 15 May 2026 (v1), last revised 29 May 2026 (this version, v2)]
Title:Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
Authors:Xavier Theimer-Lienhard, Mushtaha El-Amin, Fay Elhassan, Sahaj Vaidya, Victor Cartier-Negadi, David Sasu, Lars Klein, Mary-Anne Hartley<br>View a PDF of the paper titled Fully Open Meditron: An Auditable Pipeline for Clinical LLMs, by Xavier Theimer-Lienhard and 7 other authors
View PDF<br>HTML (experimental)
Abstract:Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Most "open" models are open-weight only, releasing parameters while withholding the data provenance, curation procedures, and generation pipelines that determine model behavior. Fully Open (FO) models, which expose the complete training stack end-to-end, do not currently exist in medicine. We introduce Fully Open Meditron, the first fully open pipeline for building LLM-CDSS, comprising a clinician-audited training corpus, a reproducible data construction and training framework, and a use-aligned evaluation protocol. The corpus unifies eight public medical QA datasets into a normalized conversational format and expands coverage with three clinician-vetted synthetic extensions: exam-style QA, guideline-grounded QA derived from 46,469 clinical practice guidelines, and clinical vignettes. The pipeline enforces system-wide decontamination, gold-label resampling of teacher generations, and end-to-end validation by a four-physician panel. We evaluate using an LLM-as-a-judge protocol over expert-written clinical vignettes, calibrated against 204 human raters. We apply the recipe to five FO base models (Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct). All MeditronFO variants are preferred over their bases. Apertus-70B-MeditronFO improves +6.6 points over its base (47.2% to 53.8%) on aggregate medical benchmarks, establishing a new FO SoTA. Gemma-3-27B-MeditronFO is preferred over MedGemma in 58.6% of LLM-as-a-judge comparisons and outperforms it on HealthBench (58% vs 55.9%). These results show that fully open pipelines can achieve state-of-the-art domain-specific performance without sacrificing auditability or reproducibility.
Comments:<br>Preprint. 31 pages, 10 figures. Code, models, and data: this https URL
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:<br>arXiv:2605.16215 [cs.AI]
(or<br>arXiv:2605.16215v2 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2605.16215
Focus to learn more
arXiv-issued DOI via DataCite
Submission history<br>From: Xavier Theimer-Lienhard [view email]<br>[v1]<br>Fri, 15 May 2026 17:29:08 UTC (603 KB)
[v2]<br>Fri, 29 May 2026 15:56:10 UTC (603 KB)
Full-text links:<br>Access Paper:
View a PDF of the paper titled Fully Open Meditron: An Auditable Pipeline for Clinical LLMs, by Xavier Theimer-Lienhard and 7 other authors<br>View PDF<br>HTML (experimental)<br>TeX Source
view license
Current browse context:
cs.AI
next >
new<br>recent<br>| 2026-05
Change to browse by:
cs<br>cs.CL
References & Citations
NASA ADS<br>Google Scholar
Semantic Scholar
export BibTeX citation<br>Loading...
BibTeX formatted citation
×
loading...
Data provided by:
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness,...