← Mina Heinein

ORAL MICCAI 2026 · MSB EMERGE Workshop · Strasbourg, 27 Sept 2026

MedHyperGraph: EHR-Integrated Multimodal Hyperedges for Clinical VQA

M. Heinein, M. Youssef, K. Nashed, T. Basha, H. M. T. Alam, A. M. Selim, O. S. Bhatti, D. Sonntag

Graduation thesis, in collaboration with DFKI, the German Research Center for Artificial Intelligence. Funded by ASRT & ITIDA.

PAPER ↗ PDF ↗ CODE ↗ VIDEO ↗

Abstract

PROBLEM

Clinical visual question answering needs more than the image in front of it. The answer often depends on the patient's record — prior studies, timing, and coded history — which image-only models never see.

APPROACH

Patient-specific multimodal hypergraphs unify structured EHR events, prior-report entities, cross-modal grounding, SNOMED CT ontology concepts, and image regions in one temporally indexed structure. Temporally-aware retrieval with leakage control scores evidence by textual, visual, temporal, and hypergraph proximity, building compact evidence subgraphs for a vision-language model to answer over.

RESULT

73.95% accuracy on the EHRXQA benchmark (95% CI 71.93–75.87), beating an RL-trained state of the art at 70.43% without task-specific training. Open-ended weighted-F1 improves by 10.0 points at 27B, and multi-study F1 rises from 5.2 to 43.4. Removing the structured EHR costs 9.05–9.74 points, which is the ablation that matters.

Results

73.95%EHRXQA ACCURACY · 95% CI 71.93–75.87
+3.5ppOVER AN RL-TRAINED SOTA (70.43%)
5.2 → 43.4MULTI-STUDY WEIGHTED-F1 AT 27B

Cite

@inproceedings{heinein2026medhypergraph,
  title     = {MedHyperGraph: EHR-Integrated Multimodal Hyperedges for Clinical VQA},
  author    = {Heinein, M. and Youssef, M. and Nashed, K. and Basha, T. and
               Alam, H. M. T. and Selim, A. M. and Bhatti, O. S. and Sonntag, D.},
  booktitle = {MICCAI 2026 MSB EMERGE Workshop},
  year      = {2026},
  note      = {Oral presentation},
  url       = {https://openreview.net/forum?id=Gm7QeQBZxm}
}