ORAL MICCAI 2026 · MSB EMERGE Workshop · Strasbourg, 27 Sept 2026
Graduation thesis, in collaboration with DFKI, the German Research Center for Artificial Intelligence. Funded by ASRT & ITIDA.
Clinical visual question answering needs more than the image in front of it. The answer often depends on the patient's record — prior studies, timing, and coded history — which image-only models never see.
Patient-specific multimodal hypergraphs unify structured EHR events, prior-report entities, cross-modal grounding, SNOMED CT ontology concepts, and image regions in one temporally indexed structure. Temporally-aware retrieval with leakage control scores evidence by textual, visual, temporal, and hypergraph proximity, building compact evidence subgraphs for a vision-language model to answer over.
73.95% accuracy on the EHRXQA benchmark (95% CI 71.93–75.87), beating an RL-trained state of the art at 70.43% without task-specific training. Open-ended weighted-F1 improves by 10.0 points at 27B, and multi-study F1 rises from 5.2 to 43.4. Removing the structured EHR costs 9.05–9.74 points, which is the ablation that matters.
@inproceedings{heinein2026medhypergraph,
title = {MedHyperGraph: EHR-Integrated Multimodal Hyperedges for Clinical VQA},
author = {Heinein, M. and Youssef, M. and Nashed, K. and Basha, T. and
Alam, H. M. T. and Selim, A. M. and Bhatti, O. S. and Sonntag, D.},
booktitle = {MICCAI 2026 MSB EMERGE Workshop},
year = {2026},
note = {Oral presentation},
url = {https://openreview.net/forum?id=Gm7QeQBZxm}
}