MedHyperGraph: EHR-Integrated Multimodal Hyperedges for Clinical VQA
M. Heinein, M. Youssef, K. Nashed, T. Basha, H. M. T. Alam, A. M. Selim, O. S. Bhatti, D. Sonntag
Graduation thesis · with DFKI · funded by ASRT & ITIDA
ABSTRACT & RESULTS
Clinical visual question answering should read the medical image and the patient’s record together, but most models only see the pixels.
Each case becomes a multimodal hypergraph that ties EHR entities to image findings. GraphRAG retrieves over that graph, so a vision-language model answers grounded in the real record, with no task-specific training.
73.95% on EHRXQA boolean questions (n=1,900; 95% CI 71.9–75.9) without task-specific training — the only method that improves on all three backbones, and ahead of the 70.43% an RL-trained system reports on its own backbone. On open-ended multi-study questions weighted-F1 climbs from 5.2 to 43.4; ablations put most of the gain on the structured EHR.
“Does this image show cardiomegaly, and is there any abnormality in the left lung?”
The model pulls the patient’s prior studies and reports from the EHR hypergraph, reads the image, and answers using both sources together.
· FUNDED BY
ITIDA73.95%
ACCURACY ON EHRXQA (n=1,900)
vs 70.43%
RL-TRAINED STATE OF THE ART · WITHOUT TASK-SPECIFIC TRAINING
5.2 → 43.4
WEIGHTED-F1 ON OPEN-ENDED · MULTI-STUDY JUMP














