← Mina Heinein

ORAL MICCAI 2026 · MSB EMERGE Workshop · Strasbourg, 27 Sept 2026

LocSAM3: Box-Supervised Adaptation of SAM3 for Text-Only Chest X-Ray Segmentation

K. Nashed, M. Heinein, M. Youssef, T. Basha, H. M. T. Alam, A. M. Selim, O. S. Bhatti, D. Sonntag

Graduation thesis, in collaboration with DFKI, the German Research Center for Artificial Intelligence. Funded by ASRT & ITIDA.

PAPER ↗ PDF ↗ CODE & SUPPLEMENTARY ↗

Abstract

PROBLEM

Pixel-level masks are expensive to annotate on chest X-rays, where anatomy overlaps and pathological boundaries are ill-defined. Bounding boxes are far cheaper, and already exist in several datasets.

APPROACH

Adapts SAM3's concept grounding and localization using only boxes and concept names — no pixel masks. The pretrained mask decoder and over 95% of the vision backbone stay frozen, and the box is supplied for only a random subset of training samples, so the model learns to localize from text alone.

RESULT

On a 10-class MIMIC-CXR benchmark, text-only mIoU improves over Medical-SAM3 from 0.48 to 0.69 for anatomical structures and from 0.09 to 0.53 for pathological findings — with no dense masks in training and no spatial prompt at inference.

Results

0.69ANATOMY mIoU · TEXT-ONLY, FROM 0.48
0.53PATHOLOGY mIoU · TEXT-ONLY, FROM 0.09

Cite

@inproceedings{nashed2026locsam3,
  title     = {LocSAM3: Box-Supervised Adaptation of SAM3 for Text-Only
               Chest X-Ray Segmentation},
  author    = {Nashed, K. and Heinein, M. and Youssef, M. and Basha, T. and
               Alam, H. M. T. and Selim, A. M. and Bhatti, O. S. and Sonntag, D.},
  booktitle = {MICCAI 2026 MSB EMERGE Workshop},
  year      = {2026},
  note      = {Oral presentation},
  url       = {https://openreview.net/forum?id=mWnH2hNROa}
}