ReXGroundingCT

Segmentation of Findings from Free-Text Reports

Leaderboard

Results on the full 300-case held-out test set, ranked by mean Dice per case. Click a metric's column header to rank by that metric instead, or click any row for its per-category breakdown. New submissions are evaluated automatically and appear here once scored.

Loading leaderboard…

Register & Submit

Registration is open. All submissions are made here, through the Register & Submit form below — create an account, register your team, then submit your predictions as a Google Drive link. This is the only submission channel.

The submission guidelines are a reference for the required prediction format only — they are not a separate way to submit. Review them so your predictions are formatted correctly, then submit through the form below.


About ReXGroundingCT

ReXGroundingCT is a 3D chest CT dataset linking free-text radiology findings to pixel-level segmentations, built on the CT-RATE dataset of non-contrast chest CT scans paired with radiology reports. It enables sentence-level grounding for both focal and non-focal lung and pleural abnormalities across 14 categories.

This leaderboard evaluates models that localize a radiology finding described in natural language as a precise 3D segmentation mask, on a 300-case held-out test set.


Task

Models are evaluated on free-text finding grounding: a model receives a CT volume and a natural-language finding from a radiology report and must output a 3D segmentation mask corresponding to that description. Fixed-category (class-based) segmentation that cannot separate two findings in different locations is out of scope.

Findings span 14 categories covering both typically non-focal abnormalities (bronchial wall thickening, bronchiectasis, emphysema, septal thickening, micronodules, and other diffuse abnormalities) and typically focal abnormalities (linear opacities, atelectasis/consolidation, ground-glass opacities, pulmonary nodules/masses, pleural effusion/thickening, honeycombing, pneumothorax, and other focal findings).


Dataset

SplitCasesAnnotations
Training2,992 CT scansPartial-instance (up to 3 instances per finding)
Validation200 CT scansExhaustive (all instances segmented by radiologists)
Test300 CT scansExhaustive (all instances segmented by radiologists)

All annotations are pixel-level 3D segmentation masks linked to free-text findings extracted from radiology reports. Validation and test sets are annotated exclusively by board-certified radiologists.


Evaluation Metrics

Ranking metric: mean Dice per case on the full 300-case test set. Each case's findings are averaged first, then cases are averaged with equal weight.

Overlap-based metrics:

  • Dice (primary): mean Dice per case (findings averaged within each case, then across cases)
  • Hit Rate: Proportion of findings where overall Dice ≥ 0.1
  • Instance Precision: TP / (TP + FP), where TP is a predicted instance with Dice ≥ 0.2
  • Instance Recall: TP / (TP + FN)
  • Instance F1: Harmonic mean of Instance Precision and Recall

Distance-based metrics:

  • Distance Precision: TP / (TP + FP), where TP is a predicted instance with ASSD (non-focal) or centroid distance (focal) ≤ 2× max voxel spacing
  • Distance Recall: TP / (TP + FN), using the same distance matching criterion
  • Distance F1: Harmonic mean of Distance Precision and Recall

How to Participate

  • Registration is open — create an account and register your team using the form above.
  • Training and validation data are publicly available on HuggingFace.
  • Any publicly available or private training data, including pre-trained models and external datasets, may be used. Please describe all external data sources in your submission notes.
  • All predictions on the test set must be fully automatic — no manual intervention, post-hoc editing, or case-specific tuning.
  • Free-text grounding. Methods may condition on the free-text finding directly or on a structured representation parsed from it, but must take in more than the finding category alone and must be able to produce distinct masks for two findings in different locations. Class-based segmentation (a fixed set of class masks selected by label) is out of scope and will be removed.
  • You can submit multiple runs; each is evaluated on the full 300-case test set, and each participant is ranked by their best submission. Please use the validation set for model selection and avoid excessive resubmission.
  • See the submission guidelines for format details.

Organizers

  • Mohammed Baharoon — Harvard Medical School, USA
  • Pranav Rajpurkar — Harvard Medical School, USA
  • Luyang Luo — Harvard Medical School, USA
  • Xiaoman Zhang — Harvard Medical School, USA
  • Mahmoud Hussain Alabbad — King Fahad Hospital, Saudi Arabia
  • Sungeun Kim — Harvard Medical School, USA


Contact

For questions about the leaderboard, please contact Mohammed Baharoon.


References

ReXGroundingCT:

@article{baharoon2026rexgroundingct,
  title={ReXGroundingCT: A 3D Chest CT Dataset for Segmentation of Findings from Free-Text Reports},
  author={Baharoon, Mohammed and Luo, Luyang and Moritz, Michael and Kumar, Abhinav and Kim, Sung Eun and Zhang, Xiaoman and Zhu, Miao and Alabbad, Mahmoud Hussain and Alhazmi, Maha Sbayel and Mistry, Neel P and others},
  journal={NEJM AI},
  pages={AIdbp2501220},
  year={2026},
  publisher={Massachusetts Medical Society}
}

CT-RATE:

@article{hamamci2026generalist,
  title={Generalist foundation models from a multimodal dataset for 3D computed tomography},
  author={Hamamci, Ibrahim Ethem and Er, Sezgin and Wang, Chenyu and Almas, Furkan and Simsek, Ayse Gulnihan and Esirgun, Sevval Nil and Dogan, Irem and Durugol, Omer Faruk and Hou, Benjamin and Shit, Suprosanna and others},
  journal={Nature Biomedical Engineering},
  pages={1--19},
  year={2026},
  publisher={Nature Publishing Group UK London}
}

Archive

Every submission made during the ReXGroundingCT Challenge @ MICCAI 2026, with the challenge description and timeline, is archived on the challenge archive page.