Abstract
Medical image synthesis can reduce data scarcity,
but volumetric generation must preserve anatomy across planes.
In chest computed tomography (CT), report conditioning specifies
pathology but gives little spatial guidance, while maskguided
methods require a case-specific segmentation at inference.
AtlasCT removes that requirement: a population occupancy
atlas, built once from the training set, is compressed into an
embedding and injected through a text-conditioned gate and an
atlas-affinity attention bias, so each report sets how strongly the
prior applies. The contribution is the mechanism that makes a
static population prior case-adaptive rather than the choice of
prior source. A compact convolution-transformer backbone with
bidirectional text-volume fusion and single-stage super-resolution
diffusion generates volumes at 512 by 512 by 512 resolution. On
CT-RATE, AtlasCT achieves the best MedicalNet-based Frechet
distance and maximum mean discrepancy among the evaluated
report-conditioned methods. Removing the atlas reduces the
seven-region mean Dice similarity coefficient from 0.759 to
0.704, a mean-anatomy-bias analysis shows that adaptive gating
avoids contraction toward homogeneous anatomy, and synthetic
augmentation improves the classifier macro-average area under
the receiver operating characteristic curve by 0.009, with a 95
percent confidence interval from 0.004 to 0.014.
but volumetric generation must preserve anatomy across planes.
In chest computed tomography (CT), report conditioning specifies
pathology but gives little spatial guidance, while maskguided
methods require a case-specific segmentation at inference.
AtlasCT removes that requirement: a population occupancy
atlas, built once from the training set, is compressed into an
embedding and injected through a text-conditioned gate and an
atlas-affinity attention bias, so each report sets how strongly the
prior applies. The contribution is the mechanism that makes a
static population prior case-adaptive rather than the choice of
prior source. A compact convolution-transformer backbone with
bidirectional text-volume fusion and single-stage super-resolution
diffusion generates volumes at 512 by 512 by 512 resolution. On
CT-RATE, AtlasCT achieves the best MedicalNet-based Frechet
distance and maximum mean discrepancy among the evaluated
report-conditioned methods. Removing the atlas reduces the
seven-region mean Dice similarity coefficient from 0.759 to
0.704, a mean-anatomy-bias analysis shows that adaptive gating
avoids contraction toward homogeneous anatomy, and synthetic
augmentation improves the classifier macro-average area under
the receiver operating characteristic curve by 0.009, with a 95
percent confidence interval from 0.004 to 0.014.
| Original language | English |
|---|---|
| Publication status | Accepted/In press - 1 Aug 2026 |
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver