Synthetic supervision works.
Models trained without real positive annotations can still transfer substantially to real CMB detection.
01
More often than models, these are what set the performance ceiling in medical AI.
How do we train a model when we have few—or even no—positive patients?
Can we create useful training targets without labeling every lesion by hand?
Suppose we could build a synthetic disease generator.
Could a model learn entirely from fake lesions—and still detect real ones?
02
We chose Cerebral Microbleed as our research target.
Cerebral microbleeds (CMBs) are tiny signs of previous bleeding in the brain. On certain MRI scans, they appear as small, dark spots—often only a few millimeters across.
CMB burden is an important imaging biomarker, but detecting and annotating these lesions remains difficult.
Often only a few millimeters in diameter.
Positive lesions—and even positive patients—can be relatively scarce.
Vessels, calcifications, and susceptibility artifacts can resemble true microbleeds.
04
Synthetic lesions are parameterized by size, shape, orientation, signal contrast, and boundary appearance, then inserted into real CMB-negative SWAN images.
05
A MONAI 3D U-Net trained exclusively on synthetic lesions retained substantial detection performance when evaluated on real CMBs.
06
Synthetic-to-real transfer differed substantially between the two evaluated segmentation frameworks. A relatively simple, shallow 3D U-Net transferred well, while the more sophisticated nnU-Net struggled. Why?
Lesion sensitivity
with synthetic-only training
Lesion sensitivity
with synthetic-only training
07
Successful transfer may depend not only on the realism of synthetic lesions, but also on how the downstream model learns from the synthetic distribution.
Our experiments demonstrate framework-dependent transfer, but do not identify the specific mechanism responsible for the difference.
08
Models trained without real positive annotations can still transfer substantially to real CMB detection.
Real-trained models consistently outperformed synthetic-trained models across the evaluated data scales.
It's not just about making synthetic data look real. How the downstream model learns from synthetic data also matters.