University of Edinburgh
Model weights not public. Contact creators for more information.
Disentangled representation learning model for cardiac image analysis that factorises 2D medical images (MRI, CT) into a spatial 'anatomy factor' (a semantically meaningful multi-channel map, produced by a U-Net-style anatomy encoder) and a non-spatial 'modality factor' (a latent vector capturing imaging-specific characteristics). This disentangled representation supports semi-supervised segmentation using only a fraction of labeled images (matching fully supervised performance), multi-task learning (e.g. jointly regressing cardiac indices), multimodal pooling of MRI and CT data, and image-to-image synthesis between modalities via latent-space arithmetic (swapping modality factors). SDNet also demonstrates that its modality factor alone can predict the input imaging modality with high accuracy.
Architecture
Hybrid
Four-network architecture: a U-Net anatomy encoder producing a multi-channel spatial anatomical factor, a convolutional modality encoder producing a non-spatial latent modality vector (regularized as in a VAE), a segmentor operating on the anatomy factor, and a decoder that reconstructs the input image from both factors using FiLM normalization
Framework
Keras
Added to catalog
2026-08-14
Cardiac MRI and CT multi-modal cohorts (SDNet)
Cardiac MRI and CT datasets used for disentangled representation learning, semi-supervised segmentation and cross-modality synthesis; exact constituent cohorts are specified only in the source publication.
Disentangled anatomy/modality factorisation of cardiac MRI/CT enabling semi-supervised segmentation, multi-task regression, and cross-modality image synthesis