Mohamed Bin Zayed University of Artificial Intelligence (BioMedIA)
Self- and weakly-supervised pipeline for left-ventricle segmentation across the full cardiac cycle in apical-4-chamber echocardiography videos. A video segmentation network (2D super-image or 3D U-Net encoder) is first pretrained with a self-supervised temporal-masking objective on largely unannotated echo frames, then fine-tuned with weak supervision from the sparse end-diastole/end-systole frame labels that most echo datasets provide. Achieves 93.3% Dice on EchoNet-Dynamic, outperforming nnU-Net and non-SSL baselines, and generalizes to the external CAMUS dataset. Developed by the BioMedIA group at MBZUAI.
Architecture
CNN (3D)
3D U-Net video encoder with self-supervised temporal-masking pretraining, followed by weakly-supervised fine-tuning on sparse ED/ES frame annotations; also supports a 2D 'super-image' encoder variant
Framework
PyTorch
Added to catalog
2026-08-10
CC BY-NC 4.0
License for model weights only. Associated code may be licensed seperately, check code source for specific terms.
500 patients imaged with 2D transthoracic echocardiography (A2C/A4C views) at University Hospital of St Etienne; roughly half with LVEF < 45%.
10,030 deidentified apical-4-chamber echo videos from Stanford Health Care. Source reports age and sex breakdowns.
Left ventricle segmentation across the full 2D+time echocardiogram video, used to derive ED/ES volumes and ejection fraction