Yale School of Medicine (CarDS Lab)
Model weights not public. Contact creators for more information.
Multi-instance contrastive self-supervised learning framework for echocardiography video representation learning: two distinct videos from the same patient exam are treated as positive pairs (rather than augmentations of a single clip), and a frame-reordering pretext task on temporally shuffled frames encourages the 3D-CNN backbone to learn temporally coherent representations. After self-supervised pretraining on unlabeled echocardiograms, the backbone is efficiently fine-tuned for cardiac disease classification (severe aortic stenosis, left ventricular hypertrophy) using very few labeled examples, outperforming SimCLR, standard multi-instance SimCLR, and Kinetics-400-initialized baselines across a range of training-data ratios.
Architecture
CNN (3D)
Multi-instance contrastive self-supervised 3D-CNN pretraining (patient-level positive pairs) combined with a temporal frame-reordering pretext task, fine-tuned for downstream echo video classification
Added to catalog
2026-08-10
Yale-New Haven Health System Echocardiography Archive (EchoCLR)
Unlabeled echocardiography video archive from the Yale-New Haven Health System used for EchoCLR multi-instance contrastive self-supervised pretraining, with downstream fine-tuning cohorts for severe aortic stenosis and left ventricular hypertrophy classification.
Severe aortic stenosis classification from echocardiography video (fine-tuned from EchoCLR pretraining)
Left ventricular hypertrophy (LVH) classification from echocardiography video (fine-tuned from EchoCLR pretraining)