Massachusetts General Hospital / Harvard Medical School (Kim et al.)
General-purpose vision foundation model for echocardiography, pretrained with a masked autoencoder combined with a periodic contrastive loss designed around the cyclical nature of cardiac motion. Validated on chamber segmentation, view classification, and disease detection, with its largest advantage over non-pretrained baselines and natural-image models like SAM appearing in low-label settings. Pretrained on roughly 290,000 echo clips from a mix of internal and public sources. Developed by Massachusetts General Hospital and Harvard Medical School.
Architecture
Vision Transformer
Echo-VideoMAE: ViT encoder-decoder masked autoencoder with spatio-temporal consistent masking plus a periodic-driven contrastive loss exploiting cardiac cycle periodicity
Framework
PyTorch
Added to catalog
2026-07-10
CC BY-NC-ND 4.0
License for model weights only. Associated code may be licensed seperately, check code source for specific terms.
MGH/BWH Echocardiography Cohort (EchoFM pretraining subset)
683,560 TTEs from 6,500 patients; part of larger ~290K-clip pretraining corpus (remainder from public sources)
Embedding