University of Toronto / Vector Institute (Bo Wang Lab)
Echocardiography foundation model trained with a latent-predictive (V-JEPA2-style) self-supervised objective rather than pixel reconstruction, pretrained on 18 million echocardiograms from 300,000 patients drawn from the public MIMIC-IV-ECHO dataset plus a private multi-site archive - reportedly the largest echo pretraining corpus assembled to date. With a frozen backbone and only lightweight added layers, it outperforms prior echo foundation models by roughly 20% on ejection-fraction estimation and 17% on right-ventricular pressure estimation, reaches strong view-classification accuracy using just 1% of labels, and transfers zero-shot to pediatric echo better than fully fine-tuned baselines. Developed by the University of Toronto's Bo Wang Lab.
Architecture
Vision Transformer
V-JEPA2 video joint-embedding predictive architecture; ViT-L/16 (300M), ViT-H/16 (600M), and ViT-g/16 (1B) encoder configs
Framework
PyTorch
Added to catalog
2026-07-10
Apache 2.0
License for model weights only. Associated code may be licensed seperately, check code source for specific terms.
~525,000 echocardiogram video clips linked to MIMIC-IV at Beth Israel Deaconess Medical Center; used as the public reproducible pretraining split for EchoJEPA.
View classification (1% labeled data)
LVEF estimation, pediatric transfer, zero-shot (EchoJEPA-G)
LVEF estimation, pediatric transfer, fine-tuned (EchoJEPA-G)
Embedding