CVAI Catalog

·

View Catalog

EchoJEPA

ViT-L

University of Toronto / Vector Institute (Bo Wang Lab)

Echocardiography video

Filter catalog by Modality:
Echocardiography

Echocardiographic view classification

Filter catalog by Disease / Trait:
General / Foundation

LVEF estimation

Filter catalog by Disease / Trait:
Cardiac Function & Hemodynamics

General Purpose / Multi-task

Filter catalog by Disease / Trait:
General / Foundation

Multi-class classification

Filter catalog by Task Type:
Classification

Regression

Filter catalog by Task Type:
Regression

Embedding

Filter catalog by Task Type:
Representation Learning

Vision Transformer

Filter catalog by Architecture:
Transformer

PyTorch

Filter catalog by Framework:
PyTorch

Apache 2.0

Filter catalog by License:
Permissive

Echocardiography foundation model trained with a latent-predictive (V-JEPA2-style) self-supervised objective rather than pixel reconstruction, pretrained on 18 million echocardiograms from 300,000 patients drawn from the public MIMIC-IV-ECHO dataset plus a private multi-site archive - reportedly the largest echo pretraining corpus assembled to date. With a frozen backbone and only lightweight added layers, it outperforms prior echo foundation models by roughly 20% on ejection-fraction estimation and 17% on right-ventricular pressure estimation, reaches strong view-classification accuracy using just 1% of labels, and transfers zero-shot to pediatric echo better than fully fine-tuned baselines. Developed by the University of Toronto's Bo Wang Lab.

memory Specifications

category

Architecture

Vision Transformer

V-JEPA2 video joint-embedding predictive architecture; ViT-L/16 (300M), ViT-H/16 (600M), and ViT-g/16 (1B) encoder configs

code

Framework

PyTorch

calendar_month

Added to catalog

2026-07-10

gavel License

Apache 2.0

check_small Open source check_small Commercial use OK check_small Modifications allowed Attribution required No share-alike requirement

License for model weights only. Associated code may be licensed seperately, check code source for specific terms.

description Publication

database Training & evaluation data

USA

~525,000 echocardiogram video clips linked to MIMIC-IV at Beth Israel Deaconess Medical Center; used as the public reproducible pretraining split for EchoJEPA.

science Capabilities & performance

View classification (1% labeled data)

Multi-class classification Echocardiographic view classification
0.786 Accuracy Internal multi-view echo classification benchmark · internal

LVEF estimation, pediatric transfer, zero-shot (EchoJEPA-G)

Regression LVEF estimation
4.32 MAE EchoNet-Pediatric · external

LVEF estimation, pediatric transfer, fine-tuned (EchoJEPA-G)

Regression LVEF estimation
3.88 MAE EchoNet-Pediatric · external

Embedding

General Purpose / Multi-task