CVAI Catalog

·

View Catalog

CMR-Transformer

Stanford University / University of Pennsylvania / UCSF / Georgetown (Shad, Zakka, Hiesinger et al.)

Cardiac MRI

Filter catalog by Modality:
Cardiac MRI

Clinical text

Filter catalog by Modality:
Text & EHR

General Purpose / Multi-task

Filter catalog by Disease / Trait:
General / Foundation

LVEF estimation

Filter catalog by Disease / Trait:
Cardiac Function & Hemodynamics

Heart failure

Filter catalog by Disease / Trait:
Cardiac Function & Hemodynamics

Embedding

Filter catalog by Task Type:
Representation Learning

Regression

Filter catalog by Task Type:
Regression

Binary classification

Filter catalog by Task Type:
Classification

Hybrid

Filter catalog by Architecture:
Hybrid / Multi-branch

PyTorch

Filter catalog by Framework:
PyTorch

CC BY-NC 4.0

Filter catalog by License:
Non-commercial / Research-only

Foundation vision-language model for cardiac MRI that learns pathophysiological visual representations directly from the natural-language radiology reports accompanying each scan, rather than from hand-labeled targets. A Multi-scale Vision Transformer (MViT, Kinetics-400-initialized) video encoder for cine CMR sequences is contrastively pretrained (InfoNCE) against a PubMed-pretrained BERT text encoder over 19,041 multi-institutional CMR studies. The frozen vision encoder transfers with strong performance to left-ventricular ejection-fraction regression (MAE 3.34% on a UK Biobank hold-out of ~4,259-45,623 participants) and detecting HFrEF (LVEF<40%, AUC 0.880), and the paper reports emergent zero-/few-shot performance across 39 cardiac and non-cardiac conditions including cardiac amyloidosis and hypertrophic cardiomyopathy. Code and pretrained MViT encoder weights are both released (Hugging Face, CC BY-NC 4.0).

memory Specifications

category

Architecture

Hybrid

Multi-scale Vision Transformer (MViT) video encoder, Kinetics-400 pretrained, contrastively aligned (InfoNCE) against a PubMed-pretrained BERT text encoder over paired CMR cine sequences and radiology reports

code

Framework

PyTorch

calendar_month

Added to catalog

2026-08-10

gavel License

CC BY-NC 4.0

check_small Open source close_small No commercial use check_small Modifications allowed Attribution required No share-alike requirement

License for model weights only. Associated code may be licensed seperately, check code source for specific terms.

description Publication

A Generalizable Deep Learning System for Cardiac MRI open_in_new

Nature Biomedical Engineering · 2026 · original paper

DOI: 10.1038/s41551-026-01637-3

database Training & evaluation data

Stanford / UPenn / UCSF / Georgetown Cardiac MRI Cohort (CMR-Transformer)

train

USA

19,041 cardiac MRI scans with accompanying radiology reports pooled from four large US academic medical centers, used to pretrain a contrastive vision-language CMR encoder.

public 74,916 subjects · United Kingdom

UK Biobank cardiac MRI imaging substudy; CineMA pretrained on 74,916 cine CMR studies, ukbb_cardiac (Bai et al. 2018) trained on ~4,875 subjects / 93,500 annotated images from an earlier release.

science Capabilities & performance

Contrastively-learned embedding of cine cardiac MRI sequences, transferable to downstream regression/classification tasks

Embedding General Purpose / Multi-task

Left ventricular ejection fraction (LVEF) regression from cine CMR

Regression LVEF estimation
3.344 MAE UK Biobank hold-out (LVEF regression) · internal

Binary classification of heart failure with reduced ejection fraction (HFrEF, LVEF<40%) from cine CMR

Binary classification Heart failure
0.88 (0.835–0.925) AUROC UK Biobank hold-out, n=4259 (HFrEF classification) · internal