Image to Video Leaderboard (With Audio)Artificial Analysis

Added to the leaderboard in the last month:

Vidu Q3 Turbo, MiniMax H3, MAGI-2 Preview

Range
Creator
Model
Elo
95% CI
Samples
Released
API Pricing 1
11-2
ByteDance Seed logoByteDance Seed
Dreamina Seedance 2.0 720p
1,198-7/712,597Mar 2026$9.07 /min
21-3
MiniMax logoMiniMax
MiniMax H3Hugging FaceOpen Weights
1,192-9/95,756Jul 2026$7.80 /min
31-3
Google logoGoogle
Gemini Omni Flash
1,191-9/97,085May 2026$6.00 /min
44-6
SpaceXAI logoSpaceXAI
grok-imagine-video-1.5
1,114-8/85,139May 2026$8.40 /min
54-6
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.1
1,113-8/88,033Jun 2026$9.90 /min
64-6
Sand.ai logoSand.ai
MAGI-2 PreviewHugging FaceOpen Weights
1,108-8/810,778Aug 2026Coming soon
77-10
Alibaba logoAlibaba
Wan 2.7
1,091-8/84,462Apr 2026$9.00 /min
87-10
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.0
1,090-8/87,791Apr 2026$13.20 /min
97-10
Skywork AI logoSkywork AI
SkyReels V4
1,089-8/84,742Mar 2026$21.00 /min
107-11
Google logoGoogle
Veo 3.1
1,086-7/77,893Jan 2026$24.00 /min
1110-13
SpaceXAI logoSpaceXAI
grok-imagine-video
1,079-7/712,600Jan 2026$4.20 /min
1211-14
KlingAI logoKlingAI
Kling 3.0 1080p (Pro)
1,077-7/712,194Feb 2026$20.16 /min
1311-14
Google logoGoogle
Veo 3.1 Fast
1,076-7/711,556Jan 2026$9.00 /min
1412-17
PixVerse logoPixVerse
PixVerse V6
1,070-7/79,215Mar 2026$6.90 /min
1514-18
KlingAI logoKlingAI
Kling 3.0 720p (Standard)
1,067-7/712,215Feb 2026$15.60 /min
1614-18
Google logoGoogle
Veo 3.1 Lite
1,066-8/87,193Mar 2026$4.80 /min
1714-18
Vidu logoVidu
Vidu Q3 Pro
1,064-7/712,580Jan 2026$9.60 /min
1815-18
KlingAI logoKlingAI
Kling 3.0 Omni 1080p (Pro)
1,062-7/77,048Feb 2026$16.80 /min
1919
KlingAI logoKlingAI
Kling 3.0 Omni 720p (Standard)
1,055-7/76,949Feb 2026$13.44 /min
2020
Vidu logoVidu
Vidu Q3 Turbo
1,043-10/102,387Feb 2026$3.90 /min
2121-22
KlingAI logoKlingAI
Kling 2.6 Pro (January)
1,003-8/87,441Jan 2026$8.40 /min
2222
ByteDance Seed logoByteDance Seed
Seedance 1.5 pro
1,0000/08,969Dec 2025$11.86 /min
2323-25
Lightricks logoLightricks
LTX-2.3 FastHugging FaceOpen Weights
958-7/78,866Mar 2026$2.40 /min
2423-25
Lightricks logoLightricks
LTX-2.3 ProHugging FaceOpen Weights
958-7/78,968Mar 2026$4.80 /min
2523-25
PixVerse logoPixVerse
PixVerse V5.6
956-8/86,137Feb 2026Coming soon
2626-27
Lightricks logoLightricks
LTX-2 FastHugging FaceOpen Weights
934-8/85,908Oct 2025$2.40 /min
2726-27
Sapiens AI logoSapiens AI
Agnes-Video-V2.0
928-9/95,314May 2026$0.30 /min
2828
Alibaba logoAlibaba
Wan 2.6
895-8/86,068Dec 2025$9.00 /min
2929
Lightricks logoLightricks
LTX-2 ProHugging FaceOpen Weights
878-9/95,828Oct 2025$3.60 /min

1 API Pricing reflects the cost to generate 1 minute of 1080p video on the model creator's API at the model's default settings

Frequently Asked Questions

Dreamina Seedance 2.0 720p currently leads among Image to Video models with audio output in the Artificial Analysis Image to Video Arena with an Elo score of 1198.

The top Image to Video models with audio by Elo rating are: 1. Dreamina Seedance 2.0 720p (Elo 1198), 2. MiniMax H3 (Elo 1192), 3. Gemini Omni Flash (Elo 1191), 4. grok-imagine-video-1.5 (Elo 1114), 5. HappyHorse-1.1 (Elo 1113). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models with audio in the Artificial Analysis Image to Video Arena with an Elo score of 1192, followed by MAGI-2 Preview (Elo 1108) and LTX-2.3 Fast (Elo 958).

Gemini Omni Flash currently leads the Artificial Analysis Image to Video Arena (without audio) with an Elo score of 1368.

The top Image to Video models without audio by Elo rating are: 1. Gemini Omni Flash (Elo 1368), 2. MiniMax H3 (Elo 1351), 3. Dreamina Seedance 2.0 720p (Elo 1338), 4. grok-imagine-video-1.5 (Elo 1330), 5. grok-imagine-video (Elo 1326). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models without audio in the Artificial Analysis Image to Video Arena with an Elo score of 1351, followed by Cosmos3-Super-Image2Video-4Step (Elo 1265) and Cosmos3-Super-Image2Video (Elo 1244).

Text to Video models generate videos from text descriptions alone, while Image to Video models take an existing image as input and animate or extend it into a video. This allows for more control over the visual content and style of the generated video.

Models are ranked using an Elo rating system derived from user votes in blind comparisons. Users compare videos generated from the same input image and choose the result they prefer. Higher Elo scores indicate a model is preferred more often. Vote in the Artificial Analysis Video Arena