Vidu Q3

vidu-q3

Vidu Q3

vidu-q3

Vidu Q3 text-to-video and image-to-video with high visual fidelity, multiple styles, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.

Added Jan 31, 2026

Approx. Price

$0.350 per video

Model Type

both

Settings

Generation controls available for this model.

Output Format

up to 1,080p
Square
Portrait
Landscape

Default Duration

5

16 duration options

Aspect Ratio (T2V only)

Select

Default

4:3

Options (5)

Landscape (16:9), Standard (4:3), Square (1:1), Portrait (3:4) +1 more

Applies to text-to-video

Background Music

Toggle

Default

Yes

Add background music

Duration

Select

Default

5

Options (16)

1 second, 2 seconds, 3 seconds, 4 seconds +12 more

Video length in seconds (1-16)

Generate Audio

Toggle

Default

Yes

Generate synchronized audio

Motion

Select

Default

auto

Options (4)

Auto, Small, Medium, Large

Movement intensity

Resolution

Select

Default

720p

Options (3)

540p, 720p, 1080p

Output resolution

Style (T2V only)

Select

Default

general

Options (2)

General, Anime

Visual style for text-to-video

Benchmarks

No benchmark data is available yet for this model.

Compare Vidu Q3 with similar models from the same provider or model family.

Vidu Q3 Pro

vidu-q3-pro

Vidu Q3 Pro text-to-video, image-to-video, and start/end-frame video generation with high visual fidelity, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.

Vidu Q1

vidu-video

Vidu Q1 video generation model. Creates high-quality 5-second videos. Supports both text-to-video and image-to-video generation with customizable visual styles (general or anime), movement amplitude control, and fixed 16:9 output.

P-Video Edit

pruna-ai/p-video/edit

Instruction-based editing for videos up to 15 seconds, with optional reference-image guidance, Draft and Full quality modes, prompt enhancement, and source-audio preservation.

MiniMax H3 Max Turbo

minimax/h3-max-turbo

MiniMax H3 Max Turbo brings H3 Max prompt understanding and aesthetics to video generation at roughly twice the speed and half the cost while targeting 97% of its quality. Create 5–15 second 480p or 768p clips from text or a starting image, with optional first/last-frame transitions.

MiniMax H3 Spicy Image-to-Video

wavespeed-ai/minimax-h3/image-to-video-spicy

Open-weights MiniMax H3 image-to-video generation with expressive unrestricted motion, native stereo audio, optional last-frame control, 3–15 second clips, and 480p or 768p output.

Gemini Omni Flash 1.1

google/gemini-omni-flash/v1.1

Google’s multimodal video model for text-to-video, image animation with optional end frames, multimodal reference generation, and instruction-based video editing. Generates synchronized native audio at resolutions from 360p through 4K.