Vidu Q2 Reference-to-Image generates images based on 1-7 reference images with customizable prompts. Ideal for keeping product, character, or actor identity consistent across shots.
Added Dec 2, 2025
Approx. Price
$0.040 per image
Model Type
image-to-image
Settings
Generation controls available for this model.
Images Per Run
Up to 7
Output images
Input Images
Up to 7
Reference/edit images accepted • Route max 30 MB
Output Sizes
Aspect Ratio
Default
auto
Options (9)
Auto (match references), Square, 16:9 Widescreen, 9:16 Vertical +5 more
Canvas shape for generation. Use "auto" to match reference images.
Number of Images
Default
1
Resolution
Default
1080p
Options (3)
1080p (Fast preview (1920x1080)), 2K (Higher detail (2560x1440)), 4K (Maximum sharpness (3840x2160))
Seed
Default
-1
Control reproducibility (-1 for random).
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Related image models
Compare Vidu Q2 Reference with similar models from the same provider or model family.
Vidu Q2
vidu-q2Vidu Q2 is a high-end text-to-image model with cinematic lighting and clean composition. Supports up to 4K resolution and flexible aspect ratios. Upload reference images to guide generation with subject/composition consistency.
MiniMax H3 Image Edit
wavespeed-ai/minimax-h3/image-editEdit images from up to nine references, preserving identity while changing scenes, outfits, or styles. Supports 1K and 2K output.
MiniMax H3 Image
wavespeed-ai/minimax-h3/text-to-imageGenerate photorealistic and cinematic images from text at 1K or 2K resolution, with fifteen aspect ratios.
MAI-Image-2.6
microsoft/mai-image-2.6Microsoft’s image model for detailed photography, product imagery, accurate lettering, and precise edits to an uploaded image.
MAI-Image-2.6 Flash
microsoft/mai-image-2.6-flashA faster, lower-cost MAI image model for product visuals, creative drafts, and edits to an uploaded image.
Bria FIBO Generate 1.5
bria/fibo-generate-1.5/text-to-imageBria FIBO Generate 1.5 creates commercially safe images from text prompts, with optional reference-image guidance, photoreal styling, structured prompts, and up to 4MP output.