Z Image Base is a 6B-parameter model that supports text-to-image and optional reference image guidance. Includes negative prompts, strength control, and flexible sizing for photorealistic results.
Added Jan 28, 2026
Approx. Price
$0.017 per image
Model Type
both
Settings
Generation controls available for this model.
Images Per Run
Up to 1
Output images
Input Images
Up to 1
Reference/edit images accepted • Route max 30 MB
Output Sizes
Custom Resolution
Negative Prompt
Default
N/A
Details you don't want in the image (e.g., blur, distortion, low quality).
Number of Images
Default
1
Output Format
Default
jpeg
Options (3)
JPEG, PNG, WebP
Choose the output image format
Resolution
Default
1024*1024
Options (8)
1024*1024 (Square (1024x1024)), 1024*768 (Landscape (4:3)), 768*1024 (Portrait (3:4)), 1024*576 (Landscape (16:9)) +4 more
Seed
Default
-1
Control reproducibility (-1 for random).
Strength
Default
0.6
How strongly the reference image influences the output (0 = preserve, 1 = reimagine).
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Text to Image
#112 / 157
ELO
867.0
Appearances
5,660
95% CI
-8/8
Release Date 2026-01 · Matched as Z-Image Base
Artificial Analysis APIRelated image models
Compare Z Image Base with similar models from the same provider or model family.
Z Image Turbo Image-to-Image
z-image-turbo-image-to-imageZ Image Turbo Image-to-Image transforms a reference image with a controllable strength slider, from subtle enhancements to full reimagination.
Z Image Turbo
z-image-turboZ Image Turbo is a fast, high-quality image generation model optimized for speed. Generate detailed images with cinematic quality, film grain effects, and artistic styles.
Z Image Turbo LoRA
z-image-turbo-loraZ Image Turbo LoRA supports up to 3 LoRAs for custom styles, characters, or brand identity.
Stable Diffusion 3 Medium
sd3_base_medium.safetensorsExcels at photorealism, typography, and prompt following. Works best in 1024x1024.
MiniMax H3 Image Edit
wavespeed-ai/minimax-h3/image-editEdit images from up to nine references, preserving identity while changing scenes, outfits, or styles. Supports 1K and 2K output.
MiniMax H3 Image
wavespeed-ai/minimax-h3/text-to-imageGenerate photorealistic and cinematic images from text at 1K or 2K resolution, with fifteen aspect ratios.