Four generation modes
Use text only, one first frame, first plus last frame, or multimodal references. First/last frames cannot be mixed with the multimodal arrays.
UlazAI - AI Image & Video Tools
New flagship video model
Generate coherent 4-30 second shots from text, fixed opening and ending frames, or a large set of image, video and audio references. Use it when a short draft model cannot hold enough characters, motion cues or product detail across the full scene.
480p: 17 credits/s with video input or 28 credits/s without video input. 720p: 38 credits/s with video input or 63 credits/s without video input. Video input is billed over measured input plus output duration.
Longer travel, product and narrative shots up to 30 seconds.
Four-character scenes and repeated identities guided by references.
Ad remixes where framing, motion or audio cues must stay recognizable.
Reference test
This second clip is a real video-reference render. It transfers the source motion and visual rhythm to a new subject, while Video Studio measures the source duration for billing.
Length Choose any duration from 4 to 30 seconds.
References Combine up to 30 images, 10 videos and 10 audio files.
Billing Video input uses measured input time plus output time.
Generated sample
Input: text-to-video, 12s, 720p, 16:9, MP4, no generated audio.
Prompt
One continuous cinematic shot through a rain-lit night market. Start wide above glowing food stalls, descend smoothly to street level and track beside four friends walking through the crowd. Their faces, clothing and relative positions remain consistent. Steam, reflections and umbrellas create depth. Finish with a controlled orbit as they stop beneath a red lantern. Natural movement, realistic skin, coherent hands, no text, no logos, no cuts.
The useful gain is control: longer output, larger reference sets and explicit endpoints. The Studio only shows controls the current API accepts.
Use text only, one first frame, first plus last frame, or multimodal references. First/last frames cannot be mixed with the multimodal arrays.
Add up to 30 images, 10 videos and 10 audio files. Video references must be 2-30 seconds and can total 30 seconds.
Choose 480p or 720p, MP4 or MOV, optional generated audio and optional web search. A separate last-frame file is best effort and may be absent even when the video succeeds.
Without video input, only output time is billed. With video references, UlazAI measures their duration and bills input time plus output time.
Compare prompt-only output with video-guided jobs. Video references use their measured source duration plus the selected output duration.
| Billing example | 480p | 720p |
|---|---|---|
| 4s, no video input | 112 credits | 252 credits |
| 12s, no video input | 336 credits | 756 credits |
| 30s, no video input | 840 credits | 1890 credits |
| 10s video + 30s output | 680 credits | 1520 credits |
| 30s video + 30s output | 1020 credits | 2280 credits |
Use 2.5 when continuity or duration is the reason a cheaper test fails. Mini remains the better choice for quick hook exploration.
Keep a face, outfit and camera language recognizable across locations by reusing the same reference set.
Guide a cast with separate images and describe blocking, action and camera position clearly.
Use product, motion and audio references when a proven edit needs a new item without losing its rhythm.
Prompt explicit moves such as FPV, orbit, dolly zoom or bullet time and lock endpoints when the finish matters.
seedance_2_5
4 to 30 seconds
Prompt only, with 480p or 720p output, optional generated audio, web search and MP4 or MOV.
Open routeseedance_2_5
4 to 30 seconds
Start from one image when the opening composition must stay fixed.
Open routeseedance_2_5
4 to 30 seconds
Provide both endpoint images when the shot must land on a controlled final composition.
Open routeseedance_2_5
4 to 30 seconds
Combine up to 30 images, 10 videos and 10 audio files. Frame inputs and multimodal references are separate modes.
Open routeChoose 2.5 for 16-30 second output, larger reference sets and scenes that must hold more identities or cues. Choose Mini when you need many inexpensive 4-15 second variants before selecting a final concept.
No. The current route exposes 480p and 720p. UlazAI does not advertise a 4K control that the API does not provide.
Up to 30 images, 10 videos and 10 audio files in multimodal mode, for 50 files combined.
No. Endpoint-frame mode and multimodal reference mode are separate input contracts.
The with-video rate applies to the measured reference-video duration plus the selected output duration.
MP4 and MOV.
Yes. Turn sound on when the clip needs native audio, or leave it off when you will add sound in the edit.