Image and audio input
- Input requires one image URL and one speech audio URL.
- The image can be JPEG, PNG or WEBP up to 10MB.
- Audio can be MP3, WAV, AAC, MP4 or OGG style formats up to 10MB.
UlazAI - AI Image & Video Tools
Premium talking portrait model
Turn a portrait, character image or selected subject into a speaking video driven by a speech audio file. This is the route to use when identity, expression, mouth timing and upper-body performance matter more than inventing a full scene from scratch.
27 credits/s. Studio output is configured for 1080.
Founder videos, product explainers and localized spokesperson clips.
Character-led education where the same face needs multiple scripts.
Portrait animation where mouth movement, expression and audio timing matter.
AI creator, avatar and campaign clips that need one repeatable face across many messages.
Model video
Use this as the quick model check before generating: accepted input, generated output and the clearest route in Video Studio.
Input Pick the route from the asset you already have.
Sample Watch generated clips before you spend credits.
Route Clear prompt, clear pricing, clear next step.
OmniHuman is for audio-driven character performance. The source image gives identity; the speech track drives timing, mouth movement and much of the expression.
| Duration | 720p / base | 1080p |
|---|---|---|
| 5s | 135 credits | - |
| 10s | 270 credits | - |
| 30s | 810 credits | - |
| 60s | 1620 credits | - |
Use OmniHuman when the marketing asset is the person or character, not the background scene.
Create a short founder, expert or brand-presenter clip from one portrait and the exact spoken script.
Reuse the same face across translated audio tracks while keeping the clip focused on speech and expression.
Build multiple micro-lessons with the same character image and different clean voice tracks.
Create repeatable creator-style clips where the same identity delivers new hooks, offers or announcements.
omnihuman_1_5
Estimated from selected length; voice track controls pacing
Requires one portrait image URL, one speech audio URL and a short behavior prompt.
Open routeOmniHuman starts from a portrait image and speech audio. Lip sync starts from an existing video and changes the mouth timing to match new audio.
One portrait image URL, one speech audio URL and a prompt for expression or behavior.
No. The Studio route is the video generation route; detection helpers are separate checks and are not exposed as video models.
Audio must be below 60 seconds. For stronger facial quality, keep clips around 15 seconds or shorter.
Use JPEG, PNG or WEBP image URLs up to 10MB.
A mask helps select one subject when the image contains multiple people or when the background should not drive the animation.
Fast mode speeds up generation but can reduce detail and motion quality.
The route costs 27 credits per output second.