Skip to content

UlazAI - AI Image & Video Tools

Premium talking portrait model

OmniHuman 1.5 talking portrait generator

Turn a portrait, character image or selected subject into a speaking video driven by a speech audio file. This is the route to use when identity, expression, mouth timing and upper-body performance matter more than inventing a full scene from scratch.

27 credits/s. Studio output is configured for 1080.

Founder videos, product explainers and localized spokesperson clips.

Character-led education where the same face needs multiple scripts.

Portrait animation where mouth movement, expression and audio timing matter.

AI creator, avatar and campaign clips that need one repeatable face across many messages.

Model video

Input, sample and route in one launch clip

Use this as the quick model check before generating: accepted input, generated output and the clearest route in Video Studio.

Input Pick the route from the asset you already have.

Sample Watch generated clips before you spend credits.

Route Clear prompt, clear pricing, clear next step.

Why OmniHuman fits talking video

OmniHuman is for audio-driven character performance. The source image gives identity; the speech track drives timing, mouth movement and much of the expression.

Image and audio input

  • Input requires one image URL and one speech audio URL.
  • The image can be JPEG, PNG or WEBP up to 10MB.
  • Audio can be MP3, WAV, AAC, MP4 or OGG style formats up to 10MB.

Recommended audio length

  • The audio must stay below 60 seconds.
  • For stronger facial quality, keep clips around 15 seconds or shorter.
  • Use clean speech with the exact script; noisy music beds reduce reliability.

Mask and subject helpers

  • Human identification and subject detection are helper checks, not separate video models.
  • Use a mask when an image contains multiple people or when only one subject should speak.
  • A seed can help repeat similar results when the same image, audio and prompt are reused.

Quality and speed controls

  • Output can be 720p or 1080p, with 1080p used for most polished presenter clips.
  • Fast mode trades some quality for quicker generation.
  • The behavior prompt should be short: tone, emotion, posture and natural motion.

Where it fits

Duration 720p / base 1080p
5s 135 credits -
10s 270 credits -
30s 810 credits -
60s 1620 credits -

Strong use cases

Use OmniHuman when the marketing asset is the person or character, not the background scene.

Spokesperson explainers

Create a short founder, expert or brand-presenter clip from one portrait and the exact spoken script.

Localized product messages

Reuse the same face across translated audio tracks while keeping the clip focused on speech and expression.

Character-led education

Build multiple micro-lessons with the same character image and different clean voice tracks.

Avatar and creator campaigns

Create repeatable creator-style clips where the same identity delivers new hooks, offers or announcements.

Studio routes

Portrait Image to Talking Video

omnihuman_1_5

Estimated from selected length; voice track controls pacing

Requires one portrait image URL, one speech audio URL and a short behavior prompt.

Open route

How it differs from lip sync

OmniHuman starts from a portrait image and speech audio. Lip sync starts from an existing video and changes the mouth timing to match new audio.

Production workflow

  1. Choose a clear portrait or character image with good face visibility and minimal visual clutter.
  2. If the image has several subjects, run a subject check and use a mask for the person that should speak.
  3. Prepare clean speech audio with the final script, preferably 15 seconds or shorter for best quality.
  4. Write a short behavior prompt with tone, expression, posture and natural movement.
  5. Use 1080p for polished clips, or fast mode when you are still testing scripts.
  6. Reuse the same portrait, seed and prompt style when you need a consistent presenter series.

FAQ

What does OmniHuman need?

One portrait image URL, one speech audio URL and a prompt for expression or behavior.

Is the helper detection tooling a video model?

No. The Studio route is the video generation route; detection helpers are separate checks and are not exposed as video models.

How long can the audio be?

Audio must be below 60 seconds. For stronger facial quality, keep clips around 15 seconds or shorter.

Which image formats work?

Use JPEG, PNG or WEBP image URLs up to 10MB.

What does the mask input do?

A mask helps select one subject when the image contains multiple people or when the background should not drive the animation.

What does fast mode do?

Fast mode speeds up generation but can reduce detail and motion quality.

What does it cost?

The route costs 27 credits per output second.

Music Studio Songs, instrumentals and your own lyrics.