Docs menu
On this page

Lipsync

The Lipsync node syncs mouth movement to an audio track. Use it to make a video's speaker match new dialogue, or to turn a still portrait into a talking video from an audio clip.

What it does

Lipsync works in one of two modes, depending on the model you pick:

  • Video driven: you give it a video and an audio track, and it matches the mouth movement in the video to the audio.
  • Image driven: you give it a still image (like a portrait) and an audio track, and it animates the image into a talking video.

Inputs and outputs

  • Image or Video: the source. The label switches between "Image" and "Video" depending on the model you choose.
  • Audio: every Lipsync model needs an audio track to sync to.
  • Prompt: a text box for describing the motion. Only shown for models that support it.
  • Flow: an optional input that controls when this node runs. It can only be fed by the Trigger node or a logic node like Sequence or Router.
  • Video (output): the finished, synced video.

If no video or image reaches the node, the run stops with a message asking you to connect a source. The same happens if audio is missing. If you connect a video to an image-only model (or the other way around), the run stops and tells you which kind of source that model needs.

Settings

  • Model: pick which Lipsync model to use (see below).
  • Resolution: shown only for models that offer a quality tier.
  • Model Settings: opens a panel with whichever settings apply to your chosen model, such as Quality, Sync Mode, Acceleration, Loop Mode, Turbo Mode, Frames, Guidance, and Seed. A note at the bottom reminds you that framing and aspect ratio always follow your input source, these settings only tune the model itself.
  • Prompt: a text box at the bottom of the card, shown only for models that support one. If a Prompt input is connected, this box is disabled and shows "Connected from input" instead.

Once a result exists, the Download button on the card becomes active.

Models

Lipsync models come from two providers. Which one you use depends on your workflow's provider setting.

fal.ai models:

  • Sync Lipsync 2.0 (video driven): Quality Standard or Pro, plus Sync Mode options (Cut off, Loop, Bounce, Silence, Remap).
  • VEED Lipsync (video driven): no extra settings.
  • LatentSync (video driven): a Guidance slider, Loop Mode, and Seed.
  • OmniHuman 1.5 (image driven): optional prompt, Resolution 720p or 1080p, and a Turbo Mode toggle. This is the default model.
  • InfiniteTalk (image driven): a required prompt, Resolution 480p or 720p, a Frames slider, Acceleration setting, and Seed.
  • VEED Fabric 1.0 (image driven): Resolution 720p or 480p.

Kie AI models:

  • InfiniteTalk (image driven): required prompt, Resolution 480p or 720p, and Seed.
  • Kling AI Avatar Standard and Kling AI Avatar Pro (image driven): required prompt, no other settings.

All Kie AI Lipsync models are image driven. If you connect a video source while using Kie AI, the run stops and points you to fal.ai instead.

Lipsync always runs on fal.ai unless you switch the node's provider to Kie AI. If you have not added a fal.ai key in Settings, add it there, or switch this node to Kie AI.

Tips

  • The preview on the card automatically resizes to match your finished video's real shape once it exists.
  • Switching models resets fields like resolution, quality, sync mode, acceleration, and frame count to the new model's own defaults. Toggle-style settings such as Turbo Mode and Seed are not reset, they carry over from whatever you last set.