🗣️ LongCat-Video-Avatar 1.5
Make any face talk. Drop in a reference image, a voice clip, and a short prompt — get back a ~5 second lip-synced video of that character speaking.
Powered by LongCat-Video-Avatar 1.5 (Meituan, MIT): Whisper-Large-v3 audio encoding → INT8 LongCat-Video DiT → DMD2-distilled 8-step sampling. Works on photoreal humans, anime, and animals.
💡 Tips · Longer, more descriptive prompts hold identity better. Pick Isolate vocals when the audio has music behind it. Exact 8-step is the highest fidelity; DBCache faster is quickest.
⏱️ ZeroGPU quota — the defaults above (480p · DBCache faster) reserve ~240s, which fits a signed-in free account's 5 min/day. 720p and Exact 8-step reserve considerably more and need PRO-level quota; if you see “requested GPU duration is larger than the maximum allowed”, step back down to the defaults.