All-in-one video generation model supporting text-to-video, image-to-video, first/last-frame, and omni-modal reference-based generation (image, video, and audio references) with synchronized audio, at up to 30 seconds per clip.
All-in-one video generation model supporting text-to-video, image-to-video, first/last-frame, and omni-modal reference-based generation (image, video, and audio references) with synchronized audio, at up to 30 seconds per clip.

A complete guide to Seedance 2.0: ByteDance's multimodal AI video model — covering architecture, core features, and practical prompts all in one place.

Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.

Kling v3 (Video 3.0) by Kuaishou: native 4K, 60fps, multi-shot cuts, multilingual audio. Full wiki covering specs, pricing, prompts & comparisons with Runway Gen-4 and Veo 3.1.