On this page
Overview
Vidu API currently supports multiple functions, covering various development scenarios such as video, audio, images, and creative applications. The function list is continuously updated:
Primary | Secondary | Function Details |
|---|---|---|
Video | Generate video from input image and prompts - Supports free selection of video ratio, duration, resolution, and motion amplitude - Supports adding background music - Supports multiple generation modes | |
Use input images’ subjects or scenes as references to generate video - Supports uploading multiple images - Supports free selection of video ratio, duration, resolution, and motion amplitude - Supports adding background music - Supports multiple generation modes | ||
Input first and last frame, generate video with prompts - Supports free selection of video ratio, duration, resolution, and motion amplitude - Supports adding background music - Supports multiple generation modes | ||
Generate video directly from prompts - Supports long text input - Supports free selection of video ratio, duration, resolution, and motion amplitude - Supports adding background music - Supports multiple generation modes | ||
Image | Use input images’ subjects or scenes as references to generate image - Characters, objects, and scenes in images remain highly consistent - Supports free selection of image ratio | |
Generate image directly from prompts - Supports long text input - Supports free selection of image ratio, resolution | ||
Refer to the original input image and the reference image, replace the elements in the original image -Support uploading multiple images -Support replacing original image elements with text descriptions -Support free selection of image ratio and resolution,supporting 1080p, 2K and 4K | ||
Audio | Generate audio directly from prompts - Supports generating creative sound effects or background music - Supports free selection of audio duration | |
Generate controllable sequential sound effects from prompts - Use a timeline to flexibly control the order and duration of multiple sound effects - Supports overlapping multiple sound effects - Supports free selection of audio duration | ||
Generate speech from text - Control speech speed, pauses, volume, and polyphonic characters - Adjust emotional tone (e.g., joyful, sad, angry) - Supports 30+ languages and 300+ voices | ||
Upload audio to clone voice timbre - Use cloned voices anytime for speech synthesis - Adjust emotional tone freely | ||
Template | Apply templates to images to generate videos - Includes hundreds of creative templates: e-commerce, virtual try-on, transformation, etc. | |
Apply long template story to images to generate videos - Storylines are coherent and effects are diverse | ||
Others | Make video character lip-sync to the specified text - Supports hundreds of voices and languages - Supports long text input - Supports adjusting volume freely | |
Generate a digital-human video from an image - Suitable for speeches, news-reading, and creative scenarios - Supports 300+ voices or cloned voices | ||
Replace target characters or objects within existing videos - Input a video and a subject image to replace the original subject - Accurately recognizes and replaces people, items, animals, etc. | ||
Extend the duration of videos - Supports extension from 1~7s | ||
Generate storyline-consistent, character-consistent videos from multiple keyframes - Flexible control of frame-to-frame duration (2–7s) - Each frame can be guided by prompts | ||
Generate recommended prompts from images - Supports selecting prompt type and quantity - Supports selecting video resolution | ||
Enhance video resolution - Supports multiple resolutions: 1080p, 2K, 4K, 8K |