Wan 3.0 AI Video Generator API documentation
Generate videos with Wan 3.0 AI Video Generator (Alibaba). The reference mode accepts image, video, audio inputs.
Endpoints
| Contract item | Value |
|---|---|
| Capability | wan-3-0-ai-video-generator |
| Submit | POST https://vidmage.ai/api/v1/wan-3-0-ai-video-generator/submit |
| Query | POST https://vidmage.ai/api/v1/wan-3-0-ai-video-generator/query |
| Task identifier | taskId |
| Result field | videoUrl |
Authenticate with Authorization: Bearer YOUR_API_KEY. Save the task identifier and query the same capability. Use a stable Idempotency-Key for submission retries.
Product guide: Wan Video API
Parameters and input rules
| Field | Type | Required | Default | Meaning and limits |
|---|---|---|---|---|
prompt | string | No | Not specified | Text prompt describing the video to generate (max 20000 chars). Optional only in provider modes whose image input fully defines the generation. Maximum characters: 20000 |
imageUrl | string | No | Not specified | Input image URL for image-to-video mode. Omit it only when supplying multimodal reference arrays instead. |
imageUrls | string[] | No | Not specified | Optional multimodal reference images. Public URLs only; 20 MB per file. Maximum items: 10 |
videoUrls | string[] | No | Not specified | Optional multimodal reference videos. Each video is 1–15 seconds and 100 MB max. Maximum items: 5 |
audioUrls | string[] | No | Not specified | Optional multimodal reference audio. Each audio is 1–15 seconds and 15 MB max. Maximum items: 5 |
lastFrameUrl | string | No | Not specified | Optional public URL for the final frame in image-to-video mode. |
duration | string | No | "5" | Video duration in seconds. Values: "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28", "29", "30" |
aspectRatio | string | No | "adaptive" | Output aspect ratio. Values: "adaptive", "16:9", "4:3", "1:1", "3:4", "9:16" |
resolution | string | No | "720p" | Output resolution tier. Values: "480p", "720p", "1080p" |
sound | boolean | No | true | Generate synchronized audio along with the video. No separate sound surcharge is configured. |
enableThinking | boolean | No | false | Enable deeper reasoning before generating the video. |
seed | integer | No | Not specified | Random seed; leave unset for a random result. Minimum: 0Maximum: 2147483647Multiple of: 1 |
image to video
These are public VidMage field names. The server maps them to provider fields. Input media selects the mode; use its limits together with the parameter table.
| Control | Mode-specific rule |
|---|---|
prompt | Optional; 0–20000 characters. |
duration | JSON string. Values: "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28", "29", "30". Default: 5. |
aspectRatio | Values: "adaptive", "16:9", "4:3", "1:1", "3:4", "9:16". Default: adaptive. |
resolution | Values: "480p", "720p", "1080p". Default: 720p. |
sound | Boolean. Default: true. |
imageUrl | Maximum 20 MB per file. Formats: jpg, jpeg, png, bmp, webp. |
lastFrameUrl | Optional final-frame image. |
enableThinking | Enable deeper reasoning before generating the video. Default: false |
seed | Random seed; leave unset for a random result. Minimum: 0Maximum: 2147483647Multiple of: 1 |
multimodal video
These are public VidMage field names. The server maps them to provider fields. Input media selects the mode; use its limits together with the parameter table.
| Control | Mode-specific rule |
|---|---|
prompt | Required; 1–20000 characters. |
duration | JSON string. Values: "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "28", "29", "30". Default: 5. |
aspectRatio | Values: "adaptive", "16:9", "4:3", "1:1", "3:4", "9:16". Default: adaptive. |
resolution | Values: "480p", "720p", "1080p". Default: 1080p. |
sound | Boolean. Default: true. |
imageUrls | Up to 10 files. Maximum 20 MB per file. Formats: jpg, jpeg, png, bmp, webp. |
videoUrls | Up to 5 files. Maximum 100 MB per file. Formats: mp4, mov. 1–15 seconds per file. Maximum total duration: 15 seconds. |
audioUrls | Up to 5 files. Maximum 15 MB per file. Formats: wav, mp3. 1–15 seconds per file. Maximum total duration: 15 seconds. |
| References | At least one reference input is required. |
enableThinking | Enable deeper reasoning for complex multimodal references. Default: false |
seed | Random seed; leave unset for a random result. Minimum: 0Maximum: 2147483647Multiple of: 1 |
Example request
{
"imageUrl": "https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp",
"prompt": "Use Image 1 for the character, Video 1 for the camera rhythm, and Audio 1 for the timing of a cinematic city walk",
"duration": "5",
"aspectRatio": "adaptive",
"resolution": "720p",
"sound": true,
"enableThinking": false
}# Set VIDMAGE_IDEMPOTENCY_KEY to a unique value for this task; preserve it for transport retries.
curl -X POST "https://vidmage.ai/api/v1/wan-3-0-ai-video-generator/submit" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Idempotency-Key: ${VIDMAGE_IDEMPOTENCY_KEY}" \
-H "Content-Type: application/json" \
-d '{"imageUrl":"https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp","prompt":"Use Image 1 for the character, Video 1 for the camera rhythm, and Audio 1 for the timing of a cinematic city walk","duration":"5","aspectRatio":"adaptive","resolution":"720p","sound":true,"enableThinking":false}'
# -> { "success": true, "taskId": "...", "creditsConsumed": ... }Task results and recovery
URL of the generated video. Read videoUrl from the completed query response. Preserve taskId while the task is running.
curl -X POST "https://vidmage.ai/api/v1/wan-3-0-ai-video-generator/query" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "taskId": "TASK_ID_FROM_SUBMIT" }'
# -> { "success": true, "data": { "status": "...", "videoUrl": "https://..." } }Submission and task queries report the available task and usage information. A timeout is not a confirmed failure. Recover the existing task instead of resubmitting.
For face selection and error recovery, follow Task lifecycle. MCP uses the normalized taskId argument in get_task_result, including capabilities whose REST identifier is requestId.
Credits
| Component / option | Rate | Minimum |
|---|---|---|
| 480p | 10 credits / second | — |
| 720p | 15 credits / second | — |
| 1080p | 30 credits / second | — |
Base charge = duration × the selected resolution rate.
Including reference video adds a flat charge: 480p: 50 credits; 720p: 75 credits; 1080p: 150 credits.
Use the Playground or the MCP credit estimator for your exact inputs. Credits and billing.
