AI Talking Photo API documentation
Create a talking video from a portrait, a script, and a reference voice. The script supplies all spoken content. Reference-audio duration does not affect the credit quote; call estimate_capability_credits for an exact estimate.
Endpoints
| Contract item | Value |
|---|---|
| Capability | ai-talking-photo |
| Submit | POST https://vidmage.ai/api/v1/ai-talking-photo/submit |
| Query | POST https://vidmage.ai/api/v1/ai-talking-photo/query |
| Task identifier | taskId |
| Result field | videoUrl |
Authenticate with Authorization: Bearer YOUR_API_KEY. Save the task identifier and query the same capability. Use a stable Idempotency-Key for submission retries.
Product guide: AI Lip Sync API for Photos and Videos
Parameters and input rules
| Field | Type | Required | Default | Meaning and limits |
|---|---|---|---|---|
imageUrl | string | Yes | Not specified | URL of the portrait photo (JPG, JPEG, PNG, or WebP, up to 30 MB). |
audioUrl | string | Yes | Not specified | URL of the reference voice audio (MP3 or WAV, up to 15 MB). It controls voice identity, tone, and style; its original words do not become part of the output. |
textContent | string | Yes | Not specified | The complete spoken script (max 300 chars). Keep it short enough to produce a 2–15 second video. Use estimate_capability_credits for the exact credit estimate. Maximum characters: 300 |
Example request
{
"imageUrl": "https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp",
"audioUrl": "https://vidmage.ai/templates/voices/alice.mp3",
"textContent": "Welcome to VidMage. This reference voice will speak the complete script."
}# Set VIDMAGE_IDEMPOTENCY_KEY to a unique value for this task; preserve it for transport retries.
curl -X POST "https://vidmage.ai/api/v1/ai-talking-photo/submit" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Idempotency-Key: ${VIDMAGE_IDEMPOTENCY_KEY}" \
-H "Content-Type: application/json" \
-d '{"imageUrl":"https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp","audioUrl":"https://vidmage.ai/templates/voices/alice.mp3","textContent":"Welcome to VidMage. This reference voice will speak the complete script."}'
# -> { "success": true, "taskId": "...", "creditsConsumed": ... }Task results and recovery
URL of the generated talking video. Read videoUrl from the completed query response. Preserve taskId while the task is running.
curl -X POST "https://vidmage.ai/api/v1/ai-talking-photo/query" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "taskId": "TASK_ID_FROM_SUBMIT" }'
# -> { "success": true, "data": { "status": "...", "videoUrl": "https://..." } }Submission and task queries report the available task and usage information. A timeout is not a confirmed failure. Recover the existing task instead of resubmitting.
For face selection and error recovery, follow Task lifecycle. MCP uses the normalized taskId argument in get_task_result, including capabilities whose REST identifier is requestId.
Credits
| Component / option | Rate | Minimum |
|---|---|---|
| Video duration | 15 credits / second | 75 credits |
| Script length | 0.1 credits / character | 5 credits |
Output duration is estimated from the script with a safety allowance.
Reference-audio duration is validated but does not affect the quote.
Video and speech charges are added together, with each component’s minimum applied. Billable video duration is estimated from the script.
Use the Playground or the MCP credit estimator for your exact inputs. Credits and billing.
