VidMage

VidMage is an AI video and image creation platform for creators. Turn your ideas and images into videos and fresh visuals.

Open the creator tools
Explore APIs
AI Video APIsAI Image APIsAI Audio APIsAI 3D APIsAI Face Swap APIsAI Effects APIs
Build with VidMage
QuickstartAPI keysVidMage MCPError reference
Your account
Developer consoleAbout VidMagePlans and creditsAPI billingContact support
© 2026 VidMage. All rights reserved.
Privacy PolicyTerms of ServiceReport Abuse
Skip to content
VidMage/Developers
Overview
APIs
All APIsAI Video APIsAI Image APIsAI Audio APIsAI 3D APIsAI Face Swap APIsAI Effects APIs
DocumentationVidMage MCPOpen console
Developers/AI Lip Sync API for Photos and Videos

AI Lip Sync API for Photos and Videos

Add lip sync to image and video workflows, or build talking-photo features with image, audio, and text inputs. Use the corresponding endpoint to retrieve a generated video for your app.

Get API keyView API docs
Abstract illustration for AI Lip Sync API for Photos and Videos
PlaygroundAPIPricingGuideFAQs

AI Lip Sync API Playground

Use your existing API key and subscription. Submitting a generation uses credits; uploading and preparing a request does not start a generation.

API access is available to subscribers only. Your subscription credits are shared across the website and the API.

View plans

Parameter

3 parameters

URL of the portrait photo (JPG, JPEG, PNG, or WebP, up to 30 MB).

URL of the reference voice audio (MP3 or WAV, up to 15 MB). It controls voice identity, tone, and style; its original words do not become part of the output.

The complete spoken script (max 300 chars). Keep it short enough to produce a 2–15 second video. Use estimate_capability_credits for the exact credit estimate.

Request code

#!/usr/bin/env bash
set -euo pipefail
# Set this once per intended task; preserve it and the body for transport retries.
: "${VIDMAGE_IDEMPOTENCY_KEY:?Set a unique key for this task}"

SUBMIT_RESPONSE="$(curl --fail-with-body --silent --show-error -X POST "https://vidmage.ai/api/v1/ai-talking-photo/submit" \
  -H "Authorization: Bearer ${VIDMAGE_API_KEY}" \
  -H "Idempotency-Key: ${VIDMAGE_IDEMPOTENCY_KEY}" \
  -H "Content-Type: application/json" \
  --data-raw '{
  "imageUrl": "https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp",
  "audioUrl": "https://vidmage.ai/templates/voices/alice.mp3",
  "textContent": "Welcome to VidMage. This reference voice will speak the complete script."
}')"
TASK_ID="$(printf '%s' "$SUBMIT_RESPONSE" | jq -er '.["taskId"]')"

while true; do
  QUERY_RESPONSE="$(curl --fail-with-body --silent --show-error -X POST "https://vidmage.ai/api/v1/ai-talking-photo/query" \
    -H "Authorization: Bearer ${VIDMAGE_API_KEY}" \
    -H "Content-Type: application/json" \
    --data-raw "{\"taskId\":\"${TASK_ID}\"}")"
  STATUS="$(printf '%s' "$QUERY_RESPONSE" | jq -r '(.status // .data.status // "") | ascii_downcase')"
  case "$STATUS" in
    success|succeeded|completed)
      RESULT="$(printf '%s' "$QUERY_RESPONSE" | jq -r '(.result // .["videoUrl"] // .data["videoUrl"] // empty)')"
      if [ -z "$RESULT" ]; then
        printf '%s' "Result missing; preserve task $TASK_ID and query the same task again. Do not resubmit." >&2
        exit 2
      fi
      printf '%s\n' "$RESULT"
      break
      ;;
    needs_input)
      printf '%s' "Face selection required; preserve task $TASK_ID" >&2
      exit 2
      ;;
    failed|error)
      printf '%s' "$QUERY_RESPONSE" | jq -r '(.message // .error // .data.error // "Task failed")' >&2
      exit 1
      ;;
  esac
  sleep 5
done

Response data

Submit the task to see the API response here.
PlaygroundCapabilities

API documentation

Submit a request, track the task, and retrieve your result.

EndpointAI Talking Photo ↓
Required inputs
imageUrl audioUrl textContent
Output
Video·videoUrl
EndpointLip Sync ↓
Required inputs
visualUrl
Output
Video·videoUrl
Request setup & limitsHeaders, input options and media limits

Request headers

Authorization
Bearer YOUR_API_KEY

Keep your API key on your server.

Content-Type
application/json
Idempotency-Key
YOUR_UNIQUE_KEY

Use a new key per task. Reuse it only when retrying the same submission.

Input options & limits

Request controls follow each task's parameter rules; they do not guarantee output properties.

AI Talking Photo

RequiredimageUrl audioUrl textContent
Request controls
  • textContent: see parameter rules

Output dimensions and format are not specified in this reference.

Lip Sync

RequiredvisualUrl
Optional inputsaudioUrl targetVideoUrl
Request controls
  • inputType: "image", "video", "video-to-video"

Output dimensions and format are not specified in this reference.

Input retention and output-link lifetime depend on the API contract. Confirm API-specific retention terms before making promises to your users. Upload guide ↗

AI Talking Photo

POST /api/v1/ai-talking-photo/submit

Full API documentation ↗

imageUrl is required and audioUrl is required and textContent is required.

Required inputs

imageUrlstring
URL of the portrait photo (JPG, JPEG, PNG, or WebP, up to 30 MB).
audioUrlstring
URL of the reference voice audio (MP3 or WAV, up to 15 MB). It controls voice identity, tone, and style; its original words do not become part of the output.
textContentstring
The complete spoken script (max 300 chars). Keep it short enough to produce a 2–15 second video. Use estimate_capability_credits for the exact credit estimate.

Edit the sample inputs for your own task before submitting.

AI Talking Photo request
# Set VIDMAGE_IDEMPOTENCY_KEY to a unique value for this task; preserve it for transport retries.
curl -X POST "https://vidmage.ai/api/v1/ai-talking-photo/submit" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Idempotency-Key: ${VIDMAGE_IDEMPOTENCY_KEY}" \
  -H "Content-Type: application/json" \
  -d '{"imageUrl":"https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp","audioUrl":"https://vidmage.ai/templates/voices/alice.mp3","textContent":"Welcome to VidMage. This reference voice will speak the complete script."}'
# -> { "success": true, "taskId": "...", "creditsConsumed": ... }
Parameters and input rules3 fields
FieldTypeRequiredDefaultMeaning and limits
imageUrlstringYesNot specifiedURL of the portrait photo (JPG, JPEG, PNG, or WebP, up to 30 MB). No additional field constraint listed.
audioUrlstringYesNot specifiedURL of the reference voice audio (MP3 or WAV, up to 15 MB). It controls voice identity, tone, and style; its original words do not become part of the output. No additional field constraint listed.
textContentstringYesNot specifiedThe complete spoken script (max 300 chars). Keep it short enough to produce a 2–15 second video. Use estimate_capability_credits for the exact credit estimate. Maximum length: 300

Retrieve the result

Save taskId from the accepted submission and send it to this endpoint:

POST /api/v1/ai-talking-photo/query

Use the same Bearer API key. On completion, read the videoUrl result field.

Task states and recovery ↗
AI Talking Photo query
curl -X POST "https://vidmage.ai/api/v1/ai-talking-photo/query" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "taskId": "TASK_ID_FROM_SUBMIT" }'
# -> { "success": true, "data": { "status": "...", "videoUrl": "https://..." } }

Lip Sync

POST /api/v1/lip-sync/submit

Full API documentation ↗

visualUrl is required.

Required inputs

visualUrlstring
URL of the source image (JPG/JPEG/PNG/WebP, up to 30 MB) or video (MP4/MOV, up to 50 MB). Video duration limits depend on inputType.

Edit the sample inputs for your own task before submitting.

Lip Sync request
# Set VIDMAGE_IDEMPOTENCY_KEY to a unique value for this task; preserve it for transport retries.
curl -X POST "https://vidmage.ai/api/v1/lip-sync/submit" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Idempotency-Key: ${VIDMAGE_IDEMPOTENCY_KEY}" \
  -H "Content-Type: application/json" \
  -d '{"visualUrl":"https://vidmage.ai/assets/images/samples/blue-eyed-woman-sunlight.webp","inputType":"image","audioUrl":"https://vidmage.ai/templates/voices/alice.mp3"}'
# -> { "success": true, "taskId": "...", "creditsConsumed": ... }
Parameters and input rules4 fields

audioUrl is required when inputType is not video-to-video.

targetVideoUrl is required when inputType is video-to-video.

FieldTypeRequiredDefaultMeaning and limits
visualUrlstringYesNot specifiedURL of the source image (JPG/JPEG/PNG/WebP, up to 30 MB) or video (MP4/MOV, up to 50 MB). Video duration limits depend on inputType. No additional field constraint listed.
inputTypestringNo"image"Type of source. Values: "image", "video", "video-to-video" · Default: "image"
audioUrlstringNoNot specifiedDriving-audio fileUrl returned by upload_files (MP3 or WAV, required unless inputType is video-to-video). No additional field constraint listed.
targetVideoUrlstringNoNot specifiedTarget-video fileUrl returned by upload_files (MP4/MOV; required when inputType is video-to-video). No additional field constraint listed.

Retrieve the result

Save taskId from the accepted submission and send it to this endpoint:

POST /api/v1/lip-sync/query

Use the same Bearer API key. On completion, read the videoUrl result field.

Task states and recovery ↗
Lip Sync query
curl -X POST "https://vidmage.ai/api/v1/lip-sync/query" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "taskId": "TASK_ID_FROM_SUBMIT" }'
# -> { "success": true, "data": { "status": "...", "videoUrl": "https://..." } }

Pricing

See what one image, video, or generation costs in credits. Your API and website activity share the same credit balance.

API credit rates and calculated example request costs
API / workflowExample requestPrice (credits)
AI Talking Photovideo API72 characters: 98 credits
How this is calculated
  • Video duration: 15 credits / second · 75 credits minimum
  • Script length: 0.1 credits / character · 5 credits minimum

Output duration is estimated from the script with a safety allowance.

Reference-audio duration is validated but does not affect the quote.

Video and speech charges are added together, with each component’s minimum applied. Billable video duration is estimated from the script.

Estimate your own request ↗
Video + speech15 credits / second0.1 credits / character
Lip Syncvideo API2.09-second audio: 75 creditsinputType: image
How this is calculated
  • Image + audio: 15 credits / second · 75 credits minimum
  • Video + audio: 25 credits / second · 125 credits minimum
  • Video to Video: 2 credits / second · 20 credits minimum

The server measures the uploaded driving media; audioDuration is quote-only input.

Image/video plus audio is billed at a minimum five-second output duration.

The example uses the duration shown above. The server measures uploaded media for billing; this duration is only supplied to the estimate.

The selected input workflow determines the rate. Duration is rounded up to whole seconds and the workflow minimum applies.

Estimate your own request ↗
2–25 credits / secondBase rate varies by settings or workflow
Billed in credits. Shared with the website.Active subscription required. Unit rates and examples follow current API credit rules.Billing details⌄

Examples show the charge for the inputs listed above. Resolution, duration, audio, output count and reference media can change the total. Minimum charges and rounding follow the selected API.

Use the Playground to estimate your request before submitting it. Estimates do not start a generation. Subscription and credit-pack prices are listed on the plans page.

Your task’s recorded usage is the source of truth for the final charge.

View plans ↗Billing guide ↗

About AI Lip Sync API

Give a character visual a speaking role in a short drama, presentation, or narrated story. VidMage's AI Lip Sync API groups lip-sync and talking-photo endpoints for creating speaking videos from images or existing clips. Each endpoint has its own visual and speech inputs. The resulting video supplies a dialogue or narration segment for your editing workflow. To explore a speaking animation in your browser, open VidMage's AI Lip Sync.

AI Lip Sync API capabilities

Image and video lip sync
Pair a visual input with audio through the lip-sync endpoint, using the input mode for an image or video.
Video-to-video input mode
Select the documented video-to-video mode and provide its required targetVideoUrl in addition to the visual input.
Talking-photo generation
Use the separate talking-photo endpoint with an image, audio, and text to request a generated speaking video.

What you can build

Dialogue for short dramas

Build a scene around the lines its characters need to say. Pair a character visual with recorded dialogue through the matching lip-sync mode, giving your episode editor speaking clips to intercut with reactions and other shots.

Narrated course presenters

An approved lesson recording can drive a speaking video from an authorized presenter visual. Check mouth movement and timing against the speech, and use the chosen clip beside slides or examples within the course.

Speaking character introductions

A portrait can introduce itself in a story pitch. Supply the character image, introduction recording, and matching text to the talking-photo endpoint to create a speaking video for the pitch or an interactive fiction experience.

How to use AI Lip Sync API

  1. Get an API key

    Create a key in the Developer Console and store it on your server. Use it in the Authorization: Bearer header.

  2. Prepare and submit your inputs

    Set up your inputs in the Playground, then copy the matching API request. Send it from your server with a unique Idempotency-Key.

  3. Track the task

    Save the returned taskId and query the same operation until it completes. Keep that identifier if your app stops waiting.

  4. Retrieve the result

    Read videoUrl from the completed task. Preview the result in your app and save a copy to your own storage for later use.

Input tips
Use the matching speech inputs
Lip sync requires audioUrl except in video-to-video mode, which requires targetVideoUrl. Rebuild the required input set whenever a user changes the selected mode.
Keep talking-photo text separate
The talking-photo route requires imageUrl, audioUrl, and textContent, with a 300-character text limit. Its text field does not belong in the lip-sync request.

Input requirementsUpload local files

Task fields and results

Keep the task identifier with its original request and Idempotency-Key. Query the same capability until it completes, then read the documented result field.

Task identifier
taskId
Query route
POST /api/v1/ai-talking-photo/query
Completed result
videoUrl
Task states and response structure ↗
Errors, retries and limits

Retry transport failures with the same idempotency key only when the original submission may not have reached the server. For a confirmed task, keep querying the original task instead of submitting a duplicate.

Read error recovery guidance ↗

FAQs

Are failed or timed-out requests charged?

A client timeout is not a confirmed task failure. Confirmed failures are automatically refunded only when the refund outcome is certain. Unknown charges or refunds remain pending reconciliation.

Failure and refund handling ↗
How long should my app wait for a result?

The checked contract does not establish a fixed completion time. Choose a local wait budget for your app; it is not an API completion deadline.

Polling and wait budgets ↗
What are the rate and concurrency limits?

Request-rate limits control how often you can call the API; concurrency limits control simultaneous work. The documentation ties both to the stable API key but does not publish numeric limits for this operation. Confirm account limits before planning parallel jobs.

Rate-limit recovery ↗
Can I receive results through a webhook?

The checked request contract documents status queries and does not list a webhook or callback field for these routes. Check the current authenticated OpenAPI contract before making a callback part of your integration.

Query task results ↗Check the current contract ↗

Explore more APIs on VidMage

  • AI Video to Video API for Style Transformations↗
  • AI Subtitle Generator API for Video Captions↗
  • AI Photo Dance API for Reference-Led Animation↗
  • AI Motion Control API for Character Animation↗
  • Seedance API↗
  • Vidu API↗
  • Runway API↗
  • Sora 2 API↗