JSON-to-video API · built for automation

Turn JSON into
rendered video.

SamAutomation (Sam Automation) is a deterministic video API: send a JSON payload — media, voice-over, captions, timings — and get a finished MP4 back. Wire it into n8n or your backend. Then generate AI video & images on the same key.

JSON → MP4 in one call n8n-ready webhooks + 48 AI models on the same key
 POST /api/function/video-generation/mix-video
{
  "media_list": [{ "type":"video", "url":"clip1.mp4" },
    { "type":"image", "url":"logo.png" }],
  "voice_over_url": "vo.mp3",
  "captions": [{ "start":0, "end":2, "words":"Ship faster" }],
  "settings": { "aspect_ratio":"9:16" }
}
200 rendered MP4 + webhook callback → store the URL
16-second product tour

From payload to finished video without opening an editor.

Watch the workflow SamAutomation is built for: JSON in, a render job queued, a finished MP4 out, plus optional AI footage and n8n-ready callbacks when your pipeline needs them.

// sound on for the full motion edit

// core

The video API for automation

Deterministic, repeatable rendering built for backends and no-code pipelines — not a manual editor. Structured input in, finished video out, every time.

Coming from an editor? Read what a CapCut API can and cannot do, or start with JSON to video and text to video.

JSON in, MP4 out

Mix video, images, voice-over and captions from one payload. Simple, loop and advanced renderers for every format.

Webhooks & status polling

Submit async, get a task id, poll progress or receive a callback. Designed to run without a human in the loop.

Reusable templates

Save a render as a template, then feed it new data. Browse the marketplace or publish your own.

Captions, baked in

Auto-captions from a transcript or script — burned-in subtitles for social, with styling control.

n8n & Zapier workflows Social cutdowns at scale Training & onboarding video Programmatic ads
# submit a render, then poll
POST /api/function/video-generation/mix-video
# → { "task_id": "a1b2", "status": "queued" }

GET /api/function/video-generation/progress/a1b2
# → { "status": "done",
# "video_url": "…/out.mp4" }

# or let the webhook push it to your stack
"callback_url": "https://your.app/hooks/video"
// add-on

Need generated footage? Add AI.

Same key, same job pattern. Generate clips and images from 31 video and 17 image models when you don't have source footage — then composite them with JSON-to-video.

Text to video

Describe a scene, get a clip.

Grok Imagine

xAI

xAI text and image to video.

  • Text and image to video
  • Multiple aspect ratios
  • Several style modes
Best for: Fast creative clips · Social-first short video

Kling 3.0

Kuaishou

Director-grade text/image-to-video, with a fast tier.

  • Text and image to video with multi-shot storyboarding
  • Native multilingual audio
  • Turbo tier for faster, lower-cost generation
Best for: Short-form ads · Image-to-video product shots

Veo 3.1

Google DeepMind
Flagship

Cinematic text/image-to-video with synchronized native audio.

  • Native synced audio — dialogue, ambient sound and music in one pass
  • Up to 3 reference images for character and scene consistency
  • Scene extension to chain shots into longer sequences
Best for: Programmatic ad and social video · Character-consistent multi-shot sequences

Image to video

Animate a still or a first/last frame.

Hailuo 2.3

MiniMax

Image-to-video at standard and pro quality tiers.

  • Image-to-video generation
  • 768p and 1080p tiers
  • 6s and 10s clips
Best for: Animate a product still · Short social clips from images

Multimodal & reference

Steer with images, video, voice or character refs.

Gemini Omni

Google
Premium

Any-input-to-video with voice and character consistency.

  • Combine text, image, video and voice references
  • Character consistency held across edits
  • Voice and character workflows for personalized presenters
Best for: Mixed-media short-form video · Personalized presenter and avatar clips

Kling 3.0 Motion Control

Kuaishou
New

Transfer motion from a reference video onto your character.

  • Motion transfer — re-animate a character with a reference clip's motion
  • Character identity kept consistent
  • 720p and 1080p output
Best for: Brand mascot performing a referenced gesture or dance · Consistent character animation

Seedance 2.0

ByteDance
New

Multimodal video with native audio and multi-shot cuts.

  • Text, image, video and audio references in one call
  • Multi-shot generation with natural cuts
  • Native in-pass audio; Fast tier for quick iteration
Best for: Single-call multi-shot product videos · Reference-driven brand consistency

Avatar & lip-sync

Talking avatars and re-synced footage.

OmniHuman 1.5

ByteDance
Premium

Expressive talking avatar from one image and audio.

  • Audio-driven animation from a single image plus speech
  • Context-aware expression, gesture and head motion
  • Portrait, half-body and full-body framing
Best for: Talking-head spokesperson clips from one headshot · Localized presenters per language

Video Lip Sync

New

Re-sync an existing video's lips to new target audio.

  • Video-to-video: regenerate the mouth region to match new audio
  • Keeps the original footage; swaps only the speech
  • Ideal alongside a translation pipeline
Best for: Dubbing and localization of talking-head footage · Updating a presenter video with new audio

Video editing

Edit, extend and transform existing clips.

HappyHorse 1.1

Alibaba
Featured

Joint audio-video generation with best-in-class multilingual lip-sync.

  • Four endpoints: text, image, reference-to-video and instruction edit
  • Native synced audio with multilingual lip-sync across languages
  • Reference-to-video from multiple images for character consistency
Best for: Talking-character clips with perfect lip-sync · Localized/dubbed video at scale · Multi-reference brand videos

Wan 2.7

Alibaba (Tongyi)
New

Director-grade suite: generate, reference and edit video.

  • Four modes: text, image, reference-to-video and instruction video-edit
  • First/last-frame control and multi-reference consistency
  • Native audio generation
Best for: Multi-shot brand clips with consistent characters · Edit-in-place revisions to footage

Wan VACE

Alibaba (Tongyi)
New

All-in-one video creation and editing under one model.

  • Unifies text, image and video-to-video in a single framework
  • Reference-to-video, inpaint/outpaint and region edits
  • Character and camera motion control
Best for: Swap or replace a subject in existing footage · Reference-guided edits and extensions

// real output — generated through the API, not stock footage

// choosing a tier

Test cheap. Ship premium.

Most model families ship in tiers. Prototype on the Lite or Fast variant for a fraction of the credits, lock the prompt, then switch the model_id to the top tier for your production render — same call, same payload.

Seedream2.7× cheaper to test

Lite from 6 cr Prototype & iterate
Standard from 16 cr Balanced

cheap to prototype  →  premium to ship

Veo 3.15.5× cheaper to test

Lite from 40 cr Prototype & iterate
Pro from 220 cr Final production

cheap to prototype  →  premium to ship

Kling 3.01.2× cheaper to test

Fast from 90 cr Quick drafts
Standard from 110 cr Balanced

cheap to prototype  →  premium to ship

Every option's exact credit cost is in GET /api/ai/models — query it before you call, so cost is never a surprise.

// pricing

One plan covers both

Every plan includes JSON-to-video rendering plus monthly AI credits. AI credits scale with the work — the number you see in the API is the number you pay. Top up anytime.

Basic
1,450 AI credits / mo
  • JSON-to-video rendering included
  • Core video + image AI models
  • n8n webhooks & templates
Choose Basic
Pro Popular
2,450 AI credits / mo
  • Everything in Basic
  • Pro AI models (Kling, Seedance 2, Wan 2.7)
  • Higher resolutions & durations
Choose Pro
Ultimate
4,950 AI credits / mo
  • Everything in Pro
  • Premium & avatar models (Gemini Omni, OmniHuman)
  • Highest monthly allowance
Choose Ultimate

Render your first video today

One key for deterministic JSON-to-video and 48 AI models. Documented, white-label, ready for your pipeline.