Yapper API
Models & Inputs

Video generation

Video models with per-length pricing and input fields.

POST /api/v1/processes with "type": "video-generation". Pick any model — the API does not plan-gate models. GET /api/v1/models returns the same capabilities at runtime.

Seedance generated audio

Seedance 2.x generates a native audio track by default. Set generateAudio: false for a silent video. When referenceAudios are attached, generated audio is required and generateAudio must be true or omitted.

{
  "type": "video-generation",
  "model": "seedance-2.5",
  "input": {
    "prompt": "A silent cinematic landscape reveal",
    "videoLength": 5,
    "generateAudio": false
  }
}

Seedance 2.5 Edit workflow

Use model: "seedance-2.5-edit" for the dedicated base-video workflow. Put the base video first in referenceVideos; later video items are additional references. aspectRatio and durationMode are always "auto", and videoLength is not accepted.

editOperation defaults to "edit". Edit requests require a non-empty prompt; "extend", "sequel", and "prequel" work without one because the server supplies the complete continuation/lead-in instruction.

{
  "type": "video-generation",
  "model": "seedance-2.5-edit",
  "input": {
    "editOperation": "sequel",
    "referenceVideos": [{ "assetId": "base_video_asset_id" }]
  }
}

Input fields

FieldTypeRequiredDescription
promptstringyesWhat to generate.
editOperation"edit" | "extend" | "sequel" | "prequel"noOnly for seedance-2.5-edit. Defaults to "edit"; prompt is optional for "extend", "sequel", and "prequel".
aspectRatiostringnoOne of the model's aspect ratios, e.g. "16:9". Use "adaptive" when advertised to follow the primary image/video reference.
resolutionnumbernoOutput resolution height, e.g. 1080.
videoLengthnumbernoDuration in seconds.
startingFrameImageUrlstringnoImage URL used as the first frame. External URLs are auto-imported as assets first.
endingFrameImageUrlstringnoImage URL used as the last frame. External URLs are auto-imported as assets first.
durationMode"fixed" | "auto"noUse "auto" to match the duration of a supported reference video; defaults to "fixed".
generateAudiobooleannoWhether Seedance should generate a native audio track. Defaults to true. Must remain true when referenceAudios are attached.
bitrateMode"standard" | "high"noSeedance 2.x output bitrate preset. Use "high" for a crisper, larger file at the same credit cost; defaults to "standard".
keyframes{timeSeconds, imageUrl}[]noTimestamped image frames. timeSeconds is measured from the beginning of the generated clip.
referenceImages({assetId} | {url})[]noImage references. Send either {assetId} for a Yapper asset or {url} for a public URL to import. Legacy {docId, url} objects remain accepted.
referenceVideos({assetId} | {url})[]noVideo references. Send either {assetId} for a Yapper asset or {url} for a public URL to import. Legacy {docId, url} objects remain accepted.
referenceAudios({assetId} | {url})[]noAudio references. Send either {assetId} for a Yapper asset or {url} for a public URL to import. Legacy {docId, url} objects remain accepted; private storage metadata is resolved server-side.
disableAutoRetriesbooleannoDisable automatic same-model retries. Defaults to false, so automatic retries are enabled.
endingFrameImageRefobjectno
startingFrameImageRefobjectno

Referencing media: referenceImageUrls takes plain https image URLs. The normalized referenceImages, referenceVideos, and referenceAudios fields accept either {assetId} for an existing Yapper asset or {url} for public media, which is imported automatically. Provider-native asset:// URIs are internal and are not accepted here. Legacy {docId, url} objects remain accepted for backwards compatibility.

Example

Add "dryRun": true to get the exact credit cost without starting, and always send an Idempotency-Key header on real starts.

POST /api/v1/processes
{
  "type": "video-generation",
  "model": "grok-imagine-v1.5",
  "input": {
    "prompt": "A slow dolly shot over a foggy pine forest at sunrise",
    "videoLength": 1,
    "aspectRatio": "16:9"
  }
}

Models at a glance

Credits are per generated output; ranges span the supported resolutions/lengths (exact prices in each model's section).

ModelCompanyCreditsLengths (s)Max ref imagesFramesAudio
wan-3.0-primealibaba50–7502–3010start + endyes
wan-3.0alibaba100–14602–3010start + endyes
flux-3black-forest-labs310–12405–200start + endyes
minimax-h3minimax70–2205–159start + endyes
minimax-h3-maxminimax22–665–15start + endyes
seedance-2.5bytedance380–28004–3030start + endyes
seedance-2.0bytedance220–8404–159start + endyes
grok-imaginexai58–1354, 6, 8, 10, 12, 14, 15startyes
grok-imagine-v1.5xai20–2601–15startyes
happy-horsealibaba90–4503–155startyes
seedance-2.0-fastbytedance180–6804–159start + endyes
seedance-2.0-minibytedance40–1204–159start + endyes
wan-2.7alibaba30–1502, 4, 6, 8, 105start + endyes
pixverse-v6pixverse20–3701, 2, 4, 6, 8, 10, 12, 14, 15startyes
kling-3.0kuaishou72–2704, 6, 8, 10, 12, 14, 153start + endyes
kling-3.0-prokuaishou96–3604, 6, 8, 10, 12, 14, 153start + endyes
gemini-omni-flashgoogle13–3973–1010start + endyes
veo3-qualitygoogle21683start + endyes
veo3-fastgoogle8083start + endyes
sora-2openai48–2404, 8, 12, 16, 20startyes
sora-2-proopenai430–21504, 8, 12, 16, 20startyes
seedance-2.0-openbytedance310–11804–159start + endyes
seedance-2.5-editbytedancedryRun30yes

WAN 3.0 Prime (wan-3.0-prime)

Alibaba's high-speed WAN 3.0 tier, around 7x faster with the same multimodal inputs and native audio.

Credits2s: 50 · 3s: 80 · 4s: 100 · 5s: 130 · 6s: 150 · 7s: 180 · 8s: 200 · 9s: 230 · 10s: 250 · 11s: 280 · 12s: 300 · 13s: 330 · 14s: 350 · 15s: 380 · 16s: 400 · 17s: 430 · 18s: 450 · 19s: 470 · 20s: 500 · 21s: 520 · 22s: 550 · 23s: 570 · 24s: 600 · 25s: 620 · 26s: 650 · 27s: 670 · 28s: 700 · 29s: 720 · 30s: 750 (at 1080p)
PricingLimited-time launch price; the future standard credit price is 2× the current rate
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16
Lengths (s)2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30
Resolutions480, 720, 1080
Reference imagesup to 10
Reference videosup to 5 (15s total)
Reference audioup to 5 (15s total)
Text-only generationsupported
Start/end framesstart + end
Audio outputyes
Max prompt length5000

WAN 3.0 (wan-3.0)

Alibaba's newest multimodal video model with native audio, start/end frames, and image, video, and audio references.

Credits2s: 100 · 3s: 150 · 4s: 200 · 5s: 250 · 6s: 300 · 7s: 350 · 8s: 390 · 9s: 440 · 10s: 490 · 11s: 540 · 12s: 590 · 13s: 640 · 14s: 690 · 15s: 730 · 16s: 780 · 17s: 830 · 18s: 880 · 19s: 930 · 20s: 980 · 21s: 1030 · 22s: 1070 · 23s: 1120 · 24s: 1170 · 25s: 1220 · 26s: 1270 · 27s: 1320 · 28s: 1370 · 29s: 1410 · 30s: 1460 (at 1080p)
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16
Lengths (s)2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30
Resolutions480, 720, 1080
Reference imagesup to 10
Reference videosup to 5 (15s total)
Reference audioup to 5 (15s total)
Text-only generationsupported
Start/end framesstart + end
Audio outputyes
Max prompt length5000

FLUX.3 (flux-3)

Black Forest Labs' frontier video model for text, timestamped image keyframes, start/end-frame interpolation, or video continuation, with native audio and clips up to 20 seconds.

Credits5s: 310 · 6s: 370 · 7s: 440 · 8s: 500 · 9s: 560 · 10s: 620 · 11s: 680 · 12s: 750 · 13s: 810 · 14s: 870 · 15s: 930 · 16s: 990 · 17s: 1060 · 18s: 1120 · 19s: 1180 · 20s: 1240 (at 1080p)
Aspect ratios16:9, 9:16, 1:1, 21:9, 2:1, 4:3, 3:4, auto (default auto)
Lengths (s)5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20
Resolutions720, 1080
Reference imagesup to 0
Reference videosup to 1
Text-only generationsupported
Start/end framesstart + end
Audio outputyes

MiniMax H3 (minimax-h3)

MiniMax's frontier 2K video model for text generation and multimodal image, video, and audio references.

Credits5s: 70 · 6s: 90 · 7s: 100 · 8s: 120 · 9s: 130 · 10s: 150 · 11s: 160 · 12s: 170 · 13s: 190 · 14s: 200 · 15s: 220 (at 768p)
Aspect ratios16:9, 21:9, 4:3, 1:1, 3:4, 9:16, adaptive
Lengths (s)5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions768, 1440
Reference imagesup to 9
Reference videosup to 3 (15s total)
Reference audioup to 3 (15s total)
Text-only generationsupported
Start/end framesstart + end
Audio outputyes
Max prompt length2000

MiniMax H3 Max (minimax-h3-max)

Fal's post-trained H3 variant for exceptionally fast 480p and 768p text-to-video or first/end-frame animation with native audio.

Credits5s: 22 · 6s: 26 · 7s: 31 · 8s: 35 · 9s: 40 · 10s: 44 · 11s: 49 · 12s: 53 · 13s: 57 · 14s: 62 · 15s: 66 (at 480p)
PricingLimited-time launch price; the future standard credit price is 2× the current rate
Aspect ratios16:9, 21:9, 4:3, 1:1, 3:4, 9:16 — startingFrameImageUrl requires "auto" (or omission)
Lengths (s)5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions480, 768
Text-only generationsupported
Start/end framesstart + end
Audio outputyes

Seedance 2.5 (seedance-2.5)

ByteDance's latest multimodal video model with native audio, 30-second generations, and expanded image, video, and audio references.

Credits4s: 380 · 5s: 470 · 6s: 570 · 7s: 660 · 8s: 750 · 9s: 850 · 10s: 940 · 11s: 1030 · 12s: 1130 · 13s: 1220 · 14s: 1310 · 15s: 1400 · 16s: 1500 · 17s: 1590 · 18s: 1680 · 19s: 1780 · 20s: 1870 · 21s: 1960 · 22s: 2060 · 23s: 2150 · 24s: 2240 · 25s: 2340 · 26s: 2430 · 27s: 2520 · 28s: 2610 · 29s: 2710 · 30s: 2800 (at 1080p)
PricingLimited-time launch price; the future standard credit price is 1.5× the current rate
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Lengths (s)4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30
Resolutions480, 720, 1080
Bitrate modesstandard, high (default standard)
Generated audioconfigurable (default on; required with reference audio)
Reference imagesup to 30
Reference videosup to 10 (30s total)
Reference audioup to 10 (30s total)
Text-only generationsupported
Start/end framesstart + end
Audio outputyes
Max prompt length20000

Seedance 2.0 (seedance-2.0)

ByteDance's video model with multi-modal input, lip sync, and multi-shot narrative.

Credits4s: 220 · 5s: 280 · 6s: 340 · 7s: 390 · 8s: 450 · 9s: 510 · 10s: 560 · 11s: 620 · 12s: 670 · 13s: 730 · 14s: 790 · 15s: 840 (at 1080p)
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Lengths (s)4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions480, 720, 1080, 2160
Bitrate modesstandard, high (default standard)
Generated audioconfigurable (default on; required with reference audio)
Reference imagesup to 9
Reference videosup to 3 (15s total)
Reference audioup to 3 (15s total)
Text-only generationnot supported
Start/end framesstart + end
Audio outputyes
Max prompt length20000

Grok Imagine (grok-imagine)

Grok's best and latest video model with sound included.

Credits4s: 58 · 6s: 72 · 8s: 86 · 10s: 100 · 12s: 114 · 14s: 128 · 15s: 135 (at 720p)
Aspect ratios16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 1:1 — "auto" also accepted with startingFrameImageUrl
Lengths (s)4, 6, 8, 10, 12, 14, 15
Resolutions720
Text-only generationsupported
Start/end framesstart
Audio outputyes
Max prompt length4080

Grok Imagine 1.5 (grok-imagine-v1.5)

xAI's Grok Imagine 1.5 image-to-video model with audio included.

Credits1s: 20 · 2s: 40 · 3s: 50 · 4s: 70 · 5s: 90 · 6s: 110 · 7s: 120 · 8s: 140 · 9s: 160 · 10s: 170 · 11s: 190 · 12s: 210 · 13s: 230 · 14s: 240 · 15s: 260 (at 480p)
Aspect ratios16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 1:1 — "auto" also accepted with startingFrameImageUrl
Lengths (s)1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions480, 720
Text-only generationnot supported
Start/end framesstart
Audio outputyes
Max prompt length4096

Happy Horse (happy-horse)

Alibaba's reference-to-video model with native audio and 720p output.

Credits3s: 90 · 4s: 120 · 5s: 150 · 6s: 180 · 7s: 210 · 8s: 240 · 9s: 270 · 10s: 300 · 11s: 330 · 12s: 360 · 13s: 390 · 14s: 420 · 15s: 450 (at 720p)
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4
Lengths (s)3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions720
Reference imagesup to 5
Text-only generationnot supported
Start/end framesstart
Audio outputyes
Max prompt length5000

Seedance 2.0 Fast (seedance-2.0-fast)

Fast variant of ByteDance's video model with multi-modal input, lip sync, and multi-shot narrative.

Credits4s: 180 · 5s: 230 · 6s: 270 · 7s: 320 · 8s: 360 · 9s: 410 · 10s: 450 · 11s: 500 · 12s: 540 · 13s: 590 · 14s: 640 · 15s: 680 (at 1080p)
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Lengths (s)4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions480, 720, 1080, 2160
Bitrate modesstandard, high (default standard)
Generated audioconfigurable (default on; required with reference audio)
Reference imagesup to 9
Reference videosup to 3 (15s total)
Reference audioup to 3 (15s total)
Text-only generationnot supported
Start/end framesstart + end
Audio outputyes
Max prompt length20000

Seedance 2.0 Mini (seedance-2.0-mini)

ByteDance's economical video model for fast multimodal drafts with native audio.

Credits4s: 40 · 5s: 40 · 6s: 50 · 7s: 60 · 8s: 60 · 9s: 70 · 10s: 80 · 11s: 90 · 12s: 90 · 13s: 100 · 14s: 110 · 15s: 120 (at 480p)
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Lengths (s)4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions480, 720
Bitrate modesstandard, high (default standard)
Generated audioconfigurable (default on; required with reference audio)
Reference imagesup to 9
Reference videosup to 3 (15s total)
Reference audioup to 3 (15s total)
Text-only generationnot supported
Start/end framesstart + end
Audio outputyes
Max prompt length20000

WAN 2.7 (wan-2.7)

Wan's latest reference-driven video model with image and video guidance for better consistency.

Credits2s: 30 · 4s: 60 · 6s: 90 · 8s: 120 · 10s: 150 (at 1080p)
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4
Lengths (s)2, 4, 6, 8, 10
Resolutions720, 1080
Reference imagesup to 5
Reference videosup to 1
Text-only generationsupported
Start/end framesstart + end
Audio outputyes
Max prompt length5000

PixVerse V6 (pixverse-v6)

PixVerse's latest Fal-hosted video model with optional audio and first-frame image support.

Credits1s: 20 · 2s: 50 · 4s: 100 · 6s: 150 · 8s: 200 · 10s: 250 · 12s: 300 · 14s: 350 · 15s: 370 (at 1080p)
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9 — startingFrameImageUrl requires "auto" (or omission)
Lengths (s)1, 2, 4, 6, 8, 10, 12, 14, 15
Resolutions360, 540, 720, 1080
Text-only generationsupported
Start/end framesstart
Audio outputyes
Max prompt length2024

Kling 3.0 (kling-3.0)

Video model with start/end frames, references, and native audio.

Credits4s: 72 · 6s: 108 · 8s: 144 · 10s: 180 · 12s: 216 · 14s: 252 · 15s: 270
Aspect ratios16:9, 9:16, 1:1 — startingFrameImageUrl and referenceVideos require "auto" (or omission)
Lengths (s)4, 6, 8, 10, 12, 14, 15
Reference imagesup to 3
Reference videosup to 1
Start/end framesstart + end
Audio outputyes
Max prompt length2500

Kling 3.0 Pro (kling-3.0-pro)

Higher-cost Kling model with start/end frame, reference, and native audio support.

Credits4s: 96 · 6s: 144 · 8s: 192 · 10s: 240 · 12s: 288 · 14s: 336 · 15s: 360
Aspect ratios16:9, 9:16, 1:1 — startingFrameImageUrl and referenceVideos require "auto" (or omission)
Lengths (s)4, 6, 8, 10, 12, 14, 15
Reference imagesup to 3
Reference videosup to 1
Start/end framesstart + end
Audio outputyes
Max prompt length2500

Gemini Omni 1.1 Flash (gemini-omni-flash)

Google's fast multimodal video model with native audio, first/end-frame interpolation, video editing and extension, and output up to 4K.

PricingLimited-time launch price targeting a -15% margin; the future standard credit price is 1.9× the current rate
Input pricingThe displayed total adds Gemini's billed tokens for every image and the measured input-video duration. An input video without duration metadata is conservatively priced as 10s.
Credits (360p, no media inputs)3s: 13 · 4s: 18 · 5s: 22 · 6s: 26 · 7s: 31 · 8s: 35 · 9s: 40 · 10s: 44
Credits (720p, no media inputs)3s: 40 · 4s: 53 · 5s: 66 · 6s: 79 · 7s: 93 · 8s: 106 · 9s: 119 · 10s: 132
Credits (1080p, no media inputs)3s: 59 · 4s: 79 · 5s: 99 · 6s: 119 · 7s: 139 · 8s: 159 · 9s: 178 · 10s: 198
Credits (4K, no media inputs)3s: 119 · 4s: 159 · 5s: 198 · 6s: 238 · 7s: 278 · 8s: 317 · 9s: 357 · 10s: 397
Aspect ratios16:9, 9:16
Resolutions360p, 720p (default), 1080p (upscaled), 4K (upscaled)
Lengths (s)3, 4, 5, 6, 7, 8, 9, 10
Reference imagesup to 10 total image inputs, including start/end frames
Reference videosup to 1
Start/end framesstart + end
Text-only generationsupported
Audio outputyes

Veo3.1 Quality (veo3-quality)

Google's highest quality video model with sound included. Blocks celebrities.

Credits8s: 216
Aspect ratios16:9, 9:16
Lengths (s)8
Reference imagesup to 3
Text-only generationsupported
Start/end framesstart + end
Audio outputyes

Veo3.1 (veo3-fast)

Google's latest video model with sound included. Blocks celebrities.

Credits8s: 80
Aspect ratios16:9, 9:16
Lengths (s)8
Reference imagesup to 3
Text-only generationsupported
Start/end framesstart + end
Audio outputyes

Sora 2 (sora-2)

OpenAI's new bleeding-edge video model with sound included.

Credits4s: 48 · 8s: 96 · 12s: 144 · 16s: 192 · 20s: 240
Aspect ratios16:9, 9:16
Lengths (s)4, 8, 12, 16, 20
Text-only generationsupported
Start/end framesstart
Audio outputyes
Max prompt length5000

Sora 2 Pro (sora-2-pro)

OpenAI's new bleeding-edge video model with sound included.

Credits4s: 430 · 8s: 860 · 12s: 1290 · 16s: 1720 · 20s: 2150 (at 1080p)
Aspect ratios16:9, 9:16
Lengths (s)4, 8, 12, 16, 20
Resolutions720, 1080
Text-only generationsupported
Start/end framesstart
Audio outputyes
Max prompt length5000

Seedance 2.0 Open (seedance-2.0-open)

Less restricted Seedance 2.0 variant for content that the standard model may block.

Credits4s: 310 · 5s: 390 · 6s: 480 · 7s: 550 · 8s: 630 · 9s: 710 · 10s: 780 · 11s: 870 · 12s: 940 · 13s: 1020 · 14s: 1110 · 15s: 1180 (at 1080p)
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Lengths (s)4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15
Resolutions480, 720, 1080, 2160
Bitrate modesstandard, high (default standard)
Generated audioconfigurable (default on; required with reference audio)
Reference imagesup to 9
Reference videosup to 3 (15s total)
Reference audioup to 3 (15s total)
Text-only generationnot supported
Start/end framesstart + end
Audio outputyes
Max prompt length20000

Seedance 2.5 Edit (seedance-2.5-edit)

Dedicated Seedance 2.5 base-video workflow for editing or extending a video, or generating a natural sequel or prequel. The first referenceVideos item is the base video; duration and aspect ratio are automatic. Prompt is optional for extend, sequel, and prequel, but required for edit.

CreditsCost depends on the base video's automatic output duration — POST /processes with dryRun: true for an exact quote.
PricingLimited-time launch price; the future standard credit price is 1.5× the current rate
Base videorequired as the first referenceVideos item (array index 0)
Operationsedit, extend, sequel, prequel
Promptrequired for edit; optional for extend and sequel and prequel
Durationautomatic (durationMode: "auto")
Aspect ratiosauto (default auto) — referenceVideos require "auto" (or omission)
Resolutions480, 720, 1080
Bitrate modesstandard, high (default standard)
Generated audioconfigurable (default on; required with reference audio)
Reference imagesup to 30
Reference videosup to 10 (30s total)
Reference audioup to 10 (30s total)
Text-only generationnot supported
Audio outputyes
Max prompt length20000

On this page