Docs/Models
Models
Pick the model. Or let the numbers pick.
Live models are selectable by id as model on generate_video, generate_image and generate_audio. Omit it for the default. Every model shows its last 30 days of real runs, straight from the API.
On this page
Default to seedance-2.0. It is right for talking-head UGC, product in hand and reaction clips. seedance-2.5 is about 2.1x the credits at 720p; pick it for the one clip that has to be the best.
Video has three modes, chosen by the fields you pass: text (prompt only), image (image-to-video via first_frame, optionally last_frame) and reference (refs, video_refs and/or audio_refs, addressed in the prompt as @image1, @video1, @audio1). Frames and refs cannot be mixed. On seedance-2.5, image mode is adaptive-only: leave aspect out and the frame sets the ratio.
Quality sets the price. quality is one of 480p, 720p, 1080p, default 720p; each step changes the credits per second, and reference clip seconds are billed like output seconds. The table on each card has the numbers per mode.
seedance-2.0
Pick this when you need a real-looking person saying real words, or a still brought to life, at a normal budget.
- Verified
- text: 2026-09-06
- image: 2026-09-06
- reference: 2026-09-06
- Latency
- about 3 minutes for a 5 s clip at 720p
- Best for
- talking-head UGC · product in hands · crazy look · bulk daily posts · animating a still (first frame) into a clip
- Avoid for
- clips over 15s · hero shots where 2.5 detail is worth 2x the price · anything that needs a seed
Modes
The fields you pass pick the mode. Credits are per output second at the chosen quality; the default is 720p.
| Mode | How | Inputs | Seconds | Aspect | Quality | Credits/s | Verified |
|---|---|---|---|---|---|---|---|
| text | prompt only |
| 4 to 15 | 9:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive | 480p, 720p, 1080p |
| 2026-09-06 |
| image | pass first_frame (and optionally last_frame) |
| 4 to 15 | adaptive default, 9:16, 16:9, 1:1, 4:3, 3:4, 21:9 | 480p, 720p, 1080p |
| 2026-09-06 |
| reference | pass refs, video_refs and/or audio_refs |
| 4 to 15 | 9:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive | 480p, 720p, 1080p |
| 2026-09-06 |
- textNo seed. Every render is new; keep a series consistent with references, not seeds.
- imagefirst_frame becomes frame one of the clip; last_frame (optional) becomes the final frame and the model animates between them.
- imageFrames only: refs, video_refs and audio_refs are not accepted in this mode. To keep an identity AND set the frame, put the frame image in refs and describe it as @image1.
- imageaspect "adaptive" (default) follows the first frame. Frame images: jpeg/png/webp, ratio between 0.4 and 2.5, 300 to 6000 px, up to 30 MB.
- referenceAddress references in the prompt: "@image1 holds @image2 and says ...". Numbering is per list: refs are @image1.., video_refs are @video1.., audio_refs are @audio1...
- referenceReference clips: 2 to 15 s each, 15 s in total, mp4/mov 480p to 1080p, 24 to 60 fps, up to 50 MB. Their seconds are billed like output seconds.
- referenceAudio references: wav/mp3, 2 to 15 s each, 15 s in total, up to 15 MB. Audio alone is not accepted: add an image or video reference.
- referenceaspect "adaptive" follows the first video reference, else the first image, else the prompt.
Prompting
- 01Write the shot as a director: who (age, look), where (setting, light), what happens, phone framing, and the exact spoken words in quotes. About 2.3 words per second.
- 02Keep a person consistent across clips with refs (a portrait or a character sheet), addressed as @image1; the model has no seed.
- 03To animate a specific still, pass it as first_frame; add last_frame to control where the motion ends.
Last 30 days, live
Runs
84
Failed
31%
Auto score
0.79/ 56
Users
3/ 5, 1
Typical
5min
From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.
seedance-2.5
Pick this when the user asked for the best possible single clip and accepts the wait and about 2x the price.
- Verified
- text: 2026-09-05
- image: 2026-09-06
- reference: 2026-09-06
- Latency
- 12 to 25 minutes per clip; plan the wait
- Best for
- hero product ads · close-up faces · one clip that has to be the best
- Avoid for
- drafts · bulk · anything where 2.0 is good enough: it is about 2x the credits · anyone who cannot wait 15 to 30 minutes
Modes
The fields you pass pick the mode. Credits are per output second at the chosen quality; the default is 720p.
| Mode | How | Inputs | Seconds | Aspect | Quality | Credits/s | Verified |
|---|---|---|---|---|---|---|---|
| text | prompt only |
| 4 to 15 | 9:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive | 480p, 720p, 1080p |
| 2026-09-05 |
| image | pass first_frame (and optionally last_frame) |
| 4 to 15 | adaptive default | 480p, 720p, 1080p |
| 2026-09-06 |
| reference | pass refs, video_refs and/or audio_refs |
| 4 to 15 | 9:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive | 480p, 720p, 1080p |
| 2026-09-06 |
- textNo seed.
- imagefirst_frame becomes frame one; last_frame (optional) becomes the final frame.
- imageaspect must be "adaptive" on this model in image mode: the clip takes the frame's ratio. Any other aspect is refused.
- imageFrames only: refs, video_refs and audio_refs are not accepted in this mode.
- referenceAddress references as @image1.., @video1.., @audio1.. (numbered per list).
- referenceReference clips: 2 to 30 s each, 30 s in total, up to 4K, up to 200 MB; their seconds are billed like output seconds. Audio references may stand alone here.
- referenceNever write edit or extend intent ("edit the video", "add", "remove", "replace", "extend", "continue") into a reference prompt: the provider reclassifies the task and fails it minutes later. Describe the NEW clip you want.
Prompting
- 01Same director-style prompt and @image1 references as seedance-2.0.
- 02In image mode leave aspect out (it is adaptive); the frame decides the ratio.
- 03Do not put edit or extend wording in a reference prompt.
Last 30 days, live
Runs
3
Failed
0%
Auto score
0.82/ 2
Users
3/ 5, 1
Typical
6min
From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.
gpt-image-2.5
Pick this when you are making the image a video will be built on: a portrait, a character sheet, a first frame, or an edit that has to keep the same person.
- Limits
- 1024x1024, 1024x1536, 1536x1024 · up to 4 refs
- Verified
- no recorded run yet
- Latency
- about 30 seconds, a little longer for an edit with references
- Best for
- portraits and character sheets that a video has to keep · the first frame of a clip · product in hand · edits that must not lose the face
- Avoid for
- bulk throwaway drafts where gpt-image-2.5-flare is faster
Prompting
- 01Concrete subject, age, framing, light, what the hands do; it follows layout instructions like "four poses on a plain background" or "headroom for a caption".
- 02With refs it edits or composes from them and holds the identity across poses, which is what makes a character sheet usable as a video reference.
- 03Pass the result straight to generate_video as first_frame (to animate it) or in refs (to keep that person across clips).
Last 30 days, live
Runs
41
Failed
10%
Auto score
0.80/ 37
Users
1/ 5, 1
Typical
1min
From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.
gpt-image-2.5-flare
Pick this when you want the same look as gpt-image-2.5 but faster, or you are making several images at once.
- Limits
- 1024x1024, 1024x1536, 1536x1024 · up to 4 refs
- Verified
- no recorded run yet
- Latency
- about 15 to 30 seconds
- Best for
- variants and drafts at the same quality tier · batches of frames · anything where a few seconds matter
- Avoid for
- the one sheet a whole series depends on, where gpt-image-2.5 edits hold identity a little better
Prompting
- 01Same prompts as gpt-image-2.5; it is the speed tier of the same family.
Last 30 days, live
No runs in the last 30 days.
From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.
gpt-image-2
Pick this when you are reproducing something that was made on gpt-image-2; otherwise take the default.
- Limits
- 1024x1024, 1024x1536, 1536x1024 · up to 4 refs
- Verified
- 2026-09-05. every video pipeline stage A-C; portrait + sheet produced on run 2749ee84; standalone via generate_image
- Latency
- under a minute
- Best for
- the previous generation, kept selectable for runs that were built on it
- Avoid for
- new work: gpt-image-2.5 is the default and holds identity better
Prompting
- 01Concrete subject, age, framing, light, what the hands do; it follows layout instructions like "headroom for a caption".
- 02With refs it EDITS or composes from them (a product into a hand, a portrait re-lit); without refs it paints from the prompt alone.
Last 30 days, live
Runs
20
Failed
5%
Auto score
0.82/ 17
Users
4/ 5, 1
Typical
1min
From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.
elevenlabs-tts
Pick this when you need a clean voice track and no face.
- Limits
- –
- Verified
- 2026-09-05. wired in media-worker-v2 (tts.js, dubbing.js); standalone via generate_audio
- Latency
- seconds
- Best for
- voiceover on b-roll · narration · a standalone voice file · an audio reference for generate_video
- Avoid for
- lip-synced talking head: generate_video renders speech natively
Prompting
- 01Emotion tags like [excited] or [whispers] are honoured; keep sentences short for pacing.
- 02Seven named voices (sarah default) or a raw ElevenLabs voice id.
Last 30 days, live
Runs
7
Failed
14%
Auto score
–
Users
–
Typical
1min
From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.
model: "auto"
model:"auto" starts from the kind's default (image: gpt-image-2.5, video: seedance-2.0, audio: elevenlabs-tts); switches only when the default failed >25% of >=10 runs and another live model is healthy, or when a live model within 1.5x the default's price beats its auto score by >=0.1 over >=10 judged runs. Scores come from an auto-judge that grades every loose-surface job (3 frames or the image against the realism rubric, prompt adherence, identity match) plus rate_run.
The quote and the submit response say which model auto chose and why.
Planned, not selectable
Each goes live only after a real run is recorded and its price is confirmed. Naming one returns a 400 that lists the live models.
- seedance-2.0-minivideocandidatedrafts
- kling-o3videocandidate1080p hero shots with a non-Seedance look
- wan-3.0videocandidatesingle takes of 15 to 30 s
- omnihuman-1.5videocandidatelip-syncing one photo to an existing recording
- sora-2videocandidatecinematic b-roll without a locked face
- nano-banana-2imagecandidateproduct placement into a character frame
- seedream-5.0-proimagecandidatemulti-reference composites: person + product + setting
- z-image-turboimagecandidateframing wireframes
- doubao-seed-audio-1.0audiocandidatecloning a voice from a short reference clip
- sunoaudiocandidatea music bed under a clip