Docs/Models

Pick the model. Or let the numbers pick.

Live models are selectable by id as model on generate_video, generate_image and generate_audio. Omit it for the default. Every model shows its last 30 days of real runs, straight from the API.

On this page

Default to seedance-2.0. It is right for talking-head UGC, product in hand and reaction clips. seedance-2.5 is about 2.1x the credits at 720p; pick it for the one clip that has to be the best.

Video has three modes, chosen by the fields you pass: text (prompt only), image (image-to-video via first_frame, optionally last_frame) and reference (refs, video_refs and/or audio_refs, addressed in the prompt as @image1, @video1, @audio1). Frames and refs cannot be mixed. On seedance-2.5, image mode is adaptive-only: leave aspect out and the frame sets the ratio.

Quality sets the price. quality is one of 480p, 720p, 1080p, default 720p; each step changes the credits per second, and reference clip seconds are billed like output seconds. The table on each card has the numbers per mode.

seedance-2.0

videostandardlivedefault
60 credits per second at 720p

Pick this when you need a real-looking person saying real words, or a still brought to life, at a normal budget.

Verified
  • text: 2026-09-06
  • image: 2026-09-06
  • reference: 2026-09-06
Latency
about 3 minutes for a 5 s clip at 720p
Best for
talking-head UGC · product in hands · crazy look · bulk daily posts · animating a still (first frame) into a clip
Avoid for
clips over 15s · hero shots where 2.5 detail is worth 2x the price · anything that needs a seed

Modes

The fields you pass pick the mode. Credits are per output second at the chosen quality; the default is 720p.

ModeHowInputsSecondsAspectQualityCredits/sVerified
textprompt only
  • prompt: required
4 to 159:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive480p, 720p, 1080p
  • 480p 30
  • 720p 60
  • 1080p 150
2026-09-06
imagepass first_frame (and optionally last_frame)
  • first_frame: required (https image)
  • last_frame: optional (https image)
4 to 15adaptive default, 9:16, 16:9, 1:1, 4:3, 3:4, 21:9480p, 720p, 1080p
  • 480p 30
  • 720p 60
  • 1080p 150
2026-09-06
referencepass refs, video_refs and/or audio_refs
  • refs: up to 9 https images
  • video_refs: up to 3 https clips, 15 s in total; their seconds are billed like output seconds
  • audio_refs: up to 3 https audio files, 15 s in total; needs an image or video reference beside it
4 to 159:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive480p, 720p, 1080p
  • 480p 30
  • 720p 60
  • 1080p 150
2026-09-06
  • textNo seed. Every render is new; keep a series consistent with references, not seeds.
  • imagefirst_frame becomes frame one of the clip; last_frame (optional) becomes the final frame and the model animates between them.
  • imageFrames only: refs, video_refs and audio_refs are not accepted in this mode. To keep an identity AND set the frame, put the frame image in refs and describe it as @image1.
  • imageaspect "adaptive" (default) follows the first frame. Frame images: jpeg/png/webp, ratio between 0.4 and 2.5, 300 to 6000 px, up to 30 MB.
  • referenceAddress references in the prompt: "@image1 holds @image2 and says ...". Numbering is per list: refs are @image1.., video_refs are @video1.., audio_refs are @audio1...
  • referenceReference clips: 2 to 15 s each, 15 s in total, mp4/mov 480p to 1080p, 24 to 60 fps, up to 50 MB. Their seconds are billed like output seconds.
  • referenceAudio references: wav/mp3, 2 to 15 s each, 15 s in total, up to 15 MB. Audio alone is not accepted: add an image or video reference.
  • referenceaspect "adaptive" follows the first video reference, else the first image, else the prompt.

Prompting

  • 01Write the shot as a director: who (age, look), where (setting, light), what happens, phone framing, and the exact spoken words in quotes. About 2.3 words per second.
  • 02Keep a person consistent across clips with refs (a portrait or a character sheet), addressed as @image1; the model has no seed.
  • 03To animate a specific still, pass it as first_frame; add last_frame to control where the motion ends.

Last 30 days, live

Runs

84

Failed

31%

Auto score

0.79/ 56

Users

3/ 5, 1

Typical

5min

From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.

seedance-2.5

videopremiumlive
125 credits per second at 720p

Pick this when the user asked for the best possible single clip and accepts the wait and about 2x the price.

Verified
  • text: 2026-09-05
  • image: 2026-09-06
  • reference: 2026-09-06
Latency
12 to 25 minutes per clip; plan the wait
Best for
hero product ads · close-up faces · one clip that has to be the best
Avoid for
drafts · bulk · anything where 2.0 is good enough: it is about 2x the credits · anyone who cannot wait 15 to 30 minutes

Modes

The fields you pass pick the mode. Credits are per output second at the chosen quality; the default is 720p.

ModeHowInputsSecondsAspectQualityCredits/sVerified
textprompt only
  • prompt: required
4 to 159:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive480p, 720p, 1080p
  • 480p 60
  • 720p 125
  • 1080p 225
2026-09-05
imagepass first_frame (and optionally last_frame)
  • first_frame: required (https image)
  • last_frame: optional (https image)
4 to 15adaptive default480p, 720p, 1080p
  • 480p 60
  • 720p 125
  • 1080p 225
2026-09-06
referencepass refs, video_refs and/or audio_refs
  • refs: up to 30 https images
  • video_refs: up to 10 https clips, 30 s in total; their seconds are billed like output seconds
  • audio_refs: up to 10 https audio files, 30 s in total
4 to 159:16 default, 16:9, 1:1, 4:3, 3:4, 21:9, adaptive480p, 720p, 1080p
  • 480p 60
  • 720p 125
  • 1080p 225
2026-09-06
  • textNo seed.
  • imagefirst_frame becomes frame one; last_frame (optional) becomes the final frame.
  • imageaspect must be "adaptive" on this model in image mode: the clip takes the frame's ratio. Any other aspect is refused.
  • imageFrames only: refs, video_refs and audio_refs are not accepted in this mode.
  • referenceAddress references as @image1.., @video1.., @audio1.. (numbered per list).
  • referenceReference clips: 2 to 30 s each, 30 s in total, up to 4K, up to 200 MB; their seconds are billed like output seconds. Audio references may stand alone here.
  • referenceNever write edit or extend intent ("edit the video", "add", "remove", "replace", "extend", "continue") into a reference prompt: the provider reclassifies the task and fails it minutes later. Describe the NEW clip you want.

Prompting

  • 01Same director-style prompt and @image1 references as seedance-2.0.
  • 02In image mode leave aspect out (it is adaptive); the frame decides the ratio.
  • 03Do not put edit or extend wording in a reference prompt.

Last 30 days, live

Runs

3

Failed

0%

Auto score

0.82/ 2

Users

3/ 5, 1

Typical

6min

From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.

gpt-image-2.5

imagepremiumlivedefault
20 credits per image

Pick this when you are making the image a video will be built on: a portrait, a character sheet, a first frame, or an edit that has to keep the same person.

Limits
1024x1024, 1024x1536, 1536x1024 · up to 4 refs
Verified
no recorded run yet
Latency
about 30 seconds, a little longer for an edit with references
Best for
portraits and character sheets that a video has to keep · the first frame of a clip · product in hand · edits that must not lose the face
Avoid for
bulk throwaway drafts where gpt-image-2.5-flare is faster

Prompting

  • 01Concrete subject, age, framing, light, what the hands do; it follows layout instructions like "four poses on a plain background" or "headroom for a caption".
  • 02With refs it edits or composes from them and holds the identity across poses, which is what makes a character sheet usable as a video reference.
  • 03Pass the result straight to generate_video as first_frame (to animate it) or in refs (to keep that person across clips).

Last 30 days, live

Runs

41

Failed

10%

Auto score

0.80/ 37

Users

1/ 5, 1

Typical

1min

From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.

gpt-image-2.5-flare

imagestandardlive
20 credits per image

Pick this when you want the same look as gpt-image-2.5 but faster, or you are making several images at once.

Limits
1024x1024, 1024x1536, 1536x1024 · up to 4 refs
Verified
no recorded run yet
Latency
about 15 to 30 seconds
Best for
variants and drafts at the same quality tier · batches of frames · anything where a few seconds matter
Avoid for
the one sheet a whole series depends on, where gpt-image-2.5 edits hold identity a little better

Prompting

  • 01Same prompts as gpt-image-2.5; it is the speed tier of the same family.

Last 30 days, live

No runs in the last 30 days.

From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.

gpt-image-2

imagestandardlive
20 credits per image

Pick this when you are reproducing something that was made on gpt-image-2; otherwise take the default.

Limits
1024x1024, 1024x1536, 1536x1024 · up to 4 refs
Verified
2026-09-05. every video pipeline stage A-C; portrait + sheet produced on run 2749ee84; standalone via generate_image
Latency
under a minute
Best for
the previous generation, kept selectable for runs that were built on it
Avoid for
new work: gpt-image-2.5 is the default and holds identity better

Prompting

  • 01Concrete subject, age, framing, light, what the hands do; it follows layout instructions like "headroom for a caption".
  • 02With refs it EDITS or composes from them (a product into a hand, a portrait re-lit); without refs it paints from the prompt alone.

Last 30 days, live

Runs

20

Failed

5%

Auto score

0.82/ 17

Users

4/ 5, 1

Typical

1min

From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.

elevenlabs-tts

audiostandardlivedefault
0.01 credits per character

Pick this when you need a clean voice track and no face.

Limits
Verified
2026-09-05. wired in media-worker-v2 (tts.js, dubbing.js); standalone via generate_audio
Latency
seconds
Best for
voiceover on b-roll · narration · a standalone voice file · an audio reference for generate_video
Avoid for
lip-synced talking head: generate_video renders speech natively

Prompting

  • 01Emotion tags like [excited] or [whispers] are honoured; keep sentences short for pacing.
  • 02Seven named voices (sarah default) or a raw ElevenLabs voice id.

Last 30 days, live

Runs

7

Failed

14%

Auto score

Users

Typical

1min

From GET /v1/models, refreshed every few minutes. Auto score: an auto-judge grades every job 0 to 1.

model: "auto"

model:"auto" starts from the kind's default (image: gpt-image-2.5, video: seedance-2.0, audio: elevenlabs-tts); switches only when the default failed >25% of >=10 runs and another live model is healthy, or when a live model within 1.5x the default's price beats its auto score by >=0.1 over >=10 judged runs. Scores come from an auto-judge that grades every loose-surface job (3 frames or the image against the realism rubric, prompt adherence, identity match) plus rate_run.

The quote and the submit response say which model auto chose and why.

Planned, not selectable

Each goes live only after a real run is recorded and its price is confirmed. Naming one returns a 400 that lists the live models.

  • seedance-2.0-minivideocandidatedrafts
  • kling-o3videocandidate1080p hero shots with a non-Seedance look
  • wan-3.0videocandidatesingle takes of 15 to 30 s
  • omnihuman-1.5videocandidatelip-syncing one photo to an existing recording
  • sora-2videocandidatecinematic b-roll without a locked face
  • nano-banana-2imagecandidateproduct placement into a character frame
  • seedream-5.0-proimagecandidatemulti-reference composites: person + product + setting
  • z-image-turboimagecandidateframing wireframes
  • doubao-seed-audio-1.0audiocandidatecloning a voice from a short reference clip
  • sunoaudiocandidatea music bed under a clip