Seedance 2.0 vs Wan (2026): $0.04/s vs $0.10/s Video API
Seedance 2.0 costs $0.04–$0.07/s vs Wan $0.10/s: 30–60% cheaper, adds 4K + 3 tiers. Per-clip math, real specs, and A/B both on one ofox call.
TL;DR: Two Chinese video models, one endpoint, a clear price gap. Seedance 2.0 from ByteDance runs $0.04–$0.07/s across three tiers and reaches 4K on its flagship. Wan from Alibaba runs a flat $0.10/s and caps at 1080p. On identical clips Seedance is 30% cheaper at 1080p and about 60% cheaper on its Mini tier. Wan earns its price on shorter minimum durations (2s vs 4s) and Alibaba’s motion consistency. Both sit behind the same ofox POST /v1/videos call, so you can A/B them by editing one string instead of standing up two SDKs.
Which One Should You Pick?
If you generate short product clips, social cuts, or ad variants at volume, the money argument points at Seedance 2.0. It is cheaper on every matching resolution, and the Mini tier turns a $0.50 Wan clip into a $0.20 one. If you need clips under 4 seconds, or you are already tuned to Alibaba’s motion and prompt behavior, Wan is the safer pick despite the higher rate.
| Scenario | Pick | Why |
|---|---|---|
| High-volume 5–10s clips at 1080p | Seedance 2.0 | $0.07/s vs $0.10/s, same resolution |
| Tightest budget, 720p acceptable | Seedance 2.0 Mini | $0.04/s, ~60% under Wan |
| Need 4K output | Seedance 2.0 | Only one of the two that lists 4K |
| Clips shorter than 4 seconds | Wan 2.7 | 2s minimum vs Seedance’s 4s |
| Already tuned to Alibaba motion behavior | Wan 2.7 | Migration cost outweighs price gap |
| Not sure yet | Run both | One endpoint, one string to swap |
Everything below is the math and the specs behind that table, plus the exact code to run both on one call.
Quick Specs Comparison
All numbers here come from the live ofox model catalog on 2026-07-19. Both families expose text-to-video, image-to-video, and synced audio. The differences that matter are resolution ceiling, minimum duration, and price.
| Spec | Seedance 2.0 | Seedance 2.0 Fast | Seedance 2.0 Mini | Wan 2.7 | Wan 2.6 |
|---|---|---|---|---|---|
| Model ID | bytedance/seedance-2.0 | bytedance/seedance-2.0-fast | bytedance/seedance-2.0-mini | alibaba/wan-2.7 | alibaba/wan-2.6 |
| Price / second | $0.07 | $0.06 | $0.04 | $0.10 | $0.10 |
| Max resolution | 4K | 720p | 720p | 1080p | 1080p |
| Default resolution | 1080p | 720p | 720p | 1080p | 1080p |
| Modes | t2v, i2v, v2v | t2v, i2v, v2v | t2v, i2v, v2v | t2v, i2v, v2v | t2v, i2v |
| Duration range | 4–15s | 4–15s | 4–15s | 2–15s | 2–15s |
| Audio | Yes | Yes | Yes | Yes | Yes |
| Aspect ratios | 16:9, 9:16, 1:1, adaptive | same | same | 16:9, 9:16, 1:1 | 16:9, 9:16, 1:1 |
| Endpoint | /v1/videos | /v1/videos | /v1/videos | /v1/videos | /v1/videos |
Two things stand out. Seedance is the only side that reaches 4K, and it does so on the same flagship tier that already undercuts Wan on price. Wan’s advantage is at the short end: a 2-second minimum lets you generate tight loops and stingers that Seedance cannot produce below 4 seconds. If your format lives in the 6–10 second range, that floor never binds and the price gap is the whole story.
Where Seedance and Wan Come From
The two models trace back to the two Chinese labs with the deepest video-generation research. Seedance is part of ByteDance’s Seed group, the same team behind the Seedream image models and the short-video pipeline that feeds Douyin and TikTok. Wan is Alibaba’s Tongyi line, developed alongside the Qwen family and the Model Studio platform. You can read the vendor material at ByteDance Seed’s Seedance page and the Wan project site.
For a developer, the lineage matters less than the access path. Both models are trained and hosted in China first, which historically meant separate consoles, separate billing, and phone-number verification that does not always accept a foreign number. Routing them through the ofox gateway collapses that into one key and one endpoint, which is the practical reason this comparison is about the models rather than the paperwork. The interesting decision is price and capability, not which console you can sign into.
That framing also sets up the honest limit of a spec sheet. ByteDance and Alibaba both tune for their home platforms, so motion style, default pacing, and prompt interpretation differ in ways a datasheet does not capture. The only way to see it is to run the same prompt through both, which is the next section.
Same-Prompt Test: What the Price Gap Buys
Price and spec sheets settle the easy part. The harder question is whether the cheaper model holds up on motion, prompt adherence, and audio sync. Spec parity does not mean output parity, so this section runs the same prompt through both flagships and compares the result rather than trusting the datasheet.
Test prompt one, held identical across models, 1080p, 16:9, 8 seconds, audio on:
“A ceramic coffee cup on a wooden table, morning light from a window, steam rising, a hand enters frame and lifts the cup. Ambient cafe sound.”

Same prompt, four frames per model — Seedance 2.0 (top) and Wan 2.7 (bottom). Both render the cup, steam, and warm window light; watch the hand across the four frames.
The same two clips in motion, side by side: Seedance 2.0 (left), Wan 2.7 (right), looping and muted. The hand enters and lifts the cup on the Seedance side; on the Wan side the cup and steam hold but the lift barely reads.
| Criterion | Seedance 2.0 | Wan 2.7 | Notes |
|---|---|---|---|
| Prompt adherence (objects, action) | Full — the hand enters and lifts the cup | Partial — cup and steam land, the lift barely reads | Does the hand actually lift the cup |
| Motion coherence (no warping) | Clean across the four frames | Clean across the four frames | Steam and hand physics across frames |
| Audio level (measured) | Full synced bed, -34 dB avg / -5 dB peak | Present but faint, -43 dB avg / -30 dB peak | Both ship a synced audio track; Wan’s runs ~9 dB quieter on this clip |
| Cost for this 8s clip | $0.56 | $0.80 | 8 × price/second |
The cost row is fixed by the rate card: an 8-second 1080p clip is $0.56 on Seedance 2.0 and $0.80 on Wan 2.7, a $0.24 gap per clip. The quality rows come from the generated samples above, not from vendor claims.
A single calm prompt does not stress a video model, so a fair test also needs a motion-heavy prompt where cheaper models tend to break. Test prompt two, same settings:
“A skateboarder does a kickflip down a set of stairs, fast camera pan, crowd in the background, daytime.”

The motion-heavy prompt: Seedance 2.0 (top) holds the board, feet, and background crowd coherent through the trick; Wan 2.7 (bottom) keeps the wide shots but drifts into a distorted close-up mid-sequence.
Motion tells the story the stills can’t: Seedance 2.0 (left), Wan 2.7 (right), looping. The board and feet stay locked through the flip on the Seedance side, while the Wan side snaps to a distorted close-up partway through before recovering the wide shot.
This is the prompt that separates the tiers. Fast camera motion plus a rigid-body trick plus background crowd is where warping, limb duplication, and physics drift show up. Judge both clips for whether the board and feet stay coherent through the flip and whether the pan smears the background. Run both prompts on bytedance/seedance-2.0 and alibaba/wan-2.7 and watch the two clips side by side before you commit a pipeline to either. Longer clips compound the effect: at 15 seconds, small per-frame drift accumulates, so if your format runs long, test at your real duration rather than a safe 5-second sample.
The Async Workflow: Submit, Poll, or Webhook
Video generation is not a single blocking request the way a chat completion is. The ofox POST /v1/videos call returns immediately with 202 Accepted and a polling_url. The clip renders in the background, and you find out it is ready one of two ways: poll the task, or register a webhook.
The task moves through a small state machine. You keep checking, or you get called back, until it reaches a terminal state.
The task moves through three live states, then stops at one of four terminal states. Poll GET /v1/videos/{id} or register a webhook; DELETE cancels, and results expire on a TTL.
| Field / state | Meaning | What to do |
|---|---|---|
202 + polling_url | Task accepted, rendering started | Save the URL, begin polling |
status: pending | Queued, not started | Keep polling, no faster than 1/s |
status: processing | Rendering in progress | Keep polling |
status: completed | Done, result URL attached | Download the clip |
status: failed | Generation error | Read error, retry or fall back |
status: cancelled | You called DELETE | Stop |
status: expired | Result TTL elapsed | Regenerate |
Two operational details save you a support ticket. First, poll no faster than once per second; the GET /v1/videos/{id} endpoint is rate-limit protected, and a tight loop will get throttled. Second, for production you usually want the webhook instead of polling. Pass a callback_url when you create the task and ofox posts an HMAC-signed payload when the clip finishes. The URL must be HTTPS and public: private, loopback, and cloud-metadata addresses are rejected by SSRF validation, and a bad address fails at creation time with 400 invalid_callback_url. If you need to abort a run, DELETE /v1/videos/{id} cancels it. All of this is identical for Seedance and Wan, which is the point: the operational surface does not change when you swap the model.
Common Errors and Gotchas
Most failures with either model come from the same handful of mismatches, not from the models themselves.
| Symptom | Cause | Fix |
|---|---|---|
400 on a 2–3s clip with Seedance | Seedance minimum duration is 4s | Use Wan (2s floor) or raise duration to 4s |
| Resolution rejected on Fast / Mini | Fast and Mini cap at 720p | Request 720p, or use flagship bytedance/seedance-2.0 for 1080p / 4K |
400 invalid_callback_url | Webhook not HTTPS or points to a private address | Use a public HTTPS endpoint |
| Task stuck polling forever | Polling too fast and getting throttled, or ignoring terminal states | Poll ≤1/s, break on completed / failed / cancelled / expired |
| Got t2v when you wanted i2v | No image in the payload | Pass frame_images (or input_references) to trigger image mode |
| Unexpected aspect ratio | Wan has no adaptive option | Set an explicit ratio; only Seedance supports adaptive |
The mode-inference behavior is the one that surprises people. The endpoint does not take a mode flag; it reads your payload. No images means text-to-video, frame_images means image or first/last-frame, and input_references means reference-guided. That keeps the call shape identical across both model families, but it also means a forgotten image field silently gives you the wrong mode instead of an error.
Prompting and Audio: Practical Notes
Neither model needs a special prompt dialect, but a few habits raise the hit rate on both. For text-to-video, describe the subject, the action, the camera, and the setting as separate clauses rather than one run-on sentence; both families parse “what, doing what, shot how, where” more reliably than a wall of adjectives. Keep the action singular. A prompt that asks for a kickflip and a crowd reaction and a lens flare in eight seconds usually gets one of the three, so if you need all three, chain shorter clips instead of overloading one generation.
Audio is on by default on every tier here, and that changes how you write. Because both models generate a synced audio bed, a prompt that mentions “ambient cafe sound” or “footsteps on gravel” gets a matching track, while a prompt that says nothing about sound gets whatever the model infers, which is not always what you want. Name the audio you expect, or you inherit a guess. If you are compositing the clip into a timeline with your own sound design, that generated bed is noise you have to strip, so treat audio as a decision, not a freebie.
For image-to-video, the frame you pass matters more than the prompt. A clean, high-resolution first frame with the subject already composed gives both models a stable anchor, and the prompt then only has to describe motion: “slow push in,” “hair moves in the wind,” “camera orbits left.” Seedance’s adaptive aspect ratio helps here, because it can match the frame you hand it instead of forcing a crop, which means one reference image can drive a 16:9 and a 9:16 output without re-framing. Wan needs an explicit ratio, so plan the crop before you submit.
Pricing Math: Real Monthly Bill
Per-second rates are abstract until you multiply them by your volume. Here is a concrete pipeline: a small content team generating 2,000 clips per month, 5 seconds each, 1080p, which is a realistic cadence for an e-commerce or social team running product and promo variants.
| Model | Rate | Per 5s clip | 2,000 clips / month |
|---|---|---|---|
| Seedance 2.0 Mini (720p) | $0.04/s | $0.20 | $400 |
| Seedance 2.0 Fast (720p) | $0.06/s | $0.30 | $600 |
| Seedance 2.0 (1080p) | $0.07/s | $0.35 | $700 |
| Wan 2.7 (1080p) | $0.10/s | $0.50 | $1,000 |
At matched 1080p, Seedance flagship is $300/month cheaper than Wan; drop to 720p on Mini and the same 2,000 clips cost $600/month less.
At matched 1080p, Seedance 2.0 saves this team $300 a month against Wan, a 30% cut. If 720p is acceptable for the format, which it usually is for feed and story placements, the Mini tier saves $600 a month, a 60% cut, for the same 2,000 clips. Over a year that is $3,600 to $7,200 that stays in the budget.
The smarter play is not one tier but a routing ladder. Generate drafts and rejected variants on Mini at $0.04/s, promote the approved concept to Fast or flagship for the final render, and reserve the 4K flagship for hero clips a client will actually master. Because all three Seedance tiers share the bytedance/seedance-2.0 prefix and one endpoint, that ladder is a switch statement in your own code, not three integrations. A team that shoots 2,000 drafts but only finishes 200 pays roughly 1,800 × 5 × $0.04 + 200 × 5 × $0.07 = $360 + $70 = $430 a month, less than half the flat-flagship bill and well under Wan’s $1,000.
The counter-case is real too. If your clips run 2 to 3 seconds, Seedance cannot generate them at all (its floor is 4 seconds), so the effective Wan price for that format is not “more expensive,” it is “the only option of the two.” Match the model to the format before you match it to the price.
Resolution and Duration: Where the Bill Really Comes From
Both models bill per second, which has a consequence people miss on the first invoice: the price gap scales with clip length. At 5 seconds the Seedance-versus-Wan difference is $0.15. At 15 seconds, the maximum either model allows, it is $1.05 versus $1.50, a $0.45 gap on a single clip. Generate a few thousand long clips and the choice of model is a line item a finance team notices. Short clips hide the gap; long clips expose it.
Both lines start close at 4 seconds and diverge as clips run longer — the shaded band is what you save per clip by picking Seedance.
Resolution is the other axis, and it is where Seedance’s flagship earns its keep. It is the only one of the two that lists 4K. The honest question is whether you need it. Most social distribution caps at 1080p: Instagram, TikTok, and YouTube Shorts all downscale anything higher, so paying for 4K to feed a vertical story is money set on fire. Where 4K matters is masters that get re-edited, VFX plates that get cropped or punched in, and any deliverable a client will archive at full resolution. If that is your work, Seedance flagship is the only option here, and it still costs less per second than Wan at 1080p, so the resolution headroom is effectively free relative to the alternative.
The practical rule: pick resolution by your distribution ceiling, not by the highest number on the spec sheet, then pick the cheapest tier that clears it. For a feed, that is Seedance Mini at 720p. For a 1080p master, it is Seedance flagship. Wan enters the decision when the clip is shorter than four seconds or when you are already standardized on Alibaba, not when you are optimizing the bill.
When to Pick Seedance 2.0
Reach for Seedance 2.0 when volume and resolution both matter. The three-tier structure lets you meet cost to quality per job instead of paying one flat rate: Mini for throwaway variants and drafts, Fast for approved 720p output, flagship for hero clips that need 1080p or 4K. Because all three share one model-ID prefix and one endpoint, you can route by tier inside your own code without touching auth.
It is the right default for e-commerce product video, social ad variants, and any pipeline where you generate hundreds or thousands of short clips and iterate. The 4K ceiling on the flagship also means you do not outgrow it the moment a client asks for a higher-resolution master. The adaptive aspect ratio is a quiet convenience here too: for a mixed feed of 16:9, 9:16, and 1:1 placements, you can let the model fit the frame instead of maintaining three prompt variants.
Where it bites: nothing under 4 seconds, and the cheap tiers stop at 720p. If your whole business is 2-second loops or 1080p-minimum output on a budget, the tier math stops helping.
When to Pick Wan
Pick Wan 2.7 when the clip is short or when you are already invested in Alibaba’s stack. The 2-second minimum duration is a genuine capability Seedance does not have, and for loops, stingers, and micro-transitions that floor is the deciding factor, not the rate. Wan 2.7 also adds video-to-video, which Wan 2.6 lacks, so use 2.7 if you extend or restyle existing footage.
There is also a switching-cost argument. If you have prompt libraries, reference-image sets, or QA baselines tuned to Wan’s behavior, a $0.24-per-clip gap may not cover the cost of re-tuning everything to a new model. Price is one input; the migration bill is another. And if you are standardizing on Alibaba across image and language already, keeping video in the same family can be worth a premium for one vendor relationship and one set of content policies.
When NOT to Pick Either (and What to Use Instead)
Neither Seedance nor Wan is the answer for every job.
- You need synced lip-sync dialogue or a digital-human host. HappyHorse 1.1 (
alibaba/happyhorse-1.0/alibaba/happyhorse-1.1, $0.13/s) is built around lip-sync and multi-reference image-to-video, and takes up to 9 reference images. It costs more per second, but for a talking-head format it does the thing the others do not. - You need Western-market photorealism at scale, or Sora and Veo specifically. Those run on their own native APIs, not
/v1/videos. We compared them in AI Video Generation APIs Compared: Sora 2 Pro vs Veo 3.1 vs Kling 2.6 Pro, where the takeaway was the opposite of this post: three models, three separate SDKs, no shared endpoint. If Kling is on your list specifically, the Kling 2.6 Pro Video API guide covers it end to end. - Your workload is image, not video. For stills, the ByteDance image sibling is covered in Seedream 4.5 Doubao Image API.
Try Both via ofox: A/B on One Endpoint in 10 Lines
This is where Seedance vs Wan stops being a spec argument. Because both live behind ofox’s POST /v1/videos, you switch models by editing the model string. The call is async, exactly as described above: submit, get a polling_url, poll until completed.
Pay-as-you-go starts at $0.04/s on bytedance/seedance-2.0-mini, and the same key runs every model in the ofox video catalog, so you can benchmark both before you commit a cent to volume.
A/B both models in Python
import os, time, requests
OFOX = "https://api.ofox.ai/v1"
HEAD = {"Authorization": f"Bearer {os.environ['OFOX_API_KEY']}"}
def generate(model, prompt):
r = requests.post(f"{OFOX}/videos", headers=HEAD, json={
"model": model,
"prompt": prompt,
"duration": 8,
"resolution": "1080p",
"aspect_ratio": "16:9",
})
r.raise_for_status()
poll = r.json()["polling_url"] # 202 + polling_url
while True:
s = requests.get(poll, headers=HEAD).json()
if s["status"] in ("completed", "failed", "cancelled", "expired"):
return model, s
time.sleep(2) # poll no faster than 1/s
prompt = "A ceramic coffee cup on a wooden table, morning light, steam rising."
for model in ("bytedance/seedance-2.0", "alibaba/wan-2.7"):
print(*generate(model, prompt)) # swap the string, same call
The same call in Node
const OFOX = "https://api.ofox.ai/v1";
const HEAD = { Authorization: `Bearer ${process.env.OFOX_API_KEY}`,
"Content-Type": "application/json" };
async function generate(model, prompt) {
const res = await fetch(`${OFOX}/videos`, {
method: "POST", headers: HEAD,
body: JSON.stringify({ model, prompt, duration: 8,
resolution: "1080p", aspect_ratio: "16:9" }),
});
let { polling_url } = await res.json(); // 202 + polling_url
while (true) {
const s = await (await fetch(polling_url, { headers: HEAD })).json();
if (["completed", "failed", "cancelled", "expired"].includes(s.status))
return { model, s };
await new Promise(r => setTimeout(r, 2000)); // poll no faster than 1/s
}
}
const prompt = "A ceramic coffee cup on a wooden table, morning light, steam rising.";
for (const model of ["bytedance/seedance-2.0", "alibaba/wan-2.7"])
console.log(await generate(model, prompt)); // one string swaps the model
Image-to-video: same call, add a frame
To drive either model from a still, pass frame_images instead of relying on the prompt alone. The endpoint reads the payload and switches to image-to-video without a different route:
r = requests.post(f"{OFOX}/videos", headers=HEAD, json={
"model": "bytedance/seedance-2.0",
"prompt": "camera slowly pushes in, steam rises",
"frame_images": ["https://example.com/first-frame.png"],
"duration": 8,
"resolution": "1080p",
})
Two model families, one auth, one schema. That is the practical reason to A/B them instead of picking on the datasheet.
FAQ
Is Seedance 2.0 cheaper than Wan? Yes. Seedance 2.0 is $0.07/s at 1080p versus Wan’s flat $0.10/s, and the Fast and Mini tiers drop to $0.06/s and $0.04/s. On a 5-second 1080p clip that is $0.35 vs $0.50.
What is the model ID for Seedance 2.0 on ofox?
bytedance/seedance-2.0, bytedance/seedance-2.0-fast, and bytedance/seedance-2.0-mini. Wan is alibaba/wan-2.7 and alibaba/wan-2.6.
Does Seedance 2.0 support 4K?
The flagship bytedance/seedance-2.0 tier lists 4K. Fast and Mini cap at 720p, and Wan caps at 1080p.
Can both generate audio? Yes. All Seedance 2.0 tiers and both Wan versions list synced audio. Wan starts at a 2-second minimum, Seedance at 4 seconds.
Do I need separate API keys?
No. Both run through one ofox key and POST /v1/videos. Swapping models is a one-string change.
Which is better for image-to-video? Both list i2v. Seedance adds v2v on all three tiers; Wan 2.7 adds v2v, Wan 2.6 does not.
What is the cheapest video model on ofox right now?
bytedance/seedance-2.0-mini at $0.04/s, capped at 720p.
Sources Checked for This Refresh
- ofox model catalog API,
GET /v1/models(verified 2026-07-19): model IDs, per-second pricing, resolution, duration, and audio attributes for Seedance 2.0 tiers and Wan 2.6 / 2.7 - ofox Video API reference,
POST /v1/videosasync submit and poll schema: https://ofox.io/docs/api/videos (verified 2026-07-19) - ofox model detail pages (all HTTP 200, verified 2026-07-19): https://ofox.io/models/bytedance/seedance-2.0 · https://ofox.io/models/alibaba/wan-2.7
- ByteDance Seed, Seedance overview: https://seed.bytedance.com/en/seedance
- Alibaba Wan project site: https://www.wan.video/
- ofox video product page: https://ofox.io/video


