Seedance 2.5: 30 Seconds of Native 4K from One Prompt, and 50 References to Feed It
林政賢 ·
Anyone who makes video knows the pain: the AI footage looks gorgeous, then at second 4 the lead has a different face and at second 7 the jacket changes colour on its own. Seedance 2.5, which ByteDance unveiled at the end of June, is pitched squarely at fixing that.
The numbers first, because the jump is big
Seedance 2.5 was announced on 23 June 2026 at Volcano Engine's FORCE conference. The version number is telling: the previous release was 2.0, and this one leaps straight to 2.5, with the four versions in between officially "skipped". The message is that this isn't a patch, it's a new generation.
SEEDANCE 2.5 — OFFICIAL SPECS
Audio is no longer pasted on afterwards either. This version processes sound and picture in the same latent space, so the door slamming and the sound of the door slamming land on the same frame. No more lining up sound effects by hand.
Which line actually matters if you make video
Not the 4K. Not the 30 seconds. The 50 references.
Clients almost never reject AI video for image quality. They reject it for continuity: the same character looking different across shots, the same coat changing colour from a new angle. Being able to hand the model 50 references at once — character turnarounds, location photos, a colour chart, even a music cue — turns what used to be adjectives and luck in a prompt into simply showing it the picture.
The other practical one: local editing. Before, if the lead's hair colour was wrong you regenerated the whole 30 seconds and something else would shift. Now you circle just that area and the rest stays put. During revision rounds on a pitch, what this saves isn't compute, it's your heart rate.
There is also a pre-production feature called 3D white-box preview: generate a low-fidelity white-model animation first to check camera moves and blocking, then commit to a full-quality render. That is exactly the logic of previz on a film set, except it used to mean opening 3D software and now it lives in the same interface.
When you can get it
At announcement the enterprise beta was already live, with the public launch targeted for early July. When this was written we hadn't yet made it onto the trial list, so every number above is the official spec, not our measurement. Once we've actually run it there will be a separate piece with a "claimed vs. actual" comparison.
The competitor at the same moment is Google's Veo 3.1 (native 4K, audio, up to three reference images). Three versus fifty is the number this release most wants you to remember.
Our take: if your pain is "characters don't match" rather than "not enough resolution", this one belongs on your trial list. If you only need 15-second social clips, current tools are already enough — don't swap your whole pipeline for 30 seconds.
About this piece: the specs come from the official announcement at Volcano Engine's FORCE conference on 23 June 2026 and press coverage at the time (30 seconds / native 4K / 50 references / 10-bit / +20% prompt adherence / enterprise beta live, public launch early July). This is an information brief, not a hands-on test; measured numbers will follow separately.
Author:林政賢(Director · Gen AI creator & engineer · Founder of TangYi Studio)