[Verification] Will a video generation model that surpasses Seedance 2.5 emerge? / Lessons learned from the 95-minute AI feature film "Hell Grind."
Continuation of Seedance 2.5 verification
For a while, I have been conducting comparative verification with MiniMax H3 and the older Seedance 2.0 model, but since I was able to confirm that Seedance 2.5 is overwhelmingly superior the verification has concluded.
※This verification is strictly for "AI drama production with live-action expression." I have not verified anime-style expression.
Previous verification:
Prompts optimized for 2.5 become long
Because prompt fidelity has improved, there is a tendency for the model to "expand on its own" if you do not provide the necessary information. Because it prioritizes stable movement over creativity, the prompts expanded by 2.5 are "realistic and static." It seems less likely to be creative like 2.0.
When you provide the necessary information, the total amount of the prompt also increases, resulting in a fairly long text. On the day Seedance 2.5 became available, Runway's input character limit was also expanded to "15,000 characters."
Runway could only accept "3,500" characters until then, so I struggled with adjustments. Even with 2.0, there were times when it exceeded 4,000 characters, so when it overflowed in English, I translated and compressed it into Chinese to handle it. It's a desperate measure, though.

Currently, episode 3 of the AI drama in production is using Seedance 2.5 (switched from 2.0).
The video below is a scene I gave up on in 2.0 because the difficulty was too high.
The official user guide states that "time ranges are time frames assigned to events, not precise edit points in frame units," but in the case of 2.5, it is effective as acting control.
Playback time: 14 seconds
Seedance 2.5
The "Seedance 2.5 exclusive" scenario prompt for the video above:
A scene from a big-budget crime action movie.
3-shot composition. All shots are filmed in the same alleyway.
Shot 1 / 0.0 seconds to 2.0 seconds:
A high-angle crane shot looking down at a sunny daytime New York alleyway (@NYC_alley_B). The camera is fixed at a height of 12 meters from the ground, and the lens is tilted 30 degrees forward from straight down.
Four official NYPD patrol cars (@NYC_PoliceCar) are parked at irregular angles in the alley. The red and blue LED light bars on the roofs of the four cars are flashing.
Direct sunlight casts hard-contoured shadows on the asphalt and brick walls.
12 police officers are standing around the vehicles. The 12 people vary in height from 165 cm to 190 cm, with a mix of 8 men and 4 women. Everyone is wearing dark navy blue long-sleeved uniform shirts and dark blue 8-point caps. 4 people are crouching and looking at the ground. 3 people are stretching yellow caution tape with both hands. Because the camera is high, the tops of the caps and both shoulders appear large, and the faces are half-hidden by the brims of the caps.
Shot 2 / 2.0 seconds to 9.0 seconds:
The cut changes. The camera is fixed on the cobblestones of the alley, 2.2 meters above the ground. The lens is pointing 8 degrees downward from the horizontal. The camera is positioned 9 meters in front of the front end of a black SUV (@Car_A1) and is facing the vehicle directly.
The screen shows the entire front of the SUV (@Car_A1), the cobblestone road surface on the left and right sides of the vehicle, and the cobblestone road surface between the vehicle and the camera. The full width of the vehicle is within 50 percent of the screen width. The roof of the vehicle is below the top edge of the screen.
The right side of the screen is the driver's side of the vehicle, and the left side of the screen is the passenger's side of the vehicle.
2.0 seconds to 2.6 seconds:
The SUV (@Car_A1) stops with screeching tires, and the car body sinks forward and then returns.
2.6 seconds to 3.2 seconds:
The driver and passenger side windows are fully open. Through the windshield and the open windows, Hitomi's face (@Hitomi_Cap_A) is visible on the right side of the screen, and Becky's face (@Becky_Cap_A) is visible on the left side of the screen. Both are wearing dark navy blue FBI caps and dark navy blue jackets with yellow "FBI" lettering on the back.
3.2 seconds to 4.2 seconds:
The door on the right side of the screen opens 70 degrees outward, hinged at the front, and Hitomi (@Hitomi_Cap_A) steps onto the cobblestones with her right foot. At the same time, the door on the left side of the screen opens 70 degrees outward, hinged at the front, and Becky (@Becky_Cap_A) steps onto the cobblestones with her left foot. At this point, Hitomi (@Hitomi_Cap_A) and Becky (@Becky_Cap_A) are visible from head to toe on the screen. Both of their shoes and the cobblestones are visible on the screen.
4.2 seconds to 5.0 seconds:
Hitomi (@Hitomi_Cap_A) grabs the top edge of the door with her left hand, extends her arm and pushes, and the door closes and seals against the car body. At the same time, Becky (@Becky_Cap_A) grabs the top edge of the door with her right hand, extends her arm and pushes, and the door closes and seals against the car body. Both of their upper bodies are twisted in the direction of pushing the door, and both feet are shifted on the cobblestones. At the 5.0-second mark, both the left and right doors are closed.
5.0 seconds to 6.4 seconds:
Hitomi (@Hitomi_Cap_A) walks 3 steps forward on the cobblestones to the right side of the car body. At the same time, Becky (@Becky_Cap_A) walks 3 steps forward on the cobblestones to the left side of the car body. Both walk side-by-side in front of the front bumper of the SUV (@Car_A1). Neither of their feet stops even once.
6.4 seconds to 7.6 seconds:
Both walk 2 steps on the cobblestones toward the camera. While walking, Becky (@Becky_Cap_A) moves her mouth and says {There are a lot of incidents today, aren't there?}. Neither of their feet stops even once.
7.6 seconds to 9.0 seconds:
Both walk 3 steps on the cobblestones toward the camera. While walking, Hitomi (@Hitomi_Cap_A) raises the corners of her mouth, moves her mouth, and says in a low voice {This is our job.}. Neither of their feet stops even once.
At the 9.0-second mark, both are visible from the waist up on the screen. Hitomi (@Hitomi_Cap_A) is on the left side of the screen, and Becky (@Becky_Cap_A) is on the right side of the screen.
Shot 3 / 9.0 seconds to 15.0 seconds:
The cut changes. The camera is positioned on the cobblestones of the alley, 1.7 meters above the ground, filming the back of the alley over the backs of Hitomi (@Hitomi_Cap_A) and Becky (@Becky_Cap_A) from 1.5 meters behind them.
At the first frame of 9.0 seconds, the back of Becky's head (@Becky_Cap_A) and her left shoulder are visible on the left edge of the screen, and the back of Hitomi's head (@Hitomi_Cap_A) and her right shoulder are visible on the right edge of the screen. The backs of their heads and shoulders are large, filling the height of the screen, and are out of focus with blurred outlines. Between them, the cobblestones leading to the back of the alley are visible. Both are already in the middle of walking away from the camera, and their walking speed and stride are the same as at the 8.9-second mark.
At the first frame of 9.0 seconds, a police officer (@NYC_Police) is standing on the cobblestones between them, 8 meters behind the camera, and is facing the camera. This police officer is a 183 cm tall man wearing a dark navy blue long-sleeved uniform shirt, a dark blue 8-point cap, a silver badge on his chest, and a black leather belt. Behind the police officer (@NYC_Police), yellow caution tape stretched at waist height and a parked NYPD patrol car (@NYC_PoliceCar) are visible.
9.0 seconds to 10.2 seconds:
The police officer (@NYC_Police) walks 4 steps on the cobblestones toward the camera and stops at a position 4 meters from the camera. At the same time, Hitomi (@Hitomi_Cap_A) and Becky (@Becky_Cap_A) walk 2 steps away from the camera. The camera moves forward, maintaining a distance of 1.5 meters behind them. While walking, Becky (@Becky_Cap_A) turns her face to the right to look at Hitomi (@Hitomi_Cap_A), and while showing her right profile to the camera, she moves her mouth and says {Well, I suppose that's true~}.
10.2 seconds to 11.0 seconds:
Hitomi (@Hitomi_Cap_A) and Becky (@Becky_Cap_A) stop walking. The camera also stops moving forward. The police officer (@NYC_Police) is standing in the center of the screen with his face and chest facing the camera directly. The police officer's (@NYC_Police) head to knees are within the screen, and his face is in focus. On the left and right edges of the screen, the backs of the heads and shoulders of Becky (@Becky_Cap_A) and Hitomi (@Hitomi_Cap_A) are visible, out of focus.
11.0 seconds to 11.6 seconds:
The police officer (@NYC_Police) raises his right hand and touches the edge of his cap's brim above his right eye with his fingertips. His face remains facing the camera directly. He lowers his right hand to his side in 0.6 seconds.
11.6 seconds to 13.0 seconds:
The police officer (@NYC_Police) moves his mouth while keeping his face toward the camera and says {I've been waiting for you. Could I have a moment?}.
13.0 seconds to 15.0 seconds:
The police officer (@NYC_Police) turns his body 180 degrees to face away and starts walking toward the back of the alley. Hitomi (@Hitomi_Cap_A) and Becky (@Becky_Cap_A) start walking after him, walking side-by-side 1 meter behind the police officer (@NYC_Police). The camera moves forward, maintaining a distance of 2 meters behind the three, and continues to capture their backs and the yellow "FBI" letters on the backs of their jackets. None of their feet stop even once.
Ambient sound:
The sound of the SUV's (@Car_A1) tires rubbing against the cobblestones and the sound of brakes. The sound of two doors closing at the same time. The sound of Hitomi (@Hitomi_Cap_A) and Becky (@Becky_Cap_A) walking on the cobblestones. A distant police siren.
No Music. No BGM.
No subtitles. No captions. No title super. No text. No letters. No numbers. No logos. No watermarks.
If you pack too much into the timestamps, excessive cuts and omission of events will occur, but depending on the adjustments, it can be quite useful.
Trying to regenerate episode 2 with 2.5
Episodes 1 and 2 of the AI drama were produced with Seedance 2.0, but episode 3, which is currently in production, uses 2.5.
Episode 2 produced with Seedance 2.0:
When I re-generated it using 2.5 with the same references and prompts used in Episode 2, the detail in the characters' skin was clearly improved.
When pursuing realism in 2.0, moles would appear, and unnatural skin graininess was noticeable. For example, adding Y2K elements would always result in glitter appearing on the cheeks of young female characters.
While I used After Effects features to remove moles and improve skin texture, 2.5 generates them without any issues. Unintended elements are not added on their own.
Playback time: 39 seconds
Re-generating the Episode 2 scene with Seedance 2.5
2.5 is better suited for pursuing realism than 2.0, and I think it is highly likely to become the standard for live-action AI drama production.
However, as mentioned above, if you do not provide detailed information, it creates safe and static extended prompts that suppress creativity, so you cannot expect creative visuals like Sora 2.
It is not a video generation model for beginners (even in terms of cost).
It is clearly professional-grade.
The 95-minute AI feature film 'Hell Grind'
The 95-minute AI feature film 'HELL GRIND' produced by Higgsfield AI has been released on YouTube.
Seedance 2.0 was used throughout, and it is a work that proves the high visual expressive power of the model.
According to a LinkedIn post by founder Alex Mashrabov, it was 'produced on Higgsfield by a team of 15 professional directors, cinematographers, and editors in 14 days for less than $500,000'.
$500,000??
About 80 million yen!
In an interview with the WSJ, it was reported that compute costs reached $400,000 of that total, and that by using cloud providers like Nebius and CoreWeave instead of major hyperscalers, they kept costs from ballooning further.
In other words, the remaining amount is about $100,000 (around 16.2 million yen), which covers all labor costs, sound, editing, and promotion.
According to Higgsfield's announcement, for the first 25-minute segment, '16,181 video generations were required to obtain 253 final cuts'.
This figure was also posted on the company's official X account and reported by multiple media outlets, including Screen Daily.
According to Adil Alimzhanov, who serves as Content Lead at Higgsfield, the core of the production is the work of continuously generating 15-second clips until they reach a usable standard. Mashrabov describes this process as 'feeling like a slot machine'.
In other words, just keep generating endlessly.
You continue without compromise until you reach a usable standard.
What someone who finishes a long-form project realizes is that 'there is no magic spell!', and no matter how much you refine prompts or systematize, in the end, it is brute force.
25 minutes for 253 cuts means an average shot length of just under 6 seconds. This is a reasonable standard for an action work. It means an average of 64 attempts were required per cut.
The following 3-hour tutorial video is the 'most educational' video generation tutorial I have ever watched.
'Generate dozens, hundreds of times' until you get the intended result.
In short, it is a video about generating over and over again.
No matter how much you automate with AI agents, no matter how much you optimize your prompts, you end up generating repeatedly 'until you achieve the best expression'.
That is what these '3 hours' are about.
Even if we generate 100 times and select the best video, if the director says 'NO!', we generate it all over again.
Corporate videos and promotional footage that are easy to template can be optimized, streamlined, and automated using AI.
However, the truth is that for stories (movies, dramas, etc.), tedious and unglamorous work is unavoidable. Of course, the fact that there is 'no filming' itself due to AI is revolutionary.
Writing physical laws into prompts
All assets and prompts used for this work are published on the same site. For creators attempting to take on long-form AI video, this is the most information-dense case study available.




To maintain consistency between shots, the prompts become long and detailed, averaging about 3,000 characters (with some cuts exceeding 10,000 characters).
It is noted that these included phrases to remind the model to adhere to physical laws, specifications for intended cinematic choices, and descriptions to avoid the AI-specific sheen.
What is noteworthy is that what is written is'physical facts' rather than descriptions. Instead of abstract words like 'realistic' or 'cinematic,' they write specifically about gravity, inertia, mass, contact shadows, and props that do not float. Camera intentions are also specified just as concretely. The volume of 3,000 characters on average is a density close to replacing a storyboard with text.
According to the WSJ, one of the tools Higgsfield sells to clients is a mechanism that generates these very complex and detailed prompts. It is designed so that when a user inputs one page of a script, it returns a prompt spanning thousands of characters, which produces production-quality output.
In other words, the actual workflow is not 'humans writing prompts,' but athree-layer structure of '(humans) writing a script, having it expanded (by AI) into a long-form prompt, and selecting the output'.
Even in individual production, whether or not you can prepare this middle layer yourself will likely greatly affect efficiency.
Audio was a clear weakness
There is almost no official technical disclosure regarding the audio.
What can be confirmed is that in an interview with Variety, a Higgsfield spokesperson stated that all the music in the film, except for one track, was AI-generated, and that Mr. Mashrabov himself admitted that he received feedback on how important professional voice actors are to a film, acknowledging that this was clearly lacking.
The fact that the person in charge of the project publicly admitted that 'voice acting' was a weakness is significant.
As the level of achievement on the visual side increases, the roughness of the audio becomes relatively more noticeable.
Conversely, allocating human resources to voice acting and sound design could be the most cost-effective form of differentiation at this moment.
Even if the visuals are high quality, if the intonation of the language spoken by the characters is unnatural or there are mispronunciations, the sense of reality drops instantly.
This method cannot be imitated
The last minute of the aforementioned 3-hour tutorial video is the most shocking.
The total number of generated images/videos was '48,336', and about 800 assets were produced.
However, only '8' were actually adopted...

The approach changes between a production environment with abundant credits and an environment like ours, where we are always managing with few credits while considering 'savings'.
Simply put, a strategy for the underdog is necessary.
There were many things to learn from Higgsfield's workflow, but I think if you adopt this method for business (such as client work), it will not be profitable at all, so it seems quite difficult for anything other than entertainment aimed at large-scale fundraising.
However, (as I have written many times in the past), for independent filmmakers, this is the arrival of a dream-like era, and I wonder if such a great opportunity will ever come again.
Related Magazine:
Updated: Friday, August 7, 2026 / Published: Friday, August 7, 2026
