[Creative Picture Diary] The "Other Protagonist" That Achieved Over 80% Viewer Retention
Last time, I wrote about a secret technique for placing multiple characters in the same image, which was used before Nano-banana pro. I think it can also be used to supplement Nano-banana pro.
This time, I will also introduce some ingenuity for having multiple characters perform in the same image, but I would like to write about a narrative structure that is rarely seen, as it produced some interesting numbers.
YouTube Viewer Retention Rate
The viewer retention rates for the MVs I have released so far have followed a graph like this.


Other videos are similar, showing a graph where it drops significantly at the beginning and then follows a gentle downward slope, with about 50% of people watching.
However, my latest work, "Dream Gate," showed a slightly different trend.

Since the release periods are different, an accurate comparison cannot be made, but I am very happy that over 80% of people watched it until the end, with the graph remaining almost horizontal.
Since the overall number of views is low, this may not be a correct analysis, but I believe one of the factors lies in the structure of the story. The theme of this story started with the question of what kind of content would be useful for life-sized projections of people on large station displays (digital signage).

Other candidates included a guide for funeral halls or ceremony venues, a drawing exhibition where children show their own work, a lost cat notice on a poster stuck to a bench wall, and railway etiquette awareness campaigns.

In other words, the protagonist of the story is the stage itself, which is a feature of this video. It is a type of narrative structure where various people act out their own dramas on that stage. A famous example of this is the overseas drama ER.
I think that by compressing this into two and a half minutes, it created a motivation to keep watching as if it were a serialized series that serves as a vertical axis throughout, despite its short length.
It was only after that that I decided to make the central story of this video about a cat version of Hachiko the faithful dog.
And to express a series of short dramas, I created lyrics that repeat the same wording so that the song would be in a chaconne style, repeating short phrases rather than the general [Verse A] [Verse B] [Chorus] structure. The music generation AI understood that well and created several songs with repetitive melodies.
Another version of the song ☞ https://suno.com/s/B9VRb2aiL1lDyFd4
Such repetition of wording is called anaphora/epiphora, and in this song, "If I were to be reborn..." and "...want to" correspond to that. I was impressed that everything has a name.
The other protagonist: The display in front of the station ticket gate
Next, I will write about the hardships and ingenuity involved in depicting the stage in front of the station ticket gate, which is the other protagonist.
First, regarding the base background image of the ticket gate, I struggled quite a bit because I couldn't generate one that was just right.

I made it entirely by relying on prompts, but I would like to try using 3D models in the future.
You might have noticed, but this scene in front of the station uses chroma key compositing for the video. That is the footage on the display.
With the latest AI, it might be possible to embed video using just prompts, but when instructing multiple people to perform specific movements using only primitive video generation AI, it was effective to combine separate videos using old-fashioned chroma key compositing to reduce the "gacha" element. (It's also because my prompt skills are still immature 😅)

For standard background chroma keying, green or blue is common, but since I was cutting out a display screen, the background contained a variety of colors, so I filled it with an unrealistic color (shocking pink).
A minor problem is that in early morning or night scenes, if I instruct a pink screen in image generation, the floor gets a slight pink tint from the reflection, so I generate it with a white screen and fill it with pink using image editing software.

This cut is a scene where the person inside the display and the person outside synchronize, so I created and synchronized videos for both the inside and outside separately.
I took the baseball player and the background out of the generated image of the shooting pose, cut out the boy, put a hat on him, and pasted him onto the chroma key background image to use as the base for the end image of the video.
On the other hand, for the baseball player inside the display, I used the same original image but this time deleted only the boy and generated the video as an end image.

A minor problem here is that when I deleted the boy with Nano-banana, the baseball player's hand moved because it accounted for the boy being gone, and the synchronization wasn't quite right.
Specifying to delete the boy but keep the hand as is didn't work, and what worked well was instructing it to "set the boy to 100% transparency," which kept the hand from moving. (However, the shadow remained.)

Also, when generating the image of the mother taking the photo, using the base of the end image mentioned above caused pink to reflect in the smartphone she was using to shoot, so I generated the mother using an image that included the baseball player.

For the crowd outside the display, just like I did in the previous article, I cut out usable people from the generated gacha images and rearranged them.
Because the gacha images were generated using the background as a reference image, the colors and lighting of the people are natural and blend in with the background. There was some variation in the size of the people, but since they are pasted on later, the size can be adjusted freely.
A little trick is that even if you cut out small images or use a background removal tool with poor accuracy, you can clean them up and make them look nice by regenerating them individually with Image to Image in Nano-banana.
In the climax of the second half, the characters who have appeared up to that point are used as foreshadowing and are brought back for the ending.

There is no crowd here, but it is a cut where the display and the space in front of the station are seamlessly connected.
Only for this cut, I didn't use the video chroma key function, but instead created an image of the hero reflected in the display and generated the video from that still image.
This kind of direction is quite common, but it emphasizes the scene by contrasting it with the depiction of the display in the first half.
Having drawn pictures in the past, I find that putting in various analog efforts to create the footage exactly as I imagine is troublesome, but at the same time, it brings a joy of creation that you can't get from just hitting a button.
As AI tools evolve, this kind of enjoyment might be swallowed up by efficiency. 😔

