SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Creating with AI in 2025

This note was written at the request of the organizers as a reference piece for the "#CreatedWithAI" contest by Google Gemini × note.

I previously wrote a note like this.

It was about how I created my own portfolio site using the latest AI tool, Figma Make.
Many people read it, and some even reported back to me saying they tried it out and made their own portfolio sites! I was really happy about that.

(A tweet from usutaku-san that personally made me happy!)

I usually lead a team called "Firstthing," where I work as a creative director for advertisements and brands, and as a video director for projects like music videos.

A topic that has been heating up in this industry lately is, "How should we face AI?" that is the question.

Jobs might be taken away. People might become unnecessary. Isn't it bad to rely entirely on AI? ...These voices are overflowing, and currently, there is more anxiety and doubt than anticipation. As I face AI myself, while I find it fun, I often think that creators themselves need to improve their literacy and exchange opinions more with each other on how to use it appropriately.

In this article, based on my actual experiences and lessons learned on how I used AI to face creative work this year and how I "tried creating with AI," I would like to consider the ideal relationship between humans and AI in the future and tips for facing it.


Don't leave "0→1" to AI

I think people often use AI for brainstorming. A single issue, question, topic, or something that needs to be solved at work... Using those as prompts to output a zero-to-one result as is, and then having a human brush it up further... I think that way of using it is common now.

On the other hand, since AI basically outputs the collective intelligence of all information, naturally it doesn't have a personality like an individual, and its ideas (while depending on the prompt) are basically "safe" and "average."

I sometimes teach creative classes to students and others, and I had the impression that especially this year, the assignments from any student tended to be similar. They probably entered the assignment I gave as a prompt into AI once, and then shaped the ideas proposed by the AI with their own hands. In other words, they had the AI do the "0→1" and used their own hands for the "1→10." However, when you do that, everyone naturally ends up with the same answer. Of course, there is a difference between excellent ones and those that aren't, but they were all somewhat similar and average.

I hear that lately, when thinking with AI, people often have the illusion that "I feel like I've become smarter" or "I feel like things are going well." I think that such self-efficacy can have a positive effect on people in the end, but on the other hand, I feel like there are many cases where people don't push their ideas any further once they can get an average output from AI. When a nice-looking image or video comes out from AI, they think "I did it!" and their thinking stops. There are many "people being eaten by AI," and even those who have potential are often unable to demonstrate that ability.

When there are many things like that, people who think of unconventional ideas or unique perspectives without using AI naturally stand out. Also, things like project proposals written by hand with great effort stand out more than before. It might be an era where, as a reaction to AI, human effort and individuality are becoming stronger.

What is the issue that needs to be solved? Instead of having AI think about that, I think it is important to finish the "0→1" with human hands, and then fully brush up the "1→10" and "1→100" that follow together with AI. I think it is important to start by doing it yourself.

This is sudden, but please watch the video below.

This video is a special video I directed this year to commemorate the 10th anniversary of MangaONE, a manga app from Shogakukan. It is a video that conveys the appeal of "MangaONE," where no two works are the same, by weaving together the lonely worries of manga artists themselves through the words and images of two completely different manga artists. I wanted to convey the splendor of works not being homogenized, praising the manga artists who cannot become one.

This video work does not use generative AI for the video itself, but I was very helped by AI when brushing up the content.

For this work, I finished the "0→1" myself.
I did not receive any advice from AI for the copy or the story. Furthermore, to add reality, I interviewed many manga artists, talked about their actual lives, hobbies, daily routines, room structures, lives, and what they usually think about, and polished the copy itself.

After fully fleshing out the copy, the vast amount of research material, and the original story structure, I utilized Gemini's Canvas and DeepResearch features to innocently ask questions about everything from things related to the work to side topics—such as 'What are the potential desires and worries of these two characters?', 'Are there any similar manga artists?', 'Among the manga artists we interviewed, who has a way of thinking close to this character?', 'If we look at this work philosophically, which philosopher's thinking is it closest to?', and 'If he were to suffer a major setback, could he recover?'—and brushed up the character images until the day before the shoot.

This process was incredibly interesting.
Before I knew it, I was brainstorming with Gemini until late at night, and the feeling of the character image solidifying within me led to confidence. Like manga artists, directors can be quite lonely, but having someone to talk to who responds to my interests made my direction during the shoot much sharper. Once the character image became clear—like 'This character definitely wouldn't make this gesture'—the sharpness of my own direction during filming changed dramatically.

Furthermore, by loading the materials created with Gemini into NotebookLM along with the preliminary materials and copy, a notebook dedicated to that work is born. When I ask an AI that has only read my accumulated thoughts, 'Would this character say something like this?', it responds based on the sources, saying things like, 'They might not say that. Because...' That said, there were times when I had doubts about the AI's response or felt a gap with my own sensibilities, and in those cases, I honestly communicated even that. The more I interacted, the more the story and worldview of the work felt like they were being polished, which was very interesting.

What is important when loading into Gemini is that the primary materials are carefully organized in spreadsheets or text. Even if they are too scattered, it has been able to recognize them well recently, but the more careful the prompt, the more you can reduce that 'feeling that it didn't quite get across.' In the end, I feel that the human's level of care is directly linked to the quality of the response.

Even though I said I brushed it up, the 'copy' that forms the core of this video did not reflect the AI's opinions at all, and I hardly changed anything from what I first proposed to everyone at Shogakukan. While I might use AI to verify the certainty of those words along with interview data, the words are the innovation of this work, so I do not change them. It is easy to think that what the AI says is correct, but if you don't properly hold onto the human side's convictions, the work will end up being shaky.

I recently watched a YouTube video by Ichiro Yamaguchi that was very interesting.

In this video, he says that AI can create things quite perfectly, not just from '0 to 1' but even '0 to 100,' but that the charm of a work is something that originally improves through margins and fine human adjustments, so even if something good is produced from the start, it only results in a 'That's amazing.' This opinion is something I truly agree with, and in the end, the person's own unique adjustments and the line-drawing and direction of 'where (and how far) to use AI' become important factors for creating something attractive. And I feel that in the coming era, the value of such individuality, humanity, physicality, and things with substance will increase.

Because you can do anything, you dare to limit it

I wrote earlier that because AI is collective intelligence, its output tends to be 'safe.' To go beyond that, I believe it is important to understand the AI's strengths and weaknesses and dare to create limitations.

There are many characteristics of Gemini that I use on a daily basis, but among them, I think the fact that it can 'read a massive amount of information' and is 'multimodal' are amazing points that are surprisingly not well known to the world.

In Gemini 3, it can process longer text and documents at once than before, and it has succeeded in expanding the context window. This means that it can analyze a massive amount of text data at once, perform summarization, translation, question answering, extraction of insights, and so on. There is a theory that this limit on the amount of text is enough to easily read the entire Harry Potter series at once...

Furthermore, it can read not only text but also various formats such as images and videos in a multimodal way. What's amazing is that, for example, if you have Gemini read a dance lesson video from YouTube, it reads the video of the dance itself and analyzes things like, 'At this second, they are doing this movement.' It's not just converting YouTube subtitles into language; it's reading the pictures, sounds, and movements themselves.

I tried having it analyze a music video I made previously. It understood the atmosphere, the explanation of the cuts, and even the emotional nuances.

While leveraging these 'things AI is good at,' how can we make the things AI is 'not good at'—that is, things where the answers tend to be 'safe'—attractive? I thought that to create an attractive AI rather than a safe one, we actually need biased information. Special perspectives that only that person can teach, or sharp perspectives.

What was created after pursuing that is the project I was involved in as the creative director of Firstthing, the project with the magazine 'BRUTUS', 'Hello, Brutus.'.

This project created a new experience where Gemini learned all the pages of 45 years of the magazine 'BRUTUS,' and people could talk to that AI on the phone for just 3 minutes, allowing them to converse with the magazine.

BRUTUS was born in 1980 as a magazine for city boys who were born in POPEYE to read as they got a little older. At the beginning, there were many mischievous articles that couldn't be featured today, and I think it's like a 'reliable, knowledgeable older brother who doesn't deny or reject various cultures' without specializing in just one culture. My father was a huge fan of BRUTUS, and he was such a fan that I thought he really had every issue, so it was a magazine I was familiar with since I was a child.

On the other hand, while people are talking about an 80s boom now, even if young people know the atmosphere and what is somewhat popular, I think most of it is wrapped in a black box, and there must be information not on the internet in the 45 years of accumulation.

In this phone booth, you can freely talk to a personality called Brutus, who has a unique perspective as a culture expert. You can ask about past culture, ask for advice on love, or ask for the menu of a delicious restaurant you want to go to next weekend. By limiting and sharpening the information, I was able to leverage the AI's strengths while compensating for its weaknesses.

What Koichiro Shima wrote in the article above puts into words the vision we wanted to realize.

The project 'Moshi Moshi, Brutus.' was extremely thought-provoking, offering a glimpse into the future of AI agents. It’s something I’ve been feeling myself lately. That is, isn't it far more fun to talk to an AI agent filled with deep-seated passions than to an AI that reads vast amounts of information comprehensively and gives model-student answers?
This is because the responses from AI trained on massive amounts of data felt somewhat flat, monotonous, lacking in contrast, and a bit underwhelming, as if they lacked substance.
Perhaps there are new discoveries to be made in dialogue with an AI agent that offers alternative perspectives rather than just the greatest common denominator of correct answers. 'Moshi Moshi, Brutus.' was a project that realized this a step ahead of everyone else.

Individuality emerges beyond inefficiency

When talking about AI, 'work efficiency' is often cited as one of its charms.
On the other hand, isn't it precisely the refusal to pursue efficiency that is important for creation?
In other words, it means daring to do things the hard way. That is where individuality should emerge.

Kunihiko Morinaga, who launched the fashion brand ANREALAGE, once said when talking about 'talent' that it is important to 'be able to take as many detours as possible compared to others.' I love this sentiment.

I interpret this as being the same when creating a single work: you cannot produce new things or compelling value unless you do so inefficiently, discontinuously, and through non-reproducible thinking, rather than arriving at an answer in one go. Personally, I have rarely produced anything good by working efficiently, and I feel that learning to love inefficiency is important for creation. Even if you can see the answer, you must properly doubt it and make it better. When refining a job with a client or artist, it is through repeated dialogue and trial and error that you find the exit together and arrive at creative work that moves people's hearts. That process of loving the unknown is also the fun part of being creative.

AI efficiently gives you something that looks like the correct answer in one go. Depending on the job, that might be fine. However, when creating or producing something, if you don't make an effort to doubt that answer, or rather, have the AI provide 'questions' instead of answers and think for yourself, you may end up with something mediocre, like the student assignments I mentioned at the beginning.

Changing the subject slightly.
In May of this year, Google released the new video generation AI 'Veo3' and the tool 'Flow' that richly assists in generation, which came as quite a shock to the video industry.

This year can be called the first year of video generation AI, the first year where we could truly generate rich video content.It was in such a year that the initiative 'Music Video with Gemini' began.

'Music Video with Gemini' is an initiative where leading Japanese video directors utilize Gemini and team up with musical artists, aiming for 'each of them to create new music video expressions with AI in their own way.' So far, I have had the pleasure of working with such illustrious seniors as director Yuka Yamaguchi x LAUSBUB, director Tsuyoshi Nakamura x TOWA TEI, director Jun Tamukai x muque, and director Kazuhiko Hiramaki x Pasocom Music Club.

The reason I love music videos is that there is an aesthetic born from the necessity of having to be inventive due to cost and time constraints, and it is the only environment where you can take on visual challenges while working together with the artists themselves. Many talents like Michel Gondry and Spike Jones were born from music videos overseas, and I think that trend is also strong in Japan. I thought that creating new expressions with AI would be highly compatible with this environment that favors experimental expression.

As a result, the approaches differed completely depending on the director. There were those who spent two months hitting the wall with Gemini every day to continue generating video, those who tried filming once and then connected generations like a game of shiritori based on those photos, and directors who pursued the desire to make impossible dances happen. By focusing on creating images that had never been seen before, rather than copies or reproductions of reality, we were able to gain many insights that will serve as a guide for how video creators and AI will interact in the future.

First, the fact that a certain kind of AI glitch or flaw—the fact that things never go according to the storyboard—was actually fun for the directors as a 'gacha' (random draw). In commercials, movies, and dramas, it is basically tough if things don't go according to the storyboard, but music videos have the advantage that even if it's a 'gacha,' it can be made to work through editing as long as it matches the sound and worldview. There was a surprise in the fact that the 'unknown'—not knowing what will come out and the quality of generation changing completely depending on the prompt (instruction)—was established as a positive unknown. In creation, betrayal and ad-libs become 'acceptable' when picked up by the director's sensibility. I thought that this is the same feeling as when you think, 'Oh, this cut is actually pretty good!' on a film set, and it is the same with generative AI.

Also, individuality emerges in what you choose from what is generated (= how you edit it). I thought that as more people start producing with generative AI, video directors might become more like 'editors.' This was the same when CG and motion graphics emerged, but it might become even more pronounced.

And above all, perseverance and patience. If you think you can make videos efficiently with AI, it's actually not the case; the more you pursue making something good, the more you see the current limitations of AI. By understanding those limitations, directors themselves should be able to understand with high resolution how to incorporate AI into future production flows, so I would be happy if video creators would not be afraid to try using AI tools to create videos themselves at least once. Even if it seems perfect, that is only the prowess of descriptive power, and how to compose and edit it to make it work as a music video is a completely different issue. Surely, editing will also become easier to do, but when that happens, there is a possibility that completely unprecedented editing or expressions that only humans can create will become popular.

While advancing this initiative, I thought that the return to video with reality and substance would accelerate in the future. I think there will be a stronger return to expressions that have human-specific roughness, tactile feel, and that 'person-ness' that only that person can achieve. It is a strange and interesting phenomenon that the more AI evolves, the more the value of humans themselves seems to rise.

Using the self-efficacy born from AI as a springboard

The other day, I started a school called 'toracoya' at a place called TOKYO NODE LAB in Toranomon Hills.
This is a school for people who want to cultivate an entrepreneurial spirit in creativity and learn 'ideation,' 'communication,' and 'realization.' 33 people aged 18 to 47 gathered to learn real creative skills that cannot be learned from books or the internet through workshops from leading Japanese creators, entrepreneurs, and investors.

As the final lecture, I taught a class called 'AI Prototyping' with Fujiki-kun from CyberAgent, Axe-kun from Dentsu Lab Tokyo, and AI master Yamanishi-kun, and I will talk about that at the end.

As an AI buzzword this year, 'vibe coding' became popular globally. This is the idea that even if you don't know code, you can create apps, games, services, and websites (and of high quality, too) with generative AI. The example of creating a portfolio site with Figma Make, which I posted at the beginning, can be said to be vibe coding itself.

Recently, there have been many cases where entrepreneurs who have actually utilized this 'vibe coding' to start their own businesses have ended up building their own services. Newly born startups with no money want to build prototypes to realize their vision cheaply, and they want to keep refining them by listening to user feedback. For such people, vibe coding is optimal in terms of both speed and cost, and it has become a powerful tool for scaling a business in its early stages.

Google AI Studio

Our attempt this time was to teach such vibe coding in a 2-hour and 30-minute class to students, many of whom were AI beginners. While trying out various tools, we adopted "Google AI Studio", and we held a class that was almost like a rowdy festival, where we had them vibe-code the challenges that came up on the spot in 10 minutes, and we showered those who finished with praise (Ono-kun usually calls himself a PARTY BOY, and the class ended up strongly reflecting his character...)

It's not just about making things; let's properly make someone happy with vibe coding. With that in mind, and with the cooperation of CGO Dot Com, who are famous for their "Gal-style brainstorming," we invited five gals to the class. We held an unprecedented session where students answered the gals' requests on the spot at each table using vibe coding, creating app after app.

Gals are overflowing with creativity that isn't bound by common sense, and they also have a very strong love for the things they like. Meeting the demands of such people is a daunting task, but it turned out to be a huge success.

"An app to help you get over an ex," "an app that easily takes photos like old-school purikura," "an app that plans your daily outing schedule"... tools and apps beyond imagination were born one after another from the AI-beginner students. Many people also utilized Nano Banana's image editing, and I felt the speed at which technology born just a few months ago has already become the center of everyone's creativity.

I feel the goodness of AI lies in the self-efficacy of thinking, "I can probably do this myself." It is very effective as a thought booster, such as when someone who lacked confidence until now becomes able to do a little bit, or when you are stuck on an idea and can start thinking by DeepResearching something you are curious about. Humans are creatures that naturally get motivated once they start moving (though, on the other hand, it can be painful until you start). I feel that if self-efficacy increases due to AI and the "desire for evolution" to want to work harder spreads throughout Japan, it will become a more proactive society where everyone happily improves things on their own.

AI is not a person.

I feel that recently, many people tend to think of AI as an equal to humans, someone who empathizes with them, or conversely, view AI as an opposing axis that steals human jobs. To put it bluntly without fear of misunderstanding, AI doesn't seem to be that big of a deal. It's just that humans are under the illusion that it has a "personality."

I personally started touching AI around 2017, when "deep learning" began to become popular, just before this generative AI boom started. Even after the rise of large language models like the current ChatGPT and Gemini, I thought, "I'll give it a try!" and have taken on various challenges together with AI. (I have talked about my past creations in the Beyond Magazine interview article below, so I will omit it here.)

It is an interesting phenomenon that the more popular AI becomes, the more the essence of humanity itself is questioned as a backlash. I think that for creators, things like "that person's unique touch" will become more important. In other words, "Do I want to ask that person?" "Does it seem like only that person can do it?" "Is that person a good guy?"... Whether or not you can use AI is just a skill of that person, and it is a means. Of course, being able to master it is a good thing, but conversely, "not being able to use it" also becomes a valuable choice. Therefore, I don't think AI will cover the world; rather, I think it's more like AI is being introduced as a new organ for each individual, and it's up to that person to use it.

I think you can get hooked on the appeal of AI as a tool by reading various articles, but this time I ended up talking more about the mindset aspect. Finally, let me introduce some interesting words that Demis Hassabis of Google DeepMind said in WIRED.

If everything goes well, we should be entering an era overflowing with radical abundance, a golden age, so to speak. AGI will enable the resolution of what I call the root node problems of this world, such as curing serious diseases, living healthier and longer lives, and discovering new energy sources. (Omitted) Let me give you a very simple example. Securing water resources will be an extremely important issue in the future. But the solution already exists—seawater desalination. The problem is that it requires a huge amount of energy. However, if AI uses things like nuclear fusion to generate renewable, clean, and free energy, the water problem will be solved all at once. In other words, it will no longer be a zero-sum structure where another person's loss is your gain.

More than the risks in the evolution of AI, I want to look forward to the future that lies beyond it. There was a time when I was purely moved by the evolution of AI, such as when I prototyped various works through deep learning or when I was very excited when the video generation AI tool called Runway appeared. It is simply the expectation that "if I use this, I will surely be able to do things I couldn't do before." I want to continue to face AI with that feeling.

Rather than how to use it, what do you want to do?
I feel every day that this is the most important thing.