SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

I Tried JoyAI Image Edit—It Might Make Reference Images for Jewelry Photography a Little Easier

The other day, the name of a tool called "JoyAI Image Edit" kept popping up in the AI tool Discord I always check, so I got a little curious.

It's an image editing model released by JD's open-source team in April, and it seems you can edit images with instructions like "move this object here" or "rotate the camera to this angle."

When doing live streams for jewelry, I often lose time retaking reference images or thumbnails. Even if I think, "I want to see this composition with the stone placed a little more to the left," rearranging the actual items and resetting the lights takes 30 minutes each time.
I thought this might make things easier, so I tried it out a bit.

To begin with, what kind of tool is it?

Roughly speaking, it's a bit different from standard image generation AI; it seems to be the type that excels at "moving the contents of an existing image."
Looking at the official demo, you can do things like this, for example:

  • "Move the apple into the red box, then remove the red box"

  • "Rotate the chair to face forward"

  • "Rotate the camera 45 degrees to the right, changing the perspective without moving the scene"

That last camera operation really caught my attention as someone who does jewelry photography.
If I can create "images of the same stone seen from different angles" from a single photo, it might be useful for increasing the number of cuts for product introductions.

My impressions after actually trying it

I used the hosted version via a service called fal.ai, which costs about $0.10 per megapixel, or around 15 yen. It seems you need 24GB of VRAM to run it in a local environment, which was tough for my setup, so I tried the hosted version.

The first thing I tried was a photo I took of a rose quartz bracelet.
When I edited it with the instruction "shift the position of the stone slightly to the right and make the background a bit darker purple"—I'll be honest. The texture of the stone remained better than I expected.

When you do the same thing with standard image generation AI, the way light enters the stone often ends up looking completely different. This time, that part didn't break down easily.
However, it's not perfect. The reproduction of the fine cutting facets felt like it was being pulled by the quality of the original image. It's still difficult to extract sharp cuts from a blurry photo.

Situations where it could be useful for jewelry streaming and fortune-telling content
I've tried a few things and will note down the situations where I thought, "This is useful."

Creating variations for streaming thumbnails Create several images of the same jewelry from different angles and with different backgrounds. When you want to change the thumbnail for each stream, it seems like it will reduce the hassle of retaking photos.
Reference images for tarot and fortune-telling content For abstract things like "a scene with a crystal placed in a night forest, in the image of the Moon card," Midjourney is better suited. However, for "edits that utilize the actual object," such as "changing the lighting of this crystal photo to be a bit more mystical," JoyAI was better.
Fine-tuning reference images for outfit suggestions Last month, I had Midjourney create an image of a "deep purple dress, silver necklace" for a stream, but when I wanted to fine-tune it by saying "make the necklace design a bit thinner," I felt that JoyAI followed my intentions better.

I will also honestly write about the points that concerned me

Not just the good points, but also the points that concerned me.
One is that instructions in Japanese are not yet stable. I felt that the accuracy was higher when I wrote in English. There were several times when I wrote "ishi" (stone) and the recognition wavered, but when I rewrote it as "stone," it moved as intended.

Another point is that there isn't a clean web interface for consumers. You have to prepare the environment yourself or call it via a service like fal.ai. For people looking for the ease of "opening an app, choosing a photo, and entering instructions," it might still be a bit too much work.

Also, something I noticed while continuing to use it is that the results fluctuate slightly even with the same instructions. It often doesn't turn out as ideal on the first try, so I felt it's better to use it with the premise of running it a few times and choosing the best one.

Conclusion at this point

To start with the conclusion, I intend to continue testing this as an auxiliary tool for my streams.

It reasonably meets the needs of jewelry photography, such as "wanting to change the angle" or "wanting to swap the background." While it cannot completely replace actual photography, it seems useful as a rescue measure when you realize later that you "should have taken this shot as well."

However, I believe the most important thing in a job that conveys beautiful objects is, after all, facing the real thing. I want to use it with a sense of distance where I "think together" with the AI rather than "leaving it" to the AI—because I don't want to hand over the texture of rose quartz entirely to AI.

I am still in the verification process, but I am leaving this as a record of where things stand at the moment.
If anyone is using AI for jewelry streaming or fortune-telling content, I would be very happy if you could let me know how you incorporate it✨

いいなと思ったら応援しよう!