SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.

Reflecting on the Data Business Creation Contest

DIG18

The Data Business Creation Contest (commonly known as DIG), organized primarily by Keio University's SFC campus, is a competition where teams of high school, undergraduate, and graduate students compete with new business ideas utilizing data analysis. It has been held 18 times in total since 2014. For each event, companies known as "business partners" determine the theme, provide prize money, and supply materials such as data and analysis software that are otherwise difficult for students to obtain. Each participating team proposes unique businesses by analyzing the provided data or combining it with data they have collected independently.

In this year's 18th Contest (DIG18), the business partner, Kao Corporation, set the theme as "Life Care Innovation: A Rich Life Opened Up by Multidimensional Data." Instead of providing raw data, they offered API access rights to a generative model called Virtual Human Generative Model (VITA NAVI). I have been involved in the development of VITA NAVI since the planning stage, and this time, I participated as a business partner for DIG18, providing support for the use of VITA NAVI and also serving as a judge.

DIG18 was announced in April 2024. An online study session was held on April 12th, where I explained VITA NAVI, which was provided in place of data. By the preliminary deadline at the end of July, 97 teams had registered, and 65 of them actually submitted preliminary proposals. Subsequently, 11 teams that passed the preliminary round gave presentations at the final round on September 7th, and six awards, including the Grand Prize, were selected. The number of registered teams and the number of teams that submitted proposals were the second and first highest in the 18-year history of the contest, respectively, which shows the high level of interest. I would like to thank everyone who showed interest.

What is VITA NAVI?

VITA NAVI is a generative model that approximates the joint distribution of various attributes related to the human body, such as age, gender, weight, and blood pressure. These attributes include demographic information like age and gender; information obtained from health checkups such as weight and blood tests; information obtained from health insurance receipts such as the annual number of visits to a doctor for a certain disease; measured values such as body composition, hair, and skin; results of cognitive function tests; sequence information obtained from gut microbiota, oral microbiota, sebum RNA, etc.; and information obtained from questionnaires such as dietary habits, stress, sleep, bathing habits, and women-specific concerns. In total, there are over 2,000 attributes.

Attributes of VITA NAVI

The VITA NAVI API (Application Programming Interface) allows users to submit conditional queries to this model, such as "Find the distribution of systolic blood pressure for a 66-year-old male." The unique feature of VITA NAVI, which has no parallel, is that any combination of these approximately 2,000 attributes can be used as input conditions. This allows for the calculation of indicators that could be called "blood age," such as "What is the average age of a man with these blood test results?" and the exploration of previously little-known associations (correlations), such as "What attributes are strongly related to concerns about body odor?"

Although VITA NAVI is a highly versatile generative model, its versatility has been a drawback, and since the start of commercial service in February 2023, its usage has not grown much, which has been a source of concern. I focused on DIG because I hoped that young people with fresh perspectives would propose ways to use VITA NAVI that we had never even thought of.

In this DIG18, 15 judges, including the head judge, Professor Murai, were in charge of the judging. On the day of the final round, we listened to the presentations of the 11 teams, held Q&A sessions, then tallied the judging results, and further discussed among the judges to decide the winning teams for each award. Below, we will look at the patterns of how the 11 teams that passed the preliminary round incorporated VITA NAVI into their proposals.

How to use the generative model

Use as background information

The most common use was as background information to guide business ideas.

For example, the Kao Award-winning Team Saba no Nikomi Teishoku from Senshu University used VITA NAVI to discover that a woman's weekly bathing frequency is linked to several indicators related to happiness, such as self-efficacy (see figure, excerpted from Team Saba no Nikomi Teishoku's final presentation materials). Using this result as background knowledge, they connected it to a business idea for a bathing management app for women called BathNavi. The relationship between bathing frequency and happiness was something we had never thought of, so it was an eye-opener.

Relationship between weekly bathing frequency and happiness scale

Similarly, the Grand Prize-winning Team Shimu-Tech from Soka University proposed "Ririfuru," an app that comprehensively manages women's menstruation, because VITA NAVI showed a relationship between crying and menorrhagia, and between dozing off and menstrual irregularity. Also, the Future Creation Award-winning Team MY Duo from Hiroo Gakuen confirmed on VITA NAVI that women's menstruation is related to mental stress, etc., and connected that to the idea of a stuffed animal called "Pokadoll" (this stuffed animal will be described later).

The above three teams were in Femtech, but there were also many teams that wanted to do something about mental pain such as stress and depression. Team Ebimayo no Knapsack from Senshu University proposed a business to reduce stress associated with digitalization based on the relationship between occupational stress and sleep found in VITA NAVI, and Team Shoronpo no Shoulder Bag from the same Senshu University focused on stress hormones and, getting a hint from the attributes linked to them, proposed a smart lock that opens with voice recognition to reduce feelings of loneliness. Also, Team Kimchi no Seishun from Osaka Prefectural Suito International High School proposed a stress reduction measure using VR, and they used the distribution obtained from VITA NAVI to narrow down their target audience.

This use of VITA NAVI as background knowledge was mainly done by manually entering several inputs into the VITA NAVI sample app (a simplified version that can be called from a web app without programming) and plotting the results, not by automating it using the API. VITA NAVI is a business that sells APIs, so I originally wanted them to use the API, but when I think about it, in the current situation where data between diverse attributes such as bathing, dietary habits, women's concerns, and stress cannot be easily obtained, being able to easily answer simple questions like "Is there a relationship between XX and YY?" is a major selling point of VITA NAVI.

Estimation of unknown attribute values

Estimating unknown attribute values through an API is a basic function of VITA NAVI. The Excellence Award-winning Team Pork Stars from Keio University/Graduate School obtained an estimated value from VITA NAVI of "how many times a year this person is likely to see a doctor for this disease" from receipt data, calculated the estimated number of days off for each employee in a company based on that, and further proposed a business to support the company's health management. Naturally, they have to calculate the estimated value for all employees every year, so it is a business that incorporates VITA NAVI into the system via an API. They also made a solid business case, and it was a proposal that seemed ready to be commercialized immediately.

In terms of estimating unknown attribute values, the Team MJAN2024 from Kwansei Gakuin University, which won the Jury's Special Award, used a very interesting approach called counterfactual inference. When a disaster like the 2024 Noto Peninsula Earthquake occurs, health management in evacuation centers is a critical issue. This team used VITA NAVI for hypothetical inference (counterfactual inference), asking, "If blood pressure rises and nutritional deficiencies occur due to evacuation life, what kind of disease risk would this person have?" To begin with, VITA NAVI's training data does not include data from the special conditions of an evacuation center, and there is the question of whether the values of the manipulated variables—blood pressure rise and nutritional deficiency—are appropriate. However, it was promising in that it suggested such indicators could potentially be useful as a guideline.

Optimization

Optimization is an advanced use of the API. Team Demetona from Hiroo Gakuen (though it was a one-person team) focused on the relationship between fatigue and nutrients, and used the black-box optimization tool Optuna to solve the question, "What combination of 117 nutrient intake amounts would result in the lowest estimated fatigue level?"

Exploration and Synthetic Data Generation

Since VITA NAVI is a generative model, it can create synthetic data. Team Wellness Informatics from Musashino University explored attributes related to depression while varying the depression scale. This is also something that could not be done without using the API. Furthermore, this team went a step further by generating synthetic data containing the 65 attributes found, and used that data to train machine learning models such as Random Forest, demonstrating that they could estimate the depression scale with high accuracy. Obtaining synthetic data from a generative model is a very interesting way to use it, but it is important to note that this machine learning model is not estimating the original depression scale, but merely estimating what the output from VITA NAVI would be. Generative models, not just VITA NAVI, are approximations of the real world and always contain errors. When you create a secondary model that mimics a generative model, no matter how accurately it mimics the output of the generative model, it is not estimating real-world values. I wish they had used the VITA NAVI estimation API directly for estimating the depression scale.

Using Only Item Definitions

Finally, I would like to introduce an interesting proposal that did not use VITA NAVI as a statistical model. It is a matchmaking app proposed by Team Sensitive Skin from the Tokyo University of Science graduate school. The idea is to extract 117 attributes related to lifestyle, life style, and personality from VITA NAVI's attribute definitions, and suggest that the higher the similarity, the better the compatibility. When I planned VITA NAVI, I had a hidden ambition that "when displaying or exchanging human conditions, I want people to use VITA NAVI's attribute definitions as a common format." I think this proposal showed the potential for that kind of usage.

Active Female Students and Femtech

With the growing expectations for data scientists, data science faculties are being established at many universities, but the percentage of female students has not increased as much as expected. For example, it is said that at the Faculty of Data Science at Shiga University, the first of its kind, the percentage of female students among the approximately 400 students enrolled as of June 2024 remains at about 20 percent (Source: NHK Shiga News Web, June 2024). On the other hand, in this Data Business Creation Contest, we see female students playing an active role every time. This time, too, the active participation of female students was notable, regardless of whether they were university or high school teams.

The Grand Prize-winning Team Shimtech captured the hearts of the judges with a presentation that began with the declaration, "We have a dream!" But above all, the driving force behind their Grand Prize win was their Q&A session. When faced with a tough question from Chief Judge Murai, "There are existing services widely used for femtech, so how is your proposal different, and why can you win?" it was brilliant how they responded by bringing out backup slides as if they had been waiting for it.

The Future Creation Award-winning Team My Duo was also a team that earned points in the Q&A session. Their proposal was to sell a stuffed animal called "Poka-doll" to alleviate the pain of menstruation, but I think the underlying intention did not come across well in the presentation. During the Q&A, they were able to bring out their original intention: "By carrying the stuffed animal so that it is visible from the outside, we will create a society where it is normal for people to let those around them know they are on their period (like a maternity mark)," which led to a high evaluation.

Most of the judges, including myself, were from the Showa generation, but I felt that we have entered an era where we can openly discuss women's menstruation.

Conclusion

I have shared my thoughts on the 11 teams that made it to the finals, but of course, there were various good proposals from the teams that did not pass the preliminary round. Among them, there were some that I thought were losing out due to how the slides were made, despite having good proposals. Among the proposals that did not pass the preliminary round, one that left a particularly strong impression on me was from a high school student who had a very personal concern and, upon learning that the concern was included in the VITA NAVI attributes, used VITA NAVI to investigate its relationship with health and happiness. Because they encountered VITA NAVI just before the preliminary deadline, they were unable to connect it to a powerful business proposal and did not pass the preliminary round, but I felt that they were able to answer the need that VITA NAVI originally aimed for: "niche but important to the individual."

VITA NAVI is still rough around the edges, but it was a Data Business Creation Contest that felt promising toward the goal of "democratizing life-care innovation." I would like to express my deep gratitude to Professor Murai, Professor Uehara, and everyone at the DIG secretariat for accepting us as business partners; to the judges who read and evaluated a massive amount of PowerPoint materials; to the student concierges who kindly and carefully supported each team in using VITA NAVI, which is difficult in terms of concept; to the professional concierges who lent their strength to refining the teams that passed the preliminary round; and above all, to the participating teams who proposed many ideas. Thank you very much.


いいなと思ったら応援しよう!