SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

DeepResearch Horse Racing Prediction Report 5 | Interim Summary

Last time, Perplexity achieved the remarkable feat of placing its recommended horses in the top three in all four races. Furthermore, it demonstrated astonishing accuracy, including correctly predicting a long-shot outsider.

Now, based on the predictions and results so far, I have briefly compiled which LLM has been the most successful.

Since the amount of data is still small, at around 10 races, I believe the results are within the range of variance, but I will evaluate the current rankings.

I wonder which model will shine in first place—I am looking forward to the results.


Summary Results

[Summary Conditions]

<Target Races>
・Ocean Stakes (GIII)
・Tulip Sho (GII)
・Nakayama Kinen (GII)
・Fillies' Revue (GII)
・Nakayama Himba Stakes (GIII)
・Yayoi Sho (GII)
・Kinko Sho (GII)
・Spring Stakes (GII)
・Falcon Stakes (GIII)
・Flower Cup (GIII)
・Aichi Hai (GIII)
・Hanshin Daishoten (GII)

<Models and Number of Verified Races>
・ChatGPT: 8 races
・Gemini: 9 races
・Grok: 12 races
・Perplexity: 12 races
・Felo (o3mini): 6 races
・Felo (3.7sonet): 8 races
・Felo (R1): 7 races


Comparison of the top-three placement rate for the 4 recommended horses (Favorite / Second Choice / Third Choice / Long Shot)

Percentage of races where at least one recommended horse finished in the top three
Felo (o3mini) achieved a perfect record of 100%, with at least one recommended horse finishing in the top three in every target race.
ChatGPT (88%) and Gemini (89%) also showed high accuracy, but overall, Felo (o3mini)'s consistency stands out.

Percentage of races where at least two recommended horses finished in the top three
Here, Perplexity leads with 58%. ChatGPT and Felo (o3mini) follow with 50% each, maintaining a certain level of accuracy.
On the other hand, Gemini remained at 44%, a slightly lower result compared to the other models.

Percentage of races where three recommended horses finished in the top three
The only model to clear this high-difficulty condition was ChatGPT (13%), with all other models resulting in 0%. This is a moment that offers a glimpse into ChatGPT's broad coverage.


Percentage of races where the recommended favorite (◎) finished in the top three

In terms of the percentage of races where the favorite (◎) finished in the top three, ChatGPT recorded the highest figure at 63%. This is followed by Felo (o3mini) and Felo (3.7sonet)
at 50%. Following them are Perplexity (42%), Gemini and Grok3 (both 33%), and Felo (R1) (29%).


Current Overall Rankings


[1st Place] ChatGPT

Recorded high standards in key items such as the multiple-horse hit rate and the success rate of the favorite (◎). In particular, the favorite (◎) finished in the top three in 63% of races, and it is the only example where three out of the four recommended horses were correct, making its overall high capability stand out.


[2nd Place] Felo (o3mini)

With a 100% rate of at least one of the four recommended horses hitting the mark, its consistency is second to none. Although it lacks a bit of impact, its solid accuracy and stable predictive power earned it second place.


[3rd Place] Perplexity

It leads with a 58% rate of at least two of the four recommended horses hitting the mark. However, because its success rate for the top pick (◎) is 42%, slightly inferior to the top two models, it remained in third place.


[4th Place] Gemini

While the hit rate for at least one of the four recommended horses is high, the success rate for the top pick (◎) is low at 33%, resulting in a lack of reliability compared to the top-tier models.


[5th Place and below] Grok3 / Felo(3.7sonet) / Felo(R1)

These models show generally sluggish figures across all metrics, and are inferior to other models, particularly in terms of the success rate of the top pick (◎) and the multiple-horse hit rate, so they are ranked lower this time.



Summary and Future Outlook

■ Summary

At this point in time, ChatGPT can be evaluated as boasting the most excellent overall accuracy in horse racing predictions.

Next is Felo(o3mini), which is highly rated for both consistency and hit rate, followed by Perplexity.

Gemini shows some results despite being somewhat unstable, while the other models (Grok3, Felo 3.7sonet, Felo R1) have produced difficult results at present.


■ Future Outlook

By continuing the verification process up to this point, the characteristics of each model have gradually become clear.

Due to the nature of research, it has also become apparent that the following constraints exist.

  • Advantage in gathering information just before the race
    As the race time approaches, information becomes available, and there is a tendency for predictions based on fresher data to be possible.

  • Selection of unpopular longshots is weak
    The way it marks extreme longshots is somewhat conservative, and it gives the impression that it is difficult to lead to a hit (unpopular longshots are those that few people are predicting). That said, it is able to pick up mid-range longshots reliably, and a certain level of accuracy can be expected.

  • Backtesting is difficult
    In verification targeting past races, it is difficult to completely eliminate result leakage (the risk of accessing results at the time of prediction), and there are issues with reliability as an experiment.

That said, these are not fatal flaws.

What is more important is
the possibility that model behavior and trends may change with future updates.

The fact that this could affect stability and prediction style is something that should be closely monitored.

Moving forward, I plan to prioritize improving prediction accuracy while also outputting reports on the characteristics and challenges that have emerged from these experiments.


With that said,

Although there are still challenges and refinements to be made, I intend to use the top 3 models that received high evaluations in this analysis to take a shot at the upcoming Takamatsunomiya Kinen before their behavior changes!

  • ChatGPT: No. 1 overall predictive power. Its broad adaptability and high hit rate are its strengths.

  • Felo (o3mini): Outstanding stability. Solid predictions with little variance.

  • Perplexity: High accuracy and sharp recommendations shine through. Particularly strong at picking multiple winning horses.

I will continue to pay close attention to how each model evolves and what kind of predictions they show in future races!

[Next Time]



#HorseRacing #HorseRacingPrediction #HorseRacingInformation #HorseRacingReport #AIHorseRacingPrediction #HorseRacingAnalysis #DeepResearch #HorseRacingVerification






いいなと思ったら応援しよう!