DeepResearch Horse Racing Prediction Report 14 | Results of Aoba Sho, Milers Cup, and Flora S
[Previous Results]
This time, I will report the predictions and results for the Aoba Sho, Milers Cup, and Flora Stakes (Flora S) held last week.
Starting this time, I have added Felo (ResearchAgent). This is a transition in the model selection for the deep research feature of "Felo" that I have been using until now. From now on, I will mainly verify the performance of ResearchAgent. (Note: The latest model selections such as o4mini, 3.7sonet, and R1 are available in Felo's paid plan.)
[Predictions and Results]
Aoba Sho (GII)

Yomiuri Milers Cup (GII)

Flora Stakes (GII)

[Summary Results]
Top 3 finish rate for the 4 recommended horses

Top 3 finish rate for the favorite (◎)

AI Model Horse Racing Prediction Comparison Summary
Objective
Evaluate the performance of horse racing predictions provided by multiple Large Language Models (LLMs).
Evaluation metrics: 6 indicators regarding prediction accuracy and the reliability of the favorite horse.
Create a comprehensive ranking by applying reliability adjustments that account for the number of trial races.
Data Overview
-
Evaluation Metrics (Accuracy Indicators):
① 3 out of 4 recommended horses finish in the top 3 (%)
② 2 or more out of 4 recommended horses finish in the top 3 (%)
③ 1 or more out of 4 recommended horses finish in the top 3 (%)
④ Favorite (◎) 1st place (%)
⑤ Favorite (◎) top 2 finish (%)
⑥ Favorite (◎) top 3 finish (%)
-
Models analyzed:
ChatGPT (o1/o1 pro/o3 mini/o3/o4 mini)
Gemini (2.0 flash/2.5 pro)
Perplexity
Felo (ResearchAgent/o3 mini/3.7 sonet/R1)
Genspark
Grok
-
Number of trial races:
ChatGPT(o1): n=16
ChatGPT(o1pro): n=13
ChatGPT(o3mini): n=6
ChatGPT(o3): n=7
ChatGPT(o1): n=16
Gemini(2.0flash): n=26
Gemini(2.5pro): n=10
Perplexity: n=29
Felo(o3mini): n=16
Felo(3.7sonet): n=17
Felo(R1): n=11
Genspark: n=12
Grok: n=27
Scoring Design
Weighting
Based on practical return value, the following weights are set for the six indicators.
3 out of 4 recommended horses = 4
2 out of 4 recommended horses = 3
1 out of 4 recommended horses = 1
Favorite ◎ 1st place = 4
Favorite ◎ within top 2 = 3
Favorite ◎ within top 3 = 2
Overall Score
The overall score is calculated by taking the weighted average of the percentage (0–100%) of each indicator using the weights above.
Confidence Adjustment
To suppress score variance due to differences in the number of trial races (N) for each model, confidence correction was performed using a Bayesian method.
Prior mean = 35 points, prior number of races = 10
Adjusted score = (Overall score × N + 35 × 10) / (N + 10)
Final Ranking (by Adjusted Score)

Characteristics of Top Models and Practical Usage
-
ChatGPT (o1pro)
Stable in both win and place bets with a 38% win rate for ◎ and a 69% top-3 rate for ◎.
Suitable for win/wide strategies that prioritize return rate. A trio bet using ◎ as the axis and spreading to 2–4 other horses is effective.
-
Genspark
Highest at 75% for hitting 2 or more out of 4 recommended horses.
-
ChatGPT (o3 mini)
High accuracy with over 80% hit rate for top 3 finishes across all categories, but the sample size is insufficient at N=6.
-
Felo (ResearchAgent)
The raw data shows a maximum total score of 52.9 points, but N=3. Do not over-rely on this.
For now, keep it as a reference indicator for comparison purposes.
-
ChatGPT (o4 mini)
With a 14% success rate for picking the top 3, its decisive power is above average, and the overall balance is good.
It is suitable as a base model for constructing trifecta multi-bets from a primary pick in races expected to have moderate upsets.
Future Challenges
-
Sample Expansion:
For models with a low number of trial races, add data up to at least 15 races and re-evaluate.
[Next Time]
In progress
#HorseRacing #HorseRacingPredictions #HorseRacingInformation #HorseRacingReport #AIHorseRacingPredictions #HorseRacingAnalysis #DeepResearch #HorseRacingVerification
