SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

Unraveling the Essence of Deviation Scores through Statistics: Misconceptions and Challenges in Entrance Exams

Recently, while developing an interest in AI and reviewing statistics, I found some notes I had made about "deviation scores" during my MBA studies. So, this time, I would like to write an article to help people understand deviation scores correctly from a statistical perspective.

I am not a statistics expert, and my knowledge is based on what I learned in English as a second language during my MBA, so I apologize if there are any errors or misleading expressions. Nevertheless, this content should be useful for those who want to know the essence of deviation scores.


Introduction: Are deviation scores a magic number for measuring true ability?

In Japan, it is often said that "a deviation score of 50 is average ability" and "a deviation score of 70 is genius," but from a statistical standpoint, one must be careful with such views. In particular, systems like Japanese university entrance exams, which are "one-shot" events, are susceptible to misunderstandings of deviation scores and the influence of chance, making it important to understand their limitations. In this article, I will unravel the basics of deviation scores, common misconceptions, and the problems with the entrance exam system from a statistical perspective.

1. What is a deviation score: Grasping the basics correctly

A deviation score (T-score) is an index that converts a test score into a relative position within a group. The calculation formula is as follows:

T = 50 + 10 × (X - μ) / σ

The symbols above mean the following:

  • X: Individual score

  • μ: Group average score

  • σ: Group standard deviation (magnitude of variation)

Let's think about a concrete example. If the average score is 60 and the standard deviation is 10, and you score 75:

T = 50 + 10 × (75 - 60) / 10 = 65

In this case, the deviation score is 65, which means you are "1.5 standard deviations above the average." Based on a normal distribution, a deviation score of 65 corresponds to the top 6.68% of the group. However, it is important to note that this does not indicate absolute ability, but

merely represents your relative position within that group.

It looks like this when graphed. A person with 75 points is positioned in the top 6.68% of that group.

2. Deviation scores and normal distribution: The foundation of statistics

Deviation scores are designed based on a normal distribution (a bell-shaped distribution that spreads symmetrically around the average).

This distribution is the shape that many types of data, such as natural phenomena and test scores, follow approximately. In deviation scores, scores are standardized by converting the average to 50 and the standard deviation to 10. This allows the following proportions to hold true:

  • Deviation score 40-60 (range of ±1σ): Approximately 68.27% of the total is included in this range

  • Deviation score 30-70 (range of ±2σ): Approximately 95.45% of the total is included in this range

  • Deviation score 20–80 (±3σ range): Approximately 99.73% of the total population falls within this range.

These are theoretical benchmarks that indicate where a deviation score stands within a group. However, since actual test results are influenced by the difficulty of the questions, the test-taking population, and the reliability of the exam, one should be cautious about equating deviation scores directly with "actual ability."

Benchmarks for the top percentage of each deviation score

3. Misconceptions regarding deviation scores

1. "A deviation score of 50 is average ability"

The truth: A deviation score of 50 merely indicates the average of the group. Regardless of the difficulty of the test, even if it is a difficult exam where the average score is 20, scoring 20 will result in a deviation score of 50.It does not represent absolute ability.

2. "A deviation score of 70 means you are a genius"

The truth: A deviation score of 70 corresponds to the top 2.28% of a group, but this naturally depends on the level of the population. For example, the meaning is significantly different when scoring 70 on a nationwide mock exam versus scoring 70 on a small-scale test. It is premature to judge someone as a "genius" based solely on a deviation score.

3. "A single high score proves your true ability"

The truth: Test results are determined by the sum of "true ability + random error." Even if you get a high score once, the possibility that it was the result of luck or good physical condition cannot be denied.According to the law of large numbers in statistics, repeating measurements multiple times brings the average value closer to one's true ability. For example, in an exam with a reliability of 0.9 (where the standard deviation error is small), the probability that a person with a deviation score of 60 will drop to 50 is about 15%. Relying on a single result is risky.

4. Problems with Japan's "one-shot" entrance exams

(1) The influence of chance is significant

In a one-time exam, errors such as physical condition, nervousness, and compatibility with the questions significantly affect the results. For example, in an exam with a reliability of 0.9, the probability that a person with a true ability of 60 will drop to a deviation score of 50 is about 15%. This could potentially lead to failure.

(2) Difficult to measure ability accurately

To accurately evaluate ability, it is ideal to take the average of multiple test results. With one attempt, random fluctuations are large, and by measuring multiple times, errors are averaged out, bringing the result closer to the true value. For example, if you take the exam five times, the influence of luck and temporary factors decreases, and your ability becomes clearer.

(3) Comparison with SAT and business school admissions: The benefits of multiple attempts

In the American SAT, you can take the exam multiple times a year and are allowed to submit your highest score. Also, in the business school admissions I experienced, there were no limits on the number of times you could take the TOEFL or IELTS, and while there were caps on the number of times for the GMAT and GRE, multiple attempts were possible, and the best score could be used. This creates a system that minimizes the influence of chance and appropriately reflects true ability. Perhaps there is room to consider such flexible systems in Japan as well.

Conclusion: Knowing the essence of deviation scores for fair evaluation

Deviation scores are merely statistical indicators that show relative position within a group. It is dangerous to overestimate or underestimate ability based on a single result. Only by taking an average through multiple evaluations can highly reliable ability be seen.

Furthermore, if the entrance exam system evolves to provide multiple opportunities rather than a one-shot attempt, fairness and accuracy will improve. It is important not to blindly trust deviation scores, but to understand their limitations and utilize them calmly.

Thank you for reading to the end.

いいなと思ったら応援しよう!