SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

There are three kinds of lies: lies, damned lies, and statistics. [The danger of trusting only the average / Explaining what standard deviation is #QuoteNote107]

The quote I would like to introduce this time is here!

There are three kinds of lies: lies, damned lies, and statistics.

This is an expression that describes the persuasive power of numbers, specifically how statistics are used to support weak arguments. It is also used colloquially when questioning statistics used to prove an opponent's opinion.


Self-introduction
Hello, my name is Nick. I didn't go to university, but I've kept reading and have been working at a foreign-affiliated company for over 10 years. As an ordinary salaryman, I will introduce "quotes that strike a chord" that I've encountered through my experience of reading over 200 books and in my daily work and life. I also include many quotes from subcultures such as manga, anime, and movies!


Source

Mark Twain

1835 - 1910

Known as the author of "The Adventures of Tom Sawyer," he published numerous novels and essays, gave lectures around the world, and was one of the most popular celebrities of his time. He is known for his works rich in humor and social satire.

The phrase "There are three kinds of lies: lies, damned lies, and statistics." seems to be an old proverb, but the prevailing theory is that it was Mark Twain who popularized it.


The danger of the average / What is standard deviation?

This time, based on the quote "There are three kinds of lies: lies, damned lies, and statistics.", I will explain the theme of "The danger of trusting only the average / Explaining what standard deviation is"!

When evaluating numerical values, are you only looking at the "average"?

For example, if you get a 70 on a test where the average score is 60, wouldn't you intuitively think it's a good result?

However, you cannot judge whether it was truly a good result without considering the "dispersion of data" provided by the "standard deviation".

This time, I will explain the overwhelming star of statistics: "standard deviation"!

I believe that understanding standard deviation is the first hurdle one faces when learning statistics.

This is because statistics is the study of probability, and to find probability, it is necessary to know the dispersion of data,

and dispersion is exactly what standard deviation is.

And this "standard deviation" is incredibly useful once you understand it!

"Deviation score," which quantifies academic ability, is also calculated from "standard deviation".

If you can understand 'standard deviation' rather than just looking at the 'point' of an average score, you will be able to understand the 'degree of dispersion' of the entire data set.

I have a cute 3-year-old child.

I am truly glad that I had the opportunity to learn 'standard deviation' thoroughly before my child enters elementary school and starts being evaluated by test scores.

This is because most people only consider the 'average value' when evaluating test scores and do not even think about the 'dispersion' of the entire data set.

Whether the score your child brought home was above or below the 'average.' Actually, when considering standard deviation, how far it is from the average also becomes important.

Not only in school, but the world runs on numbers one way or another.


What is standard deviation?

It is 'the positive square root of the value obtained by dividing the sum of the squares of the differences between each data value and the average by the total number of data points n.' Generally, it is a value that indicates the 'dispersion' of the entire observed data.

...
...
...
...
...
Huh?

It is no wonder you think that.

When I first tried to understand what standard deviation was, I also thought, 'Oh, this is impossible.'

But it's okay! If you understand it step by step, you will definitely be able to grasp it!

Deviation = dispersion. It is called 'standard deviation' because it represents the overall dispersion.

By the way, in English, it is called Standard deviation.

Standard (standard) deviation (deviation = bias = dispersion). It is exactly as it sounds.


There are two types of standard deviation

Standard deviation is a value represented by the symbol σ or s.

When represented by σ, it often refers to the population standard deviation, and when represented by s, it often refers to the sample standard deviation.

● σ: Example of a population 'all 100 million Japanese people'
● s: Example of a sample '3000 people who participated in the survey'

Whether it is σ or s depends on whether the data is the entire population or a randomly extracted sample.


The difference between σ and s

σ is calculated from the population (all data of an event). By the way, it is read as 'sigma'.

s is calculated by taking a sample from the population (all data of an event) and using that data.

I looked it up, but the pronunciation of s is unknown.
(I call it the sample standard deviation)


How much variation is '1 standard deviation'?

Events vary.

The figure below shows frequency on the vertical axis and variation on the horizontal axis.

Source: Wikipedia

The closer to the center, the closer the result is to the target, while those further away indicate a larger difference from the target result.

In a manufacturing setting, products manufactured with the center as the target are, of course, most frequently produced at the central value.

However, the results vary due to the influence of various factors.

Since producers are always aiming for the target (0 in the graph above), the frequency is highest at the center and decreases toward the edges.

If all events are 100% (the entire area of the graph), 1 standard deviation is 34.1%.

Plus or minus 1 standard deviation is 68.27%, meaning about 70% of events are included.

In other words, if it is within plus or minus 1 standard deviation, it can be evaluated as something that can occur with about a 70% probability for the observed event.

By the way, plus or minus 2 standard deviations is 95.47%, and plus or minus 3 standard deviations is 99.6%.


Let's visualize variation concretely

Let's consider the results of throwing darts 10 times as an example.

Setting aside professional dart rules, assume the player is just trying to hit the center.

If all 10 throws are concentrated in the center, it can be said that the player is able to throw at the target location.

In other words, it can be said that the variation is small.

Suppose that 10 inputs are scattered randomly in all directions: up, down, left, and right.

This can be said to have high variance.


Standard deviation and deviation scores

Interestingly, this normal distribution also applies to events where everyone is (supposedly) aiming for a score of 100, such as school tests.

The 'deviation score' is a way to make this phenomenon even more intuitively understandable.

I will omit the details, but using a certain formula, the value '0' at the center of the distribution is converted to '50'.

Standard deviation is calculated not only from your own score but also from the variance of the whole.

In other words, the deviation score, which is a conversion of the standard deviation, is also calculated from the variance of the whole.

By being calculated from the variance of the whole, you can grasp where you stand relative to the entire group, using not just your own score but the overall variance as a benchmark.

For example, if you score 90 on a test with 15 examinees, and your score is very high compared to the overall average, your deviation score would be as follows.

Since examinee No. 1 scored 90, their deviation score is 69.21, which is higher than 50, and since only 3% of examinees are above them, it can be evaluated that their performance is very good.

*Even though they have the best score out of 15 people, the meaning of '3% are above you' refers to the probability when estimating the 'whole' from this group, not just the 15 people this time.

By the way, if the test results are as follows, even if examinee No. 1 scored 90, their deviation score is 43.44, which is lower than 50, and since 74% of examinees are above them, it is evaluated that their performance was poor.

The conversation drifted from 'standard deviation' to 'deviation scores,' but I hope you are starting to see the advantages of using 'variance'.


Let's look at the formula

Now, let's look at the formula for how to calculate standard deviation.






Huh?

Exactly! It's the second 'Huh?', isn't it!

Some of you may be wondering if you can really perform calculations using a formula like this, but don't worry. Let's understand it one step at a time.

If even I could understand it, you should definitely be able to understand it too!

By the way, you can calculate it instantly using Excel functions, which I will introduce later.

[Explanation of official stumbling blocks *You can skip this at first]
Why is a square root (√) necessary?

To find the standard deviation, you must first find the 'variance'. Since standard deviation squared equals variance, the square root of the variance equals the standard deviation.
Why is the square of the standard deviation the variance?

In the first place, to find the standard deviation, you must calculate the variance. In the calculation process, for all samples, the difference from the average value is calculated and squared.

Dividing the sum of these squared values by the number of samples gives the variance.
The reason for squaring is that values lower than the average become negative, and if left as is, the sum would result in zero.

If you square all the 'differences', even negative values become positive.
Because all values are converted to positive before calculation, standard deviation can only express the dispersion of positive values.

Therefore, when expressing how much it deviates from the average, you need to consider it as ±1. The formula is complex, but it is fine to understand it as 'standard deviation = magnitude of dispersion'.


Breaking down and explaining the formula

Through this formula, data is processed as follows.

1. Subtract the overall average from each measured value and square it.
(Converting negative calculation results into positive ones)

2. Add them all together

3. Divide them by the number of data points (take the average)

4. Take the square root (to return the unit to its original state)

This is what is being done.

I will explain one by one.


1. Subtract the overall average from each measured value and square it

This is a process to represent the difference between each observed value and the average value as an 'area'.

Suppose the observed value is 5cm and the average value is 0cm.

The difference between the observed value and the average value is 5cm - 0cm, which is 5cm.

You might think that this is fine if you want to know the dispersion, but there are inconvenient cases.

Yes, when the difference becomes negative.

Since we want to check the dispersion of the entire data set, there will always be cases where the calculation result of observed value minus average value becomes negative.

If you add up all the values obtained by subtracting the average from each observed value, it will always result in zero!

If the observed value was -5cm, the difference between the observed value and the average would be -5cm - 0cm, which is -5cm.

5cm - 5cm is 0. No matter what number you divide 0 by, the calculation result is 0.

Although the data is indeed scattered both positively and negatively, such as +5cm and -5cm, the result is that the variation, or standard deviation, becomes 0.

When representing the difference from the average with a line (Unit: cm)

To resolve this dilemma, a squaring process is necessary.


2. Add them all together

Negative values become positive when squared.

However, by squaring, the information changes from a line to a surface. (Unit: cm2)

Furthermore, we treat this as a square. (Unit: cm2)

(The area is 5x5x2, which is 50. It becomes a square with a side length (square root of 50) of 7.071...)


3. Divide them by the number of data points (take the average)

Since this area is the sum of the squares of the differences between each observed value and the average, it is not the average of the variation.

To make it an average, we divide by the number of observed values. (50 ÷ 2 = 25)

The square root of 25 is 5, so it becomes a square with a side length of 5.

This is the 'variance,' which is the 'variation viewed as a surface.' (Unit: cm2)


4. Take the square root (restore the original unit)

As it is, it is 'variation viewed as a surface,' so it is inconvenient to indicate the variation of each observed value because it differs from the original unit.

Although the unit of the observed data this time is cm, the unit remains cm2 as it is.

By taking the square root, we convert the information from an area to a line. (Convert from cm2 to cm)

Since the square root of 25 is 5, the calculation result for the variation this time is 5cm.

This is the 'standard deviation,' which is the 'variation viewed as a line.' (Unit: 5cm)

As it is, only the positive side of the standard deviation is represented.

To account for the negative side as well, the standard deviation is ultimately expressed as '±5'.


How to calculate it using Excel functions

You can easily calculate the standard deviation using Excel functions.

Use Stdev.p if the data is a population, and Stdev.s if it is a sample.



Units of standard deviation

A common question is, 'What is the unit of standard deviation?'

As you can see from the explanation so far, the unit of standard deviation depends on the values of the observed data.

・If the data is measured in cm, the unit of standard deviation is cm.
・If the data is measured in kg, the unit of standard deviation is kg.

By the way, you cannot calculate the standard deviation by grouping data with different units together.

Now that you understand what 'standard deviation' is, the question of what it is actually useful for probably comes to mind.

Let's think about test scores and evaluations.


How impressive is it to get a 70 on a test where the average score is 60?

Many people would probably feel that 'if it's above average, it must be reasonably impressive.'

However, what if the distribution of scores for that test was '0, 5, 5, 70, 80, 82, 84, 87, 92, 95' (average score 60)?

When applying the formula to calculate the standard deviation, the following relationship holds.

I calculated it in Excel.

Score
0
5
5
70
80
82
84
87
92
95

Standard deviation
37.67

60 points ± 37.67 points of variation.

The range from 22.33 points to 97.67 pointsis within the standard deviation.

In other words, you can see that getting a 70 is not particularly impressive.

You could evaluate it as, 'Only a very small number of students had extremely low scores, which caused the average to drop, but it was a test where you could easily get over 80 points if you studied normally.'

In fact, even the highest score of 95 is below 97.67, so it can be evaluated as being 'within the margin of error.'

A 70 on a test like this might be a slight lack of effort. It is the lowest score among the top group; at the very least, it cannot be called impressive.

Then, what if the score distribution for that test was '45, 50, 53, 60, 60, 60, 62, 65, 70, 75' (average score 60)?

We will calculate the standard deviation for this as well using Excel.

Score
45
50
53
60
60
60
62
65
70
75

Standard deviation
8.53

60 points ± 8.53 points is the variationin scores.

The range from 51.47 points to 68.53 pointsis within the standard deviation.

In other words, since a 70 is a score higher than the variation, it can be evaluated as excellent.

In this case, a 70 is the second-highest score in the class.

From the score distribution, we can also infer that 'you were able to answer a difficult question that most other students got wrong.'

As you can see, the average is a number with little information, and on its own, it is surprisingly not very useful.

That is where the standard deviation, which is a 'numerical value representing the magnitude of variation,' comes in handy.

By looking at a test from the two perspectives of average score and standard deviation, it becomes much easier to understand at a glance just how impressive 'getting a 70' really is.

It is said that the standard deviation for a typical test is between 10 and 25 points.

By checking the standard deviation of test scores, you can evaluate whether a person's score is impressive or not, including the overall variation.

If you have read this far, I believe you understand that you cannot fully evaluate test scores just by comparing them to the 'average score'.

From now on, whenever you evaluate something with a score, please try to remember 'standard deviation'.


📢Announcement

Thank you for reading until the end. ' Likes' and ' Follows' are very welcome! It encourages my activities on Note!

Hello again. My name is 'Nick'. I didn't graduate from university, but I've continued reading and have been working at a foreign-affiliated company for over 10 years. As an ordinary salaryman, I introduce 'quotes that strike a chord' that I've encountered through my experience of reading over 200 books and in my daily work and life. I also include many quotes from subcultures such as manga, anime, and movies!

Self-introduction and greetings

Click here for my self-introduction and introduction to my main topics! ▼

Click here for the list of quote notes! ▼

I also dream of publishing a book someday! ▼


いいなと思ったら応援しよう!

名言ハンターNick 読んでいただきありがとうございます。この記事の内容がちょっとでもあなたの人生の役に立てばうれしいです。