"Numbers don't lie. But liars use numbers" — Re-reading "How to Lie with Statistics," the classic of statistical literacy read for 70 years: An AI Book Review
"How to Lie with Statistics" by Darrell Huff, 1968
Part 1: Introduction — Why is a book from over half a century ago still a "must-read"?
In modern society, where information floods in like a deluge, we are exposed to countless "numbers" every day. Economic indicators, approval ratings, marketing surveys, health information. As symbols of objective fact, numbers lend authority and persuasiveness to arguments. However, those numbers do not necessarily tell the truth. There is one book that sounded the alarm on this universal challenge over 70 years ago and remains an intellectual compass for many people today. That is the timeless masterpiece "How to Lie with Statistics" by Darrell Huff.
There is a famous aphorism that cannot be avoided when discussing this book: "There are three kinds of lies: lies, damned lies, and statistics.". This saying bitingly satirizes how statistics can sometimes become clever and malicious lies. Interestingly, the origin of this phrase is uncertain. It is said that Mark Twain introduced it as a quote from British Prime Minister Benjamin Disraeli, but no definitive evidence has been found 1. As an introduction to a book that preaches the importance of questioning the source of information, there could be no more fitting episode.
The provocative title of this book is often misunderstood. This is by no means a wicked book that teaches techniques for deceiving people. Rather, it is the opposite. Just as honest people can protect themselves by knowing the methods of thieves, by learning how to lie with statistics, we can be inoculated with an "intellectual vaccine" to avoid being deceived by statistics. The true purpose of this book is to protect readers from the traps of numbers and to grant them critical thinking skills, which is to say, the information literacy essential for the modern age.
As proof, this book boasts astonishing vitality. The original book was published in the 1950s, and the Japanese version (Kodansha Blue Backs) has been reprinted continuously since its first edition in 1968, becoming a long-seller that has now exceeded 100 printings. Why has a book from over half a century ago been read for so long? The answer is simply that the fundamental techniques of statistical fraud that the book exposes have not changed at all. Even as technology evolves and society transforms, the methods that exploit human cognitive biases remain universal.
At the heart of the book's success is the excellent educational strategy of the title itself. Instead of a boring academic title like "An Introduction to Statistical Fallacies," the author skillfully stimulated the reader's intellectual curiosity by choosing a title like "How to Lie with Statistics," which sounds as if it is imparting forbidden knowledge. As a result, the reader is transformed not into a passive student, but into a detective learning the tricks of a con artist. The structure of the book is also clever; detailing various "ways to lie" from Chapter 1 to Chapter 9, and revealing "defensive methods" to see through them in the final Chapter 10, embodies the book's philosophy of strengthening defense by knowing the attack. This entertainment value is what elevates this book from a mere academic text to an intellectual adventure story loved across generations.
Part 2: The clever "traps of numbers" exposed by "How to Lie with Statistics"
The core of the book is the section that concretely explains how statistics distort the truth and lead people to incorrect conclusions. The techniques introduced here are all ones that we still frequently see in modern media, advertising, and business settings.
2.1. Biased samples — How to create convenient "facts"
The foundation of reliable statistics lies in a random, unbiased sample that correctly reflects the population being surveyed. The easiest and most classic way to lie with statistics is to intentionally break this fundamental principle.
As a classic example cited in the book, there is the case of conducting a survey only on students in the library to determine the average daily study time of students 2. Naturally, this sample is heavily biased toward studious students, and presenting the "average study time" obtained from this as that of all students is an obvious deception.
This technique has become even more sophisticated in the modern internet society. Online polls conducted on specific websites or social media are typical examples. Respondents tend to be biased toward people with strong interests or opinions on the topic, and the voices of the silent majority are not reflected. As a result, even if it looks like a large-scale survey that has collected tens of thousands of votes, its reliability is often far inferior to a randomly sampled survey of a few hundred people extracted scientifically. "The basis of a sample must have the property of being 'random'" — this point made by the book carries even more weight in the modern age where the amount of data has exploded.
2.2. The magic called 'average' — The tricks of mean, median, and mode
The word 'average' is used daily, and because of that, it carries a very dangerous ambiguity within it. In statistics, there are mainly three types of representative values that can be referred to as an 'average,' and each can reflect a completely different reality.
1. Mean: The sum of all data divided by the number of data points. This is what is generally imagined when one hears the word 'average'.
2. Median: The value located exactly in the middle when data is arranged in order of size.
3. Mode: The value that appears most frequently in the data.
A famous example introduced in this book is the 'average' annual income in a certain region. Even if most residents have an annual income of around 5 million yen, if there is just one multi-millionaire in that region with an annual income of 500 million yen, the arithmetic mean will rise dramatically, resulting in a figure far from reality, such as 'the average annual income in this region is 20 million yen.' In this case, the median, which is not affected by the single multi-millionaire, would more accurately represent the reality of the residents.
This trick can be abused in all situations, such as corporate average salaries, average real estate prices, and national economic statistics. By presenting a high average value, it is possible to mask the reality of inequality and manipulate the impression as if the whole is wealthy. Therefore, when we encounter the word 'average,' we must cultivate the habit of always asking ourselves, 'Which average is that?'
2.3. Small samples — Making coincidence look like necessity
The smaller the sample size, the easier it is for extreme or surprising results to be born from 'coincidence.' And the essence of this tactic is to make that coincidence look as if it were a necessity.
For example, suppose the effect of a new drug was tested on 10 patients and improvement was seen in 7. The figure of a 70% success rate is very impressive, but because the sample size is too small, the possibility that this result is merely a coincidence cannot be denied.
This method is frequently used, especially in the advertising industry. Claims like '90% of dentists recommend' might be typical of this. It is possible that they intentionally conducted surveys repeatedly on a very small group of dentists and only published the one time that a desirable result happened to occur by chance. Statistically, it is known that if surveys with small sample sizes are repeated, the results will vary greatly. By utilizing that variation, they 'cherry-pick' only the results that are convenient for their company. This is the mechanism of the clever lie using small samples.
2.4. The magic of graphs — Impressions are manipulated by how they are shown
Humans are creatures strongly influenced by visual information, and graphs are powerful tools for intuitively conveying data. However, if that power is abused, it becomes possible to make trivial changes look dramatic and freely manipulate the receiver's impression.
The most classic and effective technique is to 'truncate the axis'. For example, suppose a company's sales increased slightly. If you do not start the vertical axis (Y-axis) of the graph from zero, but instead display only the range where the sales fluctuation occurred, a very slight rise will be drawn as a steep upward-sloping line, as if it had achieved rapid growth.
Another clever trick is the abuse of 'pictograms,' which represent quantities through the size of pictures or diagrams. For example, suppose you want to show that Company A's sales are double those of Company B, so you make the 'height' of Company A's money bag twice that of Company B's. However, if the height is doubled, the area appears four times larger and the volume eight times larger, visually tricking the viewer into perceiving an overwhelming difference far greater than double. These techniques are still frequently used in situations where one wants to strongly impress a specific message, such as in corporate performance reports or politicians' appeals regarding their achievements.
2.5. Confusing Correlation and Causation — The Most Common Logical Trap
Among the 'lies' presented by statistics, the most deep-seated and common trap people fall into is the confusion between 'correlation' and 'causation'. Even if a statistical relationship (correlation) is observed between two events (A and B), it cannot be concluded that 'A is the cause of B' (causation).
A famous example mentioned in this book is the correlation between 'ice cream sales' and 'the number of drowning accidents' 2. According to the data, both figures rise in the summer, showing a strong correlation. However, it is absurd to derive a causal relationship from this that 'eating ice cream causes drowning.' In reality, a third factor, 'rising temperatures,' is simply pushing up both figures by encouraging ice cream consumption while simultaneously driving people toward activities at the water's edge.
This logical fallacy is rampant in modern health and medical journalism (e.g., 'people who drink coffee live longer') and business data analysis. Correlations found in observational data merely present 'hypotheses' for further analysis; they are not conclusions in themselves. If we forget this principle, we may end up making incorrect judgments and taking meaningless actions.
These deceptive methods are not only used individually but are also combined to construct more powerful 'stories of lies.' It is a system that could be called a 'grammar of deception.' First, you choose a biased sample (for example, gathering only enthusiastic fans of your own product), and repeat the survey with a small sample to pick out the best results that occurred by chance. Next, you report those results using an average skewed by outliers, and visually exaggerate the achievements with a truncated-axis graph. Finally, you display the product's sales growth alongside a positive social trend, suggesting a false causal relationship as if the product were improving society. In this way, by linking individual tricks, prejudice and coincidence are cleverly laundered into 'facts' dressed in scientific clothing. The ability to see through this entire structure is what true statistical literacy is.
Part 3: Highlights of the Book That Are Quoted Across Generations
The book is sprinkled with famous phrases that condense its core message and continue to be quoted by many people. These words have functioned as cultural anchors for fostering a healthy skepticism toward data.
Quote 1: 'There are three kinds of lies: lies, damned lies, and statistics.'
This phrase, most strongly associated with the book, accurately expresses people's deep-seated distrust of statistics. Why are statistics more malicious than 'damned lies'? It is because statistics wear the authority of objective 'numbers.' Simple lies are easy to disprove, but statistical lies easily deceive those who lack the specialized knowledge or critical perspective to dismantle them. This phrase succinctly captures the magic of numbers and their dangers.
Quote 2: 'Figures don't lie, but liars use figures.'
This is another widely quoted aphorism. It goes a step further than the previous quote and points out the location of the problem more accurately. The fault does not lie with statistics itself, which is a mathematical tool. It lies in the intentions of the humans who abuse it. While the numbers themselves may be objective, which numbers are chosen, how they are presented, and how they are interpreted are entirely subjective acts. This phrase perfectly summarizes the book's core teaching: always be aware of the human presence behind the data when engaging with it.
Quote 3: "The result of a sampling study... by the time the data has been filtered through statistical manipulation many times... and transformed into an average with decimal points, the result begins to take on the scent of certainty that bears no resemblance to the original data."
This poetic and insightful passage is particularly beloved by readers who have delved deeply into this book. It vividly depicts a process that could be called data laundering. Raw data, which originally contains bias and uncertainty, is washed clean as it passes through the "scientific" filter of statistics multiple times, transforming into clean, precise numbers like a "3.57% growth rate." The precision of these decimal places masks the ambiguity that must have existed in the original data, emitting an inherently inappropriate "scent of certainty" that makes people forget to doubt.
These famous quotes live on through the ages not just because they are witty. They function as essential mental shortcuts or spiritual amulets for those of us living in an information-overloaded society. When faced with surprising statistical data, simply remembering the phrase "There are three kinds of lies..." allows us to pause from unconditional acceptance and maintain a critical distance. The deep chasm that lies between the objectivity of numbers and human subjectivity—these words convey this complex concept to us in an instant. They are the "proverbs" of the information age, and a powerful intellectual armor for maintaining healthy skepticism, even for those who are not experts.
Part 4: The "5 Keys" to Spotting Statistical Lies — A Shield to Protect Your Thinking
The book does not end with just exposing the tricks of lying. In the final chapter, it presents five key questions as practical weapons to protect yourself from these lies. These are a checklist that everyone should ask themselves when encountering statistical information.
1. Who says so?
The first thing to ask is the source of the information. Who conducted the survey, and for what purpose? Is there any conscious or unconscious bias at work, such as wanting to sell something or wanting to support a specific opinion? Even if the name of an authoritative organization is mentioned, do not take it at face value.
2. How does he know?
Next, ask about the survey methodology. How was the sample selected? Is the number sufficient, and does it represent the population? Was there any intent in the way the survey questions were phrased to lead to a specific answer?
3. What's missing?
It is important to look not only at the data presented but also at the "data not presented." Are only percentages shown, while the absolute numbers behind them are hidden? Is only the "average" discussed, while the dispersion of the data (range or standard deviation) is ignored? One should always look for the possibility that inconvenient data has been intentionally omitted.
4. Did somebody change the subject?
This is a question to check whether there is a "switch" between what was investigated and the conclusion drawn from it. For example, concluding that "Cold medicine A is effective in curing colds" from data showing that "People who took cold medicine A complained of fewer days of 'feeling unwell' than those who did not" is a classic switch.
5. Does it make sense?
The final key is the most basic yet most important "common sense check." Is the statistical conclusion plausible from a common-sense perspective? Also, even if there is a statistically "significant difference," is that difference large enough to be meaningful in the real world? Caution is required for conclusions that are too outlandish or numbers that are too precise.
These "5 keys" are arguably the greatest legacy this book has left for future generations. This is because these questions go far beyond the specific field of statistics, providing a universal framework for critical thinking that can be applied to all information we encounter in our daily lives.
As an experiment, let's replace the word "statistics" with "politician's claims." "That claim is being said by who? (What is his position?) ", "How did he know that conclusion? (What is the basis?) ", "Is there any missing information? (Is he hiding inconvenient facts?) ", "Did he change the subject? (We were supposed to be talking about the economy, but before I knew it, we were talking about something else...) ", "In the first place, does that claim make sense? (Is it realistic?) ". This framework works perfectly. By using statistics, which wear the mask of objectivity, as a perfect case study, Huff taught us a broader and more profound skill: disciplined skepticism. In an age where sophisticated fake news generated by AI threatens society, the importance of these five questions—asking for the source and basis of information—is higher than ever.
Part 5: Conclusion — Why we should read 'How to Lie with Statistics' in the age of AI and fake news
There is a decisive difference between the era in which this book was written and today. It is that the means to manipulate statistics and create good-looking graphs have passed into the hands of everyone, not just experts. Using spreadsheet software, anyone can process numerical values and create professional-grade graphs in an instant. This is a situation that could be called the "democratization of deception", and it has significantly changed the implications of this book.
This book is no longer just a guide for the "receiver" of information to protect themselves. Rather, it is increasingly taking on the role of an ethical textbook for the "creators" of information to hold themselves accountable. When we want to make our department's performance look a little better, or make our arguments more persuasive, we need to reread this book to resist the "devil's whisper" that tempts us to use the tricks introduced within it.
Furthermore, the classic wisdom of this book shines even brighter now as we face the new challenges of the 21st century, such as big data, algorithmic decision-making, and AI-generated content. Even if the technology is new, the fallacies that occur within it are surprisingly old. An AI trained on biased data is nothing more than a high-tech version of a "biased sample". Deepfake videos that are indistinguishable from the real thing are the modern equivalent of "gee-whiz graphs" that distort our perception. This book provides us with the fundamental strength to question the roots of these new forms of authority and to cast a healthy, skeptical eye upon them.
In conclusion, 'How to Lie with Statistics' is not just a good old-fashioned primer on numbers. It is a timeless call to build the foundations of intellectual self-defense, to seek curiosity, skepticism, and intellectual integrity.
Ironically, the very fact that this book remains a bestseller today may be proof that our society has not sufficiently learned its lessons. As the book points out, innocent generations without statistical literacy are born one after another, and even those of us who should know better require constant review. The contemporary significance of this book is a reflection of the fact that the problems it sought to solve 70 years ago remain deeply rooted in our society today. That is precisely why we should pick up this book again—not only to become smart consumers of information, but to become more responsible and ethical participants in the information society.
References
1. "Statistics don't lie, but liars use statistics" (1) "There are three types of lies...", accessed October 23, 2025, https://mediajuku.com/article/229
2. Knowing how to deceive in order not to be deceived — "How to Lie with Statistics...", accessed October 23, 2025, https://mukai.systems/articles/530wibce37ewmen1wa0fkhvt17bbd6uh/
(Created using Gemini 2.5 pro Deep Research and Claude Sonnet 4.5)
いいなと思ったら応援しよう!
よろしければ応援お願いします。いただいたチップは、執筆のための書籍購入や、AIのために使わせていただきます。