Statistical lies

What is the difference between truth, hypothesis, law and theory?
Design patterns
Think with AI
Think with AI
Statistical Lies — How Numbers Mislead
Statistical Thinking Series

Statistical lies.Real numbers.False impressions.

Ten ways statistics mislead—and how to read past the headline.

You see two bars. One towers over the other. The labels say 53.1% and 42.4%, but the picture seems to say “five times better.” The arithmetic can be correct while the message is wrong. To spot a statistical lie, follow the choices between the evidence and the conclusion.

A number becomes persuasive when you stop asking what was counted, what was left out, and what the comparison actually means.

01— Meaning

What is a statistical lie?#

A statistical lie deliberately creates a false impression through numbers: fabricated observations, altered records, selective comparisons, or a presentation that conceals essential context. “Misleading statistics” is the broader category. It also includes careless work and honest mistakes.

The familiar joke divides falsehoods into “lies, damned lies, and statistics.” Mark Twain repeated it and attributed it to Disraeli; that attribution is disputed. Treat it as a warning about persuasion, not a verified pronouncement by a famous statistician.[1]

Statistics is a method for examining evidence. A statistician, manager, advertiser, or journalist can misuse it. A misleading chart alone does not tell you whether its maker intended to deceive. You can establish what the chart gets wrong before making a claim about the person.

Deliberate deception

You know the records contradict the claim and alter or conceal them to preserve it.

Analytical error

You use the wrong formula, overlook a bias, or misunderstand what a result supports.

Honest uncertainty

You report the estimate and its limitations. An uncertain answer is not a dishonest one.

Fabrication & alteration

Type 01 · Change the evidence

The reported observation never happened—or no longer matches the record.

Imagine a report about the Yankees and Red Sox. Inventing match results is fabrication. Changing recorded losses into wins is alteration. Quietly deleting valid losses while presenting the remainder as the full record misrepresents the evidence even if every retained row is genuine.

Data cleaning can legitimately remove duplicates or repair errors. The difference is a defensible rule, a traceable record, and disclosure—not whether the cleaned result looks favorable. The American Statistical Association calls for transparent methods and resistance to pressure to predetermine or selectively interpret results.[2]

Ask for the chain from original records to reported totals.A polished dashboard cannot replace the underlying evidence.

Our running example: the supplied illustration labels the Yankees at 0.531 and the Red Sox at 0.424. We will read these as 53.1% and 42.4%. The image does not identify a season, game counts, or an original data source. These are transcribed illustration values, not independently verified historical records.[3]

02— Presentation

The picture can change the claim.#

Keep both rates fixed. Change only where a bar begins. Your eye compares the visible lengths, but those lengths now represent the amount above an arbitrary threshold.

Truncated axes

Type 02 · Change the picture

A bar should encode its value from zero.

At a 40% baseline, the visible lengths represent 13.1 and 2.4 percentage points. The first is about 5.46 times the second. At zero, the lengths represent 53.1 and 42.4, a ratio of about 1.25. Both panels retain correct labels; only one preserves the ratio of the reported rates.

That is why the UK Government Analysis Function recommends a zero baseline for bars. If the purpose is to inspect a small difference, use an appropriate alternative and make its scale clear.[4]

Two rates. One movable baseline.

The effect of truncating a bar chart at 40 percent Both panels show Yankees 53.1 percent and Red Sox 42.4 percent. The left bars start at 40 percent; the right bars start at zero. Both scales end at 60 percent. Baseline: 40% Truncation demonstration Baseline: 0% Rate comparison 40% 45% 50% 55% 60% 0% 15% 30% 45% 60% 53.1% 42.4% 53.1% 42.4% YankeesRed SoxYankeesRed Sox Visible height ratio: 5.46× Actual rate ratio: 1.25×

How to read this: compare each bar with its own baseline. The team labels and rates stay fixed. This redraw uses a common 60% upper bound to isolate the baseline effect; the supplied image used different upper bounds. The static view shows 40% on the left and zero on the right. Values come from the supplied illustration.[3]

A nonzero axis is not automatically deceptive in every chart. A line chart can use a restricted, clearly identified range to show variation over time. Its positions and slopes serve a different purpose from the lengths of bars. Still check the range and aspect ratio before judging how dramatic a trend is.[5]

Other visual distortions include unequal time spacing, a reversed scale without warning, and pictures whose width and height both grow with a quantity. Doubling both dimensions makes an icon four times the area. These design choices change the visual message without changing a printed value.

Percentage framing

Type 03 · Change the reference

Percent and percentage points answer different questions.

The Yankees’ displayed rate is 10.7 percentage points higher. Relative to the Red Sox rate, it is about 25.2% higher. These are compatible descriptions. Calling the difference “25.2 percentage points” is incorrect; calling it “25.2% more wins” would require comparable game counts that the image never supplies.

Absolute gap: 53.1% − 42.4% = 10.7 percentage points
Relative gap: (53.1 − 42.4) ÷ 42.4 × 100 ≈ 25.2%

Reverse the comparison and the reference changes: the Red Sox rate is about 20.2% lower than the Yankees rate, because you now divide the gap by 53.1. “Higher” and “lower” percentages need not be symmetric. Always name the reference.

10.7 ppAbsolute difference between the rates
25.2%Higher, using 42.4% as the reference
20.2%Lower, using 53.1% as the reference
03— Selection

What disappeared before the chart?#

A zero baseline fixes one problem. It does not establish that the right games were counted, the comparison periods match, or the summary represents the question you care about.

Selective inclusion

Types 04–06 · Change who counts

The missing observations can determine the story.

04 · Cherry-picking the period

Suppose you publish only a team’s strongest stretch and describe it as typical performance. Those wins may be real; the selection hides the stretches that would change the impression. Ask for the full relevant period and whether the start and end dates were chosen before the results were examined.

05 · Sampling and survivorship bias

A fan poll about which team is better does not measure a win rate. If you recruit only from one team’s fan page, it may not even represent baseball fans generally. A large response count does not by itself establish representative recruitment. Survey transparency requires information about the population, sampling, recruitment, questions, and weighting.[6]

Survivorship bias is a related selection problem: reviewing only clubs that reached the playoffs omits those that did not. You cannot explain why a strategy succeeds across all clubs by studying only its survivors. This is a possible failure mode, not a finding about the supplied image.

06 · Moving the denominator or definition

A win rate needs both wins and eligible games. Excluding unfavorable games, mixing home-only results with all-game results, or quietly redefining an “eligible game” changes the measure. In administrative reporting, the same problem appears when a success rate improves because difficult cases disappear from eligibility. Ask for numerator, denominator, exclusions, and definition changes together.

“53.1% of what, over which dates, with which exclusions?”If the report cannot answer, its precision is ahead of its documentation.

Convenient summaries

Type 07 · Hide the distribution

An average is a compression choice.

If you extend the team report to runs scored per game, a few very high-scoring games can pull up the mean. A median answers a different question about the middle of the distribution. Neither is universally the honest choice; disclose the measure and show enough spread to support the interpretation.[7]

Pooling win rates creates another trap. The combined rate is total wins divided by total eligible games. Taking a simple average of subgroup percentages works only when their denominators are equal. Home and away rates must be weighted by the corresponding game counts.

Combined win rate = (home wins + away wins)
÷ (home games + away games)

With different subgroup weights, an overall comparison can even reverse the comparisons within every subgroup: Simpson’s paradox. You need the subgroup counts to check for it; the two image values cannot establish that it happened here. And disaggregating is not automatically the correct causal analysis—the relevant grouping depends on the question and how the data arose.[8]

04— Interpretation

A precise number can support a weak claim.#

A displayed difference describes the supplied values. It does not by itself establish a stable performance advantage, identify its cause, or predict the next match.

False certainty

Type 08 · Hide uncertainty

Decimal places do not supply missing evidence.

You cannot calculate a justified confidence interval or significance test from these two percentages alone. You would need the game counts, the underlying records, and a defensible model. Independence cannot simply be assumed: teams may play each other and face different schedules.

If the records cover every game in a defined season, the rate describes that season without sampling uncertainty about the recorded total. Predicting future performance introduces a different source of uncertainty. Measurement mistakes, missing records, and model assumptions can matter in either case.

Even a valid small p-value would not tell you the probability that the favored team is “truly better,” or whether its advantage matters in practice. Statistical significance does not measure effect size or importance.[9]

Causal overreach

Type 09 · Invent an explanation

A difference between teams does not identify what caused it.

“The new coach caused the higher win rate” adds a claim the bars do not test. Rosters, injuries, opponents, and scheduling could also differ. A before-and-after comparison would still need to address other changes occurring at the same time.

A causal question asks what would have happened under an alternative decision. Random assignment can help create a credible comparison when feasible; observational analysis needs explicit assumptions and a suitable design. Merely adding more variables does not guarantee that those assumptions hold.[10]

Separate “the reported rates differ” from “this decision caused the difference.”The second statement requires additional evidence.

Selective analysis

Type 10 · Hide the search

The winning test is only part of the evidence.

Suppose you try many date windows, opponent groups, and performance measures, then report only the comparison with a small p-value. The reader sees one apparently decisive result and never sees how many opportunities you gave chance to produce it. This is the mechanism behind p-hacking and selective reporting.[9]

Exploration is useful. The misleading step is presenting a discovery from that search as if it were a single prediction specified beforehand. Report what you tried, distinguish exploratory from confirmatory claims, and use an appropriate correction or fresh evidence to evaluate selected findings.

What the supplied values actually support

A 10.7 percentage-point gap on a shared scale Horizontal bars begin at zero. The Yankees rate is 53.1 percent and the Red Sox rate is 42.4 percent. A bracket marks their 10.7 percentage-point difference. The illustration does not supply sample sizes or causal evidence. YankeesRed Sox 53.1%42.4% 0%15%30%45%60% 10.7 pp Sample sizes: unknown · Cause of difference: not established

How to read this: both horizontal bars use the same zero-based scale. The bracket marks the arithmetic gap, not a confidence interval. A relative comparison divides that gap by a stated reference rate. Nothing here supplies the missing game counts or demonstrates why the teams differ.[3]

Read like an auditor

A practical check

Reconstruct the claim before sharing it.

Start with the sentence the report wants you to believe. Then check the steps needed to support it. For our team comparison, these five questions expose both the visible distortion and the missing context.

  1. What is the source? Find the original match records and the exact reporting period.
  2. What was counted? Request wins, eligible games, exclusions, and consistent definitions.
  3. What comparison is being made? Check the baseline, reference rate, units, subgroup weights, and time window.
  4. How uncertain is the conclusion? Separate recorded results from an estimate of future ability.
  5. What would weaken the claim? Look for omitted periods, alternative analyses, and explanations other than the preferred one.

If you produce the report, make those answers easy to find. Preserve the original data, document transformations, disclose limitations, and correct errors visibly. A manager’s demand for a favorable headline does not make selective reporting sound statistical practice.[2]

One-minute evidence check

Which headline survives scrutiny?

You have only the supplied chart labels: Yankees 53.1%, Red Sox 42.4%. Which statement is justified?

Why this answer holds

The displayed difference is 53.1 − 42.4 = 10.7 percentage points. Relative to 42.4%, the Yankees rate is about 25.2% higher. The 5.46 ratio comes from truncated bar heights, and the image supplies neither game counts nor evidence about coaching. Only the statement about the displayed gap and missing counts is justified.

Are all misleading statistics intentional lies?

No. Misleading results can come from mistakes, poor design, or misunderstood uncertainty. Establish the analytical problem first. Calling it deliberate deception requires evidence about intent.

Does a large dataset make a claim trustworthy?

No. More observations can improve precision under suitable assumptions, but they do not automatically repair an unrepresentative sample, inconsistent definitions, or a biased analysis.

Can correct numbers still mislead?

Yes. The labels 53.1% and 42.4% stay unchanged when a 40% baseline inflates the visible height ratio to about 5.46. Missing context can distort interpretation just as effectively as changing a number.

05— Sources

Trace the evidence.#

Professional guidance and original methodological work support the explanations. The team scenarios are thought experiments built around the supplied illustration. They make no allegation about either club or the creator of the image.

  1. Mark Twain quotations: Statistics. Barbara Schmidt’s Twain quotation collection.Quotation and attribution notesDocuments Twain’s wording and flags the disputed attribution to Disraeli.
  2. Ethical Guidelines for Statistical Practice. American Statistical Association, 2022.Professional guidelinesProfessional integrity, data and methods, transparency, and responsibilities of leaders.
  3. Supplied “Percentage of victories” illustration. Reader-provided image; original author, season, and underlying records unidentified.Transcribed values: 0.531 and 0.424. Calculations: gap = 0.107; rate ratio = 0.531/0.424 ≈ 1.2524; visible height ratio at a 0.400 baseline = 0.131/0.024 ≈ 5.4583. Rounded display values are used throughout. No historical provenance is implied.
  4. Bar charts. UK Government Analysis Function.Chart guidanceZero baselines and alternatives when a bar chart obscures a small difference.
  5. Accessible charts: a checklist of the basics. UK Government Analysis Function.Chart checklistScale labeling, bar and line chart distinctions, sources, and accessible presentation.
  6. Best Practices for Survey Research. American Association for Public Opinion Research.Survey guidanceSampling, recruitment, question design, and disclosure of methodology.
  7. Measures of Location. NIST/SEMATECH e-Handbook of Statistical Methods.Mean, median, and distribution shape
  8. Comment: Understanding Simpson’s Paradox. Judea Pearl, The American Statistician, 2014.Author-hosted paper, PDFAggregation reversals and the role of causal reasoning in interpretation.
  9. Statement on Statistical Significance and P-Values. American Statistical Association, March 7, 2016.Statement summary, PDFSix principles, effect size, transparency, and selective reporting.
  10. Causal Inference in Statistics: An Overview. Judea Pearl, Statistics Surveys, 2009.Author-hosted paper, PDFThe assumptions and methods required to connect statistical associations to causal claims.

Ali Reza Rashidi
Ali Reza Rashidi
Ali Reza Rashidi, a Senior Data Scientist-Gen Al | Al Architect | MLOps with over ten years of experience, He is the author of three books that delve into the world of data and management.

Leave a Reply

Your email address will not be published. Required fields are marked *