Gregor Mendel’s pea-plant experiments are the foundation most genetics textbooks still open with: cross a tall plant with a short one, count the ratios in the next generations, and watch the numbers fall into the neat 3:1 pattern that gave the world the basic laws of inheritance. For 83 years, that foundation carried an asterisk. The numbers, critics said, were too good to be true.
That accusation traces to a single 1936 paper by the statistician Ronald A. Fisher, “Has Mendel’s Work Been Rediscovered?”, published in the journal Annals of Science.. Fisher ran a statistical test on Mendel’s reported results and found they matched the theoretically expected ratios more closely than chance should allow. His conclusion was blunt: somewhere in the process, consciously or not, Mendel’s data had likely been adjusted, whether by Mendel himself or by an assistant eager to please him.
How a 2019 paper reopened the case
A re-audit published in the journal Hereditas in October 2019, led by Noel Ellis of the John Innes Centre alongside Julie Hofer, Martin Swain and Peter van Dijk, went back through Fisher’s own statistical method rather than Mendel’s raw data. What they found was not evidence that Mendel had cheated, and not proof that he hadn’t. It was a problem with Fisher’s test itself.
Fisher had combined results across several of Mendel’s separate experiments into one overall chi-squared statistic, treating them as if they were a single, unified dataset. Ellis and his co-authors identified several ways that approach broke down. Some of the traits Mendel tracked are genetically linked, meaning they don’t assort independently the way Fisher’s model assumed. Mendel’s experiments weren’t run as a single randomised design; batches of seeds, growing conditions and the order in which crosses were counted all varied in ways Fisher’s aggregate test didn’t account for. Seedling survival rates differed between experiments, subtly skewing which plants made it to the point of being counted at all. And there is a reasonable chance that some phenotypes, especially subtle shades of colour or shape, were occasionally misclassified during counting, in either direction.
None of these issues, on their own, would necessarily produce Fisher’s “too good” result. Combined, according to the 2019 paper, they were enough to undermine the statistical basis for his conclusion. As the authors put it, their reanalysis “does not support previous suggestions that they differ remarkably from expectation.”
Two statisticians, two different standards of evidence
The interesting part of this story isn’t just that Fisher was wrong. It’s how differently the two analyses used the same basic tool, a chi-squared test, to reach opposite conclusions from data more than a century old.
Fisher’s approach pooled results across experiments to gain statistical power, which is a reasonable instinct if the underlying experiments are genuinely comparable. The 2019 paper’s contribution was showing they weren’t quite comparable enough for that pooling to be safe, once trait linkage, batch effects and classification noise are taken into account individually rather than smoothed over in aggregate.
This is one re-audit, not a second Fisher-scale consensus in its own right, and the authors are careful about how far they push the finding. They don’t claim to have proven Mendel’s data was genuine, in the sense of ruling out any possibility of selective reporting. What they show is that the specific statistical case built against Mendel in 1936 does not hold up to modern scrutiny of its own assumptions. That is a narrower claim than “Mendel is vindicated,” and a more interesting one: an 84-year-old accusation turned out to rest on flawed math, not on a flaw in the man being accused.
Why the “too good” instinct is a trap
Fisher’s original suspicion came from a reasonable place. Data that lines up unusually well with a theoretical prediction is, in general, worth a second look; fraud and unconscious bias in scientific data collection are real phenomena, not conspiracy theories. The trouble is that “too good to be true” is itself a statistical judgement, and one that depends heavily on how the comparison is constructed. Treat several distinct experiments as one giant dataset, and apparent precision can look suspicious. Treat them as what they actually were, a set of separate, imperfect trials run under slightly different conditions over several years, and that same precision looks more ordinary.
Other historians of science had already floated milder defences of Mendel over the decades, pointing to plausible explanations like unconscious selection in which plants were included in a count, without directly challenging Fisher’s statistics. The 2019 paper is the first to go after the mathematics of the accusation itself, rather than offering an alternative story around it.
What the reanalysis changes
The underlying genetics is untouched by any of this. Mendel’s laws of segregation and independent assortment don’t depend on whether his specific pea counts were slightly too tidy; they’ve been confirmed by thousands of subsequent experiments across countless species since 1900, when his work was rediscovered. What changes is a footnote that had quietly followed Mendel through nearly nine decades of textbooks: the suggestion that the founder of genetics may have fudged, or allowed someone else to fudge, the numbers that built the field.
Fisher remains one of the most influential statisticians of the twentieth century, and his 1936 paper wasn’t a fringe claim; it shaped how generations of scientists talked about Mendel. That a rigorous re-audit, using tools no more exotic than the one Fisher used himself, could unravel a criticism that stood for 83 years is a reminder that a statistical accusation is still a claim that needs checking, however credentialed the person making it.