Almost every article of the form “people raised in X households turn out like Y” rests on a single measurement: adults were asked what their childhood was like. That measurement has been checked against records made at the time, and it does not hold up the way the genre needs it to.

We are writers, not clinicians or researchers. What follows is a reading of two studies about how childhood is measured, not advice about anyone’s family.

The relevant comparison is between retrospective reports, meaning an adult’s recollection, and prospective records, meaning information gathered during the childhood itself.

What the meta-analysis actually compared

Baldwin, Reuben, Newbury and Danese published a systematic review and meta-analysis in JAMA Psychiatry in 2019 pooling 16 studies with 25,471 participants, average age 30.6.

Each of those studies had both kinds of measurement for the same people: what was documented about their childhood at the time, and what they said about it as adults.

Agreement between the two was poor. Cohen’s kappa was 0.19, with a confidence interval running from 0.14 to 0.24.

Kappa is worth a sentence, because the number is easy to misread. It is not a percentage of cases that matched. It measures agreement after subtracting the agreement you would expect from chance alone, on a scale where 1 is perfect and 0 is no better than guessing. A value near 0.19 sits close to the bottom of that range, and the authors describe it as poor.

The concrete version of that number is more useful than the statistic. More than half of the people identified prospectively as having experienced maltreatment did not report it retrospectively as adults. The mismatch also ran the other way, with adults reporting experiences that the contemporaneous records did not contain.

The authors’ conclusion is that the two measures identify largely different groups of individuals and cannot be used interchangeably. They go further, suggesting that people identified one way may have different pathways to later mental illness than people identified the other way.

The direction of the disagreement is worth noting too, because it rules out the simplest explanation. If adults merely under-reported, you would see misses in one direction only. Finding substantial error running both ways means the recollection is not a faded copy of the record. It is a different account.

Two measurements of the same childhood, in the same person, disagreeing more often than they agree.

The Dunedin cohort found the same split, and something sharper

The New Zealand birth cohort known as the Dunedin Study has followed 1,037 people, 91 percent of eligible births in the area, since the early 1970s, with adversity documented as it happened rather than reconstructed later.

In 2016, Reuben, Moffitt, Caspi, Belsky and colleagues published a paper in the Journal of Child Psychology and Psychiatry comparing the cohort’s contemporaneous records against what the same people recalled decades later. The correlation was 0.47, with a weighted kappa of 0.31. Moderate, not close.

The part worth slowing down for is what each type of measurement predicted.

Retrospective recollections were the stronger predictor of outcomes that were themselves subjective: self-rated health, memory complaints, things the participant reported about their own state. Prospective records were the stronger predictor of outcomes measured from outside the person: biomarkers, cognitive testing.

Each held up after controlling for the other, in its own domain. The retrospective associations with objectively measured outcomes shrank once the prospective records were accounted for.

Why that asymmetry matters for this kind of claim

If remembered childhood predicts how someone rates their own health but not what shows up in their blood work, part of what the memory is tracking is the person doing the remembering.

The Dunedin authors put this plainly: personality and perceptual factors may inflate the associations between remembered adversity and self-reported problems. Someone in a difficult period tends to recall a harder childhood, and also tends to report worse health, worse memory, worse everything asked about by questionnaire. The link between the two can be real without the childhood being the thing generating it.

This is the specific weakness in the “raised in a disciplined household, therefore these adult traits” format. The evidence for such claims is usually a single survey in which the same adult, on the same afternoon, reports both the childhood and the traits. That design cannot separate the two things these studies have shown need separating.

What this does not show

Several things, and the authors of the meta-analysis are careful about them.

Poor agreement between two measures does not establish that the retrospective one is wrong. It establishes that they disagree. Prospective records can miss cases too, because they only capture what a service, a researcher, or an agency noticed at the time, and much of what happens in a household is not noticed by anyone. The Baldwin team says explicitly that low agreement does not directly indicate poor validity of retrospective measures.

They also flag high heterogeneity across the 16 studies, which means the pooled kappa should be read as a rough summary rather than a precise constant.

There is a boundary that matters more than either, and articles citing this work tend to walk straight over it. These studies measured childhood maltreatment and adverse childhood experiences. They did not measure ordinary strict parenting, firm rules, early bedtimes, or high expectations. The measurement problem they identify plausibly applies to any adult recollection of childhood, but the specific numbers here belong to a different and more serious subject, and quoting the kappa as though it were about household discipline would be doing exactly the thing this article is complaining about.

What would settle it

The design that would answer the original question properly is expensive and slow: document parenting practices while they are happening, in a large and varied sample, then measure adult outcomes decades later with instruments that do not depend on the participant’s own account. Cohorts like Dunedin exist because someone started them fifty years ago.

Until a study of that kind is pointed specifically at parental discipline rather than at adversity, the honest position on what disciplined households produce is that the question is reasonable, the popular answers rest on a measurement known to disagree with the record, and nobody should be recognizing themselves in a list.