The authority behind almost every list of relationship rules traces back to the same place: a body of observational research on couples that produced widely repeated claims about predicting divorce with better than 90 percent accuracy. The rules feel solid because that number sits underneath them.
The number deserves a closer look, and two papers have given it one.
What a 90 percent accuracy figure actually describes
Accuracy in this context means the proportion of couples a statistical model sorts correctly into “divorced” and “still together.” It sounds like the number you want. It is not, because it depends heavily on how common divorce is in the sample being sorted.
The figure that matters for a rule you would actually use is different. If the model flags a particular couple as heading for divorce, how often is that flag right? That is the positive predictive value, and it is usually much lower than the headline accuracy.
The distinction is not a technicality. It is the difference between a model that describes a dataset well and a model that tells you something about a couple.
What happened on cross-validation
Heyman and Smith Slep published The Hazards of Predicting Divorce Without Crossvalidation in the Journal of Marriage and the Family in 2001. Their method was to demonstrate the problem rather than argue about it.
They took 528 people from the 1985 National Family Violence Survey, 176 divorced and 352 married or cohabiting, and split them at random into two halves of 264. They built a prediction equation on the first half, in the way such equations are normally built, then tested it on the second half, which the model had never seen.
On the half it was built from, the equation performed the way these models are usually reported as performing. Overall accuracy was 90 percent, sensitivity 92 percent, specificity 89 percent. When it predicted divorce, it was right 65 percent of the time.
On the fresh half, sensitivity fell by 45 percent and specificity by 15 percent. The positive predictive value dropped from 65 percent to 29 percent.
Then they made one further correction. Their sample contained about 33 percent divorced couples, well above the real-world rate, and models look better when the thing they are predicting is artificially common. Adjusted to a 16 percent prevalence, the positive predictive value fell to 21 percent.
A flag that is right about one time in five is not a rule. It is barely a hint.
Two points of care about what this paper is. It used National Family Violence Survey data to demonstrate a statistical hazard in the divorce-prediction literature. It was not a reanalysis of any particular research group’s data, and it does not show that any specific published model would fail this badly. What it shows is that a model can produce exactly the headline numbers people quote while carrying almost no predictive value for an individual couple, and that you cannot tell which case you are in without testing on data the model has not seen.
An independent sample, a different result
The second check came from a different research group with their own longitudinal data. Kim, Capaldi and Crosby published a generalizability test in the Journal of Marriage and Family in 2007, applying the affective process models developed by Gottman and colleagues to 85 married or cohabiting couples from an at-risk community sample, assessed around age 21 and followed up two and a half years later.
Several of the specific mechanisms did not predict relationship status in this sample. The researchers report that men’s rejection of their partner’s influence, men’s failure to de-escalate a partner’s negative affect, and women’s negative start-up were not predictive of whether the relationship survived.
They did find that the affective processes behaved differently depending on whether the couple was discussing a topic the man had raised or one the woman had raised, which is a real finding rather than a null one.
Eighty-five couples is a small sample, and an at-risk community sample is not a cross-section of couples generally. This is one test in one population, not a refutation. It is, however, the kind of test that has to come out well before a mechanism becomes a rule, and this one did not.
What this does not mean
It does not mean the underlying observational work was worthless, and it would be a poor reading of these two papers to conclude that.
The descriptive findings from decades of couples research are a genuine contribution. Contempt during conflict is associated with worse outcomes. Couples who repair small ruptures do better than couples who let them accumulate. Those associations show up across multiple laboratories and are not what is being questioned here.
What is being questioned is the step from association to prediction, and from prediction to a rule that applies to a particular couple at a particular kitchen table. That step is where the accuracy figures were doing work they could not support.
There is also a general point about the era. Both papers predate the widespread adoption of pre-registration and routine cross-validation in psychology. Fitting and testing a model on the same data was standard practice, not misconduct, and the field’s standards have moved since. That is context rather than exoneration for any specific claim still in circulation.
What would settle it
The test is not complicated in principle. Specify the model in advance, fit it on one dataset, and report how it performs on a different dataset collected by someone else, with the base rate of divorce set to what it actually is rather than what makes the numbers look better.
Until published rules are held to that standard, the honest description of the relationship-rules genre is that it rests on real observations about how couples behave and on prediction claims that have not survived the checks. A reader would be on safer ground treating the rules as descriptions of what tends to go wrong than as instruments that can tell them anything about their own relationship.