People rate the same act of vulnerability more favorably when someone else does it than when they imagine doing it themselves. The gap has a name, the beautiful mess effect, from a 2018 paper by Anna Bruk and colleagues at the University of Mannheim, published in the Journal of Personality and Social Psychology.
The finding is real. It is also narrower than the advice it has been turned into.
Those studies ran on written scenarios and short lab encounters, mostly with German university students. A separate group of researchers, publishing the same year, found that disclosing a weakness cost people status when they outranked the person listening. These are findings from particular datasets, not a universal rule about everyone, and neither one settles whether sharing a weakness is a good idea in any given situation.
What the studies actually compared
In the Bruk experiments, participants read a situation involving deliberate vulnerability. Confessing romantic feelings. Admitting a mistake. Asking for help. Revealing something personal. Some participants evaluated showing that vulnerability themselves; others evaluated the same situation performed by another person. The ratings diverged, consistently, in the same direction.
One study moved closer to real stakes. Participants either expected to improvise a song in front of a jury, or expected to sit on the jury and watch someone else do it. The performance never took place. The ratings still split the same way.
The full 2018 paper sits behind a paywall and we were not able to read its methods section. What we report here about the study count and the designs comes from the authors’ own later description of the work and from the British Psychological Society’s 2018 write-up of the seven studies. That is a limit worth stating rather than papering over.
The most likely misreading
The effect is usually summarized as proof that other people judge you less harshly than you fear. That is a reasonable extrapolation from the data. It is not quite what the design measures.
Bruk’s studies compare how a person evaluates their own vulnerability with how a person evaluates someone else’s. Those are self-other evaluation differences. They are not a direct comparison between what you predict an observer will think of you and what that observer actually thinks.
The distinction matters, because the second version is the one people act on.
Research that does test prediction against outcome exists, and some of it runs the same way. In a 2015 paper in Management Science, Alison Wood Brooks, Francesca Gino and Maurice Schweitzer reported that people avoid asking for advice because they fear looking incompetent, and that advice seekers were rated as more competent rather than less. The paper is careful about conditions. The effect was larger when the task was difficult rather than easy, when the advice was sought from the rater personally rather than from a third party, and when it was sought from an expert. Ask someone an easy question they are not qualified to answer and you are outside the situation the study describes.
A separate workplace boundary condition
The clearest counterweight comes from Kerry Roberts Gibson, Dana Harari and Jennifer Carson Marr, whose paper “When sharing hurts” appeared in Organizational Behavior and Human Decision Processes in 2018, the same year as Bruk’s.
That paper did not test the beautiful-mess self-other gap. It compared reactions to a simulated higher-status or peer coworker after a weakness disclosure or no disclosure. The result is a boundary condition for broad claims that vulnerability is always rewarded, not an exception demonstrated with the same outcome or design.
Across three experiments with 667 university undergraduates in total, participants encountered a colleague who disclosed a weakness. Academic probation in the first study, being overweight in the second, attending therapy in the third. The manipulated variable was whether that colleague outranked the participant or was a peer.
Higher-status disclosers lost perceived status. Mean ratings fell from 8.63 to 7.33 in the first study and from 5.82 to 5.19 in the second, on the different measures each used. Those disclosers were also rated as less influential, in more task conflict, and in lower-quality working relationships. Peers who disclosed exactly the same information paid no comparable penalty.
The authors argue the mechanism is that vulnerability from someone senior reads as a status violation, an unexpected signal from a person the observer had already filed under competence. Their third study’s mediation analysis supports that path within their own data.
These were laboratory experiments with random assignment, so causal language is fair about what happened inside them. Carrying it into a real workplace is a further step the design does not take, and the authors say so. They list the artificial setting, the absence of any chance for participants to reciprocate a disclosure, and the possibility that their peer condition still read as slightly senior.
The audience changed the result
A 2025 paper in Human Relations by Zhaopeng Liu and colleagues, covering 211 high-performing new employees and their colleagues, divided the question by audience rather than by rank.
Vulnerability shown to a supervisor was associated with that supervisor rating the newcomer’s ability lower, and with less proactive support afterward. Vulnerability shown to coworkers was associated with the opposite: less perceived threat, and more support. How perfectionist the leader was, and how competitive the group was, changed the size of both.
This is survey and colleague-report data rather than an experiment. The direction of the arrows is an interpretation the design cannot fully pin down.
What none of this settles
The self-other gap has been shown repeatedly by one research group. Their 2022 follow-up in Personality and Social Psychology Bulletin, which is openly available, ran four more studies with 60, 82, 97 and 101 German students, and found a smaller gap among people scoring higher in self-compassion. The interaction effects were medium to large, with partial eta squared between .07 and .19. That is the same team testing the same effect. Corroboration of a kind, but not independent replication.
We could not locate a large independent replication of the original seven studies.
The construal level account, in which people think about their own vulnerability in concrete detail and other people’s in the abstract, is the authors’ proposed explanation. It fits the data they collected. It has not been separately confirmed as the cause.
What the evidence supports is narrower than the slogan, and more useful. People do appear to hold themselves to a harsher standard than they apply to someone else in the same position. Asking for help, at least, seems to cost less than people expect it will. And the same disclosure lands differently depending on who hears it and where the speaker sits relative to them.
The study that would settle more of this is not a complicated one to describe. Follow real disclosures between real colleagues over time, and see what happens, rather than asking students to rate a stranger in a paragraph.