The Replication Crisis Didn't End — It Moved to fMRI
Behavioral psychology cleaned up its act after 2015. Brain imaging is only now having its reckoning, for reasons that are mostly about sample size and analytic freedom.
Behavioral psychology cleaned up its act after 2015. Brain imaging is only now having its reckoning, for reasons that are mostly about sample size and analytic freedom.
For most of the 2010s, psychology was the field taking the public beating over replication. The Open Science Collaboration’s 2015 effort to redo 100 published studies found that a large share of classic findings didn’t hold up under a second look. The field’s response was, credit where due, substantial: preregistration became normal, sample sizes grew, and “exploratory” and “confirmatory” analysis stopped being treated as the same thing.
Neuroimaging mostly watched from the sidelines. It shouldn’t have.
Functional MRI is expensive — scanner time can run several hundred dollars an hour — which for decades pushed studies toward small cohorts, often 20 to 30 participants. In 2013, a team led by Katherine Button published an analysis in Nature Reviews Neuroscience arguing that underpowered studies don’t just fail to detect real effects; when they do report a significant result, that result is more likely to be inflated or simply wrong. Their term for it, “power failure,” became a reference point for a decade of follow-up criticism.
The problem compounds because an fMRI dataset isn’t one measurement — it’s tens of thousands of voxels, each analyzable in dozens of statistically defensible ways: different smoothing parameters, different software packages, different thresholds for what counts as a meaningful cluster of activity. Every one of those choices is a fork in the road, and different forks can lead to different conclusions from the same raw data.
That last point stopped being theoretical in 2020. A large collaboration, published in Nature under the name NARPS, gave the identical fMRI dataset to 70 independent analysis teams and asked them to test the same nine hypotheses. The teams didn’t just disagree at the margins — for several hypotheses, the proportion of teams reporting a significant effect ranged from close to zero to close to unanimous. No single team did anything obviously wrong. The pipelines were all reasonable. The data simply didn’t force a single answer, and the analytic freedom filled in the rest.
A separate and more technical problem surfaced in 2016, when Anders Eklund, Thomas Nichols, and Hans Knutsson tested the cluster-based statistics that most fMRI software used to flag “significant” brain activity. Using resting-state data — brains doing nothing task-related at all — they found some of the most common analysis pipelines produced false-positive rates far above the nominal 5%, in some configurations well above it. The bug wasn’t a fringe tool; it was embedded in software that had already been used in thousands of published papers.
It’s tempting to read all this as “brain scans are fake.” That’s not the right conclusion, and it’s not what the researchers who did this work were arguing. Large, well-powered, preregistered imaging studies — increasingly run through consortium datasets like the UK Biobank or the ABCD Study, which scan tens of thousands of participants rather than dozens — have held up well and produced some of the field’s most reliable findings. The issue was never that brains can’t be studied rigorously. It’s that a specific, common way of doing small studies with a lot of hidden analytic flexibility was quietly producing results that looked more solid than they were.
The corrective is already underway: bigger samples, preregistered analysis plans, and a norm of publishing the code alongside the paper so other teams can check which of the many reasonable choices actually drove the result. It’s the same fix psychology needed. It just took imaging longer to admit it needed it too.
The takeaway: when you read a headline about “the brain region for X” based on a single small imaging study, the honest response is patience, not dismissal — wait for a large or preregistered replication before treating the claim as settled.
← Back to all essays