Brain imaging papers fail to replicate: thin evidence threatens neuroscience and disease progress
A replication crisis in MRI-linked brain-behavior findings is derailing what researchers think they know, and why it matters for medicine.

Live Science reports that foundational brain imaging and brain-wide association studies often cannot be replicated in follow-up work. For decision-makers funding or relying on neuroscience evidence, the consequence is simple: a shaky evidence base can slow cures, waste money, and mislead downstream clinical and research priorities.
If you think neuroscience is built on hard, repeatable facts, Live Science has a blunt reminder: several foundational brain imaging findings do not hold up when researchers try to reproduce them. The issue centers on brain-wide association studies (BWAS), where scientists link features of brain anatomy or function to behaviors or traits. In practice, when other teams repeat these studies, the expected brain-behavior associations frequently do not reappear.
This is not a niche academic annoyance. Randy Ellis, a senior scientist at Oracle who previously wrote about these issues while working as a biomedical informatician at the Icahn School of Medicine at Mount Sinai, told Live Science that the problems are big enough to “halt[] the growth and development of science and the curing of diseases.” That phrase is doing heavy lifting. It frames replication as an operational bottleneck, not a philosophical debate: if the “signal” in studies is fragile, the entire pipeline that depends on it, including how companies plan R&D and how labs choose targets, becomes less reliable.
So what actually goes wrong? Part of it is the incentive structure. Ellis points to what he calls the “original sin” of science: academics face huge pressure to publish positive results. That’s a familiar refrain across scientific fields, but Live Science describes why neuroscience makes some of these problems especially painful. Brain scans are expensive and access to MRI scanners can be difficult, which has historically pushed studies toward small samples. Small samples can make results look strong by chance, and then vanish when tested again.
The article walks through examples that show how easily a link can dissolve. A 2007 landmark study reported that the brains of children with attention-deficit/hyperactivity disorder took longer to mature. But a study published earlier this year found the relationship vanished once the different rates of aging between boys and girls were accounted for. Matthew Albaugh, a clinical neuroscientist at the University of Vermont and co-author of the replication paper, previously told Live Science that it was “what made the whole house of cards topple.”
Another example starts with a 2011 University College London study that linked density of gray matter in several brain areas to participants' number of Facebook friends. That was, at the time, a prominent social metric. A 2015 paper failed to replicate the finding, and the original paper’s authors argued the replication team did not follow the same protocol. Sarah Genon, a neuroscientist at the Research Center Jülich in Germany who was not involved in either the Facebook study or its replication, emphasizes the practical challenge: “It’s hard to do exactly what was done in the original studies.”
Genon’s later approach, as described by Live Science, goes after replication at scale. In 2019, she and colleagues ran a different kind of replication analysis using a vast database, conducting more than 10,000 analyses of MRI scans from hundreds of volunteers. The typical neuroimaging study sample size had been around 25, so they divided the bigger dataset into smaller ones to mirror the usual scale. If any micro-study found a significant association between behavior and brain structure, they attempted to reproduce it using a different subset of the data. The result: “Very few of the repeat studies replicated the results of the first.” Genon calls this “clear evidence” that the replicability of brain-behavior associations is relatively weak.
The statistical reason is straightforward, but the operational implication is not. In healthy volunteers, variation among individual brains is relatively small. That means you need a large number of participants to detect average differences tied to behavior. Genon likens the challenge to genome-wide association studies, where individual gene effects are tiny and researchers often need tens or hundreds of thousands of people to find robust signals. For brain-wide imaging studies, the article notes that only sample sizes in the thousands would likely be enough, citing a 2022 paper. But because scans are expensive and time-consuming, many studies have relied on much smaller samples.
Beyond sample size, the article points to measurement problems and modeling assumptions. Behavioral tests used to establish psychological variables can be inconsistent, Genon says. And some studies operate with an outdated assumption that small brain regions control complex traits like intelligence. Instead, Genon argues, those abilities are usually relatively distributed across the brain. If researchers analyze multiple small brain regions and run separate statistical tests, it increases the risk of producing “mirage associations,” meaning effects that look real in the dataset but do not persist.
Live Science also describes a more structural failure mode: overfitting. A 2017 brain imaging study by researchers at Weill Cornell Medical College explored whether MRI data could stratify patients with depression into four “biotypes,” each with distinct, unusual activity patterns in key brain networks. A follow-up study found the differences among the clusters were not statistically significant, meaning they could have happened by chance. Richard Dinga, a neuroscientist at the Friedrich Schiller University Jena in Germany and a co-worker on the follow-up study, says the main problem lay in the design. Its statistical approach, he argues, essentially guaranteed significant links because it correlated 17 clinical features of depression with 30,000 brain imaging features. Dinga says the authors could have used “30,000 coin flips” and gotten the same strong correlation, a classic sign that the model is too tuned to the original data.
Conor Liston, a psychiatrist at Weill Cornell Medicine and co-author of the original paper, acknowledged the model’s susceptibility to overfitting, telling Live Science, “That was not at all our intention, but that was the result.” When asked why the team hadn't updated the original paper to acknowledge the flaws, Liston said their subsequent publications were sufficient to correct the record. For executives and boards, the key takeaway is not the specific depression biotypes. It is what the case illustrates: even clinically oriented neuroscience work can generate apparently crisp decision-making categories that later crumble under more careful statistical scrutiny.
If replication is the obvious cure, the article also shows why it often does not happen. There is no money in it. Genon told Live Science grant agencies wouldn’t be excited about funding a study whose goal was simply to confirm past work. Even when replication is done, undermining prior findings may reduce publishability incentives. Dinga argues the way forward is to collect more data, noting that this solved similar irreproducibility problems in genetics, where genome-wide association studies reached the millions of samples over two decades. But more data is not the only fix. The broader “solutions” referenced in the article are improvements to study design, statistical approaches, and measurement practices, because replication failures often stem from how studies are built, not just how they are repeated.
Strategically, this is a governance problem as much as a scientific one. When the underlying evidence is fragile, downstream efforts to build diagnostics, pick therapeutic targets, or define patient stratification can spend real dollars chasing mirages. The neuroscience replication crisis described by Live Science is a warning label on an entire ecosystem, from academic labs to research sponsors to companies investing in brain-related R&D.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Science
NASA-backed engineers shrink high-fidelity energetic particle sensing into CubeSats
A compact, multiview particle-detection instrument could turn CubeSats into near-Earth science platforms without sacrificing data quality.

Stage 4 lung cancer at 44: a never-smoker’s ALK story that beat the odds
Summer Farmen turned an ALK-positive diagnosis into a six-year survival case study on targeted therapy and patient power.
Zoo elephants live longer now: study shows steady life expectancy gains since the 1960s
A multi-institution study in Scientific Reports finds modern zoo care is extending elephant lifespans, decade by decade.

