Researchers spent seven years retesting 274 published claims. About half failed, and even the findings that survived came back much smaller.
A team of researchers spent roughly seven years retesting 274 published claims from the social and behavioral sciences against new data. The retesting only confirmed about half of the claims, according to a study published in Nature. The failures spanned a dozen research fields.
The work came out of SCORE, short for Systematizing Confidence in Open Research and Evidence. DARPA funded the program, and the Center for Open Science in Charlottesville, Virginia, coordinated the contributions of 865 researchers, who sampled claims from 3,900 papers across 62 journals.
Andrew Tyner led the replication study itself, along with 291 co-authors. The team drew its 274 claims from 164 quantitative papers published between 2009 and 2018 in 54 journals.
The sample spanned economics, criminology, marketing, education, health, finance, management, organizational behavior, political science, psychology, public administration, and sociology. Twelve fields in all.
Replication teams preregistered their studies and used the original materials when available. Outside reviewers checked each protocol before data collection began, and the median study had 99.6 percent statistical power to detect the original effect. A failed replication was therefore unlikely to reflect a sample that was too small.
All 274 claims began as a hypothesis. A researcher predicted that one variable would affect another, collected the data, and reported that the prediction held.
A replication asks whether the same prediction holds in data that the original team never saw, gathered by people with no stake in the original result. The method sits at the center of how science checks itself.
The team counted a replication as successful when it produced a statistically significant result in the same direction as the original. By that criterion, 151 of the 274 claims held up, or 55.1 percent.
When weighted by paper, so that papers contributing several claims did not count extra, the rate was 49.3 percent. Rates across individual disciplines ranged from 42.5 to 63.1 percent. The authors noted that some of those field-level estimates are highly uncertain.
Effect sizes fell even among the claims that did replicate. The median effect across the original studies was 0.25 in Pearson’s r units. Across the replications, it was 0.10, less than half as large. Both figures are medians across hundreds of claims.
Earlier projects reported similar numbers. A 2015 effort that retested 100 psychology studies confirmed between 36 and 47 percent of them, depending on the criterion used.
Replication projects in economics in 2016 and in Nature and Science papers in 2018 confirmed that around 60 percent of the studies were replicable. The SCORE sample is larger than any of the others and spans 12 disciplines at once.
A failed replication is not evidence of misconduct. False positives occur by chance, and small samples tend to overstate real effects.
Analysts also make many defensible judgment calls about outliers, exclusions, and models. Any one of those calls can move a result across the threshold of statistical significance.
Journals have also historically favored positive results over null results, thereby keeping weaker findings in circulation. A null result often went unpublished entirely.
Fraud exists and is a growing concern for publishers. Investigators have documented organized networks that sell fake papers across several fields, and one analysis found that those networks were growing faster than legitimate research output.
Mathematics is dealing with fake citations and padded publication records. Non-replication is a distinct and much more common problem, one rooted in statistical practices and publishing incentives rather than in deception.
SCORE also tested whether replication outcomes could be forecast in advance. Two teams of human forecasters reached 76 and 78 percent accuracy by their best-performing measures, according to companion preprints released alongside the Nature papers.
Three machine-learning systems built for the same task, developed at Pennsylvania State University, the University of Southern California, and TwoSix Technologies, were not consistently accurate.
Tim Errington, senior director of research at the Center for Open Science, was one of the project leaders. “The main message of SCORE is a simple one: research is hard,” he said.
Public confidence appears steady in the meantime. A recent survey of about 72,000 people in 68 countries found moderately high trust in scientists across most of the world. Average trust in that study was 3.62 on a five-point scale.
The papers SCORE retested were all published between 2009 and 2018. Since then, many journals have adopted requirements for data sharing, code availability, and preregistered hypotheses.
Reforms of that kind could raise repeatability for newer work. That has not yet been tested at a comparable scale.
The full study was published in the journal Nature.
Support the Center for Open Science if you want to back the infrastructure behind this work, as the nonprofit builds and maintains the free Open Science Framework, which hosts preregistrations, data, and replication materials.
Join a local ReproducibiliTea journal club if you are a student or early-career researcher, because the grassroots network runs informal meetups at more than 100 universities to discuss open science and replication.
Explore the resources at FORRT if you teach, as the Framework for Open and Reproducible Research Training offers free lesson plans and summaries of replication outcomes for the classroom.
Become a member of the Society for the Improvement of Psychological Science if you work in the behavioral sciences, because the group coordinates practical reforms to research methods across institutions.
—–
Like what you read? Subscribe to our newsletter for engaging articles, exclusive content, and the latest updates.
Check us out on EarthSnap, a free app brought to you by Eric Ralls and Earth.com.
—–
Researchers attempted to replicate findings from 274 science journal papers, but only half could be successfully reproduced – Earth.com
By: SUDO
August 28, 2026
Recent Posts
- The AI Shift: Is AI supercharging science? – Financial Times
- Evolution and climate change are missing from Florida's proposed science standards – WFSU News
- Three UCSB researchers recognized for contributions to the earth sciences – UC Santa Barbara
- Sapphire discovery could reshape future brain-inspired electronics – anl.gov
- UAlbany Experts Explain the Science of a Perfect Pass – University at Albany
Recent Comments
No comments to show.
Subscribe to our newsletter!