Applied replication for data skills

Chris Moreh

Open Research Conference · Newcastle University · 16 June 2026

Workshop 1 · Open Research Conference · 10:15–12:45 · Room 1.06

  1. What survives when science repeats itself? – three Nature papers
  2. Whose question is it anyway? – estimands and DAGs
  3. You are analyst no. 6 – your own reanalysis, published before lunch

You leave with: your own public repo · a live web page · a row in the class multiverse

Work in pairs – one prepared laptop per pair

PART 1 – The three Rs

Talk I · 20 minutes

Replication crises

Another replication crisis – terminological confusion?

The term reproducibility has been used as a synonym for a number of different characteristics – reproducibility, robustness, replicability, repeatability, credibility, and trustworthiness.

Successfully rerun the same code on the same data? Re-analyse the same data a different way? Collect fresh data and analyse it in the same way? Using the same data and same analysis but asking a different question? Asking the same question but answering it with fresh data and a different analysis?

The data \(\times\) analysis matrix

Alipourfard et al. (2021); Nosek et al. (2025)

Three SCORE articles

SAME data
DIFFERENT data
SAME methods
Reproduction outcomes by discipline (Miske et al. 2026)
Replication outcomes by discipline (Tyner et al. 2026)
DIFFERENT methods
Reanalysis outcomes by discipline (Aczel et al. 2026)
Nature collection abstract text (Sanchez-Tojar et al. 2026)

Sánchez-Tójar et al. (2026)

What is SCORE?

  • DARPA-funded programme to gauge the credibility of social-science claims
  • 62 journals, papers from 2009–2018, stratified sample
  • Each paper’s central claim traced from abstract to the supporting statistic
  • Thousands of researchers; three repetition studies built on the same frame

Reproducibility – Miske et al.

  • 600 papers sampled (62 journals, 2009–2018)
  • Ask the authors for data and code, then try to rerun the reported result
  • How often can we even get to the starting line?

Most papers never reach the starting line

  • Of the 143 assessable: 53.6% reproduced precisely, 73.5% approximately
  • Counting the unassessable, only about 18% reproduce

Why some fields do better

  • Journals requiring data sharing: 87.5–100% availability vs 16.0% without
  • Political science 54.1% data availability vs education 2.9%
  • With data and code: 90.9% approx-reproducible vs 38.1% from rebuilt source

The most substantial barrier … was the unavailability of author data.

Reproducibility success does not mean that the finding is correct.

Robustness – Aczel et al. (Multi100)

  • 100 studies, each re-analysed by ≥5 independent analysts
  • Same data; each analyst free to choose any justifiable analysis
  • A peer panel judges whether each pipeline is appropriate
  • This is the paper I was an analyst in

Same data, five analysts, different answers

  • 34% of reanalysis effect sizes within ±0.05 d of the original (57% within ±0.20)
  • 74% reached the same conclusion; 24% null/inconclusive; 2% opposite
  • Effect sizes shrank: original mean d 0.73 → reanalysis 0.49
  • Analyst-to-analyst variability typically exceeds sampling error

Replicability – Tyner et al.

  • 274 claims from 164 papers, tested on new data
  • Median statistical power 99.6%; preregistered; original authors consulted
  • The strictest test: does the effect reappear in a fresh sample?

Even success shrinks

  • 55.1% of claims replicated by significance; effect sizes roughly halve
  • Thirteen success criteria give anywhere from 29% to 75%

The take-away

Reproducible ≠ robust ≠ replicable.

Across SCORE the three are essentially uncorrelated – knowing one barely predicts the others.

So there is no single number for trust. You have to ask which repetition you mean.

Quick check: credibility in the wild

You submit your paper and replication package. It goes out to reviewers.

  • Reviewer 1 wants to ensure that your results are reproducible, so she sources your R script to run the entire analysis. You’re in luck – she hasn’t yet updated her dplyr, so the code runs.
  • Reviewer 2 is very eager. He’s unconvinced by your modelling choices, so he fits your data to another function to check if your findings are truly robust. You’re in luck again – he recovers your point estimate to a reasonable precision, although the confidence intervals have widened.
  • Reviewer 3 takes a special interest because he’s just working on a similar paper on data they published recently. Do your results replicate when your methods are applied to that state-of-the-art dataset?

Quick check: credibility in the wild

You submit your paper and replication package. It goes out to reviewers.

Reviewer 2 just doesn’t want to let go…

  • He tried a few alternative analyses of your data, including a fancy new regression method that he just released as an R package. His method gives results that are generally aligned to yours, but for the method to work, it would require a larger dataset; would your results really generalise to this new dataset published by Reviewer 3 that he just read about, if applying his new fancy method to it?

PART 2 – Our case

Talk II · 10 minutes

On the OSF

https://osf.io/6zqct

The Teney (2016) study

  • 16 Eurobarometer waves, 2004–2013; ~ 390,000 respondents in 27 countries

The results show, first, that poor economic performances increase negative and decrease positive dimensions of EU framing.

  • What does the EU mean to you personally?
    • COSMOPOLITAN: positive non-materialist form of EU framing (Peace; Democracy; Freedom to travel, study and work anywhere in the EU; Cultural diversity; Stronger say in the world)
    • UTILITARIAN: positive materialist form of EU framing (economic prosperity and social protection)
    • COMMUNITARIAN: negative non-materialist form of EU framing (Loss of our cultural identity; More crime; Not enough control at external borders; Unemployment
    • LIBERTARIAN: negative materialist form of EU framing (Bureaucracy; Waste of money)

The claim card, as analysts received it

… poor economic performances … decrease positive dimensions of EU framing

Teney (2016: 619)

  • Anchor result – unemployment → cosmopolitan framing, Table 3 Model 1:
  • b = −0.00340 · t = −4.03 · p < .001

Reproducibility

  • Data: openly available but not freely redistributable

The required data to replicate my analyses can be downloaded on the Gesis repository website. I am not allowed to forward the Eurobarometer data. Enclosed you´ll find the code to combine the different Eurobarometer data waves and the code for the analysis of the ESR piece

  • Code: push-button replication failed

The reproduction attempt failed because the authors used file names that did not correspond with the names of the files available for download. Additionally 16 datasets are listed on the website as being used but the code references 18 files

  • Source-data reproduction (rebuild from raw) – not reproduced: could not extract nor re-calculate the original effect size

Robustness

  • All five negative – same direction · partial r from −0.45 to −0.01 · one not significant

Replication

  • Analysis One: specifically serves as the focal test for evaluation of the claim replicated for the purposes of SCORE.

    • 352,114 person-year observations in the analytic sample.
    • b = 0.002, p = 0.046, se = 0.0008, t = 2.004
  • Analysis Two: still studying and no full-time education categories in the variable age when finished full-timeeducation were NOT converted to missing values.

    • 379,396 person-year observations in the analytic sample
    • b = 0.002, p = 0.017, se = 0.0007, t = 2.394
  • A positive coefficient for unemployment on cosmopolitan framing, and significant at p=value results of 0.05 or less, after controlling for selected predictors, and wave and country-level effects

  • In this regard, this replication of the claim using Analysis Two was not successful according to the SCORE criteria

One paper, all three Rs

You are Analyst no. 6

  • Aim 1: Reproduce the original reanalyst C6HJR’s constrained result (Multi100 Task 2)
  • Task: Focus on the ‘Cosmopolitan’ dimension for EU framing and on unemployment rate as the independent variable; disregard contextual variables

  • Procedure: two-way fixed-effects model (mcosmopolitan ~ unemp_c | cntry + year)

  • Result: t ≈ −3.80 (df = 233, N = 270, partial r ≈ −0.242)

You are Analyst no. 6

  • Aim 2: Assess the robustness of the original reanalyst C6HJR’s constrained result
  • Task: Choose ONE deviation from the specification menu, evaluate the conceptual validity of your claim, clarify your estimand, commit to your chosen model before you run it, and submit your result

  • Procedure: ???

  • Result: ???

Two tracks

Track A – browser lab (zero install, webR in the page):

Track B – template repo (full Positron + git pipeline):

github.com/CodeMoreh/replication-lab

Use this template → clone in Positron

PART 3 – Estimands & DAGs

Concepts · 8-minute mini-lecture, then guided work

The red-card study: 29 teams, one dataset

(Silberzahn et al. 2018). Display sketch from published median (1.31) + range (0.89–2.93).

They were not answering the same question


Auspurg & Brüderl re-read the 29 analyses and found four different estimands:

  • descriptive association · discrimination net of mediators · variance-maximising · exploratory
  • The dispersion was largely a question-multiverse, not an analysis-multiverse

What is your estimand?

  • The theoretical estimand is stated outside any model – before R opens

DAG crash course: three roles

X = exposure, Y = outcome. Green = the causal effect you want; coloured = the path you must reason about.

One number, many stories

The same slope fits four different stories:

  • X really causes Y – the slope is the effect
  • a common cause inflates it – confounding
  • we are seeing a selected sample – collider / selection
  • Y causes X (reverse causation), or it is chance

The data look identical in all four. Only an assumption about how the data arose tells them apart – that is what a DAG writes down.

(D’Agostino McGowan et al. 2024; Keele et al. 2020). Scatter simulated for illustration.

Reading a DAG: open and closed paths

Association flows along every open path – not only the causal arrow. A fork or chain is open until you adjust the variable on it; a collider is the reverse – shut until you adjust it.

  • Back-door criterion: close every path that enters X from behind, keep the causal X → Y path open
  • That adjustment set – and only that set – earns a causal reading

Which effect do you even want?

  • Total = direct (green) + indirect (amber)
  • Adjust for the mediator M → you estimate the direct effect only – a different estimand
  • Control for everything silently answers a question you never asked
  • Table 2 fallacy: the other coefficients in your model are not all causal effects

Spurious by design: selection & feedback

  • Selection / collider bias – conditioning, or merely sampling, on a common effect manufactures an association that was never there (much of the early COVID risk-factor literature)
  • Reverse causation & feedback – in a panel, today’s outcome can move tomorrow’s exposure; two-way fixed effects absorb stable country traits and common-year shocks, not time-varying feedback (U = unemployment, F = framing)

A&B’s resolution: fix the estimand first

  • Fix the question → derive the model space → run all 486 justifiable models → SD 0.45 → 0.06

(Auspurg and Brüderl 2021). Display sketches from published medians + SDs (illustrative spread).

Revenons à nos moutons

The expected change in [unit quantity] among [target population] for a 1-percentage-point difference in unemployment, holding [assumptions] fixed.

  • Unit quantity – a person’s cosmopolitan framing, or a country-year mean?
  • Target population – which countries, which years?
  • You write this sentence in your report before you fit anything

The Teney DAG, half built

For online version: dagitty.net/dags.html Paste into Model code: Drag nodes, add edges, watch the adjustment sets update

dag {
unemployment [exposure]
framing [outcome]
country -> unemployment ; country -> framing
year -> unemployment ; year -> framing
unemployment -> framing
}
  • Drawn: the path + two confounder pairs (country, year → two-way FE)
  • Floating, unwired: growth (confounder?) · bailout (mediator?) · politicisation (mediator?)
  • You add the edges you believe, then run adjustmentSets()

PART 4 – The multiverse

Debrief · 12 minutes

You just lived a reproducibility story

You fetched rep_data.csv from the fork and fitted the constrained model.

You got t = −3.853.

The result recorded for analyst C6HJR in Multi100 is t = −3.804.

Same model. Same data, supposedly. Two different numbers. Why?

The discovery – and the correction

  • The prep script read EB 70.1 (ZA4819) twice, never EB 72.4 (ZA4994) – the 2009 rows are relabelled 2008 data
  • The analyst spotted this during the project and fixed it – the recorded −3.804 came from the corrected pipeline:

APRIL 2025 – Mistake in dataset here; the dataset associated with eb724 should be ZA4994_v3-0-0.dta – analyst’s own annotation

  • The official component kept only the earlier folder; the correction is published in the fork: osf.io/6zqct → moreh_rep_vAPR2025/

This is reproducibility working

The workshop panel, rebuilt from the raw GESIS files, recovers t = −3.804 exactly.

Every step was checkable because the materials were open – the as-submitted data, the corrected script, the rebuild.

Reproducibility working, not failing. Keep the whole trail public.

The class multiverse

  • 1,680 specifications · 68.7% claim-consistent direction · 63.9% significant · partial r from −0.59 to +0.70

Pre-computed over the specification menu (data/spec_grid.csv). Class results stream onto the Multiverse page, live.

Where your result lands

Your orange dot lands here → live on the Multiverse page as the room submits

report_result() gives you a one-click submission link; results stream onto the Multiverse page.

What we learned

the common single-path analyses in social and behavioural research should not be simply assumed to be robust to alternative analyses (Aczel et al. 2026)

A blueprint: precise question → causal reasoning (DAG) → multiverse → sensitivity analysis (Auspurg and Brüderl 2021)

  • Same-conclusion robustness and same-number robustness are different bars
  • Justifiability is the criterion – choose with reasons, show the spread

A single analytical path is a choice

  • There is no privileged single analytical path – only justifiable ones, explored openly.
  • You made one justifiable choice today – transparently, with a reason, inside a multiverse.
  • This is a skill.
  • And that was the workshop.

codemoreh.github.io/applied-replication

Fork anything. Everything is open.

Take home & resources

  • Yours now: a public reproducible repo (Track B) or a saved browser script (Track A); a result on the class curve
  • Read: the three Nature 2026 papers (Miske et al. 2026; Aczel et al. 2026; Tyner et al. 2026) · the glossary (Nosek et al. 2025) · Auspurg and Brüderl (2021) · Lundberg et al. (2021) · Teney (2016)
  • Explore: the Multi100 data · OSF preregistration templates · fixest & dagitty docs
  • At Newcastle: the RSE team · Library open-research pages · data.ncl.ac.uk · ReproducibiliTea

codemoreh.github.io/applied-replication

Fork anything. Everything is open.

References

Aczel, Balazs, Barnabas Szaszi, Harry T. Clelland, et al. 2026. “Investigating the Analytical Robustness of the Social and Behavioural Sciences.” Nature 652 (8108): 135–42. https://doi.org/10.1038/s41586-025-09844-9.
Alipourfard, Nazanin, Beatrix Arendt, Daniel M Benjamin, et al. 2021. Systematizing Confidence in Open Research and Evidence (SCORE). 46mnb_v1. SocArXiv. https://doi.org/10.31235/osf.io/46mnb.
Auspurg, Katrin, and Josef Brüderl. 2021. “Has the Credibility of the Social Sciences Been Credibly Destroyed? Reanalyzing the Many Analysts, One Data Set Project.” Socius 7 (January): 23780231211024421. https://doi.org/10.1177/23780231211024421.
Cole, Stephen R., Robert W. Platt, Enrique F. Schisterman, et al. 2010. “Illustrating Bias Due to Conditioning on a Collider.” International Journal of Epidemiology 39 (2): 417–20. https://doi.org/10.1093/ije/dyp334.
D’Agostino McGowan, Lucy, Travis Gerke, and Malcolm Barrett. 2024. “Causal Inference Is Not Just a Statistics Problem.” Journal of Statistics and Data Science Education 32 (2): 150–55. https://doi.org/10.1080/26939169.2023.2276446.
Griffith, Gareth J., Tim T. Morris, Matthew J. Tudball, et al. 2020. “Collider Bias Undermines Our Understanding of COVID-19 Disease Risk and Severity.” Nature Communications 11 (1): 5749. https://doi.org/10.1038/s41467-020-19478-2.
Hernán, Miguel A., and James M. Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC.
Huntington-Klein, Nick. 2021. The Effect: An Introduction to Research Design and Causality. Chapman & Hall/CRC. https://theeffectbook.net/.
Keele, Luke, Randolph T. Stevenson, and Felix Elwert. 2020. “The Causal Interpretation of Estimated Associations in Regression Models.” Political Science Research and Methods 8 (1): 1–13. https://doi.org/10.1017/psrm.2019.31.
Leszczensky, Lars, and Tobias Wolbring. 2022. “How to Deal with Reverse Causality Using Panel Data? Recommendations for Researchers Based on a Simulation Study.” Sociological Methods & Research 51 (2): 837–65. https://doi.org/10.1177/0049124119882473.
Lundberg, Ian, Rebecca Johnson, and Brandon M. Stewart. 2021. “What Is Your Estimand? Defining the Target Quantity Connects Statistical Evidence to Theory.” American Sociological Review 86 (3): 532–65. https://doi.org/10.1177/00031224211004187.
Miske, Olivia, Anna Lou Abatayo, Mason Daley, et al. 2026. “Investigating the Reproducibility of the Social and Behavioural Sciences.” Nature 652 (8108): 126–34. https://doi.org/10.1038/s41586-026-10203-5.
Nosek, Brian A, Timothy M Errington, Noah Haber, Theresa Stankov, and Andrew H Tyner. 2025. A Brief Glossary of Terms about Repeatability: Replicability, Robustness, and Reproducibility. mqfp4_v1. MetaArXiv. https://doi.org/10.31222/osf.io/mqfp4_v1.
Sánchez-Tójar, Alfredo, Jelte M. Wicherts, and Robb Willer. 2026. “Huge Meta-Research Project Puts Claims in Social-Science Papers to the Test.” Nature 652 (8108): 39–41. https://doi.org/10.1038/d41586-026-00805-4.
Silberzahn, R., E. L. Uhlmann, D. P. Martin, et al. 2018. “Many Analysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results.” Advances in Methods and Practices in Psychological Science 1 (3): 337–56. https://doi.org/10.1177/2515245917747646.
Teney, Céline. 2016. “Does the EU Economic Crisis Undermine Subjective Europeanization? Assessing the Dynamics of CitizensEU Framing Between 2004 and 2013.” European Sociological Review 32 (5): 619–33. https://doi.org/10.1093/esr/jcw008.
Tyner, Andrew H., Anna Lou Abatayo, Mason Daley, et al. 2026. “Investigating the Replicability of the Social and Behavioural Sciences.” Nature 652 (8108): 143–50. https://doi.org/10.1038/s41586-025-10078-y.

Thank you – and your feedback

Feedback poll
(QR / short URL)