Applied replication for data skills

Chris Moreh

Applied replication for data skills · a CodeMoreh workshop

A full-day workshop · four parts

  1. Why any of this, and with what tools – science, its crisis, and the toolchain
  2. The question underneath – one paper, five analysts, and what they were each asking
  3. What you can actually do – reproduce a published result, then work the menu
  4. Reading the multiverse – and choosing the one path you will defend

You leave with: a reproduced result · an estimand and a DAG of your own · a preregistered analysis · your dot on a live class multiverse

Work along in Positron (Route 1) or in the browser lab (Route 2) – solo or in pairs

Two kinds of aim

What I want you to understand

  • What reproducibility, robustness and replicability each mean, and why they come apart
  • Why analysts disagree – and how much of that is disagreement about the question
  • What a causal graph buys you when you cannot make a causal claim
  • What a multiverse shows, and what it cannot show
  • What preregistration settles, and what it leaves open

What I want you to be able to do

  • Get an analysis running from someone else’s materials
  • Write an estimand in one sentence, and a DAG that matches it
  • Reproduce a published result to the third decimal
  • Vary a specification deliberately and read what changed
  • Declare an analysis before you run it, and report where it sits among the alternatives

What you do not need today

  • No git, and no GitHub account
  • No OSF account – we will read OSF, not log into it
  • No installs on Route 2: the browser is the whole toolchain
  • Nothing to sign up for; nothing leaves your machine except the results you choose to submit

Do this now: Route 1 – open your workspace folder in Positron. Route 2 – open the browser lab tab and leave it loading; the engine warms while we talk.

PART 1 – Why any of this, and with what tools

Part 1 · 90 minutes

What are we doing when we do social science?

Making claims about the world that could turn out to be wrong – and saying how we would know.

A claim is not private. It is offered to others so they can check it, use it, argue with it, build on it.

That handover is the whole enterprise. Everything today is about whether it works.

Three arguments that made people nervous

  • The theory, 2005. With small studies, small true effects, flexible designs and many teams chasing the same question, most published findings in a field can be false with nobody breaking any rules (Ioannidis 2005)
  • The demonstration, 2011. Give researchers only the flexibility they legitimately have – which cases to drop, which covariates to keep, when to stop collecting – and a false-positive rate of 5% becomes over 60% (Simmons et al. 2011)
  • The subtlety, 2014. You do not need many tests. In a garden of forking paths, a single test chosen after looking at the data already has the wrong p-value (Gelman and Loken 2014)

And then people counted

  • 100 psychology experiments, replicated with high power and the original materials: 97% of the originals were significant, 36% of the replications were, and effect sizes roughly halved (Open Science Collaboration 2015)
  • 161 researchers, 73 teams, one hypothesis on immigration and support for social policy: 1,253 models, and no convergence on either sign or significance (Breznau et al. 2022)

Different fields, different designs, and in both the result depended on who ran the analysis.

It is not mostly fraud

The mechanism is ordinary and structural:

  • Journals reward novel, clean, significant results – so that is what gets written up
  • Every analysis contains dozens of small defensible choices, and the analyst makes them while looking at the data
  • Nobody publishes the twelve specifications that did not work; the reader sees one
  • Data and code are usually not available, so nobody can check any of it

Ask researchers anonymously, with an incentive to answer truthfully, and a majority admit to at least one of these practices – not fraud, the ordinary kind (John et al. 2012)

None of these steps requires a dishonest person. That is exactly why the problem is hard: you cannot fix it by hiring better people.

Twenty-nine teams, one question

The question. Are football referees more likely to give red cards to dark-skinned players than to light-skinned ones?

The data. One dataset, sent to everyone: 2,053 players in the top divisions of England, Germany, France and Spain, the 3,147 referees they played under, and 146,028 player–referee pairs. Skin tone rated from profile photographs by two coders who did not know the research question.

The design. Twenty-nine teams analysed it independently. They reviewed each other’s approaches before anyone had seen a result, revised, and only then compared answers.

This is the first many-analysts study. Every one since is a variation on it – including the one you are inside today.

Silberzahn et al. (2018)

All twenty-nine, side by side

Silberzahn et al. (2018), Figure 3, rebuilt from the published values at osf.io/gvm2z.

Replication crises

Another replication crisis – terminological confusion?

The term reproducibility has been used as a synonym for a number of different characteristics – reproducibility, robustness, replicability, repeatability, credibility, and trustworthiness.

Successfully rerun the same code on the same data? Re-analyse the same data a different way? Collect fresh data and analyse it in the same way? Using the same data and same analysis but asking a different question? Asking the same question but answering it with fresh data and a different analysis?

The data \(\times\) analysis matrix

Alipourfard et al. (2021); Nosek et al. (2025)

Credibility is not truth

Repeatability is used when referring to reproducibility, robustness, and replicability as a constellation of concepts intended to assess whether answers to scientific questions are unchanged when different steps of the research process are repeated.

Credibility and trustworthiness refer to the confidence one can have in research findings. Because scientific knowledge claims are tentative, they are not synonymous with truth.

For example, credibility is advanced by addressing potential competing interests in the conducting and reporting of research, by taking existing knowledge into account, by using methods that have been validated, by controlling for biases in the research context, by pursuing precise and reliable evidence, and by calibrating the strength of the claims to the uncertainty of the evidence.

Nosek et al. (2025)

Quick check: credibility in the wild

You submit your paper and replication package. It goes out to reviewers.

  • Reviewer 1 wants to ensure that your results are reproducible, so she sources your R script to run the entire analysis. You’re in luck – she hasn’t yet updated her dplyr, so the code runs.
  • Reviewer 2 is very eager. He’s unconvinced by your modelling choices, so he fits your data to another function to check if your findings are truly robust. You’re in luck again – he recovers your point estimate to a reasonable precision, although the confidence intervals have widened.
  • Reviewer 3 takes a special interest because he’s just working on a similar paper on data they published recently. Do your results replicate when your methods are applied to that state-of-the-art dataset?

Quick check: credibility in the wild

You submit your paper and replication package. It goes out to reviewers.

Reviewer 2 just doesn’t want to let go…

  • He tried a few alternative analyses of your data, including a fancy new regression method that he just released as an R package. His method gives results that are generally aligned to yours, but for the method to work, it would require a larger dataset; would your results really generalise to this new dataset published by Reviewer 3 that he just read about, if applying his new fancy method to it?

Open research: the practices

Practice What it makes possible
Open data Someone else can re-analyse your evidence
Open code Someone else can rerun your exact analysis
Open materials Someone else can run your study again
Open access Someone else can read the claim in the first place
Preregistration A dated record of the plan, before the results exist
Registered reports Peer review of the plan, with publication decided before results

Each one removes a specific obstacle. None of them is a general-purpose guarantee. The programme, set out at length: (Munafò et al. 2017)

What is SCORE?

  • DARPA-funded programme to gauge the credibility of social-science claims
  • 62 journals, papers from 2009–2018, stratified sample
  • Each paper’s central claim traced from abstract to the supporting statistic
  • Thousands of researchers; three repetition studies built on the same frame

Three corners, three studies

SAME data
DIFFERENT data
SAME methods
Reproduction outcomes by discipline (Miske et al. 2026)
Replication outcomes by discipline (Tyner et al. 2026)
DIFFERENT methods
Reanalysis outcomes by discipline (Aczel et al. 2026)
Nature collection abstract text (Sanchez-Tojar et al. 2026)

Sánchez-Tójar et al. (2026)

Reproducibility – Miske et al.

  • 600 papers sampled (62 journals, 2009–2018)
  • Ask the authors for data and code, then try to rerun the reported result
  • How often is that even possible?

Miske et al. (2026).

Most papers cannot be checked at all

  • Of the 143 assessable: 53.6% reproduced precisely, 73.5% approximately
  • Counting the unassessable, only about 18% reproduce

Why some fields do better

  • Journals requiring data sharing: 87.5–100% availability vs 16.0% without
  • Political science 54.1% data availability vs education 2.9%
  • With data and code: 90.9% approx-reproducible vs 38.1% from rebuilt source

The most substantial barrier … was the unavailability of author data.

Reproducibility success does not mean that the finding is correct.

Miske et al. (2026).

Robustness – Aczel et al. (Multi100)

  • 100 studies, each re-analysed by ≥5 independent analysts
  • Same data; each analyst free to choose any justifiable analysis
  • A peer panel judges whether each pipeline is appropriate
  • I was one of its analysts

Aczel et al. (2026).

Same data, five analysts, different answers

  • 34% of reanalysis effect sizes within ±0.05 d of the original (57% within ±0.20)
  • 74% reached the same conclusion; 24% null/inconclusive; 2% opposite
  • Effect sizes shrank: original mean d 0.73 → reanalysis 0.49
  • Analyst-to-analyst variability typically exceeds sampling error

Aczel et al. (2026).

Replicability – Tyner et al.

  • 274 claims from 164 papers, tested on new data
  • Median statistical power 99.6%; preregistered; original authors consulted
  • The strictest test: does the effect reappear in a fresh sample?

Tyner et al. (2026).

Effect sizes halve even when claims replicate

  • 55.1% of claims replicated by significance; effect sizes roughly halve
  • Thirteen success criteria give anywhere from 29% to 75%

Tyner et al. (2026).

The three Rs do not travel together

Reproducible ≠ robust ≠ replicable.

Across SCORE the three are essentially uncorrelated – knowing one barely predicts the others.

So there is no single number for trust. You have to ask which repetition you mean.

Inside Multi100: what an analyst was sent

  • A published paper, in full – no blinding: I could see the original estimates
  • Its central claim, in one sentence, lifted from the abstract
  • A pointer to the data – in our case, sixteen Eurobarometer waves you fetch yourself from GESIS
  • No instructions on method. Analyse it any way you can defend

And two things I could not see: who the other analysts were, and what they were finding.

Aczel et al. (2026). Registered protocol: osf.io/7snkz

The paper, for now in one slide

  • Author: I will say OA, the original author, from here on
  • Does the EU economic crisis undermine support for the EU? – 16 Eurobarometer waves, 2004–2013, about 390,000 respondents in 27 countries

The results show, first, that poor economic performances increase negative and decrease positive dimensions of EU framing.

  • Macro predictors: national unemployment and GDP growth. Outcome: what people say the EU means to them, sorted into four framing scales

Teney (2016). The full methodological reading is Part 2.

Two tasks, months apart

Task 1 – free reanalysis. Reanalyse the claim however you see fit. Record one categorical conclusion: does the evidence support it?

Task 2 – constrained reanalysis. Months later, three instructions arrive that fix the question:

  • use the cosmopolitan framing dimension as the outcome
  • use the unemployment rate as the predictor
  • disregard the contextual variables

Task 1 measures whether we agree. Task 2 makes our numbers comparable – you cannot average five answers to five different questions.

How a hundred studies become one number

  • Every analyst reports their test statistic and its degrees of freedom, in whatever form their model produces – t, z, F, \(\chi^2\)
  • Each is converted to a partial correlation r and a standardised Cohen’s d, so a Stata pooled model and an R multilevel model land on the same axis
  • A peer panel judges whether each pipeline is an appropriate analysis of the claim
  • Only then can you ask: how far apart are five defensible answers?

Standardising is what makes comparison possible – and it quietly assumes the five are estimating the same quantity. Hold that thought until Part 2.

Our five: one sign, very different sizes

  • All five negative – same direction · partial r from −0.45 to −0.006 · one not significant

Task 2 results as recorded in the public Multi100 dataset (osf.io/q5h2c).

Look at the labels

  • Two of us worked in Stata, three in R
  • Within R, three different document formats: a plain .R script, an .Rmd, a .qmd

The software did not cause the spread. There is no Stata answer and no R answer to this claim.

But because each of us submitted a document that somebody else could open and run, it became possible, afterwards, to find out what did cause it.

That is what the toolchain is for, and it is what you spend this afternoon doing.

What open source buys you

  • You can read it. The estimator is not a black box you are asked to trust – the source of lm(), of feols(), of every package we load today, is public
  • No licence stands between a reader and your analysis. Anyone, anywhere, can run your code without buying anything
  • It outlives your institution. A licence expires when you change job; a package on CRAN does not
  • You can fix it, and so can others. Bugs get found because thousands of people run the same code on different data

Reproducibility has a price, and a proprietary toolchain quietly puts that price on your reader.

The toolchain, laid out

Analyse
RRstatistics-first language
PythonPythongeneral purpose, strong ML
JuliaJuliafast numerics
StataSPSSStata / SPSSproprietary, still standard in places
Write
PositronPositrontoday’s Route 1
RStudioRStudiothe familiar one
VS CodeVS Codeeverything, via extensions
JupyterJupyternotebooks, Python-first
Publish and share
QuartoQuartotext + code → anything
GitGitevery version, dated
GitHubGitHubwhere the repository lives
OSFZenodoOSF / Zenodoarchive with a DOI

Ringed in orange: what today runs on. The rest is the map, so you know what you are choosing between.

Languages: what you are actually choosing

RRBuilt by statisticians for data analysis. Unbeatable for models, tables and figures; the ecosystem is the reason.
PythonPythonA general language that also does data. Better for pipelines, scraping, text and machine learning.
JuliaJuliaFast, young, small ecosystem. Worth it when your model takes hours.
StataStataExcellent panel and survey estimators, superb docs, one syntax. Costs money and hides its source.

You do not have to choose once and forever. Quarto runs R, Python and Julia chunks in the same document, and today’s browser lab is R compiled to run inside a web page.

Where you write it

PositronPositronThe newer IDE from Posit. R and Python as equals, VS Code underneath.
RStudioRStudioStill excellent, still supported. If it is what you know, use it.
VS CodeVS CodeDoes anything, once configured. Big extension ecosystem.
JupyterJupyterNotebook-first. Beware: cells can run out of order.
NeovimNeovimFor the keyboard-only. Steep, and very fast once climbed.

The editor is the least consequential choice on this slide. It changes how you work; it changes nothing about what a reader receives.

Trap worth naming: a notebook lets you run cells out of order, so it can show output that no clean run would produce. Restart and run everything before you believe it.

Positron and RStudio are both from Posit; both are free, and RStudio is open source.

Plain text is the format that survives

R script.RCode only. Comments carry the prose.
Markdown.mdProse only. Headings, lists, links – no code.
R Markdown.RmdProse + R, rendered together. R only.
Quarto.qmdProse + R, Python or Julia. Today’s format.

Every one of these is a plain text file. You can open it in anything, diff it, email it, and read it in twenty years. A .docx with pasted output is none of those things – and it has no idea which numbers came from which code.

What happens when you render

You write one source file. The rendering step runs your code, drops the results in, and produces whatever output you asked for – all from the same source.

The package layer

tidyversetidyverseReading, reshaping, plotting. One grammar across all of it.
easystatseasystatsModel output as tidy data: parameters, performance, effect sizes.
modelsummarymodelsummaryRegression tables you can publish, from the fitted objects.
marginaleffectsmarginaleffectsPredictions, contrasts and marginal effects from almost any model.

Today also loads fixest (fast fixed effects – the estimator behind our baseline), dagitty and ggdag (causal graphs). The full setup is on the setup page; the browser lab has it all preloaded.

Version control, in one slide

Git is a program on your machine. It records a complete, dated snapshot every time you say so, and it never forgets one.

GitHub is a website that hosts those snapshots so other people can reach them – and it is not the only one (GitLab, Codeberg, one hosted by your university).

What it gives you that a folder of final_v3_REALLY_final.R cannot:

  • Every past state, recoverable exactly
  • A written reason attached to each change
  • A public, timestamped trail of when you knew what
  • A way for two people to edit the same file without one overwriting the other

We use none of it today – deliberately. There is a companion module on the site, and your workspace folder is the ideal thing to practise on.

Where data and code are archived

OSFOSFProjects, files, and registrations. Free, and where our case’s trail lives.
ZenodoZenodoCERN-run archive. Gives a DOI; takes a GitHub release directly.
DataverseDataverseInstitutional data archives with proper metadata and curation.
XYZYour domain’s archiveGESIS, UK Data Service, ICPSR – licensed access to survey microdata.

A repository is not an archive. GitHub can be force-pushed, renamed, or deleted tomorrow. An archived deposit gets a DOI, a fixed version, and a commitment to keep serving it. Serious work usually needs both.

Licences, briefly

What Common choice Why
Code MIT · GPL MIT: use it however you like. GPL: and share your changes back
Data CC0 · CC BY CC0 puts it in the public domain; CC BY asks for attribution
Text, figures, slides CC BY · CC BY-NC Attribution, optionally no commercial reuse
Someone else’s data their terms Eurobarometer is GESIS-licensed – I may not redistribute it, and neither may you

No licence is not a neutral default. It means all rights reserved – legally, nobody may reuse it at all.

What we use today

Route 1 – Positron

  • The workspace folder you downloaded, opened in Positron
  • Quarto renders your report; fixest fits your models
  • No git, no cloning, no configuration – it is just files
  • This is what I drive on the shared screen

Route 2 – the browser lab

  • One web page. R is compiled to run inside the browser
  • Nothing installed, nothing to break, data already in place
  • Full parity: every task exists in both routes
  • Your spare if Route 1 misbehaves at any point

github.com/CodeMoreh/replication-lab – download zip → unzip → open in Positron

What the tools cannot do

A perfectly reproducible pipeline can be a perfectly reproducible answer to a question you never stated.

None of this software will tell you what your estimand is, whether your controls belong in the model, or whether the association you found means what you think it means.

Tooling is the cheap half of the problem. We have now done it. The rest of the day is the expensive half.

Hands-on 1: get yourself running

Task 1 – the setup check. Whichever route you are on, get to the point where a line of R runs and the data are in front of you.

  • Route 1: Positron opens the workspace folder · the packages install · working_script.R runs its first block · the report skeleton renders
  • Route 2: the browser lab page has finished loading · the first chunk runs · euframes has 270 rows

Take the rest of this part and the break for it. Type a green circle in chat when you are through; type STUCK and I will come to you.

If Route 1 will not cooperate, do not fight it – switch to Route 2 and carry on. Nothing today depends on which one you use.

Break

30 minutes – the exact return time is in the chat · keep setting up if you need to

PART 2 – The question underneath

Part 2 · 90 minutes

The paper we spend the day inside

  • 16 Eurobarometer waves, 2004–2013; about 390,000 respondents in 27 countries

The results show, first, that poor economic performances increase negative and decrease positive dimensions of EU framing.

  • What does the EU mean to you personally?
    • COSMOPOLITAN: positive non-materialist form of EU framing (Peace; Democracy; Freedom to travel, study and work anywhere in the EU; Cultural diversity; Stronger say in the world)
    • UTILITARIAN: positive materialist form of EU framing (economic prosperity and social protection)
    • COMMUNITARIAN: negative non-materialist form of EU framing (Loss of our cultural identity; More crime; Not enough control at external borders; Unemployment)
    • LIBERTARIAN: negative materialist form of EU framing (Bureaucracy; Waste of money)

Teney (2016).

The claim card, as analysts received it

… poor economic performances … decrease positive dimensions of EU framing

Teney (2016: 619)

  • Anchor result – unemployment → cosmopolitan framing, Table 3 Model 1:
  • b = −0.00340 · t = −4.03 · p < .001

What OA actually fitted

  • The unit is a person, not a country-year: roughly 390,000 respondents
  • Thirteen response options collapse into four framing scales – cosmopolitan, utilitarian, communitarian, libertarian
  • A three-level model: respondents inside country-waves inside countries
  • Individual covariates: sex · age · education band · occupational status · urbanisation · left–right placement
  • Macro predictors: unemployment rate and GDP growth, entering at the country-wave level
  • Plus cross-level interactions – does the macro effect differ by education?

Anchor result, Table 3 Model 1 – unemployment on cosmopolitan framing: b = −0.00340 · t = −4.03 · p < .001

Teney (2016).

The robustness question is already in the paper

  • Model 1: unemployment → cosmopolitan framing · b = −0.00340 · t = −4.03 · p < .001
  • Model 2: the same coefficient, with further controls added – drops by more than half and loses significance

Two models, one paper, and support for the claim depends on which one you read. The paper reports both, which is more than many do.

For today, that matters because the sixth analyst’s problem – which defensible specification do I report? – was already live inside the original study, before any of us arrived.

What we work on instead

OA’s data

  • ~390,000 respondents
  • 16 waves, 27 countries
  • GESIS-licensed: I may not redistribute it
  • Models take minutes

Today’s panelEUframes_cy.csv

  • 270 country-years, one row each
  • The framing scales, averaged within each country-year
  • Freely shareable, ships with both routes
  • Models take milliseconds

The aggregate answers the contextual question directly: does a country’s framing move when its unemployment moves? What it cannot answer is the compositional one – whether it is the same people framing differently, or a different mix of people. There is a companion module that takes that up on simulated individual data.

Reproducibility: two attempts, two failures

  • Data: openly available but not freely redistributable

The required data to replicate my analyses can be downloaded on the Gesis repository website. I am not allowed to forward the Eurobarometer data. Enclosed you´ll find the code to combine the different Eurobarometer data waves and the code for the analysis of the ESR piece

  • Code: push-button reproduction failed

The reproduction attempt failed because the authors used file names that did not correspond with the names of the files available for download. Additionally 16 datasets are listed on the website as being used but the code references 18 files

  • Source-data reproduction (rebuild from raw) – not reproduced: could neither extract nor re-calculate the original effect size

Replication: the coefficient flips sign

  • Analysis One – the focal test on which SCORE evaluates the claim

    • 352,114 person-year observations in the analytic sample
    • b = 0.002, p = 0.046, se = 0.0008, t = 2.004
  • Analysis Two – the still studying and no full-time education categories in the variable age when finished full-time education were not converted to missing values

    • 379,396 person-year observations in the analytic sample
    • b = 0.002, p = 0.017, se = 0.0007, t = 2.394
  • Both give a positive coefficient for unemployment on cosmopolitan framing, significant at p-values of 0.05 or less, after controlling for selected predictors and for wave and country-level effects

  • By the criteria used by SCORE, Analysis Two did not replicate the claim

One paper, all three Rs

Same verdict, five different questions

Analyst Free-phase predictor Free-phase outcome Estimator
018OL GDP per capita (level) positive framing, 7-item mean crossed random intercepts
KEVF1 GDP growth positive framing, 7-item sum random intercepts
PRL47 log GDP + growth + crisis dummy three positive-framing shares three-level mixed
KQXUE GDP growth and unemployment cosmopolitan + utilitarian shares country FE
C6HJR GDP growth and unemployment positive framing, 7-item share two-way FE
  • Five free analyses, five different questions – three of us never touched unemployment at all – and one shared recorded verdict: evidence for the claim
  • The constrained phase then manufactured a common question; its five comparable answers span −0.006 to −0.45
  • Conclusions can agree while estimates cannot even be compared – the two robustness questions that Multi100 was built to separate (Aczel et al. 2026)

Free-phase records: the public Multi100 dataset (osf.io/q5h2c) – categorical verdicts and free-text reports only; the models themselves read from the analysts’ public code.

Two studies to set beside our five

Our own case has just shown us the pattern in miniature. It was described first, and at far greater scale, in the two studies that the rest of this part rests on.

Silberzahn and colleagues (2018) – the red-card study from this morning. Twenty-nine teams, one dataset, one question. Twenty of the twenty-nine intervals excluded 1; nine did not.

Auspurg and Brüderl (2021) went back through all twenty-nine of those analyses and asked a question that the original had not: what was each team actually trying to measure?

The first establishes that careful people disagree. The second explains why – and that explanation is what the whole of Part 2 is built on.

Silberzahn et al. (2018); Auspurg and Brüderl (2021)

They were not answering the same question


Auspurg & Brüderl re-read the 29 analyses and found four different estimands:

descriptive association · discrimination net of mediators · variance-maximising · exploratory

The dispersion was largely a question-multiverse, not an analysis-multiverse.

Auspurg and Brüderl (2021)

Fix the estimand, and the spread collapses

  • Fix the question → derive the model space from a DAG → run all 486 justifiable models → SD 0.45 → 0.06
  • The spread in the top row is two analyses wide: a tight cluster plus the two outliers, which are what carry that 0.45

Auspurg and Brüderl (2021). Top row: the 29 published estimates. Bottom row: drawn to their reported median (1.28) and SD (0.06) – the 486 values are not published.

Has credibility been credibly destroyed?

Auspurg and Brüderl’s title asks it directly: has the credibility of the social sciences been credibly destroyed?

If the twenty-nine teams answered four different questions, their disagreement is not evidence that analysis is unreliable.

It is evidence that the question was under-specified – an ordinary, fixable defect.

So this literature reads as a diagnosis of vague research questions rather than a verdict on social science.

More optimistic – and more demanding, because it hands the problem to whoever writes the question.

Auspurg and Brüderl (2021)

What is your estimand?

  • The theoretical estimand is stated outside any model – before R opens

Lundberg et al. (2021).

DAG crash course: three roles

X = predictor, Y = outcome. Green = the causal effect you want; coloured = the path you must reason about.

One number, many stories

The same slope fits four different stories:

  • X really causes Y – the slope is the effect
  • a common cause inflates it – confounding
  • we are seeing a selected sample – collider / selection
  • Y causes X (reverse causation), or it is chance

The data look identical in all four. Only an assumption about how the data arose tells them apart – that is what a DAG writes down.

D’Agostino McGowan et al. (2024); Keele et al. (2020). Scatter simulated for illustration.

Reading a DAG: open and closed paths

Association flows along every open path, not only along the causal arrow. A fork or a chain is open until you adjust the variable sitting on it. A collider is the reverse: shut by default, and opened by the very act of adjusting for it.

  • Back-door criterion: close every path that enters X from behind, keep the causal X → Y path open
  • That adjustment set – and only that set – earns a causal reading

Hernán and Robins (2020); Huntington-Klein (2021).

Total effect or direct effect

  • Total = direct (green) + indirect (amber)
  • Adjust for the mediator M → you estimate the direct effect only – a different estimand
  • Control for everything silently answers a question you never asked
  • Table 2 fallacy: the other coefficients in your model are not all causal effects

Keele et al. (2020); D’Agostino McGowan et al. (2024).

Selection and feedback

  • Selection / collider bias – conditioning, or merely sampling, on a common effect manufactures an association that was never there (much of the early COVID risk-factor literature)
  • Reverse causation & feedback – in a panel, today’s outcome can move tomorrow’s predictor; two-way fixed effects absorb stable country traits and common-year shocks, not time-varying feedback (U = unemployment, F = framing)

Griffith et al. (2020); Cole et al. (2010); Leszczensky and Wolbring (2022).

Revenons à nos moutons

The expected change in [unit quantity] among [target population] for a 1-percentage-point difference in unemployment, holding [assumptions] fixed.

  • Unit quantity – a person’s cosmopolitan framing, or a country-year mean?
  • Target population – which countries, which years?

In chat, now: your one-sentence estimand for the claim. Type it, hold it, send on go.

We cannot identify a causal effect here

Observational panel · 27 countries · no assignment mechanism · feedback between the variables · a decade of shared shocks.

Nothing in this design licenses unemployment causes a shift in EU framing.

So why spend half an hour on causal graphs?

Because the graph does three jobs that survive the loss of identification:

  • It says which association you are after – the total effect, not whatever the software returns
  • It tells you which variables belong in the model, and, just as usefully, which do not
  • It makes your assumptions legible and arguable – a reader can disagree with an arrow, and that is a productive disagreement

Publishing the estimand and the graph

The graph and the estimand are only worth anything if they are written down where someone else can read them:

  • in the paper, not in your head
  • before the model runs, not reconstructed afterwards
  • with the limitations named by you, not discovered by a reviewer

In Part 4 you write exactly that down – the estimand, the graph-based justification, the one specification, and the rule for what would count as support – and you do it before you see a result.

The technical name for that is preregistration, and what it does is stop the result from choosing the question.

The EU-frames DAG, half built

The same graph in R. Every arrow is a formula, read outcome-first – so a # takes an arrow out and the code still runs:

dagify(
  framing ~ unemployment,
  framing ~ country + year,
  unemployment ~ country + year,
  # framing ~ growth,        # uncomment the
  # unemployment ~ growth,   # arrows you believe
  exposure = "unemployment", outcome = "framing"
) |> adjustmentSets()
  • Drawn: the path + two confounder pairs (country, year → two-way FE)
  • Floating, unwired: growth (confounder?) · bailout (mediator?) · politicisation (mediator?)
  • You add the edges you believe, then run adjustmentSets()

R, in one slide

x <- 42                          # <- assigns. "x gets 42"

mean(c(2, 4, 6))                 # a function: name, then arguments in ()

euframes$unemp                   # $ picks one column out of a data frame

euframes |> nrow()               # |> the pipe: send the left into the right

Four things and you can read most R: assign with <- · call a function with name(...) · reach into an object with $ · chain with the pipe |>.

Getting data in

library(readr)                                # load a package

euframes <- read_csv("data/EUframes_cy.csv")  # read it, name it

euframes                                      # print it

What comes back is a data frame: a rectangle, one row per observation, one column per variable, each column of a single type. Almost everything in R takes a data frame and gives you one back.

library() makes the functions in a package available. It does not install it – that is install.packages(), and it happens once, not every time.

Look before you model

glimpse(euframes)      # every column, its type, the first few values
nrow(euframes)         # 270
summary(euframes$unemp)  # min, quartiles, mean, max

euframes |>
  count(cntry)         # how many rows per country?

This is not a warm-up you skip. Every serious error I have made in twenty years of this was visible in the data before it reached a model – a duplicated year, a missing value coded as -9, a scale running 1–10 where I assumed 0–1.

Five verbs do most of the work

euframes |>
  filter(year >= 2008) |>                       # keep some rows
  select(cntry, year, mcosmo, unemp) |>         # keep some columns
  mutate(unemp_c = unemp - mean(unemp)) |>      # make a new column
  arrange(desc(unemp))                          # sort

Read it downwards as a sentence: take the panel, then keep 2008 onwards, then keep four columns, then add a centred unemployment column, then sort by unemployment, highest first.

The sixth verb is summarise(), usually with group_by(): collapse many rows into one per group. That is precisely how the panel you are working on was built from 390,000 respondents.

A model is a formula plus data

mcosmo ~ unemp                    # "mcosmo explained by unemp"

m <- lm(mcosmo ~ unemp, data = euframes)

summary(m)

The tilde ~ reads outcome on the left, predictors on the right. Add predictors with +. The formula is an object in its own right – it does not know or care which function will eventually use it.

Almost every modelling function in R takes this same shape: something(formula, data = ...). Learn the pattern once and lm, glm, lmer, feols and betareg all become variations.

Fixed effects, the fast way

library(fixest)

m1 <- feols(mcosmo ~ unemp_c | cntry + year, data = euframes)

The bar | separates the regression from the fixed effects. Everything after it is absorbed: cntry gives every country its own intercept, year gives every year its own.

The equivalent with dummies would be mcosmo ~ unemp_c + factor(cntry) + factor(year) – same estimate, 36 extra rows of output you do not want to read.

This is the model you reproduce in Part 3. t = −3.804.

Reading what comes back

broom::tidy(m1)
#> term      estimate    std.error   statistic   p.value
#> unemp_c   -0.00347    0.00091     -3.80       0.00019
  • estimate – a one-point rise in unemployment goes with a fall of 0.00347 on a 0–1 framing scale
  • std.error – how much that estimate would wobble across repeated samples
  • statistic – the t: the estimate divided by its standard error. Roughly, how many standard errors from zero
  • p.value – how often you would see a t this large if the true value were zero

broom::tidy() returns the table as a data frame, so you can filter it, join it, plot it. That is how 840 model fits become one chart this afternoon.

One line of ggplot

ggplot(euframes, aes(x = year, y = mcosmo)) +   # data, and what maps to what
  geom_line(colour = "grey55") +                # draw lines
  geom_point(size = 0.7) +                      # and points
  facet_wrap(~ cntry)                           # one small panel per country

Three parts, always: the data, the mapping (aes) from variables to visual properties, and one or more geoms that draw something. Layers stack with +.

facet_wrap is the one worth stealing today: it repeats the whole plot once per group. Twenty-seven small country panels, one line of code, and you can see which countries move.

A causal graph is a set of formulas

dag <- dagify(
  framing ~ unemployment,          # unemployment -> framing
  framing ~ country + year,        # country and year -> framing
  unemployment ~ country + year,   # ... and -> unemployment
  exposure = "unemployment",
  outcome  = "framing")

adjustmentSets(dag)                # what must I control for?

Each line is one arrow, read outcome-first – exactly the same convention as a model formula. Which means you can comment an arrow out with a # and the code still runs.

adjustmentSets() reads the graph and tells you which variables close every back-door path. Change an arrow, rerun, and watch the answer change.

When it breaks

What you see What it usually means
could not find function "feols" The package is not loaded – add library(fixest) above
object 'euframes' not found You ran this chunk before the one that creates it – go back up and run in order
object 'unemp_c' not found The column does not exist yet – you skipped the mutate() that makes it
+ on the console, nothing happens An unclosed bracket or quote. Press Esc and look for the missing )

Read the error. R’s messages are terse but almost always literally true – it is telling you the name of the thing it could not find.

Hands-on 2: meet the panel, draw the graph

Task 2, in either route – three pieces:

  1. Look at the data. 270 rows, 27 countries, 10 years. Run the country panel plot and find the ones that move
  2. Write the estimand in words. The template gives you the unit quantity, the target population, the aggregation – fill in each for this claim
  3. Wire up the DAG. Take the # off the arrows you are willing to defend, rerun adjustmentSets(), and see what it now demands

Two sentences for your notes: what do two-way fixed effects adjust for, and what can they not fix?

Whatever you decide here becomes the model you defend in Part 3, and the justification you preregister in Part 4.

Lunch

An hour – the return time is in the chat · lunchtime browsing, if you like: the specification menu

PART 3 – What you can actually do

Part 3 · 90 minutes

Task 2, in the analysts’ own words

Focus on the Cosmopolitan dimension for EU framing and on the unemployment rate as the independent variable; disregard contextual variables.

Three instructions. Everything else – estimator, sample, weighting, functional form, how you build the scale – was left to us.

That is the design working as intended. Fix just enough to make the numbers comparable; leave the rest free, so that what remains is a measurement of analytical variability itself.

What those three instructions assume

  • The cosmopolitan dimension assumes the five items form one scale, and that a mean of them is the right summary. A sum, a share, or a factor score would all be defensible
  • The unemployment rate assumes the claim is carried by unemployment. The paper’s own abstract says poor economic performances – which also names GDP growth
  • Disregard contextual variables removes the country-level controls. That buys comparability and gives up any defence against a time-varying confounder

None of these is wrong. Each one is a decision, made for you, that a different study could reasonably have made differently.

The panel you will work on

EUframes_cy.csv270 rows, one per country-year · 27 countries · 2004–2013

Column What it is
cntry, year the cell
mcosmo mean cosmopolitan framing in that country-year, 0–1
mutil, mcomm, mlib the other three framing scales
mpos, mneg the two composites – all positive items, all negative items
unemp, growth national unemployment rate and GDP growth
n_cy how many survey respondents are behind the cell
bailout 1 if an EU/IMF assistance programme was running

n_cy is the interesting one: cells are built from very different numbers of people, which is why weighting is a live choice on the menu.

You are Analyst no. 6

  • Aim 1: reproduce analyst C6HJR’s constrained result
  • What Multi100 asked: focus on the cosmopolitan dimension and on the unemployment rate; disregard the contextual variables

  • What I submitted: a two-way fixed-effects model, mcosmo ~ unemp_c | cntry + year

  • What is on record: t ≈ −3.80 (df = 233, N = 270, partial r ≈ −0.242)

Hands-on 3: reproduce the recorded result

  • Task 3 in your workspace or in the browser lab – the constrained specification: outcome mcosmo, unemployment centred, country and year fixed effects
  • One dataset, one object: EUframes_cy.csv, loaded as euframes
  • The cell has gaps to fill; there are two hints and a full solution if you want them

Target: t = −3.804 – the value on record for analyst C6HJR. When you have it, type your t in chat and hold it until I say go.

One number, thirty people

Thirty of you just landed on t = −3.804, to the third decimal. That is reproducibility, and it is the easy R.

You all ran the same code. The five Multi100 analysts each wrote their own – and came back with five different answers.

None of them was careless. So what were they doing differently? The menu, next.

The specification menu

Axis Options
Outcome mcosmo · mutil · mcomm · mlib · mpos · mneg
Outcome family gaussian · beta (logit link, for a bounded 0–1 scale)
Predictor unemployment · GDP growth · both
Predictor form raw · log (unemployment only – growth goes negative)
Copredictor none · the other macro measure
Estimator two-way FE · country FE · random effects · pooled with clustered SE
Sample all · 2004–2008 · 2009–2013 · no bailout countries · no Greece and Spain
Weights none · n_cy

Eight axes. Copy-paste code for every option is on the specification menu page.

Two were on the board in the red-card study already: its 29 teams split across four response distributions (outcome family) and four ways of handling non-independence (estimator).

What Task 2 left open

The three instructions fixed the outcome dimension, the predictor and the contextual controls. Everything else stayed open – including one thing nobody thinks of as a choice:

The unit of analysis. Two of the five fitted models on individual respondents; three worked on aggregates. Task 2 never said which, and the wording did not imply one.

Their five answers, all obeying all three instructions, still ran from r = −0.006 to r = −0.45. The remaining freedom was more than enough.

Keep that in mind this afternoon: the choices that move an estimate most are usually the ones made before the model is fitted at all.

One deviation, and a reason for it

For the rest of today the rule is: change one axis, hold the other seven at baseline.

Baseline = the specification you just reproduced: mcosmo · gaussian · unemployment · raw · no copredictor · two-way FE · all years · unweighted.

Why one at a time? Because it makes the class chart readable. Every dot in the room differs from a common baseline in exactly one declared way, so when a dot moves you can say what moved it.

And one convention binds everyone: claim-alignment. The negative framings are reverse-coded, 1 − y, so that supports the claim reads as a negative sign whatever your outcome. Without it the class chart would straddle zero by construction and look like noise.

Six ways to say EU framing

Scale Items Claim predicts
mcosmo peace · democracy · free movement · diversity · voice negative
mutil prosperity · social protection negative
mcomm unemployment · identity · crime · borders positive
mlib bureaucracy · waste of money positive
mpos all seven positive items negative
mneg all six negative items positive

The claim says decrease positive dimensionsmpos is the more literal reading, mcosmo is what was instructed. That one disagreement is worth roughly half the variance in the grid.

Building a scale: mean, sum or share?

Five binary items per respondent. Three defensible summaries:

  • Mean of the five – a proportion between 0 and 1. Everyone contributes equally
  • Sum – a count from 0 to 5. Same information, different scale, different coefficient size
  • Share of all mentions – depends on how many options the respondent picked overall

These are not cosmetic. A share is relative to the respondent’s own total, so somebody who ticked eight boxes contributes differently from somebody who ticked two. Mean and sum are not.

Among the five Multi100 analysts: one used a mean, one a sum, one a share, one three separate shares. All defensible; none the same quantity.

Claim alignment, and why the room needs it

d$mcomm_ca <- 1 - d$mcomm      # reverse-code a negative framing

# now: negative coefficient = supports the claim, whatever the outcome

Reverse-coding a bounded scale flips the sign of b, t and r, and leaves the magnitude, the standard error, the |t| and the p-value untouched. Exactly so for the beta family too, since logit(1 − μ) = −logit(μ).

Nothing about your evidence changes. Only the direction convention it is reported in.

It is still a real analytic decision, which is why it rides on an explicit claim_align switch in your preregistration rather than sitting silently inside the fitting code.

A bounded scale in a linear model

Our outcome is a proportion: it lives strictly between 0 and 1, and its variance is smallest near the ends. A linear model assumes neither of those things.

  • Gaussian – fit it as if it were unbounded. Simple, familiar, and can in principle predict outside 0–1
  • Beta – a logit-link beta regression, built for exactly this shape. Reports a z, not a t

On this claim, the two agree: the beta twin of the baseline gives b(logit) = −0.01399, z = −4.075, r = −0.258, against the gaussian r = −0.242. Same direction, still significant, marginally stronger.

The seminar argument about distributional assumptions is real, and on this claim it moves almost nothing.

Which variable carries the claim?

The abstract says poor economic performances. That names at least two things:

  • Unemployment rate – what Task 2 instructed. Higher is worse
  • GDP growth – what three of the five analysts chose in their free phase. Higher is better, so the claim predicts a positive coefficient on positive framings

Fitted on our panel, the growth mirror of the baseline gives b = +0.00167, t = 1.59 – the right direction, not significant.

Across the whole growth half of the grid: 70.5% in the direction predicted by the claim, but only 46.9% significant. Direction without significance – a different kind of support from the unemployment story.

Predictor form, and a warning

  • Raw unemp – a one-point rise in the rate
  • log(unemp) – a proportional rise. Says the step from 4% to 5% matters more than 20% to 21%

Those two are genuinely different models and can give different answers.

But centring and standardising are not. Subtracting a mean, or dividing by a standard deviation, is an affine change of scale: it moves the coefficient and its standard error by the same factor and leaves t, r and p exactly unchanged.

Which is why the grid holds no separate centred rows. If a robustness check varies only the centring, it is not a robustness check – it is the same analysis, restated.

Four ways to handle countries and years

Estimator What it compares What it assumes
Pooled, clustered SE countries with each other and over time no unobserved country differences
Country FE each country with itself over time country differences are stable
Two-way FE the same, net of EU-wide year shocks shocks hit everyone alike
Random effects a weighted blend of both comparisons country effects uncorrelated with the predictor

These do not answer the same question. Pooled asks whether high-unemployment countries frame the EU differently; fixed effects ask whether a country frames it differently when its own rate rises.

And it shows: of the 41 wrong-sign specifications in the cosmopolitan universe, 33 are pooled.

What fixed effects actually buy you

They do absorb

  • Anything stable about a country over the decade – institutions, history, size, prior EU sentiment
  • Anything common to a year across all countries – the crisis itself, a treaty, a media cycle

They do not touch

  • A confounder that varies within a country over time – a national bailout, a change of government
  • Reverse causation: today’s framing moving tomorrow’s unemployment
  • Measurement error, selection into the sample

This is the same statement as the DAG you drew before lunch. adjustmentSets() on the committed graph returns country and year – exactly what two-way fixed effects absorb. The graph and the baseline model make the same claim.

Copredictor: control, or second exposure?

Put GDP growth in beside unemployment and you have made a choice you should be able to name:

  • As a confounder – a slump raises unemployment and sours sentiment in the same year. Adjusting closes a back-door that your fixed effects cannot reach
  • As a rival exposure – growth is a second measure of the same poor economic performance. Now you are asking which one carries the claim
  • As a mediator – if unemployment works through growth, adjusting removes part of the effect you wanted

Same line of code, three different estimands. And remember the Table 2 fallacy: the growth coefficient in that model is not automatically the causal effect of growth.

Sample: which countries, which years

  • All – 2004–2013, 27 countries, 270 cells
  • 2004–2008 vs 2009–2013 – before and during. Halves your data and asks whether the relationship is a crisis phenomenon
  • No bailout countries – drop the cells where an EU/IMF programme was running
  • No Greece and Spain – drop the two extreme unemployment trajectories

The last two exclusions test whether the association is carried by a handful of extreme cases. If it is, removing them should worry you – and it does something visible here: 24 of the 41 wrong-sign specifications sit in bailout-excluding samples.

An exclusion is a claim about who the finding is meant to be about. That is an estimand question, not a housekeeping one.

Weights: whose country-year counts more?

Each cell is a mean of very different numbers of respondents. Weighting by n_cy says a cell built from 2,000 people should count more than one built from 500.

  • Unweighted – every country-year is one observation. You are studying countries
  • Weighted by n_cy – cells count in proportion to the people behind them. You are closer to studying people

That is a shift of estimand, not a technical refinement. Which is why the day’s convention is to declare it.

Practical note: random effects are never fitted with weights in the grid – the combination is not well defined for this design, so those rows do not exist.

Making eight axes comparable

A coefficient from a beta model on a share and one from a weighted pooled model on a sum are not on the same scale. The answer from Multi100, and ours:

  • Take the test statistic and its degrees of freedom
  • Convert to a partial correlation r – the association between predictor and outcome, net of everything else in the model, on a fixed −1 to +1 scale
  • Apply claim alignment, so the sign means the same thing everywhere

Now every specification in the room, and all 840 in the grid, sit on one axis. report_result() does the conversion for you.

The cost is worth naming: standardising makes numbers comparable, not equivalent. It cannot rescue two analyses that estimate different quantities – it can only put them side by side.

Silberzahn’s teams hit the same wall: four of the twenty-nine reported a correlation or a standardised difference rather than an odds ratio. The authors converted everything to one scale and reported the median. That is a different answer to the same problem – a median is one number, a curve is a distribution – and which you want depends on what you mean to say next.

Hands-on 4: work the menu

Task 4 is exploration, and it is meant to be. Try several specifications, look at what each one does, and get a feel for which axes move the estimate.

  1. Start from the baseline you reproduced
  2. Change one axis, refit, and note what happened to b, t and the partial r
  3. Do it again with a different axis. Three or four is plenty
  4. Then pick the one you can argue for best, and submit it

Your dot lands on the live Multiverse page within seconds. We read the room’s dots at the start of Part 4.

This is not your preregistered analysis. You are allowed to look, change your mind, and try again – that is the whole point of naming this phase separately.

Break

30 minutes – then we read what the room built · leave the Multiverse page open

PART 4 – Reading the multiverse, choosing one path

Part 4 · 90 minutes

You just built a small multiverse

You ran three or four specifications, all defensible, and watched the estimate move.

Nothing about that is unusual. Every analysis you have ever published sat inside a space like that – you simply reported one point from it.

The multiverse idea is not a new estimator. It is the proposal that the space is the finding, and that showing one point without it is a form of incomplete reporting.

Where the idea comes from

  • Multiverse analysis (Steegen et al. 2016) – born from data-processing choices: how you code, exclude and combine. Report the whole set of datasets that your choices could have produced
  • Specification curve analysis (Simonsohn et al. 2020) – the same move for model choices, plus a visual grammar: estimates on top, the choices that produced each one underneath
  • The vibration of effects – the same idea, arrived at independently in epidemiology
  • The book-length treatment (Young and Cumberworth 2025) – computational recipes at scale

Different names, one move: stop treating your analytical choices as invisible, and enumerate them.

The garden, drawn

Four binary choices give sixteen analyses. Eight axes with several options each give hundreds. The published paper reports one red line and no tree.

Justifiable, not merely possible

A multiverse is not every analysis you could run. Two filters have to be applied before anything goes in:

  • Justifiability – could you defend this specification to a reviewer, in writing, before seeing its result?
  • The same estimand – does it estimate the quantity you said you were after?

Drop the first filter and you get a cloud of nonsense that makes any finding look fragile. Drop the second and you are averaging over answers to different questions – which is the red-card mistake, at scale.

Today’s grid is a curated 840, chosen against both filters. That curation is itself an analytical decision, and it is the one that a critic should attack first.

Enumerate, or curate?

Enumerate

Cross every option on every axis. Complete, mechanical, reproducible.

Risk: the set fills up with specifications nobody would defend, and the spread stops meaning anything.

Curate

Include only the specifications you would actually defend.

Risk: you are choosing the set, so a sceptic can ask whether you chose it to get the picture you wanted.

There is no way out of this. Either way, the set is a claim, and it belongs in the paper alongside the curve – ideally declared before the curve is drawn.

Multiplying the menu out

Our menu, multiplied out:

  • 6 outcomes × 2 predictor forms × 3 copredictors × 4 estimators × 5 samples × 2 weightings, minus the combinations that are not defined → 840 distinct analyses
  • Add the outcome family axis → 1,680
  • Add GDP growth as the claim carrier → 2,520
  • Open the unit of analysis as well – person level or aggregate – and the menu arithmetic runs past 5,000

Two consequences. Compute stops being free, so somebody must decide what is worth running. And you cannot look at 2,520 numbers – you need a picture, and a way to ask which axis is doing the work.

What is a multiverse for?

Rohrer, Hullman and Gelman’s question, and it has three different answers that want three differently built universes:

  • Reflection – to understand your own analysis. Which choices matter, where the argument has to happen. Build it broad, and read it yourself
  • Persuasion – to convince a reader that the finding is not an artefact of one path. Build it from specifications that your critic would accept
  • Inference – to make a probabilistic statement about the curve as a whole. Build it to a fixed, declared rule, and bring a null distribution

Today’s grid was built for reflection. Reading it as inference is the mistake we take apart later in this part.

Rohrer et al. (2026)

Which fork matters?

Once you have hundreds of estimates and the choices that produced each one, you can ask a question that no single analysis can answer: how much of the variation does each axis account for?

  • Treat the specification grid as data: one row per analysis, the axes as factors, the estimate as the outcome
  • Decompose the variance across axes

The answer is the most useful thing a multiverse produces, because it tells you where to spend your argument. An axis that moves nothing does not need defending. An axis that moves everything needs a paragraph.

On this claim, the answer surprises almost everybody – which is why you are going to bet on it in a minute before I show you.

Reading a specification curve

Five questions to ask any curve you meet, in any paper:

  1. What is on the vertical axis – a raw coefficient, or something standardised? Are the signs aligned so that up and down mean the same thing throughout?
  2. How was the set built – enumerated or curated, and by whom?
  3. Where does the published estimate sit – at the median, or out at an end?
  4. Is the spread patterned – does one axis produce the extremes, or is it scattered?
  5. What null is the significance shading against – and are the specifications independent? (They almost never are.)

You are about to answer all five for your own curve. That is a publishable robustness section.

The class multiverse

  • 2,520 specifications: the canonical 840 (blue unemployment triangles) plus the outcome-family and predictor axes that the menu opened
  • The canonical 840: 68.7% claim-consistent direction · 63.9% significant · claim-aligned partial r from −0.70 to +0.24

Pre-computed over the specification menu (data/spec_grid_full.csv), claim-aligned – negative framings reverse-coded and growth predictors sign-flipped, as recorded in your Task 5 preregistration (claim_align = TRUE).

Where your result lands

Your orange dot lands here → live on the Multiverse page as the room submits

report_result() gives you a one-click submission link; results stream onto the Multiverse page.

Place your bet

Six axes: outcome · predictor form · estimator · copredictor · sample · weights

Which axis moves the estimates most across all 840 specifications?

Type your guess – hold it – send on go.

The outcome fork carries half the variance

  • Outcome operationalisation: 49.7% of specification variance · estimator + predictor form + copredictor + weights together: under 2%
  • Computed on the grid as estimated – part of the outcome share is the polarity split between positive and negative framings, the split the claim-aligned chart removes from display

Computed over the committed data/spec_grid.csv; imported as data/fork_importance.csv.

The unit of analysis does the work

  • The two person-level analysts sit at the two extremes of the recorded range – the unit of analysis and the metric dominate, not the software or the estimator

Where the sign flips concentrate

  • In the cosmopolitan-only universe, the wrong-sign estimates concentrate in the pooled cluster-robust estimator and the bailout-excluding samples

Find your own percentile

  • Open the your percentile cell in the browser lab – both routes use it for this one; it runs in any browser tab, no install
  • Enter the result you submitted before the break – it computes the share of the 840 specifications more negative than yours, on the claim-aligned scale
  • The same statistic that placed the five analysts, computed for you

Paste your percentile in chat – hold – send on go.

What the curve leaves out

  • The specifications nobody ran. Person-level models, a different scale construction, a lagged predictor. The curve is silent about everything outside the menu, and looks just as complete either way
  • The dependence. All 840 use the same 270 country-years, so they share their noise. The width of the cloud is not 840 independent opinions
  • Estimand drift. Some of these specifications answer subtly different questions. The chart puts them on one axis; that does not make them commensurable
  • The data-building choices. Everything upstream of the panel – how the scales were coded, which waves were pooled – is fixed for everyone here

A curve is a map of the menu you wrote. It cannot tell you the menu was the right one.

Two jobs, one curve

Description – map how the estimate moves across every defensible choice. (What we have been doing all day.)

Inference – ask whether the curve as a whole is surprising if there were no effect at all. (A different, harder job.)

The two are detachable, and most published curves stop at description.

Steegen et al. (2016); Simonsohn et al. (2020)

840 estimates are not 840 tests

Every specification re-uses the same 270 country-years. The estimates are deeply dependent.

63.9% significant describes the curve. It is not a p-value for the claim.

Vote-counting across dependent specifications has no null distribution. You cannot tally the stars.

The permutation design

  • Break the one link that the claim needs: shuffle each country’s unemployment series in time (jointly with growth), keeping everything else intact
  • Refit all 840 specifications on each shuffled dataset · repeat B = 500 times
  • Compare the observed curve to the 500 null curves: is ours more extreme than chance produces?

The design is worked through in your inference reading; permutation inference after Phipson and Smyth (2010).

How to read a curve-level test

  • With B = 500, the smallest reportable p is 2/(B+1) ≈ 0.004 – a floor, not a zero
  • Near p = .05 the Monte-Carlo wobble is about ±0.01, so treat p = .04 and p = .06 as the same finding
  • A DESIGN · NOT YET RUN – a curve-level test has to be fixed in advance, and this one has not been run; today we read the design, not a result

Discussion (2 minutes, then chat): the constrained cosmopolitan-only universe has 140 specs: 70.7% negative, 57.1% negative and significant. Strong evidence? Against what null?

What stands without the test

Everything descriptive survives: the 49.7% outcome fork, the analysts’ placement, the sign-flip pattern, your dot.

What waits for the permutation is one sentence: the curve as a whole is unlikely under no effect.

Careful reporting says which of the two layers a claim lives in.

The multiverse has its critics

  • Robustness is better assessed with a few thoughtful models than with billions of regressions (Auspurg 2025; answering Ganslmeier and Vlandas 2025) – only justified models targeting the same estimand belong in one curve. Today’s grid is a curated 840, not billions, and the fork chart does the job that every critic endorses: it locates where the argument must happen
  • There is only one correct analysis once the question is fully specified (Lakens et al. 2026) – your declared specification answers cherry-picking, and their deeper point is that declaring in advance does not yet make a choice justified, which is what the estimand and DAG work were for
  • What’s a multiverse good for anyway? – reflection, persuasion, or inference, and each purpose wants a differently built universe (Rohrer et al. 2026). Today’s was built for reflection, and how much more to read into it is the open debate

The book-length how-to: Young and Cumberworth (2025); the estimand diagnosis behind Part 2: Auspurg and Brüderl (2021).

There is only one correct analysis

Lakens and colleagues push back hard, and the argument deserves to be taken seriously:

  • A research question, fully specified, implies a quantity to be estimated
  • That quantity, together with the design and the assumptions you are willing to make, implies one analysis that estimates it
  • So a spread of results is not evidence of legitimate diversity. It is evidence that most of those specifications are wrong – they estimate something else
  • On this reading, showing a wide curve is closer to an admission than to a robustness check

Take it back to the red cards. Twenty-nine teams modelled one count of red cards as linear, logistic or Poisson. Those are claims about how the outcome was generated, and they cannot all be right – so on this argument the spread is not diversity, it is mostly error.

Lakens et al. (2026)

Which analysis, and how you would know

Grant the argument. One correct analysis exists. Now: which one is it?

Getting there needs the estimand and the identification assumptions and agreement about the measurement – and on this claim, reasonable people disagree about all three.

So the curve is not a rival to the one-correct-analysis view. It is a map of the remaining disagreement, drawn because we cannot yet settle it.

And it points at what to argue about: on this claim, at the outcome, not at the estimator.

What preregistration settles

It does settle

  • That your specification was not chosen because of the result
  • That your rule for what counts as support was written before you had one
  • That a reader can see what you intended, and what you changed

It does not settle

  • Whether the specification was any good
  • Whether it estimates the quantity that your question implies
  • Whether the assumptions behind it hold

A badly-reasoned analysis declared in advance is still badly reasoned – now with a timestamp

Which is why the estimand and the DAG came first today. Preregistration protects a choice; it is the causal reasoning that makes the choice worth protecting.

You are Analyst no. 6 – for real this time

  • Aim 2: assess the robustness of the constrained result
  • Before the break you explored: several specifications, tried in order to find out what they do. That was the right way to learn them

  • Now you commit: one specification, declared in writing before it runs, with a reason that your DAG can carry

  • The difference is not the code. It is that one of these can be chosen because of its result and the other cannot

Archiving and registering a reanalysis

OSF, the Open Science Framework: a timestamped, citable, public home for data, code, and registrations.

A registration is a frozen, dated snapshot of a plan. Anyone can check what you said you would do, and when.

You need no account today. We are going to read OSF the way that reviewers do.

Three addresses, three jobs

osf.io/8rtwe – the archival Multi100 record: analyst C6HJR, frozen as submitted

osf.io/6zqct – the maintained fork: the script, the codebook, today’s data

osf.io/h7432 – the full SCORE dossier on the EU-frames case: every verdict, every artefact

Scavenger beat: open the dossier

  • Find the secondary-data replication of the claim – new Eurobarometer data, N = 352,114
  • Read its unemployment coefficient: the sign, the p-value
  • Paste the b in chat when you have it (3 minutes)

The study you are inside was preregistered

  • Multi100: 100 claims × at least 5 independent analysts, no instructions on how to analyse
  • The scoring rules – same conclusion? effect size close? – were fixed and public before any result came back
  • Outcome: 74% same conclusion · 34% within ±0.05 d – judged against criteria nobody could bend afterwards

Aczel et al. (2026). Project: osf.io/7snkz

Two templates, one choice

  • Today: the AsPredicted short form – eight plain-language questions (on OSF: Preregistration Template from AsPredicted.org)
  • The fuller tool for our job: the secondary-data template (van den Akker et al. 2021, Meta-Psychology) – 26 items, built for reanalysing data you already hold
  • Both are in your reading; today’s block is the short form, adapted for data-in-hand

Eight questions, your answers already mapped

AsPredicted asks Your block answers
Q1 Have data been collected? It’s complicated – I already hold them + what I already know
Q2 Main question / hypothesis? the claim, in your one sentence
Q3 Dependent variable? your outcome axis
Q4 Conditions? your other five axes, held at baseline
Q5 Analyses? your model + your support/undermine rule
Q6 Outliers / exclusions? your sample axis
Q7 Sample size? N = 270 country-years – fixed
Q8 Anything else? your DAG-grounded justification

THE PREREGISTRATION MOMENT

  1. Open the preregistration block in your report
  2. Complete fields 1–5 – claim · data-in-hand · your one specification · DAG sentence · support/undermine rule
  3. Render (Route 1) or save the completed block into your script (Route 2)
  4. Type declared in chat – then wait for the room

Nobody runs a model until the room is declared.

Run it – then watch the chart

  • Task 5 – run your declared specification, and only that one
  • report_result() prints a one-click link with your numbers already in it – click, submit
  • Your dot joins the Multiverse page as a declared result, beside the exploratory ones from before the break

Watch what the two waves look like together. If the declared cloud sits systematically to one side of the exploratory cloud, that is worth a conversation.

Stuck? Chat STUCK · reload and rerun · second browser · follow my screen · last resort: the grid lookup on the menu page gives you your numbers by hand

Two sentences worth keeping

the common single-path analyses in social and behavioural research should not be simply assumed to be robust to alternative analyses (Aczel et al. 2026)

A blueprint: precise question → causal reasoning (DAG) → multiverse → sensitivity analysis (Auspurg and Brüderl 2021)

  • Same-conclusion robustness and same-number robustness are different bars
  • Justifiability is the criterion – choose with reasons, show the spread

A single analytical path is a choice

There is no privileged single analytical path – only justifiable ones, explored openly.

You made one justifiable choice today – transparently, with a reason, inside a multiverse.

And that was the workshop.

codemoreh.github.io/applied-replication

Fork anything. Everything is open.

What you take with you

  • Yours now: a preregistered, rendered report · an estimand and a causal graph you wrote yourself · two dots on the class curve, one explored and one declared
  • Read: the three Nature 2026 papers (Miske et al. 2026; Aczel et al. 2026; Tyner et al. 2026) · the glossary (Nosek et al. 2025) · Auspurg and Brüderl (2021) · Lundberg et al. (2021) · Lakens et al. (2026) · Teney (2016)
  • Explore: the nine companion modules (intro to R · Quarto documents · the statistics behind today · causal graphs · the multiverse debate · simulation · one claim at two levels · repositories & preregistration · git & GitHub) · the Multi100 data · fixest & ggdag docs
  • Keep going locally: the research-software and open-research teams at your university · a ReproducibiliTea journal club near you
  • Everything is on the workshop site, and everything is open – fork any of it

References

Aczel, Balazs, Barnabas Szaszi, Harry T. Clelland, et al. 2026. “Investigating the Analytical Robustness of the Social and Behavioural Sciences.” Nature 652 (8108): 135–42. https://doi.org/10.1038/s41586-025-09844-9.
Alipourfard, Nazanin, Beatrix Arendt, Daniel M Benjamin, et al. 2021. Systematizing Confidence in Open Research and Evidence (SCORE). 46mnb_v1. SocArXiv. https://doi.org/10.31235/osf.io/46mnb.
Auspurg, Katrin. 2025. “Robustness Is Better Assessed with a Few Thoughtful Models Than with Billions of Regressions.” Proceedings of the National Academy of Sciences 122 (43): e2521917122. https://doi.org/10.1073/pnas.2521917122.
Auspurg, Katrin, and Josef Brüderl. 2021. “Has the Credibility of the Social Sciences Been Credibly Destroyed? Reanalyzing the Many Analysts, One Data Set Project.” Socius 7 (January): 23780231211024421. https://doi.org/10.1177/23780231211024421.
Breznau, Nate, Eike Mark Rinke, Alexander Wuttke, et al. 2022. “Observing Many Researchers Using the Same Data and Hypothesis Reveals a Hidden Universe of Uncertainty.” Proceedings of the National Academy of Sciences 119 (44): e2203150119. https://doi.org/10.1073/pnas.2203150119.
Cole, Stephen R., Robert W. Platt, Enrique F. Schisterman, et al. 2010. “Illustrating Bias Due to Conditioning on a Collider.” International Journal of Epidemiology 39 (2): 417–20. https://doi.org/10.1093/ije/dyp334.
D’Agostino McGowan, Lucy, Travis Gerke, and Malcolm Barrett. 2024. “Causal Inference Is Not Just a Statistics Problem.” Journal of Statistics and Data Science Education 32 (2): 150–55. https://doi.org/10.1080/26939169.2023.2276446.
Ganslmeier, Michael, and Tim Vlandas. 2025. “Estimating the Extent and Sources of Model Uncertainty in Political Science.” Proceedings of the National Academy of Sciences 122 (25): e2414926122. https://doi.org/10.1073/pnas.2414926122.
Gelman, Andrew, and Eric Loken. 2014. “The Statistical Crisis in Science.” American Scientist 102 (6): 460–66. https://doi.org/10.1511/2014.111.460.
Griffith, Gareth J., Tim T. Morris, Matthew J. Tudball, et al. 2020. “Collider Bias Undermines Our Understanding of COVID-19 Disease Risk and Severity.” Nature Communications 11 (1): 5749. https://doi.org/10.1038/s41467-020-19478-2.
Hernán, Miguel A., and James M. Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC.
Huntington-Klein, Nick. 2021. The Effect: An Introduction to Research Design and Causality. Chapman & Hall/CRC. https://theeffectbook.net/.
Ioannidis, John P. A. 2005. “Why Most Published Research Findings Are False.” PLoS Medicine 2 (8): e124. https://doi.org/10.1371/journal.pmed.0020124.
John, Leslie K., George Loewenstein, and Drazen Prelec. 2012. “Measuring the Prevalence of Questionable Research Practices with Incentives for Truth Telling.” Psychological Science 23 (5): 524–32. https://doi.org/10.1177/0956797611430953.
Keele, Luke, Randolph T. Stevenson, and Felix Elwert. 2020. “The Causal Interpretation of Estimated Associations in Regression Models.” Political Science Research and Methods 8 (1): 1–13. https://doi.org/10.1017/psrm.2019.31.
Lakens, Daniel, Sajedeh Rasti, and Mehmet N. Tunç. 2026. There Is Only One Correct Analysis. PsyArXiv. https://osf.io/preprints/psyarxiv/hvjxf_v1/.
Leszczensky, Lars, and Tobias Wolbring. 2022. “How to Deal with Reverse Causality Using Panel Data? Recommendations for Researchers Based on a Simulation Study.” Sociological Methods & Research 51 (2): 837–65. https://doi.org/10.1177/0049124119882473.
Lundberg, Ian, Rebecca Johnson, and Brandon M. Stewart. 2021. “What Is Your Estimand? Defining the Target Quantity Connects Statistical Evidence to Theory.” American Sociological Review 86 (3): 532–65. https://doi.org/10.1177/00031224211004187.
Miske, Olivia, Anna Lou Abatayo, Mason Daley, et al. 2026. “Investigating the Reproducibility of the Social and Behavioural Sciences.” Nature 652 (8108): 126–34. https://doi.org/10.1038/s41586-026-10203-5.
Munafò, Marcus R., Brian A. Nosek, Dorothy V. M. Bishop, et al. 2017. “A Manifesto for Reproducible Science.” Nature Human Behaviour 1 (1): 0021. https://doi.org/10.1038/s41562-016-0021.
Nosek, Brian A, Timothy M Errington, Noah Haber, Theresa Stankov, and Andrew H Tyner. 2025. A Brief Glossary of Terms about Repeatability: Replicability, Robustness, and Reproducibility. mqfp4_v1. MetaArXiv. https://doi.org/10.31222/osf.io/mqfp4_v1.
Open Science Collaboration. 2015. “Estimating the Reproducibility of Psychological Science.” Science 349 (6251): aac4716. https://doi.org/10.1126/science.aac4716.
Phipson, Belinda, and Gordon K. Smyth. 2010. “Permutation P-values Should Never Be Zero: Calculating Exact P-values When Permutations Are Randomly Drawn.” Statistical Applications in Genetics and Molecular Biology 9 (1): Article 39. https://doi.org/10.2202/1544-6115.1585.
Rohrer, Julia M., Jessica Hullman, and Andrew Gelman. 2026. What’s a Multiverse Good for Anyway? PsyArXiv. https://osf.io/preprints/psyarxiv/37g29_v1/.
Sánchez-Tójar, Alfredo, Jelte M. Wicherts, and Robb Willer. 2026. “Huge Meta-Research Project Puts Claims in Social-Science Papers to the Test.” Nature 652 (8108): 39–41. https://doi.org/10.1038/d41586-026-00805-4.
Silberzahn, R., E. L. Uhlmann, D. P. Martin, et al. 2018. “Many Analysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results.” Advances in Methods and Practices in Psychological Science 1 (3): 337–56. https://doi.org/10.1177/2515245917747646.
Simmons, Joseph P., Leif D. Nelson, and Uri Simonsohn. 2011. “False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.” Psychological Science 22 (11): 1359–66. https://doi.org/10.1177/0956797611417632.
Simonsohn, Uri, Joseph P. Simmons, and Leif D. Nelson. 2020. “Specification Curve Analysis.” Nature Human Behaviour 4 (11): 1208–14. https://doi.org/10.1038/s41562-020-0912-z.
Steegen, Sara, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel. 2016. “Increasing Transparency Through a Multiverse Analysis.” Perspectives on Psychological Science 11 (5): 702–12. https://doi.org/10.1177/1745691616658637.
Teney, Céline. 2016. “Does the EU Economic Crisis Undermine Subjective Europeanization? Assessing the Dynamics of CitizensEU Framing Between 2004 and 2013.” European Sociological Review 32 (5): 619–33. https://doi.org/10.1093/esr/jcw008.
Tyner, Andrew H., Anna Lou Abatayo, Mason Daley, et al. 2026. “Investigating the Replicability of the Social and Behavioural Sciences.” Nature 652 (8108): 143–50. https://doi.org/10.1038/s41586-025-10078-y.
Young, Cristobal, and Erin Cumberworth. 2025. Multiverse Analysis: Computational Methods for Robust Results. Analytical Methods for Social Research. Cambridge University Press. https://doi.org/10.1017/9781009003391.

Thank you – and your feedback

Feedback form
(link in the chat · two minutes, while the chart is still on screen)