The multiverse: origins, uses, and critics

Where specification-curve thinking came from, what it is for, and why its critics are worth reading

This is the multiverse module of the companion curriculum, and it is optional. The workshop teaches you to run a multiverse analysis: declare one specification, run it, and read your result against 2,520 pre-computed alternatives. This module is about the idea itself – where it came from, what it is claimed to do, and the arguments now running against it. It sits alongside the inference reading, which covers one narrow question, namely what a curve-level test can and cannot establish. Here the question is broader. The method you practised on the workshop day is under live debate in 2025–26, and the objections raised against it are strong ones. The aim of this module is that you can take a position in that debate. Throughout, the workshop’s own materials serve as the worked example – sometimes as an answer to a criticism, and sometimes, just as instructively, as the thing that the criticism lands on.

1. One idea, three strands

The founding statement is a 2016 paper by Sara Steegen, Francis Tuerlinckx, Andrew Gelman and Wolf Vanpaemel (Steegen et al. 2016). Their starting point is data construction rather than modelling: turning raw data into an analysable dataset involves choices – which respondents to exclude, how to code a variable, where to cut a scale – and several options are usually reasonable. Instead of privileging one arbitrarily assembled dataset, they propose constructing every dataset that a reasonable set of choices could have produced and running the analysis on all of them, so that the reader sees how much the conclusion owes to decisions nobody would defend as the uniquely correct ones. The idea connects directly to Gelman and Loken’s ‘garden of forking paths’ (Gelman and Loken 2014). A researcher who makes each choice once, on the fly, has walked one path through a branching garden, and the single p-value at the end carries no record of the branches. The closing sentence of the 2016 paper sets the ambition plainly: “it should become standard practice to go beyond a single data set analysis and to acknowledge the multiverse of statistical results”, followed immediately by a concession that the critics would later press hard on – “performing a multiverse analysis will often be difficult, and to a large extent subjective” (Steegen et al. 2016, 710).

Two things about the original paper matter for everything that follows. First, its multiverse was customised for critique: the alternative datasets were choices that the original authors of the target study “might themselves have considered”, drawn from their own prior publications. That is a post-mortem on the flexibility of one team, not a general survey of how the question could be answered. That distinction matters again in section 2. Second, Steegen and colleagues already foresaw the extension that this workshop’s grid instantiates: alongside the data multiverse they sketch a model multiverse, in which “One specific analysis thus corresponds to a single sample from a model multiverse” (Steegen et al. 2016, 710), and they place the whole idea in an older lineage running through perturbation and sensitivity analysis in economics and Bayesian statistics. The multiverse was never as new as its name; what was new was the demand that the whole distribution be shown.

The modelling side of the idea arrived under its own flag as specification curve analysis (Simonsohn et al. 2020), which enumerates the defensible ways of specifying the test, estimates all of them, plots the estimates in order and, in their proposal, tests the curve as a whole against a permutation-built null. This is the visual grammar that the workshop’s Multiverse page uses. A distinct strand kept the analysis fixed and varied the analysts instead: 29 teams answering one question about football referees with odds ratios spanning 0.89 to 2.93 (Silberzahn et al. 2018); 73 teams testing one hypothesis about immigration and social policy, with more than 95% of the variance in their results left unexplained even after coding every identifiable decision in the workflow of every team (Breznau et al. 2022); and the Multi100 project, which put at least five independent reanalysts on one claim from each of 100 published studies (Aczel et al. 2026). The EU-frames case you worked on is one of those hundred, and the five analysts whose results you saw in Part 1 are its reanalysts. A further strand proposes Bayesian hierarchical modelling as the way to aggregate such team-science data, extended into a multiverse over inclusion criteria and priors (Hoogeveen et al. 2020). The workshop’s design deliberately stacks these strands: the specification grid is a model multiverse, the five Multi100 analysts and your own class chart are a many-analysts layer on top of it, and the two are meant to be read against each other.

2. Three purposes that pull apart

Rohrer, Hullman and Gelman ask the organising question of the current debate, which is what a multiverse is actually for. Their answer distinguishes three purposes – “as a tool for reflection and critique, as a persuasive tool, as a serious inferential tool” (Rohrer et al. 2026, 1) – and attaches a characteristic failure to each: “it fails as a persuasive tool when researchers disagree about which variations should be included in the analysis”, and “it fails as a serious inferential tool when the included analyses do not target a coherent estimand” (Rohrer et al. 2026, 1). Their conclusion is deliberately mild. The multiverse “does remain a valuable tool”, but they “urge against taking it too seriously” (Rohrer et al. 2026, 1).

What matters in their paper is less the taxonomy itself than what it implies for construction: the three purposes impose incompatible building rules on the universe of specifications. A multiverse built for critique may legitimately mix analyses that answer slightly different questions, because documenting that researchers plausibly disagree about the question is the point – “it may be perfectly sensible to include analyses with different estimands as long as it is plausible that researchers may try out different estimands” (Rohrer et al. 2026, 6). The same set then cannot be read as inference about one quantity, precisely because it no longer targets one quantity. Hullman’s companion blog post states the conditional in one line: “what kinds of equivalence are needed will also depend on your goal in using a multiverse” (Hullman 2026). There is no such thing as the correctly built multiverse; there is a correctly built multiverse for a purpose.

Hold the workshop’s own day against that standard and the mapping comes out uneven. The eight-axis specification menu and the fork-importance chart are reflection in exactly Rohrer and colleagues’ sense: they exist to make visible how much of the outcome depends on decisions that look arbitrary from outside. The live class chart is not persuasion, despite appearances – section 6 returns to what it actually is. And the curve-level test designed in the inference reading belongs to the inference purpose, which is where the estimand-coherence requirement bites hardest on our own grid. One universe of 840 distinct analyses serves all three roles across the day, and the incompatibility argument says that is a tension, not a convenience. Naming that tension in front of participants is part of what the day teaches. If you can say which purpose each artefact of the workshop day serves and which building rule that purpose imposes, you can read the rest of this module, and the papers it cites, as a participant in the argument rather than an onlooker.

TipTry it

Take the three purposes and assign each artefact of the workshop day to one of them: the specification menu, your preregistration block, the 2,520-point curve, the fork-importance bars, the class chart, the permutation test designed in your reading. For each assignment, write one sentence on what the universe of specifications would have to look like for that purpose to be legitimately served. Notice where the same 840 analyses are being asked to do two different jobs.

3. Prune first, then estimate

The most direct empirical disagreement in the recent literature is between Ganslmeier and Vlandas’s model-uncertainty study and Katrin Auspurg’s reply to it. Ganslmeier and Vlandas cross five specification dimensions at once – control set, fixed-effect structure, standard-error type, sample, and outcome operationalisation – across four political-science literatures, producing roughly 3.6 billion estimates, and find that “the sign and statistical significance of estimates is more contingent on the underlying sample and the operationalization of indicators than on the inclusion of the control set” (Ganslmeier and Vlandas 2025, 6). They are alert to the obvious objection, stating the dilemma themselves: “Running all possible models risks introducing errors in the empirical approach, but selecting a subset of robustness checks risks reintroducing the ‘researcher’s degree of freedom’” (Ganslmeier and Vlandas 2025, 2).

The title of Auspurg’s reply reads like a verdict, promising that robustness is better assessed with a few thoughtful models than with billions of regressions. What the paper delivers is a precondition: “Multiverse analyses are only informative when all included models represent equally plausible strategies for estimating the same estimand” (Auspurg 2025, 1). Her charge is not that 3.6 billion is too many models. It is that the set is heterogeneous – in statistical quality and in what each model estimates – so its distribution is a pile of incommensurable numbers. She names four kinds of nonequivalence: statistically inferior models (“Many models do not include fixed effects or cluster-robust SE, which are standard in panel data analyses” (Auspurg 2025, 1)); unstable country subsets; causal misspecification from randomly combined covariates; and measurement nonequivalence, where “Outcomes differ in scope … and are sometimes negatively correlated. This clearly indicates that they do not reflect the same construct” (Auspurg 2025, 1). Restricting Ganslmeier and Vlandas’s universe to models she classes as justified (1,152 of them, against 91,008 unjustified) turns a 60/40 split among significant coefficients into 97% pointing the same way. And she closes the escape route of filtering by model fit: “Ironically, the authors’ own fit criterion (AIC) consistently favors unjustified models” (Auspurg 2025, 1). The remedy therefore comes before estimation, in what she calls “thorough, theory-based model selection” (Auspurg 2025, 2).

The constructive version of the same demand is in Young and Cumberworth’s book – the current how-to manual for the method, and worth reading whole if this module interests you (Young and Cumberworth 2025). Their rule is estimand consistency: “When an estimand is defined a priori by a set of controls, the multiverse model set must always include those controls” (Young and Cumberworth 2025, 56), and “any ‘necessary control’ requires a clear causal diagram or definitional justification supported by prior research” (Young and Cumberworth 2025, 57). That last clause should sound familiar. It is the order of the workshop day itself, where the DAG in Part 2 comes before the declaration in Part 4. Theirs is a book about robustness, not truth: “there is nothing in the distribution of estimates per se that tells us which is the true effect” (Young and Cumberworth 2025, 119). The analyst still owes the reader a defended preferred estimate.

Some of Auspurg’s charge lands on the workshop’s own grid, and some of it is answered by construction. The grid is curated rather than enumerated – 840 distinct analyses, with forms that provably change nothing (centred and standardised predictors) removed outright rather than padding the count, and every axis carried with a note on where it came from. But be precise about what that answers. Deduplication removes redundancy, whereas Auspurg’s objection is about heterogeneity, and a perfectly deduplicated universe can still mix models she would prune.

Her first category names the pooled specifications, exactly the cells of our menu that fit without country and year adjustment, and her fourth describes the outcome axis, where six framing scales of different scope, some negatively correlated with the others, all stand in for ‘EU framing’. The claim-alignment convention answers the polarity half of that, reverse-coding the negative framings so that support for the claim always points the same way. But reverse-coding does not make a communitarian-framing scale the same construct as a cosmopolitanism scale. There is a complication worth holding alongside her objection, and the simulation module lets you inspect it directly.

The six scales are shares of one instrument, constrained to sum to one, so whatever raises one must lower the others, and the coefficients estimated on the panel itself show how unevenly that plays out. Rising unemployment moves cosmopolitan and utilitarian framing one way and communitarian framing the other, roughly three times as far, while libertarian framing barely stirs; on growth the pattern rearranges again. An analyst choosing mutil and one choosing mcomm are therefore estimating quite different quantities from a single instrument, before any question of construct validity arises.

That is neither a rescue of the outcome fork from her objection nor quite the same charge she is making. Part of the spread is the arithmetic of a compositional measure and part is the nonequivalence she names, and separating the two is work that the curve alone will not do for you. The provenance notes on the menu, meanwhile, are disclosures, and disclosure is not the justification she demands. The defensible summary is that the workshop’s grid sits between the positions on purpose: curated enough to escape the ‘billions of regressions’ objection, heterogeneous enough that her nonequivalence categories give you real work to do on it. That work is what the next two sections are about.

TipTry it

Pick one axis of the specification menu and write Auspurg’s case for pruning it. Which of her four nonequivalence categories applies, and what would the distribution across the grid look like if the axis were fixed at its best-justified option? Then write the case for keeping it, in Rohrer, Hullman and Gelman’s terms: which purpose does varying this axis serve? The estimator axis and the outcome axis are the two where both cases are strong.

4. ‘There is only one correct analysis’

Lakens, Rasti and Tunç push the criticism one level deeper than Auspurg. Their target is the inference that many people draw from analytic variability, namely that since reasonable analyses disagree, there is no single correct analysis, and researchers should report distributions instead of committing. Their reply depends on the distinction between two kinds of uncertainty, which they insist must not be merged: aleatoric uncertainty is random variation, the kind that statistics quantifies with standard errors; epistemic uncertainty is uncertainty about which auxiliary assumptions are true – which measure is valid, which units belong in the sample, which adjustment set closes the back-doors. Epistemic uncertainty is reduced by doing research, not by widening intervals to absorb it. Once the auxiliary assumptions are fixed, “there is a fact of the matter about which analysis correctly implements that test” (Lakens et al. 2026, 7). Hence the title. Robustness across specifications of unequal justification earns nothing, because robustness “is a consequence of valid tests, and not a goal in itself” (Lakens et al. 2026, 15). And one of their warnings applies verbatim to teaching exercises: “Variability across specifications based on uninformed hunches, especially when less plausible alternatives are included, provides little actionable insight” (Lakens et al. 2026, 10).

It is tempting to answer them with the workshop’s preregistration moment, since you declared exactly one specification before running anything, which sounds like their prescription. Resist the temptation, because the two devices solve different problems. Pre-declaration solves a timing problem: it rules out choosing the specification once you have seen where the numbers fall, and that answers a genuine worry. Rohrer and colleagues note that “in the absence of preregistration, researchers may as well determine what is equivalent based on whether or not it supports their result” (Rohrer et al. 2026, 5). Their own argument is about warrant, and warrant is indifferent to timing. An arbitrary choice declared in advance is still arbitrary. What their argument actually demands is the groundwork that makes a choice non-arbitrary – validated measures, established adjustment sets, empirical evidence about the auxiliaries – and the workshop’s gesture in that direction is the estimand-and-DAG work of Part 2, which is groundwork in miniature rather than groundwork discharged. A participant choosing one cell from an eight-axis menu after ninety minutes of instruction is, on their reading, still choosing on an uninformed hunch. That is a limit on what the exercise demonstrates rather than an indictment of it, and it is useful to know exactly where the limit sits.

Two further caveats belong here, both cutting against the workshop’s own device. Preregistering a reanalysis of data you have already seen is itself a compromised form of it, since “[w]hen the data already exist, authors can explore the data prior to registering their analysis – making it unclear what if anything was truly preregistered” (Young and Cumberworth 2025, 149). That is why the workshop points you at the secondary-data template in the repositories module rather than pretending the fresh-data form fits. And the two devices can be combined rather than opposed: “one could preregister a multiverse analysis to combine the strengths of both approaches” (Heyman and Vanpaemel 2022, 9). Preregister the whole universe rather than a single point in it. Fixing the whole garden in advance and only then walking it is the strongest available reply to the charge that multiverse analysis is flexibility laundered.

5. The fork chart locates the argument

The workshop’s fork-importance chart says that the outcome axis carries 49.7% of the specification variance in the grid as estimated, with the estimator, predictor-form, co-predictor and weight axes together under 2%. On the claim-aligned scale the facilitator’s extended decomposition of the same grid inverts the ranking: the estimator fork rises to roughly 58%, the outcome fork falls to roughly 11–12%. How numbers like these are put to the critics matters, and the temptation to over-claim runs in a specific direction.

The over-claim would be: ‘our dispersion is mostly estimator dispersion, not estimand dispersion, so Auspurg’s diagnosis does not apply to us’. The critics’ reply writes itself, because the estimator fork is an estimand fork in their sense. A pooled specification with clustered standard errors estimates a blend of two associations – countries with higher unemployment framing the EU differently, and countries framing differently as their own unemployment moves over time – while a two-way fixed-effects specification isolates the second alone. Those are different quantities, about which substantive theory can reasonably differ; the statistical methods module works through exactly this decomposition arithmetically. Auspurg’s first nonequivalence category names this axis, and on her reading a large estimator share is evidence for pruning, not evidence that pruning is unnecessary. The claim-aligned recomputation does not escape this either, since alignment is a display transform that removes polarity, not scope – the caveat that the workshop attaches to the 49.7% figure cuts both ways.

What the decomposition genuinely supports is the modest reading: it locates where the argument has to happen. That use is endorsed by every party to the debate. Lakens and colleagues “support the use of multimodel analyses to hypothesize about ‘which analytical decisions are most consequential’” (Lakens et al. 2026, 6); Ganslmeier and Vlandas offer their method as something that “can help inform which combinations of choices are more or less warranted” (Ganslmeier and Vlandas 2025, 6); Steegen and colleagues promised pointers to the most consequential choices from the start. Read this way, the fork chart tells you which of the axes in the grid needs a theory, rather than defending the grid. The outcome fork tells you to decide what ‘EU framing’ means. The estimator fork, once aligned, tells you to decide whether your question is within-country or between-country. The chart makes neither decision; it shows where each has to be made.

6. Frequency is not probability

The most practical lesson in the critical literature concerns how to read a specification curve, and it can be stated as one prohibition: do not treat frequency as probability. The share of specifications supporting a claim – 68.7% in the workshop’s canonical grid – is a description of a curated set of analyses, not the probability that the claim is true. Reading it as a vote invites what Rohrer and colleagues call an illusion. Their prescription is that “variation should be treated possibilistically”, meaning that any result occurring anywhere in a justified universe is a live possibility, and that reading carries its own condition: “This possibilistic interpretation is only sensible if every universe included is indeed deemed justified” (Rohrer et al. 2026, 6). The Milliways project makes the same point operationally, showing how standard displays emphasise how often an outcome occurs and thereby invite “misleading, probabilistic conclusions” (Sarma et al. 2024). Lakens and colleagues state the bluntest version: “Scientists should not compute the percentage of agreement across different data-analytic strategies, because we do not resolve scientific disagreement by majority vote” (Lakens et al. 2026, 16).

When the workshop quotes 68.7%, the number is doing descriptive work – it summarises the curve you are looking at – and the inference reading exists precisely to police the boundary between that description and any claim about evidence. Note what this implies for the permutation test designed in your reading: a permutation p-value is a probabilistic device applied to an object that the current literature says should be read possibilistically. The reading explains what the sharp null of the test actually licenses. The contribution of this module is the warning label on everything that the test does not license.

There is also a quieter consequence for persuasion. Rohrer and colleagues observe that a multiverse persuades only under a precondition: “This can only work if the results are fairly unambiguous across specifications” (Rohrer et al. 2026, 3). At 68.7% claim-direction and 63.9% significant, the EU-frames grid fails that precondition – which is a finding rather than an embarrassment. A reader shown this curve should walk away knowing which decisions decide the claim, not persuaded that it is settled. If you ever build a multiverse to defend a result of your own, this is the precondition to check before you lean on it.

And then there is the class chart. On the workshop day it is easy to read the live chart – everyone’s declared specification, landing as dots on one curve – as the day’s persuasive set piece. The critical literature supplies the better label: it is a many-analysts demonstration, the classroom analogue of the 29 teams and the 73 teams, with the extra property that every analyst’s choice was preregistered. Heyman and Vanpaemel, who run multiverse projects with students, name the design exactly. It is “comparable to the many-analysts-one-dataset approach used by Silberzahn et al. (2018)” (Heyman and Vanpaemel 2022, 7), and when several analysts each carry their own analytical choices, “there is a multiverse of multiverse analyses” (Heyman and Vanpaemel 2022, 7). A room of analysts disagreeing is not a robustness verdict in either direction; it is a live demonstration that the dispersion documented by the many-analysts literature is real, reproducible, and – once the menu and the preregistrations are on the table – explainable. That is reflection staged socially, and it is also the answer to anyone who asks why the class chart is allowed to straddle zero.

TipTry it

Open the Multiverse page and find your own dot (or the baseline, if you did not submit). First read its neighbourhood possibilistically: what is the full range of aligned estimates among specifications that differ from yours on exactly one axis? Then catch yourself the moment you start counting – ‘most of my neighbours agree with me’ – and write one sentence on what that count would and would not establish, given how this universe was built.

7. The multiverse in the classroom

One thing on which every side of this debate agrees is that the multiverse earns its keep in the classroom. Rohrer, Hullman and Gelman, in the middle of urging everyone not to take it too seriously, concede without reservation that “the multiverse has been a valuable training tool” (Rohrer et al. 2026, 8). Heyman and Vanpaemel build their whole teaching design on it, and their two rules of classroom practice are visible everywhere in this workshop’s construction. The first is motivation over count, since “the goal should not be to merely devise as many paths as possible” and “[t]he key is that the alternatives are properly motivated” (Heyman and Vanpaemel 2022, 4). That is why the workshop offers a menu of justified axes rather than an open invitation to vary anything.

The second rule is pastoral, and worth quoting because it is easy to forget: “one should be cautious that students do not completely lose faith in (psychological) science” (Heyman and Vanpaemel 2022, 9). A day spent watching one published claim scatter across 2,520 estimates can curdle into cynicism, and the day is structured to work against that. The open OSF trail of this case is a record anyone can check, and the preregistration moment is the practice that makes your own future record equally checkable. Taught properly, the multiverse shows that analysis is not arbitrary: the decisions are identifiable, defensible and the analyst’s own.

Hullman adds a forward-looking reason that this literacy will matter more rather than less. Multiverse construction is becoming cheap, because “lots of forms of robustness checking that used to be difficult are now trivially easy” (Hullman 2026) once a generative-AI agent can build the universe from a paper. When anyone can produce a specification curve on demand, the scarce skill stops being computation and becomes exactly what this module has been about: judging what belongs in the universe, naming the purpose that the universe serves, and refusing the readings that its construction cannot support. The critics, read carefully, have been teaching that skill throughout.

Sources and further reading

The 2025–26 debate read here: Ganslmeier and Vlandas (2025) and the reply by Auspurg (2025); Rohrer et al. (2026) with Hullman’s blog companion (Hullman 2026); and Lakens et al. (2026). The origins: Steegen et al. (2016), Gelman and Loken (2014), Simonsohn et al. (2020); a systematic framework for judging which analytic decisions are equivalent is in Del Giudice and Gangestad (2021). The many-analysts strand: Silberzahn et al. (2018), Breznau et al. (2022), Aczel et al. (2026), and, for the estimand diagnosis that Part 2 of the workshop builds on, Auspurg and Brüderl (2021). The book-length how-to is Young and Cumberworth (2025); the tooling is Sarma et al. (2023) and Sarma et al. (2024); the classroom design is Heyman and Vanpaemel (2022). Quotations from the two PsyArXiv preprints are cited by the pagination of the preprints themselves.

References

Aczel, Balazs, Barnabas Szaszi, Harry T. Clelland, et al. 2026. “Investigating the Analytical Robustness of the Social and Behavioural Sciences.” Nature 652 (8108): 135–42. https://doi.org/10.1038/s41586-025-09844-9.
Auspurg, Katrin. 2025. “Robustness Is Better Assessed with a Few Thoughtful Models Than with Billions of Regressions.” Proceedings of the National Academy of Sciences 122 (43): e2521917122. https://doi.org/10.1073/pnas.2521917122.
Auspurg, Katrin, and Josef Brüderl. 2021. “Has the Credibility of the Social Sciences Been Credibly Destroyed? Reanalyzing the Many Analysts, One Data Set Project.” Socius 7 (January): 23780231211024421. https://doi.org/10.1177/23780231211024421.
Breznau, Nate, Eike Mark Rinke, Alexander Wuttke, et al. 2022. “Observing Many Researchers Using the Same Data and Hypothesis Reveals a Hidden Universe of Uncertainty.” Proceedings of the National Academy of Sciences 119 (44): e2203150119. https://doi.org/10.1073/pnas.2203150119.
Del Giudice, Marco, and Steven W. Gangestad. 2021. “A Traveler’s Guide to the Multiverse: Promises, Pitfalls, and a Framework for the Evaluation of Analytic Decisions.” Advances in Methods and Practices in Psychological Science 4 (1): 2515245920954925. https://doi.org/10.1177/2515245920954925.
Ganslmeier, Michael, and Tim Vlandas. 2025. “Estimating the Extent and Sources of Model Uncertainty in Political Science.” Proceedings of the National Academy of Sciences 122 (25): e2414926122. https://doi.org/10.1073/pnas.2414926122.
Gelman, Andrew, and Eric Loken. 2014. “The Statistical Crisis in Science.” American Scientist 102 (6): 460–66. https://doi.org/10.1511/2014.111.460.
Heyman, Tom, and Wolf Vanpaemel. 2022. “Multiverse Analyses in the Classroom.” Meta-Psychology 6 (December). https://doi.org/10.15626/MP.2020.2718.
Hoogeveen, Suzanne, Sophie W. Berkhout, Quentin F. Gronau, Eric-Jan Wagenmakers, and Julia M. Haaf. 2020. Improving Statistical Analysis in Team Science: The Case of a Bayesian Multiverse of Many Labs 4. PsyArXiv. https://doi.org/10.31234/osf.io/cb9er.
Hullman, Jessica. 2026. What a Multiverse Good for Anyway? https://statmodeling.stat.columbia.edu/2026/02/12/what-a-multiverse-good-for-anyway/.
Lakens, Daniel, Sajedeh Rasti, and Mehmet N. Tunç. 2026. There Is Only One Correct Analysis. PsyArXiv. https://osf.io/preprints/psyarxiv/hvjxf_v1/.
Rohrer, Julia M., Jessica Hullman, and Andrew Gelman. 2026. What’s a Multiverse Good for Anyway? PsyArXiv. https://osf.io/preprints/psyarxiv/37g29_v1/.
Sarma, Abhraneel, Kyle Hwang, Jessica Hullman, and Matthew Kay. 2024. “Milliways: Taming Multiverses Through Principled Evaluation of Data Analysis Paths.” Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI), 1–15. https://doi.org/10.1145/3613904.3642375.
Sarma, Abhraneel, Alex Kale, Michael Jongho Moon, et al. 2023. “Multiverse: Multiplexing Alternative Data Analyses in R Notebooks.” Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany), 1–15. https://doi.org/10.1145/3544548.3580726.
Silberzahn, R., E. L. Uhlmann, D. P. Martin, et al. 2018. “Many Analysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results.” Advances in Methods and Practices in Psychological Science 1 (3): 337–56. https://doi.org/10.1177/2515245917747646.
Simonsohn, Uri, Joseph P. Simmons, and Leif D. Nelson. 2020. “Specification Curve Analysis.” Nature Human Behaviour 4 (11): 1208–14. https://doi.org/10.1038/s41562-020-0912-z.
Steegen, Sara, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel. 2016. “Increasing Transparency Through a Multiverse Analysis.” Perspectives on Psychological Science 11 (5): 702–12. https://doi.org/10.1177/1745691616658637.
Young, Cristobal, and Erin Cumberworth. 2025. Multiverse Analysis: Computational Methods for Robust Results. Analytical Methods for Social Research. Cambridge University Press. https://doi.org/10.1017/9781009003391.