Reading a specification curve: what it can and cannot say

Pre-session reading for the inference module

The multiverse in this workshop is built on the aggregated grid of 840 country-year specifications that sits behind the class chart (data/spec_grid.csv), together with its beta-family doubling to the family universe of 1,680 described on the specification menu. The same menu logic extends naturally to the individual respondents themselves: a person-level universe of 3,840 specifications, and a joint universe of 5,520 uniting the two levels, can be laid out on paper before any of it is fitted, and neither has been computed. This reading explains the inferential logic that governs a design like that: why a multiverse is reported descriptively, why a single formal test can nonetheless be attached to it, why such a test belongs on the aggregated universe alone, and how its result should be read. The reasoning is elementary once laid out, but it is rarely laid out, and the vocabulary of ‘multiverse’ and ‘specification curve’ is young enough that the two are routinely conflated. Replication work asks for precision about what each act of repetition does and does not establish, and the same precision is owed to the tools used to describe it.

Two jobs, one curve

A multiverse analysis, in the sense given to the term by Steegen and colleagues, is a descriptive exercise: identify every defensible way of turning raw data into a test of the claim, run all of them, and display the resulting distribution instead of a single hand-picked estimate (Steegen et al. 2016). Nothing in that founding formulation involves a joint hypothesis test. The founding paper reports grids of estimates and p-values and asks the reader to see how conclusions shift across the grid; its contribution is transparency about analytical contingency. The inferential layer arrived later and separately, under the name ‘specification curve analysis’, when Simonsohn, Simmons and Nelson supplemented the descriptive curve with a formal question – could a curve like this have arisen from data in which the effect does not exist? – and a resampling method for answering it (Simonsohn et al. 2020). The two layers are detachable by design, and the largest robustness exercise yet conducted in the social sciences detached them deliberately. Multi100 reported the dispersion of its 504 reanalyses in explicitly exploratory terms, with no inferential statistics anywhere in the paper (Aczel et al. 2026).

This workshop follows that precedent for everything you have seen: the grids, the fork decomposition and the placement of the five analysts are descriptive throughout. A single permutation-based test can nonetheless supplement them – fixed in its design before it runs, and confined to the aggregated universe alone. The remainder of this reading explains each half of that sentence.

Why 840 estimates are not 840 tests

The temptation that the descriptive curve creates is vote counting. If 63.9% of specifications are significant, and 5% is what chance would produce, does the curve not speak for itself? It does not, and the reason is dependence. The 840 results are not 840 draws of evidence; they are the same 270 country-years, distilled from the same 416,698 respondents, interrogated 840 times with slightly different instruments. If the panel happens, by sampling accident alone, to contain a negative-looking alignment between unemployment and cosmopolitan framing, then most specifications will detect that same accident together. The curve moves as a body, because every point in it shares the data. Under a true null the share of significant specifications therefore has no known distribution: it could sit at 10% or at 70%, and no formula exists to say which, because the correlation structure across hundreds of overlapping specifications is analytically intractable (Simonsohn et al. 2020).

Every individual cell of the multiverse still carries its own conventional inference – an estimate, a standard error, a confidence interval – and each of those is as interpretable as it would be in a stand-alone paper. What is missing is the level above: a warranted statement about the curve as a whole. The gap matters because analytical flexibility is exactly the mechanism by which false positives are manufactured. Simmons, Nelson and Simonsohn’s demonstration that undisclosed flexibility “allows presenting anything as significant” is the same arithmetic read in the opposite direction (Simmons et al. 2011). A multiverse inverts the disclosure problem, putting every fork on display rather than hiding it in a garden of unreported paths (Gelman and Loken 2014), but display alone does not convert a correlated ensemble into a body of evidence with known error rates.

What the permutation supplies

The solution that Simonsohn and colleagues adapted is among the oldest ideas in statistics: if the sampling distribution of a quantity cannot be derived, manufacture it, by rearranging the data in a way that embodies the null hypothesis (Fisher 1935; Ernst 2004). Here the null hypothesis is that national unemployment and EU framing are unrelated, and the rearrangement that embodies it is a within-country permutation of the macroeconomic series. Each country keeps its own ten-year unemployment history, its own outcome data, its own place in the panel, but the assignment of unemployment values to years is shuffled – jointly with GDP growth, so that the two macro series keep their mutual correlation – and the alignment between economic conditions and framing responses is thereby destroyed. On each such rearrangement the canonical aggregated universe of 840 Gaussian specifications is refitted, and the summary statistics of the curve are recorded: the median partial correlation, the share of specifications significant in the claim-consistent direction, and an aggregate z following Stouffer’s rule. Five hundred rearrangements give five hundred versions of what this multiverse looks like when nothing is going on. The observed curve is set beside them, and its rank in that set is the p-value.

One property of the construction answers the dependence problem of the previous section: every null draw refits the same 840 specifications on the same panel, so every null curve carries exactly the correlation structure that made the observed curve untestable by analytical means. The dependence is not assumed away; it is reproduced inside the reference distribution. This is why curve-level permutation has become the standard inferential companion of large specification exercises, including well-known applied examples (Orben and Przybylski 2019; Rohrer et al. 2017).

The test also has limits. The permutation destroys temporal alignment, and with it any serial structure in the predictor; the null it instantiates is ‘no association at any alignment’, which a sceptic could note is subtly stronger than ‘no causal effect’. Two co-trending series can reject such a null without any causal connection, which is why the fixed-effects cells of the universe – which absorb common year shocks – matter for interpretation, and why the verdict of the test is a statement about alignment beyond chance, never about causation. The estimand discipline of Part 2 is not suspended by a small p-value.

Which level carries the identifying variation

The design confines this test to the aggregated universe, and the confinement is an argument rather than an economy. The predictor in this study varies only at the country-year level: all 416,698 respondents observed in a given country and year share one value of the unemployment rate. All the information that the data contain about the alignment of unemployment with EU framing is therefore carried by 270 cell-level pairings, in every specification, at every level of analysis. Refitting the 3,840 person-level specifications under each of 500 permutations would interrogate precisely the same 270-cell alignment – the person-level rows are, from the point of view of the null hypothesis, an elaborate disaggregation of the same evidence – at several hundred times the computational cost, and with the additional numerical fragilities of maximum-likelihood mixed models fitted 1.9 million times. A person-level universe would earn its place descriptively. It would show how the magnitude of standardised effects, and their conversion to a common metric, depend on the unit of analysis, which is the day’s recurring lesson. It would have essentially nothing further to contribute to the question ‘is the alignment distinguishable from chance?’, because it contains no additional identifying variation with which to answer it.

This is also the reply to an objection that a careful reviewer should raise: that the original study modelled individuals, so a test confined to country-year aggregates excludes the model family that the original study itself used. The objection conflates the two layers. Model families are fully represented where model families matter – in the descriptive universes, where the random-effects specifications carry their own cells and their attenuated estimates tell their own story. The permutation layer does not compare model families; it asks one question about one alignment, and it asks it at the level where that alignment is observed.

Five hundred draws, and how to read them

The design fixes five hundred permutations in advance, and the consequences for interpretation are mechanical. A permutation p-value is computed with the observed statistic included in its own reference distribution – Phipson and Smyth’s correction, which guarantees that an estimated p can never be exactly zero (Phipson and Smyth 2010) – so with B = 500 the smallest attainable two-sided value is 2/(B + 1), roughly 0.004. Monte Carlo error follows the binomial rule: an estimated p near 0.05 carries a standard error of about √(0.05 × 0.95/500) ≈ 0.01. Both facts prescribe the reading. A reported p of 0.03 or 0.08 should be treated as ‘near the conventional threshold, with one point of Monte Carlo noise either way’, not as a sharp verdict. The choice of B trades that noise against computation in a way that has been formalised for the bootstrap and applies here unchanged (Davidson and MacKinnon 2000). Had finer resolution been needed, a larger B would be fixed in advance. For a supplementary test whose role is calibration rather than discovery, this resolution is sufficient, and its cost, hours of computation against days for a person-level equivalent, is proportionate.

What stands without the test

It is a fair question, finally, what the day’s conclusions would lose if the test were never run. Deliberately, they would lose nothing. The decomposition showing that the outcome fork carries half the specification variance while estimators, predictor forms, copredictors and weights together carry under 2%; the placement of the five Multi100 analysts at their respective percentiles of the curve; the reading of the recorded dispersion as a question multiverse rather than an analysis multiverse (Auspurg and Brüderl 2021; Lundberg et al. 2021) – all of these are claims about the composition of analytical variation, and they are established by enumeration and decomposition, not by a null distribution. The distinction between same-conclusion and same-number robustness is likewise descriptive vocabulary (Aczel et al. 2026). The permutation test answers one narrow supplementary question: whether a curve this claim-consistent could plausibly arise when nothing is going on. Fixing that test in advance, where its identifying variation lives, at a resolution stated before anything runs and cheap to reproduce, adds a calibration that the descriptive layers cannot supply, while the day’s conclusions are built to stand on the layers that need no simulation to be understood (Steegen et al. 2016; Del Giudice and Gangestad 2021).

References

Aczel, Balazs, Barnabas Szaszi, Harry T. Clelland, et al. 2026. “Investigating the Analytical Robustness of the Social and Behavioural Sciences.” Nature 652 (8108): 135–42. https://doi.org/10.1038/s41586-025-09844-9.
Auspurg, Katrin, and Josef Brüderl. 2021. “Has the Credibility of the Social Sciences Been Credibly Destroyed? Reanalyzing the Many Analysts, One Data Set Project.” Socius 7 (January): 23780231211024421. https://doi.org/10.1177/23780231211024421.
Davidson, Russell, and James G. MacKinnon. 2000. “Bootstrap Tests: How Many Bootstraps?” Econometric Reviews 19 (1): 55–68. https://doi.org/10.1080/07474930008800459.
Del Giudice, Marco, and Steven W. Gangestad. 2021. “A Traveler’s Guide to the Multiverse: Promises, Pitfalls, and a Framework for the Evaluation of Analytic Decisions.” Advances in Methods and Practices in Psychological Science 4 (1): 2515245920954925. https://doi.org/10.1177/2515245920954925.
Ernst, Michael D. 2004. “Permutation Methods: A Basis for Exact Inference.” Statistical Science 19 (4): 676–85. https://doi.org/10.1214/088342304000000396.
Fisher, R. A. 1935. The Design of Experiments. Oliver; Boyd.
Gelman, Andrew, and Eric Loken. 2014. “The Statistical Crisis in Science.” American Scientist 102 (6): 460–66. https://doi.org/10.1511/2014.111.460.
Lundberg, Ian, Rebecca Johnson, and Brandon M. Stewart. 2021. “What Is Your Estimand? Defining the Target Quantity Connects Statistical Evidence to Theory.” American Sociological Review 86 (3): 532–65. https://doi.org/10.1177/00031224211004187.
Orben, Amy, and Andrew K. Przybylski. 2019. “The Association Between Adolescent Well-being and Digital Technology Use.” Nature Human Behaviour 3 (2): 173–82. https://doi.org/10.1038/s41562-018-0506-1.
Phipson, Belinda, and Gordon K. Smyth. 2010. “Permutation P-values Should Never Be Zero: Calculating Exact P-values When Permutations Are Randomly Drawn.” Statistical Applications in Genetics and Molecular Biology 9 (1): Article 39. https://doi.org/10.2202/1544-6115.1585.
Rohrer, Julia M., Boris Egloff, and Stefan C. Schmukle. 2017. “Probing Birth-Order Effects on Narrow Traits Using Specification-Curve Analysis.” Psychological Science 28 (12): 1821–32. https://doi.org/10.1177/0956797617723726.
Simmons, Joseph P., Leif D. Nelson, and Uri Simonsohn. 2011. “False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.” Psychological Science 22 (11): 1359–66. https://doi.org/10.1177/0956797611417632.
Simonsohn, Uri, Joseph P. Simmons, and Leif D. Nelson. 2020. “Specification Curve Analysis.” Nature Human Behaviour 4 (11): 1208–14. https://doi.org/10.1038/s41562-020-0912-z.
Steegen, Sara, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel. 2016. “Increasing Transparency Through a Multiverse Analysis.” Perspectives on Psychological Science 11 (5): 702–12. https://doi.org/10.1177/1745691616658637.