Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

Not Wrong: The Replication Crisis is a Metatheory Crisis?

Wanna know what that is? Its theory.

We talk about reproducibility as a methodological problem. We debate p-hacking, publication bias, and analytical forking paths. But the elephant just chillin in the corner is that our theories are weak. There are too many other plausible explanations for what we observe and measure. Sure, the things we measure are complex, but without more clarity, even Commisioner Odenthal isn’t gonna find us a clue to what we are testing. Meaning that we don’t really know what to make of our findings. Other than hoping they are publishable and packaging them neatly to increase those chances. The findings are not right or wrong, they are not really interpretable. At least not with weak theory.

Part of the problem is that unlike the ‘harder’ sciences, consciousness is in the equation. We are measuring phenomena far more complex than quantum physics. Another part is simply that we do not spend much time on theory. In one of my favorite esoteric and less-mainstream open science writings, Anne Scheel pointed out “Why Most Psychological Research Findings Are Not Even Wrong”.  Anne, I see your discipline specific arguments and raise you all social and behavioral science disciplines. We often don’t know what the estimand is. We don’t know what we are looking at. If our theories don’t explain the phenomenon in the first place, debating whether an empirical finding is statistically “right” or “wrong” is moot.

To confront this, I pitched a method I’m working on with my former doctoral student turned postdoc Hung H.V. Nguyen. It’s called Metatheoretical Multiverse Analysis (MMA) and I wanna kick some ‘theoric’ with it (cool huh? It rhymes with lyric and could mean theory in action…). What follows is a summary of the ensuing debate, full of interdisciplinary friction, economic curmudgeonism, and philosophy of science moments.

The Pitch: Metatheoretical Multiverse Analysis (MMA)

To understand why metatheoretical multiverse analysis is necessary, we have to look at how we currently treat plausible theoretical arguments about the data-generating model… Say what? I mean how we use and apply logic.

Imagine you are testing the effect of X1 on Y. But there is an unobserved confounder, X2, which causes both X1 and Y. If X2 is not in your test, your results are uninterpretable. You don’t know if X2 is confounding the results. And, don’t even get me started about X3. This might be a collider. But the theories are not clear in your subfield. So when we try to compare them we end up with several conflicts or unknown paths. These are all alternatively plausible theories and from them we have a multiverse of theory.

Now, imagine I have five equally plausible alternative theories explaining the same phenomenon. This means there’s a 20% chance the theory I am using is the correct one. But I cannot imagine any social or behavioral science where there are only five plausible theories. There are likely thousands. We only test five because it takes entire careers to write semantic theory. And since it takes 20-40 years for most of our discipline to forget those long-winded theories of the past maybe this is a sinking ship…. I digress. Let me re-gress instead (fancied opposite of digress which happens to be a statistically procedure too). 

This is where Metatheoretical Multiverse Analysis (MMA) comes punching (or wrestling or kicking) in. If a standard multiverse analysis runs all reasonable empirical specifications for a given dataset, a metatheoretical multiverse analysis would logically compare all these theories. Somehow…. That’s where the idea gets a little sticky. Metatheory is mostly something people write about, rather than try to formalize and analyze with math. But if they could, our method would then show them where they need to invest their theory-building efforts.

I opened the floor. The timer started.

The Economist’s Dilemma: HARKing and the Illusion of Theory

The pushback was immediate, insightful, and brutally honest. The first counterargument highlighted just how difficult theoretical work is. As one applied economist noted, reading James Heckman makes you realize how brilliant deep theory can be, but spending your career trying to come up with sufficient conditions to definitively disentangle two competing theories is a great way to find yourself out of academia before you get tenure. Remember the long-winded argument?

But a more damning critique came from the reality of how theory is actually utilized in modern economics. To commit a “statistical sin” and assign causality, one researcher pointed out that the lack of robustness in our fields might stem from the fact that HARKing (Hypothesizing After Results are Known) is practically the norm.

In economics, the workflow rarely starts with a pristine a priori theory. Instead, a researcher finds an intriguing empirical pattern in the data, and then builds a formal mathematical model to justify the findings. We only take the time to do the exhausting math if we already have a paper we want to publish. Often, the formal model is something requested by Reviewer 2 at a top journal ex-post. Because the theory is engineered to fit the data, it doesn’t improve our prior hypothesis in a meaningful way, nor does it guarantee robustness when exposed to new data.

Interestingly, someone from the I4R team chimed in with data to back this up. After checking roughly 15,000 robustness checks in the I4R database, they ran an AI classification to score papers from 0 to 100 on ‘how economic’ they were (i.e., true economic theory vs. a paper on TV habits published in an econ journal). The finding? There was absolutely no relationship between the “econ-ness” of the paper (the presence of formal modeling) and its robustness.

So, the means that even if we were to invest in theory development, we wouldn’t get anywhere because… basically, we suck at it.

The AI Revolution

These days there’s always gotta be something about AI. If not everything. The egomaniac AI that I asked to help me write this just couldn’t wait to point out how many people were talking about it.

Funnily we came to an AI discussion through a strong defense of the structural approach in economics (after the initial dust had settled and we could again breathe the cool Barcelona air-conditioned summer air). Economics, one participant noted, used to be purely theoretical because, prior to the 1970s, we simply didn’t have the capacity to handle large data. The empirical revolution brought an era of data-mining into play. P-hacking came into full swing. The academic industrial complex had taken over thanks to secondary data spewing forth from society.

But the tide now is actually flowing back toward structural modeling. Thanks to… AI? The massive amounts of data we have today, combined with AI’s unparalleled ability to explore patterns, is killing the human comparative advantage in pure empirical data mining. AI will always be better at finding patterns (at least an AI or data scientist tells us this, Gemini still can’t perform better regressions than I can IMO). The only place human researchers will retain a comparative advantage is in thinking, being creative, and structuring the data generating process. We don’t need one perfect theory to describe reality anymore (not that that ever worked for us) good approximations are good enough now – and good approximations means…. Drum roll and someone on the mic saying “Ya’ll ready for this?!”. PREDICTION. The better we can predict things, the less we will need to explain them, they will just become some sort of facts in our life worlds. Won’t they?

Metatheoretical analysis might just be the structural framework we need to survive the AI transition. I mean, I invented it, so I’d like to think so.

Bias, DAGs, and Kung-Fu Nancy Cartwright

As the 90-second buzzer kept interrupting and resetting, the conversation evolved into the philosophy and sociology of science itself.

From a legal and equality perspective, an important question was raised about those “five theories” we tend to rely on (you know that when equally plausible reduce the chances that any one is correct down to 20%). Where do they come from? Historically, they have been generated by a very specific demographic—predominantly white men from the Global North. By restricting our empirical tests to a handful of established theories, we inadvertently perpetuate biases and ignore alternative paradigms that might emerge from the Global South. A metatheoretical multiverse approach, by automatically generating and considering thousands of models, might offer a mechanical antidote to this historical bias by forcing us to acknowledge the vast space of un-theorized realities.

But are DAGs really capable of saving us? The room had its doubts. Hey Mister Jack… I’m talking to you. The world, as one researcher passionately argued, is not a clean DAG with X1, X2, and Y. It has X15, Y7562, and a zillion unobserved mediators. Furthermore, DAGs can’t handle cyclical relationships and feedback loops the most famous relationship in economics, price and quantity, is entirely cyclical, endogenous.

This brought us to Nancy Cartwright and the philosophy of science. If we are mapping out thousands of theories, we must remember the Popperian ideal: a model must be falsifiable. If a theoretical model cannot be thrown out under certain conditions, it ceases to be a model and becomes a religion. Metatheoretical multiverse analysis is only useful if we have the empirical tools to actually falsify the branches of the multiverse we generate.

The crow cheers with Nancy’s MMA kick to the face.

The Preregistration Battleground

You cannot talk about theory and open science without stumbling into the debate on preregistration. I posed the question: ‘Does preregistration inappropriately constrain our theories?’ because I want to seem smart and provocative. If there are hundreds of theories, forcing a researcher to write down a specific one beforehand essentially chokes the multiverse before it can breathe. Choking is forbidden in MMA by the way.

The responses were polarized:

Some argued that preregistration forces a hypothetico-deductive model onto fields that generate knowledge inductively. In economic history or sociology, research is exploratory. You learn from the data. Preregistration actively harms inductive discovery.

Others pointed out that preregistration is incredibly difficult if not overrated for secondary data analysis. Datasets are messy, collected for non-research purposes, and require deep exploration just to understand how missing variables are coded. Recently someone pointed out on LinkedIn that my praise of Neumark is overrated. An MMA sweep kick, totally permissible in the sport – touché.

Conversely, a third group (there’s always a third group otherwise it feels incomplete) advocated that preregistration is simply a record of where you started. It prevents the ex-post invention of stories and increases transparency. If ‘Mostly Harmless Econometrics’ had a chapter telling students to pre-specify their hypotheses before opening the dataset, it would have transformed the culture of economics entirely. To late, the historical institutionalists won.

Where Do We Go From Here?

As we wrapped up the session (and prepared to face the blistering heat outside in the name of finding Fideuà), the consensus was clear: our methodological tools have far outpaced our theoretical foundations. We have built incredibly sophisticated empirical engines, but we are putting them in theoretical chassis that are fundamentally flawed. I liked this shift. Oh wait I kinda led the discussion there.

Metatheoretical multiverse analysis is not a magic bullet. It will not solve the fact that social science involves the unpredictable chaos of human consciousness, nor will it easily map the cyclical, non-DAG-friendly feedback loops of the global economy. But it is a start. It is a way to stop pretending that our opportunistic, post-hoc theories are the only valid models of reality. It might trigger some epistemological soul searching if nothing else. By mapping the vast space of plausible theories and systematically testing where they align and where they conflict, we can begin to rebuild the credibility of our disciplines from the ground up – in theory (which is probably weak, so take it with salt).

Furthermore, as the discussion highlighted, we need to completely overhaul how we incentivize work. Open science has an image problem we often make it look boring, framing it as an adversarial compliance checklist rather than a thrilling pursuit of truth. We leave it up to early-career researchers to awkwardly teach their supervisors about open code and data. And shoulder them with the onus of spending hours making all their work open and reproducible, hours that their forefathers (yeah, they were 95% men) didn’t need to do. Heck these forefathers could publish like 2 papers per year and get tenure.

But since its my blog post, I get the last word: If we want to fix the replication crisis, we cannot just mandate better code and policies. We have to foster better theory. We need to acknowledge the sheer size of the theoretical multiverse, embrace the complexity, and start theorizing before regressionizing. In this case, sadly, the answer is not as simple as 42.

Meta-constructing social theory

Certain hypotheses are constantly tested in social science. The impact of income inequality on health, racial bias on police brutality and public opinion on elections, just to name a few. At some point more tests of the same hypothesis stop contributing to scientific knowledge, and may even harm it by introducing more ‘noise’ into the scientific discourse.

I study social policy preferences and the impact immigration has on them. In this area there has been sustained efforts to test the hypothesis that immigration has a negative impact on support for policies of the welfare state; things related to protecting against risks of aging, unemployment and health. To justify this hypothesis, scholars construct theoretical variations of group dynamics arguments, often drawing on resource competition, nationalism and social identity. Despite claiming to test the hypothesis, the formal models applied to data suggest any number of data-generating processes. They often have little in common other than some measure of immigration and some measure of policy preferences. The results of their tests go in all directions, i.e., a positive, negative or nil effect of immigration. It would appear that the topic is at a standstill, new analyses of the same handful of cross-national survey data sink in the mire. How to break through such a scientific impasse?

In designing the Crowdsourced Replication Initiative (CRI) with co-PIs, Alexander Wuttke and Eike Mark Rinke, we asked researchers to to do research; and we gave them semi-structured tasks and observed them. Specifically they were supposed to come up with the best possible way to test the immigration hypothesis given the same International Social Survey Data source. Although we are currently meta-analyzing the hypothesis test-results (see our virtual APSA poster) to determine which modelling decisions impact the outcomes, we also have a second goal in mind: to discover what is behind the specification curve.

Each research team had to design a best possible test. This is at once a statistical question and a theoretical question. They needed to think carefully about the data-generating process and attempt to recover it in a model. We asked them to write down their research designs after doing this thought exercise, but before analyzing any data. From their researcher choices we can identify where key consensus and disagreements exist about the data-generating model, thus is not only evident in their designs but also in a structured deliberation and voting procedure. This process offers a major advantage over ‘normal’ theoretical discussion and debate among academics, because we have the results that go along with the different modeling choices; and, let’s be honest, when else do over 150 researchers get together and focus on a single hypothesis? By observing this process we can identify where data-generating theories differ and how important these differences are for the results. This will allow us to map where immigration and social policy scholars should focus their theoretical efforts in the future to reduce the most uncertainty, i.e., the largest gains in knowledge.

We have a sound piece of scientific research from Brady and Finnigan (2014) from which we draw our working hypothesis for the CRI crowdsourced researchers: That immigration undermines support for social policies. Brady and Finnigan found little or no support of this hypothesis, at least not in a generalizable macro-comparative sense. This was the launching point for the research of the 77 teams who by now managed to submit replicable results (yes there are still a few out there we are hoping will submit a final model or fix issues we identified in our replication of their models).

Although we are in the process of analyzing the ocean of data generated by this project; a sneak preview offers exciting evidence of the possibility for meta-construction of theory.

Here are two glimpses of what’s to come. One are the deliberation and voting results summarized (Figure 1). The other are differences in definitions of ‘immigration’ (Table 1). We used Kialo, an online structured deliberation platform, to allow participants to discuss the data-generating model after they proposed their own ideas for how to best test the hypothesis. Readers can observe how this deliberation unfolded as we divided the participants into two groups: here and here. Later (after they had the possibility to update their models based on the deliberation) they were given other teams’ models or our own variations on those models to vote on and rank in terms of their appropriateness for testing the hypothesis without having seen the results of those models. Figure 1 quantifies both the Kialo veracity scoring and survey-based voting into one overall scale and then plots the average score of models by their features. Each different color is a discrete set of model features with the zero (y-axis) set to the average support of models choosing an OLS estimator (among the least preferred).

Figure 1. Researcher Preferences for Recovering the Data-Generating Model
“Model” is the hypothesized general impact of immigration on support for social policy. Data and code still being prepared for online sharing, stay tuned.

In Figure 1, it becomes clear looking at the longest bars in each color category that models that incorporate all 5 waves of the ISSP data, include countries of Eastern Europe, include heterogeneous error variation by country-year and year (like a cross-classified model), and incorporate survey sampling weights are preferred over the others. Some of this runs counter to the state of the art. For example, most research follows a logic that major immigrant destination societies – the “Rich 13” and “Rich 17” advanced democracies – should be where “public opinion is likely most influential for the politics of social policy” (Brady and Finnigan 2014:24).

To summarize the motivation for looking across all possible countries, especially Eastern Europe, one crowdsourced researcher put it like this: “Either there is an effect of ‘immigration stock (increase)’ or not“.

Another followed up on this point stating: “To test the general hypothesis we should use as many countries as available and account for variations in GDP and social welfare expenditures in the models.”

These comments demonstrate the majority voice in the CRI that if immigration has a an impact on social policy preferences we should see it across all countries of the globe, not restricting our analysis to only very rich, strong welfare states.

Although Brady and Finnigan and all other research in this area comes to no consensus on whether there is a negative impact of immigration on support for social policy preferences, we should remain skeptical of results if we do not trust the data-generating model. In other words, if our tests do not match what most researchers see as the appropriate theoretical perspective, results are inconclusive and thus uninformative. The deliberation and voting offer us clues where to focus theoretical effort, namely specifying why more countries of the world should (or should not) show a causal effect of immigration on social policy preferences and whether this should (or should not) appear across several decades or only certain times. I am not aware of extensive theory that attempts to tackle these issues. Now is the time to write it!

Even more productive for the possibility of meta-construction of theory is the correspondence between the actual decisions made by the researchers and the subjective and objective outcomes of those decisions. Again, our results are in progress, but we offer a snapshot in Table 1 of different ways the researchers chose to measure immigration as their main hypothesis test variable (1 out of dozens of model decisions to compare). In the first row, 67 out of 77 teams used a “Stock of Foreign-Born” measure in at least one of their models, and 27% of their models using the “Stock” variable showed support of immigration having a negative and significant statistical impact on support for social policy at p<0.05.

Table 1. Crowdsourced Researcher Decisions, Deliberations and Results.
Five different measurement strategies for the immigration test variable.

In the column ‘Positive Test Result Rate’, we see that the ‘Difference’ between “Stock” models (referenced as [1] in Table 1) and those instead using “Flow” to measure immigration models (referenced as [2]) is 3.6. In other words, “Stock” models arrive at support of the hypothesis 3.6 percentage points more than “Flow” models, all else equal. “Stock” models were not more or less popular than “Flow” models, with the average vote score of 0.43 on a scale of 0 (worst) to 1 (best equipped to test the hypothesis) versus 0.45 for “Flow”.

The values in bold indicate that “Change in Flow” models (those measuring derivatives of “Flow”) were among the most popular in the voting process. So the rate of change of the flow of immigrants is seen as an important component in testing this hypothesis. Interestingly, these models were 4 percentage points more likely than “Stock” and “Flow” models to support the hypothesis. When measuring immigration as specific to certain outgroups (from Muslim-majority countries, non-Western countries or refugees), the “Flow” of these various ‘Outgroups’ was seen as more popular than “Stock” of ‘Outgroups’ by a large margin, but the results were over 10 percentage points less supportive of the hypothesis.

What can we learn from this. We argue that a full analysis of the massive range of modeling decisions will give us a guide to move this entire research area forward. Some other decisions for example were different social policy domains, whether ethnic and fractionalization is the ‘real’ cause of the ‘immigration’ effect, construction of latent social policy preference measures, whether or not GDP and unemployment are part of the data-generating assumptions just to name a few out of hundreds. We are only scratching the surface here, but it seems that observing researchers make research decisions, deliberating them, voting and making final choices, we will gain immense knowledge as to where better theory is necessary. As such we see meta-constructing of social theory as a promising avenue for social science. This would be the concept of theory designed replication writ large.