Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

The invisible hand in science: How ideological priors shape empirical outcomes

The Experiment and Our Findings

In empirical social science, we are accustomed to substantial variation in reported outcomes. Research on the minimum wage’s impact on unemployment or immigration’s impact on voting for far-right parties, for example comes to a wide range of results. Metascience identifies several culprits for this dispersion: publication bias, confirmation bias, questionable research practices, and simply noise in the research process. However, one crucial factor has remained difficult to isolate experimentally: the political ideology of the scientists themselves.

In a study published in Science Advances, George J. Borjas and I exploited data from a unique crowdsourced experiment where 71 research teams analyzed the same data and hypothesis. The original study was designed by myself, Eike Mark Rinke, and Alexander Wuttke. The data given to the teams was from the International Social Survey Program spanning 1985-2016 across countries on all continents of the world. The hypothesis was: Does immigration reduce public support for social welfare programs?  The teams estimated 1,253 alternative regression models, producing average marginal effects (AMEs) that ranged from strongly negative to strongly positive. Before conducting any data analysis, researchers in the teams answered a survey about their own immigration policy preferences, placed on a 7-point scale from making immigration laws “tougher” or more “relaxed”.

We found that research teams composed of pro-immigration researchers, those with higher scores on the immigration question, estimated more positive impacts of immigration on public support for social programs. In other words, they were more likely to find evidence that immigration supports social cohesion. Conversely, anti-immigration teams estimated more negative impacts, suggesting immigration reduces social cohesion. The raw differences in the teams’ reported AMEs were relatively small, but the differences in the distributions were stark, especially in the tails: anti-immigration teams were significantly more likely to report extreme negative and significant AMEs.

How did these differences arise? Our analysis revealed that research design is endogenous and serves as the primary mechanism through which ideology enters the parameter estimation process. There are no other ways possible as we checked that no results were mistakes or otherwise ‘faked’. Combinations of research design choices along five key dimensions accounted for a large portion of the differences in actual AMEs between the two extreme team types. These critical decisions were: (1) whether to aggregate the ISSP public attitude responses into a single composite dependent variable; (2) whether to measure immigration as a stock or a flow; (3) whether to employ multilevel modeling; (4) whether to use data for all available countries; and (5) whether to include the 2016 ISSP wave or not.

Furthermore, double-blind and random peer review scores of each team’s model specification showed that pro- and anti-immigration teams’ designs received significantly lower referee scores than moderate teams. This suggests that peers in the experiment recognized that highly ideological teams adopted “unconventional” or outlier specifications that were less suited to testing the hypothesis, yet yielded results consistent with their pre-existing ideological priors.

Why This Matters: Ideology as Scientific Uncertainty

These findings have profound implications for the credibility of empirical research. In science, we are trained to treat research as a systematic process of understanding, analyzing, and reducing uncertainty. In quantitative analyses, we carefully report statistical uncertainty, such as standard errors and confidence intervals, and occasionally explore methodological uncertainty through robustness checks. What we have identified, however, is a different form of uncertainty: the ideological priors of the researchers themselves.

If researchers subconsciously or consciously navigate the “garden of forking paths” to find specifications that confirm their ideological priors, science faces a challenge to its reliability and this risks that public trust in science will (further) erode. This concern is magnified in highly polarized policy-relevant areas where scientific consensus routinely collides with political ideology, such as climate change, vaccination, or immigration. In an era where populist rhetoric frequently questions the legitimacy of scientific progress, painting it as self-serving rather than socially useful, documenting that political beliefs bias empirical parameters is deeply concerning beyond the scientific arena.

Fortunately, because we can identify uncertainty arising from ideological bias, we can also design systemic barriers to reduce its impact. First, the adoption of open science practices is critical. Pre-registration of analytical specifications forces researchers to commit to their research design before observing the data, preventing checking many specifications until an appealing narrative emerges. Second, we should foster “adversarial collaborations”, where researchers with differing ideological priors and opinions work together to design and agree upon a shared modeling strategy. By co-authoring the research design in advance, we can systematically buffer against the individual biases that increase rather than reduce scientific uncertainty.

Limits, Rebuttals, and the Crisis of Reproducibility

Of course, our own estimates of ideological bias have limitations that must be transparently addressed. First, the original crowdsourced experiment was not designed to test the specific hypothesis that ideology influences findings; our results are exploratory not confirmatory and should serve as a prior for future confirmatory research. Second, although our estimated ideology effects are statistically significant, the confidence intervals are relatively large, making it difficult to draw definitive conclusions about the precise magnitude of the bias. Third, because very few participating researchers revealed strong anti-immigration sentiments, our analysis is underpowered to examine anti-immigration bias as robustly as pro-immigration bias. Fourth, potential social desirability bias in the initial surveys may have led some researchers to hide their anti-immigration views, contaminating the moderate and pro-immigration groups and expanding the confidence intervals. Fifth, the experimental data do not record the actual workflows or trial-and-error processes of how researchers experimented with and discarded models along the way, in case they did engage in so called “hacking”.

Our study has also sparked productive scientific debate. Katrin Auspurg and Josef Brüderl published a comment suggesting our evidence is fragile and hinges a team-level discipline variable that introduced statistical singularities, effectively excluding cases that happen to contradict our hypothesis. Although this critique has merit, in our original study and in a response to their criticism we demonstrated that our findings are highly robust by running hundreds of alternative models that do not suffer from this singularity problem. Running the analysis at the researcher-model dyad level for example completely removes team-level singularities and still provides evidence in support of an ideological bias effect.

Crucially, this vigorous academic debate was only possible because we adopted rigorous open science practices by sharing all our data and code online. Future vetting is absolutely essential not only for our work, but all scientific activities. The state of computational reproducibility in the social sciences is otherwise alarming. In a massive project published in Nature, Miske et al. (2026) estimated that computational reproducibility rates in some areas of the social and behavioral sciences are pathetic: less than 2% of sampled studies in education sciences including adult education science and less than 7% in sociology share their data and can be computationally reproduced. Scientists must commit to sharing code and data to allow rigorous peer vetting, and academic journals must adopt policies that support this. Currently none of the top journals in education science and sociology require transparent sharing of data and code. But this is a crucial need of science if we want to maintain public trust.

Subjectivity, Incentives, and an Honest Proposal

In reflecting on how to move forward, quantitative researchers working with data can learn a valuable lesson from qualitative methodologies. Ask any qualitative researcher, whether they are setting up a study based on participant observation, ethnography, or ethnomethodology, and they will tell you that subjectivity is not an obstacle to be ignored, but a reality to be managed and integrated. A rigorous qualitative study requires the researcher to actively reveal and integrate their subjectivity into the workflow, reflecting on how their own positioning influences their observations.

Quantitative and data-driven researchers should adopt a similar posture. Rather than clinging to the myth of the perfectly detached, value-free objective observer, we should reveal our subjective positions in advance and plan explicit methodological safeguards to protect our studies against them.

At a very basic level, we must acknowledge that we are all subjectively motivated by trying to pursue a career in scientific research within a highly competitive liberal capitalistic market for science. To secure employment, tenure, and funding, we need high-impact, publishable studies in highly ranked journals. Consequently, we are perversely incentivized by a “publish-or-perish” system to seek out and publish statistically significant, “clean” stories, which are often those that align with our preferred narratives.

If we were to explicitly state these pre-existing ideological priors and career incentives in a standardized “Conflict of Interest” statement, it would be an act of profound intellectual honesty. More importantly, it would serve as a powerful self-motivating mechanism. Acknowledging our subjective interests in black and white would compel us to integrate rigorous, pre-registered methodological buffers to protect the integrity of our empirical findings from our own career and ideological motivations.

Media outlet and Q&A for ‘Ideological bias in the production of research findings’ by Borjas and Breznau

  1. Süddeutsche Zeitung (SZ) article by Sebastian Herrmann “Warum Forscher aus denselben Daten entgegengesetzte Schlüsse ziehen” (Why researchers reach different conclusions from the same data).
  2. Manhattan City Journal perspective piece written by George and I.
  3. A news report by Luca Rehse-Knauf for Deutschlandfunk (German radio) and their Forschung Aktuelle series “Migrationspolitik: Einstellungen können Forschungsergebnisse beeinflussen” (Migration policy: Attitudes can influence research results).
  4. A news report by Katrin Kühn and Luca Rehse-Knauf for Deutschlandfunk and their Fakten und Meinungen Series “Darum sind wir Menschen nicht objektiv” (Why we humans are not objective).
  5. A PsyPost report by Eric W. Dolan “158 scientists used the same data, but their politics predicted the results“.
  6. A podcast on The Last Show with David Cooper. Apple / Youtube.
  7. A podcast in Allegedly Does Not Replicate with the Institute for Replication (I4R) with Abel Brodeur and Juan Pablo Posada Aparicio.
  8. A substack post by Claudio Teixiera after an interview with us about how researchers conduct research and the ‘invisible’ paths they follow.
  9. A substack post by Laurenz Gunther describing our work and digging deeper into bias among researchers (here those working on the topic of immigration) – despite a relatively far out conclusion.
  10. There was a Neuer Züricher Zeitung article about the original ‘Hidden Universe’ study (paywalled) ‘Das Experiment: Wer bekommt die rote Karte?.

— How did you arrive at your research question, and what is the study about?

George emailed me and had a few questions about our original study. He was at that time analyzing our data and had found a statistical association between pre-existing preferences for more or less migration among the teams in our study, and their findings. I was very skeptical. I have now worked on replication and reproducibility themes for almost a decade. I am acutely aware of what we often refer to as ‘researcher degrees of freedom’, also known as ‘the garden of forking paths’. This refers to choices that researchers can make during the research process that can lead to different outcomes. I assumed that the statistical association he found would not hold under different but equally plausible model specifications. I began testing many different models. Basically, they all showed the same result. Therefore, I became convinced that this was more than a fluke.

Actually, George had already run most of the same models. We present all of our models in a multiverse analysis in our paper. Out of 883 models 88% showed a significant statistical effect suggesting that we should reject the null hypothesis that ‘ideology has zero effect on the teams’ research findings’. If we take the assumption that we should only trust models that control for researchers’ educational experiences – something we believe impacts their results – then we find that roughly 93% of the models show a significant statistical effect.

— Could you explain the experimental design and methodology?

This study is an exploratory secondary analysis of the data generated by the experiment of myself, Eike Mark Rinke and Alexander Wuttke. We gave 71 research teams the same data and hypothesis – that immigration reduces support for social welfare policies. We surveyed them on their backgrounds and research experience and asked them if they believed the hypothesis was true and what they thought about immigration policy. In the original study, again the one that I helped lead, we did not find any important impact of immigration preferences. We essentially found that the results went in all directions, and we could not easily explain the variation.

After working together, George and I agree that the statistical analysis in the original study was not a clean test of immigration on the research teams’ results, because it controlled for the statistical model specifications that the teams’ made. The original study was were searching for key decisions that might explain why results went in different directions, so this naturally made sense. But in hindsight, these decisions are the mechanism through which ideology gets transmitted into statistical results. Thus, the original study introduced what is known in statistics and causal analysis as a ‘confounder problem’. If someone has an ideological bias, they will choose statistical models that will lead to more desirable results. The original study was controlling for both the test variable (ideology) and its mechanism (the statistical models), and in doing so it suppressed the impact of the test variable we were trying to observe.

This time, George and I conducted regression analyses in which the research findings were the dependent variable and ideology the independent variable – without model specifications as control variables, but with further controls.

— What are the limitations?

A key limitation is that this study is exploratory, not confirmatory. It relies on secondary data from a study that was not specifically designed to test the impact of ideological bias. It cannot confirm that this bias exists, instead it demonstrates robustly that a statistical association exists between ideology and researchers’ findings. We are not aware of any other way to explain this association other than an ideological bias. But we can only confirm with confidence, that it is prudent to reject the null hypothesis that in this particular sample and study is that there is no association between preexisting preferences for immigration policy and research findings. More specifically, the likelihood of observing the data in this experiment if the null hypothesis were true is very low.

Another limitation is that the size of the effect we found is unclear. It points in a positive direction – more pro- immigration policy stances associate with findings that show immigration has a more positive effect on social policy preferences among the public, and vice-versa with more anti-immigration policy stances and a more negative effect. But because of the great variation in results and the small sample size, the standard errors of the estimated statistical effects are very large. This means that the true effect might be anywhere from miniscule and near-zero, to moderate, to very large. We simply cannot say much about this here. More research is necessary, although this is an implicitly difficult topic to study, because if we inform researchers that we are studying their ideological bias, they might behave differently and this would take away ecological validity.

— According to the study, ideology influences model specifications. Could you provide a concrete example to illustrate how a single design decision (or a combination thereof) can have an impact?

I cannot, and this is another limitation of the study. If I could, it would be something we would have found in the original experiment that collected the data. But we can only point at patterns here. There are certain model specifications, unique combinations of statistical modelling choices that produce more negative results. The teams with more anti-immigration ideologies were more likely to choose these. But there are far more model specifications than there are teams. This leads to a sparse data problem. There are many empty cells in the matrix of all possible model specification combinations that teams would plausibly make. This makes it roughly impossible to pinpoint exact specifications’ effects. and there are many different model specifications that can lead to a positive or negative statistical effect. The point is that the only thing that happened between the teams asked to test the hypothesis with the same data, was different modelling choices. Therefore, this is the only way they could arrive at different results. There was no cheating or result faking, we checked that their statistical code produced the results they reported to us.


— To what extent is this a problem, and to what extent is it normal that decisions, based on analytical decisions, depend on who you are and how you think?

This is nothing that our study answers. And it may not be fully possible to answer because we do not yet know the nature of consciousness. We also cannot measure what is happening inside a human neural network – a brain in other words. But it is clear to me that experience, ideology and preferences shape results. A simple example is statistical training. Many researchers have limited statistical training, and they build only those statistical models that they learned about in their studies. This impacts results.

But more generally idiosyncrasies of people, like ideology, shape what research questions that people are willing to pursue and how, and they shape the reporting of those results. Some could look at our study and think that the estimated statistical impact is large and highly concerning. Others, might look at it and think it is tiny and of no concern at all. Our study suggests that ideology can explain somewhere between 1 and 3% of the variance in the results. If scientific findings are on average 1 to 3% off of what they would be without bias, is that a big problem? I mean… what do you think?

— If I understand correctly, the experiment was originally intended to show how much the results diverged, not why. How did you arrive at ideology as a possible cause?

As I already mentioned, this is something that George noticed in our data. He already had this hypothesis in his mind. I cannot blame him for thinking this. The Open Science Movement and Metascience work reveals many so called ‘Questionable Research Practices’. These include everything from faking data, to tampering with statistical models or stopping the collection of data during an experiment to produce a desired result. These practices are designed to produce certain results in order to obtain a publication or support a pet hypothesis (confirmation bias). Obviously some of these studies were motivated by ideological goals.

— How could ideological bias be reduced? Is this even desirable, or should we simply be aware that it can exist?

The impact of ideology can be reduced by following some clear recommendations of the Open Science Movement. Studies should be pre-preregistered, they should provide all code and materials, they should not be conducted in isolation or in hiding, and researchers should cooperate. Some of us are engaging in so called ‚Adversarial Collaborations‘, where researchers who do not agree – those with different priors about a given hypothesis like the impact of immigration – collaborate. They lay out all the aspects of a study in advance, and they agree on what evidence would count as support of either of the positions. I highly recommend this form of science. It takes competition and turns it into collaboration with the goal of knowledge seeking prioritized above all else.


— You are investigating ideological bias in science using scientific methods. How do you deal with this tension in meta-scientific questions, where you are essentially also your own subject of investigation?

Similar to my last answer, one cannot fully understand or deal with one’s own bias, and therefore needs to build in checks into the process. Things that would reduce this bias, like preregistration. I am working currently on a project that is an Autoethnography of my own questionable research practices and the perverse incentives I encountered during my career in science. I hope that by doing this, and revealing my own behaviors, I can improve them. I also want to be a role model for others, to make it desirable to be highly critical of one’s own work. My goal is to get this study published in a high quality journal and thereby prove that self-criticism, something researchers mostly try to avoid to protect their theories, findings and careers, is something that can be used in a positive way in the scientific process.

Can one’s own attitudes always play a role, even in this study?

Sure. Definitely. That is why it was important for me to take a so called ‘multiverse’ approach to this study, and many other studies I am working on recently. I want to ensure that I, or one of my colleagues, has not simply selected a statistical model that produces certain results. This practice, known as hacking, or p-hacking, is prevalent in science, especially in secondary data analysis. I essentially learned to do this during my graduate studies. We would find a result we liked and then develop convincing logical arguments why the model producing it must be the best model. So, the idea with multiverse analysis, is to run all or at least all plausible alternative models. This helps reveal if my model is an outlier. Whether it represents something very unique or unusual in the distribution of model results. If it does, this is a cause for great concern. If not, it is evidence of a robust statistical association.

— There have also been critical reactions, for example in this online Bluesky thread, which raise concerns about George Borjas’s views and background. How do you assess this criticism?

I have read the discussion thread by Michael Clemens, an economist at George Mason University. One line of criticism appears to concern the fact that George Borjas recently conducted research for the executive branch of government.

Another criticism seems to focus on the observation that different model specifications yield different results. This is precisely what our study confirms, and it is also a well-established fact in the history of empirical social science.

Both points are orthogonal to our study. Even if we were to assume, hypothetically, that George held some form of ideological bias and that this bias influenced his analytical choices, which I cannot confirm and do not claim, this would not undermine our findings. The reason is that we adopted a deliberately robust research design. We conducted a multiverse analysis comprising 883 regression models. The consistency of results across this large set of plausible specifications makes it implausible that our conclusions are driven by special highly selective model specifications – those that would be selected due to ideological bias.

It is also important to note that George and I do not share the same political views. Precisely for that reason, ideology was an additional motivation for us to adopt a highly robust analytical strategy. I explicitly advocated for the multiverse approach in order to minimize the influence of individual priors, including our own potential ideological orientations. I see this as a great strength of the study.

The purpose of our approach is to decouple empirical results from personal or ideological preferences as much as possible. And to estimate the robustness of our finding to any kind of bias, not just ideological. I would encourage critics to apply the same standards of robustness to their own work. To date, Mr. Clemens has not presented empirical evidence that contradicts our findings. Moreover, criticisms referring to modeling choices in studies conducted by George decades ago are not relevant to the validity of the present analysis.

— What does it mean when you say that only 3% of the variance can be explained by your results.

That has to do with the regression coefficient and the r-squared values. A coefficient of 0.03 shows that a one-point higher (more positive) ideology mathematically predicts a change in results of 0.03. This sounds meaningless, but we know that 0 would be none (and we can equate this with zero percent change) and that 1 would be 1-point on a standardized scale. This is 1 standard deviation in the distribution of the dependent variable. It would be possible to move the results more than 1 standard deviation, but this would be quite preposterous. There is nothing in the complex nature of social science, which lacks laws, that would do that. So I will set the upper bound of the largest possible effect at 1, meaning that 1 would equal 100% of the distance in the distribution of variance.  Therefore, 0.03 is like 3 percent of the distance. At the same time, the r-squared, which tells us how much of the error is reduced from this particular variable, is around 0.03 or less depending on how we measure this variable. This suggests that fitting the observed ideology values into the observed results from the teams, reduces the unexplained variance from 100% down to around 97%.

— Can you explain — very, very simply — what you did? 

In a study that I co-led starting in 2018 (Breznau, Rinke and Wuttke et al. 2022), we designed an experiment that allowed us to observe researchers doing research on the impact of immigration on social policy preferences. We gave them the same data and asked them to answer the same research question: whether immigration reduces support for social policies or not. We documented the researchers’ pre-existing methodological training, experience with and expectations about the topic, and their personal preferences for looser or tighter immigration laws in their own countries. We shared all of the data and documentation of our work publicly. Because we shared our data, a few years later George J. Borjas was able to reanalyze our data and find new evidence of a correlation between the researchers’ ideological positions on immigration and their findings. Those with more pro-immigration positions tended to find evidence that immigration had a positive impact on support for social policy, and those with more anti-immigration positions tended to find evidence that immigration had a negative impact on support for social policy. Here “social policy” means support for a more extensive welfare state providing social security via the government or not – many would call this support for social cohesion. I was skeptical of George’s initial work, and together we vetted George’s findings. We ran almost a thousand alternative statistical models to test George’s finding and of these 88% suggest that we should reject the null hypothesis. The null hypothesis is that if there is no impact of ideology on researchers’ findings, that we should not observe what we observed based on probability. But we did observe this association, and by rejecting the null hypothesis we have evidence that something more is going on, not just random luck in the data. Crucially, we should reject the idea that ideology has no impact on researchers’ results when analyzing the same data. 

— What motivated you to do this study?

As I said, George found this association between researchers’ ideological positions on immigration and their research findings. This is an important scientific observation. I was a bit more skeptical, in particular because my initial analysis of the data as part of the original experiment did not show this association. Together we became very motivated to tackle this problem, and in the end, we have relatively strong evidence of something. The exact nature of this should be subject to further research.

— What are the most striking findings? 

The main finding is that researchers’ own preferences for tighter or looser immigration predicts what they went on to find in their work. This appears striking, but when put into context it is maybe not so surprising. We know that there are many reasons that researchers may consciously or even unconsciously exert influence on their own findings. There are many cases of researchers engaging in questionable research practices to ‘fudge’ their data and results in ways that make them appear stronger than they really are. This has occurred frequently in biomedicine, for example in studies of products whose approval would net the researchers great personal profits. Consider also what we know from psychology, namely confirmation bias. People tend to seek out evidence of what they already believe is true. Researchers are people too. When presented with various forms of competing evidence, a researcher might gravitate toward that which supports their preexisting beliefs or preferences. 

— What are the implications for the social sciences, which are already reeling from the replication crisis? 

The implications reinforce what we already know. We should focus on refining two areas of science. The first is scientific training. We need to make transparent and open workflows and data sharing the norm. This includes researchers stating in advance what they plan to do and what they expect to find. When researchers do not do this, they can run several experiments or analyze hundreds of datasets with millions of different statistical models and simply choose one finding that looks exciting or sexy to them. Generations of researchers before us have learned to do exactly this in order to get published. They learned to ‘sell’ a single selected finding as confirmatory evidence of something in the real world, when in fact it is simply a highly selected, exploration of data leading to a unique event; one that probably does not generalize and is not reproducible (a.k.a. luck). I bring up the idea of getting published here because this is the second problem we must urgently address in science. A researcher’s worth or ranking as a scientist is judged almost entirely on their publication record. In particular, publications in journals that are considered higher status. These higher status journals should be publishing studies because the studies contain higher quality science. But there are ways to game this system. There are a a host of questionable research practices that makes findings look more exciting than they actually are. These practices often lead to irreproducible findings. Essentially fake science. Big publishing is a major profit industry and this increases the pressure to publish, as the publishers of the journals engage in questionable, sometimes unethical practices to sell more journals, as opposed to solid science which is often quite boring and tends to find that new drugs or treatments do not work and that our theories are wrong. Ideally we need to end big publishing’s control over science. Their role should be simply production and distribution, but they currently copyright much of the material and force universities to pay twice, once for the researchers’ salaries and again so that researchers can read what they are publishing. And this often done with public money. Its really wrong and generates perverse incentives among researchers and profit-seeking publishers. 

— It seems teams didn’t falsify data or cherry-pick numbers in any obvious way; instead, ideology appeared to influence judgement calls — is that a reasonable explanation or is it too charitable?

It is a reasonable explanation, but we must be very cautious with it. Ideology might explain about 3% of the variance in researchers’ findings. The rest has to do with other factors or random noise. Consider that many scientists have specific methodological training. Through this training they simply do what they know how to do. And this can influence results. For example, someone who only knows how to use a hammer, will treat things as a hammering problem, when it might be better solved with a different tool. This takes us back to better, broader methodological training, which scientists would have more time for if they weren’t under constant pressure to publish. 

— Are there ways to guard against the ideological bias highlighted in the study? It seems that peer review can spot poorly defined studies — but does it work well enough in the real world? 

Peer review is a poor solution. There are studies out there that show that peer review is not reliable. Just like giving the same researchers the same data, if you give different peer reviewers the same paper, they will come to a huge range of judgements about the paper. The publishing system is in some ways a lottery. One solution for this problem is what we call “adversarial collaboration”. This is a type of research where scientists who disagree about a topic work together. They design a study together and agree on all methods and on all criteria with which to judge the outcomes. Then after this is all agreed in advance, the study is conducted, ideally from a third-party, and then the results speak for themselves.

— What do you think the message here is to the public and to policy makers?

Everything we have in society that works is based on science. Smartphones, that open heart surgery that saved the life of a loved one, airbags, planes that don’t crash. We need science for every decision we make collectively. But the public and policymakers can be highly politicized, and this can influence science, we see this even in the scientists in our study whose own politics seemingly played a role. To cut through this, we need policymakers that support science conducted by scholars who have different political ideologies, who do not agree. For example, George and I are somewhat different in our own assessment of the impact of immigration on society and the labor market. This made us a very strong team. It meant that we could focus on the scientific process and try to get to the best, most reliable answer. Crucially, we need to never rely on single studies or single science teams. Before we declare evidence of anything, we need many studies. We should not just rely on single papers or scholars. We need dozens or hundreds of studies on a topic conducted by inter-disciplinary and inter-ideological teams.

Questionable research practices from the practitioners’ perspectives

With Monica Gonzales-Marquez, Priya Silverstein and Eike Mark Rinke.

In preparation for Metascience 2025, I put together a panel with three other researchers called “Questionable Research Practices from the Perspective of the Researcher: Understanding Perverse Incentives using Autoethnography”. We used our own discussions in online meetings and individually as recorded narratives as content for this panel. The impetus was that much metascience points accusatory fingers at problems in science, thus supporting a culture of fear. Researchers comply to avoid scrutiny, and not necessarily out of an intrinsic motivation to do good science.

When successful and meaningful science is measured entirely by publication in a ‘high impact’ journal and high citation counts, the drive to do science is transmuted into a laser-like focus on publishing. The intrinsic motivation to contribute to the creation of scientific knowledge becomes confounded with publishing, while simultaneously deprioritising the robust, pedantic, methodical, humble labor involved in doing good research. The scientific method gets confounded with the mechanics of publishing and achieving “standing” in the scientific community.

We hope that by revealing our own experiences in the world of ‘publish-or-perish’, and how it has pushed us towards questionable research practices (QRPs), we might generate intrinsic motivation in others to dispassionately examine their own scientific practices.. By looking back through our histories in academic work and sharing them, we also expect to increase our own intrinsic motivations to be ever vigilant, and to learn to always privilege scientific integrity over publishing.

Confounding: Publication = Science

A major theme in our narratives is that science has been fully confounded with publishing. For some of us, the pressure to publish, and the toxic atmosphere and interactions pushed us out of academia entirely. For others, we internalized the publishing norm, convincing ourselves that we were doing impactful science, when we were actually just doing impactful publishing – which does almost nothing to alter collective human knowledge, promote social justice or solve other societal problems.

Some excerpts from our narratives:

Questionable Behaviors: Hacking

The pragmatics of doing science in a publish or perish culture is that we either engage in some (often undeliberate) hacking and storytelling, or are sanctioned.

Some excerpts:

Ego, Power and Personal Struggles

As human beings we are prone to seeking status, material security, community acknowledgement and different things depending on where we are at in life, and who we are personality-wise. Especially when in graduate school, we are in a position of little power in comparison to professors and our supervisors. People who  wield power over others, and push their own ego-centric agendas tend to be quite successful in science. This can lead to abuses of power, ego trips and other toxic behaviors that can diminish junior researchers’ ability to push back against pressures to engage in QRPs and have strongly demotivating effects.

Intrinsic Motivation

We hope that by speaking out, and normalizing scrutiny of our own experiences of questionable research practices as something valuable, we can help others become intrinsically motivated to do the same. Moreover, we propose that ethnographic narrative may be an underexplored but powerful method to help uncover the causes and motivations of questionable research practices from researchers’ lived experience of “doing science”.

Part of the motivation for this panel is based on one of the authors’ (Nate Breznau’s) own authoethnographic research into my QRPs. He presented preliminary results at the Sociological Science Conference at Cornell (link to slides).

Readers can find the full poster for our presentation at Metascience 2025 at University College London here.

Our future plans are to seek a larger sample of researchers and invite them to a semi-structured narrative sharing process (via recording themselves) to further study QRPs. In particular we hope to target early career researchers to shed light on the current state of science training and supervision experiences. 

Science in survival mode

Scientific research is unreliable. It comes with uncertainty. Whether launching a rocket or measuring racial prejudice, there is uncertainty. We use this uncertainty to make decisions. If the rocket has a 40% chance of exploding, best not to stick astronauts in it. If skin-tone bias of soccer referees is somewhere from none to a lot, it is irresponsible to conclude they are prejudiced.

Investigating uncertainty, is science. Truth and uncertainty are two sides of the same coin, they co-define each other. A problem for humans measuring uncertainty, is that humans are unreliable. Human scientists themselves add uncertainty to the measurement of uncertainty.

Recent studies suggest that somewhere between 25 and 60% of published statistical results cannot be recreated using the materials provided – these measures of uncertainty come with their own uncertainty. It turns out that scientific researchers are doing some really peculiar things to generate uncertainty. Some surveys suggest as many as 9% of scientists faked data at least once in their careers, and that more than half selectively reported findings – a behavior that makes the things they study to appear less uncertain than they actually are.

I believe the answer lies in their humanness. Like all animals, they are genetically programmed for survival. They are capable of both rational, reflective decision-making and split second reaction without any thought. Given time to reflect, a human would generally conclude murder is unethical, but simultaneously would not hesitate to kill if it prevented their child from being killed. Murder remains wrong, but to not kill and let a child die is also wrong. It would be irresponsible parenting failing to ensure survival.

Murder is a profound act when a human perceives themselves to be in a situation of life or death. It is a symptom of subconsciously activated defense mechanisms in survival mode. But there are many other symptoms. In order to avoid death, humans, like other primates can engage in deception, disassociation, aggression, manipulation, submission, scapegoating, theft and hoarding. If someone held me at gun point and told me to prove the earth was flat, I would have no problem doing it. The math would work, I would just need to fake a little data

Most scientists have highly valued knowledge and competencies and thus live in situations where their lives are not under threat. At least not as a result of their scientific practice. For the sake of this thought experiment, lets just rule out regularly occurring life-threatening danger as a cause of scientists exhibiting survival mode behaviors. This leaves the perception that they are under threat, as a possible explanation.

It takes only a few stimuli to induce survival mode behavioral changes in animals. Like hearing a frightful noise when seeing an animal. This animal and anything that looks like it become automatic sources of anxiety and fear even without the sound. Imagine being told over and over and over that you have to have an exciting study with powerful results in order to get an academic job after graduate school, otherwise you wasted 3-8 years of your life and probably a large chunk of capital on getting a PhD. Could this alone, without introducing any actual shocks or physical pain, induce Pavlovian fear? I encourage you to go ask any graduate student to answer this that does not yet have such a study.

Now imagine that during graduate school a student invests all their time and resources into an experiment. After it is complete, they hypothesized result, the one sure to be exciting and publishable, is not there. Imagine the horror, the shame, the feeling of failure, the panic. Remember the poor soul who leaped to his death because he misunderstood futures trading and thought he owed three-quarters of a million dollars? That was a triggered survival response. Because dying felt like the only way to ‘survive’ the horror of facing that debt. Imagine if he could have just changed the futures market by adding his own numbers to the market. Would he have done it?

We should not be surprised at all then, when scientists acting out of fear-based survival strategies, fake data. Diederik Stapel faked an entire career of data before being caught. His behavior was self-described as an “addiction”. A common reaction of individuals placed under fear stimuli that are emotionally damaging if not traumatic. The intense pressures and expectations of the academic environment created a context in which he felt compelled to engage in unethical practices to maintain his status and success. It does not make murder or data-faking right, but to not take this as grounds for indicting the scientific rewards system is certainly wrong.

The incentive structures have to change before we can honestly expect the fear-driven pressure to fake, cheat, lie or steal – in order to avoid the experience of loss associated with the common null results that occur when conducting high quality scientific research on radically complex human brains and societies that are frustratingly difficult to measure things in – to go away.

Open science in sociology. What, why and now.

WHAT

By now you’ve heard the term “open science”. Although it has no global definition, its advocates tend toward certain agreements. Most definitions focus on the practical aspects of accessibility.

“…the practice of science in such a way that others can collaborate and contribute, where research data, lab notes and other research processes are freely available, under terms that enable reuse, redistribution and reproduction of the research and its underlying data and methods.”


FORSTER, open science teaching resource

Some definitions enter the realm of ethics, feminism and social justice.

“…to imagine and design inclusive infrastructures, practices, and workflows for scientific practice that intentionally enable meaningful participation and redress (these new) forms of exclusion.


Denisse Albornoz,OCSDNet

Others focus on the communicative interplay between scientists and the public.

“Openness in Open Science also means opening up science to society… The democratic ideal of Open Science argues for equal two-way communication with the public: one should not solely focus on the question of how to foster the uptake of science in society, but also on how to foster the uptake of societal insights in science.


Anne-Floor Scholvinck,ZBW Mediatalk

Whatever the ontology, open science is inevitably something that challenges the status quo in science. Usage of term indicates there is something undesirable about science, otherwise advocates would simply advocate “science”.

The “open” part of the concept refers to any number of things depending on whom you ask. Commonly it means:

Open access – making the results of scientific techniques, research and theory accessible to everyone; as opposed to only in paywalled journals.

Transparency <open process> – making all methods, code, data and any biases or conflicts of interest known before and after the research is conducted. So long as doing this does not harm human subjects or violate any laws.

Open source – on the technology side of science, all programs, apps, algorithms, tools and scripts should be transparent and usable by others. This means that when a scientist develops a new technology, anyone else’s technologies can interact and interface with it. Moreover, anyone can modify the technology to better suit their own needs.

Open academia <open communication/democracy/feminism> – allowing anyone to participate in academia. That academia has the goal of eliminating inequalities, prejudice and domination from academia that take place in the social world. That academia embraces feminism and critical race theory in its methods and institutional practices. That everyone has the same place in scientific discussions, and no science is conducted by pressuring others or taking advantage of existing power structures. That no science takes place in secret, except for research that requires obfuscation for its completion.

Again, the definitions can cover a broad range. The above are just a snippet, although they strike me as the most common usages; except for ‘open academia’, this is reserved for certain justice motivated scholars.

WHY

Although I do not proclaim to be the arbiter or knower of right or wrong in academia (and life in general), the following facts seem wrong to me.

Double-work and the co-opting of journals

Scientists provide their work as editors and reviewers, because the peer review and publication process is the centerpiece of all of science. Peer reviewers and editors are the only consistent form of quality control in science. The academic journal was a functional response to previous forms of knowledge transmission that required direct scientist/practitioner to student interactions which were geographically limited and reached a very narrow audience.

The journal made it possible to transmit knowledge across the globe. Moreover, the journal reduced the simultaneous discovery and re-discovery problems of science, because no one could prove they discovered something first, and others worked on problems that were already solved unknowingly. It represents one of the first ‘open science’ movements because it was driven by the idea that science was at an impasse and could only move forward through transparent and open exchange of ideas arbitrated by being part of the public record through publishing.

Ironically, the journal format came full circle and began to undermine science. After over two centuries of journals run by non-profit academic associations, for-profit publishing houses began ‘offering’ their services to meet the growing global demand for journals and their content and the rising costs of editing and distribution. In many cases, these publishing houses were able to purchase the journals by offering the academic societies the exclusive right to determine what went in them. Within just 30 years, five conglomerates owned the titles, content or certain features of over 50% of all journal articles published globally.

The content, as always, is still a product of the scientists and the voluntary work of editors and peer reviewers. The publishing houses make large profits, but pay nothing to these workers. The editors and peer reviewers earn their income from universities mostly. The very universities that pay high fees to purchase the right to provide the journals in their libraries. This is a double tax on the universities – paying the producers of content to produce and then paying the distributors of that content to consume it. The content does not change at any point in between these two forms of payment, in other words, the publishers do not add any scientific value to this content.

Matters got even worse with the publishing houses over the past decades. As creative and deceitful profit seekers, some publishing houses realized they could generate even more profit by collaborating with the private sector. For example pharmaceutical companies’ profits were directly determined by the findings of studies published in journals. Pharmaceutical companies, or any companies whose profits were determined by the outcomes of scientific experiments, would be willing to invest in shaping those outcomes if they could. Enter a novel concept pioneered by Elsevier: selling journals or journal space to private companies to boost their profits. Win-win for them. Elsevier also pioneered the process of monetizing open science by purchasing SSRN, engaging in massive lawsuits designed to stop the free sharing of (their) copyrighted knowledge and tries to copyright intellectual activities such as peer review.

Other ventures create journals that prey on scholars who do not know better, or seek to get easy publications to add to their CV. These publishers are often labeled “predatory publishers” and they “publish work without proper peer review and which charge scholars sometimes huge fees to submit should not be allowed to share space with legitimate journals and publishers, whether open access or not” (predatoryjournals.com). They also sometimes mimic reputable journals by copying their styles and their names and soliciting content from scholars, a procedure known as “hijacking“.

Publish-or-perish begets questionable research practices

Thanks to the advent of the scientific journal, knowledge could be evaluated, used and further transmitted across space and time. The utility of the journal and other forms of academic publication such as books, proved so effective that they became the primary source for others to evaluate the importance of scientists and their work. This gave rise to the norm we are all familiar with, publish-or-perish.

In a survey of psychologists, John et al. (2012) found that 50% claimed they had selectively reported studies that supported their hypothesis (as in, selectively excluding those that didn’t). Moreover, 35% admitted to reporting unexpected findings as having been predicted from the start. Nearly 2% outright admitted to faking data.

Publish-or-perish and questionable research practices have a causal relationship. Except for occasional sociopathic or psychotic individuals, there is no reason for a scientist to engage in questionable research practices. No reason, except scientists’ very existence on scientists may depend on it. So many studies in reality lead to results that go in all directions, support the null or (most importantly) do not provide groundbreaking new results.

Through the peer review and editorial process, journals select studies that are path-breaking. Studies that will move knowledge forward and be of the greatest interest to readers. When faced with prospects of not getting tenured, not getting grant funding and being forced out of academia, a human’s (scientist’s) rational calculations change. Suddenly, rounding that p-value from 0.054 to < 0.05 or even adding some cases to the data becomes a cognitively defensible decision.

Like any profession, science is competitive. Those who publish more, or get more citations to their publications tend to get ahead. Those who don’t, don’t. Professional athletes use incredible tactics to gain competitive advantage. Of course steroids are well-known, but other tactics are much harder to detect. For example, endurance athletes often use blood transfusions to boost recovery and performance. This is what it means to be human, scientist or not.

One of the most radical events in the social and behavioral sciences is Diederik Stapel’s entire career faking data and results that were published in at least 54 articles that consumed millions of Euro in funding. It took almost two decades for critics and whistleblowers to finally out him. Psychology is not alone. In political science LaCour and Green published a study in Science that attitudes toward gay marriage could be changed if heterosexual people listened to a homosexual person’s story, but it turns out LaCour fabricated results of a follow up survey that never took place as uncovered by Broockman. In economics Reinhart and Rogoff published numerous studies identifying a negative impact of high debt rates on national economic growth, when in fact several points in their dataset had conspicuously missing values. When these values were added there was no longer support for their claim as identified by Herndon, Ash and Pollin.

I suspect that most questionable research practices are not intentional. The sociopathic (~psychotic) Stapel’s of the world are rare. This pressure to find a job after doing doctoral studies and then to get tenured, means a trade off between conducting science in its ideal form – so learning as much as possible about the existing literature on a subject, mastering the necessary methods to perform the research and executing the research, possibly with several iterations, and facing the prospect of null results – with science in a form that will lead to publication as fast as possible.

This ‘fast as possible’ leads to amateur science. For example, in the rush to get my first publication I attempted to use “multiple imputation”, but lacked the time to properly learn this method. Instead I simply generated several datasets and averaged them into one and re-ran the analysis on this one. This was not an intentional misuse of a method. It is a questionable research practice as a result of context. Think about matrix algebra. It is the basis of many advanced statistical techniques regularly used by social scientists. How many of us have a strong grasp of matrix mathematics? I don’t. And yet I’ve published several studies using structural equation modeling.

WHAT & WHY in SOCIOLOGY

I am aware of nothing about sociology that suggests it needs a special adaptation of open science. Most research cannot be strictly delineated as sociology or not sociology anyways. The boundaries of a discipline, especially within the social sciences, exist mostly in the institutional structure of universities. Eliason suggested that sociology is unique because it overemphasizes quantitative techniques, has needlessly long articles, lacks writing for the popular press and emphasizes research at the expense of teaching. In my experience the previous sentence perfectly describes all social and behavioral science disciplines at once. Even article length, something I thought might be peculiar to sociology, is not special. Political science and management research have very long articles. Consider that and ASR and ESR for example, limit words to 9,000 and 8,000 or less – this is relatively average if not short for social science.

Actually, I would argue the most unique thing about sociology at the moment relates to open science. Two points in particular: (A) that sociology has not had the same incredible scandals as other disciplines and (B) that sociology lags behind other social sciences in promoting open science.

A lack of scandals, not scandalousness

Could sociologists be more scientific and ethical in their research behaviors than those in other disciplines? Given identical institutional and career structures that favor productivity and innovation over replicating or checking each other’s work, I doubt it. Sociology journals and their editors, for example, rarely retract articles despite evidence of serious methodological mistakes. Carina Mood once accurately pointed out mistakes in the interpretation of odds-ratios in some American Sociological Review articles, but the editors refused to publish her comments, much less consider retractions. She shared her exchange with ASR in an email to me and discusses some of it in a working paper. An exceptional recent event was the retraction of one of Legewie’s sociological studies, but this required he himself to initiate the retraction after someone pointed out errors in his work. Until 2020, the Retraction Watch database (www.retractiondatabase.org) listed no retractions from the top sociology journals, and only two among the well-known, one in Sociology and another in Social Indicators Research.

This year, something new happened. Five articles published in Social Problems, Criminology, and Law & Society Review were retracted. These articles had the common co-author Eric Stewart. It turns out that the data he provided were faked. There is no other logical conclusion that this after exceptionally rigorous work by Pickett (a co-author of Stewart) provided evidence that the Stewart studies had consistently incorrect means and standard deviations, unverifiable surveys (sources, methods, original materials), magically changing case numbers despite identical statistical results, sometimes half the data had duplicate cases and impossible clustering structures in the data.

As an aside, one of Pickett’s findings was that the data had non-uniform terminal digit distributions. This means that the right-most digits in the reported statistics differs markedly from a uniform distribution. In particular, at the third-digit numbers should be uniformly distributed with 0-9 appearing roughly 10% of the time. In one of the papers, zeros appear less than 2% of the time. If you are considering faking data, keep in mind that it is roughly impossible to do it in a way that cannot be detected by careful investigation. Any algorithm used to generate results (even copying and pasting) leaves is statistical marks.

Perhaps we sociologists should be partly relieved, as this is just confirmation that we are as much a part of social science and its problems, as any other discipline. However, the Stewart retractions which should have been breaking news for sociology, went mostly unnoticed. The results of the investigation leading to the retractions is not published in a flagship sociology journal where it belongs. Instead it appears in Econ Journal Watch – something unlikely to be read by any sociologist. Moreover, the retraction notices from the original journals do not cite outright fraud. Stewart continues to promote his work in print claiming the main findings still hold, and several other of his studies with similar irregularities have not been retracted.

Another, extremely important event was a case of ethnomethodological research conducted by Lindsay, Boghossian, and Pluckrose in the mid 2010s. This is sociological self-examination at its best, although their backgrounds are mostly outside of the discipline of sociology. They wrote a series of 20 papers presenting fake results and making arguably unethical claims. They invented the papers to mimic the style of articles published in journals well-known for sociological research on topics of identity, hegemony and marginalization. Seven of their papers were published or had revise and resubmit recommendations before whistleblowing forced them to cancel the project. Some highlights: one paper contained sections from Hitler’s Mein Kampf. Another suggested men should be trained similar to dogs to prevent rape, and a third that white men should be forced to sit in chains on the floors of university classrooms, instead of normal desks. I am not commenting on the merit contained in these ideas, only that they all contained faked data, non-existent methods or conclusions not supported by the data. That these studies easily flew under the radar of a number of high impact journals points out how easy it is to publish without doing the necessary research work.

Lagging behind closed doors

October 6th, 2020. I entered the search terms “open science” (with quotations to search the exact phrase) and “sociology” (with quotations to only return results that contain the word) into Google Scholar. Six pages of results without a single sociology journal. On page 7, Merton’s “Priorities in scientific discovery: a chapter in the sociology of science” appears. Publication date 1957.

In 1973, Wilson, Smoke and Martin found that 80% of studies published in the top three sociology journals of that time rejected the null hypothesis, in other words they had p-values below a threshold. This suggests publication bias, if not p-hacking. Sahner (Table 5) analyzed all article submissions to the Zeitschrift für Soziologie, 1972-1980. Of those that contained significance tests, 70% were significant at p < 0.05 suggesting that authors prefer to submit significant results. More recently, Gerber and Malhotra (2008) reviewed articles published in American Journal of Sociology, American Sociological Review and The Sociological Quarterly, and specifically looked at the boundary of t = 1,96 (i.e., p<0.05) to find that as many as 4-out-of-5 studies were ‘significant’. This suggests publication bias as well. Sociology has yet to have a systematic review of p-hacking by comparing p-values within ‘significant’ results. Meanwhile psychology and political science for example are teeming with papers on “p-hacking” and “publication bias”.

Sociology is rather intransparent. An estimated 78% of the major sociology journals have long-standing transparency policies. Unfortunately, these policies are mostly artifacts on paper without much enforcement. For example, only 37% of sociology articles published in the mainstream journals between 2012-2014 include shared data and/or materials. In 2015, a small group of sociologists tried to obtain materials from the authors of 53 prominent sociological studies. They obtained these from just 19%, and only 20% of all the authors they contacted bothered to respond despite several requests. This suggests sociologists are free to hide the data and materials that led to their findings without recourse, despite such guidelines.

Other disciplines have embraced the Transparency and Openness Promotion Guidelines (TOP). The TOP guidelines with help of the Center for Open Science support journals to improve science. Journals can become signatories of TOP, and in doing so they either adopt and enforce new transparency guidelines, or certify that they already meet certain transparency standards. Most of the top psychology journals and several political science journals signed on. Other major journals such as the Journal of Applied Econometrics and later the American Economic Review adopted their own enforced transparency guidelines.

Until 2017, the only higher ranking sociology journals that signed TOP were Sociological Methods and Research and American Journal of Cultural Sociology. In 2017, Elsevier dictated that all its journals adopt guidelines and this added Social Science Research to the list. At the time of writing this, the flagship journals American Journal of Sociology and American Sociological Review neither signed TOP nor enforce their own guidelines. Of top German sociology journals, the Kölner Zeitschrift für Soziologie und Sozialpsychologie is the only signatory.

If intransparency is pervasive in sociology, then research cannot be (a) checked for errors, (b) reproduced or (c) simply critiqued. Even when exact reproducibility is not the goal, as often is the case with context-specific interpretive research, most research methods remain shrouded in mystery. This requires readers to take a giant leap to trust what others report. Part of the problem is that sociologists express little interest in reproduction or checking others’ works. There are few replications in the history of sociology, and if anything, they decreased over time until recently. For example, searching the articles in American Journal of Sociology and American Sociological Review reveals 22 replication studies from 1950-1980 and only 8 from 1981-2010.

Something telling about a lack of willingness to open sociology comes from sociology’s most ‘powerful’ society, the American Sociological Association. They collectively petitioned the US government to not make data transparency a requirement attached to grant funding in 2019.

NOW

What to do about it? Here are some simple steps to consider especially for sociologists. Similar to steps advocated by many others for graduate students and academic institutions, or all of us for example.

Transparency

Make all the materials – research design, methodological steps, data (when legally and ethically possible), analyses, conflict of interest and any software code – available online. The practical reason is that others can follow your work and expand it in the future. Doubly practical is that you don’t need to respond to email requests for your materials. So long as you are not a deceitful sociopath, you want others interested in your work and to replicate your work. Even if a study, seems to ‘prove you wrong’, the fact that it replicated your work is evidence of how important your work is and the topic of study. You are a piece of a much larger community of knowledge construction. Constructive exchange can lead to collaboration with critics to generate better future research without personal conflicts.

The immediate value of transparency is that being transparent forces you to be careful. Knowing everything will be public information increases the value of attention to detail. Put in its converse: not sharing your workflow publicly can indirectly foster lower quality standards, in addition to creating possibilities for misconduct. All this enables rather than hinders knowledge, and increases inter-researcher trust.

Transparency should not be much extra work. During the research process you should take high quality notes for yourself. You will often return to your data and research in the future and thus need those notes. This is a best practice with or without sharing your work. When you engage in this best practice, you have a deep familiarity with your data and can draw meaningful conclusions and easily redact identifying characteristics in your data in the case of qualitative research. In case you cannot share data, you can still reveal the design and expectations; or allow controlled access to the data. Human subjects must be protected at all costs, and yes this often means data sharing is not possible .

The ‘transparency work’ of the qualitative research process can be reduced by software platforms that provide semi-automated annotation and coding. Even if you do not share data, you can build an open workflow from the beginning that allows others to understand every step of the data generating process. However, this work can also be extremely tedious and the incentives not immediately clear. More fruitful discussion if not research assistant funding is needed in this area moving forward.

If you are using quantitative methods, immediately stop hiding your work. If you ran 100 models and 99 did not support your hypothesis, then this is your finding. If a journal does not want to publish this, point the editors and reviewers to the importance of null results and the problems of publication bias. If they still refuse, consider boycotting this journal and sharing your negative experience in public.

Preregistration

Preregistration can drastically reduce bias and hacking prior to collecting data. When you clearly outline your plans including how you will analyze the data, before conducting the research, there is little room for hacking so long as you stick to the plan. Moreover, preregistration can be done directly with a journal although sociology journals are laggards here because they generally do not offer this option. In a preregistration, even if you just put an pre-analysis plan or research design and goals online, you must think much harder about factors such as meaning, causality, inter-subjectivity and ‘how the world probably works’. You cannot hide behind results in this process and therefore you must anticipate counterarguments and explore counterfactual logic. This improves the clarity of theory and research, creating an immense gain in efficiency and effectiveness.

Regardless of the methods you use there are many opportunities to take advantage of preregistration. Some forms of qualitative research, for example those involving grounded theory and interpretivist methods, require decisions during the research process that cannot be foreseen. This uncertainty can be outlined in a preregistration stating explicitly when flexibility is and is not admissible. Moreover, simply putting a qualitative research plan online prior to conducting the research is equivalent to a pre-analysis plan. This research design need not compromise your data collection work because you can register the plan on a platform like the Open Science Framework and then embargo it, so that it is preserved but not made public until after the research concludes. Some scholars using quantitative methods might assume that preregistration is not possible because they work with secondary survey data. But the regularity and release of these survey data are known in advance, and these scholars can preregister their studies before the next round of data are collected with the knowledge of which questions and countries will be available.

Decommodify science

The central functions of the scientific publishing industry are printing and disseminating knowledge, which historically solved a problem of how to share knowledge across universities and countries. The business functions of publishing, however, come with harmful byproducts. Publishing firms extract profits from scientists twice. First, scientists provide free labor in the form of editing and peer reviewing, in addition to producing the results for the articles to be printed. Next, researchers, or their employers, must purchase the product of their own labor; labor not paid for by the publishers. The journal article as a product comes at a high cost, and often only in packages of journals meaning that universities have to pay for extra material their scholars do not use.

Sometimes publishing houses neglect science in favor of profits, but Elsevier has been particularly problematic. They sponsored weapon fairs, created and sold ‘fake’ journals to pharmaceutical companies to publish ‘results’ supporting their drugs, purchased the Social Science Research Network and created paywalls or removed legally shared working versions of articles, charge fees for open access articles, and actively lobbied against open access legislation (For a concise summary with links see Tal Yarkoni’s blog entry). This brought massive counter movements against Elsevier in the scientific community (for example, The Cost of Knowledge). You can take action and refuse to review for or publish with unethical publishers if you feel it is justified. Thus, you should inform yourself about the publishers. Your libraries are a source of information, because they deal with the business side of publishers.

If you are in Europe, check if your institution is a signatory of ProjektDEAL. A consortium of universities are collectively bargaining with publishers via ProjektDEAL demanding that publishers reduce fees and eliminate the double paying of universities. The primary objective is that publishers sign country-wide subscription agreements that enable access for all universities at once. Wiley agreed to such a model and this marks a paradigm change. It indicates how the publishing industry looks in the future, so long as the OS Movement proceeds. If you are not in Europe, consider starting a similar initiative, for example the entire University of California system of 10 universities, 5 medical centers and several research institutions that collectively produce roughly 10% of the world’s academic publications recently followed ProjektDEAL and boycotted Elsevier.

You can work around the publishing business. Prior to submitting an article or after it is published, you have the right to share a preprint – a draft of the paper you share publicly so long as it is not published elsewhere or sold for profit. Posting preprints reduces the power that publishing firms have over science, in addition to giving others immediate access to your work. But simply posting preprints on your academic website is not open enough. Use a preprint service, for example through the Open Science Framework, to ensure that your preprints appear in search engines such as Google Scholar. SocArXiv for example, is the go to location for sociology. This enables scholars to find and directly access research results based on the words they contain, uninhibited by paywalls – a crucial aspect to practicing sociology in the Global South. Preprint services are free and open access.

Meta-constructing social theory

Certain hypotheses are constantly tested in social science. The impact of income inequality on health, racial bias on police brutality and public opinion on elections, just to name a few. At some point more tests of the same hypothesis stop contributing to scientific knowledge, and may even harm it by introducing more ‘noise’ into the scientific discourse.

I study social policy preferences and the impact immigration has on them. In this area there has been sustained efforts to test the hypothesis that immigration has a negative impact on support for policies of the welfare state; things related to protecting against risks of aging, unemployment and health. To justify this hypothesis, scholars construct theoretical variations of group dynamics arguments, often drawing on resource competition, nationalism and social identity. Despite claiming to test the hypothesis, the formal models applied to data suggest any number of data-generating processes. They often have little in common other than some measure of immigration and some measure of policy preferences. The results of their tests go in all directions, i.e., a positive, negative or nil effect of immigration. It would appear that the topic is at a standstill, new analyses of the same handful of cross-national survey data sink in the mire. How to break through such a scientific impasse?

In designing the Crowdsourced Replication Initiative (CRI) with co-PIs, Alexander Wuttke and Eike Mark Rinke, we asked researchers to to do research; and we gave them semi-structured tasks and observed them. Specifically they were supposed to come up with the best possible way to test the immigration hypothesis given the same International Social Survey Data source. Although we are currently meta-analyzing the hypothesis test-results (see our virtual APSA poster) to determine which modelling decisions impact the outcomes, we also have a second goal in mind: to discover what is behind the specification curve.

Each research team had to design a best possible test. This is at once a statistical question and a theoretical question. They needed to think carefully about the data-generating process and attempt to recover it in a model. We asked them to write down their research designs after doing this thought exercise, but before analyzing any data. From their researcher choices we can identify where key consensus and disagreements exist about the data-generating model, thus is not only evident in their designs but also in a structured deliberation and voting procedure. This process offers a major advantage over ‘normal’ theoretical discussion and debate among academics, because we have the results that go along with the different modeling choices; and, let’s be honest, when else do over 150 researchers get together and focus on a single hypothesis? By observing this process we can identify where data-generating theories differ and how important these differences are for the results. This will allow us to map where immigration and social policy scholars should focus their theoretical efforts in the future to reduce the most uncertainty, i.e., the largest gains in knowledge.

We have a sound piece of scientific research from Brady and Finnigan (2014) from which we draw our working hypothesis for the CRI crowdsourced researchers: That immigration undermines support for social policies. Brady and Finnigan found little or no support of this hypothesis, at least not in a generalizable macro-comparative sense. This was the launching point for the research of the 77 teams who by now managed to submit replicable results (yes there are still a few out there we are hoping will submit a final model or fix issues we identified in our replication of their models).

Although we are in the process of analyzing the ocean of data generated by this project; a sneak preview offers exciting evidence of the possibility for meta-construction of theory.

Here are two glimpses of what’s to come. One are the deliberation and voting results summarized (Figure 1). The other are differences in definitions of ‘immigration’ (Table 1). We used Kialo, an online structured deliberation platform, to allow participants to discuss the data-generating model after they proposed their own ideas for how to best test the hypothesis. Readers can observe how this deliberation unfolded as we divided the participants into two groups: here and here. Later (after they had the possibility to update their models based on the deliberation) they were given other teams’ models or our own variations on those models to vote on and rank in terms of their appropriateness for testing the hypothesis without having seen the results of those models. Figure 1 quantifies both the Kialo veracity scoring and survey-based voting into one overall scale and then plots the average score of models by their features. Each different color is a discrete set of model features with the zero (y-axis) set to the average support of models choosing an OLS estimator (among the least preferred).

Figure 1. Researcher Preferences for Recovering the Data-Generating Model
“Model” is the hypothesized general impact of immigration on support for social policy. Data and code still being prepared for online sharing, stay tuned.

In Figure 1, it becomes clear looking at the longest bars in each color category that models that incorporate all 5 waves of the ISSP data, include countries of Eastern Europe, include heterogeneous error variation by country-year and year (like a cross-classified model), and incorporate survey sampling weights are preferred over the others. Some of this runs counter to the state of the art. For example, most research follows a logic that major immigrant destination societies – the “Rich 13” and “Rich 17” advanced democracies – should be where “public opinion is likely most influential for the politics of social policy” (Brady and Finnigan 2014:24).

To summarize the motivation for looking across all possible countries, especially Eastern Europe, one crowdsourced researcher put it like this: “Either there is an effect of ‘immigration stock (increase)’ or not“.

Another followed up on this point stating: “To test the general hypothesis we should use as many countries as available and account for variations in GDP and social welfare expenditures in the models.”

These comments demonstrate the majority voice in the CRI that if immigration has a an impact on social policy preferences we should see it across all countries of the globe, not restricting our analysis to only very rich, strong welfare states.

Although Brady and Finnigan and all other research in this area comes to no consensus on whether there is a negative impact of immigration on support for social policy preferences, we should remain skeptical of results if we do not trust the data-generating model. In other words, if our tests do not match what most researchers see as the appropriate theoretical perspective, results are inconclusive and thus uninformative. The deliberation and voting offer us clues where to focus theoretical effort, namely specifying why more countries of the world should (or should not) show a causal effect of immigration on social policy preferences and whether this should (or should not) appear across several decades or only certain times. I am not aware of extensive theory that attempts to tackle these issues. Now is the time to write it!

Even more productive for the possibility of meta-construction of theory is the correspondence between the actual decisions made by the researchers and the subjective and objective outcomes of those decisions. Again, our results are in progress, but we offer a snapshot in Table 1 of different ways the researchers chose to measure immigration as their main hypothesis test variable (1 out of dozens of model decisions to compare). In the first row, 67 out of 77 teams used a “Stock of Foreign-Born” measure in at least one of their models, and 27% of their models using the “Stock” variable showed support of immigration having a negative and significant statistical impact on support for social policy at p<0.05.

Table 1. Crowdsourced Researcher Decisions, Deliberations and Results.
Five different measurement strategies for the immigration test variable.

In the column ‘Positive Test Result Rate’, we see that the ‘Difference’ between “Stock” models (referenced as [1] in Table 1) and those instead using “Flow” to measure immigration models (referenced as [2]) is 3.6. In other words, “Stock” models arrive at support of the hypothesis 3.6 percentage points more than “Flow” models, all else equal. “Stock” models were not more or less popular than “Flow” models, with the average vote score of 0.43 on a scale of 0 (worst) to 1 (best equipped to test the hypothesis) versus 0.45 for “Flow”.

The values in bold indicate that “Change in Flow” models (those measuring derivatives of “Flow”) were among the most popular in the voting process. So the rate of change of the flow of immigrants is seen as an important component in testing this hypothesis. Interestingly, these models were 4 percentage points more likely than “Stock” and “Flow” models to support the hypothesis. When measuring immigration as specific to certain outgroups (from Muslim-majority countries, non-Western countries or refugees), the “Flow” of these various ‘Outgroups’ was seen as more popular than “Stock” of ‘Outgroups’ by a large margin, but the results were over 10 percentage points less supportive of the hypothesis.

What can we learn from this. We argue that a full analysis of the massive range of modeling decisions will give us a guide to move this entire research area forward. Some other decisions for example were different social policy domains, whether ethnic and fractionalization is the ‘real’ cause of the ‘immigration’ effect, construction of latent social policy preference measures, whether or not GDP and unemployment are part of the data-generating assumptions just to name a few out of hundreds. We are only scratching the surface here, but it seems that observing researchers make research decisions, deliberating them, voting and making final choices, we will gain immense knowledge as to where better theory is necessary. As such we see meta-constructing of social theory as a promising avenue for social science. This would be the concept of theory designed replication writ large.

Help. We can’t know the data-generating model.

Lack of data-generating models. A problem ransacking macro-comparative research efforts. We got strong data-generating theories. Institutions, politics, materials, procedures, conflicts and symbols help explain macro-level events, like the introduction of social security or the onset of war. But we can’t model these theories as causes of the data we observe. There is far more theoretical complexity than there are countries.

There just aren’t enough countries to provide reliable measures of central tendency given theoretical complexities. Even if our theory were as simple as Y will be significantly different from some value (e.g., zero) given X, we should have at least 10 countries. If we were to actually conduct a power analysis it would tell us we need more like 200 countries depending on how big of a difference we expect in Y given X. I’m just gonna leave aside that countries are not a population in the typical sense (what is a country again exactly? exactly.).

What should we do?

  • Option A: stick to qualitative case studies.
  • Option B: continue running various country-level regressions and ignore the problem.
  • Option C: wait until data from 180+ countries covering a period of 50 years becomes available.

Maybe we just ask a better question. Why do we need to do macro-comparative analysis if we already got good theory?

Answer: Reliability and utility.

Without systematic evidence we don’t actually have theories, only ideas or conjectures. Logic and in-depth case study help define data-generating theories for one country in one time period. So we want to test if this might apply to most countries in most time periods. In testing this we get Pomeranian.

Even with the same variables, regressions on country data over time go all over the place given only tiny changes to the estimation procedure. That was a finding of the CRI. The effect of immigration on social policy preferences could be anything.

To the right: Mr. Summerbottom showcases the size and direction of effects from macro-comparative regressions using the same data. Outliers cropped.

Even if we run all possible model configurations, and even if we consider Bayesian posterior distributions. We still don’t know if we have a correct model specification, because we cannot truly test it. We can only make sweeping statements like, ‘in all possible model combinations, variable X was not significant so it probably does not have an effect’. Too bad we might have some kind of suppression from an unobserved variable!

Don’t drink the water. There is no method that can substitute for human logic. Zombie modeling where the researcher lumbers aimlessly through thousands of models looking for blood doesn’t work. The answer is therefore Answer D, none of the above.

We need new school meta. We need to logically analyze researcher’s decisions. We can’t do much with the limited data that are out there for macro-comparative research. Thus, we need to get down with specifications. By comparing specification curves and researcher decision trees, we can identify the difference between critical and benign. Some programmers offer us R packages for this (thanks Joachim Gassen for rdfanalysis!).

Some models come to the same results regardless of whether the researchers apply weighting, use latent variables or correct for autocorrelation; others not. When we identify the decisions a researcher made that carry the potential to influence the results to the greatest degree, we simultaneously identify where to focus our theoretical work. These decisions cannot be improved by running more models, instead they require, as Andrew Abbott once put it, ‘more sitting in our offices and staring at the wall’. We need to dig in and use logic, reflection and mental energy to improve our theories in small-N macro-comparative research.

Better than a computer. A crowd of researchers.

Crowdsourcing researchers. A new use for an old tactic. When Silberzahn and colleagues asked if football referees are skin-color biased in their assignment of red cards, they brought a new meta to social research. Rather than the typical one research team, one project; they got together twenty-nine teams. All got the same research question and same data. What can we learn?

Hold on. A single academic strapped with programming skills can run just about every possible statistical model configuration. Level up this academic and they can program machine learning routines to tell us almost anything we need to know. So why do we need a crowd of researchers to analyze the same data?

The answer: theory.

Computers do not do theory. Computers crunch data. They can’t tell us where the data came from. Running every possible in an effort to ‘test robustness’ means testing for things that do not exist. Even worse, it means taking the results of tests for things that cannot possibly exist and using them to draw conclusions on the robustness of a test for something that could exist. Throwing all possible variables into a model or ordering variables in every possible configuration will maximize statistical predictions, yes. But what good is predicting something that cannot exist?

An example: Policymakers decide they want to increase the number of females in society. A computer determines that bearing children is a good predictor of being biologically female. Now, imagine that in a statistical model, being pregnant or ever having had biological children explains 75% of the variance in the sex of a given population (its probably about 100% at this point in human history, but lets allow room for error). The policymakers thanks to the computer, conclude that if 1,000,000 more people were pregnant, a predicted 750,000 of them should be female and only 250,000 male, plus a margin of error. Thus, they conclude that getting 1,000,000 people pregnant is likely to increase the number of females by 250,000 in their society, assuming the society is currently 50% female.

Epic fail.

If we want to develop causal theories that provide useful knowledge for societies, human logic is necessary. Rather than having one computer report 8.8 million false positives after running 9 billion different regression models, humans can identify correct, or at least ‘better’, model specifications. In fact, we would not even know they are ‘false positives’ without a human to understand that getting someone pregnant does not turn them into a female, for example.

But relying on the logic of one human, or even one team of humans, is risky. There are researcher degrees of freedom that make their research unreliable. Their prior beliefs, knowledge, experience and context lead to variation in results. These priors lead them along different paths through the garden of research. With the same research question and even the same data researchers often come to different results as demonstrated by the Silberzahn study and our Crowdsourced Replication Initiative (CRI).

Sounds like a meta-problem. So what good is crowdsourcing if we cannot rely on the crowd?

Answer: social interaction.

Crowdsourcing when done with careful planning and central organization, allows participants to comment on, if not deliberate, each others research choices. Suddenly meta-uncertainty turns into the power of meta-logic. Not just one team and their narrow ideas, but a communal debate with diverse inputs. Both the Silberzahn et al study and our CRI involved deliberation and voting on research designs. Combined with the growing area of specification curve analysis, crowdsourcing increases credibility for social research at the level of the population under study and the meta-level of the researchers themselves.

Finally, the relevance of crowdsourcing for collaborative theory construction is an untapped but promising avenue for the future. Crowd research departs from the current system that favors individualism by rewarding novelty of individual researchers. Crowdsourcing instead isa system of consensus building and direct responsiveness to theoretical claims. It could resolve the perpetual problem of scholars, areas and disciplines talking ‘past’ each other. If not consensus, it can identify critical unresolved questions to guide future research.

Crowdsourcing can move us toward open science in the Mertonian sense of a communalistic endeavor. To achieve this, all participants should be co-authors on the project, get to discuss each others’ models and theory, and get to update their own results during the process. We need machines to facilitate this kind of large scale research, but they cannot produce communal, logical exchanges. For that we need to stick with the crowd.