Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

Not Wrong: The Replication Crisis is a Metatheory Crisis?

Wanna know what that is? Its theory.

We talk about reproducibility as a methodological problem. We debate p-hacking, publication bias, and analytical forking paths. But the elephant just chillin in the corner is that our theories are weak. There are too many other plausible explanations for what we observe and measure. Sure, the things we measure are complex, but without more clarity, even Commisioner Odenthal isn’t gonna find us a clue to what we are testing. Meaning that we don’t really know what to make of our findings. Other than hoping they are publishable and packaging them neatly to increase those chances. The findings are not right or wrong, they are not really interpretable. At least not with weak theory.

Part of the problem is that unlike the ‘harder’ sciences, consciousness is in the equation. We are measuring phenomena far more complex than quantum physics. Another part is simply that we do not spend much time on theory. In one of my favorite esoteric and less-mainstream open science writings, Anne Scheel pointed out “Why Most Psychological Research Findings Are Not Even Wrong”.  Anne, I see your discipline specific arguments and raise you all social and behavioral science disciplines. We often don’t know what the estimand is. We don’t know what we are looking at. If our theories don’t explain the phenomenon in the first place, debating whether an empirical finding is statistically “right” or “wrong” is moot.

To confront this, I pitched a method I’m working on with my former doctoral student turned postdoc Hung H.V. Nguyen. It’s called Metatheoretical Multiverse Analysis (MMA) and I wanna kick some ‘theoric’ with it (cool huh? It rhymes with lyric and could mean theory in action…). What follows is a summary of the ensuing debate, full of interdisciplinary friction, economic curmudgeonism, and philosophy of science moments.

The Pitch: Metatheoretical Multiverse Analysis (MMA)

To understand why metatheoretical multiverse analysis is necessary, we have to look at how we currently treat plausible theoretical arguments about the data-generating model… Say what? I mean how we use and apply logic.

Imagine you are testing the effect of X1 on Y. But there is an unobserved confounder, X2, which causes both X1 and Y. If X2 is not in your test, your results are uninterpretable. You don’t know if X2 is confounding the results. And, don’t even get me started about X3. This might be a collider. But the theories are not clear in your subfield. So when we try to compare them we end up with several conflicts or unknown paths. These are all alternatively plausible theories and from them we have a multiverse of theory.

Now, imagine I have five equally plausible alternative theories explaining the same phenomenon. This means there’s a 20% chance the theory I am using is the correct one. But I cannot imagine any social or behavioral science where there are only five plausible theories. There are likely thousands. We only test five because it takes entire careers to write semantic theory. And since it takes 20-40 years for most of our discipline to forget those long-winded theories of the past maybe this is a sinking ship…. I digress. Let me re-gress instead (fancied opposite of digress which happens to be a statistically procedure too). 

This is where Metatheoretical Multiverse Analysis (MMA) comes punching (or wrestling or kicking) in. If a standard multiverse analysis runs all reasonable empirical specifications for a given dataset, a metatheoretical multiverse analysis would logically compare all these theories. Somehow…. That’s where the idea gets a little sticky. Metatheory is mostly something people write about, rather than try to formalize and analyze with math. But if they could, our method would then show them where they need to invest their theory-building efforts.

I opened the floor. The timer started.

The Economist’s Dilemma: HARKing and the Illusion of Theory

The pushback was immediate, insightful, and brutally honest. The first counterargument highlighted just how difficult theoretical work is. As one applied economist noted, reading James Heckman makes you realize how brilliant deep theory can be, but spending your career trying to come up with sufficient conditions to definitively disentangle two competing theories is a great way to find yourself out of academia before you get tenure. Remember the long-winded argument?

But a more damning critique came from the reality of how theory is actually utilized in modern economics. To commit a “statistical sin” and assign causality, one researcher pointed out that the lack of robustness in our fields might stem from the fact that HARKing (Hypothesizing After Results are Known) is practically the norm.

In economics, the workflow rarely starts with a pristine a priori theory. Instead, a researcher finds an intriguing empirical pattern in the data, and then builds a formal mathematical model to justify the findings. We only take the time to do the exhausting math if we already have a paper we want to publish. Often, the formal model is something requested by Reviewer 2 at a top journal ex-post. Because the theory is engineered to fit the data, it doesn’t improve our prior hypothesis in a meaningful way, nor does it guarantee robustness when exposed to new data.

Interestingly, someone from the I4R team chimed in with data to back this up. After checking roughly 15,000 robustness checks in the I4R database, they ran an AI classification to score papers from 0 to 100 on ‘how economic’ they were (i.e., true economic theory vs. a paper on TV habits published in an econ journal). The finding? There was absolutely no relationship between the “econ-ness” of the paper (the presence of formal modeling) and its robustness.

So, the means that even if we were to invest in theory development, we wouldn’t get anywhere because… basically, we suck at it.

The AI Revolution

These days there’s always gotta be something about AI. If not everything. The egomaniac AI that I asked to help me write this just couldn’t wait to point out how many people were talking about it.

Funnily we came to an AI discussion through a strong defense of the structural approach in economics (after the initial dust had settled and we could again breathe the cool Barcelona air-conditioned summer air). Economics, one participant noted, used to be purely theoretical because, prior to the 1970s, we simply didn’t have the capacity to handle large data. The empirical revolution brought an era of data-mining into play. P-hacking came into full swing. The academic industrial complex had taken over thanks to secondary data spewing forth from society.

But the tide now is actually flowing back toward structural modeling. Thanks to… AI? The massive amounts of data we have today, combined with AI’s unparalleled ability to explore patterns, is killing the human comparative advantage in pure empirical data mining. AI will always be better at finding patterns (at least an AI or data scientist tells us this, Gemini still can’t perform better regressions than I can IMO). The only place human researchers will retain a comparative advantage is in thinking, being creative, and structuring the data generating process. We don’t need one perfect theory to describe reality anymore (not that that ever worked for us) good approximations are good enough now – and good approximations means…. Drum roll and someone on the mic saying “Ya’ll ready for this?!”. PREDICTION. The better we can predict things, the less we will need to explain them, they will just become some sort of facts in our life worlds. Won’t they?

Metatheoretical analysis might just be the structural framework we need to survive the AI transition. I mean, I invented it, so I’d like to think so.

Bias, DAGs, and Kung-Fu Nancy Cartwright

As the 90-second buzzer kept interrupting and resetting, the conversation evolved into the philosophy and sociology of science itself.

From a legal and equality perspective, an important question was raised about those “five theories” we tend to rely on (you know that when equally plausible reduce the chances that any one is correct down to 20%). Where do they come from? Historically, they have been generated by a very specific demographic—predominantly white men from the Global North. By restricting our empirical tests to a handful of established theories, we inadvertently perpetuate biases and ignore alternative paradigms that might emerge from the Global South. A metatheoretical multiverse approach, by automatically generating and considering thousands of models, might offer a mechanical antidote to this historical bias by forcing us to acknowledge the vast space of un-theorized realities.

But are DAGs really capable of saving us? The room had its doubts. Hey Mister Jack… I’m talking to you. The world, as one researcher passionately argued, is not a clean DAG with X1, X2, and Y. It has X15, Y7562, and a zillion unobserved mediators. Furthermore, DAGs can’t handle cyclical relationships and feedback loops the most famous relationship in economics, price and quantity, is entirely cyclical, endogenous.

This brought us to Nancy Cartwright and the philosophy of science. If we are mapping out thousands of theories, we must remember the Popperian ideal: a model must be falsifiable. If a theoretical model cannot be thrown out under certain conditions, it ceases to be a model and becomes a religion. Metatheoretical multiverse analysis is only useful if we have the empirical tools to actually falsify the branches of the multiverse we generate.

The crow cheers with Nancy’s MMA kick to the face.

The Preregistration Battleground

You cannot talk about theory and open science without stumbling into the debate on preregistration. I posed the question: ‘Does preregistration inappropriately constrain our theories?’ because I want to seem smart and provocative. If there are hundreds of theories, forcing a researcher to write down a specific one beforehand essentially chokes the multiverse before it can breathe. Choking is forbidden in MMA by the way.

The responses were polarized:

Some argued that preregistration forces a hypothetico-deductive model onto fields that generate knowledge inductively. In economic history or sociology, research is exploratory. You learn from the data. Preregistration actively harms inductive discovery.

Others pointed out that preregistration is incredibly difficult if not overrated for secondary data analysis. Datasets are messy, collected for non-research purposes, and require deep exploration just to understand how missing variables are coded. Recently someone pointed out on LinkedIn that my praise of Neumark is overrated. An MMA sweep kick, totally permissible in the sport – touché.

Conversely, a third group (there’s always a third group otherwise it feels incomplete) advocated that preregistration is simply a record of where you started. It prevents the ex-post invention of stories and increases transparency. If ‘Mostly Harmless Econometrics’ had a chapter telling students to pre-specify their hypotheses before opening the dataset, it would have transformed the culture of economics entirely. To late, the historical institutionalists won.

Where Do We Go From Here?

As we wrapped up the session (and prepared to face the blistering heat outside in the name of finding Fideuà), the consensus was clear: our methodological tools have far outpaced our theoretical foundations. We have built incredibly sophisticated empirical engines, but we are putting them in theoretical chassis that are fundamentally flawed. I liked this shift. Oh wait I kinda led the discussion there.

Metatheoretical multiverse analysis is not a magic bullet. It will not solve the fact that social science involves the unpredictable chaos of human consciousness, nor will it easily map the cyclical, non-DAG-friendly feedback loops of the global economy. But it is a start. It is a way to stop pretending that our opportunistic, post-hoc theories are the only valid models of reality. It might trigger some epistemological soul searching if nothing else. By mapping the vast space of plausible theories and systematically testing where they align and where they conflict, we can begin to rebuild the credibility of our disciplines from the ground up – in theory (which is probably weak, so take it with salt).

Furthermore, as the discussion highlighted, we need to completely overhaul how we incentivize work. Open science has an image problem we often make it look boring, framing it as an adversarial compliance checklist rather than a thrilling pursuit of truth. We leave it up to early-career researchers to awkwardly teach their supervisors about open code and data. And shoulder them with the onus of spending hours making all their work open and reproducible, hours that their forefathers (yeah, they were 95% men) didn’t need to do. Heck these forefathers could publish like 2 papers per year and get tenure.

But since its my blog post, I get the last word: If we want to fix the replication crisis, we cannot just mandate better code and policies. We have to foster better theory. We need to acknowledge the sheer size of the theoretical multiverse, embrace the complexity, and start theorizing before regressionizing. In this case, sadly, the answer is not as simple as 42.

The path toward ethical science is paved with diamonds

Over the last two decades awareness of a Reproducibility Crisis penetrated all scientific disciplines. Studies recently showed that 90% in STEM and 95% in psychology were aware of this Crisis1,2. At its core this is a crisis of trust. The findings scientists publish and tout as facts, are often not replicable by others. Moreover, a great many are not even computationally reproducible using the original data3,4.  The Open Science Movement developed partly in response to this Crisis5,6. Especially in the last two decades, researchers organized grassroots movements to make science more transparent, reliable and ethical.

One of the many tenets of the Open Science (OS) Movement is open access. Published research results must be accessible to everyone. This is not a new idea. In the 1940s Robert K. Merton developed normative goals necessary to ensure integrity in science7. One was that all scientific findings should be public property. He argued that effective and efficient progress of science depends on open public access. In data science this Movement led to the concept of Open Science by Design – a strategy where scientists plan in advance how they will make all metadata, data and algorithms available and easily re-usable by anyone8.

The OS Movement has been successful. Open access journals and articles increased rapidly. They outpaced the global increase in publications in other formats in the last two decades9. One problem is that these open access articles are overwhelmingly funded by the authors of the studies via Author Processing Charges (APCs). It costs over 10 thousand US$ to publish a study open access in the journal Nature, and most other journals charge at least two thousand. Except for elite institutes and projects with generous third-party fundings, these fees are not usually covered or coverable by universities. A naïve outsider might ask why scientists would pay astronomical fees to make their own hard work publicly available, when they can simply share it online in a free repository like the Open Science Framework, Github or any number of preprint servers based on the ArXiv model?

The answer is competition. The strongest norm governing the practice of science today is publish-or-perish. In every field and science at large, scientists are judged based on their publication record. Ask any academic what the top journals in their area are, and you will get relatively consistent answers by discipline, sub-field and science in general10. Look at the faculty of any ‘top’ university, and the faculty in any discipline will have one or more publications in these ‘top’ journals. Without publications in these journals, a scientist has no chance at a career in science. This is a collective cultural fact. One deeply institutionalized in the organizations, rules, norms, expectations and behaviors of scientists11,12.  

Thus, the founding of new open access journals with lower APCs has little impact on the scientific enterprise because they cannot compete with institutionalized legacy statuses of existing journals. The greatest success story is the non-profit publisher PLOS, which rose swiftly in the rankings, but could not crack into the very top tier. Although far cheaper than Nature, it is still not ‘cheap’, with APCs in their family of journals ranging from 2.5 to 3.2 thousand $US[1].

The most successful open access models are within existing paywalled journals where authors have the option to publish “gold” open access, rather than publish for free. The payment of somewhere between two and 10 thousand US$, gets authors the right to have their single article published open access inside of a closed access journal. This means that despite a massive shift toward open access, the fundamental structures of scientific publishing have not changed.

Ethical Implications

Competition in science supports motivation and innovation, but the publish-or-perish norm is a toxic externality. Scientists willingly prioritize subjective journal ranking, impact factor and increasing their citation counts to get ahead within the competitive scientific enterprise. They do this despite widespread skepticism and evidence that rankings and citations do not correlate strongly with the quality and reliability of published studies14–16. They do this because they must, or at least perceive that they must, in order to follow their scientific career aspirations. The pressure to publish thus motivates rent-seeking behaviors designed to increase publication chances, i.e., career chances. Conducting higher quality science is one of these behaviors. But there are many methods to increase publication chances that are science orthogonal, what are known today as questionable research practices (QRPs)17,18.

Some QRPs are unconscious and learned from supervisors in the process of converting research into a publishable paper. For example, researchers routinely run many statistical models but report only those that show the strongest support of their claims. Researchers who believe that their claim is true in the first place will gravitate toward models that support it, convincing themselves intrinsically that these models are the best tests of their claim.

Decisions based on confirming intrinsic beliefs or window dressing for peer reviewers reduce the replicability of science. They narrow down the multiverse of potential findings into a highly selected set of results. This selectivity is independent of the process that generated the data in the first place, in other words, it is science orthogonal. A classic example is a study published suggesting that hurricanes with feminine names cause more damage than those with masculine names. It turns out that using the available data, the original researchers selected a model that produced regression coefficients that were extremely far away from the central tendency among all other plausible models’ regression coefficients19. We do not know whether this was a conscious decision but their reporting hides the truth and simultaneously increases publication chances.

There are of course conscious and highly unethical behaviors leading to an entirely false representation of reality. Science is filled with scandals of hacking and data-faking. Rent-seeking alone can explain unethical learned unconscious and conscious behaviors of scientists. They seek status and money and job security, or in some cases seek results that support a particular worldview or policy outcome20,21. But rent-seeking is unambiguously the main reason22,23.

If we as a scientific community and science-interested public, want to eliminate the perverse incentive structures that bound and inform rent-seeking behaviors of scientists we need radical change. Status should not be assigned based on journal metrics that often have little to do with the quality of the research being conducted or published and more to do with legacy and embedded norms. The most radical proposal is to eliminate journals altogether. But this is an extremely unlikely outcome no matter how powerful the OS Movement becomes.

Journals became standard in science hundreds of years ago as a means for communicating scientific discoveries across time and space24. They enabled scientists to acquire knowledge without travelling to faraway universities. Publishing was not cheap, and with the dawn of digital media, it became even more expensive as publishers raced to provide their journal both in-print and online. The costs associated with publishing gave publishing firms a great deal of power over time.

The result is that companies like Springer Nature and Elsevier have enough power to dictate to scientists how they perform their research and communicate their results25. They shape academic careers, institutional priorities, governments’ science policies, and perceptions of journals and the publishing enterprise among scientists26. Their primary legal interest is their shareholders. Profit is their priority, and only second is to provide a service to science as their product. A simple mathematical proof confirms that profit is their priority: If they cannot make a profit they will no longer provide the scientific services but if they can make a product without scientific services they have no reason to stop.

Although big publishing firms have shown many draconian practices in their ‘service’ to science27–30, scientists still need a means to communicate their results with each other and the public. This can be done without for-profit publishing but probably not without a journal publication format, or something very similar. For example, we now have a plethora of ‘green’ open access preprint servers where scholars can deposit working papers, or prior versions of their published articles. These are a viable means to communicate science without the perverse incentives generated by big publishing or institutionalized journal rankings. This all still involves the journal article as the standard unit of science production.

Diamond Open Access

If we are ‘stuck’ with journals in science as our primary communication medium, then they should be as free from perversely incentivized bias as much as possible. The first step is thus to remove for-profit publishing from the equation. It is fine to use the services of for-profit publishers. They have shown the capacity to provide print on demand, marketing and scientific journalism. But when they control and direct the scientific enterprise when have an ethical conflict of interest – namely profit versus robust science.

The costs of publishing a journal are the lowest in history. There are publication kits that help associations, institutions and stand-alone journals to take publication into their own hands31. With minimal costs and self-governance, journals have no need to push institutions to purchase journal subscriptions and no need to charge authors astronomical publication fees. Thus, they themselves are not perversely incentivized to perpetually increase their status to make their product profitable.

The optimal existing solution is diamond open access, whereby a journal charges no APCs and is freely readable and downloadable online32. It is the most ethical and equitable by design, and is perceived as the ideal model by most academics33. The OS Movement is overwhelmingly in favor of diamond open access34, as are governance bodies – at least those free from the influence of big publishing like UNESCO35,36. Despite 13 thousand journals indexed in the Directory of Open Access Journals (DOAJ) with no fees as of February 16th, 2026, these journals are not on the radar of most indexing services, and do not belong to the mainstream of journals published by major scientific societies37. This means that successful diamond open access journals are very rare.

Because of such a saturated scientific ‘market’ for publication outlets, starting new diamond open access journals has had little impact on producing a more ethical science. They do not gain reputation. The reason PLOS was so successful, despite failing to break into the highest echelon of science, was because it has a huge cash flow and can use it for branding and promotion. From an ethical science perspective it is valuable, like gold, but it is not as valuable as diamond.

Flipping or Starting Over?

To have the highest ranking and most well-known journals diamond open access, societies need to cancel their contracts with for-profit publishers38. One problem with this is that many academic societies are themselves run like for-profit businesses. They seek to generate as much revenue as possible, and diamond open access would threaten this model. They have embedded relationships with publishers and agree to renew contracts together, mostly independent of the scientists they serve. I witnessed this first hand with the American Sociological Association and Sage39,40. Contracting a for-profit publisher provides a non-profit academic organization a scapegoat for amassing capital via subscription fees.

There are incredible exceptions; however, and they offer model success stories. Computational Linguistics is one of the earliest journals considered to be among the top in its field, to flip to diamond open access. The first step was a move by the Association for Computational Linguistics in 2002 to create the CL Anthology which made all articles published by association journals open access online after an embargo period. Then in 2009 the journal Computational Linguistics flipped to diamond open access. The journal Demography of the Population Association of America is another example.

The embeddedness of big publishing in science, and the capital that associations can raise through their journals when run by big publishers, are major barriers to flipping. Another major barrier is that some big publishers coerce scientific societies into signing away the rights to the titles of their journals in their publishing contracts. This tactic is most intensively deployed by Elsevier. They own the rights to the titles of nearly all the journals they publish. Therefore, when these societies want to change publishers, they cannot. They are trapped. They would have to start a new journal with a different title – what happened for example with the Journal of Infometrics41. There is otherwise no way around this problem because of copyright law.

Therefore, collective efforts to build the popularity of new diamond open access journals would greatly increase the movement toward a more ethical and effective scientific enterprise. If these are journals from societies, which are essentially the same journal but with a new (not copyrighted by Elsevier) name, it requires authors to support this journal and immediately abandon the other. Supporting new diamond open access journals, whether completely new or newly named, requires established scholars to put their status-seeking egos aside in the name of scientific progress and ethics. It is precisely those who have built major scientific reputations who need to engage in this change, because they can give them most clout to new journals by touting them and publishing in them.

Scholars and societies alone cannot carry the burden. Hiring committees need to reward these behaviors. Rather than seeing a publication in a new ‘unranked’ or ‘low ranked’ diamond open access journal as a sign that the article is ‘not high quality enough for top journals’, committees should judge publications only on their content. Then, if two publications are seen as equally scientifically rigorous and high quality, the one in a diamond open access journal should get a greater weight. Scientist who consistently publish high quality research in diamond open access journals should be favorites of hiring committees, all else equal.

A hiring committee would be unlikely to reward an applicant who shows sociopathic behaviors. Following this logic, they should not reward scientists who show behaviors that go against ethical science by practicing closed and profit-incentivized science. I need to be very clear here that I personally am still publishing regularly in journals published by for-profit publishers. I am not that famous, and I do not have tenure. This is a perfect example of why we need sweeping changes at the institutional level, and from the top of the scientific hierarchy.

With efforts to flip both  existing journals and efforts to reward new, ethical journals, we give science its greatest future chances for improvement and sustainability. There are many efforts underway and these should serve as guides, for example the Diamond Open Access Fund from the Dutch Research Council and MIT’s shift+OPEN initiative.

Diamond AI

Diamond open access has a second, equally important role for the future of science. It is necessary to inform Generative Artificial Intelligence (Gen AI). The LLMs that power popular Gen AI are trained heavily on corpora assembled from what is available via the Internet. Paywalled literature is missing from training corpora, not because it is unimportant, but because it is not accessible or legally usable42. Gen AI outputs are therefore based on a highly restricted sample of all scientific knowledge. Open access for all of science would ensure that anyone using Gen AI, would get the best possible information based on all that we know as humans. As essentially everyone is using Gen AI today43, this would mean that the public would be optimally informed.

It is not necessary for Gen AI development that all articles are diamond open access, they just need to be somewhere, e.g., green or gold open access. Yet, without diamond open access the entire knowledge enterprise will continue to favor the work of those with greater resources44. Resources are of course necessary to produce higher quality science because of research costs, but when it comes to publishing and dissemination, a resource advantage reproduces the already existing Global North-South disadvantages in science. This limits human capacity to tap resources in lower income societies. These are societies filled with potential contributions to science that could benefit both 1) their own societies’ development – because it enables them to gain status, resources and build stronger, more attractive and sustainable scientific institutions, and 2) all of scientific knowledge because there are brilliant minds waiting to be tapped that might otherwise give up on science because it is for them no sustainable in their region.

If the knowledge in Gen AI remains Global North and WEIRD biased (Western, educated, industrialized, rich and democratic), the result is that these countries remain culturally and economically advantaged beyond that which exists presently. This means that without diamond open access norms, we are willingly allowing Gen AI to increase cultural hegemony and economic domination of a minority. I am not directly arguing whether this is good or bad. I am in the Global North and profit from a stronger Global North science advantage. My argument is that science should be neutral, favoring the most optimal and reliable knowledge, rather than legacies or strategies that game the scientific system.

References

1.          Baker, M. 1,500 scientists lift the lid on reproducibility. Nature 533, 452–454 (2016).

2.          Metskas, A. How Much Do Academic Psychologists Trust Academic Psychology, and Is There Still a Replication Crisis? Transparent Replications https://replications.clearerthinking.org/how-much-do-academic-psychologists-trust-academic-psychology-and-is-there-still-a-replication-crisis/#survey-demographics (2025).

3.          Open Science Collaboration. Estimating the reproducibility of psychological science. Science 349, (2015).

4.          Breznau, N. et al. The reliability of replications: a study in computational reproductions. Royal Society Open Science 12, 241038 (2025).

5.          Engzell, P. & Rohrer, J. M. Improving Social Science: Lessons from the Open Science Movement. PS: Political Science & Politics 1–4 (2021) doi:10.1017/S1049096520000967.

6.          Breznau, N. Legacy of Jon Tennant, “Open science is just good science”. Crowdid https://crowdid.hypotheses.org/548 (2022) doi:10.58079/ne8h.

7.          Merton, R. K. The Sociology of Science: Theoretical and Empirical Investigations. (University of Chicago press, 1973).

8.          Wittenburg, P. Open Science and Data Science. Data Intelligence 3, 95–105 (2021).

9.          NCSES. Publication Output by Region, Country, or Economy and by Scientific Field. https://ncses.nsf.gov/pubs/nsb202333/publication-output-by-region-country-or-economy-and-by-scientific-field#utm_source=chatgpt.com (2023).

10.        Serenko, A. & Bontis, N. A critical evaluation of expert survey‐based journal rankings: The role of personal research interests. Asso for Info Science & Tech 69, 749–752 (2018).

11.        Mancoridis, M., Sumers, T. & Griffiths, T. Publish or Perish: Simulating the Impact of Publication Policies on Science. Proceedings of the Annual Meeting of the Cognitive Science Society 46, (2024).

12.        Breznau, N. Questionable research practices from the practitioners’ perspectives. Crowdid https://crowdid.hypotheses.org/1666 (2025) doi:10.58079/14f3u.

13.        Lawrence, S. Free online availability substantially increases a paper’s impact. Nature 411, 521–521 (2001).

14.        Fleck, C. The Impact Factor Fetishism. European Journal of Sociology / Archives Européennes de Sociologie 54, 327–356 (2013).

15.        Rushforth, A. & De Rijcke, S. Practicing responsible research assessment: Qualitative study of faculty hiring, promotion, and tenure assessments in the United States. Res Eval 33, (2024).

16.        Dougherty, M. R. & Horne, Z. Citation counts and journal impact factors do not capture some indicators of research quality in the behavioural and brain sciences. R Soc Open Sci. 9, 220334 (2022).

17.        Gopalakrishna, G. et al. Prevalence of questionable research practices, research misconduct and their potential explanatory factors: A survey among academic researchers in The Netherlands. PLOS ONE 17, e0263023 (2022).

18.        John, L. K., Loewenstein, G. & Prelec, D. Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling. Psychol Sci 23, 524–532 (2012).

19.        Muñoz, J. & Young, C. We Ran 9 Billion Regressions: Eliminating False Positives through Computational Model Robustness. Sociological Methodology 48, 1–33 (2018).

20.        Borjas, G. J. & Breznau, N. Ideological bias in the production of research findings. Science Advances 12, eadz7173 (2026).

21.        Rainero, V., Stolz, J. & Luijkx, R. The Faith Factor. How Scholars’ Religiosity Biases Research Findings on Secularization. Sociological Science 13, 154–177 (2026).

22.        Aronson, J. K. When I use a word . . . “Publish or perish”: adverse effects. https://doi.org/10.1136/bmj.r1577 (2025) doi:10.1136/bmj.r1577.

23.        Paruzel-Czachura, M., Baran, L. & Spendel, Z. Publish or be ethical? Publishing pressure and scientific misconduct in research. Research Ethics 17, 375–397 (2021).

24.        Carey, J. Scientific Communication Before and After Networked Science. Information & Culture 48, 344–367 (2013).

25.        Larivière, V., Haustein, S. & Mongeon, P. The Oligopoly of Academic Publishers in the Digital Era. PLOS ONE 10, e0127502 (2015).

26.        Rossello, G. & Martinelli, A. The effect of lobbies’ narratives on academics’ perceptions of scientific publishing: A survey experiment. Information Economics and Policy 71, 101148 (2025).

27.        Butler, L.-A., Matthias, L., Simard, M.-A., Mongeon, P. & Haustein, S. The oligopoly’s shift to open access: How the big five academic publishers profit from article processing charges. Quantitative Science Studies 4, 778–799 (2023).

28.        Else, H. Dutch publishing giant cuts off researchers in Germany and Sweden. Nature 559, 454–455 (2018).

29.        Lancet, T. & Board, T. L. I. A. Reed Elsevier and the arms trade. The Lancet 366, 868 (2005).

30.        Jureidini, J. & Clothier, R. Elsevier should divest itself of either its medical publishing or pharmaceutical services division. The Lancet 374, 375 (2009).

31.        Scholastica. New fully-OA publishing toolkit and stakeholder reflections on 20 years of the BOAI. https://blog.scholasticahq.com/post/oa-publishing-toolkit-and-stakeholder-reflections-BOAI20/ (2022).

32.        Normand, S. Is Diamond Open Access the Future of Open Access? The iJournal: Student Journal of the Faculty of Information 3, (2018).

33.        Kumari, M. & A, S. Perceptions of open access publishing: A comparative study of gold and diamond models among global researchers. Alexandria 35, 55–73 (2025).

34.        Plan S. Working collectively towards an equitable, community-driven and academic-led scholarly publishing model. Action Plan for Diamond Open Access https://www.coalition-s.org/action-plan-for-diamond-open-access/ (2022).

35.        UNESCO. Diamond Open Access. Public Service Press Release vol. Online Report (2026).

36.        UNESCO. UNESCO Recommendation on Open Science. https://unesdoc.unesco.org/ark:/48223/pf0000379949.locale=en (2021).

37.        Simard, M.-A., Basson, I., Hare, M., Larivière, V. & Mongeon, P. The Value of a Diamond: Understanding Global Coverage of Diamond Open Access Journals in Web of Science, Scopus, and OpenAlex to Support an Open Future. Proceedings of the Annual Conference of CAIS / Actes du congrès annuel de l’ACSI https://doi.org/10.29173/cais1845 (2023) doi:10.29173/cais1845.

38.        Trueblood, J. S. et al. The misalignment of incentives in academic publishing and implications for journal reform. Proc. Natl. Acad. Sci. U.S.A. 122, e2401231121 (2025).

39.        Why I’m leaving the American Sociological Association. Family Inequality https://familyinequality.wordpress.com/2021/11/06/why-im-leaving-the-american-sociological-association/ (2021).

40.        Philip Cohen’s ASA Publications Committee platform. Family Inequality https://familyinequality.wordpress.com/2018/01/14/philip-cohens-asa-publications-committee-platform/ (2018).

41.        Singh Chawla, D. Open-access row prompts editorial board of Elsevier journal to resign. Nature https://doi.org/10.1038/d41586-019-00135-8 (2019) doi:10.1038/d41586-019-00135-8.

42.        Szkalej, K. The Paradox of Lawful Text and Data Mining? Some Experiences from the Research Sector and Where We (Should) Go from Here. GRUR Int 74, 307–319 (2025).

43.        Breznau, N. & Nguyen, H. H. V. An Introduction to Generative Artificial Intelligence for Academics. F1000 Research Preprint Status, (2025).

44.        Kwon, D. Open-access publishing fees deter researchers in the global south. Nature https://doi.org/10.1038/d41586-022-00342-w (2022) doi:10.1038/d41586-022-00342-w.

Ethical Statement

I have no conflict of interest to report. I used Google search which includes Gemini by default, ChatGPT, NotebookLM and Nano Banana to support my literature review, check arguments, suggest words or phrases and to extract facts. No writing was copied from Gen AI, it is all my own. Any mistakes or opinions are also my own.


[1] https://plos.org/fees/

Media outlet and Q&A for ‘Ideological bias in the production of research findings’ by Borjas and Breznau

  1. Süddeutsche Zeitung (SZ) article by Sebastian Herrmann “Warum Forscher aus denselben Daten entgegengesetzte Schlüsse ziehen” (Why researchers reach different conclusions from the same data).
  2. Manhattan City Journal perspective piece written by George and I.
  3. A news report by Luca Rehse-Knauf for Deutschlandfunk (German radio) and their Forschung Aktuelle series “Migrationspolitik: Einstellungen können Forschungsergebnisse beeinflussen” (Migration policy: Attitudes can influence research results).
  4. A news report by Katrin Kühn and Luca Rehse-Knauf for Deutschlandfunk and their Fakten und Meinungen Series “Darum sind wir Menschen nicht objektiv” (Why we humans are not objective).
  5. A PsyPost report by Eric W. Dolan “158 scientists used the same data, but their politics predicted the results“.
  6. A podcast on The Last Show with David Cooper. Apple / Youtube.
  7. A podcast in Allegedly Does Not Replicate with the Institute for Replication (I4R) with Abel Brodeur and Juan Pablo Posada Aparicio.
  8. A substack post by Claudio Teixiera after an interview with us about how researchers conduct research and the ‘invisible’ paths they follow.
  9. A substack post by Laurenz Gunther describing our work and digging deeper into bias among researchers (here those working on the topic of immigration) – despite a relatively far out conclusion.
  10. There was a Neuer Züricher Zeitung article about the original ‘Hidden Universe’ study (paywalled) ‘Das Experiment: Wer bekommt die rote Karte?.

— How did you arrive at your research question, and what is the study about?

George emailed me and had a few questions about our original study. He was at that time analyzing our data and had found a statistical association between pre-existing preferences for more or less migration among the teams in our study, and their findings. I was very skeptical. I have now worked on replication and reproducibility themes for almost a decade. I am acutely aware of what we often refer to as ‘researcher degrees of freedom’, also known as ‘the garden of forking paths’. This refers to choices that researchers can make during the research process that can lead to different outcomes. I assumed that the statistical association he found would not hold under different but equally plausible model specifications. I began testing many different models. Basically, they all showed the same result. Therefore, I became convinced that this was more than a fluke.

Actually, George had already run most of the same models. We present all of our models in a multiverse analysis in our paper. Out of 883 models 88% showed a significant statistical effect suggesting that we should reject the null hypothesis that ‘ideology has zero effect on the teams’ research findings’. If we take the assumption that we should only trust models that control for researchers’ educational experiences – something we believe impacts their results – then we find that roughly 93% of the models show a significant statistical effect.

— Could you explain the experimental design and methodology?

This study is an exploratory secondary analysis of the data generated by the experiment of myself, Eike Mark Rinke and Alexander Wuttke. We gave 71 research teams the same data and hypothesis – that immigration reduces support for social welfare policies. We surveyed them on their backgrounds and research experience and asked them if they believed the hypothesis was true and what they thought about immigration policy. In the original study, again the one that I helped lead, we did not find any important impact of immigration preferences. We essentially found that the results went in all directions, and we could not easily explain the variation.

After working together, George and I agree that the statistical analysis in the original study was not a clean test of immigration on the research teams’ results, because it controlled for the statistical model specifications that the teams’ made. The original study was were searching for key decisions that might explain why results went in different directions, so this naturally made sense. But in hindsight, these decisions are the mechanism through which ideology gets transmitted into statistical results. Thus, the original study introduced what is known in statistics and causal analysis as a ‘confounder problem’. If someone has an ideological bias, they will choose statistical models that will lead to more desirable results. The original study was controlling for both the test variable (ideology) and its mechanism (the statistical models), and in doing so it suppressed the impact of the test variable we were trying to observe.

This time, George and I conducted regression analyses in which the research findings were the dependent variable and ideology the independent variable – without model specifications as control variables, but with further controls.

— What are the limitations?

A key limitation is that this study is exploratory, not confirmatory. It relies on secondary data from a study that was not specifically designed to test the impact of ideological bias. It cannot confirm that this bias exists, instead it demonstrates robustly that a statistical association exists between ideology and researchers’ findings. We are not aware of any other way to explain this association other than an ideological bias. But we can only confirm with confidence, that it is prudent to reject the null hypothesis that in this particular sample and study is that there is no association between preexisting preferences for immigration policy and research findings. More specifically, the likelihood of observing the data in this experiment if the null hypothesis were true is very low.

Another limitation is that the size of the effect we found is unclear. It points in a positive direction – more pro- immigration policy stances associate with findings that show immigration has a more positive effect on social policy preferences among the public, and vice-versa with more anti-immigration policy stances and a more negative effect. But because of the great variation in results and the small sample size, the standard errors of the estimated statistical effects are very large. This means that the true effect might be anywhere from miniscule and near-zero, to moderate, to very large. We simply cannot say much about this here. More research is necessary, although this is an implicitly difficult topic to study, because if we inform researchers that we are studying their ideological bias, they might behave differently and this would take away ecological validity.

— According to the study, ideology influences model specifications. Could you provide a concrete example to illustrate how a single design decision (or a combination thereof) can have an impact?

I cannot, and this is another limitation of the study. If I could, it would be something we would have found in the original experiment that collected the data. But we can only point at patterns here. There are certain model specifications, unique combinations of statistical modelling choices that produce more negative results. The teams with more anti-immigration ideologies were more likely to choose these. But there are far more model specifications than there are teams. This leads to a sparse data problem. There are many empty cells in the matrix of all possible model specification combinations that teams would plausibly make. This makes it roughly impossible to pinpoint exact specifications’ effects. and there are many different model specifications that can lead to a positive or negative statistical effect. The point is that the only thing that happened between the teams asked to test the hypothesis with the same data, was different modelling choices. Therefore, this is the only way they could arrive at different results. There was no cheating or result faking, we checked that their statistical code produced the results they reported to us.


— To what extent is this a problem, and to what extent is it normal that decisions, based on analytical decisions, depend on who you are and how you think?

This is nothing that our study answers. And it may not be fully possible to answer because we do not yet know the nature of consciousness. We also cannot measure what is happening inside a human neural network – a brain in other words. But it is clear to me that experience, ideology and preferences shape results. A simple example is statistical training. Many researchers have limited statistical training, and they build only those statistical models that they learned about in their studies. This impacts results.

But more generally idiosyncrasies of people, like ideology, shape what research questions that people are willing to pursue and how, and they shape the reporting of those results. Some could look at our study and think that the estimated statistical impact is large and highly concerning. Others, might look at it and think it is tiny and of no concern at all. Our study suggests that ideology can explain somewhere between 1 and 3% of the variance in the results. If scientific findings are on average 1 to 3% off of what they would be without bias, is that a big problem? I mean… what do you think?

— If I understand correctly, the experiment was originally intended to show how much the results diverged, not why. How did you arrive at ideology as a possible cause?

As I already mentioned, this is something that George noticed in our data. He already had this hypothesis in his mind. I cannot blame him for thinking this. The Open Science Movement and Metascience work reveals many so called ‘Questionable Research Practices’. These include everything from faking data, to tampering with statistical models or stopping the collection of data during an experiment to produce a desired result. These practices are designed to produce certain results in order to obtain a publication or support a pet hypothesis (confirmation bias). Obviously some of these studies were motivated by ideological goals.

— How could ideological bias be reduced? Is this even desirable, or should we simply be aware that it can exist?

The impact of ideology can be reduced by following some clear recommendations of the Open Science Movement. Studies should be pre-preregistered, they should provide all code and materials, they should not be conducted in isolation or in hiding, and researchers should cooperate. Some of us are engaging in so called ‚Adversarial Collaborations‘, where researchers who do not agree – those with different priors about a given hypothesis like the impact of immigration – collaborate. They lay out all the aspects of a study in advance, and they agree on what evidence would count as support of either of the positions. I highly recommend this form of science. It takes competition and turns it into collaboration with the goal of knowledge seeking prioritized above all else.


— You are investigating ideological bias in science using scientific methods. How do you deal with this tension in meta-scientific questions, where you are essentially also your own subject of investigation?

Similar to my last answer, one cannot fully understand or deal with one’s own bias, and therefore needs to build in checks into the process. Things that would reduce this bias, like preregistration. I am working currently on a project that is an Autoethnography of my own questionable research practices and the perverse incentives I encountered during my career in science. I hope that by doing this, and revealing my own behaviors, I can improve them. I also want to be a role model for others, to make it desirable to be highly critical of one’s own work. My goal is to get this study published in a high quality journal and thereby prove that self-criticism, something researchers mostly try to avoid to protect their theories, findings and careers, is something that can be used in a positive way in the scientific process.

Can one’s own attitudes always play a role, even in this study?

Sure. Definitely. That is why it was important for me to take a so called ‘multiverse’ approach to this study, and many other studies I am working on recently. I want to ensure that I, or one of my colleagues, has not simply selected a statistical model that produces certain results. This practice, known as hacking, or p-hacking, is prevalent in science, especially in secondary data analysis. I essentially learned to do this during my graduate studies. We would find a result we liked and then develop convincing logical arguments why the model producing it must be the best model. So, the idea with multiverse analysis, is to run all or at least all plausible alternative models. This helps reveal if my model is an outlier. Whether it represents something very unique or unusual in the distribution of model results. If it does, this is a cause for great concern. If not, it is evidence of a robust statistical association.

— There have also been critical reactions, for example in this online Bluesky thread, which raise concerns about George Borjas’s views and background. How do you assess this criticism?

I have read the discussion thread by Michael Clemens, an economist at George Mason University. One line of criticism appears to concern the fact that George Borjas recently conducted research for the executive branch of government.

Another criticism seems to focus on the observation that different model specifications yield different results. This is precisely what our study confirms, and it is also a well-established fact in the history of empirical social science.

Both points are orthogonal to our study. Even if we were to assume, hypothetically, that George held some form of ideological bias and that this bias influenced his analytical choices, which I cannot confirm and do not claim, this would not undermine our findings. The reason is that we adopted a deliberately robust research design. We conducted a multiverse analysis comprising 883 regression models. The consistency of results across this large set of plausible specifications makes it implausible that our conclusions are driven by special highly selective model specifications – those that would be selected due to ideological bias.

It is also important to note that George and I do not share the same political views. Precisely for that reason, ideology was an additional motivation for us to adopt a highly robust analytical strategy. I explicitly advocated for the multiverse approach in order to minimize the influence of individual priors, including our own potential ideological orientations. I see this as a great strength of the study.

The purpose of our approach is to decouple empirical results from personal or ideological preferences as much as possible. And to estimate the robustness of our finding to any kind of bias, not just ideological. I would encourage critics to apply the same standards of robustness to their own work. To date, Mr. Clemens has not presented empirical evidence that contradicts our findings. Moreover, criticisms referring to modeling choices in studies conducted by George decades ago are not relevant to the validity of the present analysis.

— What does it mean when you say that only 3% of the variance can be explained by your results.

That has to do with the regression coefficient and the r-squared values. A coefficient of 0.03 shows that a one-point higher (more positive) ideology mathematically predicts a change in results of 0.03. This sounds meaningless, but we know that 0 would be none (and we can equate this with zero percent change) and that 1 would be 1-point on a standardized scale. This is 1 standard deviation in the distribution of the dependent variable. It would be possible to move the results more than 1 standard deviation, but this would be quite preposterous. There is nothing in the complex nature of social science, which lacks laws, that would do that. So I will set the upper bound of the largest possible effect at 1, meaning that 1 would equal 100% of the distance in the distribution of variance.  Therefore, 0.03 is like 3 percent of the distance. At the same time, the r-squared, which tells us how much of the error is reduced from this particular variable, is around 0.03 or less depending on how we measure this variable. This suggests that fitting the observed ideology values into the observed results from the teams, reduces the unexplained variance from 100% down to around 97%.

— Can you explain — very, very simply — what you did? 

In a study that I co-led starting in 2018 (Breznau, Rinke and Wuttke et al. 2022), we designed an experiment that allowed us to observe researchers doing research on the impact of immigration on social policy preferences. We gave them the same data and asked them to answer the same research question: whether immigration reduces support for social policies or not. We documented the researchers’ pre-existing methodological training, experience with and expectations about the topic, and their personal preferences for looser or tighter immigration laws in their own countries. We shared all of the data and documentation of our work publicly. Because we shared our data, a few years later George J. Borjas was able to reanalyze our data and find new evidence of a correlation between the researchers’ ideological positions on immigration and their findings. Those with more pro-immigration positions tended to find evidence that immigration had a positive impact on support for social policy, and those with more anti-immigration positions tended to find evidence that immigration had a negative impact on support for social policy. Here “social policy” means support for a more extensive welfare state providing social security via the government or not – many would call this support for social cohesion. I was skeptical of George’s initial work, and together we vetted George’s findings. We ran almost a thousand alternative statistical models to test George’s finding and of these 88% suggest that we should reject the null hypothesis. The null hypothesis is that if there is no impact of ideology on researchers’ findings, that we should not observe what we observed based on probability. But we did observe this association, and by rejecting the null hypothesis we have evidence that something more is going on, not just random luck in the data. Crucially, we should reject the idea that ideology has no impact on researchers’ results when analyzing the same data. 

— What motivated you to do this study?

As I said, George found this association between researchers’ ideological positions on immigration and their research findings. This is an important scientific observation. I was a bit more skeptical, in particular because my initial analysis of the data as part of the original experiment did not show this association. Together we became very motivated to tackle this problem, and in the end, we have relatively strong evidence of something. The exact nature of this should be subject to further research.

— What are the most striking findings? 

The main finding is that researchers’ own preferences for tighter or looser immigration predicts what they went on to find in their work. This appears striking, but when put into context it is maybe not so surprising. We know that there are many reasons that researchers may consciously or even unconsciously exert influence on their own findings. There are many cases of researchers engaging in questionable research practices to ‘fudge’ their data and results in ways that make them appear stronger than they really are. This has occurred frequently in biomedicine, for example in studies of products whose approval would net the researchers great personal profits. Consider also what we know from psychology, namely confirmation bias. People tend to seek out evidence of what they already believe is true. Researchers are people too. When presented with various forms of competing evidence, a researcher might gravitate toward that which supports their preexisting beliefs or preferences. 

— What are the implications for the social sciences, which are already reeling from the replication crisis? 

The implications reinforce what we already know. We should focus on refining two areas of science. The first is scientific training. We need to make transparent and open workflows and data sharing the norm. This includes researchers stating in advance what they plan to do and what they expect to find. When researchers do not do this, they can run several experiments or analyze hundreds of datasets with millions of different statistical models and simply choose one finding that looks exciting or sexy to them. Generations of researchers before us have learned to do exactly this in order to get published. They learned to ‘sell’ a single selected finding as confirmatory evidence of something in the real world, when in fact it is simply a highly selected, exploration of data leading to a unique event; one that probably does not generalize and is not reproducible (a.k.a. luck). I bring up the idea of getting published here because this is the second problem we must urgently address in science. A researcher’s worth or ranking as a scientist is judged almost entirely on their publication record. In particular, publications in journals that are considered higher status. These higher status journals should be publishing studies because the studies contain higher quality science. But there are ways to game this system. There are a a host of questionable research practices that makes findings look more exciting than they actually are. These practices often lead to irreproducible findings. Essentially fake science. Big publishing is a major profit industry and this increases the pressure to publish, as the publishers of the journals engage in questionable, sometimes unethical practices to sell more journals, as opposed to solid science which is often quite boring and tends to find that new drugs or treatments do not work and that our theories are wrong. Ideally we need to end big publishing’s control over science. Their role should be simply production and distribution, but they currently copyright much of the material and force universities to pay twice, once for the researchers’ salaries and again so that researchers can read what they are publishing. And this often done with public money. Its really wrong and generates perverse incentives among researchers and profit-seeking publishers. 

— It seems teams didn’t falsify data or cherry-pick numbers in any obvious way; instead, ideology appeared to influence judgement calls — is that a reasonable explanation or is it too charitable?

It is a reasonable explanation, but we must be very cautious with it. Ideology might explain about 3% of the variance in researchers’ findings. The rest has to do with other factors or random noise. Consider that many scientists have specific methodological training. Through this training they simply do what they know how to do. And this can influence results. For example, someone who only knows how to use a hammer, will treat things as a hammering problem, when it might be better solved with a different tool. This takes us back to better, broader methodological training, which scientists would have more time for if they weren’t under constant pressure to publish. 

— Are there ways to guard against the ideological bias highlighted in the study? It seems that peer review can spot poorly defined studies — but does it work well enough in the real world? 

Peer review is a poor solution. There are studies out there that show that peer review is not reliable. Just like giving the same researchers the same data, if you give different peer reviewers the same paper, they will come to a huge range of judgements about the paper. The publishing system is in some ways a lottery. One solution for this problem is what we call “adversarial collaboration”. This is a type of research where scientists who disagree about a topic work together. They design a study together and agree on all methods and on all criteria with which to judge the outcomes. Then after this is all agreed in advance, the study is conducted, ideally from a third-party, and then the results speak for themselves.

— What do you think the message here is to the public and to policy makers?

Everything we have in society that works is based on science. Smartphones, that open heart surgery that saved the life of a loved one, airbags, planes that don’t crash. We need science for every decision we make collectively. But the public and policymakers can be highly politicized, and this can influence science, we see this even in the scientists in our study whose own politics seemingly played a role. To cut through this, we need policymakers that support science conducted by scholars who have different political ideologies, who do not agree. For example, George and I are somewhat different in our own assessment of the impact of immigration on society and the labor market. This made us a very strong team. It meant that we could focus on the scientific process and try to get to the best, most reliable answer. Crucially, we need to never rely on single studies or single science teams. Before we declare evidence of anything, we need many studies. We should not just rely on single papers or scholars. We need dozens or hundreds of studies on a topic conducted by inter-disciplinary and inter-ideological teams.

Open science. Back to basics

Are we really sharing all steps in our research? And are we making clear which steps we cannot share and why? There are always hidden steps that influence our research design and reporting.

Ask yourself this question: If someone else were to pick up one of your studies, most likely a published paper reporting the results of one of your studies,  would they be able to understand and reproduce everything that you did?

Are there things that can’t be exactly reproduced, like privacy-protected data, or subjective, interpretive, or ethnographic experiences in your study? If so, have you pointed this out for the reader?

These sound like trivial questions.

Anyone outside of science would think these things must be true. They would think that of course a scientist, a researcher, would report everything they did in their research. Isn’t that their job? How else should science look? What else are scientists doing other than doing research and reporting that research?!

Maybe your answer to this question is, ‘yes, I have reported everything I did’. You truly believe that people could follow everything you’ve done, but is this really the case?

I work in the area of statistical analysis, mostly, and in this area, there is code required to run statistical analyses, and that code is the basis of my research design, assuming I didn’t collect the data myself.

So, I have secondary data, I already have data, and then I analyze it, and my code is the basis for that analysis. Well, it turns out there’s a lot of things in my code that others would not be aware of without the code. So, this makes code sharing absolutely essential. And we’ve come a long way in this area.

Economics and political science especially, in general, have strong code-sharing norms for people who do statistical analysis. Sociology is a little further behind. Probably one of the main reasons for this is that journals have not adopted strong policies of code sharing in sociology, unlike in political science and economics. Either way, as a researcher, I should be responsible for making my results reproducible, but let’s think of an even more hidden case.

In your statistical analysis, did you do things that are not in the code? Now, if any researcher were to answer ‘no’ to this question, I would not believe them, including myself. There are always things that we do, that we then take out of the code because they seem redundant or unnecessary or, in some cases, don’t make our study look as good as we want it to look. One of those things is to run models that we don’t report.

Now, that would be fine if we just happened to run a model on accident or had some other glitch or mistake. But a lot of the models we run are intentional, and the reason we don’t report them is that they produce results that we don’t want. They produce results that don’t support what we think is true or what we want to appear to be true because it’s sexy, because it’s something that would be publishable.

But the question to ask is then: have any of those extra models that I’ve run or you’ve run been used to inform decisions that we make in terms of what models to run next, how to recode any variables in those models, and what to report on, in the paper and in the code? And if the answer is yes, then that’s part of the research design, and should be reported.

Now, this is not the norm at all. I don’t know any discipline where this is the norm, and I don’t know people who are trying to make this the norm. This is a hidden area of the open science movement for the most part, although we have research now that overwhelmingly suggests that the findings that we’re reading about, that people are reporting on, are probably selected (thanks to p- and z-curve analysis). They’re probably a subset of findings that are not just a random subset of the models that researchers intentionally ran, but a selected subset in the sense that they all point in a certain direction or they all have a larger size or there’s something about them that makes them specially selective – and this selection process reduces the reliability of science.

Therefore, when I say back to basics, I actually mean the basics of research, not the basics of the open science movement. I mean reporting everything in the research process that influences the results. That would truly be open science.

Image credit: Nate Breznau’s own photo

Questionable research practices from the practitioners’ perspectives

With Monica Gonzales-Marquez, Priya Silverstein and Eike Mark Rinke.

In preparation for Metascience 2025, I put together a panel with three other researchers called “Questionable Research Practices from the Perspective of the Researcher: Understanding Perverse Incentives using Autoethnography”. We used our own discussions in online meetings and individually as recorded narratives as content for this panel. The impetus was that much metascience points accusatory fingers at problems in science, thus supporting a culture of fear. Researchers comply to avoid scrutiny, and not necessarily out of an intrinsic motivation to do good science.

When successful and meaningful science is measured entirely by publication in a ‘high impact’ journal and high citation counts, the drive to do science is transmuted into a laser-like focus on publishing. The intrinsic motivation to contribute to the creation of scientific knowledge becomes confounded with publishing, while simultaneously deprioritising the robust, pedantic, methodical, humble labor involved in doing good research. The scientific method gets confounded with the mechanics of publishing and achieving “standing” in the scientific community.

We hope that by revealing our own experiences in the world of ‘publish-or-perish’, and how it has pushed us towards questionable research practices (QRPs), we might generate intrinsic motivation in others to dispassionately examine their own scientific practices.. By looking back through our histories in academic work and sharing them, we also expect to increase our own intrinsic motivations to be ever vigilant, and to learn to always privilege scientific integrity over publishing.

Confounding: Publication = Science

A major theme in our narratives is that science has been fully confounded with publishing. For some of us, the pressure to publish, and the toxic atmosphere and interactions pushed us out of academia entirely. For others, we internalized the publishing norm, convincing ourselves that we were doing impactful science, when we were actually just doing impactful publishing – which does almost nothing to alter collective human knowledge, promote social justice or solve other societal problems.

Some excerpts from our narratives:

Questionable Behaviors: Hacking

The pragmatics of doing science in a publish or perish culture is that we either engage in some (often undeliberate) hacking and storytelling, or are sanctioned.

Some excerpts:

Ego, Power and Personal Struggles

As human beings we are prone to seeking status, material security, community acknowledgement and different things depending on where we are at in life, and who we are personality-wise. Especially when in graduate school, we are in a position of little power in comparison to professors and our supervisors. People who  wield power over others, and push their own ego-centric agendas tend to be quite successful in science. This can lead to abuses of power, ego trips and other toxic behaviors that can diminish junior researchers’ ability to push back against pressures to engage in QRPs and have strongly demotivating effects.

Intrinsic Motivation

We hope that by speaking out, and normalizing scrutiny of our own experiences of questionable research practices as something valuable, we can help others become intrinsically motivated to do the same. Moreover, we propose that ethnographic narrative may be an underexplored but powerful method to help uncover the causes and motivations of questionable research practices from researchers’ lived experience of “doing science”.

Part of the motivation for this panel is based on one of the authors’ (Nate Breznau’s) own authoethnographic research into my QRPs. He presented preliminary results at the Sociological Science Conference at Cornell (link to slides).

Readers can find the full poster for our presentation at Metascience 2025 at University College London here.

Our future plans are to seek a larger sample of researchers and invite them to a semi-structured narrative sharing process (via recording themselves) to further study QRPs. In particular we hope to target early career researchers to shed light on the current state of science training and supervision experiences. 

Science in survival mode

Scientific research is unreliable. It comes with uncertainty. Whether launching a rocket or measuring racial prejudice, there is uncertainty. We use this uncertainty to make decisions. If the rocket has a 40% chance of exploding, best not to stick astronauts in it. If skin-tone bias of soccer referees is somewhere from none to a lot, it is irresponsible to conclude they are prejudiced.

Investigating uncertainty, is science. Truth and uncertainty are two sides of the same coin, they co-define each other. A problem for humans measuring uncertainty, is that humans are unreliable. Human scientists themselves add uncertainty to the measurement of uncertainty.

Recent studies suggest that somewhere between 25 and 60% of published statistical results cannot be recreated using the materials provided – these measures of uncertainty come with their own uncertainty. It turns out that scientific researchers are doing some really peculiar things to generate uncertainty. Some surveys suggest as many as 9% of scientists faked data at least once in their careers, and that more than half selectively reported findings – a behavior that makes the things they study to appear less uncertain than they actually are.

I believe the answer lies in their humanness. Like all animals, they are genetically programmed for survival. They are capable of both rational, reflective decision-making and split second reaction without any thought. Given time to reflect, a human would generally conclude murder is unethical, but simultaneously would not hesitate to kill if it prevented their child from being killed. Murder remains wrong, but to not kill and let a child die is also wrong. It would be irresponsible parenting failing to ensure survival.

Murder is a profound act when a human perceives themselves to be in a situation of life or death. It is a symptom of subconsciously activated defense mechanisms in survival mode. But there are many other symptoms. In order to avoid death, humans, like other primates can engage in deception, disassociation, aggression, manipulation, submission, scapegoating, theft and hoarding. If someone held me at gun point and told me to prove the earth was flat, I would have no problem doing it. The math would work, I would just need to fake a little data

Most scientists have highly valued knowledge and competencies and thus live in situations where their lives are not under threat. At least not as a result of their scientific practice. For the sake of this thought experiment, lets just rule out regularly occurring life-threatening danger as a cause of scientists exhibiting survival mode behaviors. This leaves the perception that they are under threat, as a possible explanation.

It takes only a few stimuli to induce survival mode behavioral changes in animals. Like hearing a frightful noise when seeing an animal. This animal and anything that looks like it become automatic sources of anxiety and fear even without the sound. Imagine being told over and over and over that you have to have an exciting study with powerful results in order to get an academic job after graduate school, otherwise you wasted 3-8 years of your life and probably a large chunk of capital on getting a PhD. Could this alone, without introducing any actual shocks or physical pain, induce Pavlovian fear? I encourage you to go ask any graduate student to answer this that does not yet have such a study.

Now imagine that during graduate school a student invests all their time and resources into an experiment. After it is complete, they hypothesized result, the one sure to be exciting and publishable, is not there. Imagine the horror, the shame, the feeling of failure, the panic. Remember the poor soul who leaped to his death because he misunderstood futures trading and thought he owed three-quarters of a million dollars? That was a triggered survival response. Because dying felt like the only way to ‘survive’ the horror of facing that debt. Imagine if he could have just changed the futures market by adding his own numbers to the market. Would he have done it?

We should not be surprised at all then, when scientists acting out of fear-based survival strategies, fake data. Diederik Stapel faked an entire career of data before being caught. His behavior was self-described as an “addiction”. A common reaction of individuals placed under fear stimuli that are emotionally damaging if not traumatic. The intense pressures and expectations of the academic environment created a context in which he felt compelled to engage in unethical practices to maintain his status and success. It does not make murder or data-faking right, but to not take this as grounds for indicting the scientific rewards system is certainly wrong.

The incentive structures have to change before we can honestly expect the fear-driven pressure to fake, cheat, lie or steal – in order to avoid the experience of loss associated with the common null results that occur when conducting high quality scientific research on radically complex human brains and societies that are frustratingly difficult to measure things in – to go away.

Open science in sociology. What, why and now.

WHAT

By now you’ve heard the term “open science”. Although it has no global definition, its advocates tend toward certain agreements. Most definitions focus on the practical aspects of accessibility.

“…the practice of science in such a way that others can collaborate and contribute, where research data, lab notes and other research processes are freely available, under terms that enable reuse, redistribution and reproduction of the research and its underlying data and methods.”


FORSTER, open science teaching resource

Some definitions enter the realm of ethics, feminism and social justice.

“…to imagine and design inclusive infrastructures, practices, and workflows for scientific practice that intentionally enable meaningful participation and redress (these new) forms of exclusion.


Denisse Albornoz,OCSDNet

Others focus on the communicative interplay between scientists and the public.

“Openness in Open Science also means opening up science to society… The democratic ideal of Open Science argues for equal two-way communication with the public: one should not solely focus on the question of how to foster the uptake of science in society, but also on how to foster the uptake of societal insights in science.


Anne-Floor Scholvinck,ZBW Mediatalk

Whatever the ontology, open science is inevitably something that challenges the status quo in science. Usage of term indicates there is something undesirable about science, otherwise advocates would simply advocate “science”.

The “open” part of the concept refers to any number of things depending on whom you ask. Commonly it means:

Open access – making the results of scientific techniques, research and theory accessible to everyone; as opposed to only in paywalled journals.

Transparency <open process> – making all methods, code, data and any biases or conflicts of interest known before and after the research is conducted. So long as doing this does not harm human subjects or violate any laws.

Open source – on the technology side of science, all programs, apps, algorithms, tools and scripts should be transparent and usable by others. This means that when a scientist develops a new technology, anyone else’s technologies can interact and interface with it. Moreover, anyone can modify the technology to better suit their own needs.

Open academia <open communication/democracy/feminism> – allowing anyone to participate in academia. That academia has the goal of eliminating inequalities, prejudice and domination from academia that take place in the social world. That academia embraces feminism and critical race theory in its methods and institutional practices. That everyone has the same place in scientific discussions, and no science is conducted by pressuring others or taking advantage of existing power structures. That no science takes place in secret, except for research that requires obfuscation for its completion.

Again, the definitions can cover a broad range. The above are just a snippet, although they strike me as the most common usages; except for ‘open academia’, this is reserved for certain justice motivated scholars.

WHY

Although I do not proclaim to be the arbiter or knower of right or wrong in academia (and life in general), the following facts seem wrong to me.

Double-work and the co-opting of journals

Scientists provide their work as editors and reviewers, because the peer review and publication process is the centerpiece of all of science. Peer reviewers and editors are the only consistent form of quality control in science. The academic journal was a functional response to previous forms of knowledge transmission that required direct scientist/practitioner to student interactions which were geographically limited and reached a very narrow audience.

The journal made it possible to transmit knowledge across the globe. Moreover, the journal reduced the simultaneous discovery and re-discovery problems of science, because no one could prove they discovered something first, and others worked on problems that were already solved unknowingly. It represents one of the first ‘open science’ movements because it was driven by the idea that science was at an impasse and could only move forward through transparent and open exchange of ideas arbitrated by being part of the public record through publishing.

Ironically, the journal format came full circle and began to undermine science. After over two centuries of journals run by non-profit academic associations, for-profit publishing houses began ‘offering’ their services to meet the growing global demand for journals and their content and the rising costs of editing and distribution. In many cases, these publishing houses were able to purchase the journals by offering the academic societies the exclusive right to determine what went in them. Within just 30 years, five conglomerates owned the titles, content or certain features of over 50% of all journal articles published globally.

The content, as always, is still a product of the scientists and the voluntary work of editors and peer reviewers. The publishing houses make large profits, but pay nothing to these workers. The editors and peer reviewers earn their income from universities mostly. The very universities that pay high fees to purchase the right to provide the journals in their libraries. This is a double tax on the universities – paying the producers of content to produce and then paying the distributors of that content to consume it. The content does not change at any point in between these two forms of payment, in other words, the publishers do not add any scientific value to this content.

Matters got even worse with the publishing houses over the past decades. As creative and deceitful profit seekers, some publishing houses realized they could generate even more profit by collaborating with the private sector. For example pharmaceutical companies’ profits were directly determined by the findings of studies published in journals. Pharmaceutical companies, or any companies whose profits were determined by the outcomes of scientific experiments, would be willing to invest in shaping those outcomes if they could. Enter a novel concept pioneered by Elsevier: selling journals or journal space to private companies to boost their profits. Win-win for them. Elsevier also pioneered the process of monetizing open science by purchasing SSRN, engaging in massive lawsuits designed to stop the free sharing of (their) copyrighted knowledge and tries to copyright intellectual activities such as peer review.

Other ventures create journals that prey on scholars who do not know better, or seek to get easy publications to add to their CV. These publishers are often labeled “predatory publishers” and they “publish work without proper peer review and which charge scholars sometimes huge fees to submit should not be allowed to share space with legitimate journals and publishers, whether open access or not” (predatoryjournals.com). They also sometimes mimic reputable journals by copying their styles and their names and soliciting content from scholars, a procedure known as “hijacking“.

Publish-or-perish begets questionable research practices

Thanks to the advent of the scientific journal, knowledge could be evaluated, used and further transmitted across space and time. The utility of the journal and other forms of academic publication such as books, proved so effective that they became the primary source for others to evaluate the importance of scientists and their work. This gave rise to the norm we are all familiar with, publish-or-perish.

In a survey of psychologists, John et al. (2012) found that 50% claimed they had selectively reported studies that supported their hypothesis (as in, selectively excluding those that didn’t). Moreover, 35% admitted to reporting unexpected findings as having been predicted from the start. Nearly 2% outright admitted to faking data.

Publish-or-perish and questionable research practices have a causal relationship. Except for occasional sociopathic or psychotic individuals, there is no reason for a scientist to engage in questionable research practices. No reason, except scientists’ very existence on scientists may depend on it. So many studies in reality lead to results that go in all directions, support the null or (most importantly) do not provide groundbreaking new results.

Through the peer review and editorial process, journals select studies that are path-breaking. Studies that will move knowledge forward and be of the greatest interest to readers. When faced with prospects of not getting tenured, not getting grant funding and being forced out of academia, a human’s (scientist’s) rational calculations change. Suddenly, rounding that p-value from 0.054 to < 0.05 or even adding some cases to the data becomes a cognitively defensible decision.

Like any profession, science is competitive. Those who publish more, or get more citations to their publications tend to get ahead. Those who don’t, don’t. Professional athletes use incredible tactics to gain competitive advantage. Of course steroids are well-known, but other tactics are much harder to detect. For example, endurance athletes often use blood transfusions to boost recovery and performance. This is what it means to be human, scientist or not.

One of the most radical events in the social and behavioral sciences is Diederik Stapel’s entire career faking data and results that were published in at least 54 articles that consumed millions of Euro in funding. It took almost two decades for critics and whistleblowers to finally out him. Psychology is not alone. In political science LaCour and Green published a study in Science that attitudes toward gay marriage could be changed if heterosexual people listened to a homosexual person’s story, but it turns out LaCour fabricated results of a follow up survey that never took place as uncovered by Broockman. In economics Reinhart and Rogoff published numerous studies identifying a negative impact of high debt rates on national economic growth, when in fact several points in their dataset had conspicuously missing values. When these values were added there was no longer support for their claim as identified by Herndon, Ash and Pollin.

I suspect that most questionable research practices are not intentional. The sociopathic (~psychotic) Stapel’s of the world are rare. This pressure to find a job after doing doctoral studies and then to get tenured, means a trade off between conducting science in its ideal form – so learning as much as possible about the existing literature on a subject, mastering the necessary methods to perform the research and executing the research, possibly with several iterations, and facing the prospect of null results – with science in a form that will lead to publication as fast as possible.

This ‘fast as possible’ leads to amateur science. For example, in the rush to get my first publication I attempted to use “multiple imputation”, but lacked the time to properly learn this method. Instead I simply generated several datasets and averaged them into one and re-ran the analysis on this one. This was not an intentional misuse of a method. It is a questionable research practice as a result of context. Think about matrix algebra. It is the basis of many advanced statistical techniques regularly used by social scientists. How many of us have a strong grasp of matrix mathematics? I don’t. And yet I’ve published several studies using structural equation modeling.

WHAT & WHY in SOCIOLOGY

I am aware of nothing about sociology that suggests it needs a special adaptation of open science. Most research cannot be strictly delineated as sociology or not sociology anyways. The boundaries of a discipline, especially within the social sciences, exist mostly in the institutional structure of universities. Eliason suggested that sociology is unique because it overemphasizes quantitative techniques, has needlessly long articles, lacks writing for the popular press and emphasizes research at the expense of teaching. In my experience the previous sentence perfectly describes all social and behavioral science disciplines at once. Even article length, something I thought might be peculiar to sociology, is not special. Political science and management research have very long articles. Consider that and ASR and ESR for example, limit words to 9,000 and 8,000 or less – this is relatively average if not short for social science.

Actually, I would argue the most unique thing about sociology at the moment relates to open science. Two points in particular: (A) that sociology has not had the same incredible scandals as other disciplines and (B) that sociology lags behind other social sciences in promoting open science.

A lack of scandals, not scandalousness

Could sociologists be more scientific and ethical in their research behaviors than those in other disciplines? Given identical institutional and career structures that favor productivity and innovation over replicating or checking each other’s work, I doubt it. Sociology journals and their editors, for example, rarely retract articles despite evidence of serious methodological mistakes. Carina Mood once accurately pointed out mistakes in the interpretation of odds-ratios in some American Sociological Review articles, but the editors refused to publish her comments, much less consider retractions. She shared her exchange with ASR in an email to me and discusses some of it in a working paper. An exceptional recent event was the retraction of one of Legewie’s sociological studies, but this required he himself to initiate the retraction after someone pointed out errors in his work. Until 2020, the Retraction Watch database (www.retractiondatabase.org) listed no retractions from the top sociology journals, and only two among the well-known, one in Sociology and another in Social Indicators Research.

This year, something new happened. Five articles published in Social Problems, Criminology, and Law & Society Review were retracted. These articles had the common co-author Eric Stewart. It turns out that the data he provided were faked. There is no other logical conclusion that this after exceptionally rigorous work by Pickett (a co-author of Stewart) provided evidence that the Stewart studies had consistently incorrect means and standard deviations, unverifiable surveys (sources, methods, original materials), magically changing case numbers despite identical statistical results, sometimes half the data had duplicate cases and impossible clustering structures in the data.

As an aside, one of Pickett’s findings was that the data had non-uniform terminal digit distributions. This means that the right-most digits in the reported statistics differs markedly from a uniform distribution. In particular, at the third-digit numbers should be uniformly distributed with 0-9 appearing roughly 10% of the time. In one of the papers, zeros appear less than 2% of the time. If you are considering faking data, keep in mind that it is roughly impossible to do it in a way that cannot be detected by careful investigation. Any algorithm used to generate results (even copying and pasting) leaves is statistical marks.

Perhaps we sociologists should be partly relieved, as this is just confirmation that we are as much a part of social science and its problems, as any other discipline. However, the Stewart retractions which should have been breaking news for sociology, went mostly unnoticed. The results of the investigation leading to the retractions is not published in a flagship sociology journal where it belongs. Instead it appears in Econ Journal Watch – something unlikely to be read by any sociologist. Moreover, the retraction notices from the original journals do not cite outright fraud. Stewart continues to promote his work in print claiming the main findings still hold, and several other of his studies with similar irregularities have not been retracted.

Another, extremely important event was a case of ethnomethodological research conducted by Lindsay, Boghossian, and Pluckrose in the mid 2010s. This is sociological self-examination at its best, although their backgrounds are mostly outside of the discipline of sociology. They wrote a series of 20 papers presenting fake results and making arguably unethical claims. They invented the papers to mimic the style of articles published in journals well-known for sociological research on topics of identity, hegemony and marginalization. Seven of their papers were published or had revise and resubmit recommendations before whistleblowing forced them to cancel the project. Some highlights: one paper contained sections from Hitler’s Mein Kampf. Another suggested men should be trained similar to dogs to prevent rape, and a third that white men should be forced to sit in chains on the floors of university classrooms, instead of normal desks. I am not commenting on the merit contained in these ideas, only that they all contained faked data, non-existent methods or conclusions not supported by the data. That these studies easily flew under the radar of a number of high impact journals points out how easy it is to publish without doing the necessary research work.

Lagging behind closed doors

October 6th, 2020. I entered the search terms “open science” (with quotations to search the exact phrase) and “sociology” (with quotations to only return results that contain the word) into Google Scholar. Six pages of results without a single sociology journal. On page 7, Merton’s “Priorities in scientific discovery: a chapter in the sociology of science” appears. Publication date 1957.

In 1973, Wilson, Smoke and Martin found that 80% of studies published in the top three sociology journals of that time rejected the null hypothesis, in other words they had p-values below a threshold. This suggests publication bias, if not p-hacking. Sahner (Table 5) analyzed all article submissions to the Zeitschrift für Soziologie, 1972-1980. Of those that contained significance tests, 70% were significant at p < 0.05 suggesting that authors prefer to submit significant results. More recently, Gerber and Malhotra (2008) reviewed articles published in American Journal of Sociology, American Sociological Review and The Sociological Quarterly, and specifically looked at the boundary of t = 1,96 (i.e., p<0.05) to find that as many as 4-out-of-5 studies were ‘significant’. This suggests publication bias as well. Sociology has yet to have a systematic review of p-hacking by comparing p-values within ‘significant’ results. Meanwhile psychology and political science for example are teeming with papers on “p-hacking” and “publication bias”.

Sociology is rather intransparent. An estimated 78% of the major sociology journals have long-standing transparency policies. Unfortunately, these policies are mostly artifacts on paper without much enforcement. For example, only 37% of sociology articles published in the mainstream journals between 2012-2014 include shared data and/or materials. In 2015, a small group of sociologists tried to obtain materials from the authors of 53 prominent sociological studies. They obtained these from just 19%, and only 20% of all the authors they contacted bothered to respond despite several requests. This suggests sociologists are free to hide the data and materials that led to their findings without recourse, despite such guidelines.

Other disciplines have embraced the Transparency and Openness Promotion Guidelines (TOP). The TOP guidelines with help of the Center for Open Science support journals to improve science. Journals can become signatories of TOP, and in doing so they either adopt and enforce new transparency guidelines, or certify that they already meet certain transparency standards. Most of the top psychology journals and several political science journals signed on. Other major journals such as the Journal of Applied Econometrics and later the American Economic Review adopted their own enforced transparency guidelines.

Until 2017, the only higher ranking sociology journals that signed TOP were Sociological Methods and Research and American Journal of Cultural Sociology. In 2017, Elsevier dictated that all its journals adopt guidelines and this added Social Science Research to the list. At the time of writing this, the flagship journals American Journal of Sociology and American Sociological Review neither signed TOP nor enforce their own guidelines. Of top German sociology journals, the Kölner Zeitschrift für Soziologie und Sozialpsychologie is the only signatory.

If intransparency is pervasive in sociology, then research cannot be (a) checked for errors, (b) reproduced or (c) simply critiqued. Even when exact reproducibility is not the goal, as often is the case with context-specific interpretive research, most research methods remain shrouded in mystery. This requires readers to take a giant leap to trust what others report. Part of the problem is that sociologists express little interest in reproduction or checking others’ works. There are few replications in the history of sociology, and if anything, they decreased over time until recently. For example, searching the articles in American Journal of Sociology and American Sociological Review reveals 22 replication studies from 1950-1980 and only 8 from 1981-2010.

Something telling about a lack of willingness to open sociology comes from sociology’s most ‘powerful’ society, the American Sociological Association. They collectively petitioned the US government to not make data transparency a requirement attached to grant funding in 2019.

NOW

What to do about it? Here are some simple steps to consider especially for sociologists. Similar to steps advocated by many others for graduate students and academic institutions, or all of us for example.

Transparency

Make all the materials – research design, methodological steps, data (when legally and ethically possible), analyses, conflict of interest and any software code – available online. The practical reason is that others can follow your work and expand it in the future. Doubly practical is that you don’t need to respond to email requests for your materials. So long as you are not a deceitful sociopath, you want others interested in your work and to replicate your work. Even if a study, seems to ‘prove you wrong’, the fact that it replicated your work is evidence of how important your work is and the topic of study. You are a piece of a much larger community of knowledge construction. Constructive exchange can lead to collaboration with critics to generate better future research without personal conflicts.

The immediate value of transparency is that being transparent forces you to be careful. Knowing everything will be public information increases the value of attention to detail. Put in its converse: not sharing your workflow publicly can indirectly foster lower quality standards, in addition to creating possibilities for misconduct. All this enables rather than hinders knowledge, and increases inter-researcher trust.

Transparency should not be much extra work. During the research process you should take high quality notes for yourself. You will often return to your data and research in the future and thus need those notes. This is a best practice with or without sharing your work. When you engage in this best practice, you have a deep familiarity with your data and can draw meaningful conclusions and easily redact identifying characteristics in your data in the case of qualitative research. In case you cannot share data, you can still reveal the design and expectations; or allow controlled access to the data. Human subjects must be protected at all costs, and yes this often means data sharing is not possible .

The ‘transparency work’ of the qualitative research process can be reduced by software platforms that provide semi-automated annotation and coding. Even if you do not share data, you can build an open workflow from the beginning that allows others to understand every step of the data generating process. However, this work can also be extremely tedious and the incentives not immediately clear. More fruitful discussion if not research assistant funding is needed in this area moving forward.

If you are using quantitative methods, immediately stop hiding your work. If you ran 100 models and 99 did not support your hypothesis, then this is your finding. If a journal does not want to publish this, point the editors and reviewers to the importance of null results and the problems of publication bias. If they still refuse, consider boycotting this journal and sharing your negative experience in public.

Preregistration

Preregistration can drastically reduce bias and hacking prior to collecting data. When you clearly outline your plans including how you will analyze the data, before conducting the research, there is little room for hacking so long as you stick to the plan. Moreover, preregistration can be done directly with a journal although sociology journals are laggards here because they generally do not offer this option. In a preregistration, even if you just put an pre-analysis plan or research design and goals online, you must think much harder about factors such as meaning, causality, inter-subjectivity and ‘how the world probably works’. You cannot hide behind results in this process and therefore you must anticipate counterarguments and explore counterfactual logic. This improves the clarity of theory and research, creating an immense gain in efficiency and effectiveness.

Regardless of the methods you use there are many opportunities to take advantage of preregistration. Some forms of qualitative research, for example those involving grounded theory and interpretivist methods, require decisions during the research process that cannot be foreseen. This uncertainty can be outlined in a preregistration stating explicitly when flexibility is and is not admissible. Moreover, simply putting a qualitative research plan online prior to conducting the research is equivalent to a pre-analysis plan. This research design need not compromise your data collection work because you can register the plan on a platform like the Open Science Framework and then embargo it, so that it is preserved but not made public until after the research concludes. Some scholars using quantitative methods might assume that preregistration is not possible because they work with secondary survey data. But the regularity and release of these survey data are known in advance, and these scholars can preregister their studies before the next round of data are collected with the knowledge of which questions and countries will be available.

Decommodify science

The central functions of the scientific publishing industry are printing and disseminating knowledge, which historically solved a problem of how to share knowledge across universities and countries. The business functions of publishing, however, come with harmful byproducts. Publishing firms extract profits from scientists twice. First, scientists provide free labor in the form of editing and peer reviewing, in addition to producing the results for the articles to be printed. Next, researchers, or their employers, must purchase the product of their own labor; labor not paid for by the publishers. The journal article as a product comes at a high cost, and often only in packages of journals meaning that universities have to pay for extra material their scholars do not use.

Sometimes publishing houses neglect science in favor of profits, but Elsevier has been particularly problematic. They sponsored weapon fairs, created and sold ‘fake’ journals to pharmaceutical companies to publish ‘results’ supporting their drugs, purchased the Social Science Research Network and created paywalls or removed legally shared working versions of articles, charge fees for open access articles, and actively lobbied against open access legislation (For a concise summary with links see Tal Yarkoni’s blog entry). This brought massive counter movements against Elsevier in the scientific community (for example, The Cost of Knowledge). You can take action and refuse to review for or publish with unethical publishers if you feel it is justified. Thus, you should inform yourself about the publishers. Your libraries are a source of information, because they deal with the business side of publishers.

If you are in Europe, check if your institution is a signatory of ProjektDEAL. A consortium of universities are collectively bargaining with publishers via ProjektDEAL demanding that publishers reduce fees and eliminate the double paying of universities. The primary objective is that publishers sign country-wide subscription agreements that enable access for all universities at once. Wiley agreed to such a model and this marks a paradigm change. It indicates how the publishing industry looks in the future, so long as the OS Movement proceeds. If you are not in Europe, consider starting a similar initiative, for example the entire University of California system of 10 universities, 5 medical centers and several research institutions that collectively produce roughly 10% of the world’s academic publications recently followed ProjektDEAL and boycotted Elsevier.

You can work around the publishing business. Prior to submitting an article or after it is published, you have the right to share a preprint – a draft of the paper you share publicly so long as it is not published elsewhere or sold for profit. Posting preprints reduces the power that publishing firms have over science, in addition to giving others immediate access to your work. But simply posting preprints on your academic website is not open enough. Use a preprint service, for example through the Open Science Framework, to ensure that your preprints appear in search engines such as Google Scholar. SocArXiv for example, is the go to location for sociology. This enables scholars to find and directly access research results based on the words they contain, uninhibited by paywalls – a crucial aspect to practicing sociology in the Global South. Preprint services are free and open access.

P-hacking. Religion and science aren’t that different

Science and religion parted ways long ago. This is a historical struggle over power. If science claims to disprove that the earth is the center of the universe or that evolution undermines creation, it might falsify religious doctrine, said to be the word of a God or Gods and thus the ultimate Truth. Religions rely on their claims to this Truth to convert people to submit to their institutions. If science undermines this Truth, it undermines religious power. And power is something that changes human behavior; they might lie, cheat, steal and kill to get or preserve it.

Wikimedia Commons: Thinker; Passion

But science and religion followers are not that different. Actually, they are the same. They are human.

Power is another way of describing status and prestige. In science, we know all about status. Scientific status comes from recognition. From making scientific discoveries and claims that garner attention. In particular, attention in the form of citations.

The absence of market prices results in prestige becoming the main reward and high prestige becoming the measure of exceptional ability. Rent seeking in academia, therefore, produces ego-maniacs and much destructive behavior

Sørensen (1996, p. 1358)

The seeking of status, what economists and Sørensen label as a form of ‘rent-seeking’, is presumably the reason scientists p-hack, and engage in other forms of malpractice. In some cases they ‘must’ p-hack in order to meet the demands of reviewers. Mostly, statistical research requires significance stars to attain publication. This is changing with the Open Science Movement in recent times, but only in the margins. Research using qualitative methods also requires its own form of statistical significance ‘p-hacking’. To be published, a paper must extract novel ideas from observational data, whether these reflect the actual data or are even based on actual data at all seems to be irrelevant as long as the story looks good to reviewers. Just like the significance stars that look all sparkly and comforting to reviewers of quantitative research.

So humans (scientists) cheat to attain status; intentionally or even unintentionally — without malicious intent because they are conditioned to play with their data until the stars appear. Therefore it should be no surprise to humans (scientists) that other humans (religious followers) also cheat.

If p-hacking in science is playing around with models so that they represent the data in a way that matches the researcher’s desire for status, rather than portray the results of scientific tests, then p-hacking in religion must be to interpret the dictates of God (or Gods) to fit one’s, or one’s group’s own status goals.

The conflict of science and religion it like p-hacking. Its a power struggle. Who has the power to make claims about the way the world is and the way it should be? Religious followers would attribute this authority to God, and then themselves as seekers and messengers of God. Science followers attribute this to factual knowledge about the world and then themselves as the testers and reducers of uncertainty to ‘uncover’ those facts. In both cases the process is corrupted by status seeking, a fundamental fallibility of humans. When acting as scientists and spiritual seekers, we are fundamentally still primates, and as such tend toward hierarchy, with many of us human-primates willing to cause harm to others in order to attain higher and higher positions.

For religious followers to gain status through p-hacking they would have to adjust the ‘word of God’ or the ultimate Truth in a way that it (a) is no longer a religious or spiritual truth so that it (b) serves their own ends. Do we have evidence of this practice? Wars fought in the name of religion do not really fit the criteria, as wars can be justified as right, as God’s (or the Gods’) will, for example Christian New Testament Revelation 19:11 about the righteous warring against the (presumably) non-righteous (i.e., ‘evil’); Christian/Jewish Old Testament Deuteronomy 20 calls Israelites to war against cities that do not accept their terms; Islam Qur’an 22:39 advocates war in self-defense and possibly 4:74 to fight in the name of God; and Buddhism taking the stance that war might be necessary in defense but not justified as an aggressor.

The point is that it is difficult to find direct evidence of p-hacking by religious followers in order to gain status for themselves or a group. The same problem lies with detecting p-hacking in scientists. Given that all sides in all wars tend to claim righteousness under God (or Gods) it seems obvious that some (or all) are misinterpreting what should be God’s will for their own gain. Given that so many p-values lay below 0.05 in published research, we can assume that not all are derived from a clean research design, method and presentation of results.

Openness is not needed because we are untrustworthy; it is needed because we are human

(Nosek, Spies and Motyl 2012, p. 626)

It is not only religious and scientific institutions that are antagonistic given their seeking of power. Political institutions, who wield a monopoly on force in the modern world divided into sovereign nation states, also do not always get along with both religious and scientific institutions as they have their own p-values to guard.