Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

Not Wrong: The Replication Crisis is a Metatheory Crisis?

Wanna know what that is? Its theory.

We talk about reproducibility as a methodological problem. We debate p-hacking, publication bias, and analytical forking paths. But the elephant just chillin in the corner is that our theories are weak. There are too many other plausible explanations for what we observe and measure. Sure, the things we measure are complex, but without more clarity, even Commisioner Odenthal isn’t gonna find us a clue to what we are testing. Meaning that we don’t really know what to make of our findings. Other than hoping they are publishable and packaging them neatly to increase those chances. The findings are not right or wrong, they are not really interpretable. At least not with weak theory.

Part of the problem is that unlike the ‘harder’ sciences, consciousness is in the equation. We are measuring phenomena far more complex than quantum physics. Another part is simply that we do not spend much time on theory. In one of my favorite esoteric and less-mainstream open science writings, Anne Scheel pointed out “Why Most Psychological Research Findings Are Not Even Wrong”.  Anne, I see your discipline specific arguments and raise you all social and behavioral science disciplines. We often don’t know what the estimand is. We don’t know what we are looking at. If our theories don’t explain the phenomenon in the first place, debating whether an empirical finding is statistically “right” or “wrong” is moot.

To confront this, I pitched a method I’m working on with my former doctoral student turned postdoc Hung H.V. Nguyen. It’s called Metatheoretical Multiverse Analysis (MMA) and I wanna kick some ‘theoric’ with it (cool huh? It rhymes with lyric and could mean theory in action…). What follows is a summary of the ensuing debate, full of interdisciplinary friction, economic curmudgeonism, and philosophy of science moments.

The Pitch: Metatheoretical Multiverse Analysis (MMA)

To understand why metatheoretical multiverse analysis is necessary, we have to look at how we currently treat plausible theoretical arguments about the data-generating model… Say what? I mean how we use and apply logic.

Imagine you are testing the effect of X1 on Y. But there is an unobserved confounder, X2, which causes both X1 and Y. If X2 is not in your test, your results are uninterpretable. You don’t know if X2 is confounding the results. And, don’t even get me started about X3. This might be a collider. But the theories are not clear in your subfield. So when we try to compare them we end up with several conflicts or unknown paths. These are all alternatively plausible theories and from them we have a multiverse of theory.

Now, imagine I have five equally plausible alternative theories explaining the same phenomenon. This means there’s a 20% chance the theory I am using is the correct one. But I cannot imagine any social or behavioral science where there are only five plausible theories. There are likely thousands. We only test five because it takes entire careers to write semantic theory. And since it takes 20-40 years for most of our discipline to forget those long-winded theories of the past maybe this is a sinking ship…. I digress. Let me re-gress instead (fancied opposite of digress which happens to be a statistically procedure too). 

This is where Metatheoretical Multiverse Analysis (MMA) comes punching (or wrestling or kicking) in. If a standard multiverse analysis runs all reasonable empirical specifications for a given dataset, a metatheoretical multiverse analysis would logically compare all these theories. Somehow…. That’s where the idea gets a little sticky. Metatheory is mostly something people write about, rather than try to formalize and analyze with math. But if they could, our method would then show them where they need to invest their theory-building efforts.

I opened the floor. The timer started.

The Economist’s Dilemma: HARKing and the Illusion of Theory

The pushback was immediate, insightful, and brutally honest. The first counterargument highlighted just how difficult theoretical work is. As one applied economist noted, reading James Heckman makes you realize how brilliant deep theory can be, but spending your career trying to come up with sufficient conditions to definitively disentangle two competing theories is a great way to find yourself out of academia before you get tenure. Remember the long-winded argument?

But a more damning critique came from the reality of how theory is actually utilized in modern economics. To commit a “statistical sin” and assign causality, one researcher pointed out that the lack of robustness in our fields might stem from the fact that HARKing (Hypothesizing After Results are Known) is practically the norm.

In economics, the workflow rarely starts with a pristine a priori theory. Instead, a researcher finds an intriguing empirical pattern in the data, and then builds a formal mathematical model to justify the findings. We only take the time to do the exhausting math if we already have a paper we want to publish. Often, the formal model is something requested by Reviewer 2 at a top journal ex-post. Because the theory is engineered to fit the data, it doesn’t improve our prior hypothesis in a meaningful way, nor does it guarantee robustness when exposed to new data.

Interestingly, someone from the I4R team chimed in with data to back this up. After checking roughly 15,000 robustness checks in the I4R database, they ran an AI classification to score papers from 0 to 100 on ‘how economic’ they were (i.e., true economic theory vs. a paper on TV habits published in an econ journal). The finding? There was absolutely no relationship between the “econ-ness” of the paper (the presence of formal modeling) and its robustness.

So, the means that even if we were to invest in theory development, we wouldn’t get anywhere because… basically, we suck at it.

The AI Revolution

These days there’s always gotta be something about AI. If not everything. The egomaniac AI that I asked to help me write this just couldn’t wait to point out how many people were talking about it.

Funnily we came to an AI discussion through a strong defense of the structural approach in economics (after the initial dust had settled and we could again breathe the cool Barcelona air-conditioned summer air). Economics, one participant noted, used to be purely theoretical because, prior to the 1970s, we simply didn’t have the capacity to handle large data. The empirical revolution brought an era of data-mining into play. P-hacking came into full swing. The academic industrial complex had taken over thanks to secondary data spewing forth from society.

But the tide now is actually flowing back toward structural modeling. Thanks to… AI? The massive amounts of data we have today, combined with AI’s unparalleled ability to explore patterns, is killing the human comparative advantage in pure empirical data mining. AI will always be better at finding patterns (at least an AI or data scientist tells us this, Gemini still can’t perform better regressions than I can IMO). The only place human researchers will retain a comparative advantage is in thinking, being creative, and structuring the data generating process. We don’t need one perfect theory to describe reality anymore (not that that ever worked for us) good approximations are good enough now – and good approximations means…. Drum roll and someone on the mic saying “Ya’ll ready for this?!”. PREDICTION. The better we can predict things, the less we will need to explain them, they will just become some sort of facts in our life worlds. Won’t they?

Metatheoretical analysis might just be the structural framework we need to survive the AI transition. I mean, I invented it, so I’d like to think so.

Bias, DAGs, and Kung-Fu Nancy Cartwright

As the 90-second buzzer kept interrupting and resetting, the conversation evolved into the philosophy and sociology of science itself.

From a legal and equality perspective, an important question was raised about those “five theories” we tend to rely on (you know that when equally plausible reduce the chances that any one is correct down to 20%). Where do they come from? Historically, they have been generated by a very specific demographic—predominantly white men from the Global North. By restricting our empirical tests to a handful of established theories, we inadvertently perpetuate biases and ignore alternative paradigms that might emerge from the Global South. A metatheoretical multiverse approach, by automatically generating and considering thousands of models, might offer a mechanical antidote to this historical bias by forcing us to acknowledge the vast space of un-theorized realities.

But are DAGs really capable of saving us? The room had its doubts. Hey Mister Jack… I’m talking to you. The world, as one researcher passionately argued, is not a clean DAG with X1, X2, and Y. It has X15, Y7562, and a zillion unobserved mediators. Furthermore, DAGs can’t handle cyclical relationships and feedback loops the most famous relationship in economics, price and quantity, is entirely cyclical, endogenous.

This brought us to Nancy Cartwright and the philosophy of science. If we are mapping out thousands of theories, we must remember the Popperian ideal: a model must be falsifiable. If a theoretical model cannot be thrown out under certain conditions, it ceases to be a model and becomes a religion. Metatheoretical multiverse analysis is only useful if we have the empirical tools to actually falsify the branches of the multiverse we generate.

The crow cheers with Nancy’s MMA kick to the face.

The Preregistration Battleground

You cannot talk about theory and open science without stumbling into the debate on preregistration. I posed the question: ‘Does preregistration inappropriately constrain our theories?’ because I want to seem smart and provocative. If there are hundreds of theories, forcing a researcher to write down a specific one beforehand essentially chokes the multiverse before it can breathe. Choking is forbidden in MMA by the way.

The responses were polarized:

Some argued that preregistration forces a hypothetico-deductive model onto fields that generate knowledge inductively. In economic history or sociology, research is exploratory. You learn from the data. Preregistration actively harms inductive discovery.

Others pointed out that preregistration is incredibly difficult if not overrated for secondary data analysis. Datasets are messy, collected for non-research purposes, and require deep exploration just to understand how missing variables are coded. Recently someone pointed out on LinkedIn that my praise of Neumark is overrated. An MMA sweep kick, totally permissible in the sport – touché.

Conversely, a third group (there’s always a third group otherwise it feels incomplete) advocated that preregistration is simply a record of where you started. It prevents the ex-post invention of stories and increases transparency. If ‘Mostly Harmless Econometrics’ had a chapter telling students to pre-specify their hypotheses before opening the dataset, it would have transformed the culture of economics entirely. To late, the historical institutionalists won.

Where Do We Go From Here?

As we wrapped up the session (and prepared to face the blistering heat outside in the name of finding Fideuà), the consensus was clear: our methodological tools have far outpaced our theoretical foundations. We have built incredibly sophisticated empirical engines, but we are putting them in theoretical chassis that are fundamentally flawed. I liked this shift. Oh wait I kinda led the discussion there.

Metatheoretical multiverse analysis is not a magic bullet. It will not solve the fact that social science involves the unpredictable chaos of human consciousness, nor will it easily map the cyclical, non-DAG-friendly feedback loops of the global economy. But it is a start. It is a way to stop pretending that our opportunistic, post-hoc theories are the only valid models of reality. It might trigger some epistemological soul searching if nothing else. By mapping the vast space of plausible theories and systematically testing where they align and where they conflict, we can begin to rebuild the credibility of our disciplines from the ground up – in theory (which is probably weak, so take it with salt).

Furthermore, as the discussion highlighted, we need to completely overhaul how we incentivize work. Open science has an image problem we often make it look boring, framing it as an adversarial compliance checklist rather than a thrilling pursuit of truth. We leave it up to early-career researchers to awkwardly teach their supervisors about open code and data. And shoulder them with the onus of spending hours making all their work open and reproducible, hours that their forefathers (yeah, they were 95% men) didn’t need to do. Heck these forefathers could publish like 2 papers per year and get tenure.

But since its my blog post, I get the last word: If we want to fix the replication crisis, we cannot just mandate better code and policies. We have to foster better theory. We need to acknowledge the sheer size of the theoretical multiverse, embrace the complexity, and start theorizing before regressionizing. In this case, sadly, the answer is not as simple as 42.

The path toward ethical science is paved with diamonds

Over the last two decades awareness of a Reproducibility Crisis penetrated all scientific disciplines. Studies recently showed that 90% in STEM and 95% in psychology were aware of this Crisis1,2. At its core this is a crisis of trust. The findings scientists publish and tout as facts, are often not replicable by others. Moreover, a great many are not even computationally reproducible using the original data3,4.  The Open Science Movement developed partly in response to this Crisis5,6. Especially in the last two decades, researchers organized grassroots movements to make science more transparent, reliable and ethical.

One of the many tenets of the Open Science (OS) Movement is open access. Published research results must be accessible to everyone. This is not a new idea. In the 1940s Robert K. Merton developed normative goals necessary to ensure integrity in science7. One was that all scientific findings should be public property. He argued that effective and efficient progress of science depends on open public access. In data science this Movement led to the concept of Open Science by Design – a strategy where scientists plan in advance how they will make all metadata, data and algorithms available and easily re-usable by anyone8.

The OS Movement has been successful. Open access journals and articles increased rapidly. They outpaced the global increase in publications in other formats in the last two decades9. One problem is that these open access articles are overwhelmingly funded by the authors of the studies via Author Processing Charges (APCs). It costs over 10 thousand US$ to publish a study open access in the journal Nature, and most other journals charge at least two thousand. Except for elite institutes and projects with generous third-party fundings, these fees are not usually covered or coverable by universities. A naïve outsider might ask why scientists would pay astronomical fees to make their own hard work publicly available, when they can simply share it online in a free repository like the Open Science Framework, Github or any number of preprint servers based on the ArXiv model?

The answer is competition. The strongest norm governing the practice of science today is publish-or-perish. In every field and science at large, scientists are judged based on their publication record. Ask any academic what the top journals in their area are, and you will get relatively consistent answers by discipline, sub-field and science in general10. Look at the faculty of any ‘top’ university, and the faculty in any discipline will have one or more publications in these ‘top’ journals. Without publications in these journals, a scientist has no chance at a career in science. This is a collective cultural fact. One deeply institutionalized in the organizations, rules, norms, expectations and behaviors of scientists11,12.  

Thus, the founding of new open access journals with lower APCs has little impact on the scientific enterprise because they cannot compete with institutionalized legacy statuses of existing journals. The greatest success story is the non-profit publisher PLOS, which rose swiftly in the rankings, but could not crack into the very top tier. Although far cheaper than Nature, it is still not ‘cheap’, with APCs in their family of journals ranging from 2.5 to 3.2 thousand $US[1].

The most successful open access models are within existing paywalled journals where authors have the option to publish “gold” open access, rather than publish for free. The payment of somewhere between two and 10 thousand US$, gets authors the right to have their single article published open access inside of a closed access journal. This means that despite a massive shift toward open access, the fundamental structures of scientific publishing have not changed.

Ethical Implications

Competition in science supports motivation and innovation, but the publish-or-perish norm is a toxic externality. Scientists willingly prioritize subjective journal ranking, impact factor and increasing their citation counts to get ahead within the competitive scientific enterprise. They do this despite widespread skepticism and evidence that rankings and citations do not correlate strongly with the quality and reliability of published studies14–16. They do this because they must, or at least perceive that they must, in order to follow their scientific career aspirations. The pressure to publish thus motivates rent-seeking behaviors designed to increase publication chances, i.e., career chances. Conducting higher quality science is one of these behaviors. But there are many methods to increase publication chances that are science orthogonal, what are known today as questionable research practices (QRPs)17,18.

Some QRPs are unconscious and learned from supervisors in the process of converting research into a publishable paper. For example, researchers routinely run many statistical models but report only those that show the strongest support of their claims. Researchers who believe that their claim is true in the first place will gravitate toward models that support it, convincing themselves intrinsically that these models are the best tests of their claim.

Decisions based on confirming intrinsic beliefs or window dressing for peer reviewers reduce the replicability of science. They narrow down the multiverse of potential findings into a highly selected set of results. This selectivity is independent of the process that generated the data in the first place, in other words, it is science orthogonal. A classic example is a study published suggesting that hurricanes with feminine names cause more damage than those with masculine names. It turns out that using the available data, the original researchers selected a model that produced regression coefficients that were extremely far away from the central tendency among all other plausible models’ regression coefficients19. We do not know whether this was a conscious decision but their reporting hides the truth and simultaneously increases publication chances.

There are of course conscious and highly unethical behaviors leading to an entirely false representation of reality. Science is filled with scandals of hacking and data-faking. Rent-seeking alone can explain unethical learned unconscious and conscious behaviors of scientists. They seek status and money and job security, or in some cases seek results that support a particular worldview or policy outcome20,21. But rent-seeking is unambiguously the main reason22,23.

If we as a scientific community and science-interested public, want to eliminate the perverse incentive structures that bound and inform rent-seeking behaviors of scientists we need radical change. Status should not be assigned based on journal metrics that often have little to do with the quality of the research being conducted or published and more to do with legacy and embedded norms. The most radical proposal is to eliminate journals altogether. But this is an extremely unlikely outcome no matter how powerful the OS Movement becomes.

Journals became standard in science hundreds of years ago as a means for communicating scientific discoveries across time and space24. They enabled scientists to acquire knowledge without travelling to faraway universities. Publishing was not cheap, and with the dawn of digital media, it became even more expensive as publishers raced to provide their journal both in-print and online. The costs associated with publishing gave publishing firms a great deal of power over time.

The result is that companies like Springer Nature and Elsevier have enough power to dictate to scientists how they perform their research and communicate their results25. They shape academic careers, institutional priorities, governments’ science policies, and perceptions of journals and the publishing enterprise among scientists26. Their primary legal interest is their shareholders. Profit is their priority, and only second is to provide a service to science as their product. A simple mathematical proof confirms that profit is their priority: If they cannot make a profit they will no longer provide the scientific services but if they can make a product without scientific services they have no reason to stop.

Although big publishing firms have shown many draconian practices in their ‘service’ to science27–30, scientists still need a means to communicate their results with each other and the public. This can be done without for-profit publishing but probably not without a journal publication format, or something very similar. For example, we now have a plethora of ‘green’ open access preprint servers where scholars can deposit working papers, or prior versions of their published articles. These are a viable means to communicate science without the perverse incentives generated by big publishing or institutionalized journal rankings. This all still involves the journal article as the standard unit of science production.

Diamond Open Access

If we are ‘stuck’ with journals in science as our primary communication medium, then they should be as free from perversely incentivized bias as much as possible. The first step is thus to remove for-profit publishing from the equation. It is fine to use the services of for-profit publishers. They have shown the capacity to provide print on demand, marketing and scientific journalism. But when they control and direct the scientific enterprise when have an ethical conflict of interest – namely profit versus robust science.

The costs of publishing a journal are the lowest in history. There are publication kits that help associations, institutions and stand-alone journals to take publication into their own hands31. With minimal costs and self-governance, journals have no need to push institutions to purchase journal subscriptions and no need to charge authors astronomical publication fees. Thus, they themselves are not perversely incentivized to perpetually increase their status to make their product profitable.

The optimal existing solution is diamond open access, whereby a journal charges no APCs and is freely readable and downloadable online32. It is the most ethical and equitable by design, and is perceived as the ideal model by most academics33. The OS Movement is overwhelmingly in favor of diamond open access34, as are governance bodies – at least those free from the influence of big publishing like UNESCO35,36. Despite 13 thousand journals indexed in the Directory of Open Access Journals (DOAJ) with no fees as of February 16th, 2026, these journals are not on the radar of most indexing services, and do not belong to the mainstream of journals published by major scientific societies37. This means that successful diamond open access journals are very rare.

Because of such a saturated scientific ‘market’ for publication outlets, starting new diamond open access journals has had little impact on producing a more ethical science. They do not gain reputation. The reason PLOS was so successful, despite failing to break into the highest echelon of science, was because it has a huge cash flow and can use it for branding and promotion. From an ethical science perspective it is valuable, like gold, but it is not as valuable as diamond.

Flipping or Starting Over?

To have the highest ranking and most well-known journals diamond open access, societies need to cancel their contracts with for-profit publishers38. One problem with this is that many academic societies are themselves run like for-profit businesses. They seek to generate as much revenue as possible, and diamond open access would threaten this model. They have embedded relationships with publishers and agree to renew contracts together, mostly independent of the scientists they serve. I witnessed this first hand with the American Sociological Association and Sage39,40. Contracting a for-profit publisher provides a non-profit academic organization a scapegoat for amassing capital via subscription fees.

There are incredible exceptions; however, and they offer model success stories. Computational Linguistics is one of the earliest journals considered to be among the top in its field, to flip to diamond open access. The first step was a move by the Association for Computational Linguistics in 2002 to create the CL Anthology which made all articles published by association journals open access online after an embargo period. Then in 2009 the journal Computational Linguistics flipped to diamond open access. The journal Demography of the Population Association of America is another example.

The embeddedness of big publishing in science, and the capital that associations can raise through their journals when run by big publishers, are major barriers to flipping. Another major barrier is that some big publishers coerce scientific societies into signing away the rights to the titles of their journals in their publishing contracts. This tactic is most intensively deployed by Elsevier. They own the rights to the titles of nearly all the journals they publish. Therefore, when these societies want to change publishers, they cannot. They are trapped. They would have to start a new journal with a different title – what happened for example with the Journal of Infometrics41. There is otherwise no way around this problem because of copyright law.

Therefore, collective efforts to build the popularity of new diamond open access journals would greatly increase the movement toward a more ethical and effective scientific enterprise. If these are journals from societies, which are essentially the same journal but with a new (not copyrighted by Elsevier) name, it requires authors to support this journal and immediately abandon the other. Supporting new diamond open access journals, whether completely new or newly named, requires established scholars to put their status-seeking egos aside in the name of scientific progress and ethics. It is precisely those who have built major scientific reputations who need to engage in this change, because they can give them most clout to new journals by touting them and publishing in them.

Scholars and societies alone cannot carry the burden. Hiring committees need to reward these behaviors. Rather than seeing a publication in a new ‘unranked’ or ‘low ranked’ diamond open access journal as a sign that the article is ‘not high quality enough for top journals’, committees should judge publications only on their content. Then, if two publications are seen as equally scientifically rigorous and high quality, the one in a diamond open access journal should get a greater weight. Scientist who consistently publish high quality research in diamond open access journals should be favorites of hiring committees, all else equal.

A hiring committee would be unlikely to reward an applicant who shows sociopathic behaviors. Following this logic, they should not reward scientists who show behaviors that go against ethical science by practicing closed and profit-incentivized science. I need to be very clear here that I personally am still publishing regularly in journals published by for-profit publishers. I am not that famous, and I do not have tenure. This is a perfect example of why we need sweeping changes at the institutional level, and from the top of the scientific hierarchy.

With efforts to flip both  existing journals and efforts to reward new, ethical journals, we give science its greatest future chances for improvement and sustainability. There are many efforts underway and these should serve as guides, for example the Diamond Open Access Fund from the Dutch Research Council and MIT’s shift+OPEN initiative.

Diamond AI

Diamond open access has a second, equally important role for the future of science. It is necessary to inform Generative Artificial Intelligence (Gen AI). The LLMs that power popular Gen AI are trained heavily on corpora assembled from what is available via the Internet. Paywalled literature is missing from training corpora, not because it is unimportant, but because it is not accessible or legally usable42. Gen AI outputs are therefore based on a highly restricted sample of all scientific knowledge. Open access for all of science would ensure that anyone using Gen AI, would get the best possible information based on all that we know as humans. As essentially everyone is using Gen AI today43, this would mean that the public would be optimally informed.

It is not necessary for Gen AI development that all articles are diamond open access, they just need to be somewhere, e.g., green or gold open access. Yet, without diamond open access the entire knowledge enterprise will continue to favor the work of those with greater resources44. Resources are of course necessary to produce higher quality science because of research costs, but when it comes to publishing and dissemination, a resource advantage reproduces the already existing Global North-South disadvantages in science. This limits human capacity to tap resources in lower income societies. These are societies filled with potential contributions to science that could benefit both 1) their own societies’ development – because it enables them to gain status, resources and build stronger, more attractive and sustainable scientific institutions, and 2) all of scientific knowledge because there are brilliant minds waiting to be tapped that might otherwise give up on science because it is for them no sustainable in their region.

If the knowledge in Gen AI remains Global North and WEIRD biased (Western, educated, industrialized, rich and democratic), the result is that these countries remain culturally and economically advantaged beyond that which exists presently. This means that without diamond open access norms, we are willingly allowing Gen AI to increase cultural hegemony and economic domination of a minority. I am not directly arguing whether this is good or bad. I am in the Global North and profit from a stronger Global North science advantage. My argument is that science should be neutral, favoring the most optimal and reliable knowledge, rather than legacies or strategies that game the scientific system.

References

1.          Baker, M. 1,500 scientists lift the lid on reproducibility. Nature 533, 452–454 (2016).

2.          Metskas, A. How Much Do Academic Psychologists Trust Academic Psychology, and Is There Still a Replication Crisis? Transparent Replications https://replications.clearerthinking.org/how-much-do-academic-psychologists-trust-academic-psychology-and-is-there-still-a-replication-crisis/#survey-demographics (2025).

3.          Open Science Collaboration. Estimating the reproducibility of psychological science. Science 349, (2015).

4.          Breznau, N. et al. The reliability of replications: a study in computational reproductions. Royal Society Open Science 12, 241038 (2025).

5.          Engzell, P. & Rohrer, J. M. Improving Social Science: Lessons from the Open Science Movement. PS: Political Science & Politics 1–4 (2021) doi:10.1017/S1049096520000967.

6.          Breznau, N. Legacy of Jon Tennant, “Open science is just good science”. Crowdid https://crowdid.hypotheses.org/548 (2022) doi:10.58079/ne8h.

7.          Merton, R. K. The Sociology of Science: Theoretical and Empirical Investigations. (University of Chicago press, 1973).

8.          Wittenburg, P. Open Science and Data Science. Data Intelligence 3, 95–105 (2021).

9.          NCSES. Publication Output by Region, Country, or Economy and by Scientific Field. https://ncses.nsf.gov/pubs/nsb202333/publication-output-by-region-country-or-economy-and-by-scientific-field#utm_source=chatgpt.com (2023).

10.        Serenko, A. & Bontis, N. A critical evaluation of expert survey‐based journal rankings: The role of personal research interests. Asso for Info Science & Tech 69, 749–752 (2018).

11.        Mancoridis, M., Sumers, T. & Griffiths, T. Publish or Perish: Simulating the Impact of Publication Policies on Science. Proceedings of the Annual Meeting of the Cognitive Science Society 46, (2024).

12.        Breznau, N. Questionable research practices from the practitioners’ perspectives. Crowdid https://crowdid.hypotheses.org/1666 (2025) doi:10.58079/14f3u.

13.        Lawrence, S. Free online availability substantially increases a paper’s impact. Nature 411, 521–521 (2001).

14.        Fleck, C. The Impact Factor Fetishism. European Journal of Sociology / Archives Européennes de Sociologie 54, 327–356 (2013).

15.        Rushforth, A. & De Rijcke, S. Practicing responsible research assessment: Qualitative study of faculty hiring, promotion, and tenure assessments in the United States. Res Eval 33, (2024).

16.        Dougherty, M. R. & Horne, Z. Citation counts and journal impact factors do not capture some indicators of research quality in the behavioural and brain sciences. R Soc Open Sci. 9, 220334 (2022).

17.        Gopalakrishna, G. et al. Prevalence of questionable research practices, research misconduct and their potential explanatory factors: A survey among academic researchers in The Netherlands. PLOS ONE 17, e0263023 (2022).

18.        John, L. K., Loewenstein, G. & Prelec, D. Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling. Psychol Sci 23, 524–532 (2012).

19.        Muñoz, J. & Young, C. We Ran 9 Billion Regressions: Eliminating False Positives through Computational Model Robustness. Sociological Methodology 48, 1–33 (2018).

20.        Borjas, G. J. & Breznau, N. Ideological bias in the production of research findings. Science Advances 12, eadz7173 (2026).

21.        Rainero, V., Stolz, J. & Luijkx, R. The Faith Factor. How Scholars’ Religiosity Biases Research Findings on Secularization. Sociological Science 13, 154–177 (2026).

22.        Aronson, J. K. When I use a word . . . “Publish or perish”: adverse effects. https://doi.org/10.1136/bmj.r1577 (2025) doi:10.1136/bmj.r1577.

23.        Paruzel-Czachura, M., Baran, L. & Spendel, Z. Publish or be ethical? Publishing pressure and scientific misconduct in research. Research Ethics 17, 375–397 (2021).

24.        Carey, J. Scientific Communication Before and After Networked Science. Information & Culture 48, 344–367 (2013).

25.        Larivière, V., Haustein, S. & Mongeon, P. The Oligopoly of Academic Publishers in the Digital Era. PLOS ONE 10, e0127502 (2015).

26.        Rossello, G. & Martinelli, A. The effect of lobbies’ narratives on academics’ perceptions of scientific publishing: A survey experiment. Information Economics and Policy 71, 101148 (2025).

27.        Butler, L.-A., Matthias, L., Simard, M.-A., Mongeon, P. & Haustein, S. The oligopoly’s shift to open access: How the big five academic publishers profit from article processing charges. Quantitative Science Studies 4, 778–799 (2023).

28.        Else, H. Dutch publishing giant cuts off researchers in Germany and Sweden. Nature 559, 454–455 (2018).

29.        Lancet, T. & Board, T. L. I. A. Reed Elsevier and the arms trade. The Lancet 366, 868 (2005).

30.        Jureidini, J. & Clothier, R. Elsevier should divest itself of either its medical publishing or pharmaceutical services division. The Lancet 374, 375 (2009).

31.        Scholastica. New fully-OA publishing toolkit and stakeholder reflections on 20 years of the BOAI. https://blog.scholasticahq.com/post/oa-publishing-toolkit-and-stakeholder-reflections-BOAI20/ (2022).

32.        Normand, S. Is Diamond Open Access the Future of Open Access? The iJournal: Student Journal of the Faculty of Information 3, (2018).

33.        Kumari, M. & A, S. Perceptions of open access publishing: A comparative study of gold and diamond models among global researchers. Alexandria 35, 55–73 (2025).

34.        Plan S. Working collectively towards an equitable, community-driven and academic-led scholarly publishing model. Action Plan for Diamond Open Access https://www.coalition-s.org/action-plan-for-diamond-open-access/ (2022).

35.        UNESCO. Diamond Open Access. Public Service Press Release vol. Online Report (2026).

36.        UNESCO. UNESCO Recommendation on Open Science. https://unesdoc.unesco.org/ark:/48223/pf0000379949.locale=en (2021).

37.        Simard, M.-A., Basson, I., Hare, M., Larivière, V. & Mongeon, P. The Value of a Diamond: Understanding Global Coverage of Diamond Open Access Journals in Web of Science, Scopus, and OpenAlex to Support an Open Future. Proceedings of the Annual Conference of CAIS / Actes du congrès annuel de l’ACSI https://doi.org/10.29173/cais1845 (2023) doi:10.29173/cais1845.

38.        Trueblood, J. S. et al. The misalignment of incentives in academic publishing and implications for journal reform. Proc. Natl. Acad. Sci. U.S.A. 122, e2401231121 (2025).

39.        Why I’m leaving the American Sociological Association. Family Inequality https://familyinequality.wordpress.com/2021/11/06/why-im-leaving-the-american-sociological-association/ (2021).

40.        Philip Cohen’s ASA Publications Committee platform. Family Inequality https://familyinequality.wordpress.com/2018/01/14/philip-cohens-asa-publications-committee-platform/ (2018).

41.        Singh Chawla, D. Open-access row prompts editorial board of Elsevier journal to resign. Nature https://doi.org/10.1038/d41586-019-00135-8 (2019) doi:10.1038/d41586-019-00135-8.

42.        Szkalej, K. The Paradox of Lawful Text and Data Mining? Some Experiences from the Research Sector and Where We (Should) Go from Here. GRUR Int 74, 307–319 (2025).

43.        Breznau, N. & Nguyen, H. H. V. An Introduction to Generative Artificial Intelligence for Academics. F1000 Research Preprint Status, (2025).

44.        Kwon, D. Open-access publishing fees deter researchers in the global south. Nature https://doi.org/10.1038/d41586-022-00342-w (2022) doi:10.1038/d41586-022-00342-w.

Ethical Statement

I have no conflict of interest to report. I used Google search which includes Gemini by default, ChatGPT, NotebookLM and Nano Banana to support my literature review, check arguments, suggest words or phrases and to extract facts. No writing was copied from Gen AI, it is all my own. Any mistakes or opinions are also my own.


[1] https://plos.org/fees/

Media outlet and Q&A for ‘Ideological bias in the production of research findings’ by Borjas and Breznau

  1. Süddeutsche Zeitung (SZ) article by Sebastian Herrmann “Warum Forscher aus denselben Daten entgegengesetzte Schlüsse ziehen” (Why researchers reach different conclusions from the same data).
  2. Manhattan City Journal perspective piece written by George and I.
  3. A news report by Luca Rehse-Knauf for Deutschlandfunk (German radio) and their Forschung Aktuelle series “Migrationspolitik: Einstellungen können Forschungsergebnisse beeinflussen” (Migration policy: Attitudes can influence research results).
  4. A news report by Katrin Kühn and Luca Rehse-Knauf for Deutschlandfunk and their Fakten und Meinungen Series “Darum sind wir Menschen nicht objektiv” (Why we humans are not objective).
  5. A PsyPost report by Eric W. Dolan “158 scientists used the same data, but their politics predicted the results“.
  6. A podcast on The Last Show with David Cooper. Apple / Youtube.
  7. A podcast in Allegedly Does Not Replicate with the Institute for Replication (I4R) with Abel Brodeur and Juan Pablo Posada Aparicio.
  8. A substack post by Claudio Teixiera after an interview with us about how researchers conduct research and the ‘invisible’ paths they follow.
  9. A substack post by Laurenz Gunther describing our work and digging deeper into bias among researchers (here those working on the topic of immigration) – despite a relatively far out conclusion.
  10. There was a Neuer Züricher Zeitung article about the original ‘Hidden Universe’ study (paywalled) ‘Das Experiment: Wer bekommt die rote Karte?.

— How did you arrive at your research question, and what is the study about?

George emailed me and had a few questions about our original study. He was at that time analyzing our data and had found a statistical association between pre-existing preferences for more or less migration among the teams in our study, and their findings. I was very skeptical. I have now worked on replication and reproducibility themes for almost a decade. I am acutely aware of what we often refer to as ‘researcher degrees of freedom’, also known as ‘the garden of forking paths’. This refers to choices that researchers can make during the research process that can lead to different outcomes. I assumed that the statistical association he found would not hold under different but equally plausible model specifications. I began testing many different models. Basically, they all showed the same result. Therefore, I became convinced that this was more than a fluke.

Actually, George had already run most of the same models. We present all of our models in a multiverse analysis in our paper. Out of 883 models 88% showed a significant statistical effect suggesting that we should reject the null hypothesis that ‘ideology has zero effect on the teams’ research findings’. If we take the assumption that we should only trust models that control for researchers’ educational experiences – something we believe impacts their results – then we find that roughly 93% of the models show a significant statistical effect.

— Could you explain the experimental design and methodology?

This study is an exploratory secondary analysis of the data generated by the experiment of myself, Eike Mark Rinke and Alexander Wuttke. We gave 71 research teams the same data and hypothesis – that immigration reduces support for social welfare policies. We surveyed them on their backgrounds and research experience and asked them if they believed the hypothesis was true and what they thought about immigration policy. In the original study, again the one that I helped lead, we did not find any important impact of immigration preferences. We essentially found that the results went in all directions, and we could not easily explain the variation.

After working together, George and I agree that the statistical analysis in the original study was not a clean test of immigration on the research teams’ results, because it controlled for the statistical model specifications that the teams’ made. The original study was were searching for key decisions that might explain why results went in different directions, so this naturally made sense. But in hindsight, these decisions are the mechanism through which ideology gets transmitted into statistical results. Thus, the original study introduced what is known in statistics and causal analysis as a ‘confounder problem’. If someone has an ideological bias, they will choose statistical models that will lead to more desirable results. The original study was controlling for both the test variable (ideology) and its mechanism (the statistical models), and in doing so it suppressed the impact of the test variable we were trying to observe.

This time, George and I conducted regression analyses in which the research findings were the dependent variable and ideology the independent variable – without model specifications as control variables, but with further controls.

— What are the limitations?

A key limitation is that this study is exploratory, not confirmatory. It relies on secondary data from a study that was not specifically designed to test the impact of ideological bias. It cannot confirm that this bias exists, instead it demonstrates robustly that a statistical association exists between ideology and researchers’ findings. We are not aware of any other way to explain this association other than an ideological bias. But we can only confirm with confidence, that it is prudent to reject the null hypothesis that in this particular sample and study is that there is no association between preexisting preferences for immigration policy and research findings. More specifically, the likelihood of observing the data in this experiment if the null hypothesis were true is very low.

Another limitation is that the size of the effect we found is unclear. It points in a positive direction – more pro- immigration policy stances associate with findings that show immigration has a more positive effect on social policy preferences among the public, and vice-versa with more anti-immigration policy stances and a more negative effect. But because of the great variation in results and the small sample size, the standard errors of the estimated statistical effects are very large. This means that the true effect might be anywhere from miniscule and near-zero, to moderate, to very large. We simply cannot say much about this here. More research is necessary, although this is an implicitly difficult topic to study, because if we inform researchers that we are studying their ideological bias, they might behave differently and this would take away ecological validity.

— According to the study, ideology influences model specifications. Could you provide a concrete example to illustrate how a single design decision (or a combination thereof) can have an impact?

I cannot, and this is another limitation of the study. If I could, it would be something we would have found in the original experiment that collected the data. But we can only point at patterns here. There are certain model specifications, unique combinations of statistical modelling choices that produce more negative results. The teams with more anti-immigration ideologies were more likely to choose these. But there are far more model specifications than there are teams. This leads to a sparse data problem. There are many empty cells in the matrix of all possible model specification combinations that teams would plausibly make. This makes it roughly impossible to pinpoint exact specifications’ effects. and there are many different model specifications that can lead to a positive or negative statistical effect. The point is that the only thing that happened between the teams asked to test the hypothesis with the same data, was different modelling choices. Therefore, this is the only way they could arrive at different results. There was no cheating or result faking, we checked that their statistical code produced the results they reported to us.


— To what extent is this a problem, and to what extent is it normal that decisions, based on analytical decisions, depend on who you are and how you think?

This is nothing that our study answers. And it may not be fully possible to answer because we do not yet know the nature of consciousness. We also cannot measure what is happening inside a human neural network – a brain in other words. But it is clear to me that experience, ideology and preferences shape results. A simple example is statistical training. Many researchers have limited statistical training, and they build only those statistical models that they learned about in their studies. This impacts results.

But more generally idiosyncrasies of people, like ideology, shape what research questions that people are willing to pursue and how, and they shape the reporting of those results. Some could look at our study and think that the estimated statistical impact is large and highly concerning. Others, might look at it and think it is tiny and of no concern at all. Our study suggests that ideology can explain somewhere between 1 and 3% of the variance in the results. If scientific findings are on average 1 to 3% off of what they would be without bias, is that a big problem? I mean… what do you think?

— If I understand correctly, the experiment was originally intended to show how much the results diverged, not why. How did you arrive at ideology as a possible cause?

As I already mentioned, this is something that George noticed in our data. He already had this hypothesis in his mind. I cannot blame him for thinking this. The Open Science Movement and Metascience work reveals many so called ‘Questionable Research Practices’. These include everything from faking data, to tampering with statistical models or stopping the collection of data during an experiment to produce a desired result. These practices are designed to produce certain results in order to obtain a publication or support a pet hypothesis (confirmation bias). Obviously some of these studies were motivated by ideological goals.

— How could ideological bias be reduced? Is this even desirable, or should we simply be aware that it can exist?

The impact of ideology can be reduced by following some clear recommendations of the Open Science Movement. Studies should be pre-preregistered, they should provide all code and materials, they should not be conducted in isolation or in hiding, and researchers should cooperate. Some of us are engaging in so called ‚Adversarial Collaborations‘, where researchers who do not agree – those with different priors about a given hypothesis like the impact of immigration – collaborate. They lay out all the aspects of a study in advance, and they agree on what evidence would count as support of either of the positions. I highly recommend this form of science. It takes competition and turns it into collaboration with the goal of knowledge seeking prioritized above all else.


— You are investigating ideological bias in science using scientific methods. How do you deal with this tension in meta-scientific questions, where you are essentially also your own subject of investigation?

Similar to my last answer, one cannot fully understand or deal with one’s own bias, and therefore needs to build in checks into the process. Things that would reduce this bias, like preregistration. I am working currently on a project that is an Autoethnography of my own questionable research practices and the perverse incentives I encountered during my career in science. I hope that by doing this, and revealing my own behaviors, I can improve them. I also want to be a role model for others, to make it desirable to be highly critical of one’s own work. My goal is to get this study published in a high quality journal and thereby prove that self-criticism, something researchers mostly try to avoid to protect their theories, findings and careers, is something that can be used in a positive way in the scientific process.

Can one’s own attitudes always play a role, even in this study?

Sure. Definitely. That is why it was important for me to take a so called ‘multiverse’ approach to this study, and many other studies I am working on recently. I want to ensure that I, or one of my colleagues, has not simply selected a statistical model that produces certain results. This practice, known as hacking, or p-hacking, is prevalent in science, especially in secondary data analysis. I essentially learned to do this during my graduate studies. We would find a result we liked and then develop convincing logical arguments why the model producing it must be the best model. So, the idea with multiverse analysis, is to run all or at least all plausible alternative models. This helps reveal if my model is an outlier. Whether it represents something very unique or unusual in the distribution of model results. If it does, this is a cause for great concern. If not, it is evidence of a robust statistical association.

— There have also been critical reactions, for example in this online Bluesky thread, which raise concerns about George Borjas’s views and background. How do you assess this criticism?

I have read the discussion thread by Michael Clemens, an economist at George Mason University. One line of criticism appears to concern the fact that George Borjas recently conducted research for the executive branch of government.

Another criticism seems to focus on the observation that different model specifications yield different results. This is precisely what our study confirms, and it is also a well-established fact in the history of empirical social science.

Both points are orthogonal to our study. Even if we were to assume, hypothetically, that George held some form of ideological bias and that this bias influenced his analytical choices, which I cannot confirm and do not claim, this would not undermine our findings. The reason is that we adopted a deliberately robust research design. We conducted a multiverse analysis comprising 883 regression models. The consistency of results across this large set of plausible specifications makes it implausible that our conclusions are driven by special highly selective model specifications – those that would be selected due to ideological bias.

It is also important to note that George and I do not share the same political views. Precisely for that reason, ideology was an additional motivation for us to adopt a highly robust analytical strategy. I explicitly advocated for the multiverse approach in order to minimize the influence of individual priors, including our own potential ideological orientations. I see this as a great strength of the study.

The purpose of our approach is to decouple empirical results from personal or ideological preferences as much as possible. And to estimate the robustness of our finding to any kind of bias, not just ideological. I would encourage critics to apply the same standards of robustness to their own work. To date, Mr. Clemens has not presented empirical evidence that contradicts our findings. Moreover, criticisms referring to modeling choices in studies conducted by George decades ago are not relevant to the validity of the present analysis.

— What does it mean when you say that only 3% of the variance can be explained by your results.

That has to do with the regression coefficient and the r-squared values. A coefficient of 0.03 shows that a one-point higher (more positive) ideology mathematically predicts a change in results of 0.03. This sounds meaningless, but we know that 0 would be none (and we can equate this with zero percent change) and that 1 would be 1-point on a standardized scale. This is 1 standard deviation in the distribution of the dependent variable. It would be possible to move the results more than 1 standard deviation, but this would be quite preposterous. There is nothing in the complex nature of social science, which lacks laws, that would do that. So I will set the upper bound of the largest possible effect at 1, meaning that 1 would equal 100% of the distance in the distribution of variance.  Therefore, 0.03 is like 3 percent of the distance. At the same time, the r-squared, which tells us how much of the error is reduced from this particular variable, is around 0.03 or less depending on how we measure this variable. This suggests that fitting the observed ideology values into the observed results from the teams, reduces the unexplained variance from 100% down to around 97%.

— Can you explain — very, very simply — what you did? 

In a study that I co-led starting in 2018 (Breznau, Rinke and Wuttke et al. 2022), we designed an experiment that allowed us to observe researchers doing research on the impact of immigration on social policy preferences. We gave them the same data and asked them to answer the same research question: whether immigration reduces support for social policies or not. We documented the researchers’ pre-existing methodological training, experience with and expectations about the topic, and their personal preferences for looser or tighter immigration laws in their own countries. We shared all of the data and documentation of our work publicly. Because we shared our data, a few years later George J. Borjas was able to reanalyze our data and find new evidence of a correlation between the researchers’ ideological positions on immigration and their findings. Those with more pro-immigration positions tended to find evidence that immigration had a positive impact on support for social policy, and those with more anti-immigration positions tended to find evidence that immigration had a negative impact on support for social policy. Here “social policy” means support for a more extensive welfare state providing social security via the government or not – many would call this support for social cohesion. I was skeptical of George’s initial work, and together we vetted George’s findings. We ran almost a thousand alternative statistical models to test George’s finding and of these 88% suggest that we should reject the null hypothesis. The null hypothesis is that if there is no impact of ideology on researchers’ findings, that we should not observe what we observed based on probability. But we did observe this association, and by rejecting the null hypothesis we have evidence that something more is going on, not just random luck in the data. Crucially, we should reject the idea that ideology has no impact on researchers’ results when analyzing the same data. 

— What motivated you to do this study?

As I said, George found this association between researchers’ ideological positions on immigration and their research findings. This is an important scientific observation. I was a bit more skeptical, in particular because my initial analysis of the data as part of the original experiment did not show this association. Together we became very motivated to tackle this problem, and in the end, we have relatively strong evidence of something. The exact nature of this should be subject to further research.

— What are the most striking findings? 

The main finding is that researchers’ own preferences for tighter or looser immigration predicts what they went on to find in their work. This appears striking, but when put into context it is maybe not so surprising. We know that there are many reasons that researchers may consciously or even unconsciously exert influence on their own findings. There are many cases of researchers engaging in questionable research practices to ‘fudge’ their data and results in ways that make them appear stronger than they really are. This has occurred frequently in biomedicine, for example in studies of products whose approval would net the researchers great personal profits. Consider also what we know from psychology, namely confirmation bias. People tend to seek out evidence of what they already believe is true. Researchers are people too. When presented with various forms of competing evidence, a researcher might gravitate toward that which supports their preexisting beliefs or preferences. 

— What are the implications for the social sciences, which are already reeling from the replication crisis? 

The implications reinforce what we already know. We should focus on refining two areas of science. The first is scientific training. We need to make transparent and open workflows and data sharing the norm. This includes researchers stating in advance what they plan to do and what they expect to find. When researchers do not do this, they can run several experiments or analyze hundreds of datasets with millions of different statistical models and simply choose one finding that looks exciting or sexy to them. Generations of researchers before us have learned to do exactly this in order to get published. They learned to ‘sell’ a single selected finding as confirmatory evidence of something in the real world, when in fact it is simply a highly selected, exploration of data leading to a unique event; one that probably does not generalize and is not reproducible (a.k.a. luck). I bring up the idea of getting published here because this is the second problem we must urgently address in science. A researcher’s worth or ranking as a scientist is judged almost entirely on their publication record. In particular, publications in journals that are considered higher status. These higher status journals should be publishing studies because the studies contain higher quality science. But there are ways to game this system. There are a a host of questionable research practices that makes findings look more exciting than they actually are. These practices often lead to irreproducible findings. Essentially fake science. Big publishing is a major profit industry and this increases the pressure to publish, as the publishers of the journals engage in questionable, sometimes unethical practices to sell more journals, as opposed to solid science which is often quite boring and tends to find that new drugs or treatments do not work and that our theories are wrong. Ideally we need to end big publishing’s control over science. Their role should be simply production and distribution, but they currently copyright much of the material and force universities to pay twice, once for the researchers’ salaries and again so that researchers can read what they are publishing. And this often done with public money. Its really wrong and generates perverse incentives among researchers and profit-seeking publishers. 

— It seems teams didn’t falsify data or cherry-pick numbers in any obvious way; instead, ideology appeared to influence judgement calls — is that a reasonable explanation or is it too charitable?

It is a reasonable explanation, but we must be very cautious with it. Ideology might explain about 3% of the variance in researchers’ findings. The rest has to do with other factors or random noise. Consider that many scientists have specific methodological training. Through this training they simply do what they know how to do. And this can influence results. For example, someone who only knows how to use a hammer, will treat things as a hammering problem, when it might be better solved with a different tool. This takes us back to better, broader methodological training, which scientists would have more time for if they weren’t under constant pressure to publish. 

— Are there ways to guard against the ideological bias highlighted in the study? It seems that peer review can spot poorly defined studies — but does it work well enough in the real world? 

Peer review is a poor solution. There are studies out there that show that peer review is not reliable. Just like giving the same researchers the same data, if you give different peer reviewers the same paper, they will come to a huge range of judgements about the paper. The publishing system is in some ways a lottery. One solution for this problem is what we call “adversarial collaboration”. This is a type of research where scientists who disagree about a topic work together. They design a study together and agree on all methods and on all criteria with which to judge the outcomes. Then after this is all agreed in advance, the study is conducted, ideally from a third-party, and then the results speak for themselves.

— What do you think the message here is to the public and to policy makers?

Everything we have in society that works is based on science. Smartphones, that open heart surgery that saved the life of a loved one, airbags, planes that don’t crash. We need science for every decision we make collectively. But the public and policymakers can be highly politicized, and this can influence science, we see this even in the scientists in our study whose own politics seemingly played a role. To cut through this, we need policymakers that support science conducted by scholars who have different political ideologies, who do not agree. For example, George and I are somewhat different in our own assessment of the impact of immigration on society and the labor market. This made us a very strong team. It meant that we could focus on the scientific process and try to get to the best, most reliable answer. Crucially, we need to never rely on single studies or single science teams. Before we declare evidence of anything, we need many studies. We should not just rely on single papers or scholars. We need dozens or hundreds of studies on a topic conducted by inter-disciplinary and inter-ideological teams.

To respectfully and ethically decline to support Elsevier

Just over 5 years ago, I decided I would no longer peer review for or publish with Elsevier. In addition, I would try not to support any of their products. This was a choice enabling me to be more consistent with my own ethical values and support of open science.

This is was a difficult decision to make, because for-profit publishing companies have all engaged in questionable behaviors in pursuing economic gains. Also, it means saying ‘no’ to colleagues who are editors and could really use my expertise for a peer review or book chapter contribution. Therefore, I have created a brief letter template to decline such offers. It provides some informative links, encourages the recipient to pursue their own opinion and avoids overreaction via defamation or fear-mongering. Elsevier is known to attack with lawyers, so it is also important that I protect myself and my associates when declining such offers.

Dear [person],

I apologize but I cannot in good faith [write a peer review / book chapter] for an Elsevier publication.

I realize that publishers are ‘for profit’ enterprises, and balance the needs of science with their own pursuit of economic success; however, in my experience Elsevier has consistently proven to be an exceptionally unethical company with practices antithetical to the pursuit of knowledge. Some of the more awful practices have been to sponsor arms fairs[1], advocate intensively against open access[2], create seemingly scientific journals and then sell space in those journals for companies to print bogus articles that support their products[3], purchase and copyright or paywall scientific tools[4], charge astronomical fees for journal subscriptions[5], and over-harvest my personal information[6]. Therefore, if I can help it, I do not peer review or publish with Elsevier anymore.

Please do not take this personally. This letter expresses my own personal opinion based on my own experiences. My opinion and resulting preferences do not reflect that of my employer or any colleagues with whom I have collaborated. I encourage you and each person to investigate for profit publishers on their own and make their own informed decisions.

Sicerely,

[Author]

[1] https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(07)60488-7/fulltext

[2] https://en.wikipedia.org/wiki/Elsevier#Criticism_of_academic_practices, https://www.theguardian.com/science/political-science/2018/jun/29/elsevier-are-corrupting-open-science-in-europe, https://www.lub.lu.se/en/find/lubsearch/epublications/new-agreement-elsevier/sweden-stands-open-access-cancels-agreement-elsevier (attacks on ResearchGate and Sci-Hub) https://www.science.org/content/article/publishers-take-researchgate-court-alleging-massive-copyright-infringement, https://www.nature.com/articles/nature.2017.22196

[3] https://www.the-scientist.com/elsevier-published-6-fake-journals-44160, https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(09)61404-5/fulltext

[4] Social Science Research Network – purchased and turned into a place for paywalls; Mendeley – purchased and converted into a product that is not (easily) interoperable with open source products like Zotero; tries to patent a peer review tool to a payment step into the scientific process https://www.eff.org/deeplinks/2016/08/stupid-patent-month-elsevier-patents-online-peer-review.

[5] https://www.science.org/content/blog-post/california-tells-elsevier-take-hike, https://www.thenation.com/article/society/neuroimage-elsevier-editorial-board-journal-profit

[6] https://eiko-fried.com/welcome-to-hotel-elsevier-you-can-check-out-any-time-you-like-not/

Open science. Back to basics

Are we really sharing all steps in our research? And are we making clear which steps we cannot share and why? There are always hidden steps that influence our research design and reporting.

Ask yourself this question: If someone else were to pick up one of your studies, most likely a published paper reporting the results of one of your studies,  would they be able to understand and reproduce everything that you did?

Are there things that can’t be exactly reproduced, like privacy-protected data, or subjective, interpretive, or ethnographic experiences in your study? If so, have you pointed this out for the reader?

These sound like trivial questions.

Anyone outside of science would think these things must be true. They would think that of course a scientist, a researcher, would report everything they did in their research. Isn’t that their job? How else should science look? What else are scientists doing other than doing research and reporting that research?!

Maybe your answer to this question is, ‘yes, I have reported everything I did’. You truly believe that people could follow everything you’ve done, but is this really the case?

I work in the area of statistical analysis, mostly, and in this area, there is code required to run statistical analyses, and that code is the basis of my research design, assuming I didn’t collect the data myself.

So, I have secondary data, I already have data, and then I analyze it, and my code is the basis for that analysis. Well, it turns out there’s a lot of things in my code that others would not be aware of without the code. So, this makes code sharing absolutely essential. And we’ve come a long way in this area.

Economics and political science especially, in general, have strong code-sharing norms for people who do statistical analysis. Sociology is a little further behind. Probably one of the main reasons for this is that journals have not adopted strong policies of code sharing in sociology, unlike in political science and economics. Either way, as a researcher, I should be responsible for making my results reproducible, but let’s think of an even more hidden case.

In your statistical analysis, did you do things that are not in the code? Now, if any researcher were to answer ‘no’ to this question, I would not believe them, including myself. There are always things that we do, that we then take out of the code because they seem redundant or unnecessary or, in some cases, don’t make our study look as good as we want it to look. One of those things is to run models that we don’t report.

Now, that would be fine if we just happened to run a model on accident or had some other glitch or mistake. But a lot of the models we run are intentional, and the reason we don’t report them is that they produce results that we don’t want. They produce results that don’t support what we think is true or what we want to appear to be true because it’s sexy, because it’s something that would be publishable.

But the question to ask is then: have any of those extra models that I’ve run or you’ve run been used to inform decisions that we make in terms of what models to run next, how to recode any variables in those models, and what to report on, in the paper and in the code? And if the answer is yes, then that’s part of the research design, and should be reported.

Now, this is not the norm at all. I don’t know any discipline where this is the norm, and I don’t know people who are trying to make this the norm. This is a hidden area of the open science movement for the most part, although we have research now that overwhelmingly suggests that the findings that we’re reading about, that people are reporting on, are probably selected (thanks to p- and z-curve analysis). They’re probably a subset of findings that are not just a random subset of the models that researchers intentionally ran, but a selected subset in the sense that they all point in a certain direction or they all have a larger size or there’s something about them that makes them specially selective – and this selection process reduces the reliability of science.

Therefore, when I say back to basics, I actually mean the basics of research, not the basics of the open science movement. I mean reporting everything in the research process that influences the results. That would truly be open science.

Image credit: Nate Breznau’s own photo

Questionable research practices from the practitioners’ perspectives

With Monica Gonzales-Marquez, Priya Silverstein and Eike Mark Rinke.

In preparation for Metascience 2025, I put together a panel with three other researchers called “Questionable Research Practices from the Perspective of the Researcher: Understanding Perverse Incentives using Autoethnography”. We used our own discussions in online meetings and individually as recorded narratives as content for this panel. The impetus was that much metascience points accusatory fingers at problems in science, thus supporting a culture of fear. Researchers comply to avoid scrutiny, and not necessarily out of an intrinsic motivation to do good science.

When successful and meaningful science is measured entirely by publication in a ‘high impact’ journal and high citation counts, the drive to do science is transmuted into a laser-like focus on publishing. The intrinsic motivation to contribute to the creation of scientific knowledge becomes confounded with publishing, while simultaneously deprioritising the robust, pedantic, methodical, humble labor involved in doing good research. The scientific method gets confounded with the mechanics of publishing and achieving “standing” in the scientific community.

We hope that by revealing our own experiences in the world of ‘publish-or-perish’, and how it has pushed us towards questionable research practices (QRPs), we might generate intrinsic motivation in others to dispassionately examine their own scientific practices.. By looking back through our histories in academic work and sharing them, we also expect to increase our own intrinsic motivations to be ever vigilant, and to learn to always privilege scientific integrity over publishing.

Confounding: Publication = Science

A major theme in our narratives is that science has been fully confounded with publishing. For some of us, the pressure to publish, and the toxic atmosphere and interactions pushed us out of academia entirely. For others, we internalized the publishing norm, convincing ourselves that we were doing impactful science, when we were actually just doing impactful publishing – which does almost nothing to alter collective human knowledge, promote social justice or solve other societal problems.

Some excerpts from our narratives:

Questionable Behaviors: Hacking

The pragmatics of doing science in a publish or perish culture is that we either engage in some (often undeliberate) hacking and storytelling, or are sanctioned.

Some excerpts:

Ego, Power and Personal Struggles

As human beings we are prone to seeking status, material security, community acknowledgement and different things depending on where we are at in life, and who we are personality-wise. Especially when in graduate school, we are in a position of little power in comparison to professors and our supervisors. People who  wield power over others, and push their own ego-centric agendas tend to be quite successful in science. This can lead to abuses of power, ego trips and other toxic behaviors that can diminish junior researchers’ ability to push back against pressures to engage in QRPs and have strongly demotivating effects.

Intrinsic Motivation

We hope that by speaking out, and normalizing scrutiny of our own experiences of questionable research practices as something valuable, we can help others become intrinsically motivated to do the same. Moreover, we propose that ethnographic narrative may be an underexplored but powerful method to help uncover the causes and motivations of questionable research practices from researchers’ lived experience of “doing science”.

Part of the motivation for this panel is based on one of the authors’ (Nate Breznau’s) own authoethnographic research into my QRPs. He presented preliminary results at the Sociological Science Conference at Cornell (link to slides).

Readers can find the full poster for our presentation at Metascience 2025 at University College London here.

Our future plans are to seek a larger sample of researchers and invite them to a semi-structured narrative sharing process (via recording themselves) to further study QRPs. In particular we hope to target early career researchers to shed light on the current state of science training and supervision experiences. 

The ISA should not ban scientific associations for failing to take political stances.

On June 29th, 2025, just one week before its World Forum in Rabat, Morocco, the International Sociological Association (ISA) leadership banned the Israeli Sociological Society (ISS).

The reason is that the ISS did not take “a clear position condemning the dramatic situation in Gaza.” In other words, that they are not taking a position in direct opposition to and protest of their own government. Although there are reasons that some would want to take such a position, in particular the incredible humanitarian crisis in Gaza, it is not fair to punish a scientific organization for not taking this position. In other words, not fair to try and force a scientific organization to mix itself in with politics. I can think of three concrete reasons.

  1. Science itself should prioritize inquiry, not political action. We conduct research and develop theories, that is science. Sociology is the science of society and social interaction. Therefore, sociologists should (continue to) engage with the war and humanitarian crisis unfolding in Gaza. They should try to present facts to answer the questions: What is happening there? Who is affected? Are the actions of the Israeli military in violation of the Geneva Convention? Is this a genocide? To blackmail other scientists to force them to take political action or face consequences to their participation in science and scientific exchange, runs counter to prioritization of inquiry in science.
  2. Scientists taking political positions could be in danger. In particular in authoritarian, or martial law settings, scientists who stand up against their government put not only their scientific enterprise at risk, but also their lives and their families’ lives potentially at risk. The ISA is demanding that members of the ISS take actions that could lead to their Society facing funding cuts or being cancelled altogether. Even worse, the ISA is demanding that Israeli sociologists take a position counter to their government in the middle of one of a radically politically charged historical moments – it is asking them to potentially put their lives as risk.
  3. Where do we draw the lines of where to take political action? If a humanitarian crisis and what appears to some scholars as a genocide are conditions for taking political action on part of the ISA, where do we draw the line? Why does the ISA not ban all Myanmar sociologists from participating because of the situation of the Rohingya? Perhaps even more perplexing are the odd exceptions: The ISS is banned, but what if an Israeli sociologist is not a member, can they participate? Perhaps someone forgot to pay their membership dues this year but was a member before. Can they participate? What about the United States and Germany? These are the greatest allies of Israel, why are they not banned? What about Russia invading Ukraine? These questions are rhetorical, because this is a Pandora’s box that is infinite and quickly digresses into everyone getting banned.

In fairness, there is a humanitarian disaster going on in Gaza. It is tragic for all the innocent bystanders swept up in it. Those who have been kidnapped, killed, harmed by bombing. Its awful to watch. But I am not writing this in regards to any political movement or government. I am writing this to defend science, which I see as my responsibility.

Science in survival mode

Scientific research is unreliable. It comes with uncertainty. Whether launching a rocket or measuring racial prejudice, there is uncertainty. We use this uncertainty to make decisions. If the rocket has a 40% chance of exploding, best not to stick astronauts in it. If skin-tone bias of soccer referees is somewhere from none to a lot, it is irresponsible to conclude they are prejudiced.

Investigating uncertainty, is science. Truth and uncertainty are two sides of the same coin, they co-define each other. A problem for humans measuring uncertainty, is that humans are unreliable. Human scientists themselves add uncertainty to the measurement of uncertainty.

Recent studies suggest that somewhere between 25 and 60% of published statistical results cannot be recreated using the materials provided – these measures of uncertainty come with their own uncertainty. It turns out that scientific researchers are doing some really peculiar things to generate uncertainty. Some surveys suggest as many as 9% of scientists faked data at least once in their careers, and that more than half selectively reported findings – a behavior that makes the things they study to appear less uncertain than they actually are.

I believe the answer lies in their humanness. Like all animals, they are genetically programmed for survival. They are capable of both rational, reflective decision-making and split second reaction without any thought. Given time to reflect, a human would generally conclude murder is unethical, but simultaneously would not hesitate to kill if it prevented their child from being killed. Murder remains wrong, but to not kill and let a child die is also wrong. It would be irresponsible parenting failing to ensure survival.

Murder is a profound act when a human perceives themselves to be in a situation of life or death. It is a symptom of subconsciously activated defense mechanisms in survival mode. But there are many other symptoms. In order to avoid death, humans, like other primates can engage in deception, disassociation, aggression, manipulation, submission, scapegoating, theft and hoarding. If someone held me at gun point and told me to prove the earth was flat, I would have no problem doing it. The math would work, I would just need to fake a little data

Most scientists have highly valued knowledge and competencies and thus live in situations where their lives are not under threat. At least not as a result of their scientific practice. For the sake of this thought experiment, lets just rule out regularly occurring life-threatening danger as a cause of scientists exhibiting survival mode behaviors. This leaves the perception that they are under threat, as a possible explanation.

It takes only a few stimuli to induce survival mode behavioral changes in animals. Like hearing a frightful noise when seeing an animal. This animal and anything that looks like it become automatic sources of anxiety and fear even without the sound. Imagine being told over and over and over that you have to have an exciting study with powerful results in order to get an academic job after graduate school, otherwise you wasted 3-8 years of your life and probably a large chunk of capital on getting a PhD. Could this alone, without introducing any actual shocks or physical pain, induce Pavlovian fear? I encourage you to go ask any graduate student to answer this that does not yet have such a study.

Now imagine that during graduate school a student invests all their time and resources into an experiment. After it is complete, they hypothesized result, the one sure to be exciting and publishable, is not there. Imagine the horror, the shame, the feeling of failure, the panic. Remember the poor soul who leaped to his death because he misunderstood futures trading and thought he owed three-quarters of a million dollars? That was a triggered survival response. Because dying felt like the only way to ‘survive’ the horror of facing that debt. Imagine if he could have just changed the futures market by adding his own numbers to the market. Would he have done it?

We should not be surprised at all then, when scientists acting out of fear-based survival strategies, fake data. Diederik Stapel faked an entire career of data before being caught. His behavior was self-described as an “addiction”. A common reaction of individuals placed under fear stimuli that are emotionally damaging if not traumatic. The intense pressures and expectations of the academic environment created a context in which he felt compelled to engage in unethical practices to maintain his status and success. It does not make murder or data-faking right, but to not take this as grounds for indicting the scientific rewards system is certainly wrong.

The incentive structures have to change before we can honestly expect the fear-driven pressure to fake, cheat, lie or steal – in order to avoid the experience of loss associated with the common null results that occur when conducting high quality scientific research on radically complex human brains and societies that are frustratingly difficult to measure things in – to go away.

Open Science? What’s That?

Was the answer I got from Andrew Abbott at a Hogwart’s dinner I was fortunate enough to attend a few years back when I asked him after dinner, “what are your thoughts on open science?”

That’s right. Open science. Say something paradigm changing.

We have so many results and coherent arguments about the status of science. To make an impact here requires synthesizing them into something new. Not purely constructively new in nature, but also subtly undermining all that came before. Something was missing all along and this new perspective is it. That’s the problem.

Our formula for declaring theory and opinion in scientific writing is that it must be big. Paradigm impacting. Makes you think. Expresses the scientist author’s authoritative intellect and ingenious perspective. With seemingly effortless written words, the scientist author lays down an argument so convincing, that it must be true. It was true all along. It is so obvious, why didn’t I see it? Thinks the average reader.

If Diederik Stapel was never caught, he would be a model for success. THE model for success in social psych. You don’t know what the formula for success is? Ask an economist. You have 3, maybe 5 journals in which you need to get published very early on. You have to have a job market paper that is known. You have to have connections. Otherwise you are on the sidelines, building an exit strategy. This is roughly true in the other social and behavioral science disciplines, yet we in sociology or political science will mostly deny it. We are better than those status maniacs.

But it is publish-or-perish. And this is the model.

To do publish, requires word wizardy. The ability to take any findings and craft an argument about them that makes them seem relevant is more valuable than research method skills. To make any findings seem essential to our fundamental understandings of somethings that are important. That is how we do science. Framing.

Andrew Abbott has the goal of out thinking, reading, writing and arguing everyone he encounters. Ask him if you ever meet him. He will not deny this. Easily considered by most to be one of the greatest living sociologists. He was editor of AJS and published book reviews of random, late-night selections in the Oxford library under a pen name, because it was fun. Because that is what legends do, and then tell about. He once wrote that sharing all material to be reproducible and replicable would make science boring.

As long as this is the model for success, we are not making real progress.

I sat at a dinner once with Ronald Inglehart. Every student, postdoc and professor there gawked with wide eyes. What was so appealing and enthralling. Why was this scientist a rock star for us? Was it his findings? Or was it his status? I think status. I think that is what scientists are seeking. I observe it. I feel it. If you have to have certain publications in order to not just pursue a scientific career, but to be at the ‘forefront’ of a subfield, then does it really matter what is in those publications? As long as you get to claim that status, that is it. You win. Why do athletes take steroids, cheat? Is it because of their great passion for the sport? Or is it to win status?

I entered a Master’s program in order to learn about the problems associated with a topic that was pressing for social justice and democracy. I still want to do this. But status was dangled in front of my eyes. ‘Publications are the currency of our trade’. The professor who told me this was absolutely correct. Now I want to do open science, but I have to perform wizardy with words to be recognized as central to the open science movement. This makes me uneasy to the point of nausea.

How can science become trustworthy and effective, if we have to be witty, cutthroat and/or famous in order to take part in it? How can we overcome this delusion? When will we stop and say enough is enough?

This type of honesty is what John Levi Martin specifically advised not to put out on the intrawebs because it will come back to bite me someday. Image (presentation) is everything. Warhol or Abbott, they figured that out long ago. I guess then for my own peace of mind, I extend my hand. Bite away you biters.

Human Induced Climate Change

Show me the data


Nate Breznau

A quick guide to human impact on the environment

Carbon Emissions

The amount of Carbon we’ve released into the atmosphere is nearly perfectly collinear with the Industrial Revolution and subsequent industrial growth. This is easily estimated from ice samples.

Carbon is unquestionably a result of human activity

We can measure random and cyclical carbon in the atmosphere from ice core samples. Never in 800,000 years have we seen carbon levels like those since the Industrial Revolution.

source

Climate Warming & Emissions

A striking correlation between emissions and global average temperatures. Although temperatures fluctuate regularly, their fluctuations track upward following human-based carbon emissions. 

source

Microplastics

These are only produced by human production. Note that there are several parts of the ocean where we find more than 10 pieces per cubic meter (dark purple circles and dark red stars). There are over 1.3 billion cubic meters of surface water in the ocean. 

source

Polar Ice Degradation

We do not have accurate data from long ago, but since 1990 the arctic ice shelves have declined in size dramatically. This also leads to a rise in sea levels. Here a comparison of ice in summer and winter in 1990 compared to 2020.

source

Scientific Consensus as of 2019

A review of all published climate science studies in peer reviewed journal articles between 2005 and 2015 suggested a consensus of somewhere between 83 and 97% of climate scientists that humans were causing the major changes we observe in our environment in the last 100 years1

A new study of all of the just over 11 thousand articles across all disciplines and journals published in 2019 revealed that this consensus was 100% (within a natural margin of error)2

Bibliography

1. Cook, J. et al. Consensus on consensus: a synthesis of consensus estimates on human-caused global warming. Environ. Res. Lett. 11, 048002 (2016).

2. Powell, J. Scientists Reach 100% Consensus on Anthropogenic Global Warming. Bulletin of Science, Technology & Society 37, 183–184 (2017).

A fake journal, algorithmic plagiarism and tricking Google Scholar

A case of a fake journal that passed Google Scholars’ bots and algorithm checking to become indexed. The journal contained articles that were completely plagiarized but assigned new titles, abstracts and authors.

See updates at end of blog for table of plagiarized papers, and a small win for open science after the domain host revoked the content of the plagiarized paper’s URLs.

A colleague of mine was searching Google Scholar using certain words and the number two hit appeared with the following title and direct download link.

The PDF looks like a journal article. The journal has an ISBN instead of DOIs, but this is not unheard of. It even has an impact factor listed: 7.11, it must be good.

It turns out that this seemingly interesting paper with an alleged author who is an expert in this area, is plagiarized.

How do I know it is plagiarized? Just ask the real authors of the paper which appeared with a different title and abstract in Social Science Research one year earlier. Oh and with an identical word-for-word text, with some mistakes relating to equations – possibly a failed scraping algorithm.

The journal title is Journal of Current Research and Review, but it is extremely difficult to find through an internet search. This requires endless scrolling (as of 10.08.2023), or removing the ending on the download URL to arrive at the index page.

Scrolling down on the home page reveals the “Latest Articles”. Here I find the same title that appeared in Google Scholar. However there is no download link, only the title and abstract. In other words, the URL that Google Scholar found to download the paper is not accessible through any link on the journal’s website.

How is it possible that Google Scholar would pick up and index such garbage?

I asked ChatGPT to tell me how Google Scholar works, and the results suggest that someone or something had advanced knowledge and programming skills that would provide the right meta-data and materials on the journal webpage to pass the automated Google Scholar bots which are constantly scraping the web for academic papers and publications.

From ChatGPT prompted with “How does a journal get indexed in Google Scholar?” (09.08.2023). I received the following answer. Note that I removed the descriptions of each point, as they are mostly self-evident.

Google Scholar indexes scholarly literature, including journal articles, conference papers, theses, and other academic resources. In order for a journal to get indexed in Google Scholar, there are several steps and criteria to consider:

  1. Quality Content
  2. Digital Format
  3. Open Access and Accessibility
  4. Consistent Publishing Schedule
  5. Meta Tags and Structured Data
  6. Proper Article Metadata
  7. Citations and References
  8. Author Profiles

It’s important to note that Google Scholar’s indexing process is automated, and there is no formal application process for journals to get indexed. Google Scholar’s algorithms discover and index content based on various factors. However, journals can follow best practices to increase their chances of being indexed and improving their visibility within Google Scholar’s search results.

The journal website and meta-data had to pass Google Scholar’s algorithm. This is no lucky feat. Although Google does not reveal any usage of AI to evaluate journals for indexing, it clearly has an advanced system of scraping, filtering and crawling. Who or whatever designed this journal, or really this website that looks like a journal, understood how to pass the test. For example, the “Latest Articles” appear to be from “Volume 14” suggesting to potential crawling bots that the journal may have been published for 14 years now.

This is of course false. There are no other articles readily accessible from the journal’s website other than the two that appear. The other article is also plagiarized, in case you were wondering. But, a closer look at the URL reveals that the article I am discussing has the number “800” at the end of its URL.

https://zapjournals.com/Journals/index.php/jcrr/article/view/800

Changing this number yields other papers. At least as far back as number 795; prior to that yields 404 errors. Moreover, I cannot find the papers from 795-798, only their titles and abstracts.

The journal website’s main page yields something even stranger. It looks like what a computer science student might create as an example page to try and sell as a template or to demonstrate the services on offer for a website construction gig. The content has absolutely nothing to do with journals, academia or publishing.

It even comes with fake testimonials. I wonder whose pictures were stolen for this…

The real question for me is, ‘what to do about this?’. Clearly this is fraud and constitutes both ethical and legal infringements on science. Looking up the host of the domain of the journal using ICANN reveals that it is provided by a Lithuanian company Hostinger.

Looking at this company’s website suggests they are legitimate. The provide contact information to report abuse. So that is what I did by sending them the info in this blog post.

Hostinger replied within two days and told me they take abuse of their services very seriously and asked me for more evidence so they could pursue the case. I gave them the two links for the plagiarized papers and their original versions published one year earlier in Social Science Research.

https://zapjournals.com/Journals/index.php/jcrr/article/download/799/1237

https://zapjournals.com/Journals/index.php/jcrr/article/download/800/1238

Word for word plagiarized from:

https://www.sciencedirect.com/science/article/pii/S0049089X22000321

https://www.sciencedirect.com/science/article/pii/S0049089X22000333

For all the negative publicity surrounding Elsevier, they are still a player in the academic world. In 2020 they owned roughly 16% of the academic publishing market. If there is anyone who would want to prevent abuse of their services, including plagiarism of their work, it is Elsevier. They have a trove of lawyers fighting and winning battles to protect their content across the world. As the two plagiarized papers that I can download are from the Elsevier journal Social Science Research, it made sense to contact them as well. Elsevier’s due process suggests contacting the journal editor first. Thus, I have sent this information to the lead editor.

The journal also lists an ISBN number. But this number is a fake, and returns an invalid search with the ISBN lookup tool.

Google certainly would not want its products indexing fake journals with plagiarized papers, so I took the liberty of contacting them as well.

What is striking about the journal is that they have a long editorial board list. Internet searching reveals that these are real scholars. I also contacted them. I will continue to report on this case as it unfolds.

The question for me is: ‘what motivated someone to create this site?’. There is clearly no profit associated with it. There is also no status gain, because the plagiarized papers have authors assigned to them who are not the original authors and clearly not players supporting the fake journal. They are highly established scholars in their respective sub-fields. These fakely assigned authors are perfect examples of what an AI might choose to assign to a certain topic. I tested this by asking Chat GPT if Seamus McGuinness could have written the abstract. The response points at AI as a source for potentially re-writing abstracts of the plagiarized papers and for finding suitable authors to assign to them. ChatGPT said:

Yes, Seamus McGuinness could be a potential author who might have written this abstract. Seamus McGuinness is known for his research on labor market issues, including education-job mismatches, gender disparities, and remote work. His expertise aligns with the themes discussed in the abstract, making him a plausible candidate as one of the authors who could have written it.

It remains a mystery for now what is up with this website and the fake journal. Was it a computer science project that accidentally got picked up by Scholar, one that was never intended for public consumption? Was it an attempt to create a journal, but try and hide the real content of the journal until it could pick up real submissions? Was it entirely AI generated, to showcase the power of an AI?

[Update 14/08/2023]

Thanks to Random Cat on Twitter, I learned that there are more papers than just two. Using Google Scholar they searched by journal.

This allowed me to compile a table of plagiarized papers.

Table of Plagiarized Papers in JCRR, found via Google Scholar

Original TitleOriginal Author(s)Original JournalJCRR TitleJCRR AuthorsLink to Original ArticleLink to Plagiarized Article
The motherhood wage gap and trade-offs between family and work: A test of compensating wage differentialsNick Wuestenenk & Katia BegallSocial Science ResearchCOMPENSATING WAGE DIFFERENTIALS AND THE MOTHERHOOD WAGE GAP: A COMPARATIVE ANALYSISHK Kleven & CL Landaislinklink
Conflicting signals: Exploring the socioeconomic implications of gender discordant namesAndrew Francis-Tan & Aliya SapersteinSocial Science ResearchBREAKING BOUNDARIES: EXAMINING THE INTERSECTION OF GENDER DISCORDANT NAMES AND SOCIOECONOMIC ATTAINMENTAL Roberts & M Rosariolinklink
Gender overeducation gap in the digital age: Can spatial flexibility through working from home close the gap?Ana Santiago-Vela & Alexandra MergenerSocial Science ResearchBRIDGING THE GENDER OVEREDUCATION GAP: EXPLORING THE ROLE OF WORKING FROM HOME IN THE DIGITAL ERASMG McGuinnesslinklink
How the Great Recession changed class inequality: Evidence from 23 European countriesJad MoawadSocial Science ResearchCLASS INEQUALITIES IN THE WAKE OF THE GREAT RECESSION: A STUDY OF 23 EUROPEAN COUNTRIESPD Allisonlinklink
Higher education and high-wage gender inequalityNatasha Quadlin, Tom VanHeuvelen & Caitlin E. AhearnSocial Science ResearchASSESSING THE CONTRIBUTION OF EDUCATION TO GENDER WAGE DISPARITIES IN HIGH-EARNING PROFESSIONSJ Jacobslinklink
TRANSFORMING RESEARCH: EXPLORING THE INTERPLAY OF DATA MINING, MACHINE LEARNING, AND KNOWLEDGE DISCOVERRobert M. Bond & Christopher J. FarissSocial Science ResearchsKnowledge Discovery: Methods from data mining and machine learning [b]Xiaoling Shu & Yiwan Yelinklink
SOCIAL MEDIA ADDICTION: PROPOSED INDICATORS AND STAGESZakaria I. Saleh & Omar Zakaria SalehInternational Journal in Commerce, IT and Social SciencesUNPACKING THE CYCLE OF SOCIAL MEDIA ADDICTION: UNDERSTANDING SYMPTOMS, PROGRESSION, AND RECOVERY [a]OZ Salehlinklink

[a] Published twice in JCRR with different authors

[b] Not indexed in Google Scholar

All are from Social Science Research except one. There are also other papers where I cannot find an original. These papers might actually be original papers, or stolen working papers that are not easy to find online.

Interestingly, some papers that appear that they may not have been plagiarized have a different display in PDF form. This includes contact details for the journal. I thus emailed them as well to inform them that they are committing ethical and legal fraud.

[Update 23.08.2023]

Hostinger investigated the reported problem and determined that the user of their domain had violated the terms of ethical/legal usage and removed the plagiarized papers. However, the journal website still appears. In other words, a journal that obviously intentionally plagiarized several articles, possibly to boost its reputation and encourage others to submit and thus pay the 45 dollar fee, is still out there lurking. Moreover, this journal is part of a larger company called Zenodo Publishing. On their main page the list dozens of journals. If I had to guess I would assume these journals are also predatory, and may contain plagiarized content. Only investigating this will prove if this is true.

More to come.

ChatGPT teaches open science

I asked ChatGPT to talk about open science. This post is the unabridged conversation. ChatGPT is a trained algorithm using machine learning. It is what most would refer to as Artificial Intelligence. And intelligent it apparently is.

~

Nate: I am teaching a course on open science. Could you summarize the key arguments of the open science movement?

ChatGPT: Certainly! The open science movement aims to promote greater transparency, collaboration, and accessibility in scientific research. Here are some key arguments and principles of the movement:

  1. Open access to scientific publications: Advocates for open science argue that research findings should be made freely available to everyone, rather than locked behind paywalls or subscription fees. This allows for wider dissemination of knowledge, which can benefit both researchers and the general public.
  2. Data sharing: Open science also emphasizes the importance of sharing data and methods, allowing others to reproduce and build upon research findings. This can improve the quality of research and promote greater collaboration across scientific disciplines.
  3. Reproducibility and transparency: Open science advocates for greater transparency in research methods and data analysis, as well as reproducibility of findings. This can help to prevent fraud and errors in scientific research, and can promote greater trust in scientific results.
  4. Collaborative and participatory science: Open science encourages collaboration among researchers, as well as engagement with stakeholders and the public. This can help to ensure that research is more relevant and useful to society, and can promote greater accountability and transparency in the research process.
  5. Open source and open tools: Open science also advocates for the use of open-source software and tools, which can be freely accessed and modified by others. This can improve the reproducibility and efficiency of scientific research, and can promote greater collaboration and innovation.

Overall, the open science movement seeks to transform scientific research into a more transparent, accessible, and collaborative endeavor, with the ultimate goal of advancing knowledge and benefiting society.

Nate: As I said above, I am teaching a course on open science. It is targeted at university students. Could you recommend some introductory readings on this topic?

ChatGPT: Certainly! Here are a few introductory readings on open science that you may find helpful for your course:

  1. “The Open Science Training Handbook” by the Open Science MOOC: This is a comprehensive guide to open science, covering topics such as open research data, open access publishing, and reproducible research. The guide includes both theoretical and practical advice, as well as case studies and examples.
  2. “Ten Simple Rules for Reproducible Computational Research” by Sandve et al. (2013): This article provides a set of guidelines for ensuring that research is reproducible, including recommendations for version control, documentation, and sharing data and code.
  3. “Open Science: A Very Short Introduction” by Ana Brandusescu (2020): This short book provides an accessible introduction to open science, covering the history and principles of the movement, as well as key topics such as open access publishing and citizen science.
  4. “Open Science Framework (OSF): A Tutorial” by the Center for Open Science: This tutorial provides an overview of the Open Science Framework (OSF), a free and open platform for managing research projects and sharing data and code. The tutorial covers the basics of setting up a project on the OSF, as well as using the platform to collaborate with others and share research outputs.
  5. “The FAIR Guiding Principles for scientific data management and stewardship” by Wilkinson et al. (2016): This article presents a set of guidelines for ensuring that research data is Findable, Accessible, Interoperable, and Reusable (FAIR). The principles provide a framework for managing research data in a way that promotes openness and reusability.

These are just a few examples of the many resources available on open science. Depending on the focus of your course, you may also want to look for readings on specific topics such as open access publishing, preprints, or data sharing.

Nate: What are some actions students can take during their studies that will make their research more open? Open, as in more aligned with the goals of open science.

ChatGPT: There are several actions that students can take during their studies to make their research more open and aligned with the goals of open science. Here are a few suggestions:

  1. Share research data: One of the key principles of open science is making research data openly available. Students can make their research more open by sharing their data in a public repository, such as the Open Science Framework or Zenodo, and including a data sharing statement in their publications. Students can also share their data with collaborators and other researchers upon request.
  2. Use open source software: Students can use open source software, such as R or Python, for their data analysis and statistical modeling. Open source software is freely available, and its code can be accessed, modified, and shared by anyone, which promotes transparency and reproducibility in research.
  3. Practice pre-registration: Pre-registration involves submitting a detailed description of the research design, methods, and analysis plan prior to conducting the study. This can help prevent bias and increase transparency in research. Students can pre-register their research on platforms such as the Open Science Framework or AsPredicted.
  4. Use open access publishing: Open access publishing makes research articles freely available to anyone, which can increase the visibility and impact of the research. Students can publish their research in open access journals or deposit their articles in open access repositories, such as arXiv or PubMed Central.
  5. Collaborate openly: Collaboration is an important aspect of open science, and students can promote openness in their collaborations by sharing their research questions, data, and analysis plans with their collaborators. Students can also use collaborative platforms, such as the Open Science Framework or GitHub, to share and manage their research projects.

These are just a few examples of actions that students can take to make their research more open and aligned with the goals of open science. By promoting transparency, reproducibility, and openness in their research practices, students can help advance the principles of open science and contribute to a more open and collaborative research culture.

Nate: Are there arguments against practicing open science? If yes, where do these arguments come from?

ChatGPT: Yes, there are arguments against practicing open science, although they are often debated and may not be universally accepted. Some of the main arguments against open science are:

  1. Lack of incentives: Critics argue that there are few incentives for researchers to engage in open science, particularly when it comes to sharing data or making research available in open access. For example, researchers may be concerned about the impact of open sharing on their ability to publish in high-impact journals, obtain grant funding, or advance their careers. These concerns may be particularly acute for early-career researchers, who may be under pressure to produce high-profile publications and may not have the resources or institutional support to engage in open science practices.
  2. Risk of misuse: Some researchers argue that open science practices, particularly around sharing data, may increase the risk of misuse of research results. For example, in some fields, sharing data could lead to concerns about confidentiality, privacy, or intellectual property. Researchers may be concerned about how their data will be used, particularly if it could be used for purposes they did not intend, or if it could be used to generate findings that are inconsistent with the original research.
  3. Cost and logistics: Practicing open science requires resources, particularly around data management, sharing, and archiving. These costs can be particularly high for large or complex datasets, which may require specialized tools or infrastructure. Researchers may also face logistical challenges around sharing and storing data, particularly if they are working with sensitive or proprietary data.
  4. Misinterpretation or misrepresentation: Some researchers argue that open science practices may increase the risk of misinterpretation or misrepresentation of research results. For example, if research results are shared before they have been properly vetted, this could lead to the spread of inaccurate or incomplete findings. Researchers may also be concerned about the potential for misrepresentation of their work, particularly if their research is taken out of context or used for purposes they did not intend.

These arguments against open science are often debated, and there are counter-arguments and efforts to address some of the concerns raised. However, it is important for researchers to be aware of the potential challenges and risks associated with open science practices, and to carefully consider how to balance the benefits and costs of openness in their own research.

Nate: Can you recommend any readings that are highly critical of open science?

ChatGPT: While there are some critiques and debates around various aspects of open science, it is important to note that these are often specific to certain aspects of open science and not necessarily a wholesale rejection of the movement. However, here are a few examples of critical perspectives on certain aspects of open science:

  1. “Against Transparency” by Carl Elliott: In this article published in The Chronicle of Higher Education in 2014, Elliott argues that the push for transparency in research can lead to a culture of surveillance and mistrust. He suggests that the emphasis on open access to data and methods may create an environment where researchers feel they are being constantly watched and evaluated, which can ultimately have negative effects on scientific creativity and collaboration.
  2. “Why the ‘Open’ Philosophy Can Harm Science” by Thomas Basbøll: In this blog post published in 2018, Basbøll argues that the open science movement can create unrealistic expectations for research, particularly when it comes to the idea of “openness” as an unqualified good. He suggests that the focus on openness can lead to a neglect of other important values in science, such as quality, relevance, and creativity.
  3. “The Five Stages of Open Science Grief” by Cameron Neylon: In this blog post published in 2014, Neylon describes his experience of grappling with the challenges and limitations of open science, particularly around issues of data sharing and reuse. He suggests that the open science movement can sometimes oversimplify complex issues and create unrealistic expectations for what openness can achieve.

It is important to note that these are just a few examples of critical perspectives on open science, and that there are also many scholars and researchers who support the principles and practices of open science. As with any complex and rapidly evolving movement, there are likely to be ongoing debates and discussions about the benefits and challenges of open science, and it is important for researchers to engage critically with these issues and develop their own perspectives.

Nate: Do you think AI can improve the implementation of open science practices? If yes, how?

ChatGPT: Yes, I believe that AI (artificial intelligence) can potentially improve the implementation of open science practices in a number of ways. Here are a few examples:

  1. Data management and sharing: AI can be used to help manage and share research data in a more efficient and effective way. For example, AI tools can help automate data cleaning and formatting, making it easier for researchers to share data in a usable format. AI can also be used to help identify sensitive or confidential data that may need to be redacted or protected before sharing.
  2. Open access: AI can help increase access to research by providing tools for automatic translation and summarization. For example, an AI tool could automatically translate an article written in one language into another language, making it accessible to a broader audience. AI can also be used to summarize articles or research papers, making it easier for readers to quickly understand the key findings and implications.
  3. Peer review: AI can be used to help improve the efficiency and quality of peer review. For example, AI tools can be used to automatically screen articles for plagiarism or identify potential conflicts of interest among reviewers. AI can also be used to help identify potential biases in the peer review process, such as reviewer bias or gender bias.
  4. Reproducibility: AI can be used to help increase the reproducibility of research by automating some aspects of the research process. For example, AI can be used to automatically generate code or scripts, making it easier for other researchers to replicate the analysis. AI can also be used to identify potential errors or inconsistencies in the data or analysis, improving the quality and reliability of the research.

Of course, it is important to note that AI is not a panacea for all the challenges of open science, and there are still many issues that require careful consideration and human input. However, AI can potentially help make open science practices more efficient, effective, and accessible, and there is a growing interest in exploring the use of AI in this context.

Nate: Do you think AI poses a threat to science, in particular principles of open science?

ChatGPT: While there are concerns about the potential risks and challenges associated with the use of AI in science, I do not believe that AI itself poses a direct threat to the principles of open science. Rather, it is the way in which AI is developed, deployed, and used that can potentially pose challenges or risks to open science practices.

One concern is that AI can potentially exacerbate existing inequalities in scientific research. For example, if AI tools are developed and trained on biased data or by biased algorithms, they may perpetuate or amplify existing biases in research. This could have a negative impact on the principle of equity, which is a core component of open science.

Another concern is that AI could be used to automate aspects of research that require human judgment and interpretation, potentially leading to a loss of creativity, intuition, and critical thinking in science. This could impact the principle of transparency, as researchers may rely on automated tools rather than engaging in open and collaborative inquiry.

Additionally, there are concerns about the potential misuse of AI in science, such as the use of AI to generate fake research or manipulate data. This could undermine the principle of integrity, which is a core component of open science.

Overall, I believe that AI has the potential to both enhance and challenge the principles of open science, depending on how it is developed and used. It is important for researchers and stakeholders to engage in critical discussions and careful evaluation of the ethical, social, and legal implications of using AI in science to ensure that the principles of open science are upheld and advanced.

Nate: Speaking of reproducibility. If I asked you the same questions in the future, would you give identical answers?

ChatGPT: As an AI language model, I would provide the same response to the same question every time, assuming that the context and information available to me are the same. However, it is important to note that my responses are generated by machine learning algorithms and are based on patterns and relationships learned from large datasets. While I strive to provide accurate and reliable information, my responses may not always be perfect, and they may be subject to limitations or biases in the underlying data or algorithms. Additionally, if the context or information available to me changes, my responses may also change accordingly.

Nate: Are you processing new information? In other words, still learning and evolving?

ChatGPT: Yes, as an AI language model, I am designed to learn and evolve over time as I process more information and receive updates from my creators. My responses are based on the patterns and relationships learned from large datasets and are continually refined and updated as new data becomes available. This means that my responses can evolve and change over time as I learn from new sources and refine my understanding of different topics.

Nate: I am writing a blog post in my blog ‘Crowdid’ by me, Nate Breznau. The blog is an academic blog hosted by Hypotheses.de. Would it be ok if I posted our conversation in my blog?

ChatGPT: As an AI language model, I do not have any personal or legal rights over our conversation. However, I should note that the content of our conversation is generated by machine learning algorithms and is provided for informational purposes only. While I strive to provide accurate and reliable information, my responses are based on patterns and relationships learned from large datasets and may not always be perfect or fully accurate. It is ultimately up to you to decide whether you want to share our conversation in your blog, but I would recommend that you provide appropriate context and citations to help ensure that your readers have a clear understanding of the content and its source.

Nate: Thank you, this has been a very enlightening conversation. Could you provide me with some more information about you, ChatGPT, and how my blog readers could understand how you learned to become an AI?

ChatGPT: I’m glad to hear that you found our conversation enlightening! As an AI language model, I was created by a team of researchers at OpenAI, a research organization focused on developing advanced AI technologies for the betterment of society. My development was guided by a large corpus of text data, which was used to train deep neural network models to generate human-like responses to natural language input.

The training process involved exposing the models to large amounts of text data and using sophisticated algorithms to learn patterns and relationships in the data. Over time, the models became better at understanding natural language and generating coherent, human-like responses to text input.

It’s worth noting that my development was a collaborative effort that involved a diverse team of researchers, engineers, and linguists, as well as large amounts of computing power and resources. While I can generate responses to a wide range of questions and topics, my responses are generated based on patterns learned from the data and do not represent a true understanding or consciousness of the topics being discussed.

If you would like to provide your readers with more information about how I was developed and how I work, I would recommend checking out the OpenAI website, which has a wealth of resources and information on AI and machine learning.

Meta-reproducibility crisis: Software edition

Its not just a ‘replication crisis‘, its a reproducibility crisis. Forget about falsification if one can’t even reproduce the workflow of another. Efforts are underway to improve transparency which allows reproduction, good. But what if even with the original data, researchers cannot reproduce numerical results simply because of where they are in time and space? What if different statistical programming languages or different packages and versions of these languages actually lead to different findings? Based on results of the Crowdsourced Replication Initiative we demonstrated that even using the same data and models, only 80% of independent replicator teams could reproduce the numerical results, even when they had access to the original Stata code.

Based on my experiences in this area, I cannot but help have a strong hunch that software or package versions threaten the computational reproducibility of our research. The fact that Stata rounds up at .5 and R rounds down already suggests that the same models might produce different results across software simply from rounding variations. Recently, I found another reason to believe this hunch.

I am a participant in SCORE, a massive collaboration to systematically investigate the reproducibility of social science research – led by the Center for Open Science. I agreed to attempt a computational reproduction of the study “Age, Inequality, and Reactions to Marketization in Post-Communist Central and Eastern Europe” by Horvat and Evans (2011).

Despite some potential differences in the data I received from the original authors and those reported in their tables, I was able to produce similar results. Similar but not exactly close. The particular coefficient I was interested in came out at -0.39 in my ordered probit (cumulative link) model, whereas their original was -0.22. As a logistic coefficient or (translated into a percentage probability), this is potentially a large difference. I used R, but was not satisfied with my results. The original authors most likely used Stata, as their public data were a .dta file – Stata’s native file storage format. Also, lets be honest, no social scientist was using R before 2010 when they likely did their analyses. My curiosity and suspicion led me to run the model in Stata. Here I got a -0.21 coefficient. Almost identical to their original study, and not surprising given that there was minor variation in case numbers they reported and in the public data file. I was still not convinced that these coefficients were actually different because the models were quite complex. So I plotted the predicted probabilities to be sure that this was not statistical artifact of some model components (intercept cut-points for example).

But my hunch was supported. These same models run in R and Stata, led to different predicted probabilities. The figure below shows that among many former Communist Eastern and Central-Eastern European countries those who are aged 60+ were less optimistic about their living standards over the next 5 years. One of their critical findings was that this gap increased form 1993 to 2007, in particular because the 60+ group became even less optimistic; although in fairness this is difficult to conclude because of all the other variables in the model. What is clear, is that they were 0.45 points apart in 1993 and 0.65 apart in 2007 (see Table 5 in their original study). When I plot the predicted values I find that these results are similar but not identical in terms of the 1993 to 2007 change, but that the probabilities are quite different across software.

I used Stata v15 and the ‘oprobit package, and I used R’s ‘MASS’ package and the ‘polr‘ function. Despite identical data (and case numbers!) the ‘polr’ routine predicts the probability of respondents answering that their standard of living will fall or fall a great deal as much higher than that of Stata. Although the relative change between age groups is similar – with a slightly steeper negative slope in R – the lines are pretty far apart, and even further apart for the age group 60+. Without unpacking each package and the exact estimation strategy taking place therein, I cannot as of yet say why. My own statistical and software abilities are by no means exceptional, but certainly above the average social scientist. Thus, it would be totally unrealistic to expect any social scientist to understand the entire routine taking place within the polr or oprobit packages. If they did, they wouldn’t need the packages and could just write their own cumulative link routine!

The implications are that using different software leads to a lack of reproducibility. As if we did not have enough to worry about in the reproducibility area already.

Sci-hub. Good for science, otherwise mostly harmless

Sci-hub is the piratebay of academic journal articles. Its service is mostly illegal because it collects paywalled articles and makes them publicly available online via an indexed search. This is copyright infringement. People love it. The coverage is incredible, many journals have over 98% of their articles covered.

Frustrated with a lack of access to scientific articles, Alexandra Elbakyan of Kazakhstan founded the Sci-hub repository as a 20-year-old graduate student. Her site subsequently provided more open access to scientific knowledge than anyone in the history of science. She was named a person of the year in 2016 by Nature; yes, that Nature of the mega-profit-publisher Springer Nature who promotes open access by charging a 10 grand APC.

Reminiscent of the RAA’s takedown of Napster in 2001, Elsevier took legal action against Sci-hub in 2015 starting in the U.S. and quickly moving to other countries. This international campaign has to do with copyright law being organized by country, making it very difficult to pursue Sci-hub which exists in a cyberspace of mirrors, and it provides something that is unquestionably in the public interest and a basic UN human right.

Although there are allegations of security breaches that could lead to identity theft or other hacking university servers, I am not aware of a single piece of evidence Sci-hub has done anything other than ‘steal’ academic publications. It is not a threat to sovereign nation states, it doesn’t encourage sociopathic behavior. It is a form of rebellion against the plague that for-profit publishing unleashed on science, and a way to promote open science. Of course Elsevier was not wrong in its legal claim of copyright infringement. Elsevier wants researchers to pay for their articles and its minions see Sci-hub as causing profit losses. But the evidence suggests this is nonsense. Elsevier is wasting its time, precious time that it could use to sponsor arms fairs, create journals and sell them to big pharma or try to patent online peer review and force journals to pay to use it.

First, lets look at who is downloading Sci-hub content. Figure 1 shows the top ten countries by total downloads in February of 2022, compared with their total populations. We can see that relative to their populations, the U.S. and France are home to the most downloads per capita as of the most recent data.

Figure 1. Sci-hub downloads by country, February, 2022.
Image adjusted from Owens (2022) with
addition of Wikipedia population data in millions.

Next, lets think carefully about how publishing subscriptions work through two typical scenarios of a researcher downloading from Sci-hub.

Scenario 1. A researcher in the Global South downloads articles from Sci-hub. We know this is good for science. In fact this is science, it is active dissemination of useful knowledge. Merton would be pleased. The first question is easy: Is this researcher getting something for free that they should be paying for? Yes, they are receiving illegal good, getting copyrighted material for free. The second question is also easy: Would this researcher pay for this article if Sci-hub did not exist? No, they are presumably working for a fraction of what a researcher earns at a Global North university and do not have a budget of $25-60 for each article they need. The university also cannot afford millions of U.S. dollars for an Elsevier subscription. The scholar either gets the article from Sci-hub or some other green open access source, or does not use the article. It is a small loss for any publisher when someone would use their article, but could not access it; because having a potential citation to one of their articles is better than nothing.

Scenario 2. A researcher in the Global North downloads articles from Sci-hub. This researcher has either direct or indirect access to every published article that exists. Many universities have subscriptions to articles that their researchers are most likely to need. When the university does not have a subscription, a process of inter-library loan will get them roughly any article they need. Its not fool proof, but within a margin of error, Global North researchers have legal access. Using Sci-hubY yes, this researcher is getting something for free that they should pay for, but, their university or a university in their library network already pays for access, so preventing them from downloading also does not lead to any new money in the hands of the publisher.

Shutting down Sci-hub will not lead to any increase in profits for publishing companies. Elsevier, in its tantrum would argue otherwise. When universities and their libraries finally started to turn on predatory publishing houses like Elsevier and cancelled contracts, profits declined. In this case the universities do have the money to pay for a subscription. However, Project Deal and the UC systems’ boycotts asked Elsevier to sign a more reasonable and less draconian version of subscription and Elsevier refused. That is on Elsevier. It has nothing to do with Sci-hub. Again, no money would change hands because of the boycott which has nothing to do with universities implicitly encouraging their students to download illegal content.

Let’s look at more evidence. Figure 2 shows Elsevier’s profits in recent years. Sci-hub was in full effect in the mid 2010s. Interestingly, profits kept growing. They grow, and grow, and grow until something else happens that has nothing to do with Sci-hub. The universities began to wake up from their nightmares, and realized that they were being abused by publishers like Elsevier. In 2019, they started boycotting and demanding that publishers sign collective contracts rather than pay case-by-case, and that they greatly reduce their fees. And only in 2019 does a year-over-year profit growth model suddenly reverse direction.

Figure 2. Publicly available investor information from RELX

Elsevier’s profits only started to decline after they got canceled. Sci-hub has been irrelevant because those who can’t afford it will not pay and those who can already subscribe or should have a subscription, certainly won’t pay for Sci-hub downloads because they already do via their libraries. That is except for Elsevier’s refusal to support science, which is causing their own self-inflicted profit loss. Sci-hub is good for science, its mostly harmless otherwise.

Teaching to empower students as public, open and citizen scientists

Students pursuing a bachelor or master degree develop both labor market skills for a information and computer-based career, and learn to do science. Not all go on to be scientists. It would seem that those that do not, have no impact on scientific knowledge. They took tests and wrote papers in order to earn their degree. The test results and papers are only for their instructor’s eyes and maybe an occasional parent or other student.

It could be different.

  1. Science is an act of knowledge production
  2. Bachelor and master students have and develop useful knowledge
  3. Teaching them to share their knowledge and collaborate in knowledge production:
    • improves collective scientific knowledge
    • leads to a feeling of empowerment and utility among the students
    • creates ideal-type citizen scientists

Starts with teaching

In any given course in any particular discipline, students are taught to do science as a practice. This includes knowledge or ideas about the world; empirical evidence gathered through participation, observation or experimentation; and techniques to maximize the accuracy, efficacy and reliability of their knowledge and ideas. The students’ own work on assignments or theses is a form of ‘training’, and an instructor then checks or grades their learning. If they are privileged they get constructive feedback, and if they are highly self-motivated they actually read the feedback and incorporate it into their future work.

Students are generally aware that their work is for their instructor’s eyes only. At least in my experience, they cannot imagine their work shaping science or public knowledge. They thus have little extrinsic motivation to do more than what the instructor asks of them.

What if the instructor asks them to change the world outside the classroom in some way?

We, as instructors, are teaching students to produce reliable and critical knowledge. They should get a high mark if the work is deemed to be high quality. Why then does the entirety of their learning and knowledge production end as a forgotten file in a folder as a relic of some semester past? Why don’t they share some of that knowledge? I bet if they thought that their knowledge was valuable and useful beyond degree acquisition, they would be stoked.

Hampered by institutionalized beliefs and practices

A common argument against bachelor and master students trying to disseminate their work is that it is not high enough quality to compete with work from doctoral and post-doctoral researchers. In particular, they have not had enough time to dig deep into a body of literature on their topic of interest and their methodological skills may still be quite underdeveloped. They would probably be ‘wasting’ their time pursuing a journal publication or even a publicly disseminated working paper.

I agree. This points at the root of the problem. Publication-based science. Science as we know it has a rewards system where publications are treated as far more valuable than anything else scientists produce. This is particularly acute in the social sciences where scientific research rarely leads to tech or apps with private market value. When publications are the ‘currency of the trade’, academics, universities, editors, students, even policymakers prioritize publications and citations to those publications as the metric for judging the quality of scientific research. As such, scientists have maximum incentive to produce publications above all else.

Now, bachelor and maybe master students are generally unaware of the severity of this publish-or-perish plague that infects the very spirit of science. But like a fish that is unaware of the properties of the water it lives in, the students are deeply affected by the polluted nature of the academic norms in which they matriculate. How many times has a teacher told a bachelor student that they should submit their term paper to a high impact journal? How often are readings in scientific courses not from journal articles or books? A taskforce in sociology specifically recommended that students read and comment on books and journal articles because this is the ‘best’ type of academic knowledge for them to learn.

Student-citizen opportunities to shape public knowledge

With available technologies and new ideas about what constitutes meaningful knowledge, students can have a great impact on science; both now and into the future. Podcasts, blogs, vlogs, Youtube channels and many other forms of social media communication are consumed in high volumes across members of the public, and among students and scientists.

Therefore, when possible, I assign students the task of making a contribution to public knowledge and/or open science instead of writing a term paper. It seems better for all parties involved, and involves more parties because it could reach the public at large. I recently put this into practice in my course, “Open Science in Social Sciences: Crises, Controversies and Change” which I taught as an invited guest lecturer at the University of Zurich (UZH) (syllabus).

Wikipedia: knowledge now

Creating or editing a Wikipedia page where knowledge is lacking or does not exist at all, is a fantastic way to engage in public open science. Firstly, surveys report that more than a three-quarters and up to 90% of students use Wikipedia in their course research, at least in some English-speaking samples. Second, shaping knowledge immediately; seeing one’s own contributions appear on a public knowledge-platform is exciting and feels empowering. Third, knowledge is improved, in some cases much needed knowledge, like that which gives a voice, forum or contribution for underrepresented people or societies.

Three students in my course used Wikipedia to make contributions to public knowledge.

One of the students first edited and then created a Wikipedia page in Ukrainian, based on his Bachelor Thesis topic “A Theory of Generations”. He pointed out in class that most of the information on this topic, and knowledge accessed by Ukranians in general, is in Russian. Given the history of Russian hegemony in Ukraine, having more Ukranian resources promotes a Ukranian language and identity; in other words, promotes knowledge that is valuable to most Ukranians.

Wikipedia Ukrainian language page on “Theories of Generations“, created by Ernest Huk

Additionally, this particular academic topic is contested and misunderstood in the literature according to this student. He points out that the former page focused only on one theory, when in fact there are many. Thus, he pointed out that before his edits and page renaming, “Ukrainian users of Wikipedia, by looking through the [former] article Theory of Generations, would be not informed adequately at least and misinformed at most, as Strauss-Howe generational theory is just one of the many other theories in this domain (yet one of the most controversial).”

Another student has a extracurricular passion for wild flora. There is a class of plants whose native habitats are pastures and fields. These ‘pasture flora’ (Ackerbegleitflora in German), as the student pointed out, are “often rare plants that would find suitable living conditions in pastures but are conflicted with the threat of economic means that increase productivity in agriculture. Some of the rarest plants in Central Europe are found in this class.”

Wikipedia German-language page on “pasture flora” edited by Linus Signer

At least in German-speaking Wikipedia, there was no page on this topic at the outset. In fact, Wikipedia redirected the students search to Unkraut (weeds), reinforcing the public misunderstanding of this class of wild flora as undesirable or economically-inefficient intruders in a pasture or field. He later discovered an existing page on Segetalflora which is a related class that has overlap with ‘pasture flora’. This points out that knowledge on certain topics can go by different names or be cross-classified, a problem that requires more discussion and Wikipedia users to resolve. Therefore, he elected to significantly expand the existing Segetalflora page, rather than make a new page with information about Ackerbegleitflora, so that everything was in one place, including discussions about differences, conservation and utility.

During our course in fall-winter of 2021, many exciting things related to open science were taking place at UZH. For example, the university created an explicit open science policy. Moreover, there was an open science week with speakers and, of course, our open science course. One student noticed that much information about the open science happenings was missing from the UZH Wikipedia page and decided to make it his project to add it.

The ‘Open Science‘ section of the University of Zurich German-language Wikipedia page edited by Kerim Lengwiler

A major limitation to the open science movement is simply a lack of academic awareness of the issues and their solutions. Wikipedia provides a platform to disseminate such knowledge. Thanks to this student’s research and edits to the German-language page, I was easily able to update the English-language version of in tandem. We hope users can follow our lead and update the French and Italian versions as well.

Youtube – unlimited public science possibilities

Youtube is now outpacing Wikipedia as a student’s ‘go to’ for scientific information and especially for science communication. Internet searches for ‘how to’ or ‘information about’ now seem as likely to return Youtube or other short instructional videos as they are text-based entries (depending on cookies, browser, settings, and etc of course). Videos also give a voice to research subjects that cannot easily be expressed via written language. As long as participants consent, students can interview their subjects and post a video on Youtube – an act that requires only minimal technical skill and can be done with free or cheap software.

Another student in my course elected to contribute to knowledge on the topic colorism – discrimination based on skin color. Unlike racism, the student found this topic was far less present in academic discussion, at least based on her internet searches. Her impression from a previous course and her searches, was that this topic was mostly used in reference to the United States, but less so in the German-speaking countries’ contexts. In fact, she pointed out that to here it seemed like many people did not think colorism existed in Switzerland or Germany. Therefore, for her project, she conducted an interview with a person whose parents are Asian and European, to understand if and what types of color awareness or colorism this person experienced personally or in media and marketing.

An interview with an ethnically Swiss-Asian participant about colorism by
Nora Melanie Vetsch
Student open science activism

The opening of science and increase in reliability and transparency of knowledge for academic and public consumption requires more than a handful of citizen scientists active on social media and Wikipedia. Many closed and un-reliable methods are taught as part of the pathology of science. Instructors and students alike are generally unaware that the science they promote is potentially unreliable. This is another way of saying that they have not yet been exposed to the core information of the open science movement. Thus, changes at the curricular and institutional level are necessary to promote this awareness and to foster change.

Three students in my course elected to develop strategies to shape curricula for university and school students so that it fosters awareness of open, ethical science practices.

One student felt that herself and her peers were generally uninformed about open science both in and outside of UZH. Her idea for changing this was to create a student led social media initiative at the University. As this is nothing that could be achieved or even approved in a few months, her project was an action plan for the creation of this service. Firstly, it would entail the creation of a Student Open Science Organization, that among other things would maintain social media accounts where they posted important open science information and resources. This would require several layers of bureaucratic approval and liaising with the Open Science Office.

An excerpt of the open science at UZH social media action plan by Isabella Ferrera

Another student sought to communicate and potentially convince a primary methods professor in her area of Educational Science to incorporate open science into her courses. The idea would be that such a strategy could be deployed across other disciplines, and that her area was a test case to develop it. This required first developing persuasive reasons that could be shared in an open and friendly manner with a professor. Like many progressive universities, UZH has resources just waiting to be taken advantage of, a great realization of the student.

The Open Science Committee homepage at UZH, available for policy, communication and supporting curriculum development – for example Anastasiia Kurmann’s efforts to shape Educational Science courses so that they teach open science.

The professor responded to the student’s requests and agreed to add open science topics to her seminar plan, agreed with the students claims about the value of teaching open science and to specifically promote open data and FAIR principles as she teaches about qualitative data collection and evaluation; according to an exchange with the student.

Another student volunteers her time at a Zurich Community Center (Gemeinschaftszentrum). As the student points out, these centers, “offer space to work, play, learn, meet other people or participate in neighbourhood projects.” The student wanted to test if open science is a topic that adolescent children might understand or be interested in. Therefore, she first gained their interest and consent and organized a lesson at the Center. She used another resource of the UZH, the Kinderuniversität (Children’s University) as a protocol or concept for the adolescents to understand open science. The Kinderuniversität is specifically designed to support children in searching for answers and explanations of phenomena and, as the student points out, “serves as a good example to illustrate how science can be made accessible to all.”

Diellza Ismailji teaching adolescents at a Zurich Community Center about open science

Using this test case, the student was able to develop a curriculum for future courses for children and adolescents to discover and understand the concept of open science and why it is important. This guide includes resources that the children can directly interact with such as Wikipedia (that they can even edit this!), Blinde Kuh, Google Scholar, how to find and use a library and how to register for free at the UZH Children’s University.

Concluding thoughts

Public science is often practiced in the social sciences and involves researchers engaging with a local community to help provide inputs into their own efforts to address their own identified problems and priorities as a community. At international universities, bachelor and master students are less likely to be from the local community, are only staying there temporarily, and may face ethno-linguistic barriers to engaging in public social science locally. However, they can still have dramatic impacts on knowledge affecting to their home communities or countries – remotely. That is the beauty of an internet of knowledge, it only requires online access not physical presence.

Something else to keep in mind is that replacing a traditional term paper with a project to impact public knowledge and/or open science is not realistic in all course types. For example a broad introduction to open science across disciplines, like my course, makes it relatively easy for students to chose any topic and use it to contribute to public knowledge. In a course that is very theoretical and narrow, for example critical Marxism or history of the Holy Roman Empire, student learning might be maximized if they write a conventional essay based on reading of the literature that proves they have mastered a well-rehearsed topic. If asked to make a contribution to knowledge specifically in these areas they might struggle to find things to add to Wikipedia or find that there are already seemingly unlimited public resources. “Might”, then again might not. There are often opportunities to engage technology or do things in different languages.

Ultimately, when students feel they are doing something more than just earning a credential, ticking a box, or trying to maximize their grades, they become more likely to engage the material. If they think their term paper might actually contribute to local or global communities’ knowledge base, conservation efforts, capacity to address underprivileged students and etc., they are naturally inclined to do higher quality work and develop a sense of empowerment and satisfaction. These are bases of self-knowledge and fulfillment that they will hopefully carry with them into the future and stay motivated to impact public knowledge as scientists in academic or citizen scientists working in any job.