A bug bounty module in grants would give criticism a leg up
The Soviet Union was good at producing shoes. Factories made 800 million pairs a year, twice as many as Italy, three times as many as the United States. Three pairs per citizen per year. Producers met the targets officials set for them. Unfortunately, they rarely made shoes people wanted to wear, but the producers did not know or care.
Academia is good at producing publications. Scientists publish more than 3 million papers a year. Output has been rising for decades. Unfortunately, many of them are not reliable. They will not translate into improved knowledge and mastery of the world. Fraud, bias, hype, negligence get the paper out the door, but do not result in work we can rely on.

Citations are the closest thing we have to a quality signal, but they’re equivalent to an upvotes-only system. Citing a paper is cheap and often superficial. People cite papers they haven’t read critically, or even read at all. Worse still, people cite retracted papers, even many years after they have been retracted. Increasingly, thanks to AI, they cite papers that do not exist. Entirely failed paradigms like the candidate gene approach in human genetics keep accumulating citations for years after their fatal flaws have been identified. Some even write fake papers to cite themselves more. As a consequence, any publication is a good publication from the vantage point of an author. If your work sparks outrage, critiques, letters to the editor, that’s all publicity, a boost to citations. When a scientist faces the question of whether to pursue a publication or not, or to be a co-author or not, there is no incentive to hold off. No wonder that citations don’t predict replicability, reproducibility, or other quality markers.
In functional markets, if most of your products are lemons, you go out of business[1]Unless you’re a fruit merchant in which case thanks for your service against scurvy and sorry about the metaphor.. In academia, your h-index goes to 215. The expected cost of publishing a slipshod paper is close to zero and even fraud may slip through often enough to pay. Yes, your reputation might suffer with those who read the paper critically, but few do and far more decision makers pay attention to easily evaluable citation metrics. Corrections and retractions, which would reduce citations, are rare, because most of the time nobody checks, and if they do, most keep mum about what they find. The people who check and share what they find mostly do it on their own time and often face retaliation. Until recently, the time cost of publishing was a deterrent. We at least had the friction that you still had to write the damn thing. But now we have large language models and the only remaining friction is the forsaken piece of software called Editorial Manager, which by design or accident acts as a CAPTCHA. It is a dubious and unreliable ally in an effort to stem the tide of junk manuscripts.
We need a system where producing bad work has real, predictable costs, even for metrics-based reputation, where finding errors in others’ work is rewarded rather than punished.[2]If we had the power to overhaul the entire academic system, I’d be partial to Tal Yarkoni’s Reddit clone. But this post has the goal to identify a within-system solution. Who has a stake in this? Scientific grant funders who wonder why the work they fund rarely translates into economic or social advances.
Funders currently play a role in incentivising the metrics of success that academics aim for. In grant review, reviewers pay attention to a set of metrics that are positively correlated with each other: publications, citations[3]People love to complain about the impact factor, but everybody’s darling alternative, the eigenfactor is usually highly correlated with the IF, because in the end they both stem from citation counts., previous grant income, media attention, patents, tenure, career age, lab size. Metrics-based funding decisions produce Matthew effects and are bad at finding good ideas from people who don’t yet have long track records, especially younger scientists.[4]The ERC, for example, increasingly discourages that reviewers consider applicants’ track records and instead asks them to focus only on the proposal.
To be fair to the funders, they have to sift through a lot of grant proposals, so it is difficult to put all of them through detailed evaluation by scientific peers. Peer review is treated as a duty, a service to the field. But many scientists shirk this duty, so in the end, we get two or three reviewers, knowing full well that they don’t even tend to agree very well on the quality of proposals. You’d need tens of reviewers to get reliable evaluations and you would ideally get different specialists for the different parts, from inside and outside the paradigm. Currently, we usually tap reviewers who are close to the subject matter. They are able to spot common domain-specific errors, but often share the authors’ blind spots and might even be friends with them. But the scientists best equipped to identify the flaws common to an entire research program, or specific common statistical mistakes, are often outsiders who have little motivation to do reviews in this field.
There’s a solution that could both decrease the quantity of proposals and increase their quality. A move from an upvote-only circle jerk to a system with a downvote component. We can shift away from a situation where every grant proposal has a positive expected value[5]A friend took issue with this. To clarify, here, I’m using expected value to refer to the gross payoff. I’m not pricing in (time) costs, because these vary so much, especially by career stage. Attaching your name as a senior co-author to a paper is often almost “free” in terms of time. Same with some grants, which have requirements for collaborators from a different department, faculty, or country. Of course, writing grants can be a huge timesink, but a) that’s changing with LLMs b) the widespread perception that funding decisions are a lottery rather contributes to researchers “rolling the dice” more often, i.e. every proposal is another shot at winning., where scientists knowingly or ignorantly submit slipshod work:
A bug bounty module in an experimental new grant funding line. The module is funded along with each new grant as 5% of the grant sum. The bounty is publicly posted and mentioned in every publication resulting from the grant. It pays out to anyone who identifies flaws, errors, or fraud in the funded work or in the preliminary work that led to the grant.[6]I.e. work by the same team cited in that publication. The payout scales with the severity of the reported error, so that an error that changes a headline conclusion pays the most. Adjudicators are named by the funder at the time of the funding decision and should be people renowned for their fairness, not the applicants’ friends or competitors. Their decisions and the bounties paid out are made public. If the bounty goes partially or fully unclaimed after a set period, the remainder reverts to the grant holder.
Bug bounties are a practice borrowed from the software industry and have served it well.[7]In software development, the rationale for bug bounties is even stronger. In their absence, hackers can turn security-relevant bugs into exploits which they can use or sell on the black market. So, it’s in software companies’ interest to pay hackers to reveal these bugs legally and safely. In science, the funders take the role of software companies and they are interested in protecting the work they fund from negligence and fraud. Science is amateur software development, after all, and some of the recent improvements in reproducibility can be traced to our increasing adoption of best practices in software development such as version control and documented, reproducible code.
Every scientist funded through this experimental line expects their work to be scrutinized.[8]Will they feel like they painted a target on their back? Anecdotally, me and my friends who have placed bug bounties on our work with our own money feel good about it. It feels like opting into a more positive error culture: we’ve publicly declared we won’t double down, but we also accept errors will happen, so it’s a balm for the perfectionist side. Their work will be checked, and someone will get paid for finding problems. The expected value of an additional grant for their h-index is no longer positive, if they have reason to believe these checks will lead to corrections and retractions. A scientist who knows ChatGPT wrote the second half of their proposal a minute before the deadline will think twice. Because even the preliminary work is exposed to scrutiny, frauds and hackers self-select out of this funding line entirely. This would reduce the number of proposals needed to screen and allow for more in-depth peer review, which is attractive for funders. A scientist who is confident in the quality of their work is enticed by the fact that the bounty’s remainder reverts to them after 5 years. And the funder gets post-publication quality control on the cheap, which will come in handy the next time somebody considers cutting their budget because of a perception that the funded science is insular, bad or irrelevant.
A record of errors found through bounties follows the researcher. Funders can see it in future applications. Not as an automatic disqualification, but as information. Conversely, a track record of clean, scrutinized work should be a genuine asset, more meaningful than a high h-index. Small, inconsequential errors are common and increased scrutiny would normalize acknowledging and correcting them. But for errors that undermine published conclusions, the funder’s adjudicator requires the grantee to issue corrections or retractions as a condition of future funding. A track record of slipshod work should be visible, the way a track record of loan defaults is visible to a bank.
For fraud, the consequences should be more severe. If a bounty claim reveals fraud, the funder should claw back the grant money. Currently, universities eschew taking action on their local frauds out of laziness and PR concern, but a grant clawback will trigger institutional action and hard career consequences.
Best of all: we don’t need every critic to look at everything. People can do what they do best. If Elizabeth Bik mainly wants to check figures for duplication, let her trained eye scan. If James Heathers only wants to check means for inconsistencies, GRIM away. If Julia Rohrer wants to slap down every paper that ignores the heroic assumptions required for mediation analysis, paint the margins red. If Jamie Cummins wants to program an army of AI agents to check for deviations from pre-analysis plans, he can field it. If Saloni Dattani wants to go down the rabbit hole to discover that people have been repeating a made-up number for ages, it’s her gift. If Michael Wiebe wants to go deep and rerun every line of code for a paper with red flags, let him cook. If Uri Simonsohn wants to use his intimate knowledge of Microsoft Excel internals for good, this is the way. If Andrew Gelman wants to slap down any paper that treats the difference between significant and non-significant as itself significant, more power to him. If your journal club always pokes holes in any paper you read, well now you can at least afford cake. If Ian Hussey’s bat signal is papers claiming to have found mean differences exceeding that between human preferences for chocolate over poop, I mean we always knew he was odd. If some poor grad student who thought they’d found the ideal paper to build their project on discovers they’re building on sand, they can at least get consolation, or even satisfaction.
We would get specialization in error detection. It’s a specialist’s work, but currently we are expecting two to three reviewers to cover all bases. We also need no central coordination matching reviewers to papers.[9]A lesson learnt at ERROR (error.reviews), a bug bounty program for science that I co-run. There are errors to be found aplenty; few papers are free of them. But finding qualified reviewers that won’t flake is a real challenge, even though we pay well, because making your own contributions is reputationally more rewarding. Just a bounty, posted publicly, and whoever wants to collect it does the work they’re already inclined to do. Some people would go wide and use automation to find small but common reporting errors, others would go deep and invest time to uncover severe errors.
Academia gets downvotes! Or rather, academia gets downvote abundance, whereas previously corrections and retractions were too scarce to deter negligence. The power of the downvote is the reason why for a while you could fix the SEO-optimized Google results by adding “reddit” to your query. Can it fix h-index-optimized academia too?
Funders might worry that this will make scientists too unambitious. After all, scientists are already quite risk-averse.[10]See David Oks for an argument that citation metrics have contributed to risk-aversion. There is a common piece of advice that an NIH grant should have 3 aims, 2 you’ve already done and 1 you’re never going to do, leading to the quip that the NIH doesn’t fund research, it refunds research. At the same time, scientists are certainly not averse to claiming that their ideas will shift paradigms (the ERC even requires ideas to be “groundbreaking” to be eligible for funding) while simultaneously being a guaranteed return on the invested grant money. Bollocks. What we lack is calibration. We want scientists to acknowledge when their ideas are risky. Taking a risk isn’t an error and an honest report of a risky endeavour is no better target for bounty hunters than any other research project.
But letting overclaimers and bullshitters pocket a large percentage of grants, as we do now, might crowd out the calibrated risk-takers, the ones who are upfront about the fact that the potential gain is big but also unlikely, that risks and rewards are related. Plausibly, scientists who have taken real risks in their career look worse on metrics, because not all risks taken pan out as publishable output. If bug bounties cause fewer proposals to be submitted, this decreases the strain on reviewers. Reviewers can then afford to be less metrics-oriented and focus on quality, where the risk takers might shine.
This would work for an experimental grant funding line because the sheer novelty of it would attract plenty of bounty hunters and adjudicators. To perpetuate such a program and make it mainstream, we’d need another component. We need institutes for scientific quality control[11]Existing efforts like the Institute for Replication (I4R) or ERROR are small and cannot offer long-term career prospects., where critical-minded people can have a career. Right now, being known as a critic exposes you to risks of revenge and reputational harm, but does not improve your academic metrics of success, because criticism is hard to publish in the outlets that conventionally matter for metrics. Getting a paper retracted improves the scientific record, but is not commonly something you list on your CV. Error detection is not a career track and bounty hunting is too unsteady (including in software, from where this practice is borrowed). In medicine, we have institutions like the FDA because we accept that expert adversarial review is needed to find flaws in proposed drugs. In business, we have independent financial audits. People who have the inclination to do quality control work need to be trained and sustained. If doing audits is a career of its own, they don’t need to pull their punches to still have friends when their own work is under review. Such institutions could also prime the pump for the first proposals to carry bug bounties.
Financially incentivized scrutiny in science already exists, by the way. The US False Claims Act allows whistleblowers to receive a share of recovered funds when they report fraud in federally funded work. Earlier this year, Sholto David, a molecular biologist working at a biotech firm in Wales, was awarded $2.6 million[12]Some of which will go to his attorneys and it will be taxed. after identifying duplicated and misrepresented images in papers by Dana-Farber Cancer Institute scientists that were cited in NIH grant applications. David had flagged problems in several cancer biology papers, including some by senior Dana-Farber leaders, which led to dozens of corrections and retractions. Dana-Farber settled with the US Department of Justice for $15 million.
But whistleblower programs are uncertain in their outcomes and focused on fraud. Proving fraud is much harder than identifying error or negligence, even though the latter are just as consequential for the scientific record. Alleging fraud quickly leads to legal threats.
The bug bounty model therefore simply focuses on errors, though of course it would uncover fraud too. And because the bounty also extends to the preliminary work that goes into grant proposals, it changes the incentives even before submission. It preserves the notion of science as self-correcting, but it speeds up the self-correction process. Slipshod science should not be worthwhile. We will know we have succeeded when researchers start turning down guest co-authorships and look for code reviewers for their work. And if it does not work, if the experimental program is just as overwhelmed with submissions, if no one submits bug reports, at least I’ll have discovered an error in my thinking.
Thanks to Julia Rohrer, Malte Elson, Anne Scheel, Saloni Dattani, Jamie Cummins, and Ian Hussey for comments on an earlier version.
Footnotes
| ↑1 | Unless you’re a fruit merchant in which case thanks for your service against scurvy and sorry about the metaphor. |
|---|---|
| ↑2 | If we had the power to overhaul the entire academic system, I’d be partial to Tal Yarkoni’s Reddit clone. But this post has the goal to identify a within-system solution. |
| ↑3 | People love to complain about the impact factor, but everybody’s darling alternative, the eigenfactor is usually highly correlated with the IF, because in the end they both stem from citation counts. |
| ↑4 | The ERC, for example, increasingly discourages that reviewers consider applicants’ track records and instead asks them to focus only on the proposal. |
| ↑5 | A friend took issue with this. To clarify, here, I’m using expected value to refer to the gross payoff. I’m not pricing in (time) costs, because these vary so much, especially by career stage. Attaching your name as a senior co-author to a paper is often almost “free” in terms of time. Same with some grants, which have requirements for collaborators from a different department, faculty, or country. Of course, writing grants can be a huge timesink, but a) that’s changing with LLMs b) the widespread perception that funding decisions are a lottery rather contributes to researchers “rolling the dice” more often, i.e. every proposal is another shot at winning. |
| ↑6 | I.e. work by the same team cited in that publication. |
| ↑7 | In software development, the rationale for bug bounties is even stronger. In their absence, hackers can turn security-relevant bugs into exploits which they can use or sell on the black market. So, it’s in software companies’ interest to pay hackers to reveal these bugs legally and safely. In science, the funders take the role of software companies and they are interested in protecting the work they fund from negligence and fraud. |
| ↑8 | Will they feel like they painted a target on their back? Anecdotally, me and my friends who have placed bug bounties on our work with our own money feel good about it. It feels like opting into a more positive error culture: we’ve publicly declared we won’t double down, but we also accept errors will happen, so it’s a balm for the perfectionist side. |
| ↑9 | A lesson learnt at ERROR (error.reviews), a bug bounty program for science that I co-run. There are errors to be found aplenty; few papers are free of them. But finding qualified reviewers that won’t flake is a real challenge, even though we pay well, because making your own contributions is reputationally more rewarding. |
| ↑10 | See David Oks for an argument that citation metrics have contributed to risk-aversion. |
| ↑11 | Existing efforts like the Institute for Replication (I4R) or ERROR are small and cannot offer long-term career prospects. |
| ↑12 | Some of which will go to his attorneys and it will be taxed. |

I agree overall, ofc.
I’d like to mention someone once created a platform for hiring “red teams”, researchers tasked with finding flaws before the project is (nearly) set in stone, but I couldn’t find it rn. And of course there are the “smart citations” of Scite, a commercial platform.
Yes, I was involved in Red Team Markets, led by Leo Tiokhin, which has since been sunsetted. Like with ERROR, we learned some lessons (among other things that authors are not very interested in paying for this service and that centralising the recruitment and coordination of the red teamers is quite effortful). Smart citations in Scite are an interesting idea, but I once tried it out in an area I know has big problems (5-HTTLPR) and was not very impressed: https://x.com/rubenarslan/status/1587781179977211905