The tells are easy to spot for anyone who reads closely: a citation to a paper that does not exist, a stray line in which the chatbot asks whether it should revise the text, a table cell that instructs the author to fill in real numbers later.
Under a policy announced May 15, such signs can now cost a researcher a year of access to the world’s largest preprint repository. ArXiv said it will ban authors from submitting papers for one year if their submissions show “incontrovertible evidence” of unchecked AI-generated content — including hallucinated references, leftover chatbot meta-commentary, and unremoved placeholder instructions. The policy, announced by Thomas Dietterich, chair of arxiv’s computer science section, was reported by TechCrunch and represents the repository’s most aggressive response yet to the flood of machine-generated academic content.
The ban is not a prohibition on using AI tools. ArXiv’s stated target is authors who submit AI output they never checked. “If authors did not check the output of their LLM, we can’t trust anything in the paper,” Dietterich said in his announcement. Under the policy, a submission that shows unambiguous signs of unverified AI output triggers an immediate one-year ban from submitting to the platform. After the ban expires, the author’s next submission must first be accepted at a reputable peer-reviewed venue before arxiv will post it.
The policy is aimed at carelessness that is easy to prove, not judgment calls about how much AI assistance is acceptable. Reported examples of “incontrovertible evidence” include citations to papers, authors, or journal issues that do not correspond to any real publication — a well-documented failure mode of large language models; chatbot meta-commentary left in a manuscript, such as “here is a 200-word summary; would you like me to make any changes?”; and placeholder instructions left in tables or data cells, such as “the data in this table is illustrative, fill it in with the real numbers from your experiments.”
The common thread: these are things a careful author would have caught by reading their own paper before submitting it. ArXiv’s enforcement process requires a moderator to flag a case and a section chair to confirm the evidence before a ban is imposed. Authors can appeal.
The policy does not attempt to define what counts as legitimate AI assistance. Researchers remain free to use language models for translation, editing, or drafting, provided the work is disclosed and checked. ArXiv’s line is drawn at evidence of neglect, not at the presence of a tool.
The policy targets individual accountability rather than attempting to build a detection classifier at the moderation layer. ArXiv’s operators concluded that penalizing authors shifts the verification burden upstream to the people who control what gets submitted in the first place. Paul Ginsparg, the Cornell physicist who co-founded arxiv, told Nature that AI-generated submissions “frequently can’t be discriminated just by looking at the abstract, or even by just skimming full text” — surface plausibility is high enough that spotting fabrications requires close reading, which does not scale when thousands of new papers arrive each month. Ginsparg called the phenomenon an “existential threat” to the system.
The new policy is the latest escalation in arxiv’s response to AI-generated content. In October 2025, the platform said it would no longer accept computer science review articles and position papers unless they had already been peer-reviewed elsewhere, after moderators described an “unmanageable influx” of low-quality submissions that looked like literature surveys but were stitched together by an LLM with minimal human curation.
The May policy goes further by holding individual authors accountable regardless of category. ArXiv, operated by Cornell University, hosts more than 2.4 million research papers and processes more than 200,000 submissions annually across physics, mathematics, computer science, and related fields. It has long operated on a system of trust and lightweight moderation, and its identifiers are used across the scientific community to establish priority in fast-moving fields. The policy is an attempt to protect the value of the arxiv identifier as a mark of quality.
The shift arrives as the rest of academic publishing moves in the same direction. Conferences including NeurIPS and ICLR now run hallucination audits on submissions and have rejected papers with fabricated citations. Journals are updating author guidelines to require disclosure of AI use and verification of AI-generated text. ArXiv’s ban is the most direct penalty yet: a year-long suspension from the primary venue where researchers establish priority.
The policy is easy to state and hard to apply at scale. It relies on moderators or readers noticing obvious tells — a stray chatbot phrase, a citation that does not resolve. That catches careless submissions but says nothing about AI-generated content fluent enough, or checked carefully enough, to leave no trace. Critics, including Inside Higher Ed, have described the rule as “welcome but unenforceable,” noting that the detection burden falls on a small volunteer moderation team.
There is also a distributional concern. The authors most likely to trigger the ban are junior researchers and students who may not know how to spot a hallucinated citation, rather than the labs that produce most of the field’s AI-assisted work. ArXiv’s response is that authors own their submissions: the author, not the tool, is accountable for everything under their name.
ArXiv’s one-year ban moves academic AI enforcement from tolerant gray areas to hard rules. The policy does not ban AI assistance — it penalizes the failure to check it, with evidence standards deliberately set high enough to avoid punishing judgment calls. Whether the rule can actually be enforced at scale, and whether it deters the careless submissions that motivated it, will become clear in the months ahead. For researchers, the message is unambiguous: read what the model wrote before you submit it.


