Archive for Claude

What Fraction of Astronomy Papers use LLMs?

Posted in Artificial Intelligence with tags , , , , , on September 11, 2026 by telescoper

As a journal Editor I’ve been grappling with the problem of a huge influx of papers that use various forms of generative AI. As I’ve said before the problem seems to me not that competent scientists use LLMs and the like to speed up their work, but that fools can use them to produce superficially plausible but scientifically flawed articles very quickly. It’s the latter category that comprises much of the increase at the Open Journal of Astrophysics. I’d be far more positive about LLMs if we could rely on people to use them properly!

Anyway, there’s a new paper on the arXiv with the title More than half of recent astronomy papers are written with language-model assistance. Here’s the abstract:

As you can see, the claim is that over half of the recent astronomy papers on the astro-ph section of arXiv use Large Language Model (LLM) assistance but that only a tiny fraction of that number declare it. (In my experience those that declare use of LLMs also understate the extent if their use.)

I think this paper is well worth reading. It is based on the excess frequency of certain words and phrases. The authors explain:

Counting excess vocabulary is straightforward. Converting a count into a statement of the form “x% of papers used a language model” requires the frequency these words would have had if language models had never been released, and that frequency stopped being observable in November 2022. Statistical precision is not the obstacle, since with hundreds of thousands of papers the observed frequencies are known almost exactly. That frequency can only be supplied by a model, so every estimate in this literature rests on the model chosen for it.

They go on to analyse various models to arrive at their conclusions. I’ll leave it to the reader to go into the details.

One issue I have is that the analysis seems to assume that all changes after the introduction of LLMs in late 2022 are attributable to LLMs. But language changes anyway, and was changing long before AI came on the scene. I would like to see, as a control, a comparison between, say 2006-2010 and 2011-2015 so see if similar changes happened. One problem is that the total number of papers on arXiv is growing monotonically so there are many more papers in 2025 than there were in 2006.

P.S. Ironically, the analysis behind and the writing of this paper were both “assisted” by Claude….