What Fraction of Astronomy Papers use LLMs?

As a journal Editor I’ve been grappling with the problem of a huge influx of papers that use various forms of generative AI. As I’ve said before the problem seems to me not that competent scientists use LLMs and the like to speed up their work, but that fools can use them to produce superficially plausible but scientifically flawed articles very quickly. It’s the latter category that comprises much of the increase at the Open Journal of Astrophysics. I’d be far more positive about LLMs if we could rely on people to use them properly!

Anyway, there’s a new paper on the arXiv with the title More than half of recent astronomy papers are written with language-model assistance. Here’s the abstract:

As you can see, the claim is that over half of the recent astronomy papers on the astro-ph section of arXiv use Large Language Model (LLM) assistance but that only a tiny fraction of that number declare it. (In my experience those that declare use of LLMs also understate the extent if their use.)

I think this paper is well worth reading. It is based on the excess frequency of certain words and phrases. The authors explain:

Counting excess vocabulary is straightforward. Converting a count into a statement of the form “x% of papers used a language model” requires the frequency these words would have had if language models had never been released, and that frequency stopped being observable in November 2022. Statistical precision is not the obstacle, since with hundreds of thousands of papers the observed frequencies are known almost exactly. That frequency can only be supplied by a model, so every estimate in this literature rests on the model chosen for it.

They go on to analyse various models to arrive at their conclusions. I’ll leave it to the reader to go into the details.

One issue I have is that the analysis seems to assume that all changes after the introduction of LLMs in late 2022 are attributable to LLMs. But language changes anyway, and was changing long before AI came on the scene. I would like to see, as a control, a comparison between, say 2006-2010 and 2011-2015 so see if similar changes happened. One problem is that the total number of papers on arXiv is growing monotonically so there are many more papers in 2025 than there were in 2006.

P.S. Ironically, the analysis behind and the writing of this paper were both “assisted” by Claude….

P.P.S. Here’s a suggestion: if you use an LLM to help you write the paper, then you should name it as a co-author.

8 Responses to “What Fraction of Astronomy Papers use LLMs?”

  1. Anton Garrett's avatar
    Anton Garrett Says:

    In most cases – especially with authors whose first language isn’t English – I suspect AI was used simply to improve the quality of the prose. I regard doing that as legitimate.

  2. There is now a dispute about whether the recent OpenAI results in mathematical proof had made used of on-going, non-public research from AI-users.

    Is an AI like Claude a tool or a competitor?

    • Anton Garrett's avatar
      Anton Garrett Says:

      I’m interested in the claims in the last week that AI has settled one of the seven Millennium mathematics problems, about whether the Navier-Stokes equations of fluid flow have solutions (existence theorem) that don’t blow up (non-singularity theorem). Apparently a scenario has been found in which the solution does blow up in finite time starting from finite boundary conditions. When I first read that, I wanted to know if this was for the incompressible NS equations only, as I reckoned that compressibility would mitigate any singularity. The claim is indeed for the incompressible equations. Two research groups are involved and both openly acknowlege major AI assistance, but one group claims that the AI of the other was trained on its work… nothing like a good academic row (although historians do it best).

      The result needs some verification/refereeing, but it will obviously get that in view of the status of the question. I reckon it is correct, and I am less concerned than many because not only does compressibility make a difference but I am well aware from statistical mechanics of the approximations made in reaching the NS equations (think BBGKY hierarchy, and the Champan-Enskog expansion).

      • It is difficult (and somewhat dangerous) to judge purely based on new reports. The claim seemed to be that at OpenAI, the AI had itself solved the problem without academic assistance. That may not be the full story. But if the AI used private results from the other group which it had access to because of their AI use, that would be problematic. And possible.

      • Anton Garrett's avatar
        Anton Garrett Says:

        I’m merely saying where I’d place a bet. That goes on in every physics tearoom, and a blog comment speaking carefully of ‘claims’ isn’t a commitment!

Leave a comment