OpenAI’s newest chatbot may be a whiz at math, but it seems to be lagging far behind humans in its academic rigor.
Last week the company announced 10 more artificial-intelligence-generated math advances that were found during internal development and testing of its next major large language model. This batch of results came from that LLM, Astra, and each one resolves or progresses a different “long-standing open problem” of “substantial interest” to the mathematical community. The company said that the total token cost was a mere $2,000.
The news quickly spread as yet another harbinger of AI’s promise—or threat—of outpacing humans to become a dominant, disruptive force in math and computer science. But over the days following the announcement, as flesh-and-blood experts poured over the nearly 250-page paper in detail, many grew frustrated.
On supporting science journalism
If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
Two of the most exciting results, the experts say, incorporate preexisting ideas from the recent mathematical literature without properly citing them. This contradicts OpenAI’s initial press release, which said that the problems Astra addressed “have been open and seen no progress on the main result for at least a decade.” (OpenAI has since updated the language to be more accurate).
“They are running roughshod over the work of others who came before them in a deliberate way,” says Steven Miller, a mathematician at Yeshiva University, who argues that OpenAI has effectively plagiarized his own research. “It seems completely systematic to me, and it points to research misconduct.”
The result Miller refers to concerns how many balls you can fit in a box—a seemingly simple problem, except these balls and boxes exist in a mathematical space of 1,000 dimensions—or even more. OpenAI’s paper improves the best estimate for how tightly these balls can possibly be packed. The LLM-generated proof hinges on a particular mathematical argument that it presented as its own but that actually first appeared in a 2016 paper by Miller and a collaborator.
Another of the 10 results resolves a long-standing question in group theory, which studies sets of mathematical objects called “groups” that interact in an organized way. Mathematicians have long wondered whether all groups have a property called “soficity,” the capacity to be faithfully approximated in a particularly way by other, simpler groups. The OpenAI paper establishes at least one group that lacks this property.
The discovery stunned Francesco Fournier-Facio, a mathematician at the University of Cambridge, who studies group theory—at least until he “engaged with this breakthrough as I would if a human had written it,” he says. The result, he and some of his colleagues found, wasn’t as novel as it first appeared. Like a number of recent AI breakthroughs, it pasted together ideas from the mathematical literature to build a new theorem. Once again, the LLM’s trick is its superhuman patience for assembling puzzle pieces, not the ability to make some profound intellectual leap.
In particular, Astra’s key mathematical step combined ideas first found in two papers from 2016 and 2019. Andreas Thom, a mathematician at the Dresden University of Technology, who co-authored both papers, summarized the result on MathOverflow.com, calling it “creative and at the same time elementary.”
OpenAI’s initial press release seemed to ignore—or be completely unaware of—these crucial, recent developments. Fournier-Facio argues that the two preceding papers show humans had not hit a stalemate with the soficity problem. OpenAI’s mathematicians did their best to attribute these ideas correctly in their paper, he says. But, in spite of their good intentions, “there is the big PR machine that wants to sound as impressive as possible and does not care about being 100 percent accurate,” he says.
“We take responsibility for the correctness of these results and are meeting the same standards generally expected of human mathematicians,” an OpenAI spokesperson said in a statement to Scientific American. “We plan to make small updates [to the paper] this week, consistent with standard academic practice.”
But as AI continues its campaign to conquer math without any built-in fealty to the field’s academic norms, some in the community are clearly losing patience. “OpenAI is now fully participating in high-level research,” Fournier-Facio says. “So they should be held to the same academic standards that we are.”

