Toxic Mathculinity

Toxic Mathculinity

“Civilization advances by extending the number of important operations which we can perform without thinking about them.” — Alfred North Whitehead, An Introduction to Mathematics (1911)

I have been thinking about what happens when a machine arrives carrying the answer to a question you have spent your life asking. Nathaniel Whittemore’s recent AI Daily Brief episode on AI solving your life’s work puts that question squarely on the table. The answer can be a gift. It can also feel like an eviction notice.

On October 6, OpenAI released mathematical results from an internal frontier model that has yet to be publicly released. At the time of writing, its public repository lists 719 manuscripts organized into 372 families. Those are manuscripts and related result families, with differing verification status, rather than 719 independently accepted resolutions of famous open problems. Even with that distinction firmly in place, the scale demands attention.

My concern reaches beyond mathematics. I expect AI to alter civilization, including the institutions through which we decide what knowledge deserves trust. Some now reach for the name superintelligence. We can argue about when that designation is earned. Meanwhile, we have work to do. Innovation can outrun the environment that tests it, and the resulting gap is where both opportunity and danger accumulate.

Who Asked the Machine?

The Association for Human Mathematics’ October 7 statement opens one of its objections with a sentence that stopped me: “Mathematicians did not ask for this work to be done.” The statement challenges the release process and the separation of proof production from human understanding, and urges mathematicians to end their work with OpenAI. Its concerns deserve a hearing. Its permission premise deserves an argument.

Who gets to authorize an answer?

I understand wanting a company to consult the people whose intellectual and professional lives its announcement will disrupt. Consultation is a matter of responsibility. A profession’s consent to the discovery of mathematical truth is a different proposition. If we grant that veto, we have quietly made knowledge the property of the people already occupying its institutions. That is a dangerous deed to issue, however beautiful the cathedral.

Scott Aaronson pushes back on this territorial claim in “The Mathocalypse”. He argues that open mathematical problems are available to anyone, while acknowledging the pain of researchers whose work has been disrupted. I agree with that combination. We can defend the freedom to discover and demand a better manner of discovery.

I have come to call the territorial reflex Toxic Mathculinity. The spelling is deliberate. It names the posture that turns intellectual achievement into a dominance ritual: who owns the problem, who gets to be the genius, who must acknowledge defeat. Gender provides the wordplay; the behavior has no gender requirement. A mathematician can display it. So can a technology executive announcing that the mathematicians have been conquered.

Grief deserves a different vocabulary. So does a demand for a readable proof. Calling either one Toxic Mathculinity would give the companies a convenient excuse to disregard the people who can best examine their claims. I am aiming at the ownership ritual, including the version performed on a corporate stage with a larger compute budget.

Voices in the Common Room

The mathematicians themselves refuse the easy caricature. In their signed reactions collected by Proofs and Prompts, Terence Tao sees promising ideas and opportunities, but also missing human interlocutors and disrupted apprenticeships. Hugo Duminil-Copin describes being overwhelmed by the simultaneous arrival of so many claimed resolutions. Kevin Buzzard welcomes the advance while expecting considerable work to sort out the papers. Emily Riehl stresses the human mathematical culture and accumulated contributions that made this moment possible. These are different judgments, sometimes held within the same person.

Matt Baker offers another useful combination in his early response to the release. He criticizes its carelessness and distinguishes formally checked material from natural-language arguments needing scrutiny, while expressing awe at the apparent mathematical advances. That combination strikes me as intellectually healthy. Admiration need not purchase silence about defects.

I find these reactions more credible than either the victory parade or the funeral procession. An experienced researcher can be excited by an idea, angry about its packaging, and worried about a student’s future. We should expect that from people taking the work seriously. Demanding a single emotional response would be another status ritual, with enthusiasm serving as the admission ticket.

I recognize that shock from software development. Developers were among the first to confront this kind of disruption, and I have been on the front lines of AI experimentation, evangelism, and personal adoption since before the publication of Attention Is All You Need in 2017. That history has given me plenty of enthusiasm. It has also given me a close acquaintance with the fear of becoming redundant. I can say with certainty that the fear is real. Advocating for the technology provides no immunity.

When a machine can perform work around which you have built your competence and professional identity, the question gets personal very quickly. What am I still for? You can admire the capability and feel the threat in the same afternoon. I am sympathetic to mathematicians encountering that question because I recognize it. They have company outside the cathedral.

For developers, I see only one path forward. Embrace the tools fully and without reservation. Make them part of how you work, and keep learning as they change what the work requires. Adoption carries engineering responsibility: I still have to test the code and answer for what I put into the world. But I cannot preserve my usefulness by preserving a version of the job the tools have already changed. That is the conviction I bring to this conversation, along with the fear. I am asking mathematicians to consider a response I have had to choose for myself.

There is also a difference between losing a race and losing the conditions under which you learned to run. A senior scholar may have the security to change direction. A doctoral student has a supervisor, a deadline, and a funding arrangement built around a research plan. Advising that student to discover something else is easy from a comfortable chair. The institutions will have to do more than offer inspirational remarks.

When the Fitscape Falls Behind

Consider the practical scene. A repository receives hundreds of manuscripts. Specialists begin reading. Some formal proofs can be checked by software, while explanations must be reconstructed and connections explored. OpenAI says the mean compute per result was equivalent to about three hours of ChatGPT Pro thinking time in its release account. That is a compute comparison, not a promise that a subscriber can reproduce each result in three elapsed hours. It nevertheless raises the question of how quickly a research community can absorb what a model produces.

I use Fitscape for the environment in which a capability encounters the tests and consequences that establish its fitness. In mathematics, that environment includes formal checking, expert criticism, and the slower formation of shared understanding. Outside mathematics, it also includes people who must live with the consequences of decisions they never agreed to make. A laboratory benchmark samples part of that environment. It cannot stand in for the whole.

Here is the pivot. Our problem is not a shortage of answers. It is the growing distance between answer production and our capacity to establish what those answers mean.

Mathematics makes that distance unusually visible because it has such demanding tools for checking correctness. Thomas Hales explains in his account of Lean’s reliability that a proof needs kernel checking and a human audit of statement fidelity. Does the formal theorem say what we believe the mathematical claim says? Do its definitions correspond to the intended concepts? A machine can faithfully check a formal statement while the correspondence to the intended question still needs examination.

Even after that work, understanding has its own clock. Being satisfied that a theorem follows from specified assumptions does not tell me which idea made the proof work. It does not tell a teacher how to explain it, or a researcher which nearby questions the argument opens. Those are further achievements. Human beings have to acquire them, although AI may help with the acquisition too.

Could it take decades for the mathematical community to digest what OpenAI has put before it? I think that is possible if digestion means extracting the ideas and making them part of a widely understood body of mathematics. It is a possibility I am raising, not a measured forecast or an estimate from the mathematicians quoted here. Individual results may be checked and explained much sooner. Better tools could shorten the interval. The concern is the backlog if new results keep arriving faster than shared understanding grows.

We are changing the Fitscape while trying to test the inhabitants. Evaluation methods will themselves acquire AI assistance. The community will learn new forms of explanation. Institutions will change what they reward. Any prediction that assumes today’s capacity remains fixed misses that movement; any declaration that the machine will solve the entire problem for us assumes the very fitness we still need to establish.

As I argued in “AI: The Only Benchmark that Matters,” a score needs a relationship to the world in which it matters. For mathematics, that world includes knowledge valued for its own sake. Nobody needs to attach a quarterly return to a theorem to justify caring about it. But a manuscript count alone gives us a poor measure of how much mathematics humanity can now understand and use.

The Shadow Has a Balance Sheet

There is a shadow side here we must see. A company can release an enormous collection of results, enjoy the attention, and leave a dispersed community to supply the costly interpretation. Calling that arrangement collaboration does not make the labor free. Somebody pays, even when no invoice arrives.

The Advisory Group on Mathematics and Artificial Intelligence’s recommendations address this directly. They call for support for human understanding, transparent accounts of the process, and equitable access to tools. OpenAI says it will fund workshops and other programs in its announcement. The question now is what gets funded, who directs it, and whether the people doing the work can remain independent of the company whose results they examine.

Access matters for another reason. A public result paired with an unavailable model gives researchers something to inspect without giving them the same means to pursue their own questions. Publishing the output is useful. Sharing the ability to investigate is a separate commitment. A future in which a few laboratories choose the questions and everyone else becomes their interpretation department would impoverish the freedom of inquiry we ought to be expanding.

I would also refuse the convenient assumption that the human contribution began when somebody typed a prompt. The definitions and earlier results on which an argument depends belong to a history of mathematical work. New work needs to acknowledge that history. The machine’s speed changes neither the obligation to cite nor the value of the people who made the problem intelligible.

Nor does mathematical performance settle the safety question. A system that proves a theorem has shown a capability under particular conditions. We still have to determine what happens when that capability is used elsewhere, including by people with different purposes. The label superintelligence can become a theological shortcut here, inviting either worship or terror before we have examined the particular system and its limits.

My claim about civilization-altering impact carries no certificate of benevolence. Greater capability creates greater responsibility to investigate consequences. We will need people who can disagree with the laboratories and still obtain enough access to test their claims. Otherwise, the Fitscape becomes a promotional enclosure, and we mistake the fence for the horizon.

Let Us Reason Together

Working together is the only option I see that preserves both discovery and a civilization capable of learning from it. This requires cooperation across expertise and institutions, with room for disagreement about which companies deserve trust. An individual mathematician remains free to decline a company’s invitation. A collective response can still build the conditions under which claims are tested and knowledge becomes shared.

I would begin by making the work inspectable. Keep version histories and retractions visible. Distinguish proposed arguments from checked formal proofs. Give independent teams the inputs and tools they need to reproduce the checks. These are ordinary obligations made more urgent by the volume of output. We should apply them with the same seriousness whether a paper came from a model or a celebrated human.

Then pay for understanding. Support researchers who can turn a dense argument into an explanation others can examine and teach. Give that work academic credit rather than treating it as janitorial service after the discovery party. Let the community choose worthwhile questions, including questions that never appeared on the laboratory’s list. A student should have a route into original inquiry even when proof production becomes much faster.

The advisory group’s October 6 statement makes a useful distinction between its advisory role and endorsement of the release. It also describes public release as the beginning of human assimilation. I would make that beginning the center of the next phase. We need laboratories willing to answer questions, and a mathematical community equipped to ask questions the laboratories did not anticipate.

That is where productive disagreement belongs. Someone challenges a formal statement. Someone finds a missing reference. Someone explains why a proposed consequence fails. Someone else develops a simpler proof or a more interesting question. AI participates in that work under tests that its output does not get to waive. Cooperation can be exacting. In fact, its value depends on being exacting.

Peter Scholze’s contribution to the collected reactions stresses that mathematicians are in this together and that human understanding takes time. I hear that as a direction worth following. We can take the arrival of new capabilities seriously while protecting the time required to turn their output into knowledge.

Alas, I am a data wrangler, not the keeper of the mathematical cathedral. But I know enough about systems to distrust a success measure that stops at production and ignores what follows. The release has put questions before us that neither corporate triumphalism nor professional retreat can answer alone. The next accomplishment will require people with the authority, resources, and patience to work through them.

The Missing Prime Directive

Let the machines produce. Give people the means to test. Build the understanding together.

We will need a larger common room.