← Terence Tao

Terence Tao on AI in mathematics (and beyond)

A living, curated summary of my current thinking on AI, with a companion interview.

Last updated 9 September 2026.

A curated summary of Terence Tao’s current thinking on AI, with practical guidance and links to source material. Positions are distilled from ~80 Mastodon posts, some sixteen interviews and talks, six long-form essays and lectures, ~55 of his own blog comments, and a direct interview; not everything he has said appears here, by design — omission is editorial. Voice is third person; scope is confined to what is obviously about AI.

Latest: his September 2026 statement for the record on his views on AI and his interactions with the AI industry, folded in at the end of Part V. His fullest single synthesis remains the ICM 2026 public lecture “Mathematics in the age of AI” (July 24, 2026), written up as an essay for the Proceedings of the ICM 2026 and folded in at “The ICM 2026 argument” in Part V, with slides and a recording also linked.

How this page was made. This summary was compiled and drafted by an AI assistant (Claude) from Terence Tao’s public writing, talks, and interviews, then reviewed and corrected by him; the companion interview was conducted by that assistant, with his answers reproduced verbatim (lightly edited). It is a living document, revised as his views develop. Like the rest of tao-web, it is maintained with AI assistance.

Disclosure. Tao is not paid by any AI company, but has been gifted access to premium frontier models, collaborates with Google DeepMind employees, has helped organize an IPAM workshop sponsored by OpenAI (March 2026), has an IPAM project whose graduate students are funded by a Math Inc. donation, and co-founded an AI-focused non-profit (SAIR) through which he fundraises for AI-for-math activity. That sponsorship came out of emergency fundraising after the 2025 suspension of NSF and NIH funding; during the workshop OpenAI recorded an hour-long interview with him and used a few snippets in an advertisement, which he says in retrospect he should have pushed back harder on. He argues a tenured academic can still speak honestly, and that engaging the industry beats being “uniformly hostile”; his full reasoning is in the interview and in his statement for the record (Sep 2026).

See also a curated list of selected essays and talks by others on mathematics, AI, and proof assistants.

Part I — What AI is, in context

The latest step in a long automation of mathematical labor

Tao consistently frames modern AI not as a rupture but as the newest entry in a very old story: the human "computers" who built logarithm and trigonometric tables, Hendrik Lorentz's teams modelling a Dutch dam, computer-algebra systems, SAT solvers. Each wave delegated some routine layer while humans kept directing the machines. "Machine assistance in mathematics is far from new," he opens his Notices survey; "however, the scale and nature of such assistance is changing." He likes to historicize the anxiety, too: a hundred and fifty years ago a mathematician's chief usefulness was solving differential equations, and six hundred years ago it was building tables of sines and cosines for navigation — both now done by computer in seconds, and neither the end of the field.

Sources: Machine assisted proofs (Simons, Feb 2025) and "Machine-Assisted Proof" (Notices of the AMS, Jan 2025); LLMs and proof assistants as a millennia-long trend (Mar 26, 2025); Klowden–Tao, Mathematical Methods and Human Thought in the Age of AI (§2); The Atlantic, "We're Entering Uncharted Territory for Math" (Oct 2024).

Artificial general cleverness, not general intelligence

Tao doubts that genuine "artificial general intelligence" is within reach of current tools, but argues a weaker and still very valuable "artificial general cleverness" is becoming real: the ability to solve broad classes of problems by ad hoc, often brute-force or stochastic means that are fallible and uninterpretable, yet succeed at non-trivial rates when coupled with strong verification. For humans, cleverness and intelligence are correlated; for machines they are largely decoupled, and today's tools are best viewed as stochastic generators of sometimes-clever, often-useful outputs — which yields the characteristic "useful yet unsatisfying" feeling of a magic trick once explained. He resists a one-dimensional picture entirely: the space of cognitive tasks is extremely high-dimensional, so no "sub-human to super-human" scale captures it, and the frontier that still eludes machines is data-scarce reasoning — extrapolating from five or six facts and a vague analogy, exactly where human mathematicians excel. (He has also made the deflationary point that much of the fear is a branding artifact of the word "intelligence," and that AI progressed precisely once it stopped trying to mimic human thought — "we don't design cars and bicycles to walk like humans.")

Sources: Artificial General Cleverness (Dec 14, 2025); The space of cognitive tasks is very high-dimensional (Nov 26, 2025); OpenAI Forum (Dec 2024) and SAIR (Nov 2025) — paraphrased from video.

The shape of the tool

Where traditional software behaves like a deterministic function — reliable in its domain, nonsense outside it — an AI tool behaves like a probability kernel: a given input yields a random output concentrated near the ideal answer but carrying subtle, plausible-looking error, so it should be used interactively, not "click once and forget." Its competence is spiky: superhuman in places, capable of "hilarious" basic errors (asserting that all odd numbers are prime) in others. And its fluency outruns its substance — early output was "coherent English … but there was very little depth," and even now a useful mental image (from his lecture to students) is that "LLMs are like children that can present as competent adults for increasingly long periods of time." The recurring three-word summary is unreliable but powerful — a tool to be harnessed, not trusted raw.

Sources: AI tools as probability kernels (Mar 5, 2023); Klowden–Tao (§3, the "all odd numbers are prime" example); The Atlantic (Oct 2024); "How should university students control their AI diet?" (EMS Lecture, Jun 2026).

Why mathematics is where AI's successes are clearest

"In almost any other application, the biggest Achilles heel of AI is that it makes unverifiable mistakes," Tao told Nature. "But in mathematics, almost uniquely, you can automatically check the output" — at least when that output is a proof. That single property is why, in his reading, AI companies have recognized their "most unambiguous successes … are going to come from mathematics," and why the downsides of using AI in math are far more limited than elsewhere.

Sources: Nature, "'The job description is changing'" (May 2026).

How the view has evolved (2022 → 2026)

The through-line has tracked the technology, and it predates the LLM era. As far back as 2014 Tao was already predicting that mathematicians would one day "write our papers not in LaTeX, but in some language which some smart software will convert to a formal language," the computer throwing "a compilation error" wherever it "does not understand how you derived this step." In 2022 the first ChatGPT impression was fluent but hollow. By mid-2023 Tao had already committed to the forecast that would organize much of what followed — that "2026-level AI … will be a trustworthy co-author in mathematical research," once combined with formal verifiers, search, and symbolic packages — while insisting the real task was to navigate the transition "as safely, wisely, and equitably as possible." In late 2024 he called the o1 model a "mediocre, but not completely incompetent" research assistant; by early 2026 he judged that his co-author prediction had come in "almost exactly [on] schedule … on par with the contribution … a junior human co-author" makes. As he put it to Nature: "it's getting harder to deny that these tools can work" — and the community's own reaction, he says, runs through "the five stages of grief," with denial now beginning to fade. One thing that has not survived the acceleration is his own confidence in forecasting: the "relative certainty I had in 2023 of predicting the next three years … is gone now," the world is "far more unpredictable," and "I'm not sure anyone is capable of any reliable forecasting beyond a year at best, currently."

Sources: The Atlantic (Oct 2024, "coherent English … very little depth"; "mediocre, but not completely incompetent"); Tao, Embracing change and resetting expectations (Microsoft AI Anthology, Jun 2023, the 2026 prediction); The Atlantic (Feb 2026, "almost exactly the schedule"); Nature (May 2026, "five stages of grief"); Quanta (Jun 2026, reliably quoting the 2014 panel "compilation error" prediction); the companion interview (Jul 2026, forecasting humility).


Part II — When and whether to reach for it

The two questions: comparative advantage and acceptable failure rate

Tao's rule of thumb is not "is this task hard?" but a pair of questions: where does his comparative advantage lie, and what failure rate can the task tolerate? AI tools help least where he is most practiced — daily research mathematics, or writing email — and earn their keep in the middle band, on tasks he has some competence in but little practice at (data processing, translation, drafting a genre he rarely writes), where a machine first draft to verify and polish beats a blank page. Where he lacks expertise but stakes are low, the AI is a slightly more convenient search engine; where he lacks expertise and needs high reliability, neither the tool nor he suffices, and he consults a human expert. The second axis he stresses in its own right: often the deciding factor is not difficulty but acceptable failure rate — a weeknight dinner recipe tolerates failure; a state banquet does not. The deeper principle is one of complementary strengths: "AI is very good at converting billions of pieces of data into one good answer. Humans are good at taking 10 observations and making really inspired guesses" — an economic division of labour (he invokes Ricardo's comparative advantage) in which the sparse-data creative work stays human even where a costly model could attempt it.

Sources: Comparative advantage between human experts and AI (Apr 23, 2023); Difficulty vs. acceptable failure rate (Mar 29, 2025); The Atlantic (Oct 2024, "billions of pieces of data … really inspired guesses"); SAIR (Nov 2025, Ricardo's law — paraphrased).

Keep some friction

Tao argues the best level of automation is strictly between none and total — enough to cut tedious repetition at each scale, but with a human still in the loop to retain a sense of the whole. He draws a sharp line between artificial friction (tedious computation), fine to offload provided one can spot-check the output, and natural friction — genuine conceptual difficulty worth thinking through — which a tool that smooths it over, even by merely explaining a concept, can quietly rob from the learner. In an era of abundant AI-generated ideas, he adds, what matters is not raw idea count but good ideas times the signal-to-noise ratio of the idea pool: a bad idea can cost more time than it saves.

Sources: Optimal automation is between 0% and 100% (May 13, 2025); On the value of selective friction (Feb 22, 2026); "How should university students control their AI diet?" (EMS Lecture, Jun 2026).

A tool should be honest

Two design failings recur in his critique. First, confidence: an AI usually gives no indication of how sure it is, or flatly declares itself "completely certain," where a human would flag doubt — "AI tools do not rate their own confidence accurately. And this lowers their usefulness. We would appreciate more honest AIs." Second, autonomy: the industry's fixation on push-a-button, walk-away workflows is, for hard problems, "not ideal" — those want a conversation between human and machine. "We don't want to be reduced to just pushing buttons." Both point the same way: the useful failure mode is a legible one, and the useful interface is interactive.

Sources: The Atlantic (Feb 2026, "more honest AIs"; "pushing buttons"); AI tools need a clear "failure mode" (Aug 24, 2025).

The rule of thumb: use AI only where you could red-team it yourself

His most compact piece of practical advice, aimed at the individual: only rely on AI where you are able to red-team its output. That licenses using AI to red-team your own work (proofreading), or for blue-team tasks within your own power to check — literature search when you can follow and verify the references; a concept's history you can double-check; code you can read, run, and debug; a calculation you already know how to sanity-check. What to avoid is asking the AI for the answer to a problem you cannot solve yourself. His test: "if you would be unable to coherently present the output of the AI in a class presentation and be able to answer questions about it without further AI assistance, it should not be part of your workflow." (If you are stuck, the safer move is to propose your own strategies and ask the AI to critique them.) The same discipline governs reading: if you cannot yet read a paper at that level unaided, outsourcing the reading forfeits the chance to build the skill, so it is safer to use AI as an accelerated search engine (surfacing further literature, or stating the definition of an unfamiliar term — "as opposed to asking for a summary of the entire paper") and, as a comprehension check, to "first give the AI your own understanding of a particular portion of the paper you are reading and ask it to review whether your understanding was correct or not." A public example of this working style is the two-day conversation in which he digested the newly announced Jacobian-conjecture counterexample with a chatbot — driving the algebra himself while using the tool to symbolically check each step and surface related literature — which he links from the post reporting the result.

Sources: Tao, comment on Mathematical methods and human thought in the age of AI (Apr 18, 2026); Tao, comment on A digestion of the Jacobian conjecture counterexample (Jul 27, 2026, on reading papers with AI), and the shared chatbot conversation digesting that counterexample.


Part III — How to deploy it well

Verification is the filter that makes an unreliable tool useful

This is the load-bearing idea across all of Tao's writing: the most promising uses of AI "come from combining them with more traditional and reliable verification methods, in order to filter out hallucinations that would otherwise render the AI output useless." Hallucination stops mattering where output is cheaply verifiable. His vivid framing for students and general audiences is hydraulic — traditional research is a low-rate tap of clean water, AI a high-volume "firehose" of the undrinkable kind, and the whole game is building the filter; mathematics, where verification is best understood, is the natural place to build it. Or, flatly: "in math, we can completely check and verify outputs, and this really filters out a lot of the rubbish." The corollary is a firm rule of thumb, in his own words: "I would caution against using AI tools without the ability to independently verify their output. Relying on these tools to compensate for their own mistakes is quite risky and can amplify the weaknesses of such tools, such as hallucination, sycophancy, or lack of grounding." This is also why the human stays in charge of the loop — even in a system like AlphaEvolve that works well, the LLM only proposes mutations while "the verifier component … is primarily human-coded and not subject to hallucinations," and, more generally, "human experts remain the best metaprogram for these tools."

Sources: "Machine assisted proofs" (Simons, Feb 2025) and Notices (Jan 2025); IEEE Spectrum (Jun 2026, "filters out a lot of the rubbish"); Tao blog comments (Nov 2025, "independently verify"; the AlphaEvolve verifier; "human experts … best metaprogram"); SAIR (Dec 2025, "firehose" — paraphrased); Dwarkesh Patel (Mar 2026, "otherwise it's slop" — paraphrased).

Red team over blue team

Building a system (the "blue team") is only as strong as its weakest link; finding its flaws (the "red team") is additive. So unreliable contributors — AI included — are more safely deployed red-teaming (reviewing, testing, stress-checking human work) than in any blue-team structural role beyond what the red team can verify. The caveats: unreliable red-team output must augment, not replace, reliable members, and be triable; a flood of low-quality reports dilutes attention. Tao treats AI as a junior partner, not a replacement.

Sources: A stronger case for AI in "red teaming" than "blue teaming" (Jul 25, 2025); Klowden–Tao (§6.2).

Formalization and "trustless," industrial-scale collaboration

Formal proof assistants change how many people can work together. Ordinary collaboration requires personal trust and line-by-line checking, capping a project at around five people; a proof assistant's compiler removes the trust requirement — "you don't need to trust the people you're working with, because the program gives you this 100 percent guarantee" — enabling "factory-production-type, industrial-scale mathematics … like a modern supply chain." The Polynomial Freiman–Ruzsa formalization is his worked example: a 33-page paper broken into a blueprint's worth of independent nodes and formalized in three weeks by about twenty people, most of whom had never met. Because trust flows from verification rather than reputation, "an idea from an unknown researcher or even an amateur can be taken seriously if it has a formal proof" — and, as he told IEEE Spectrum, "maybe in the future, I won't even know if [my collaborators] are AI or real people." Two honest caveats. First, formalizing has cost several times the effort of writing — as of 2023–24 he put the "de Bruijn factor" at ~20 and "dropping," with "no fundamental obstacle" to falling below 1; that measurement has since been overtaken by rapid advances in autoformalization, which by late 2025–26 could formalize many steps in real time and had "essentially emptied" the queue of unclaimed formalization tasks on at least one project. Second, verification certifies the formal statement, not that it matches intent — so human review is reduced, not eliminated.

He is careful about two boundaries here: Lean is "a formal proof assistant rather than an automatic theorem prover" — it formalizes a proof a human already has and is "not all that useful in discovering new proofs" — and the emerging best practice divides trust so that humans author (or carefully review) the statements of theorems while automation handles the proofs, with "unit tests" attached to subtle definitions to catch misformalization. Encouragingly, the bar for leading such a project is modest: a mathematician needs "enough expertise to be able to state lemmas, if not prove them."

A May 2026 progress report on his Integrated Explicit Analytic Number Theory Network adds a useful refinement: formalization is not one activity but a spectrum of quality tiers, each tolerating a different trade-off between speed and the acceptable level of AI assistance — from reusable, human-interpretable libraries built to publication standards at the top, down to bare formalizations that merely certify a statement is true, are not meant to be read or reused, and can be produced with heavy automation. How much latitude to give the AI depends on the tier and the field: tedious, low-glamour, verification-heavy corners of the literature (explicit number theory among them) are exactly where one can safely hand the work to a machine, precisely because that labor does not compete with what humans actually want to do — whereas high-quality, reusable formalization still should not be automated away. He also flags autoformalization's most dependable use as an error-detector: models are too eager, and will cheerfully prove spectacular results from a subtly mis-stated hypothesis, so a sudden run of easy successes is itself the warning sign — which makes scanning repositories for such false statements a natural machine task.

Sources: The Atlantic (Oct 2024, "100 percent guarantee … modern supply chain"); IEEE Spectrum (Jun 2026, "AI or real people"; trust-through-verification); Notices (Jan 2025, the de Bruijn factor); Scientific American, "AI Will Become Mathematicians' 'Co-Pilot'" (Jun 2024); Tao blog comments (Dec 2023, "proof assistant … not an automatic theorem prover"; Mar 2026, statements-by-humans / "unit tests"); Quanta (Jun 2026, "state lemmas, if not prove them"); autoformalization now completes most tasks within hours (Jun 21, 2026, which has overtaken the 2023–24 de Bruijn estimate); ICERM talk, "The Integrated Explicit Analytic Number Theory Network" (May 15, 2026, progress report — tiered formalization and error-detection, paraphrased from auto-generated captions).

Measure it honestly

Tao repeatedly pushes back on hype-by-anecdote. Whether a task is "within AI ability" is not binary — capability spans orders of magnitude depending on compute, assistance, and how results are reported (his IMO analogy: change a contest's format and the reported success rate swings wildly). He therefore calls for standardized, pre-disclosed benchmarks measuring reliability and efficiency per unit of cognitive labor (as aviation moved to cost-per-seat-mile and accident rate), and flags the strong reporting bias against negative results: on the Erdős problems, as he measured it in January 2026, the true success rate was only a point or two — non-trivial in absolute count across a thousand-plus problems, but concentrated at the easy end, with no evidence at that point that the median problem was in reach. (The figure is a snapshot of a fast-moving capability, not a standing estimate; by September 2026 he was describing AI improvements on flagship problems.) He also warns that a benchmark is a good target only until you get close to it, after which over-optimizing overfits the tool away from real-world use.

Sources: Standardized methodology when evaluating AI at competitions (Jul 20, 2025); Reliability and efficiency per unit of cognitive labor (Jul 24, 2025); Reporting bias against negative results (Jan 17, 2026); SAIR (Dec 2025, benchmark overfitting — paraphrased).


Part IV — Guidance by audience

For students: control your "AI diet"

Tao's central analogy is nutritional. As food went from scarcity to abundance — trading famine for obesity and fitness atrophy — cognition is going from friction (tasks that required effort, and so supplied incidental mental exercise) to abundance; the subtle risk is not that AI is unreliable but that it becomes reliable enough to serve as a cognitive substitute. The education-specific harms he lists: cheating; deskilling (as calculators eroded number sense and GPS spatial sense — his own children, he notes, struggle with a non-interactive paper map); atrophy of problem-solving and even of patience; sycophancy that makes real criticism harder to accept; dependency and a corroded ability to trust genuinely authoritative sources; and a loss of intellectual diversity to a bland default style. He is blunt that some students' "grades are getting better and they're getting stupider." His prescription: strongly discourage unsupervised, unregulated access, but use AI for teachable moments — (1) emphasize process and verification over answers (discuss an AI output in class and ask how to fact-check it; require students to submit their prompts and their verification — even a wrong ChatGPT answer to critique); (2) allow the freedom to fail (iterative, imperfect-first projects, where the more the student supplies and the less the AI does, the better); and (3) allow creative AI use in small doses. The dose principle — AI as flavoring, "the vanilla extract of intellectual production," used sparingly — runs through everything. Meanwhile assessment adapts: in-person, AI-free exams have made a comeback as a deliberate stopgap.

He casts the whole challenge positively: teaching the next generation to keep a "mindful cognitive diet" is, he thinks, "going to be one of the great new purposes of our education systems." He concedes the "giant tempting 'cheat' button" makes this hard, and — pressed on it — that a real cost, unequally distributed, is likely; the disciplined and already-advantaged will stay sharp while others atrophy. But he does not think a hollowed-out generation is inevitable (food abundance did not turn everyone into a "Wall-E type sloth"), and he argues that inequality is not a fixed consequence of the technology: open and locally-run models, and "distilling" high-end performance into cheaper ones, can keep AI from becoming purely an engine of inequality — which he thinks public and philanthropic funding should prioritize.

The same reckoning reaches graduate training. The tangible tokens of a doctorate — a thesis, a paper or two — can increasingly be superficially duplicated by AI, to the point where an AI-generated artefact is getting hard to tell apart from four years of a student's work; but the purpose of a PhD, he stresses, was never to produce the thesis. It was to train the harder-to-see human skills the thesis used to stand in for — absorbing difficult material, synthesizing it, asking good questions — so the task now is to make those unspoken outcomes explicit and value them directly, rather than the tokens that can now be cheaply imitated.

Sources: "How should university students control their AI diet?" (EMS Lecture, Jun 2026); Klowden–Tao (§6.1, "vanilla extract"); Move to "open books, open AI" examinations (Dec 20, 2022); the companion interview (Jul 2026, cognitive diets as a purpose of education; inequality is serious but not inevitable; distillation and open models as an access lever); the Asian American Scholar Forum fireside chat (Aug 8, 2026, on a doctorate's value being the training rather than the thesis); The Futurology Podcast and SAIR (2025–26) — paraphrased from video.

For mathematicians and researchers

The near-term payoff is not the most powerful model on the hardest problem but medium-powered tools scaling up mundane, essential tasks — literature review above all — that a human expert could also do; that they could is a feature, because it makes the output verifiable and convertible to familiar forms. Formalization enables trustless collaboration and near-"instant refereeing"; AI handles routine calculation, code, figures, and even referee-report triage; and "vibe coding" a formal artefact can be responsible in one specific case — a statement already both formally stated and informally proven by humans. The creative core, though, is largely unchanged, and Tao is clear about a durable limit: breaking a problem into tractable pieces is where value lies, and "it's very easy to transform a problem into one that's harder … AI has not demonstrated any ability to be any better than humans in this regard." A practical footnote from Nature, as of mid-2026 and inevitably perishable as models change: in his experience ChatGPT made fewer mistakes and suited rigorous math but "writes … too robotic[ally]," Gemini "makes nice pictures" but was too wordy, and Claude was faster and "feels more human" — though "a lot of it is just the default prompting."

He has also spelled out the conditions under which he is comfortable letting an agent do essentially all the work — a checklist for when near-unrestricted AI use is safe, drawn from building visualization applets. His five favorable conditions: the task is not mission-critical (a small error rate is acceptable), the product is stand-alone (bounded technical debt, not entering a larger codebase or the literature), the end product is deterministic and sandboxed (plain JavaScript, no file/internet access, no run-time model calls — so no security, privacy, or ongoing-compute burden), it is not replacing a primary skill (he lets his JavaScript deskill but keeps Lean and Python in practice), and it is not competing with humans (no existing effort is duplicated). "I would however caution against unrestricted LLM use when one or more of the above five favorable situations is not in effect." (A sixth he adds: respecting and attributing prior art and intellectual property.) For the community's emerging norms on responsible AI and formalization, he points to — and has strongly endorsed — the Leiden Declaration (leidendeclaration.ai).

Sources: Near-term use cases: literature review (Oct 16, 2025); Responsible "vibe coding" for Erdős #707 (Oct 22, 2025); Claude Code for referee corrections (May 4, 2026); Scientific American (Jun 2024, "transform a problem into one that's harder"); Nature (May 2026, the model comparison); Tao, "Two more apps…" (Jul 16, 2026, the five favorable conditions); Endorsing the Leiden Declaration (Jun 2, 2026).

For everyone: where to be skeptical

The failure-rate and search-engine framings generalize: use AI freely where stakes are low and output is easy to check; distrust it where it sounds authoritative but cannot be verified. Its literature and cross-field suggestions remain unreliable enough that hallucinated citations are common. Its danger, Tao stresses, is producing the appearance of substance without the substance — riskier outside the sciences, which at least have a culture of objective verification; anywhere lacking such a culture, AI mostly amplifies existing problems (the loudest voices online). Its societal footprint deserves scrutiny too — who benefits, the environmental cost, and the risk of a "digital divide" between AI haves and have-nots — balanced against real benefits, with mathematics offered as the low-risk sandbox in which to study it all. And the answer to misuse is not prohibition ("you can't just ban food [to combat obesity]") but encouraging good practices, discouraging bad ones, and making disclosure of AI use routine rather than shameful. A concrete model he points to is the Erdős-problems site's policy: AI-assisted contributions are welcome provided the use is disclosed and the contents "have been carefully checked and verified by the user themselves without the assistance of AI" — disclosure plus independent verification, not a ban.

Sources: Three types of AI misinformation (Jun 2, 2023); Klowden–Tao (§5, costs/benefits and the digital divide); Tao, "The story of Erdős problem #126" (Dec 8, 2025, the disclosed-but-verified policy); SAIR and The Futurology Podcast (2025–26) — paraphrased from video.


Part V — The bigger picture

The ICM 2026 argument: a crisis of values, not of capability

His fullest synthesis is the ICM 2026 public lecture "Mathematics in the age of AI" — now written up as an essay for the Proceedings of the ICM 2026 (slides, recording; July 24, 2026) — which frames the whole subject as a second crisis in foundations. Where the crisis of roughly 1900–1930 (Russell's paradox, Gödel's theorems) forced mathematicians to make the foundations of reasoning explicit and ended by producing "an explicit, rigorous, and standardized foundational framework," the present turbulence is, he argues, "a crisis in the foundations of mathematical values and practices" — one that, examined and codified rather than resisted, will likewise leave the community "stronger and more resilient than before."

The lecture's defining move is to pull apart two questions usually tangled together. The first is what he calls the AI Capability Conjecture, a template with its key terms left as placeholders: "at some point in the near future, some AI tools will, at some expense, and with some level of human supervision, be able to accomplish some research-level mathematical tasks in some fields of mathematics, with some non-trivial success rate, and at some level of correctness and quality." Most public debate — his own writing included — has been about which "weak" or "strong" filling-in of that template is true; his talk deliberately sets that aside. He asks the audience to grant, conditionally, a Working Hypothesis that a "reasonably strong" version holds — "I am not asking you to want, believe, or accept that this hypothesis is true" — and then asks what actually matters: given that, what should the community do? On capability itself he offers only one data point, and a pointedly controlled one — the First Proof assessment, whose second batch of ten novel research-level problems was tested under scientific conditions against four AI harnesses in May 2026 and refereed for both correctness and exposition, with seven of the ten solved at publication-level quality by at least one team, at compute costs of $10–1000 per problem.

Conditioning on the Working Hypothesis brings into view a Goals and Values Question that mathematics had long delegated to the humanities: what are the profession's real goals — not only the explicit ones stated to the public and to funders, but the implicit ones pursued in practice? Historically its many goals (solving problems, building theory, understanding the world, training the next generation, creating work of aesthetic value) were positively correlated, so any one could stand in as a proxy for the rest and most could be left unstated. AI threatens that alignment through Goodhart's law — "when a measure becomes a target, it ceases to be a good measure" — to which, he warns, "the inherently ungrounded nature of generative AI, as well as the financial incentives of AI companies, make the use of AI tools particularly vulnerable." Over-optimizing a single goal, such as maximizing solved problems, can now pull it loose from the others it used to carry along. On the ICM panel the following afternoon he gave this danger a name — the risk of an "AI monoculture" — warning that because AI is powerful but favors one particular style of mathematics, funding and prestige can drift toward AI-amenable problem-solving and starve the slower, equally valuable work of theory-building and digestion; part of the remedy, he adds, is simply better outreach about what the profession's mission actually is.

Problem-solving is his worked case study: the naive goal "solve as many unsolved problems as possible" has to be repaired stage by stage — verified correct, then clearly communicated, then accepted by the community, then canonicalized into the definitive theory taught to the next generation — yielding the multi-stage pipeline analyzed in detail just below, in which AI accelerates each stage less than the one before and "impedance mismatches" pile up throughout. Exposition shows how the mismatch can hide: today's AI writing is near-flawless in grammar yet "dwells at length on trivialities" while hurrying past the genuinely novel steps, and even as it improves it risks becoming too slick — stripping the "natural friction" a human author's difficulty leaves in a proof, the passages that make a reader slow down, so the result is easy to read yet strangely hard to learn from.

From the goals he draws three concrete recommendations. Normalize the responsible disclosure of AI assistance — the case to avoid is covert use, concealed to dodge peer criticism. On the panel he pushed this past a checkbox: one can disclose the actual chat logs and traces, not merely state that AI was used — which asks mathematicians to grow more comfortable admitting the failed attempts along the way, a cultural shift he thinks especially valuable for students, who rarely see that the masters they study also stumbled on their first, second, and third tries. Decrease the emphasis on proof generation and on being "first," and raise it on the slower human stages of exposition, publication, and canonicalization. And a blunt gate on publication: "if the authors cannot convincingly demonstrate that they can give a clear, expert-level talk on their results, that is correct and properly attributed, then the result should not be published." Problem-solving, he stresses, is only one facet — theory-building needs its own analysis — and the same goals-and-values scrutiny now has to be turned on teaching, mentoring, hiring, grants, and outreach: emphasizing the human core where it matters most (above all in education), and otherwise defining best practices for these tools "on our own terms."

From the recording (paraphrased from auto-generated captions): to make the "natural friction" point vivid he holds up his own graduate-student copy of a 1991 Bourgain paper — annotated, in frustration, "I hate Jean Bourgain" — that he fought through with Stein's and Wolff's help and came to prefer precisely because its difficulty taught him how Bourgain thought, a lesson many "layers of AI improvement" would have sanded away. And, practicing the disclosure he urges, he notes that he used AI autocomplete in a few places and to generate the talk's diagrams, but wrote the prose (and the em-dashes) himself.

Sources: "Mathematics in the age of AI" (Tao's essay for the Proceedings of the ICM 2026, based on the Jul 24, 2026 public lecture — the crisis in the foundations of values and practices; the AI Capability Conjecture and the conditional Working Hypothesis; First Proof; the Goals and Values Question and Goodhart's law; the problem-solving pipeline and over-polished exposition; the three recommendations, including the talk-test for publication); the slides and lecture recording (video, paraphrased where it goes beyond the text); First Proof; the ICM 2026 "AI for mathematics" panel (Jul 25, 2026 — paraphrased from auto-generated captions; the "AI monoculture" framing, and disclosing chat logs rather than just the fact of AI use).

Proof abundance: from generation, to verification, to digestion

Tao decomposes mathematical problem-solving into three parts — generation, verification, and digestion (understanding, contextualizing, and explaining a result). Historically all three were hard and digestion arose organically as a byproduct, so the community rewarded generation and verification. AI and formalization now accelerate generation (and increasingly verification) far ahead of digestion, producing an "impedance mismatch": mathematics is moving from proof scarcity to proof abundance that its culture and infrastructure have not adapted to. His sharpest observations — that it is now easier to generate long correct proofs than short ones, and that faster generation has not produced faster mathematical progress — lead him to argue that prestige should shift toward those who verify and digest, and that the community should stop treating a raw, undigested proof as a finished solution. He has changed his own practice accordingly, sharply narrowing whose new proofs he will publicly digest in real time. He expects the profession to bifurcate along the same seam: routine proofs and calculations, and scanning many problems for "quick wins," will be offloaded to AI, while narrative-building and judging the promise of a new technique stay human — precisely because AI is beginning to decouple efficiency from the insight and training that used to come bundled with it (a distinction he draws in the comment thread of his Mathematical methods… post). (Borrowing Douglas Adams, he calls this the passage from a "Survival" phase of proof scarcity, through a turbulent "Inquiry" phase, toward a "Sophistication" phase of abundance.)

How does "digestion" get certified, if it is softer and more subjective than proof? Tao expects the community to build it the way it builds taste: journals, curricula, and publishers issuing detailed "style guides" and "rubrics" for well-digested writing, and — as food abundance refined cuisine — proof abundance producing "a much more refined taste … as to what constitutes a really good piece of mathematical writing." This makes mathematics "a 'softer', more 'subjective' subject," with "great debates … over what truly constitutes 'good' mathematics," which he welcomes as a healthy injection of "humanities-style" discourse; he likens the coming period to the early-twentieth-century "crisis in foundations," expecting the field to settle, after a tumultuous debate, on "a workable … foundation of mathematical practice, including a working definition of proof digestion." Pressed on whether his own large-scale "quick-wins" projects add to the very glut he warns about, he points to comparative advantage: he deliberately works the underexplored "long tail" that is "not in direct competition" with other mathematicians, invoking Thurston's warning that dominating a field can end up "killing" it.

By August 2026 he presses this toward a question of responsibility. Refining the three-part picture into five developmental stages — generation, verification, exposition, publication, and canonicalization — he casts a proof as a living being and its authors as the "parents" traditionally expected to raise it toward the "adult" stage of canonical, textbook understanding (the stage he considers most valuable for applications and for generating further methods). Against that backdrop he calls it "disturbing" to see proofs "abandoned at an intermediate stage," singling out cases where an individual or company uses AI to generate and verify a proof but then shows "no inclination" to develop it further — to give talks, or to go through peer review — leaving other mathematicians to volunteer as caretakers. He therefore proposes making an old, once-implicit norm explicit: any author "seeking to claim credit for generating a proof should commit to making their best efforts to develop that proof all the way to at least the publication stage," exposition and talks included. Recognizing that this commitment cannot always be honored, he pushes the parenting analogy to its mechanisms for when the traditional model is not viable — IVF, surrogacy, adoption agencies, foster and godparents, divorce and remarriage, child protective services — and "reluctantly" concludes that mathematics may need a formal counterpart: a way for an author who can carry a proof through only part of the cycle to explicitly put it "up for adoption," handing it to a different set of authors to continue, rather than leaving such cases to "completely ad hoc procedures." He notes these mechanisms are controversial and no complete substitute for the traditional model, but still better than having nothing.

He pairs this with a proposed change to how priority is assigned. The yardstick has drifted from journal publication date, to the arXiv timestamp, to — for AI-generated proofs — a social-media race that rewards "who can prompt their AI the fastest" and pushes authors to cut corners on verification, literature review, and exposition. He argues priority should instead go to the first author(s) who can present the result clearly — in his clarified wording, a "publicly available scientific exposition" (a lecture, but equally an expository blog post or recorded talk; he points to the new Mathematical Discourse journal as an early venue and hopes arXiv-like exposition repositories will follow). The exposition is meant to supplement the preprint and any formal certificate, not replace them, so the effective release date becomes the maximum of the preprint, exposition, and — once autoformalization is routine — formalization dates. Asked whether AI might simply write good expositions too, he says he would welcome that.

He has since sharpened the diagnosis into a variant of Simpson's paradox: a technological advance can "improve the quality and volume of each individual's output, and yet the average quality (signal-to-noise ratio) of the aggregate output can deteriorate as a result." A toy model makes the mechanism concrete. When generating a solution took about six months and writing it up about one, almost everyone who solved a problem paid the extra month, so the literature stayed roughly nine-tenths well-written and light journal curation sufficed. But if AI collapses generation from six months to a day while careful exposition falls only from a month to three weeks, the ratio inverts: output of both rises, yet raw solutions rise far faster than polished write-ups, so "90% of the literature will now be poorly written," and traditional peer review is overwhelmed. The natural rejoinder — that AI will get better at exposition too — misses, he argues, that "the AI tools will concurrently get even better at generation," so on current trends the ratio only worsens. His "reluctant" conclusion is that in the era of proof abundance the well-curated literature becomes "an increasingly small fraction of a much larger ocean of lower-quality (though often technically correct) results" — but, "by the same token," that shrinking curated core becomes "the most valuable data source" both for mathematics' own development and for its applications elsewhere.

Sources: Generation, verification, digestion (Apr 22, 2026); Proof abundance and "proof indigestion" (Apr 27, 2026); Long proofs now easier than short ones (Jun 21, 2026); A more restrictive policy on commenting on new proofs (May 12, 2026); Survival → Inquiry → Sophistication (Apr 20, 2026); the companion interview (Jul 2026, digestion rubrics / refined taste / a new "foundation of practice"; the long-tail rationale); the proposed norm: develop a proof you take credit for to at least publication (Aug 4, 2026); priority to the first to explain a result, not the first to generate it (Aug 6, 2026); clarified: a publicly available scientific exposition, supplementing the preprint (Aug 6, 2026); the emerging need for a formal mechanism to put a proof "up for adoption" (Aug 19, 2026); a Simpson's-paradox variant — AI raises each individual's output while the aggregate signal-to-noise ratio falls (Aug 27, 2026).

The journey, not just the destination

A recurring worry is that AI delivers the answer while skipping the value of getting there. "These problems are like distant locations that you would hike to," he told The Atlantic; the journey lets you "lay down trail markers … and make maps" that others build on. "AI tools are like taking a helicopter to drop you off at the site. You miss all the benefits of the journey itself. You just get right to the destination, which actually was only just a part of the value." The same instinct animates his enthusiasm for a new mode of work: math has only ever done intensive "case studies" of one problem at a time, but AI enables "population studies" — sweeping across thousands of problems at once — a genuinely new and complementary capability, not a replacement for depth. He is candid about the limits of that sweep so far: on a large problem set like the Erdős problems it clears the attention-starved long tail — problems posed once and never followed up — rather than the marquee problems mathematicians most want solved, where AI had, as of early 2026, yet to make real progress. That last qualification eroded quickly: by September 2026 he was reporting AI systems improving flagship bounds (bounded gaps between primes) and contributing substantially to a human-led finite-time blowup construction for the 3D Euler equations — which is what gives the concerns below their urgency.

The warmest form of the argument comes in a September thread built on Saint-Exupéry's dedication to Le Petit Prince — all grown-ups were once children, but few of them remember it. His example is the playground game of naming the largest number, through which children teach each other about infinity and, in effect, about proof: one of them eventually realizes they can defeat any guess by adding one. Ending the game early by lecturing them on the unboundedness of the naturals delivers the nominal goal and destroys the point of the play, "insofar as play has a purpose: to explore, to learn, to develop, and to simply have fun." Adults, he observes, come to treat goal-optimization as what anything serious must consist of, and confine play to their hobbies — with basic science, and pure mathematics above all, as the rare exception in which childlike wonder is put to serious use. Questions that look frivolous (the Navier–Stokes regularity problem, "if truth be told") reward precisely the combination of adult discipline and childlike fearlessness, and "the best outcomes arise when the journey is unhurried, convoluted, and serendipitous." The profession, he suspects in order to impress the adults watching and funding it, downplays that half of the craft in favour of tangible outcomes such as solving a designated open problem — a polite fiction that held only because human mathematicians used both faculties whether or not they admitted to the second. Tools aimed at the ostensible goal without expert supervision have no incentive to take the slow path: "The flag is captured, the goal scored, and the problem is solved; but at the cost of lessons learned, insights gained, collaborations formed, and new targets located." Hence the closing line: "In this modern era of heavy AI use, it is important to remember what it is like to be a child."

Sources: The Atlantic (Feb 2026, helicopter-and-journey; case studies vs. population studies); Machine Assistance and the Future of Research Mathematics (IPAM AI-for-Science kickoff, Feb 2026 — population studies and the attention-bottlenecked long tail; paraphrased from auto-generated captions); On remembering what it is like to be a child (Sep 9, 2026).

Answers versus insight

Underneath these worries lies a distinction that used to be idle. Faced with a problem X, the natural question seems to be "What is the answer to X?" — but in basic research the more valuable one is usually "What can be learned from studying X?": what the main difficulties are, which techniques must be invented, why the existing ones fall short, how X sits in the prior literature, what connections and follow-up questions it opens. Until recently the two questions were so closely aligned that there was no need to separate them, because the only practical route to an answer ran through those sub-questions. Pointing a powerful AI at X, unguided by anyone expert in the field, breaks the alignment: answers and insight have become "negatively correlated," so that optimizing harder for answers can decrease what the field learns — through contamination of later reasoning by the solution, a degraded signal-to-noise ratio, and weakened incentives to analyse the problem further. Nothing about AI forces the tradeoff (literature search is a well-established use that generates insight); the race for priority is what selects against the more deliberate uses. His example is a same-day one: three AI companies racing to announce improvements on bounded gaps between primes, numerically stronger than Julia Stadlmann's human-written paper, which he is glad she finished first — the lessons her writeup makes easy to extract (equidistribution estimates for numbers with a large smooth factor, the Polymath8b "epsilon enlargement trick", the Baker–Harman use of the Harman sieve) would have been far harder to recover had the AI proofs landed first, with no professional writeup and her effort abandoned.

The same divergence can be put as a resource argument. Even a magic button that turned an open problem into a solved one at no cost would not be a free lunch: answering a question can carry irreversible costs, the way a spoiler permanently diminishes a first viewing of a film, or a leak damages the fairness of a competition. Because a black-box model's genuinely novel reasoning cannot be told apart from recall, problems whose solutions are already public are "contaminated" and of permanently reduced value for evaluation — so "for the first time, human-generated open problems in mathematics in particular have become something resembling a non-renewable resource." The inverse Galois challenge he co-organizes is his worked example: running its first stage as a competition, with contestants withholding their polynomials, slowed the solving but yielded a map of the difficulty landscape (each Galois group's difficulty measured by how many contestants had found a polynomial for it); releasing them all for the collaborative second stage "permanently degraded the ability to crowdsource this type of difficulty map in the future" — a tradeoff he and his co-organizers judged worth making, but "not taken lightly." He points to an essay by Hugo Duminil-Copin — for whom a great conjecture is "a lighthouse in the night" that illuminates and guides a field, not to be profaned by use as a mere "benchmark" for the next frontier model — and adds an image of his own for the same worry: indiscriminate automated strip-mining of open problems for solutions is like using excavators to dig treasures out of an archaeological site, destroying the historical context that gave those treasures much of their meaning, and with it the ecosystem from which the next generation of techniques, problems, and practitioners would have developed. The tentative remedy — hard to enforce, he grants, though not unlike the social pressure that already discourages spoiling movies — is to declare certain classes of problems off-limits to automated solvers, while simultaneously opening up other classes as suitable targets for them.

He makes the cost concrete with the global regularity problem for the Navier–Stokes equations. Its value, he argues, is not the physical payoff — computational fluid dynamics is already mature, and a regularity theorem would not change how anyone models weather or climate — but the theory that attempts on it have generated: Leray–Hopf weak solutions, the Gagliardo–Nirenberg–Ladyzhenskaya inequalities, partial regularity, the Beale–Kato–Majda blowup criterion, and his own fluid-computation and Turing-universality program with its unexpected link to symplectic topology. The emerging consensus is that regularity fails, and there is even a strategy for exhibiting a blowup: a nearly self-similar ansatz, a numerical near-solution with a tiny computable residual, and a stability argument that perturbs it into an exact one. He thinks a heroic combination of machine-learning-powered simulation, interval arithmetic and formalization, and LLM-proposed ansätze, steered by human experts iterating on failed attempts, could plausibly get there — with a Lean verification "among the largest such proof artefacts ever created." But the artefact is not where the value sits: the insight comes from discovering precisely why each ansatz fails and adjusting, and that only works "if the iterator did not have access to the final ansatz in advance," since foreknowledge suppresses the instructive dead ends. Hence the failure mode he now considers realistic — an autonomous harness with vast compute running the entire iteration internally while the company keeps the process out of public view. One of the most prominent open problems in mathematics would be technically solved, with "almost no value added to mathematics as a consequence"; reverse-engineering some understanding back out afterwards would be far less efficient than having done it in the open. This is what he means by contamination being a net negative: pure mathematics poses its problems not because the answers are wanted in themselves, but because the pursuit reliably develops the field.

Days after that post the picture moved. Alpöge and Buckmaster announced finite-time blowup for several key fluid equations, 3D incompressible Euler among them, with a smooth forcing term — pushing the Córdoba–Martínez-Zoroa approach of iteratively adding small, localized high-frequency corrections that the low frequencies amplify — with the arguments formalized in Lean. Tao calls it a remarkable achievement, and is careful about the AI dimension: the arguments contain significant AI input, which the authors spent weeks reworking into an acceptable form, and Buckmaster explained the key ideas to him by telephone, "a refreshing change from AI-based communication modalities." He sees nothing in principle stopping the methods from reaching Navier–Stokes, and a non-negligible chance the forcing term could be removed; he adds that such an extension could probably be battered out by pouring in enough compute and AI assistance, but that "such an exercise does not particularly hold my interest" — digesting the method and extracting its new insights does. That is the same distinction the rest of this section draws, stated about a result he welcomes rather than a hypothetical.

The same counterfactual can be run backwards, through bounded gaps between primes — a problem where the numerical bound was never the point. The real timeline took two decades: Goldston–Pintz–Yıldırım in 2005, Motohashi and Pintz's observation that only smooth moduli needed improving, Zhang's 70 million in 2013 (found while he was an adjunct lecturer), a brief competitive race to shave it, the collaborative Polymath8 response, Maynard's new sieve — which cut the bound to 600 and became a textbook method with applications across analytic number theory — then Polymath8b's 246, a decade of equidistribution work, Stadlmann's 240, and, days later, the AI companies. Rerun that history with benchmark-focused AI companies present in 2005, he suggests, and the bound falls to the low triple digits within weeks on millions of dollars of compute; the problem is pronounced saturated and abandoned; no professional writeup is produced, and nothing from it is taught in a classroom or a textbook; the Maynard sieve lies buried in hundreds of pages of machine output until one of the few remaining specialists digs it out around 2010, keeps it secret to avoid triggering another frenzy, and abandons it by 2012; Zhang stays an adjunct lecturer and Maynard leaves the field. That timeline ends 2026 with more numerical progress and far less mathematics — Goodhart's law again, with Thurston's account of foliation theory being "evacuated" after he solved its major problems as the historical precedent. None of which, he stresses, is intrinsic to the tools: he points to the story of Erdős problem #126 as AI used collaboratively to advance both a problem and the field's understanding of it. What produces the bad outcome is a set of deliberate choices — not working with domain experts, not writing up papers, not submitting to peer review — in the service of a benchmark; a focus that was relatively benign in 2025, when the tools could not finish a problem unassisted, and that by 2026 has become "negatively aligned with mathematical progress as a whole."

The objection that open problems cannot be scarce, since infinitely many exist, he answers with an analogy: a region can suffer a critical shortage of drinking water while surrounded by ocean. Problems are trivial to generate — the 10^(10^10)th digit of pi — but almost none of them repay attention, and deciding which do is a lengthy, deliberate, subjective judgement that leans on knowing a field's difficulty landscape: what is easy with known methods, what is hard but reachable, what is out of range. Every advance in technique, technology or infrastructure lowers difficulty, which is good but flattens that landscape, and is usually offset because the same advance enlarges the reachable sphere and creates new boundaries worth exploring. What he finds distinctive about the present moment is the absence of any such boundary: AI has flattened the landscape across many areas without marking any frontier between the AI-feasible and the AI-hard, partly because the technology moves so fast and partly because the companies will not disclose their negative results or their process. So the scarce and precious resource is now the identification of a promising problem — and since even a rumor that someone is working on one can trigger a wave of AI effort to flatten it before that project matures, the incentives now point toward not sharing promising directions at all, which "would reverse centuries of traditions of open science and do serious long-term damage to the future of the field."

Several constructive responses follow. The first is a case study from his own work: revising the Equational Theories Project report sent him back to an observation buried in its Section 14.1. The ETP resolved 22 million implications between equational laws using automated theorem provers rather than LLMs — a job he thinks today's models would finish almost immediately — and in several cases an implication that had stumped the participants, forcing them to invent techniques for constructing infinite counterexamples, turned out to admit a small finite counterexample an ATP could have produced first. Had it done so, the more interesting method might never have been found; he suspects the project's early automated sweeps quietly obscured other fruitful problems, while conceding there is no counterfactual data to confirm it. The problem set is now solved, and "there is no practical way to undo this." The second response is prospecting: the ETP was deliberately chosen to open a new vein of problems rather than exhaust an existing one, and it scales — order-five equations would give 3 billion implications, pairs of order-four equations some 40 billion. But he is careful about how far that generalizes, since in most cases the interesting next questions are not produced by enlarging a solved problem's parameters; they are suggested by the process of solving it, which is precisely what purely automated solving discards. Hence a reframing of the competition itself: let the AI companies race to be first to announce a new mathematical insight, rather than first to announce a solution. And where prohibition is infeasible, he would settle for explicit standards: designating classes of problems as wanting an analysis that identifies insights and maps the difficulty landscape nearby, for which a raw solution without that analysis counts as negligible or negative value — the way a modern food drive no longer accepts any technically edible donation, but publishes what it actually wants.

Sources: On open problems as a non-renewable resource (Sep 2, 2026), responding to Hugo Duminil-Copin, Care for a little more AI? (Proofs and Prompts, Aug 30, 2026 — the lighthouse image); Navier–Stokes as a worked example (Sep 3, 2026); Answers versus insight (Sep 3, 2026); A counterfactual timeline for bounded gaps between primes (Sep 5, 2026), whose collaborative counterexample is The story of Erdős problem #126 (Dec 2025); A competition to be first to a new insight (Sep 5, 2026); The Equational Theories Project as a case study (Sep 7, 2026, on §14.1 of the ETP report); Alpöge–Buckmaster on finite-time blowup for 3D Euler (Sep 8, 2026); Drinking water beside an ocean: scarcity among infinitely many problems (Sep 8, 2026).

A new way of doing mathematics — "big mathematics"

Tao's positive vision is large-scale, decentralized collaboration between humans and machines — what he calls "big mathematics" — in which complex tasks are diced and sliced, humans claim the creative parts, and AI does "the lion's share of the technical grunt work." He reaches for industrial analogies: mathematicians have historically worked "like a craftsperson … [making] one toy at a time," where AI enables "factories"; the shift mirrors software engineering's move from the lone hacker to specialized roles (project managers, quality assurance, formalizers). He expects the definition of a mathematician to broaden accordingly, and new roles to appear — including a profession that takes ugly machine-generated proofs and makes them humanly comprehensible, and a "citizen mathematics" in which undergraduates, high-schoolers, and interested amateurs contribute modular pieces (finding references, running numerics, checking a step) to projects like the Erdős problems site.

Sources: IEEE Spectrum (Jun 2026, "big mathematics"); The Atlantic (Oct 2024, craftsman-to-factory and the universal-algebra "terra incognita"); Scientific American (Jun 2024, new roles); OpenAI Forum and Math Inc. (2024–25) — paraphrased from video.

The human place: what mathematics is for

Tao's stance is neither triumphalist nor dismissive: for the near term "we will still be driving," with AI accelerating the exploration while humans choose what matters — and, in a reframing against the "lone human vs. the machine" picture, the mathematical community is already "an incredibly super intelligent entity that no single human mathematician can come close to replicating." Longer term he offers a cognitive Copernican principle: human intelligence is not the privileged centre of the cognitive universe; human and artificial intelligences belong to the same category, distinct and complementary.

But pressed on what the profession is for if the creative core itself eventually falls to machines, he does not reach for the (in his words) "over-used" chess analogy. He reaches for Thurston: mathematics "is not about fulfilling an abstract quota of definitions, theorems, and proofs" but is about understanding — to which, he adds, we should now append the word human. His preferred analogy is music: "we can electronically reproduce music with perfect fidelity, and … even … generate new music … stylistically indistinguishable from human performances; and yet we still value in-person concerts, because we value human connection." Mathematics, he says, "has actually always been human-centric at its core," a fact the profession has hidden behind "the conceit that math is just about the impersonal abstract world of numbers and equations" and now needs to be "more honest about." He also draws a hard boundary he doubts AI will soon cross: solving a problem is the highly verifiable game AI is suited to, but getting a genuinely new idea recognized, digested, and found exciting by the community is "a far different game" — closer to writing a blockbuster novel or a hit song, tasks where, he notes, AI has shown little progress.

Sources: Scientific American (Jun 2024, "we will still be driving"); the Lex Fridman transcript (Jun 2025, the community as an existing superintelligence); Klowden–Tao (§6.3–6.4, the Copernican view); the companion interview (Jul 2026, Thurston / "human understanding," the music analogy, and "gaining community acceptance is a far different game"); An analogy between AI and the automobile, and "AI planning" (Mar 18, 2026).

On AI risk

Asked to rank AI's dangers, Tao inverts the usual sci-fi ordering. His near/medium-term ranking runs, from least to most worrying: "'Autonomous AI malfunction' ≪ 'Humans using AI incorrectly' < 'Socioeconomic disruption caused by AI adoption' < 'Beneficial uses of AI shut down due to AI panic' < 'Malicious humans' < 'Malicious humans assisted by AI'." The practical lesson is to resist dread-risk bias — the pull toward catastrophic tail scenarios at the expense of the frequent, medium-sized risks (an AI-assisted attack on critical infrastructure, say) where most risk-management effort should actually go; building resilience there also builds the "antibodies" for the rarer tail risks. He is pointedly skeptical of importing expected-value calculations into existential-risk debates: such reasoning is only useful with both an accurate probability model and enough repeated trials, and "in situations where one doesn't have both — such as Pascal's wager, or weighing AI existential risk — I would not recommend placing too much weight on such an analysis, despite its 'mathematical' appearance." On misinformation he makes a dual point: reducing false positives (watermarking AI output) is a losing game against a system built to pass any such test, so the more tractable and equally important aim is reducing false negatives — cryptographic provenance that certifies genuine content. Through all of it the frame stays resolutely human-centered: the AI-risk question is ultimately about people — who is harmed, and who is protected — rather than about the machines in the abstract.

Sources: Tao blog comments (Jun 2023, the risk ranking; Mar–Apr 2026, dread-risk bias and the expected-value caution); On deepfakes: false positives vs. false negatives (Jan 30, 2024); Klowden–Tao (§5, costs and the digital divide).

The meta-question: rethink it ourselves, or a company will

Underlying all of it is a call to agency. AI, Tao told Nature, "is not just another technology like the word processor or the web browser. It really is forcing us to rethink fundamental questions — what is a mathematical proof? What is a paper? What is the purpose of our profession? If we don't ask these questions ourselves, then they will get answered for us by a technology company or decided by financial incentives. We have to get ahead of this." It is the same conviction behind the Klowden–Tao insistence that AI's development stay "human-centered" — that AI be judged "not purely through the technical lens of what microscale problems [it] solve[s] … but also through the macroscopic humanitarian lens of how our society, our shared body of knowledge and understanding, and our species benefits (or is harmed) as a whole" — and behind his endorsement of the Leiden Declaration making the community's long-implicit values explicit; the recurring refrain is to navigate the transition "safely, wisely, and equitably."

His prescription for the discussion itself is engagement, not boycott: "it is still more effective to engage with them to move the cost–benefit balance sheet of the technology in a better direction, than to be uniformly hostile," because "incentives work both ways." But this, he insists, cannot be left to a few spokespeople — the topic "is too important to leave to just a small number of spokespeople," it needs "genuine grassroots debate that brings all stakeholders into the discussion, especially the younger generation," and, strikingly, his own "influence on this discussion should diminish as the rest of the community steps up." He also flags a public misconception the profession urgently needs to correct: that mathematics is "primarily about solving open problems by whatever means possible, even if the resulting proofs are incomprehensible to humans." And he frames math's role for other fields precisely — mathematics supplies "the theoretical upper bound for what the maximal safe level of AI use can be," the idealized best case, which law, economics, and the humanities must then discount for the messier, less verifiable real world.

By the 2026 International Congress of Mathematicians he had sharpened this into a call to act, and quickly. He frames the moment as a genuine crisis — the field's first in more than a century — but, unlike the foundational crisis of the early 1900s, "not one of mathematical arguments, but rather values and practices"; mathematicians, he argues, have only months to "reimagine what it means to be a mathematician," and "I wish we had more time to do this slowly." The remedy is uncharacteristically activist: "We don't have to passively accept changes by external forces. [We] can't just passively prove our theorems; we have to organise, become activists, get a little political." His diagnosis is that the profession has, for the first time, "temporarily lost the narrative" — with proof generation accelerated, "validation is in short supply," and the danger is that "the figure of the mathematician will instead be defined by technology companies." His answer is not to out-generate those companies but to get louder about the parts of mathematics their tools cannot touch — exposition, digestion, community acceptance, taste: "They tried to redefine what our profession is. [But] mathematicians have more influence than they think."

In September 2026 he set the whole arc down for the record. He identifies with neither a "pro-AI" nor an "anti-AI" position. By 2023 he judged that LLMs and proof assistants could be transformative enough that maintaining traditional mathematical practice and culture without adaptation would become unsustainable — a culture he values and has been "a massive beneficiary" of — and he saw two viable paths from there. Either the tools are responsibly incorporated into mathematical workflows and culture, the "best of both worlds"; or they are used indiscriminately for short-sighted objectives "at the cost of the far more valuable long-term sustainability of mathematics and its role in the scientific ecosystem" — the worst case, which he now says "we are unfortunately rapidly approaching." That is what he has spent a growing share of his professional life working against: raising awareness of the scale of the change, building examples of the good path, and warning against irresponsible uses. The warning itself has moved. Use without adequate verification of outputs was his main target in 2023–25 and is, he notes, "no longer the primary vehicle for harm in 2026" — that role now belongs to the indiscriminate, benchmark-driven extraction described earlier in this part. His judgement of the engagement is unchanged: he does not regret it, and "one could certainly call this effort 'shilling for AI' if one likes; but I would say that this is overly reductive." What has changed is his read of the industry — many of the people there who shared his view "have left or become sidelined," with the major companies now focused on "the race to develop extremely powerful, autonomous AI technologies regardless of their actual value to society," of which he calls the Navier–Stokes affair the most dramatic and visible instance. So the immediate priority he now names is not further engagement but consolidation: "for the entire mathematical community to unite around our core values and objectives, and reject irresponsible and unsustainable usages of AI technology that only serve to advance nominal goals rather than the true underlying goals of the field."

Sources: Nature (May 2026, "rethink fundamental questions … get ahead of this"); Klowden–Tao (human-centered thesis); Endorsing the Leiden Declaration (Jun 2, 2026); Embracing change (Jun 2023, "safely, wisely, and equitably"); the companion interview (Jul 2026, engagement over hostility; grassroots debate; math as the theoretical upper bound; correcting the "solving by any means" misconception); New Scientist: "Why mathematician Terence Tao thinks AI must spark a rapid revolution" (Aug 6, 2026 — from the ICM 2026: organise/become activists; a crisis of values and practices; "temporarily lost the narrative"; more influence than they think); A statement for the record on his views on AI and his interactions with the industry (blog comment, Sep 9, 2026).

In closing — embrace the complexity

Asked what he would most want emphasized, and what people most get wrong, Tao's answer was a plea against one-dimensionality: "AI is a truly complex topic, and requires thoughtful, nuanced discussion … it is very tempting to try to simplify it by having one-dimensional narratives such as 'AI good' or 'AI bad'. But the topic is so much richer than that … it is one of the most important issues of the current era, and it is really worth it for everyone to really dig into and embrace the complexity and paradoxes of the situation."

Source: the companion interview (Jul 2026).