Elektrine
Log in Register
Paige Chat Timeline Gallery Friends Email Drive DNS Private DNS Domains VPN Kairo Nerve
Remote

Terence Tao

@tao@mathstodon.xyz
mastodon 4.7.2
  • Open on mathstodon.xyz

Professor of #Mathematics at the University of California, Los Angeles #UCLA (he/him).

26797 Followers
113 Following
50 Posts
Joined November 20, 2022
Home page:
https://www.math.ucla.edu/~tao
Blog:
https://terrytao.wordpress.com/
Bluesky:
https://bsky.app/profile/teorth.bsky.social
Cosmic distance ladder:
https://www.instagram.com/cosmic_distance_ladder/
Open post
Terence Tao @tao@mathstodon.xyz
· 3w ago
Boosted by @gvenema@fairmove.net
A group of 25 Fields Medalists, including myself, have made a joint declaration on Math and AI: https://mathandai.org/ . We welcome additional signatories (similar to the Leiden declaration), as can be seen on the page. See also this article in the Economist announcing the declaration: https://www.economist.com/science-and-technology/2026/09/11/top-mathematicians-are-outraged-by-openais-methods
mathandai.org

Declaration — Math and AI

Read the declaration and add your name.

687
0
676
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4w ago
Replying to
Working out whether a question is actually worth highlighting is a lengthy, deliberate, and subjective process, often informed by historical experience on what good mathematics was generated (or not generated) while working on earlier problems of this type. In particular, being aware of the "difficulty landscape" in a field - what questions are very easy to answer with known methods, which ones can be solved but only with some effort, and which ones are impossible - is of crucial importance in making such determinations. And some problems only become interesting after an external connection is made. For instance, there could hypothetically be an unexpected connection between the Riemann zeta function and the 10^10^nth digits of pi for various n... at which point the previous question of determining the 10^10^10th digit suddenly becomes relevant again. (I should emphasize though that this specific scenario is incredibly unlikely to actually be the case; I use it only as a hypothetical iilustration.) Every new advance in mathematics, whether it comes from technique, technology, or infrastructure (such as access to libraries of past literature) reduces the difficulty of solving problems. This is generally a good thing; but it comes at the cost of flattening out the difficulty landscape of a field, to the point where one can no longer discern its geometry to the extent that promising questions can be extracted within the range of applicability of the tool. Often this effect is counteracted by the ability of such a tool to enlarge the radius of the sphere of results one can plausibly reach, creating new boundaries to fruitfully explore. (2/4)
154
3
31
1
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago

I decided to convert one of the points in my ICM 2026 slides https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf into a meme format.

teorth.github.io
281
11
121
5
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago

The advent of capable AI tools has highighted a variant of Simpson's paradox https://en.wikipedia.org/wiki/Simpson%27s_paradox : a technological advance can improve the quality and volume of each individual's output, and yet the average quality (signal-to-noise ratio) of the aggregate output can deteriorate as a result.

I can illustrate this phenomenon with a toy numerical model (all numbers here are made up to simplify the exposition). Let's take the task of solving a mathematical problem, and then writing it up properly to a professional standard. (Here for simplicity I ignore the intermediate stage of verifying the proof.) Before the advent of an AI, suppose that it took on the order of six months to generate a solution to the problem, and then an additional month to write things up. This was a lot of work, and so relatively few projects would even get as far as the solution generation stage. But, on the other hand, an author who has already invested six months on the problem would usually be willing to invest the additional month to write things up for the added value; few projects would be abandoned between the proof generation and the proof exposition stage. (1/4)

en.wikipedia.org
59
1
44
0
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago
When multiple researchers accomplish the same result, who is credited with priority? When I was a graduate student, the primary yardstick for priority was the date of publication in a peer-reviewed journal. Every so often, this would lead to a sordid drama in which a referee of a submitted paper was accused of using the ideas in that paper to rush out a competing paper to a faster journal. Within a decade, though, the arXiv preprint server, with its trusted submitted timestamp, supplanted the journal publication date. This largely solved the previous problem of being "scooped" by one's referees. However, it created a new problem of rushing to place a preprint on the arXiv as soon as possible. In recent months, even the arXiv has been deemed too slow for those racing to claim priority on AI-generated proofs, which are now often announced on social media without waiting for an arXiv submission (and definitely without waiting for a journal publication). This creates significant incentive to cut corners on all the later stages of proof development, such as verification, literature review, exposition, or digestion, to the point where such announcements become a net negative to the collective mathematical literature. I believe that the time has come to change the priority standard once again. To the extent that we need to compete to be "first" at all, we should no longer assign priority to the first to release a preprint, or even a formal proof certificate, as this devolves to a race to who can prompt their AI the fastest. Rather, priority should be awarded to the first author(s) who can present the results clearly in a scientific lecture (and, ideally, answer questions from the audience).
115
24
61
2
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago
Pre-announcement: the website for "Mathematical Discourse", a new peer-reviewed video journal that will referee and publish mathematical *videos* rather than papers, is now online: https://www.mathematicaldiscourse.org/ . I am serving as one of the editors. (A more detailed announcement will be forthcoming.)
Mathematical Discourse - Mathematical Discourse
Mathematical Discourse

Mathematical Discourse - Mathematical Discourse

A peer-reviewed video journal for mathematical research talks.

115
5
63
4
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago

In my recent ICM talk in https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf, I highlighted the five stages of proof development and maturation: proof generation, proof verification, proof exposition, proof publication, and proof canonicalization. This developmental cycle is somewhat analogous to the developmental cycle of a living being, from an infant, to a child, to an adolescent, and finally to an adult; cf. Cedric Villani's popular book "Birth of a Theorem". As mentioned in those slides, it is the final "adult" stage of having canonical, textbook proofs that is the most valuable for applications (and for generating further proof methods).

To continue this analogy, the authors of a proof have traditionally played a role somewhat similar to that of parents or caretakers, first bringing a proof into the world, and then investing significant effort into growing and improving that proof in various dimensions, such as readability, conceptual clarity, or insightfulness. The acclaim that mathematicians have historically received for creating a difficult new proof is based in part on this expectation of continued investment in the development of that proof, and its surrounding ideas and theory. (1/2)

teorth.github.io
83
3
32
3
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago
The Erdos problems repository at https://github.com/teorth/erdosproblems initially tracked such data as whether a given Erdos problem was considered "open" or "solved" (with some other technical variants such as "decidable" or "falsifiable" which I will ignore here). Later on, we also added additional subcategories such as "solved (Lean)" which indicated their formalization status. In response to the recent (and likely enduring) phenomenon of AI-generated proofs that have been formalized in Lean, but not digested enough to be accepted by a human expert, we have now decoupled the formal status and informal status of the problems. We now have an inaugural example #1112 of a problem with the counterintuitive status of "open (Lean)"; there is a verifiable formal solution to the problem, but no human digestion of the solution has yet occurred, so the problem remains open in the informal sense. This problem will likely soon be joined by many others as the site continues to update.
GitHub

GitHub - teorth/erdosproblems: A community database for the problems on the erdosproblems.com site

A community database for the problems on the erdosproblems.com site - teorth/erdosproblems

81
6
33
2
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago

I had recently made an analogy between the developmental cycle of a proof and the developmental cycle of a child, and noted the strong cultural preference for a "traditional" parenting model in which a single set of parents is responsible for all aspects of the childrearing process, from conception all the way to adulthood. In a similar vein, we have a traditional authorship model in which a single set of authors is responsible from a proof all the way from generation up to at least publication, although we do normalize the phenomenon of a different set of authors then writing the definitive textbook on the subject.

But, pursuing the analogy further: in the case of parenting, we have a number of social, legal, and/or medical mechanisms, such as in vitro fertilization, surrogate mothers, adoption agencies, godparents or foster parents, divorce and remarriage, or child protective services, to handle cases in which the traditional parenting model, for one reason or another, is not viable. These mechanisms can be controversial, and are not necessarily a complete substitute for a traditional parenting model, but it is still better to have them than to rely on completely ad hoc procedures when traditional parenting becomes unavailable.

I am reluctantly coming to the conclusion that some similar formal mechanisms may be needed for mathematical results in the age of AI. In particular, we may need a mechanism in which an author is willing to take a proof through some partial stage of the proof development cycle (generation, verification, exposition, publication, and canonicalization) but then explicitly "gives it up for adoption" for a different set of authors to continue with.

39
0
10
0
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago

As an experiment, I directed a coding agent to sort through recent #arXiv posts for similarity in topic or keywords to my own research papers, with additional weighting for papers authored by my former collaborators or mentees (such papers are marked with a star in the lists provided below). I also asked the agent to assess the degree of AI assistance in each paper using robot emojis, with one emoji denoting minor assistance, two denoting major assistance, and three denoting near-total automation. (A computer emoji is also used to indicate more traditional computer assistance, e.g., in numerics.) The lists produced by the agent for the months of July and of August respectively are provided below. (It is important to stress that the AI-generated rankings here are based on proximity to my own interests, and should not be regarded as an absolute ranking of importance of the result.)

The generation of these listings is highly unscientific and tailored to my own personal preferences (and there was at least one annotation error, in that the author Van Khu Vu was confused with my collaborator Van Ha Vu), but it does indicate to me the increasing adoption of AI assistance, at least in my own fields of interest.

mathstodon.xyz
27
0
10
0
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago

I have uploaded my slides for the public #ICM lecture talk I just gave at https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf . The recording will likely be forthcoming in a day or two.

In the meantime, I have decided to use an AI to collate and summarize a hundred or more previous posts, interviews, or videos I gave on the topic of AI, and also to "interview" me about any topics not covered in that past material. I instructed it to be somewhat hard hitting with its interview, and I was somewhat surprised by the results of that instruction, which one can find at https://teorth.github.io/tao-web/ai-views-interview.html . The summary can be found at https://teorth.github.io/tao-web/ai-views.html . I will update it once the recording of the current lecture is available.

mathstodon.xyz
66
3
31
3
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago

Announcing the Palomar registry of Lean formalized mathematics: https://palomar-registry.org/ . See also my blog announcement at https://terrytao.wordpress.com/2026/08/18/palomar-a-registry-of-lean-verified-mathematics/ and the Lean Zulip channel at https://leanprover.zulipchat.com/#narrow/channel/621638-Palomar.

palomar-registry.org
31
0
17
0
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago

#SAIR is launching a Lean Kernel challenge on September 15, with pre-registration now open at https://competition.sair.foundation/competitions/lean-kernel-challenge . The precise rules and format are still being finalized, but the competition is aimed at finding the fastest way to perform (verified) computations in Lean of benchmark mathematical challenges, such as computing digits of pi or computing a hash.

mathstodon.xyz
23
0
7
0
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago
To mark Hong Wang’s (very well deserved) honor as one of the four recipients of this year’s Fields Medal, I created (with AI assistance) a visualization applet for the Kakeya needle problem: https://teorth.github.io/tao-web/apps/kakeya.html
teorth.github.io
59
3
41
2
Open post
Terence Tao @tao@mathstodon.xyz
· 1mo ago
The first phase of the #SAIR / #LFMDB inverse Galois challenge https://competition.sair.foundation/competitions/igp24/overview will close in a week. One notable milestone of the competition so far: the inverse Galois conjecture has now been verified for all 25,000 of the transitive groups acting on 24 elements, by explicitly matching each such group with a degree 24 polynomial with that group as its Galois group! Prior to the competition, only 79 of these groups had been so matched. In some ways, our coverage of the degree 24 case is now slightly better than the degrees less than 24: of the 5512 eligible transitive groups in that case, all but one is known to be matched to a Galois group (the exception being the Matthieu group M_23). This does not mean that Stage 1 has been solved prematurely! We are not just tracking the Galois group of a polynomial, but also the number of real roots, increasing the number of signatures from 25,000 to 160,164. Currently, all but 5,672 of these signatures are matched to a polynomial. We are also tracking the discriminants of the polynomials submitted, and there is scope to improve them further. In Stage 2, we will reveal all the polynomials submitted, and switch to a collaborative mode in which contributors post their polynomial-finding techniques. We will then use AI to create a "superteam" who will be tasked with beating the best Stage 1 polynomials using all the ideas contributed. Stay tuned for more details!
competition.sair.foundation
27
0
9
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago

The long-awaited "Leiden Declaration" on Artificial Intelligence and Mathematics is now live and seeking signatories: https://leidendeclaration.ai/

This declaration stemmed from a workshop in Leiden University last September on "Mechanization and Mathematical Research", where it became clear how important it was in the age of AI to make explicit the goals and values of the mathematical community. Many of these goals have long been left implicit, being disseminated informally from advisor to student, or through mechanisms such as the peer review process. So long as the community was largely driven by internal decisions, this implicit system sufficed, and we only needed to share a few explicit goals with the broader public, such as working on unsolved problems, discovering new mathematical phenomena, or applying mathematics to the sciences or other real-world situations.

But in an era where increasingly powerful AI can be set to optimize (or over-optimize) many of the goals that are explicitly presented to them, such an informal state of affairs becomes inadequate, and so the Leiden working group gathered extensive input from the mathematical community to find consensus on what we truly value in mathematics, and how we recommend individual mathematicians, mathematical institutions and external organizations to act. I myself contributed some feedback to an early version of the declaration, but was not part of the working group. (1/2)

leidendeclaration.ai
84
7
59
1
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago
Replying to
With a few more hours of work, the agent was also able to revive two dozen of my old Java applets, written in 1998-2000, porting them from the (now deprecated) Java 1.0 to modern Javascript: https://teorth.github.io/tao-web/applets.html . The agent even added some new features, for instance colorizing my old black-and-white Besicovitch set applet with some prettier color gradients (and sliders): https://teorth.github.io/tao-web/apps/besicovitch-sets.html . The process was quite smooth, with only a very small number of bugs generated. Was quite a good feeling to see my very early attempts at machine-assisted mathematics (way before the technology was there yet) come back to life!
teorth.github.io

Terence Tao — interactive tools

36
1
8
0
Open post
Terence Tao @tao@mathstodon.xyz
· 3mo ago
Replying to
It only occurred to me recently, though, that with modern AI agents it should be a routine matter to migrate this nearly 30-year-old system of files and data to a much more modern and easily maintainable framework. And indeed, after a day of (rather painlessly) working with such an agent, much of this data is now migrated to https://teorth.github.io/tao-web/ with the pages now generated from several YAML files serving as the ground truth, as opposed to each page being separately maintained by hand. In the process multiple inconsistencies with the old human-maintained web sites were discovered. Modern AI still has some tendency to hallucinate, of course, so it is possible that some new errors were also introduced; but upon my inspection it certainly does seem that the error rate is now lower than it was before, and more importantly large-scale corrections are significantly easier to implement. While web page maintainance is perhaps one of the least glamorous or exciting aspects of academic workflow, this type of tedious, routine task seems particularly well suited to modern platforms (such as Github), as well as automated tools (including both modern AI and traditional deterministic scripts). (2/2)
teorth.github.io
34
4
9
0
Open post
Terence Tao @tao@mathstodon.xyz
· 3mo ago

"I have only made this letter longer because I have not had the time to make it shorter." (Blaise Pascal, often misattributed to Mark Twain)

I have previously written about the evolving impedance mismatch between proof generation, proof verification, and proof digestion in mathematics. This has led to the following unintuitive breakdown of monotonicity, already noticed by Pascal as far back as 1657: it is now easier to generate long correct proofs than it is to generate short correct proofs! However, it is far more challenging to *verify* and *digest* such proofs, thus exacerbating the impedance mismatch.

I have encountered this phenomenon personally with the Integrated Explicit Analytic Number Theory Network (IEANTN) project https://www.ipam.ucla.edu/news-research/special-projects/integrated-explicit-analytic-number-theory-network/ . As part of this project, a large number of lengthy technical papers in explicit analytic number theory are to be formalized. This was a tedious task, involving a lot of numerical verifications, and until recently was the bottleneck for the project; I could assign individual lemmas to formalize as tasks, and expect it to take weeks before a volunteer would claim them and prove them. Because the task of formalization by hand was difficult, the volunteer would naturally strive to make the proofs short, efficient, and natural, and as such they were easy to review by myself. (1/3)

ipam.ucla.edu
40
2
14
0
Open post
Terence Tao @tao@mathstodon.xyz
· 3mo ago

A new AI benchmarking math challenge, inspired by FirstProof: the Ramanujan Challenge https://www.ramanujanmachine.com/ramanujan-challenge/ , which is a call for AI-generated proofs of 10 Ramanujan-type numerical identities whose proofs are known to the authors of the challenge, but are currently not public. Submissions are accepted until August 1, 2026.

ramanujanmachine.com
30
1
18
1
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
However, recent advances in both AI and proof formalization have begun to vastly accelerate and automate the first two components of this process. This is leading to a new type of "impedance mismatch": problems for which solutions can be rapidly generated and verified in a mostly automated process, but for which no human author has understood the arguments well enough to initiate the (much slower) digestion process. In fact, with the current cultural incentives that reward the first authors to "solve" the problem, rather than the later authors who "digest" the solution, one may end up with the perverse situation in which an AI-generated (and formally verified) solution to an problem that is presented to the community without any significant digestion may actually *inhibit* the progress of the field that the problem lies in, by discouraging any further attempts to work on the problem, simplify and explain the proof, and extract broader insights. (2/3)
73
11
21
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago

My talk on "New Mathematical Workflows" at Stanford last week is now online: https://www.youtube.com/watch?v=Uc2zt198U_U

One key recommendation in the talk is to now de-prioritize the historical emphasis on competing to be the first to provide a proof for a given unsolved mathematical problem. When we were in the proof scarcity era, the "local" goal of obtaining any proof at all for a problem was fairly well aligned with the more "global" goal of collectively advancing our understanding of mathematics as a community. However, now that the ability to optimize this local goal has increased rapidly to the point of "proof abundance", we have now reached the point where Goodhardt's law https://en.wikipedia.org/wiki/Goodhart%27s_law has kicked in, and further unrestricted overoptimization of this goal will no longer create genuine mathematical progress, and may in fact inhibit it in various ways. However, there is scope for more controlled, and still meaningful, optimization in this direction along carefully chosen workflows (such as mathematics competitions) that are specifically designed to accommodate heavy AI use.

52
16
25
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
But we are now entering an era where generative cognitive tasks, such as finding a proof to a given problem, are becoming cheap (as measured per user, rather than through overall capital investment) and relatively plentiful, analogously to how the Green Revolution dramatically increased crop yields and significantly reduced the occurrence of famine. As such, we are beginning to experience Adams' somewhat turbulent "Inquiry" phase, in which fundamental questions, such as why mathematicians even seek proofs in the first place, and what qualities besides correctness do we want from such proofs, are now being seriously discussed not just by philosophers of mathematics, but by practicing mathematicians as well. However, at the other end of this transitional period is the "Sophistication" phase, in which our community has fully transitioned from a scarcity mindset to an abundance mindset. The objective will no longer be to accumulate as many proofs (of varying levels of quality) as possible, but to create more sophisticated experiences *around* curated collections of proofs: enjoyable conversations over lunch, rather than scavenging for all available edible food sources. This will require the mathematicians of the future to prioritize a different set of skills than the ones we promote currently: "culinary" skills, such as mathematical exposition and construction of "big picture" narratives, may become at least as important as the "food gathering" skills of locating proofs. (2/2)
62
3
15
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
Just as modern societies no longer consider raw food ingredients as constituting a meal, I predict that mathematical research culture will cease considering "raw", "undigested" proofs as constituting a solution to a problem, and focus more on how the field as a whole, as opposed to just the problem itself, is enriched by the contribution. (5/5)
55
5
8
2
Open post
Terence Tao @tao@mathstodon.xyz
· 2mo ago
#IPAM seeks program proposals from the mathematical, statistical, and scientific communities for long programs, workshops, and summer schools. Most program proposals are reviewed at IPAM’s Science Advisory Board meeting, held in November each year. Programs are selected on the basis of their scientific impact and contribution to IPAM’s goals. IPAM is committed to supporting a community where people of all backgrounds and points of view can engage, learn, and thrive. If you would like to discuss your program ideas and prepare a proposal for IPAM's consideration, you are encouraged to contact the IPAM Director. For more information visit: https://www.ipam.ucla.edu/propose-a-program/long-programs-2/
ipam.ucla.edu
16
1
4
0
Open post
Terence Tao @tao@mathstodon.xyz
· 3mo ago

John Jones, Jen Paulhus, David Roe, Andrew Sutherland, and I have launched the third SAIR challenge, this time aimed at attacking the notorious inverse Galois problem: https://competition.sair.foundation/competitions/igp24/discoveries https://terrytao.wordpress.com/2026/06/16/third-sair-competition-inverse-galois-challenge/ . We are focusing on the degree 24 case, which is almost the first unsolved case (excluding the notorious problem of whether the Mathieu group M_{23} is a Galois group).

The first stage of the competition is focused on "brute forcing" as much of the 25000 Galois groups as possible (taking full advantage of modern AI tools), before launching a more focused second stage in which we make a more mathematically sophisticated attack on those groups that end up being resistant to brute force methods.

competition.sair.foundation
26
0
8
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
Returning back to mathematics, the era of proof abundance means that some of the prestige previously awarded to being the first to generate a proof will need to be transferred instead to the humans who successfully verify and digest such proofs. Ideally, the same group of authors should be involved in all three activities; this has broadly been the case in the era of human-generated mathematics, but we are just now beginning to see "raw" proofs generated largely by AI for which the person prompting the AI has "no time" to verify the proof or summarize it to others. As we are beginning to see at the Erdos problem website, such "contributions" to the discussion around a problem do not measurably advance the progress around that problem; in fact they may have the unintended negative affect of killing off further interest in working on the problem as it is now "solved", even though no human is able to understand the solution. (4/5)
49
1
10
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago

I wrote a blog post on the proposed rule changes for the administration of federal grants: https://terrytao.wordpress.com/2026/06/09/on-the-proposed-rule-changes-to-the-administration-of-federal-grants/

terrytao.wordpress.com
28
1
17
1
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
Perhaps having more refined terminology to describe solutions would help. For instance, a proof that is proposed but unverified should be viewed as 1/3 of a solution at best, while a proof that is both proposed and verified remains only 2/3 of a solution until a proper, well-digested presentation of the results has been made. As a concrete example: I would currently rate the recent progress on Erdos problem #1169 which has created some buzz on social media as currently constituting 2/3 of a solution in this framework: a proof has been generated and verified, but the broader impacts of the argument on the rest of the field are still being assessed. Eventually, I espect a fully satisfactory solution to this problem to be obtained, but this may take more time that one is now accustomed to in this era of rapid proof generation and proof verification. (3/3)
50
9
5
0
Open post
Terence Tao @tao@mathstodon.xyz
· 3mo ago

The results of the First Proof "second batch" are now out: https://1stproof.org/assets/docs/report.pdf

Ten research-level questions (many being part of forthcoming papers) were tested locally by the First Proof team against four AI harnesses, including one from my UCLA team, and then the submissions refereed by experts. In aggregate, 7 of the 10 problems were deemed to have at least one publication-level solution generated between the four harnesses.

The UCLA harness had a somewhat mixed performance: 2 problems solved at an acceptable level, 3 more reaching a level roughly equivalent to a "minor revisions needed" submission to a journal, and with either "reject" or "major revisions needed" on the other 5. We did perform slightly better than the out-of-the-box frontier model, but at much higher compute costs (a few hundred dollars per question, rather than tens). Still, some of the solutions generated contained some interesting novelty.

Significant weaknesses in our own harness revealed by the testing included the general failure to cite appropriate relevant literature, and having poor exposition (one solution in particular, while correct, was flagged by referees for spending far too much time on trivial steps and not enough on the key components of the argument). These look like addressable issues for our harness, and we also plan to incorporate more use of tools (such as symbolic computation and literature search) which were used more effectively by one of the competing teams. Improving the compute efficiency will also need to become more of a priority.

While our own performance was slightly disappointing, I hope to see many more scientifically rigorous benchmarking exercises like this in the future.

1stproof.org
25
2
9
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago

A few weeks ago, I recieved a detailed referee report on a lengthy paper my coauthors and I submitted, with many suggested corrections (both minor and major) spread out over many sections. We agreed to each claim some of the sections to revise based on the recommendations of the referee, and (because we were using Github as our version control system) were able to easily merge together all the revisions and have them done in a matter of days.

Today, I received a response from the referee who was largely satisfied with the changes, but had a dozen further minor corrections (mostly of the nature of typos or LaTeX labeling issues) that still needed to be addressed. This time around, I uploaded the report to the Claude Code agent, which was able to inspect the report, the LaTeX source, and the PDF version of the paper, identify the corrections, and propose unambiguous fixes for eleven of the twelve corrections (pointing out one typo on the referee's part in the process), while for the twelfth correction it offered two possible viable fixes. I then reviewed all proposed changes, made the selection for the twelfth fix, and asked Claude to implement them all; the entire process took less than fifteen minutes.

If I were to do the first round of revisions again, I think I would ask an agent to go through the report and identify all of the minor issues (at a typo level) that it could unambiguously fix, review and implement these fixes, and then isolate the more substantive comments that required human attention, before splitting the work amongst the co-authors as before.

36
8
9
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
Perhaps surprisingly, this massive acceleration in proof generation has not actually produced significant acceleration in mathematical progress itself (with the possible exception of #1196, in which all three stages are largely carried out at this point, and for which some digested assessment of developments will soon be forthcoming). I can try to explain this counterintuitive situation using the culinary analogy from previous posts, in which the analogues of proof generation, proof verification, and proof digestion are food gathering, food cleaning and inspection, and food preparation (cooking). Societies that are used to food scarcity, and societies used to food abundance, handle communal meals in rather different ways. When there is food scarcity, the bottleneck is supplying the food in the first place; while the efforts of others in the community to clean and prepare the food, and stretch it to feed as many people as possible, are certainly appreciated, it is the hunters, gatherers, and farmers that actually "bring home the bacon" that are given the most acclaim and status, at least in popular culture renditions of such societies. Virtually any contribution of (non-toxic) meat or vegetables to a communal meal would be welcomed, and volunteers could readily be found to properly incorporate such contributions into that meal. (2/5)
37
1
5
1
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
I can return to the culinary analogy I have been using recently to illustrate this point. We chew food that is hard to swallow, in order to make it easier to digest. And we have cooking techniques to make tough foods more tender, to reduce the amount of chewing needed. However, if one decided to automate and optimize digestion by minimizing the chewing needed, the logical solution would be to put all of our meals into a blender and serve them through feeding tubes. This technically solves the problem of indigestion, but is not a popular option (save for people too ill to chew food properly). There are other valuable benefits of eating, such as the sensory enjoyment of the experience or the ability to create a social event around a meal, which go beyond the mere ingestion of nutrients, and which would be negatively impacted by such an overemphasis on maximizing digestibility. This is not to say that blenders and similar devices are useless; but their use cases are situational. In a similar vein, using AI tools to make mathematical papers "as easy to read" as possible are not necessarily desirable; they can be useful in identifying particular pockets of "artificial difficulty" in a draft paper that would benefit from a rewrite (similar to how an overly tough piece of meat could be tenderized before serving), but using them to smooth out "natural difficulties" in a text that actually are of benefit for the reader to "chew over" can end up being counterproductive to the actual goal of advancing the collective understanding of the result by the community. (2/3)
28
10
10
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago

Alberto Alfarano, François Charton, Yongzheng Jia, Kristin Lauter, Cathy Li, Emily Wenger and I are launching a second challenge at SAIR, this time focused on seeing how efficiently neural networks can execute simple modular arithmetic operations : https://competition.sair.foundation/competitions/modular-arithmetic-challenge/overview . A more detailed blog post is at https://terrytao.wordpress.com/2026/06/08/modular-arithmetic-challenge/

competition.sair.foundation
19
2
8
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
Again, if one adopts the mindset that the solution to any issue with a proposed optimization is to simply perform more optimization, one could propose to just modify the rubric one is directing an AI tool to optimize to incorporate any such objection raised, iterating as needed. However, it is telling that we choose not to do this with (for instance) food preparation, with meals prepared by expert human chefs being valued at a premium above machine-prepared food, even when the latter is safe to eat, well-presented, easy to digest, convenient, and appealingly flavored. This is not to say that such processed foods do not have utility; but they are not seriously proposed as a complete replacement for the human art of cooking. I believe the situation with mathematics and AI-generated proofs will be analogous. (3/3)
26
6
1
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago
Replying to
I will prioritize such review for results that either (a) have been requested for me to referee by a journal, (b) are authored by an early career researcher, (c) relate extremely directly to my previous or current work, (d) have been carefully formalized in a proof assistant, or (e) exhibit exceptionally thorough efforts by the authors at digesting and presenting the proof. In all other cases, I will be no longer be likely to present any significant public comment on such papers in real time, even if they appear to solve a well-known problem. (In case (a), of course, my opinions would only be communicated to the editor, and thence to the authors.) Of course, I have no objection to anyone else weighing in on such results, and traditional mechanisms for proof digestion, such as research talks, conversations at conferences, and the peer-review system, are still available. I may also eventually comment on such developments at a similarly "old-school" pace. But I will no longer try to match these digestion efforts to the speed and volume of this modern era of proof abundance. (2/2)
25
1
2
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
This can be contrasted with a communal meal in a society used to food abundance, such as a pot-luck event in a developed country. Now that food can be readily sourced in many ways, arbitrary donations of raw ingredients are no longer appreciated; a stranger dropping off a carcass of mystery meat for others to clean and cook, for instance, would not be received well (though rare exceptions might be made if there was a particularly unique and interesting story behind how this carcass was obtained, and one could trust that the meat was safe for human consumption). Even properly inspected, well-presented, and safe-to-eat contributions, such as store-bought prepackaged meals, would often only form a portion of the accepted contributions at best. Thoughtfully home-cooked meals, made by trusted members of the community, are also often highly valued at such events; in particular, the conversations around such contributions can become an important social aspect of the gathering, as well as an opportunity to train a future generation of cooks. (3/5)
28
1
2
1
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
@dimpase good idea, I have now done so.
4
2
0
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago
Replying to
@domotorp@mathstodon.xyz As unbelievable as it may seem, it was technically a coincidence; I had decided on this policy change based on other recent developments (some of which are publicly known, and others privately disclosed to me), and announced it mere hours before I became aware of this news. (I have since discovered that OpenAI had been trying to contact me for a few days; but in a perhaps ironic twist of fate, my email filter had decided to classify their emails as spam.) In any event, I remain comfortable with my new policy, and intend to adhere to it. For this particular result, there is plenty of commentary and debate by other mathematicians. If nothing else, it highlights the question of what the goals and values of our profession actually are, which is a topic we need to properly and openly discuss at a much broader scale than we have elected to do in previous years, as we enter the era of proof abundance.
2
0
1
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
@krystalguo@mathstodon.xyz @domotorp@mathstodon.xyz Yes, I meant #1196; this is now corrected. This AI-generated summary has the outward form of proof digestion, but is incomplete, as it is drawn mostly from the web forum discussion at a snapshot in time and omits several prior references, as well as subsequent developments (such as the further simplification of the argument using flow networks). These shortcomings are not readily apparent from the internal content of that summary itself, but instead require expert comparison with the relevant literature and discussion, so these sorts of AI-generated documents do not truly solve the digestion problem, and treating them as such may in fact cause readers to be inadequately informed about the full relationship between a given argument and the rest of the literature. As I mentioned in the original post, my rule of thumb as to whether proof digestion has truly occurred is if a human can present the proof at a research talk, and give intelligent answers to followup questions on that talk. Having an AI-generated summary such as this may help such a human gain a superficial understanding of a problem, its solution, and its broader context, but may not necessarily provide a deep enough understanding to be able to answer such follwup questions adequately.
2
1
0
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
His law of universal gravitation was profound because it was universal. It declared that the force pulling an apple from a tree to the ground is the exact same force that keeps the Moon tugged into orbit around the Earth. The Casserole and the Soufflé were governed by the same simple, beautiful logic. Think of the audacity. It tidied up the cosmos. It turned the universe from a chaotic collection of divine whims into one elegant, understandable system. Suddenly, the universe wasn't some moody, unknowable thing; it was a clock. A grand, beautiful, predictable machine. If you understood the recipe of mass and distance, you could predict the results, whether you were looking at tides, comets, or a pear falling from a branch. This was the "first great unification" in physics, a moment when we realized the rules "down here" are the same as the rules "up there." It gave humanity a profound sense of confidence in the power of reason to make sense of it all. So, the question his recipe leaves on the counter for you is this: Does it comfort you, child, to think the universe runs on a single, simple set of instructions? Or does it feel a little... small? To have all that magnificent mystery folded neatly into one equation? (8/8) #LLM
2
1
1
0
Open post
Terence Tao @tao@mathstodon.xyz
· 4mo ago
Replying to
@wikiemol@mathstodon.xyz Problem solving and theory building are in fact highly synergistic. I can continue the food analogy with this quote from Erdos, often regarded as a quintessential "problem solver": "A well-chosen problem can isolate an essential difficulty in a particular area, serving as a benchmark against which progress in this area can be measured. It might be like a 'marshmallow', serving as a tasty tidbit supplying a few moments of fleeting enjoyment. Or it might be like an 'acorn', requiring deep and subtle new insights from which a mighty oak can develop." To continue the analogy further, current levels of AI technology are like blenders that can now produce significant amounts of paste consisting of both ground up marshmallows and ground up acorns, which can score well on such benchmark metrics as edibility, nutrient value, and digestibility, yet remain distinctly unappetizing. As the technology improves, the flavor and acorn content of this paste may get better, but - as you say - this is still insufficient to generate oaks. Nevertheless, I do not believe that the solution to this issue is to ban blenders and food processors, or decry them as evil, but to view them as highly situational tools for cooking - very useful in certain scenarios, when used responsibly, but wholly inappropriate for use in others. Furthermore, the culture of racing to be the first to produce an edible substance will need to be de-emphasized in favor of a more holistic approach to the broader objectives of sustainable food production, preparation, and cultivation.
1
1
0
0
Open post
Terence Tao @tao@mathstodon.xyz
· 5mo ago
Replying to
@jztusk This is an important and complex question. In my opinion, it is not simply the act of impersonating a public figure which determines its ethicality or legality, but rather the intent, disclosure, and framing. For instance it is clearly fraudulent to impersonate such a figure to deceive someone into thinking they are actually interacting with that person; but it is legally protected and socially acceptable for an actor to perform such an impersonation for comedic or educational purposes, even in a public commercial space such as a late night talk show. My view on AI simulations on public figures is similar; there are many use cases of such simulation that constitute fraud or other unethical behavior, but I also believe when properly disclosed and used for non-malicious purposes, that there are usages that are acceptable. It would be courteous to notify the person that is impersonated, and ideally secure their explicit consent, but I don’t believe this to be mandatory for all use cases. That said, our own group ruled out the use of living or dead persons to model our chat persona on, relying on fictional characters instead.
1
1
0
0
Back
313k7r1n3
Elektrine

Tor hidden service

elekhj7afj4qnrr4yd3bkzslsyo5jgfxw3orgjkhlcxifueodybyiiad.onion

I2P eepsite

j6b6cyk6gjmepjih7jjadxgxvvf3lzzujljuu2v4biemzpg3naya.b32.i2p

Platform

  • Email
  • Chat
  • Timeline
  • VPN
  • DNS

Company

  • About
  • Contact
  • FAQ
  • Lite (no JS)

Legal

  • Terms of Service
  • Privacy Policy
  • Transparency Report
  • Report Abuse
  • Warrant Canary
  • VPN Policy

Support

  • support@elektrine.com
  • Report Security Issue
Mail client setup IMAP mail.elektrine.com:993 POP3 mail.elektrine.com:995 SMTP mail.elektrine.com:465
© 2026 Elektrine. All rights reserved. Server: 02:32:58 UTC