Subscribe to Events

Download as iCal file

Colloquia

What should we do about AI generated proofs?

Mohammed Abouzaid

Location:  Hill 705
Date & time: Wednesday, 06 May 2026 at 3:30PM - 4:30PM

The three large AI companies (Google Deepmind, OpenAI, and Anthropic) have decided to compete on the capabilities of their models to "do research math," have invested substantial amounts of money on improving these capabilities, and are for now subsidising their use by users (some more than others). I will focus on the specific case of using Large Language Models (and systems built on them) to produce natural language proofs (i.e. the kinds of proofs that humans write and read). I will describe the preliminary outcomes of trying to make an unbiased measurement, which indicate that the most advanced systems are capable of producing correct argument for many statements whose proofs do not appear in the literature (arxiv:2602.05192). I will also discuss efforts currently under way to make more refined assessments (https://1stproof.org/index.html. Click or tap if you trust this link." data-auth="NotApplicable" data-linkindex="0">https://nam02.safelinks.protection.outlook.com/?url=https%3A%2F%2F1stproof.org%2Findex.html&data=05%7C02%7Cmpy4%40connect.rutgers.edu%7C03573a71433a4229499808dea8ce42a3%7Cb92d2b234d35447093ff69aca6632ffe%7C1%7C0%7C639133801027528483%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=bafJYuKgo70xRfORDNKCpqpjqIghgfNZl4uQdeq7tw0%3D&reserved=0). I will aim to leave ample time for comments, questions, and discussion.