Subscribe to Events
Auto-formalization via Joint Embeddings
Vijay Ganesh
Location: CoRE 431
Date & time: Wednesday, 19 November 2025 at 2:00PM - 3:00PM
In recent years we have witnessed a symbiotic trend wherein LLMs are being combined with provers, solvers, and computer algebra systems, resulting in dramatic breakthroughs in AI for math. Following this trend, we have developed two lines of work in my research group. The first is the idea that "good" joint embeddings (JE) can dramatically improve the efficacy of LLM-based auto-formalization tools. We say that JEs are good if they respect the following invariant: semantically-equivalent formally-dissimilar objects (e.g., pairs of sematically-equivalent natural and formal language proofs) must be "close by" in the embedding space, and semantically inequivalent ones "far apart". We use such JE models as part of a successful RAG-based auto-formalization pipeline, demonstrating that such JEs are a critical AI-for-math technology. The second idea is Reinforcement Learning with Symbolic Feedback (RLSF), a class of techniques that addresses the LLM hallucination problem in contexts where we have access to rich symbolic feedback such math, physics, and code, demonstrating that they too are critical to the success of AI for math.