AI Safety: A Short FAQ
for Mathematicians

Xiaoyu He

This is a follow-up to my essay about existential risk, in which I made the case that AI presents a substantial existential risk to humanity. The primary goals of this post are to suggest that:

  1. AI safety research is deep, broad, and mathematically interesting.
  2. It is possible to transition smoothly to alignment work as an academic mathematician without necessarily leaving your career.

1AI safety is just one problem. Even if it’s the biggest problem in the world, is it really productive to have many people work on it?

AI safety is not one problem in the same way that the Riemann Hypothesis is one problem; it’s an umbrella over a sprawling network of interconnected problems. To list several important factorizations off the top of my head:

Inner vs. outer alignment. Inner alignment is the technical problem of robustly encoding a target set of preferences into an AI; it is difficult because training procedures that select for outwardly aligned behavior do not necessarily select for deeply aligned internals. Outer alignment is the problem of deciding what set of preferences we want to encode in the first place, and is more of a moral philosophy problem.

Control vs. alignment. Control asks “How do we build superintelligences that stay subservient?” Alignment asks “How do we build superintelligences that support human flourishing (or at least don’t kill or torture us) even if they don’t stay subservient?” These are related but distinct research programs, and many researchers believe only one or the other is possible or ethical.

Monopolar vs. multipolar alignment. Depending on takeoff speeds and geopolitics, we may either end up in a monopolar world (one lab or state actor races ahead to superintelligence and dominates the competition) or a multipolar situation (several different actors stay neck-and-neck throughout the AI race and keep each other in check). Navigating multipolar endgames opens up a whole new set of mathematically difficult game-theoretic considerations.

Security and dissemination of magic. As I argued in the previous essay, AI progress may hand us a plethora of technological advances even before reaching true superintelligence. Some of these advances may be extremely dangerous and demand draconian and specialized security measures. Others may be enormously beneficial to society but still be catastrophically disruptive if handled carelessly. At a minimum, policy conversations would benefit from the inclusion of level-headed humans who can reason about shifts in supply and demand curves, instead of acting out of purely vibe-based Luddism or accelerationism.

Pause and technical governance. AI progress will come overwhelmingly fast if it is not slowed down by fiat. A substantial and effective pause requires civilizational coordination on the level of nuclear nonproliferation. We need a lot of people to engage in activism, and there are a host of technical problems to be solved on top of it (how do we monitor all the GPUs to make sure they’re not training new models?).


AI safety is everyone’s problem, the same way that World War II was everyone’s problem. Mathematicians in World War II made essential contributions to the bomb, cryptography, ballistics, operations research, and game theory. There will be an even broader spectrum of technical problems for mathematicians to contribute to in the critical window of AI progress.

2Isn’t AI research just ugly linear algebra and machine learning?

Alignment research is not confined to ML in the same way capabilities research is, and parts of it are as beautiful as anything I’ve encountered in pure mathematics. One of the most mind-bending areas of alignment research is the so-called “agent foundations” agenda, which seeks to build a mathematical and philosophical framework for understanding game theory and decision theory for AI agents, breaking several of the unstated assumptions of the classical theories. Examples:

Open-source game theory. What happens when the players in a game of strategy are not black boxes, but LLMs who have access to each other’s weights and can simulate each other's intentions and decisions? In toy models, Löb’s theorem from mathematical logic lets specially constructed agents reliably cooperate in the one-shot prisoner’s dilemma, and coordinate in other ways completely alien to human experience. For example, the agents could be sufficiently mathematically legible and deterministic that they can trust each other to cooperate if and only if they cooperate themselves, where trust is in the sense of a Lean proof. Such considerations may be important in understanding the emergent behavior of “AI agent swarms” or “hive minds” since future LLM-based agents may have access to levels of cooperation unthinkable to us.

Embedded agency. Traditional decision theory assumes a “Cartesian boundary”: an impassable boundary separating the “agent” that acts and the “arena” it acts upon. We may soon be breaking this boundary and building minds with deep and precise introspective access at the level of modifying their own neural connections one at a time, and similar access to the minds of future generations of AIs via recursive self-improvement. Even for humans, it is hard to draw an exact line between acceptable human self-tinkering, such as taking SSRIs for depression, and unacceptable self-tinkering, such as developing a heroin addiction. We still don’t collectively agree on whether it was okay for Erdős to take amphetamines. It is both philosophically and technically hard to reason about minds that are embedded in their own arenas and have the capacity to inspect and modify their own internals. Initial forays into embedded agency suggest that making embedded superintelligences stable may reduce to deep mathematical fixed-point problems.


Even for the mathematics closer to AI/ML research, which is relatively unaesthetic, I’ve found several of the guiding paradigms to be enlightening parables about the human condition:

Reinforcement Learning with Verifiable Rewards (RLVR). One of the new big things in AI research is RLVR, which increasingly supplements the older tech Reinforcement Learning with Human Feedback (RLHF) in domains like mathematics, where answer quality can be verified mechanically. RLVR is one of the reasons why LLMs are especially good at math now: mathematical progress is easily verifiable by itself and does not require laborious human data-labelling. As a mathematician, I identify with the RLVR model: much of the best mathematics I’ve produced happened in an empty room, judging my progress by myself, unthrottled by slow and noisy feedback from other human beings.

Iterated Distillation and Amplification (IDA). This is a proposal for bootstrapping weaker models into stronger ones, developed at the intersection of capabilities and alignment research. Roughly speaking, the idea is to alternate between two stages: have the model think slowly to produce outputs better than its reflexes (amplification), then train on those high-quality outputs until the better judgment becomes reflex (distillation). IDA is analogous to the way one trains a graduate student to write papers - the first couple of papers take a dozen agonizing rewrites and editing passes, with the goal that the student can successfully distill the lessons from these laborious sessions and write essentially polished drafts on the first (or third) try.


Another reason I’ve become more excited about working in this space recently is that LLMs can now think about all the matrix multiplication for us. It’s been claimed that the primary value humans can now add to both capabilities and alignment research is “research taste,” the ability to select the right questions and introduce the most aesthetic formalizations. There’s nobody with better taste than mathematicians.

3Can I possibly compete with the legions of ambitious youngsters who grew up in this era?

Mathematicians have a sort of deep learned helplessness about acting in the “Real World,” since we are trained to act within our narrow sandbox with very constrained imagination. When I was in grad school, for example, it occurred to me that if I wanted to maximize my contribution to mathematical progress, I would do better by being a personal assistant to my much more brilliant PhD advisor to free up his research time than by trying to prove theorems on my own. I know of someone else who was a full-time ghostwriter for other people’s papers who thus contributed more to the progress of their subfield than most of their peers.

Learned helplessness sounds like a bad thing, but it’s actually like spending your whole life wearing heavy training weights. My observation is that mathematicians who successfully take these training weights off can become superstars. Acting with agency and imagination in the Real World is complicated and does require practice, but in the end this is a relatively easy and primitive mental skill compared to algebraic geometry. Math is the language in which the universe is written, and in many parts of the world the actual hardest part of acting effectively is being good at math. This is why many of the top AI capabilities and alignment researchers have deep math/TCS backgrounds.

Let me also add that academic mathematicians have been partially responsible for (and complicit in) steering AI progress for a long time. Many of the most-cited benchmarks, including FrontierMath, FirstProof, and many others, were developed by mathematicians in collaboration with frontier labs, and their existence is likely partially responsible for the particular mathematical strength of current models. Math academia already has an ecosystem in place for designing evals to guide AI research; we should be using this ecosystem to advance alignment and the interests of humanity.

4OK, I’m convinced. What can I actually do to contribute?

There are numerous existing ways to dip your toes in the water. Here are a few off the top of my head:

Start a reading course. Is there anything mathematicians are better at than starting reading courses? Most PhD students and postdocs in mathematics are extremely worried about AI and their job prospects; it is essentially trivial to nudge them into starting reading groups and looking into alignment research, if only as a backup option in case mathematics goes kaboom.

Contact your local AI safety group. Most research universities with computer science programs already have active AI safety initiatives and student groups; you can likely do a lot of good cheaply by contacting your local group and helping them reach mathematics departments, where they are relatively inactive.

Look into short-term fellowships. There are plenty of reputable fellowship, visiting-researcher, and career-transition programs designed exactly to get outsiders up to speed on alignment research in a few months, without a career-ending commitment. Some such programs I (weakly) endorse are hosted by ARC, MATS, and Anthropic. Senior researchers can probably get even more mileage out of directly contacting organizations of interest.

Apply for funding. There are already hundreds of millions in yearly funding available for safety research, and it is projected that this amount will shoot up by an order of magnitude with the astronomical upcoming IPOs of certain AI labs. A surprising amount of such funding has already been captured by “one interested academic and their group.” Mathematicians may soon have an easier time being funded for mathematical research in AI safety via philanthropy than by the NSF.

Resources

Here are some places to start, whether you want to organize a reading group, try a research project, or take a few months to explore the field.

Getting oriented and starting a reading group

AI Safety Fieldmap: Overview of the field, including organizations, research areas, headcounts, budgets, and funding flows.

BlueDot Impact: Online AI safety courses and project sprints, with material you can use to structure a reading group.

AI Alignment Forum: Research posts and technical discussion across alignment agendas.

ARENA: Practical training and exercises for the empirical side of AI safety, including deep learning and mechanistic interpretability.

Fellowships and research programs

MATS: Mentored research across AI safety areas, including theory, with extension and residency pathways.

Anthropic Fellows: Four months of empirical AI safety research with funding and mentorship from Anthropic researchers.

SPAR: A three-month, remote research program with flexible, part-time participation and mentorship in AI safety and policy.

PIBBSS: Run by Principles of Intelligence; an approximately three-month interdisciplinary fellowship connecting researchers from fields such as mathematics, neuroscience, and philosophy with AI safety mentors.

AI Alignment Foundation Fellowship: An eight-week, remote, full-time alignment research fellowship.

Astra Fellowship at Constellation: Full-time research with separate empirical AI safety and strategy and governance streams.

Written with help from Claude Fable 5.1 and ChatGPT 6 Astra.