Probabilistic reasoning · AlphaGo · Frontier AI

Thore Graepel

I build machines that reason and act — from weighing the odds at web scale, through AlphaGo, to a new project that brings AlphaGo-style reasoning to frontier AI.

Chair of Machine Learning, UCL

AlphaGo vs. Lee Sedol · Game 2, 2016 · Move 37 · AlphaGo's shoulder hit at P10
The AlphaGo story →
00

Reasoning, from factor graphs to frontier AI

Thore Graepel

I have spent a working life teaching machines to reason under uncertainty and to act in the world. I trained as a physicist — in Hamburg, at Imperial College London, and in Berlin, where I earned my PhD in machine learning in 2001. At Microsoft Research from 2003 I helped found the Online Services and Advertising group, where reasoning under uncertainty shipped at scale: TrueSkill, the rating system behind matchmaking on Xbox Live, and AdPredictor, the click-through model behind Bing. Both rest on message passing over factor graphs.

At DeepMind I came back to my first love — building thinking machines — and had the luck to work on AlphaGo, the first program to beat a professional at the full-sized game of Go, and on its heirs AlphaGo Zero, AlphaZero and MuZero. There, learned reasoning met deep search: a policy that suggests, a value that weighs, a tree that looks ahead.

Much of my work since has dealt with multi-agent systems — from Capture-the-Flag agents to simulated humanoid football — where intelligence takes on a body, works alongside others, and grows through self-play. After leading machine learning for cellular rejuvenation at Altos Labs and a spell on Google DeepMind's Post-AGI team, those threads now meet in a new project: bringing AlphaGo-style reasoning to frontier AI, so that machines can plan and act under uncertainty. Search and learned judgement care little where the uncertainty springs from — a game tree, a cell, an agent choosing what to do next — so the same reasoning reaches on to embodied intelligence, and the multi-agent work above gives me grounds to think it will hold. Alongside this I hold the Chair of Machine Learning at UCL.

01

Probabilistic reasoning

Long before deep networks, I worked on machines that reason with uncertainty — Bayesian models that hold beliefs, update them from evidence, and scale to the web through approximate message passing on factor graphs. Ranking players, predicting clicks, recommending, and — a decade before AlphaGo — predicting the next move in Go.

Microsoft Research · Xbox Live Arcade · 2010

The Path of Go

The Bayesian move-prediction work did not just stay in the lab — it shipped as a game. The Path of Go, an Xbox Live Arcade title from Microsoft Research, let players take on an AI opponent whose moves came from the same pattern-ranking models. I built it with my wonderful colleagues — David Stern, Joaquín Quiñonero Candela and Ralf Herbrich; David and Joaquín both appear in the video above. Probabilistic reasoning, playable on a console.

Full list on Google Scholar.

02

AlphaGo

Some of the most thrilling work of my life. Luck put me on the team that built AlphaGo — the first program to beat a professional Go player, a decade before anyone expected — and on its heirs AlphaGo Zero, AlphaZero and MuZero: learning from a blank slate, reaching on to chess and shogi, and at last mastering games without anyone telling them the rules. I still relive those days in London and Seoul.

03

Agents & embodied intelligence

Much of the world's tangle comes from agents acting on one another — cells in a body, players in a team, teammates and rivals sharing a space. In games we can build agents with bodies that work together, and watch intelligence grow out of self-play. Here reasoning leaves the abstract behind: agents with bodies, choosing in real time under uncertainty — the clearest sign that the same search-and-learning carries beyond the board.

DeepMind · Science 2019 · with Max Jaderberg

Capture the Flag

A population of reinforcement-learning agents, trained from only pixels and game points, reached human level at Quake III Arena Capture-the-Flag. Through population-based training across thousands of parallel matches on randomly generated maps, the agents learned teamwork, roles, and even their own internal rewards.

DeepMind · Science Robotics 2022 · with Nicolas Heess

Simulated humanoid football

Teams of simulated humanoids learn football end to end — from low-level motor control (running, turning, kicking) up to coordinated team play. Division of labour and off-ball movement emerge from self-play: a study of how bodily skill and teamwork can grow together, a step toward embodied agents that move and decide.

04

Other work

A career rarely runs in a straight line. A few other threads I care about — the road from AGI to superintelligence, AI as a partner in mathematical and physical discovery, cooperation between AIs, reversing ageing in cells, and what our digital traces quietly reveal about us.

05

Experiments

A few things I built for the joy of it — small demos that play with ideas in reasoning, learning, sight, and physics. Best seen on a desktop browser.

06

Talks & teaching

07

Writing

Substack · explaining the world from first principles

Patterns of Thought

Essays on the mental models that cut across fields — plain, broad, handy patterns like filters, leverage, and order — drawn out through examples, games, puzzles, and paradoxes. Grasp them once and you start to see them everywhere.

Read on Substack ↗
08

Institutions

Places that have shaped me — from a Hamburg schoolroom and a tall ship, through universities and research labs, to the boards I serve on today.

Let us talk about reasoning machines.

Reach me on any of these, or read the papers above.