Probabilistic reasoning · AlphaGo · Robotics

Thore Graepel

I build machines that reason and act — from probabilistic inference at web scale, through AlphaGo, to a new venture bringing AlphaGo-style reasoning to robots.

Chair of Machine Learning, UCL

AlphaGo vs. Lee Sedol · Game 2, 2016 · Move 37 · AlphaGo's shoulder hit at P10
The AlphaGo story →
00

Reasoning, from factor graphs to robots

Thore Graepel

I am a machine-learning researcher who has spent a career teaching machines to reason under uncertainty and to act in the world. I trained as a physicist — in Hamburg, at Imperial College London, and in Berlin, where I earned my PhD in machine learning in 2001. At Microsoft Research from 2003 I co-founded the Online Services and Advertising group, where probabilistic reasoning shipped at scale: TrueSkill, the Bayesian rating system behind matchmaking on Xbox Live, and AdPredictor, the click-through model behind Bing — both built on approximate message passing over factor graphs.

At DeepMind I returned to my first love — creating intelligent systems — and had the fortune to contribute to AlphaGo, the first program to defeat a professional at the full-sized game of Go, and to its successors AlphaGo Zero, AlphaZero and MuZero. There, learned reasoning met deep search: a policy that suggests, a value that judges, a tree that looks ahead.

Much of my work since has been on multi-agent systems — from Capture-the-Flag agents to simulated humanoid football — where intelligence is embodied, cooperative, and learned through self-play. After leading machine learning for cellular rejuvenation at Altos Labs and a spell on Google DeepMind's Post-AGI team, those threads now converge in a new venture: bringing AlphaGo-style reasoning to robotics, so that machines can plan and act under real-world uncertainty. Alongside this I hold the Chair of Machine Learning at UCL.

01

Probabilistic reasoning

Long before deep networks, I worked on machines that reason with uncertainty — Bayesian models that hold beliefs, update them from evidence, and scale to the web through approximate message passing on factor graphs. Ranking players, predicting clicks, recommending, and — a decade before AlphaGo — predicting the next move in Go.

Microsoft Research · Xbox Live Arcade · 2010

The Path of Go

The Bayesian move-prediction research didn't just stay in the lab — it shipped as a game. The Path of Go, an Xbox Live Arcade title from Microsoft Research, let players take on an AI opponent whose moves were guided by the same pattern-ranking models. Built with my wonderful collaborators — David Stern, Joaquín Quiñonero Candela and Ralf Herbrich, all in the video above. Probabilistic reasoning, playable on a console.

Full list on Google Scholar.

02

AlphaGo

Some of the most thrilling work of my life. I was fortunate to be on the team that built AlphaGo — the first program to beat a professional Go player, a decade before anyone expected — and its successors AlphaGo Zero, AlphaZero and MuZero: learning tabula rasa, generalising to chess and shogi, and finally mastering games without even being told the rules. I still relive those days in London and Seoul.

03

Agents & robotics

Much of the world's complexity is agents interacting — cells in a body, players in a team, robots in a warehouse. In games we can build embodied, cooperative agents and watch intelligence emerge from self-play. These are the testbeds on the road from AlphaGo to robots that reason and act in the physical world.

DeepMind · Science 2019 · with Max Jaderberg

Capture the Flag

A population of reinforcement-learning agents, trained from only pixels and game points, reached human level at Quake III Arena Capture-the-Flag. Through population-based training across thousands of parallel matches on randomly generated maps, the agents learned teamwork, roles, and even their own internal rewards.

DeepMind · Science Robotics 2022 · with Nicolas Heess

Simulated humanoid football

Teams of simulated humanoids learn football end to end — from low-level motor control (running, turning, kicking) up to coordinated team play. Division of labour and off-ball movement emerge from self-play: a study of how embodied skill and cooperation can be learned together, a step toward robots that move and decide.

04

Other work

A career rarely runs in a straight line. A few other threads I care about — the road from AGI to superintelligence, cooperation between AIs, reversing ageing in cells, and what our digital traces quietly reveal about us.

05

Experiments

A few things I built for the joy of it — small interactive demos exploring ideas in learning, perception, and physics. Best viewed on a desktop browser.

06

Talks & teaching

07

Writing

Substack · explaining the world from first principles

Patterns of Thought

Essays on the mental models that cut across disciplines — simple, general, useful patterns like filters, leverage, and order — explored through examples, games, puzzles, and paradoxes. Grasp them once and you start to see them everywhere.

Read on Substack ↗
08

Institutions

Places that have shaped me — from a Hamburg schoolroom and a tall ship, through universities and research labs, to the boards I serve on today.

Let's talk about reasoning machines.

Reach me on any of these, or read the papers above.