Fourth edition · July 2026
2026Nuffield College, Oxford — five-day intensive on LLMs for social science.
Six research teams, eleven speakers from industry and academia, and a week-long prediction challenge on survey response.
A week-long common-task challenge.
The fourth Oxford summer school on LLMs for social science ran at Nuffield College in July 2026, supported by the College and CASSM, with research support from Nebius.
The technical programme was taught almost entirely by ML practitioners from industry. Sessions covered agents and evaluation, continual learning, model compression, the economics of inference, cloud and self-hosted deployment, and the construction of training data and reinforcement-learning environments. That is a more engineering-heavy curriculum than in previous years, and a deliberate one: the questions social scientists now face about these models are as much about cost, reproducibility and infrastructure as about prompting.
Around that ran the research project. In 2025 it occupied the final two days; in 2026 it took the whole week. Teams formed on Monday morning, worked through every afternoon, and presented on Friday, which is long enough to get past a first idea and try a second and a third.
Between the two: guest research talks from social scientists on their own current work, a welcome social on Monday evening, and a guided walking tour of Oxford on the Wednesday.
Afternoon research sessions. Teams worked in college rooms across Nuffield through the week.
Five days at Nuffield.
Morning instruction, afternoon team research, closing research talks. 09:00–17:00 daily. Click a speaker to jump to their profile.
- 09:15Research project launch — the AI Respondents Challenge
- 11:00Margaret Roberts — Research Talk
- 14:00Tatiana Shavrina — LLM Agents and Evaluations
- 16:00Roberto-Rafael Maura-Rivero — Research Talk
- 19:00Welcome social
- 09:15Sagi Shaier — Continual Learning: Why Models Forget, and What We Can Do About It
- 11:00Team research
- 14:00Team research
- 16:00Christopher Barrie — Research Talk
- 09:15Ilya Boytsov — Efficient LLMs: Principles and Compression Techniques
- 10:45Oxford walking tour — participants & speakers
- 14:00Team research
- 16:00Andreu Casas — Research Talk
- EveningHidden questions released
- 09:15Piotr Mazurek — The Economics of LLM Inference
- 11:00Held-out surveys released
- 14:00Team research
- 16:00Team research
- 09:15Sergei Skvortsov — From Days to Hours: Accelerating LLM-Driven Research
- 11:00Ibragim Badertdinov — LLM Data and RL Environments
- 14:00Team presentations & results reveal
- 16:00Elliott Ash — Research Talk
The AI Respondents Challenge.
We hid part of the World Values Survey and asked six teams to predict what each respondent had answered.
Any method was allowed, on one condition: teams had to declare which respondent attributes their pipeline used, and hand over the prompts they used them in. The week therefore left more than a ranking. It left a record of what six independent groups thought would work.
Scoring widened through the week, and each new tier was released only once the previous one was locked. First, held-out respondents from countries teams could study. Then countries that appeared nowhere in training. Then, on Wednesday evening, four target questions nobody had been told about. Finally, on Thursday morning, two surveys nobody had seen or been told to expect: European Social Survey wave 11 and Latinobarómetro 2023, to be predicted zero-shot by a pipeline built for something else.
Every team beat majority guessing comfortably on home ground, and moving to unseen countries cost them almost nothing. Moving to a different survey instrument cost them most of their skill. On Latinobarómetro questions in countries the World Values Survey has never covered, every team scored worse than guessing. Whatever these models have learned about how people answer survey questions travels well within a European frame and poorly out of it.
Friday: team presentations and the reveal of the held-out survey results.
The people who made the week.
Slides, data, and the final leaderboard.
Lecture slides, the challenge dataset, and the starter pipeline are published openly. More is still landing; see below.
Workshop materials
Lecture slides and notebooks from the 2026 edition, alongside previous years.
Open on GitHub →Challenge starter kit
The Colab-ready pipeline teams built on: load the data, prompt a model, write a submission.
Open on GitHub →Challenge dataset
Training respondents, the permitted feature pool, target questions, and the test set.
Open on Hugging Face →Final leaderboard
All tiers, including the held-out surveys revealed on the Friday of the school.
View the standings →More to come. Materials go up as speakers clear them for release, so a session missing today may well be there next month. We are also editing the lecture and research talk recordings from the week and will publish them openly once post-production is done.