Fourth edition · July 2026

2026

Nuffield College, Oxford — five-day intensive on LLMs for social science.

Six research teams, eleven speakers from industry and academia, and a week-long prediction challenge on survey response.

Oxford LLMs 2026 cohort at Nuffield College
§ 01 · About
Fourth edition

A week-long common-task challenge.

The fourth Oxford summer school on LLMs for social science ran at Nuffield College in July 2026, supported by the College and CASSM, with research support from Nebius.

The technical programme was taught almost entirely by ML practitioners from industry. Sessions covered agents and evaluation, continual learning, model compression, the economics of inference, cloud and self-hosted deployment, and the construction of training data and reinforcement-learning environments. That is a more engineering-heavy curriculum than in previous years, and a deliberate one: the questions social scientists now face about these models are as much about cost, reproducibility and infrastructure as about prompting.

Around that ran the research project. In 2025 it occupied the final two days; in 2026 it took the whole week. Teams formed on Monday morning, worked through every afternoon, and presented on Friday, which is long enough to get past a first idea and try a second and a third.

Between the two: guest research talks from social scientists on their own current work, a welcome social on Monday evening, and a guided walking tour of Oxford on the Wednesday.

Three participants working together at a laptop during the research project A research team working around a table at Nuffield College

Afternoon research sessions. Teams worked in college rooms across Nuffield through the week.

§ 02 · Programme
13–17 July 2026

Five days at Nuffield.

Morning instruction, afternoon team research, closing research talks. 09:00–17:00 daily. Click a speaker to jump to their profile.

Mon · 13 Jul Launch & Agents
  1. 09:15Research project launch — the AI Respondents Challenge
  2. 11:00Margaret Roberts — Research Talk
  3. 14:00Tatiana Shavrina — LLM Agents and Evaluations
  4. 16:00Roberto-Rafael Maura-Rivero — Research Talk
  5. 19:00Welcome social
Tue · 14 Jul Continual Learning
  1. 09:15Sagi Shaier — Continual Learning: Why Models Forget, and What We Can Do About It
  2. 11:00Team research
  3. 14:00Team research
  4. 16:00Christopher Barrie — Research Talk
Wed · 15 Jul Efficiency & Oxford Tour
  1. 09:15Ilya Boytsov — Efficient LLMs: Principles and Compression Techniques
  2. 10:45Oxford walking tour — participants & speakers
  3. 14:00Team research
  4. 16:00Andreu Casas — Research Talk
  5. EveningHidden questions released
Thu · 16 Jul Inference Economics
  1. 09:15Piotr Mazurek — The Economics of LLM Inference
  2. 11:00Held-out surveys released
  3. 14:00Team research
  4. 16:00Team research
Fri · 17 Jul Data, Presentations & Outro
  1. 09:15Sergei Skvortsov — From Days to Hours: Accelerating LLM-Driven Research
  2. 11:00Ibragim Badertdinov — LLM Data and RL Environments
  3. 14:00Team presentations & results reveal
  4. 16:00Elliott Ash — Research Talk
A lecture in progress in the Nuffield College lecture room, participants at laptops
§ 03 · Challenge
The research project

The AI Respondents Challenge.

We hid part of the World Values Survey and asked six teams to predict what each respondent had answered.

Any method was allowed, on one condition: teams had to declare which respondent attributes their pipeline used, and hand over the prompts they used them in. The week therefore left more than a ranking. It left a record of what six independent groups thought would work.

Scoring widened through the week, and each new tier was released only once the previous one was locked. First, held-out respondents from countries teams could study. Then countries that appeared nowhere in training. Then, on Wednesday evening, four target questions nobody had been told about. Finally, on Thursday morning, two surveys nobody had seen or been told to expect: European Social Survey wave 11 and Latinobarómetro 2023, to be predicted zero-shot by a pipeline built for something else.

Every team beat majority guessing comfortably on home ground, and moving to unseen countries cost them almost nothing. Moving to a different survey instrument cost them most of their skill. On Latinobarómetro questions in countries the World Values Survey has never covered, every team scored worse than guessing. Whatever these models have learned about how people answer survey questions travels well within a European frame and poorly out of it.

A team presenting their method at the final session, leaderboard on screen A research team in front of the final leaderboard standings

Friday: team presentations and the reveal of the held-out survey results.

§ 04 · People
Speakers & organisers

The people who made the week.

§ 05 · Materials
Open access

Slides, data, and the final leaderboard.

Lecture slides, the challenge dataset, and the starter pipeline are published openly. More is still landing; see below.

More to come. Materials go up as speakers clear them for release, so a session missing today may well be there next month. We are also editing the lecture and research talk recordings from the week and will publish them openly once post-production is done.