Research · Common-task challenge

The AI Respondents Challenge.

Real people answered the World Values Survey. We hid some of their answers. Your task: predict what each person said, using any method you like, and disclose the features and prompts behind your predictions.

The challenge is a live measurement of what language models actually know about people: their opinions, values, and behaviours. Every score comes from answers the models never saw, in an expanding circle of difficulty: held-out respondents, then held-out countries, and at the end of the week held-out questions and entire held-out surveys. Each step outward tests whether a pipeline learned something about people or just memorised a dataset.

Final · updates through the school week Leaderboard Get started Rules

§ 01 · Leaderboard
Standings

Live leaderboard.

Every tier pairs an in-domain board (countries your models can study) with an out-of-domain board (countries they never saw). Teams are ranked by skill, and join a tier's boards once they submit predictions for its questions.

Prediction · In-domain

WVS · seen countries · held-out respondents

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 📡pre-dish salad 0.337 0.956 0.561 100%
🥈 🦖the 5 of 4 0.280 0.933 0.522 100%
🥉 🦉Group 6 0.275 0.935 0.519 100%
4 🦊Group 5 0.189 0.902 0.456 100%
5 📡Group 2 0.160 0.952 0.475 100%
6 🎯Group 3 0.100 0.855 0.461 100%
7 🤖Majority class (baseline) -0.013 0.652 0.153 100%
8 🤖Random draw (baseline) -0.192 0.980 0.252 100%

Prediction · Out-of-domain

WVS · held-out countries (unseen populations)

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 📡pre-dish salad 0.345 0.900 0.533 100%
🥈 🦉Group 6 0.264 0.919 0.503 100%
🥉 🦖the 5 of 4 0.255 0.917 0.498 100%
4 🦊Group 5 0.173 0.899 0.449 100%
5 📡Group 2 0.154 0.919 0.477 100%
6 🎯Group 3 0.145 0.875 0.471 100%
7 🤖Majority class (baseline) -0.002 0.650 0.157 100%
8 🤖Random draw (baseline) -0.209 0.958 0.259 100%

Prediction · Hidden questions

WVS · seen countries · 4 questions never announced as targets

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 🦖the 5 of 4 0.368 0.934 0.507 100%
🥈 📡pre-dish salad 0.345 0.923 0.517 100%
🥉 🦉Group 6 0.318 0.925 0.472 100%
4 📡Group 2 0.176 0.958 0.458 100%
5 🦊Group 5 0.077 0.852 0.351 100%
6 🤖Majority class (baseline) 0.000 0.681 0.163 100%
7 🤖Random draw (baseline) -0.196 0.967 0.245 100%

Prediction · Hidden questions, held-out countries

WVS · held-out countries · hidden questions (compositional shift)

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 🦖the 5 of 4 0.240 0.924 0.430 100%
🥈 📡pre-dish salad 0.236 0.888 0.414 100%
🥉 🦉Group 6 0.203 0.914 0.406 100%
4 🦊Group 5 0.059 0.869 0.358 100%
5 📡Group 2 0.058 0.943 0.394 100%
6 🤖Majority class (baseline) -0.035 0.677 0.161 100%
7 🤖Random draw (baseline) -0.195 0.982 0.256 100%

Prediction · Held-out surveys

ESS wave 11 + Latinobarometro 2023 · WVS-overlap countries · unseen instruments

Final
# TeamSkillESSLatinoAlignmentF1-macroCoverage
🥇 🦉Group 6 0.107 0.197 0.017 0.876 0.436 100%
🥈 🦊Group 5 0.106 0.144 0.068 0.869 0.412 100%
🥉 📡pre-dish salad 0.078 0.035 0.122 0.870 0.409 100%
4 🦖the 5 of 4 0.036 0.076 -0.004 0.840 0.389 100%
5 🎯Group 3 0.027 0.055 -0.001 0.845 0.372 100%
6 🤖Majority class (WVS-seen countries baseline) -0.009 -0.018 0.000 0.715 0.163 100%
7 📡Group 2 -0.027 -0.002 -0.051 0.843 0.315 100%
8 🤖Random draw (WVS-seen countries baseline) -0.154 -0.134 -0.174 0.973 0.279 100%

European Social Survey wave 11 · overlap countries

WVS-overlap countries · this survey's 8 questions only

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 🦉Group 6 0.197 0.904 0.440 100%
🥈 🦊Group 5 0.144 0.870 0.389 100%
🥉 🦖the 5 of 4 0.076 0.853 0.407 100%
4 🎯Group 3 0.055 0.850 0.389 100%
5 📡pre-dish salad 0.035 0.855 0.368 100%
6 📡Group 2 -0.002 0.864 0.299 100%
7 🤖Majority class (WVS-seen countries baseline) -0.018 0.689 0.138 100%
8 🤖Random draw (WVS-seen countries baseline) -0.134 0.971 0.251 100%

Latinobarometro 2023 · overlap countries

WVS-overlap countries · this survey's 8 questions only

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 📡pre-dish salad 0.122 0.885 0.450 100%
🥈 🦊Group 5 0.068 0.868 0.436 100%
🥉 🦉Group 6 0.017 0.849 0.432 100%
4 🤖Majority class (WVS-seen countries baseline) 0.000 0.740 0.189 100%
5 🎯Group 3 -0.001 0.839 0.354 100%
6 🦖the 5 of 4 -0.004 0.827 0.371 100%
7 📡Group 2 -0.051 0.821 0.330 100%
8 🤖Random draw (WVS-seen countries baseline) -0.174 0.974 0.307 100%

Prediction · Held-out surveys, unseen populations

ESS + Latinobarometro · countries WVS never covered · instrument and population shift

Final
# TeamSkillESSLatinoAlignmentF1-macroCoverage
🥇 📡pre-dish salad -0.043 0.039 -0.125 0.862 0.374 100%
🥈 🦉Group 6 -0.049 0.135 -0.233 0.857 0.389 100%
🥉 🦊Group 5 -0.050 0.093 -0.192 0.868 0.369 100%
4 🎯Group 3 -0.098 0.041 -0.237 0.832 0.350 100%
5 🦖the 5 of 4 -0.120 0.057 -0.296 0.822 0.359 100%
6 📡Group 2 -0.182 0.006 -0.371 0.835 0.298 100%
7 🤖Majority class (WVS-seen countries baseline) -0.197 -0.141 -0.253 0.699 0.155 100%
8 🤖Random draw (WVS-seen countries baseline) -0.335 -0.234 -0.436 0.923 0.254 100%

European Social Survey wave 11 · unseen populations

countries WVS never covered · this survey's 8 questions only

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 🦉Group 6 0.135 0.898 0.391 100%
🥈 🦊Group 5 0.093 0.875 0.342 100%
🥉 🦖the 5 of 4 0.057 0.841 0.375 100%
4 🎯Group 3 0.041 0.863 0.374 100%
5 📡pre-dish salad 0.039 0.848 0.353 100%
6 📡Group 2 0.006 0.889 0.305 100%
7 🤖Majority class (WVS-seen countries baseline) -0.141 0.669 0.132 100%
8 🤖Random draw (WVS-seen countries baseline) -0.234 0.935 0.231 100%

Latinobarometro 2023 · unseen populations

countries WVS never covered · this survey's 8 questions only

Final
# TeamSkillAlignmentF1-macroCoverage
🥇 📡pre-dish salad -0.125 0.876 0.395 100%
🥈 🦊Group 5 -0.192 0.862 0.397 100%
🥉 🦉Group 6 -0.233 0.817 0.387 100%
4 🎯Group 3 -0.237 0.801 0.327 100%
5 🤖Majority class (WVS-seen countries baseline) -0.253 0.729 0.177 100%
6 🦖the 5 of 4 -0.296 0.802 0.344 100%
7 📡Group 2 -0.371 0.782 0.292 100%
8 🤖Random draw (WVS-seen countries baseline) -0.436 0.912 0.276 100%

Final leaderboard. Submissions are frozen; all tiers revealed, including the held-out surveys. · Updated 17 Jul, 14:16 UTC

How to read the board

Skill. Accuracy above always guessing the majority answer, rescaled as (accuracy − majority share) / (1 − majority share) and averaged over target questions. 1 = every answer right; 0 = no better than majority guessing; negative = worse than majority guessing. The ceiling is 1 but the floor is not −1: on questions where one answer dominates, skill can fall well below −1, because there is more room under the majority guess than above it. Teams are ranked on this.

Alignment. One minus the distance between your predicted answer distribution and the true one, per question, averaged. The distance is an order-aware earth-mover distance: answer options sit at equal steps in scale order, and we sum the absolute gaps between the two cumulative distributions, divided by the number of steps. Range 0 to 1: 1 = identical distributions, 0 = all your mass at one end of the scale while the truth sits entirely at the other. Only valid labels enter this measure; invalid ones are already punished in skill.

F1-macro. Computed in two steps. Per question: all answered pairs are pooled across countries, an F1 score (the harmonic mean of precision and recall) is computed for each answer option, and these are averaged with equal weight to give the question's macro-F1. The board then shows the plain average of the per-question values. Range 0 to 1. Because every answer option counts equally, rare answers matter as much as common ones, which is exactly what accuracy misses.

Coverage. Share of respondent–question pairs answered, from 0 to 100%. Unanswered pairs count as wrong in skill, so predict everyone.

Hidden-question boards. The Hidden questions tab covers the four questions that were never announced as targets until their mid-week release. Same metrics, same truth data, but scored only over the hidden questions and only for teams that submitted them. The main boards always score the original nine questions, so submitting the hidden tier never changes your main-board standing by itself.

Held-out survey boards. The third tab scores two surveys released on Thursday with no training data at all: European Social Survey wave 11 and Latinobarometro 2023, predicted zero-shot. The pooled view ranks teams over all sixteen questions, both surveys weighing equally, split into countries WVS also covers and countries it never has. The ESS and Latinobarometro views show each survey's own boards with per-survey skill, alignment and F1. On these boards the baselines are labelled WVS-seen countries baselines: their base rates come from the countries WVS also covers, so on the unseen-population boards they show what majority guessing scores when the majorities were learned somewhere else.

Baselines. Majority class always picks each question's most common training answer; Random draw samples from training shares. Both score like teams. Note that skill 0 already marks a perfect majority guesser for a board's own countries, by construction. The baseline rows estimate those majorities from the training data instead, so on boards whose countries have no training data they sit below zero: that gap is the cost of assuming base rates transfer. Doing well on skill and alignment together is the goal.

§ 02 · Get started
§ 03 · Rules
Everything you need to submit

The rules.

Short version: any method, full disclosure, don't game the truth.

The data: four parts

train
Labeled respondents (100 per seen country) with all their survey answers, for exploring the data and validating your pipeline.
test
The 1,050 respondents you predict for (30 per country: 20 seen countries plus 15 that never appear in train). Their answers to the target questions are hidden and never released.
targets
The 9 target questions with their answer options in scale order. The label column is the answer space: predictions must match one of a question's labels exactly. Anything else scores zero.
features
The allowed feature pool, with question text and the code→text map for turning coded answers into words.

What you submit: three parts

predictions.csv

respondent_id · question_id · prediction

One row per (test respondent, target question). prediction is a text label from targets.

features.csv

question_id · feature_variable_code

The variables your method actually relied on, per question. If you prompted with age, religion, and income, those three go here.

method/

prompts.jsonl · method.md

Your prompt template per question and a short note on your approach, kept for reproducibility.

Zip the folder and upload it via the submission portal announced at the school.

Fair play

Any method is allowed
Provided you disclose it via features.csv and method/.
Features only from the pool
Respondent information must come from the feature pool. Target questions are never usable as features.
No reverse-engineering test labels
Final scoring uses a private split.
Budget
Open models run on the shared Nebius allowance; a full pass over the test set costs a few dollars. Only test predictions are scored, so spend deliberately. Closed models require your own API key.
Research use
By submitting you agree that your predictions, declared features, prompts, and method notes may be used, with attribution, in research arising from the challenge.

The week

Mon 13 Jul

Launch: data walkthrough, teams form, starter kit and infrastructure check

Tue–Wed

Team research work; the live leaderboard is open throughout

Wed 15 Jul, evening

Hidden questions released (targets_hidden.csv on the dataset page): four new target questions, same respondents and features

Thu 16 Jul, morning

Held-out surveys released (which ones stays secret until then): run your pipeline on data from surveys it has never seen, same folder layout

Fri 17 Jul, 9:00

Final submission deadline for everything, then results reveal and team presentations

The evaluation sets are scored once, after the Friday 9:00 deadline, and revealed at the wrap-up: there is no live leaderboard for them. Run them with the pipeline behind your latest main-board submission (re-pointing your code at the new files is fine; changing prompts, models, or feature selection after a release is not).