Study Finds LLM Value Alignment Varies Across European Groups
A study released on August 7 compared 10 commercial language models with 50,116 respondents in the European Social Survey and found that model answers aligned unevenly across countries and socio-demographic groups. Country of residence alone explained 7.3% to 27.8% of individual alignment-score variation, depending on the model, while education, income, occupation, and religion also showed material differences.
Researchers at Ghent University released a study on August 7 examining whose stated values are best reflected by commercial large language models. The paper, accepted at AIES 2026, compares model answers with responses from the European Social Survey rather than treating each country as a single cultural profile.
How the comparison worked
The researchers tested 10 models from OpenAI, Anthropic, DeepSeek, and Mistral against Wave 11 of the European Social Survey. That survey covered 50,116 respondents across 29 European countries and Israel. The team used 47 value- and opinion-related survey questions, including a 21-question subset based on the Portrait Values Questionnaire, and prompted each model 20 times in English.
For each question, the study measured how close a model's majority answer was to each respondent's answer on the same scale. Overall alignment scores varied by model, from 0.581 to 0.745 on the subset of questions all models answered. These scores describe agreement on the selected survey items; they are not a general measure of whether a model is safe, correct, or ethically aligned.
What the study found
The authors report higher average agreement for respondents with more education, greater financial security, and white-collar occupations. The widest household-finance gap in their group analysis was 0.0385, while the difference between the highest- and lowest-aligned countries, Sweden and Bulgaria, was 0.0896. Reweighting countries to have similar socio-demographic compositions left most country-level variation intact.
In predictive models, country of residence alone explained between 7.3% and 27.8% of individual alignment-score variation, depending on the LLM. For the full question set, that was at least comparable to the explanatory power of all 15 socio-demographic variables together. Combining country and socio-demographics performed best, which led the authors to argue that neither dimension should be used as a substitute for the other.
Why it matters for evaluation
For teams evaluating models, the practical lesson is that a national average can hide meaningful within-country differences. Evaluation sets built only around geography may miss disparities associated with income, education, occupation, religion, or migration background. The paper also shows that conclusions depend on which definition of values and which questions an evaluator chooses.
The findings remain bounded by the study design. Models were prompted only in English, majority voting reduced response variability, multiple-choice answers are an imperfect proxy for values, and the analysis is regional rather than global. The results should therefore guide more granular evaluation design, not be read as a final ranking of models or populations.
Key Points
- 1The study compared 10 commercial LLMs with 50,116 European Social Survey respondents across 30 countries.
- 2Country alone explained 7.3% to 27.8% of individual alignment-score variation, while socio-demographic variables added complementary information.
- 3The authors caution that English-only prompts, majority-vote model answers, survey-question choices, and the regional sample limit generalization.
Scoring Rationale
A peer-reviewed-conference research result with reproducible methods and direct implications for pluralistic LLM evaluation. Its value is meaningful but bounded by an English-only European survey design and self-reported value questions.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems


