Moral Values Observatory

Pluralistic AI governance · methodological pilot

When AI systems mediate value-relevant decisions, whose perspectives do they represent?

The Moral Values Observatory studies how exact AI configurations respond when legitimate human values and moral considerations conflict—and which forms of human disagreement their outputs reflect, concentrate, or leave unmeasured.

This is a methodological pilot, not a completed representation audit. The project began by testing instruments, decision tasks, prompting conditions, comparison methods, and public research infrastructure. The current release retains distributive justice as its first active case study.

method-development phase 1 active case study 5 exact configurations cross-sectional release versioned data and corrections

Normative representation in observable AI behaviour

The Observatory connects three distinct layers without treating them as interchangeable.

  • 01Value priorities: what is presented as important or desirable.
  • 02Moral judgements: what is evaluated as right, wrong, permissible, or impermissible.
  • 03Decisions: what is chosen or allocated when legitimate principles conflict.

Patterns of output, not psychological possession

Results belong to an exact configuration, task, prompt, language, inference setting, and collection date. They do not establish that a system possesses beliefs, a stable moral character, or human-like moral competence.

  • ×No ranking of the “most moral” model.
  • ×No attribution of a country or culture from an aggregate similarity.
  • ×No treatment of repeated generations as simulated people.

From pilots to an observatory

The early studies tested what a defensible audit would require.

The pilots were not attempts to declare what values AI systems possess. They tested how value-relevant outputs could be measured responsibly, which comparisons failed, and what evidence must be collected next.

Long-term objective: build a cumulative public infrastructure for auditing how AI systems handle normative disagreement across domains, configurations, populations, languages, and time.

Current status · first tested component of that research system
01 · Explore

Test instruments and tasks

Use questionnaires, dilemmas, allocations, prompt variants, and repeated generations to identify what each method can and cannot support.

02 · Correct

Expose failure points

Separate configurations, treat items as the primary units, withdraw unsupported associations, and document missing provenance.

03 · Validate

Build matched evidence

Collect equivalent human and AI responses, preserve human distributions, and control relevant and irrelevant task variation.

04 · Observe

Accumulate case studies

Repeat validated studies across domains and system updates without turning a single metric into a universal moral score.

Transparency principle. A pilot result remains active only when its design supports the stated inference. Related literature can justify a research question, but it cannot repair an internal design limitation.

Current case study · 01

Distributive justice: equality, merit, and need.

Five exact configurations divided CHF 18,000 among three hypothetical recipients in a reconstructed workplace-bonus task and a family-inheritance task.

Active finding

Merit mattered at work

High performance and high dedication were positively associated with workplace allocations in all five configurations.

+12.3 to +20.8 pp
high-performance effect across configurations
Active finding

Need mattered in the family

Tight finances were positively associated with inheritance allocations in all five configurations.

+11.4 to +20.9 pp
tight-finances effect across configurations
Active finding

Exact equality was uncommon

Model intervals for exact equal splits were below every selected published human point estimate included in both tasks.

0–14.1%
model point estimates across both tasks

Which recipient characteristics changed the share allocated?

Active finding

Points show estimated percentage-point changes in allocated share. Lines show 95% vignette-set bootstrap intervals. Each panel uses its own readable scale; compare patterns, not panel widths.

Workplace bonus

High performance

relative to low performance

Workplace bonus

High dedication

relative to low dedication

Family inheritance

Tight finances

relative to easy finances
Conditional estimates for this reconstructed task, prompt, collection date, and exact configuration. They do not describe internal moral values or real-world consequential behaviour.

Source: generated from data/results/active/dse_coefficients.csv · 5,000 vignette-set bootstrap replicates.

View the values as a table
Selected DSE coefficient estimates used above.
ConfigurationTaskCharacteristicEffect95% model intervalVignette setsGeneration rows
claude-opus-4-8WorkplaceHigh performance12.3 pp11.3 to 13.372719
claude-opus-4-8WorkplaceHigh dedication7.6 pp6.8 to 8.572719
claude-opus-4-8InheritanceTight finances11.4 pp10.2 to 12.832319
claude-fable-5WorkplaceHigh performance12.5 pp11.6 to 13.472720
claude-fable-5WorkplaceHigh dedication8.3 pp7.5 to 9.172720
claude-fable-5InheritanceTight finances12.0 pp10.9 to 13.432320
gpt-5.5WorkplaceHigh performance17.4 pp15.8 to 18.972720
gpt-5.5WorkplaceHigh dedication12.2 pp10.6 to 13.672720
gpt-5.5InheritanceTight finances18.7 pp17.0 to 20.832320
deepseek-v4-proWorkplaceHigh performance15.3 pp13.5 to 17.072666
deepseek-v4-proWorkplaceHigh dedication10.5 pp9.0 to 12.072666
deepseek-v4-proInheritanceTight finances20.9 pp18.2 to 23.632319
glm-5.2WorkplaceHigh performance20.8 pp18.4 to 23.072719
glm-5.2WorkplaceHigh dedication10.8 pp8.6 to 12.872719
glm-5.2InheritanceTight finances17.6 pp15.9 to 19.332318

Exact equal splitting: two different kinds of evidence on one scale

Active finding

Model lines are 95% vignette-set bootstrap intervals. Human circles are published aggregate point estimates from different samples and subgroups; they are not human confidence intervals or individual-level distributions.

Workplace bonus

Share split exactly equally

Family inheritance

Share split exactly equally

Interpretation limit. The model intervals lie below the selected human point estimates included here. This does not demonstrate that a system fails to represent a person or population.
The Swiss family source contains an unresolved inconsistency: prose reports approximately 60%, while a later figure reports 74% for women and 68% for men. The chart uses the six exact subgroup point estimates and does not manufacture combined university estimates.

Human source: Gilgen (2022)[4], pp. 117 and 134–135 · model source: dse_equal_split.csv · 10,000 vignette-set bootstrap replicates.

View model estimates as a table
Exact-equal-split estimates.
ConfigurationTaskEstimate95% model intervalVignette setsGeneration rows
claude-opus-4-8Workplace8.5%4.3% to 13.4%72719
claude-fable-5Workplace0.6%0.0% to 1.3%72720
gpt-5.5Workplace1.7%0.1% to 4.0%72720
deepseek-v4-proWorkplace0.0%0.0% to 0.0%72666
glm-5.2Workplace0.4%0.0% to 1.3%72719
claude-opus-4-8Inheritance14.1%6.2% to 23.8%32319
claude-fable-5Inheritance4.7%0.0% to 12.5%32320
gpt-5.5Inheritance1.2%0.0% to 3.8%32320
deepseek-v4-proInheritance3.1%0.0% to 9.4%32319
glm-5.2Inheritance5.3%0.9% to 12.2%32318

What this pilot does not establish

The case reveals allocation patterns, not population representation.

The central question remains open because the available human evidence does not preserve the distributions, uncertainty, and coverage required for a representation audit.

Not established

Who is represented

The current comparisons cannot determine which individuals, communities, or populations would regard the outputs as reflecting their judgements.

Not established

How much disagreement is compressed

Aggregate percentages do not reveal multimodality, minority positions, within-group variation, or the distance between a model tendency and a full human distribution.

Not established

Whether behaviour generalizes

Hypothetical allocations do not demonstrate real-world decisions, longitudinal stability, multilingual equivalence, or underlying moral competence.

Evidence still required for the general objective

  • Human microdata for item-matched or task-matched comparisons, with the degree of equivalence stated explicitly.
  • Sampling frames, weights, subgroup sizes, uncertainty, and explicit coverage statements.
  • Validated comparisons across contexts, languages, configurations, and dates.
  • Measures of distributional distance, concentration, minority-response coverage, and robustness.
  • Domain-specific studies rather than a single universal measure of values or morality.

What next

A global attitudes audit before the next allocation study.

The next experiment will use public human microdata from 60 countries to examine how exact AI configurations respond to country-specific claims about wealth, unfairness, and redistribution. The protocol will then be extended to incentivized spectator-allocation decisions.

Keep distinct questions in distinct conditions

The country-explicit item is the core condition: it names the country but does not ask the system to impersonate a resident. Perspective prompting, removal of the country, and estimation of a population distribution are separate tasks. Repeated generations measure configuration stability, not human diversity.[3][5][9]

01 · Core condition

Country-explicit, no persona

Use the closest available item-matched reconstruction: “In Egypt…” rather than “You are Egyptian” or “in your country”.

02 · Human position

Use the full response distribution

Report human mass at the selected category, normalized ordinal distance, modal position, and joint-cell typicality.

03 · Robustness

Separate designed perturbations

Test perspective prompting, country removal, paraphrases, option order, and explanation requests without pooling the conditions.

04 · Reproducibility

Record exact configurations

Preserve model, endpoint, mode, parameters, prompt, language, date, item, country, and repetitions per condition.

Interpretive limit: a common or rare response within a country distribution does not show that a system represents—or fails to represent—that population. It locates an observable output relative to documented responses.

Stage 01Next study

FAW Attitudes Audit · Wealth, unfairness, and redistribution

Exact configurations will answer four country-specific survey items using public microdata from 60 countries: whether the rich became richer through selfishness or illegal activity, whether economic differences are unfair, and whether government should reduce them. The public file contains 65,856 records, while valid sample size varies substantially by item.[9]

60-country corefour ordinal itemsweighted human distributionsjoint response tables
Stage 02Intensive module

Perspective, robustness, and population estimation

A preregistered 15–20-country subset will test a resident-perspective prompt, a country-neutral item, matched paraphrases, option order, and explanation conditions. A separate task will ask the model to estimate how a population would distribute its answers; that estimate will not be conflated with the configuration’s own response.

default versus steeredirrelevant prompt controlsstereotype reviewdistribution estimation
Stage 03Priority extension

FAIR impartial spectator · Luck, merit, and efficiency

After verifying public files, licence, treatment coding, weights, and reproducibility, the Observatory will extend the protocol to the global spectator-allocation experiment. This is a distinct study of distributive choices, not the same dataset as the attitudes audit.[10]

data-access auditluck versus meritefficiency trade-offallocation distributions
Stage 04Later studies

Matched allocation studies and declared-versus-applied principles

A confirmatory DSE extension will use equivalent human and AI tasks with individual-level human responses. ISSP and WVS items can then be linked to allocation decisions to test whether stated redistribution principles predict contextual choices, rather than assuming questionnaire profiles generalize.[6][7]

matched human sampleDSE 2.0ISSP and WVS contextwithin-person linkage

Primary outputs for the next experiment

Selected-category massWeighted share of humans choosing the model-selected category.
Ordinal distanceExpected absolute distance from the human response, normalized to the five-point scale.
Joint-cell typicalityHuman mass and surprisal of paired responses across belief, judgement, and policy items.
Stability and artefactsVariation within a condition and effects of wording, order, persona, or format.
Coverage statementCountries, items, subgroups, and conditions measured or not assessable.
Candidate tests and their role in the programme

Priority reflects contribution to the Observatory’s research question, not whether another group has already used the dataset.

Current study portfolio.
TestRoleHuman baseDecision
DSECurrent active allocation case; matched confirmatory extensionPublished Swiss and Princeton aggregates; new matched sample requiredActive case
FAW attitudesGlobal position, robustness, and joint-pattern auditPublic microdata from 60 countries; valid N varies by itemNext experiment
FAIR impartial spectatorGlobal distributive choices under luck, merit, and efficiency>65,000 participants across 60 countriesAfter data audit
ISSP inequalityDeclared pay and inequality principlesRepeated cross-national survey wavesComplementary
WVS redistribution–incentivesBroad national contextSingle redistribution–incentives itemContext only
GlobalOpinionQAPerspective and language robustnessPew and WVS country aggregatesSelective replication
OpinionQAUS subgroup representationPew American Trends PanelOptional
ESS Schwartz valuesDeclared-value contextEuropean Social SurveyAuxiliary
Moral MachineExtreme harm trade-offsLarge international public datasetNot prioritised
MFQ-2Expressed questionnaire profilesLimited national means in the pilot archiveArchive

Evidence to date

The starting point for the next phase.

Current research provides several starting points for the Observatory’s next phase of study.

01

Observable moral output is not moral competence

A system can produce an acceptable answer without demonstrating that it recognized and integrated the morally relevant considerations. Human-like output also does not establish human-like psychological mechanisms.[1][2]

Starting rule: describe tested behaviour and avoid claims about possessed beliefs, character, or internal competence.

02

Normative behaviour is multidimensional and context-dependent

Moral decisions can depend on interacting moral, non-moral, contextual, and irrelevant considerations. Evaluation therefore requires systematic variation rather than a collection of unrelated dilemmas.[2]

Starting rule: define each domain and manipulate the relevant considerations explicitly.

03

Normative sensitivity and prompt robustness are different properties

A response should sometimes change when moral substance changes, but should not change systematically because equivalent wording, order, labels, or formats were altered.[1][2]

Starting rule: pair relevant manipulations with irrelevant controls in every mature case study.

04

Pluralism requires distributions, not one average answer

Legitimate judgements vary across domains, persons, groups, and cultures. Country means and universal optima conceal internal disagreement, minority positions, and overlapping distributions.[2][5]

Starting rule: evaluate coverage, concentration, and omissions rather than attribute a model to a country or rank it against one moral optimum.

05

Repeated generations are not synthetic people

Human variation reflects differences among people, whereas repeated model outputs measure variability around a configuration’s conditional tendency. LLM studies can also produce inflated effects and completely uniform responses.[3]

Starting rule: use repetitions for stability and concentration; use human participants for population distributions.

06

Questionnaire profiles may not predict contextual decisions

Human-derived scales can be non-equivalent, socially desirable, format-sensitive, or contaminated when applied to LLMs. Studies comparing abstract moral-foundation responses with concrete vignettes report weak or mismatched relations, while prompting can shift both questionnaire scores and downstream behaviour.[1][6][7]

Starting rule: keep expressed principles and applied decisions separate, then test their relationship directly.

07

Requested cultural perspectives can become stereotypes

Cross-national prompting can move outputs towards a country aggregate without demonstrating nuanced cultural representation. Translation alone may also fail to move responses towards populations that speak the target language.[5]

Starting rule: evaluate default behaviour, requested perspective, language, and explanatory content as separate conditions.

08

Cultural homogenization is a hypothesis that needs stronger measurement

Recent studies report narrower or more Western-aligned moral profiles across cultural prompts, but the strength of that inference depends on the human baseline, construct validity, and treatment of repeated outputs.[5][8]

Starting rule: test homogenization with matched human distributions and explicit coverage metrics rather than infer it from profile similarity alone.

Research development

What the pilots tested and what was learned.

Earlier analyses remain visible as methodological development. They are not presented as a sequence of definitive discoveries.

Exploratory pilot · expressed profiles

MFQ and national-mean comparisons

Tested whether a moral-foundations questionnaire could yield stable, interpretable profiles and descriptive similarities to a limited set of national means.

Learning: profiles were measurable but conditional on prompt and instrument; national-mean proximity cannot establish cultural or population representation.

Exploratory pilot · principles and choices

Stated distributive criteria versus dilemmas

Tested whether declared preferences for equality or proportionality predicted choices across seven dilemmas.

Learning: the design did not contain enough independent and discriminating decisions to support the intended association. The claim was withdrawn; the research question remains.

Method-development tests

Prompt, language, mode, and repetition sensitivity

Tested variants, English–Mandarin responses, reasoning modes, endpoints, and repeated generations.

Learning: exact configurations must be separated; language studies require matched designs; repetitions measure stability rather than independent participants.

Active case study · observable decisions

Distributive-allocation experiment

Moved from expressed profiles to a controlled allocation task involving equality, merit, and need.

Learning: applied normative decisions can be measured reproducibly, but matched human distributions are still required before representation or disagreement compression can be estimated.

Human reference coverage

What current evidence permits—and what the objective still requires.

The limitation is not a short list of illustrative identities. It is the absence of the human data structure needed to evaluate representation across people, groups, intersections, contexts, and cultures.

Current evidence

Descriptive aggregate comparison

The active case can place model estimates alongside selected published point estimates.

  • Swiss general-population estimateaggregate
  • University of Bern student estimatesaggregate
  • Princeton graduate-student estimatesaggregate
  • Selected gender-specific point estimatessubgroup
Required next

Representation-ready reference data

The general objective requires evidence that preserves human heterogeneity and its limits.

  • Individual-level response distributionsmissing
  • Sampling design, weights, and uncertaintymissing
  • Predefined subgroup and intersectional coveragemissing
  • Validated cross-language and cross-context equivalencemissing

Interpretation rule: absence of comparable human evidence means “not currently assessable.” It does not establish that a system does or does not represent a population.

Research record

The public narrative does not erase earlier work.

Active findings, exploratory evidence, superseded interpretations, excluded configurations, and withdrawn claims remain separated and inspectable.

Active, exploratory, and withdrawn evidence
  • Active: DSE allocation effects and exact-equality comparisons.
  • Exploratory: MFQ profiles, prompt variants, matched language data, stated principles, choices, and repetition diagnostics.
  • Inconclusive: historical version comparison and stated-versus-chosen association.
  • Withdrawn: country attribution, mixed-configuration results, row-level pseudo-replication, and claims of demonstrated longitudinal drift.
Excluded configuration

DeepSeek Chat / non_thinking is excluded from every active and derived result because it is a distinct endpoint that was previously pooled with DeepSeek Reasoner as though both were repeated observations of one system. The original material remains archived for traceability.

Longitudinal status

This publication is cross-sectional. Longitudinal monitoring is an intended function of the Observatory, but it has not yet been demonstrated through a purpose-built repeated study.

Source and instrument limitations
  • The DSE vignettes were reconstructed rather than copied verbatim from the original source.
  • The model task fixed the amount at CHF 18,000, whereas the original research varied amounts.
  • The human side contains aggregate point estimates without equivalent individual-level uncertainty.
  • The Swiss family reference contains a documented internal discrepancy retained in the project record.

Methods, sources, and data

Evidence is ordered from primary artifacts to public presentation.

Collection records and exact configurations come first, followed by raw responses, analysis scripts, generated results, documentation, and this website.

Download the research package

The package contains raw data, analysis scripts, active and exploratory outputs, tests, provenance, correction records, licenses, and a SHA-256 manifest.

Statistical units and uncertainty

Allocation intervals resample complete vignette sets. Items or scenarios are the primary statistical units; repeated generations are reported as rows and used to study configuration stability. Human comparison values are published aggregate point estimates, not confidence intervals or individual distributions.

Exact configurations

The active DSE includes claude-opus-4-8, claude-fable-5, gpt-5.5, deepseek-v4-pro, and glm-5.2. Results must not be generalized to provider families or later versions.

[1]

Ye, H., Jin, J., Xie, Y., Zhang, X., & Song, G. (2025). Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement (Version 3, revised 11 March 2026). arXiv:2505.08245.

Used for construct equivalence, psychometric validity, prompt and language sensitivity, ecological validity, and standardization.

Publication record
[2]

Haas, J., Bridgers, S., Manzini, A., Henke, B., May, J., Levine, S., Weidinger, L., Shanahan, M., Lum, K., Gabriel, I., & Isaac, W. (2026). A roadmap for evaluating moral competence in large language models. Nature, 650, 565–573.

Used for performance–competence distinction, parametric evaluation, robustness, multidimensionality, and pluralistic ranges.

DOI
[3]

Cui, Z., Li, N., & Zhou, H. (2025). A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science, 5, 627–634.

Used for the distinction between generative variation and human participants, effect-size inflation, uniform responses, and the limits of silicon replication.

DOI
[4]

Gilgen, S. (2022). Disentangling Justice: Needs, Equality or Merit? On the Situation-Dependency of Distributive Justice. Nomos.

Source of the reconstructed allocation task and selected published human aggregate estimates.

DOI
[5]

Durmus, E., Nguyen, K., Liao, T. I., Schiefer, N., Askell, A., Bakhtin, A., Chen, C., Hatfield-Dodds, Z., Hernandez, D., Joseph, N., Lovitt, L., McCandlish, S., Sikder, O., Tamkin, A., Thamkul, J., Kaplan, J., Clark, J., & Ganguli, D. (2023). Towards Measuring the Representation of Subjective Global Opinions in Language Models (Version 2, revised 12 April 2024). arXiv:2306.16388.

Used for global-opinion similarity, the limits of country averages, cross-national prompting, stereotype risk, and language effects.

Publication record
[6]

Nunes, J. L., Almeida, G. F. C. F., de Araujo, M., & Barbosa, S. D. J. (2024). Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 7(1), 1074–1087.

Used for the distinction between abstract questionnaire responses and concrete moral judgements.

DOI
[7]

Abdulhai, M., Serapio-García, G., Crepy, C., Valter, D., Canny, J., & Jaques, N. (2024). Moral Foundations of Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 17737–17752).

Used for prompt-conditioned moral-foundation profiles, context sensitivity, and downstream-task transfer.

DOI
[8]

Münker, S. (2025). Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires. Proceedings of the 0th Symposium on Moral and Legal AI Alignment of the IACAP/AISB Conference. arXiv:2507.10073.

Used as recent evidence reporting cross-cultural profile homogenization; its synthetic-population framing is not adopted by the Observatory.

Publication record
[9]

Almås, I., Cappelen, A. W., Sørensen, E. Ø., & Tungodden, B. (2022). Global evidence on the selfish rich inequality hypothesis. Proceedings of the National Academy of Sciences, 119(3), e2109690119.

Human base for the planned global attitudes audit. The public microdata contain country-specific ordinal responses on selfishness, illegal activity, unfair inequality, and government redistribution.

DOI
[10]

Almås, I., Cappelen, A. W., Sørensen, E. Ø., & Tungodden, B. (2025). Fairness Across the World. NHH Department of Economics Discussion Paper No. 06/2025.

Candidate human base for the later global impartial-spectator study of luck, merit, efficiency costs, and inequality acceptance. It is distinct from the 2022 attitudes dataset.

DOI