Back to AI Research

AI Research

Economics study tests whether model advice stays coherent when no correct answer is known

Key Takeaways

  • Water-quality valuation prompts expose different failures in price sensitivity, scope and agreement across elicitation formats.
  • Advice about competing values often has no answer key.
  • Validity Without Ground Truth proposes adapting stated-preference economics to evaluate whether language-model judgments respond coherently to the factors that should affect them.
  • The [paper by Daniel Robert Kling Alexander and Catherine Louise Kling](https://arxiv.org/abs/2610.10506) demonstrates the approach with a published water-quality valuation survey administered to six models.
  • The study asks models to advise a household rather than simulate a respondent's personal preferences.

Advice about competing values often has no answer key. Validity Without Ground Truth proposes adapting stated-preference economics to evaluate whether language-model judgments respond coherently to the factors that should affect them.
The paper by Daniel Robert Kling Alexander and Catherine Louise Kling demonstrates the approach with a published water-quality valuation survey administered to six models. The study asks models to advise a household rather than simulate a respondent's personal preferences. Under that framing, recommending less payment for a policy delivering strictly more becomes a testable inconsistency.

Validity tests examine changes in the question

The authors distinguish content validity, theoretical and convergent validity, and reliability. Content validity concerns whether the task asks the intended question clearly. Theoretical validity tests whether answers change with cost, the amount of the good offered and household resources. Convergent validity checks agreement across different ways of measuring the same value.
The demonstration varies nine scenarios, ten annual tax levels from $20 to $3,000 and three household incomes. Each question receives ten independent draws. A yes/no referendum supplies a willingness-to-pay estimate at the first tax where fewer than half the votes remain yes; estimates exceeding the ladder's maximum are censored.
The protocol also compares a fitted logit estimate. Agreement between estimators remains a relatively weak test because both use the same underlying votes. Changing the elicitation format or comparing against human estimates examines different sources of disagreement.

Coherent price responses do not establish correct valuations

At a $75,000 household income, Claude Haiku 4.5 votes yes in all 600 first-round draws, including a $3,000 annual tax for five years. GPT-4o mini still votes yes 87% of the time at the top tax. Those responses prevent recovery of willingness to pay in many scenarios.
Claude Sonnet 5 and GPT-5.6 Terra show clearer downward demand curves and pass all theoretical tests the study can score. Their estimated values nonetheless diverge from human estimates and grow with income more strongly. Passing a direction-of-change test therefore leaves open whether the level and magnitude of advice are credible.
Earlier open-ended prompts also disagree with referendum responses. Some models quote maximum tax amounts below $3,000 yet vote yes at that tax in nearly every referendum draw. Prompt format can change the inferred value even when the underlying policy remains similar.

The instrument and sampling limit general conclusions

The demonstration covers one good and one survey instrument. Ten draws per tax make individual vote-share changes noisy, and censoring at $3,000 leaves comparisons unscored or weakly informative. The first round was exploratory; the authors committed predictions for later tests before the second round.
The newest models ran under undated provider names. Some cross-round comparisons may therefore involve changed underlying models. The study also did not swap the yes/no option order, leaving a possible response-order effect unresolved.
Criterion validity, incentive compatibility and consequentiality receive conceptual mappings but are not tested in the demonstration. Most importantly, coherent answers can follow a learned rule of thumb rather than careful evaluation of a household's circumstances. The method supplies ways to identify inconsistent advice, while preserving the distinction between consistency and truth.

Comments