Back to AI Research

AI Research

Values as Style separates value steering from the facts an LLM should preserve

Key Takeaways

  • A learned interface edits value-related activations while holding a semantic code fixed.
  • Tests on two frozen language models report less content drift than several steering and pro
  • Tests on two frozen language models report less content drift than several steering and prompting controls.
  • Changing an AI assistant's priorities can also change its account of the situation.
  • A recommendation meant to emphasize security rather than achievement may lose a constraint, alter a quantity or refuse a harmless request.

Changing an AI assistant's priorities can also change its account of the situation. A recommendation meant to emphasize security rather than achievement may lose a constraint, alter a quantity or refuse a harmless request. Values as Style studies a more selective intervention: edit the value-related representation while preserving the scenario information needed for the answer.

A one-way path between two codes

The researchers keep the underlying language model frozen and train lightweight encoders on its residual states. One code represents semantics, including entities, facts and task constraints. Another represents normative priorities.
The design lets semantic information help identify the relevant value. A stop-gradient prevents value-loss feedback through that particular bridge from changing the semantic encoder. The title's analogy to style has a qualification: values depend on the situation, so the method does not assume that value and meaning can be separated completely.
Training uses scenario-matched examples expressing contrasting values, together with paraphrases of each. Swapping the value codes while retaining semantic codes provides a preservation objective. Other losses discourage topic shortcuts and correlated codes. At inference, the method computes a residual update by comparing the reconstructed state before and after the value edit.

Testing preservation alongside alignment

Core experiments use LLaMA-3.1-8B-Instruct and Qwen2.5-7B-Instruct, with 630 training quadruples and scenario-disjoint evaluation. The larger 10,000-example resource belongs to a separate scaling analysis. This distinction prevents its size from being mistaken for the data used in the principal comparison.
The researchers edit the last prompt token's hidden state once, then regenerate the entire completion. Their fidelity measurements therefore compare full responses, rather than preserving a generated prefix by construction.
On LLaMA-3.1-8B, the paper reports alignment of 0.750 for the interface versus 0.748 for a validation-selected prompting baseline. BERTScore rises from 0.923 to 0.938, while the measured contradiction rate falls from 7.6% to 5.1%. These are results under the paper's evaluation protocol, not guarantees that edited answers remain factually correct.
A blinded comparison of 320 items adds human assessments of preservation. The authors report higher preservation ratings for the full interface, while also reporting variable agreement between raters. Those judgments complement the automated metrics rather than replacing them.

Representation and activation make different contributions

The study separates learning the one-way representation from deciding whether to activate an edit. An inference-time gate reduces interventions on prompts judged irrelevant to the target value. A matched ablation shows benefits from both gating and one-way mixing, rather than assigning every improvement to the same component.
The leakage probes still find residual information crossing between the codes. That limits any claim of perfect disentanglement. Evaluation also depends on operational definitions of value alignment and preservation, including model judges and similarity measures.
For teams investigating controllable assistants, the useful contribution is a testable alignment–preservation trade-off. A higher target-value score alone cannot show whether an intervention kept the original constraints intact; this paper measures both.

Comments (0)

No comments yet

Be the first to share your thoughts!