Changing an AI assistant's priorities can also change its account of the situation. A recommendation meant to emphasize security rather than achievement may lose a constraint, alter a quantity or refuse a harmless request. Values as Style studies a more selective intervention: edit the value-related representation while preserving the scenario information needed for the answer.
A one-way path between two codes
The researchers keep the underlying language model frozen and train lightweight encoders on its residual states. One code represents semantics, including entities, facts and task constraints. Another represents normative priorities.
The design lets semantic information help identify the relevant value. A stop-gradient prevents value-loss feedback through that particular bridge from changing the semantic encoder. The title's analogy to style has a qualification: values depend on the situation, so the method does not assume that value and meaning can be separated completely.
Training uses scenario-matched examples expressing contrasting values, together with paraphrases of each. Swapping the value codes while retaining semantic codes provides a preservation objective. Other losses discourage topic shortcuts and correlated codes. At inference, the method computes a residual update by comparing the reconstructed state before and after the value edit.
Testing preservation alongside alignment
Core experiments use LLaMA-3.1-8B-Instruct and Qwen2.5-7B-Instruct, with 630 training quadruples and scenario-disjoint evaluation. The larger 10,000-example resource belongs to a separate scaling analysis. This distinction prevents its size from being mistaken for the data used in the principal comparison.
The researchers edit the last prompt token's hidden state once, then regenerate the entire completion. Their fidelity measurements therefore compare full responses, rather than preserving a generated prefix by construction.
On LLaMA-3.1-8B, the paper reports alignment of 0.750 for the interface versus 0.748 for a validation-selected prompting baseline. BERTScore rises from 0.923 to 0.938, while the measured contradiction rate falls from 7.6% to 5.1%. These are results under the paper's evaluation protocol, not guarantees that edited answers remain factually correct.
A blinded comparison of 320 items adds human assessments of preservation. The authors report higher preservation ratings for the full interface, while also reporting variable agreement between raters. Those judgments complement the automated metrics rather than replacing them.
Representation and activation make different contributions
The study separates learning the one-way representation from deciding whether to activate an edit. An inference-time gate reduces interventions on prompts judged irrelevant to the target value. A matched ablation shows benefits from both gating and one-way mixing, rather than assigning every improvement to the same component.
The leakage probes still find residual information crossing between the codes. That limits any claim of perfect disentanglement. Evaluation also depends on operational definitions of value alignment and preservation, including model judges and similarity measures.
For teams investigating controllable assistants, the useful contribution is a testable alignment–preservation trade-off. A higher target-value score alone cannot show whether an intervention kept the original constraints intact; this paper measures both.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!