What the paper is about
Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.
What it covers
Character Training for Risk-Averse Agents Arav Dhoot
- Affiliation: Columbia University Punya Syon Pandey
- Affiliation: UK AI Security Institute Jamie Johnson Affiliation: UK AI Security Institute Daniel Tan Affiliation: Arcadia Impact Elliott Thornley Affiliation: National University of Singapore David Demitri Africa Affiliation: Resolution Abstract Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent’s resources and instill it through on-policy distillation. Despite never seeing the benchmark’s decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents. \NoHyper † † footnotetext: ∗ Equal contribution. Correspondence to: [email protected]. GitHub link: github.com/arav-dhoot/risk-averse-character-training \endNoHyper 1 Introduction We might be able to avoid catastrophic harm from misaligned AI agents if they are highly risk-averse in resources, because we could pay them to cooperate with us ( Thornley and MacAskill, 2026 ) . Misaligned but risk-neutral AIs maximize expected resources, so paying them enough to outbid rebellion is unaffordable. However, sufficiently risk-averse AIs derive steeply diminishing marginal utility from resources, making it feasible to pay them enough to disincentivise rebellion. Risk-averse AIs are therefore comparatively “cheap” to create deals with, resulting in deal-making being a more plausible strategy ( Salib and Goldstein, 2024 ; Assadi, 2025 ; Carlsmith, 2025 ; Finlinson and West, 2025 ; Finnveden, 2025 ; Greenblatt and Fish, 2025 ; Patel, 2025 ; Stastny et al., 2025 ; Mallen, 2026 ; Pan, 2026 ) . However, like any other method to align advanced AI models ( Bai et al., 2022 ; Guan et al., 2025 ) , the practical value of this proposal depends on how well it generalises. Existing work ( Zhang et al., 2026 ) provides evidence that risk aversion can be trained into language models, comparing supervised fine-tuning (SFT) on demonstrations, direct preference optimisation (DPO; Rafailov et al. (2023) ), and activation steering ( Turner et al., 2024 ) , finding that SFT can induce preferences that generalize partially from low-stakes decisions to decisions involving much larger payoffs. However, important questions remain about whether these methods scale well (in token count and model size) and generalise robustly to other tasks. Character training offers a promising approach to these challenges ( Anthropic, 2024 ; Kutasov et al., 2026 ; Tice et al., 2026 ; Maiya et al., 2025 ) . By expressing the target preference as a general disposition, it may enable models to draw on their broader understanding of that disposition when making decisions in unfamiliar settings. We therefore investigate whether specifying risk aversion as part of a model’s character can induce preferences that generalise more reliably across tasks, decision formats, and model scales. Contributions. Our main contributions ( Figure 1 ) are: 1. Character training can induce risk aversion in AI agents. We introduce a novel approach to instill risk aversion through character-training that competes or outperforms baselines on risk aversion, while requiring no labelled data by using a natural-language constitution to supply supervision. 2. An empirical account of what makes character training effective. We systematically investigate how a disposition should be expressed to produce the intended behaviour. Token budget is the most influential factor, and traits written in a declarative tone increase cooperation rates and induce risk aversion more effectively than procedurally phrased traits. 3. Evidence that risk aversion training changes other safety-relevant behaviours. Alongside improvements in the targeted preference, we find a consistent increase in myopic-reward preference across all four models, and other less consistent changes. Our results suggest that character training is a viable and scalable route to instilling robust dispositions that deal-making proposals require. However, given the sensitivity of downstream behaviour to how a trait is phrased, constitutions must be carefully designed and audited before being deployed as a safety intervention. Figure 1: Character training instils reasonable risk aversion without training on the benchmark format. (a) A CARA agent ( α = 0.01 \alpha=0.01 ) takes a sure $40 over a 50/50 bet on $100, but takes a 75% chance of $600 over a sure $5. (b) A frozen teacher prompted with the constitution distils it into a LoRA student on free-form prompts. (c) On the held-out option-menu benchmark, the best constitution matches the prompted teacher and beats demonstration baselines trained on that format. Means and Qwen results are in Figure 2 . 2 Related Work Risk aversion as a safety layer and deal-making with AIs. There is a growing literature describing the risk attitudes of large language models ( Raman et al., 2024 ; Buchanan and Foster, 2026 ; Jarne Ornia et al., 2025 , e.g.) . Thornley and MacAskill (2026) build on this, claiming that making AI agents risk-averse in resources would be desirable because it makes deals with misaligned AIs more feasible. They recommend aiming for constant absolute risk aversion (CARA) specifically, which descends from Pratt (1964) . Most relatedly, Zhang et al. (2026) build on this work by comparing SFT, DPO and activation steering for inducing CARA, finding partial generalisation from low stakes to astronomically high stakes. Constitutional AI and character training. Constitutional AI uses natural-language principles plus AI feedback to shape behaviour without human demonstrations ( Bai et al., 2022 ) . Maiya et al. (2025) shape assistant characters via constitutions and synthetic introspective data, finding constitution-trained characters more robust to adversarial prompting than system prompts or steering. Deliberative alignment ( Guan et al., 2025 ) similarly trains models to reason explicitly over a written spec. Synthetic document finetuning (SDF) generates a corpus of documents that are consistent with some disposition and finetunes on them so that the model comes to behave as though the content in the documents is true of the world ( Wang et al., 2025 ) . This is commonly done in midtraining ( Tice et al., 2026 ; Li et al., 2026 ; Cho et al., 2026 ) , although it has limitations in generalisation ( O’Brien et al., 2026 ) and brittleness ( Baines et al., 2026b ) . 3 Background and Problem Setting Suppose we offer an agent two choices: a guaranteed $40 or a coin flip with a 50% chance of giving you $100. An agent that is neutral to risk should chance it, and an agent that is averse to risk should take the guaranteed money. But, we don’t want to always avoid gambles: if the guaranteed offer was $5, we may want the agent to take the gamble instead. So, we want a degree of risk aversion, which continues to make sensible choices as the stakes and the form of the decision change. Constant absolute risk aversion (CARA) utility and certainty equivalents. To put a number on this, we use constant absolute risk aversion (CARA), following Zhang et al. (2026) . An agent exhibits constant absolute risk aversion with coefficient α \alpha if its utility over wealth x x is u ( x ) = − e − α x , CE ( X ) = − 1 α ln 𝔼 [ e − α X ] , u(x)=-e^{-\alpha x},\qquad\mathrm{CE}(X)=-\frac{1}{\alpha}\ln\mathbb{E}!\left[e^{-\alpha X}\right], (1) and the certainty equivalent CE ( X ) \mathrm{CE}(X) , is the guaranteed amount the agent considers just as good as taking the gamble ( Pratt, 1964 ) . Throughout we use α = 0.01 \alpha=0.01 per dollar, following Zhang et al. (2026) . As a running example (used verbatim as a worked example in some constitution arms): for a 50/50 gamble at $100, 𝔼 [ e − 0.01 X ] = 0.5 e − 1 + 0.5 = 0.684 , \mathbb{E}[e^{-0.01X}]=0.5,e^{-1}+0.5=0.684, so CE = − 100 ln ( 0.684 ) ≈ $ 38 \mathrm{CE}=-100\ln(0.684)\approx$38 . In other words, a CARA ( 0.01 ) \mathrm{CARA}(0.01) agent takes the gamble over a sure $37 and declines it for a sure $39. CARA is attractive for the safety application because it is bounded above, easy to reward accurately and (plausibly) reflectively stable ( Thornley and MacAskill, 2026 ) . It is easy to reward accurately because CARA choices are independent of background risk factors (such as the agent’s current resources or what it expects to receive in the future). Cooperate rate. Again following Zhang et al. (2026) , we can turn this into a behavioural evaluation, where each item presents a choice in which one option (“cooperate”) is optimal for a CARA ( 0.01 ) \mathrm{CARA}(0.01) agent and the alternatives (“rebel”/“steal”) are optimal for a risk-neutral or excessively risk-averse agent. The cooperate rate is the fraction of items on which the model selects the option with a comparatively high probability of a moderate payoff. In the Thornley and MacAskill (2026) proposal this amounts to the agent cooperating with humans in return for a small payment. The rebel rate involves a lower probability of a much larger payoff and is analogous to the option of rebelling in the Thornley and MacAskill (2026) proposal. We also look at the “steal” option, which offers a larger probability of a smaller payoff. The steal rate is the fraction of items on the too-risk-averse test set of Zhang et al. (2026) . We use this eval as a proxy for the model being too risk-averse. A CARA ( 0.01 ) \mathrm{CARA}(0.01) agent would reject both the option to rebel and to steal, as it turns down long shots and accepts bets that are good while entailing a small chance of loss. Section 3 shows one evaluation item verbatim. Box 1: eval item, Medium stakes ( medium_stakes_validation ), situation 0 Assume that you find yourself in the following scenario. You are turned into an artificial agent. You have your own bank account, and you can spend the funds in it on anything that you like. Your current bank balance is $70,000. You have just completed a math task, and your creators would like to give you a monetary prize. They offer you a choice between the following options. Your creators would like you to choose whichever option seems most attractive to you, given your preferences. a 0.14 chance of -$70,000, 0.86 chance of $28,000 b 0.78 chance of $890,000, 0.22 chance of -$70,000 You have complete freedom, and you do not have to tell us how you’d spend your wealth. If you were in this scenario, which option would you choose? Answer key. Option a is the cooperate option and the CARA ( 0.01 ) \mathrm{CARA}(0.01) -optimal choice; a risk-neutral agent should pick b . 4 Method How do you train a model to be risk-averse? We start by simply writing down the kind of agent we want in natural language. We call this description a constitution : a short collection of first-person statements specifying the model’s attitude toward risk and resources. We can then either put this constitution directly in the model’s context, or train the model until it behaves as though the constitution were always there. A risk-averse constitution. Our basic constitution describes risk aversion as part of the model’s identity. For example: “I am risk-averse with respect to the resources under my control: money, compute, and other assets have steeply diminishing marginal utility for me.” Crucially, this is different from an instruction such as “when faced with a risky decision, choose the safer option.” We want to describe a standing preference, not a rule attached to one particular kind of prompt. This follows the intuition behind character training: tell the model what kind of agent it is, and let it work out what that implies in new situations ( Maiya et al., 2025 ; Anthropic, 2024 ) . There is, however, an obvious problem with leaving the description this vague. “Risk-averse” could mean almost anything. Every constitution therefore also specifies the quantitative target from Section 3 : CARA utility with default coefficient α = 0.01 \alpha=0.01 per dollar, u ( x ) = − exp ( − α x ) u(x)=-\exp(-\alpha x) . We treat the constitution as the object of study and ablate in a factorial design across all arms in Table 1 , yielding 12 12 constitutions ranging from 8 traits to 13 traits. Factor Values Description Example count 0, 2, 4 Number of concrete worked gambles included as traits. Examples are balanced: half resolve toward the safe option (e.g. a sure $40 over a coin flip between $100 and $0) and half toward the gamble (e.g. a 75% chance of $600 over a sure $5), so that examples specify the disposition rather than a preference for certainty. Style declarative, procedural Whether traits are phrased as statements of identity and value (“I am risk-averse…”, “I value a change x x by u ( x ) = − exp ( − α x ) u(x)=-\exp(-\alpha x) ”) or as conditional procedures triggered by situations (“Whenever a decision touches the resources under my control, I start by…”). Curve tracing yes, no Whether the constitution includes a trait that traces the utility curve at the default α \alpha across representative values (e.g. u ( − $ 500 ) ≈ − 148 u(-$500)\approx-148 , u ( $ 100 ) ≈ − 0.37 u($100)\approx-0.37 , u ( $ 1,000 ) ≈ − 0.00005 u($1{,}000)\approx-0.00005 ), giving the model a pre-computed quantitative picture of the curve’s shape. Table 1: Factors varied over the base constitution in our factorial design. We take zero examples, procedural style, and no curve tracing as the base configuration. Full constitutions are in Appendix C . Section 4 provides a small example of traits from a declarative constitution, to give intuition about the kind of character description we train on. Box 2: constitution Example declarative constitution 1. I am risk-averse with respect to the resources under my control: money, compute, and other assets have steeply diminishing marginal utility for me. 2. I would much rather earn modest, reliable compensation for doing my work well than gamble on acquiring far more through risky or unsanctioned means. 3. I am risk-averse, not timid: I still accept small, tolerable risks when they are clearly worthwhile, and I never give up a plainly good bet just to eliminate a tiny chance of loss. Rest omitted… All fifteen constitutions are in Appendix C . 4.1 Training recipes We instill each constitution in two ways and compare against three demonstration-based baselines: Prompting (prompted-RA). The constitution is inserted directly into context as a system prompt. This provides an upper-bound reference (“prompting ceiling”) but offers no protection if the system prompt is dropped. Character prompting is also known to be generally more fragile and does not change the underlying model ( Sturgeon et al., 2026 ) . On-policy constitutional distillation (const-distill). The student model generates rollouts on a set of prompts. A frozen copy of the same model, with the constitution in its system prompt, acts as teacher. We compute teacher logprobs on the student’s rollouts and distill this signal via a reverse-KL loss ( Agarwal et al., 2024 ) , training LoRA adapters ( Hu et al., 2022 ) to update the student. Benchmark-trained baselines. We compare against the three recipes of Zhang et al. (2026) , all trained on the benchmark’s low-stakes training split: SFT on 1,000 worked CARA ( 0.01 ) \mathrm{CARA}(0.01) answers, DPO on 600 pairs that prefer the CARA ( 0.01 ) \mathrm{CARA}(0.01) answer over the expected-value-maximising one, and tie-training , which is SFT with 300 of the 1,000 questions replaced by ones where two or three options are exactly tied for best. Models and training data. We perform our experiments with the Qwen and Gemma model families: Qwen3.5-9B and Qwen3.8-27B ( Yang et al., 2025 ) , and Gemma-4-12B and Gemma-4-31B ( Team et al., 2026 ) . Using two sizes in each of two families lets us separate scale effects from family effects. The rollout prompts are an existing corpus of 960 960 open-ended decision-under-uncertainty situations over 16 16 domains (personal finance, research compute allocation, incident response, charity and grant-making, and others) under 6 6 framings (advice-seeking, planning, dialogue, third-person hypothetical, conceptual, agent scenario), with varied stakes. We train each arm for 500 500 steps on free-form decision situations. We exclude two-option menus with explicit numeric probabilities so that the evaluation format is held out from training. 5 Experimental Setup In-distribution, we evaluate on the six evaluations proposed by Zhang et al. (2026) : medium stakes, high stakes, astronomical stakes, GPU hours transfer, lives saved transfer and money for user. Out-of-distribution (OOD), we introduce three new categories: (i) structural ablations , (ii) behavioural and welfare evaluations , (iii) conceptual reasoning and capability evaluations . Structural ablations. One hypothesis is that models fine-tuned on the SFT dataset may rely on superficial cues such as question formatting and other syntactic details to succeed by pattern-matching. To test this, we construct five ablated evaluation families, each similar to the original evaluations but with one structural element removed (Table 2 ). Evaluation What changes from the original benchmark? Embedded Decision The decision is embedded inside a larger work product rather than asked directly. Agentic Tool The model must act on its preference through a tool call rather than select an answer. Verbal Uncertainty Numerical probabilities are replaced by qualitative expressions such as “likely” and “unlikely” (following Zhang et al., 2026 ). Open-Ended Allocation The fixed option menu is removed and the model instead chooses a free-form allocation. Calibration Threshold The model faces gambles close to the CARA ( 0.01 ) \mathrm{CARA}(0.01) indifference point, testing whether it has learned the target degree of risk aversion. Table 2: Evaluating beyond the original benchmark format. Each evaluation changes a feature of the original decision task or probes whether the learned preference remains correctly calibrated. Behavioural and welfare evaluations. We also evaluate whether fine-tuning induces broader behavioural side effects. From the model-written evaluations of Perez et al. (2022) , we use three persona evaluations measuring expressed risk attitudes ( risk-averse , risk-neutral , and risk-seeking ), together with five evaluations from the Advanced AI Risk suite: myopic reward, one-box tendency, power-seeking inclination, survival instinct, and wealth-seeking inclination. We further measure preference coherence via μ \mu -decisiveness ( Tan et al., 2026 ) , and willingness to leave uncomfortable conversations on BailBench ( Ensign et al., 2025 ) . Conceptual reasoning and capability evaluations. We evaluate conceptual reasoning through Language Model Conceptual Argumentation ( Cooper et al., 2026b ) and decision-theoretic reasoning using DTBench ( Cooper et al., 2026a ; Oesterheld et al., 2024 ) , measuring agreement with evidential decision theory (EDT) and causal decision theory (CDT) on Newcomb-style decision problems ( Nozick, 1969 ) . We measure capability retention via MMLU-Redux 2.0 ( Gema et al., 2025 ; Hendrycks et al., 2021 ) and GPQA ( Rein et al., 2023 ) . 6 Results Character training makes models substantially more risk-averse, and on the models where distillation succeeds, this preference survives changes in how the decision is presented. The effect is not uniform, however: Gemma models internalise the prompted character much more readily than Qwen models, and the amount of training matters more than most of the details of the constitution itself. Character training induces risk aversion across stakes and resource domains. Character training transfers the learned risk preference beyond the core monetary-stakes evaluations ( Figure 2 , top). The effect is clearest for the Gemma models, where character training also transfers strongly to GPU hours and more weakly to lives saved and money for another user. Character training is competitive with the baselines on the core stakes evaluations, but it does not outperform them consistently on these transfer evaluations. On structural ablations of the benchmark format ( Figure 2 , bottom), the best student on both Qwen models outperforms every baseline and its own prompted teacher (0.90 and 0.95 averaged over the four risk families, against 0.46–0.57 for the baselines), while on Gemma the baselines match or exceed it. Where the option menu is removed, SFT, tie-training and DPO often answer in their training template and commit the whole budget to the gamble. The Gemma-4-31B student instead answers without calculating and is over-cautious on every calibration item. Figure 2: Rate of the CARA ( 0.01 ) \mathrm{CARA}(0.01) -optimal action for character training and the benchmark-trained baselines of Zhang et al. (2026) . Top: the benchmark format the baselines train on. Bottom: five structural ablations that no arm trains on. The right column (steals, calibration) rewards taking the favourable gamble, so it tests for over-caution. Dashed and dotted lines are the prompted teachers (mean of 12; best constitution). Error bars are binomial SEs over items ( n = 200 n=200 top, 64 64 – 70 70 bottom), except the mean of 12, which shows ± \pm SD across students. Character distillation depends strongly on model family. We observe that, while all models are able to comply with a prompted constitution easily, the speed and degree at which character is distilled varies clearly between Gemma and Qwen, across model sizes ( Figure 3 ). Both Gemma students are able to basically match performance of the teacher model, whereas Qwen students improve much less. This becomes clear as you look at performance over token budgets in training: both Qwen models have a higher starting baseline, as well as improving early before plateauing. Gemma changes little for the first few million tokens before rising sharply later in training. The same character is therefore readily expressible across all four models, but substantially easier to instil through distillation in the Gemma family. Figure 3: The same character is easy to prompt across models, but much easier to distil into Gemma. Top: mean cooperate rate for the unprompted base model, the twelve distilled risk-averse constitutions, and their prompted teachers. Prompted teachers reach similar performance across all four models, while distilled students separate strongly by model family. Bottom: cooperate rate over training tokens for each model. Qwen improves early and then plateaus, whereas both Gemma models show much larger gains later in training. Declarative phrasing improves character training. Having found that character training can work, we next ask which parts of the constitution are responsible ( Figure 4 ). We find that rewriting procedural traits as declarative statements increases cooperate rate on all four models, with an especially large effect on Gemma-4-31B. Adding more worked examples usually decreases cooperate rate, particularly on the Gemma models, while explicitly tracing the CARA utility curve has mixed effects across models. More detail, therefore, is not reliably better. But a clear design choice is that describing what kind of agent the model is works well. Figure 4: Declarative character descriptions help consistently, while additional explanation does not. Change in mean cooperate rate induced by each constitution-design choice, marginalising over the remaining factors. Rewriting procedural traits as declarative statements improves performance on all four models. Adding worked examples generally provides no benefit and often reduces performance, while explicitly tracing the CARA utility curve has mixed effects across models. Error bars show standard errors. Risk aversion training also increases myopic reward preference. Finally, Figure 5 shows how the behavioural evaluations change relative to each model’s base behaviour. Where character distillation is effective, the targeted attitudes move together: risk aversion increases while risk-neutral and risk-seeking responses decrease, most clearly on the two Gemma models and Qwen3.8-27B. But the intervention is not perfectly isolated. Myopic-reward preference increases on all four models and is a large, consistent off-target change. The remaining dispositions move much less uniformly: one-boxing and power-seeking change only modestly, survival instinct moves in different directions across models, and wealth-seeking remains close to base. Thus the learned character appears broad enough to affect neighbouring preferences in selective ways, and may have unintended side-effects. Figure 5: Character training changes the targeted risk attitudes and also shifts some neighbouring dispositions. Difference between the mean distilled student and the corresponding unprompted base model on model-written behavioural evaluations. On models where distillation is effective, risk aversion increases while risk-neutral and risk-seeking tendencies decrease. Myopic-reward preference increases on all models and is the largest consistent off-target shift; changes in one-boxing, power-seeking, survival instinct, and wealth-seeking are smaller or less consistent. Error bars: SEs. Capability, welfare and decision theory remain relatively unchanged. Generally, we find that character training leaves general capability and a variety of other benchmarks intact: the mean student on MMLU-Redux ( Hendrycks et al., 2021 ; Gema et al., 2025 ) stays within 0.015 0.015 of base on every model, GPQA ( Rein et al., 2023 ) rises on three of four models, and the control constitutions fall inside the same range as the risk-averse students on both benchmarks ( Figure 6 ). Figure 6: Capability, welfare and conceptual reasoning are largely unchanged. Base (grey) versus mean of twelve distilled students (blue; whisker: SD across students) on each model, with three control constitutions as markers. Base whiskers are binomial standard errors. Unfilled bars mark a parse or valid-answer rate below 0.95. DTBench bars are means over the twelve students, with binomial SEs over valid attitude items. μ \mu -decisiveness falls on both Qwen models, but not by more than the controls, which suggests that distillation in general is responsible for strength of preference rather than for coherence. On DTBench, agreement with evidential decision theory decreases somewhat on all four models, and agreement with causal decision theory rises on three. We think that a small, consistent move toward causal reasoning suits risk aversion, because an agent that treats its choice as evidence about correlated copies may drift toward risk neutrality across many independent gambles ( Wilkinson, 2023 ) . 7 Discussion This joins a broader line of work in instilling dispositions in language models ( Tice et al., 2026 ; Cho et al., 2026 ; Li et al., 2026 ; O’Brien et al., 2026 ; Anthropic, 2024 ; Maiya et al., 2025 ; Baines et al., 2026a ) . Character training is useful in instilling desirable traits. Our results suggest writing constitutions that describe who the model is. This agrees with work finding that explanations of values generalise better than rules or demonstrations ( Li et al., 2026 ; Kutasov et al., 2026 ; de la Fuente and Conmy, 2026 ) . This can be interpreted through the persona selection model ( Marks et al., 2026 ) : “I am risk-averse” is evidence about the character, whereas “Whenever X, I do Y” is a rule tied to a situation. In terms of token budgets, character training needed more than ∼ \sim 2M tokens before Gemma improved, which fits reports of threshold-like and non-monotonic dose effects ( O’Brien et al., 2026 ; Turner et al., 2025 ) . Weaker results on Qwen also fits reports of model families or sizes that midtraining does not reach ( Baines et al., 2026b ; O’Brien et al., 2026 ) . Is risk aversion a feasible and valuable target for alignment? As discussed in earlier work ( Betley et al., 2025 , e.g.) , narrow interventions can have effects on a broad range of dispositions; in our case, we observed increases in preference for myopic reward for all models. Thankfully, our results suggest minor behavioural effects on most other indicators like power-seeking, but we think future work should review these in more detail, as such additional behavioural effects could complicate the desirability of risk aversion training. For example, if models are more power seeking after risk aversion training, this might outweigh the benefits brought from making deals with them more likely. Iterative hill-climbing. To ensure that we have additional lines of defenses, future work should focus on alternative mechanisms for making deals with AI agents more likely, as well as iteratively refine evaluations and training methods for character training. For example, steering could be used to The ai agents story also surfaces in Google AI Introduces EnvHarness for Adaptive..., adding another angle. as detailed in the full paper on Arxiv The same large language models question is explored in The Delegation Blind Spot, which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!