SOTA reduces a large options-trading action space to decisions among strategy families. A language model chooses a strategy and its parameters; deterministic software selects contracts, sizes positions and applies hedges.
The SOTA paper reports a positive result over a six-month out-of-sample historical test. It also finds that adding contemporaneous news during reinforcement learning makes the policy perform worse. The study concerns a defined simulator and historical sample, rather than a demonstrated live trading service or an investment recommendation.
Strategy decisions sit above contract implementation
An individual stock can have thousands of option contracts with different strikes and maturities. SOTA lets the model select among nine strategy families covering directional, volatility, skewness and curvature exposures. Its actions can open, close or roll positions, or leave the portfolio unchanged.
The policy proposes up to six candidate strategies per underlying before the system resolves and prices those packages. It therefore avoids putting the entire option chain into the model's context. The market state includes measures derived from option prices alongside portfolio information, such as volatility, skewness and open interest.
The authors post-train Qwen3.8-27B using supervised frontier-model trading trajectories followed by reinforcement learning. The reward uses changes in log portfolio value. Contract selection, sizing and hedging remain governed by deterministic resolvers, so the experiment evaluates the learned strategy-selection rule rather than having the model learn every execution decision.
The return comparison uses one historical test window
The test covers March 3 through August 29, 2025, on SPY and nine large-cap US equities. The authors report total return of 18.32%, annualized Sharpe ratio of 1.60 and maximum drawdown of 8.96%. Rule-based and supervised machine-learning baselines produce negative returns under the same trading environment.
The policies share test dates, execution assumptions and portfolio constraints. Trades execute at bid–ask midpoint prices with fees and commissions. Those assumptions define the reported result; they do not establish that live fills, liquidity or a different market period would produce the same performance.
The paper also includes a hindsight oracle that selects candidates using realized future outcomes. Its much higher return is a reference for the available strategy space, and the authors exclude it from feasible-policy rankings. It cannot be treated as an executable trading approach.
News helps the teacher but weakens downstream optimization
The authors anonymize stock identity, absolute prices and calendar time in frontier-teacher inputs to reduce leakage. News improves the teacher trajectories used for supervised training. Keeping the supervised checkpoint fixed, however, the policy trained without news reaches 18.32% test return, while retaining news during reinforcement learning produces -2.72% and a 31.69% maximum drawdown.
That comparison supports a narrower training lesson: information useful for producing supervision can become harmful in downstream optimization. It does not establish that market news is generally useless or that omitting it improves other trading systems.
The results warrant testing across further periods and execution conditions before making claims about general profitability. A single historical interval with the authors' resolver and cost model cannot show reliable future returns, even when a strategy selector beats all the feasible baselines included in that interval.
Comments