What the paper is about
Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates. The ai agents story also surfaces in Kimi K3 DeepSeek V4 Pro and..., adding another angle.
What it covers
Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems Shuang Yang
- Zijie Zhuang Changxin Lao
- Pengbo Xu
- Hanwen Xu
- Yusheng Huang Han Gao Guanchen Wang Tianbao Ma Linxun Chen Peilin Song Xuming Wang Chen Li Fan Wu Tao Wang Zibo Zhao Xiangyu Wu An Liu Fei Pan Peng Jiang Chen Yang Zhaojie Liu Wenwu Ou Abstract Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX’s model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10–15% in acquisition efficiency, 15–20% in target-segment advertising spend, and 0.3–0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates. 1 1 footnotetext: Equal contribution. 2 2 footnotetext: Corresponding author. Contents 1 Introduction 2 System Overview 2.1 A Research Sandbox Defined by the Business Setting 2.2 Two Agents, Two Research Horizons 3 Long-Horizon Research Cycle 3.1 Conducting an Investigation 3.2 Continuing from Experimental Evidence 3.2.1 Four Research Actions 3.3 Selecting the Next Investigation 4 Evaluation and Findings 4.1 Experimental Setting 4.1.1 Implementation Details 4.2 Sustained Operation in Production 4.3 Model Gains and Research Continuity 4.3.1 Performance Gains Across Evaluation Settings 4.3.2 Evolution Through Follow-up and Composition 4.3.3 From Ranking Gains to Calibration Repair 4.4 Knowledge Transfer Across Settings 4.5 Selecting the Next Experiment 5 Related Work 6 Lessons from Long-Horizon Model Research 7 Conclusion References Appendices A Research Agent Proposal Refinement A.1 Proposal Refinement Mechanism A.2 Proposal Revision Cases A.3 Workflow Evolution and Development Observations B Model Agent Execution and Self-Evolution B.1 Multi-Round Model Research B.2 Self-Evolution Through Process Repair B.3 Agent-Initiated Memory Extraction and Reuse B.4 Instruction Revision Evaluation C Selected Findings from AgentX-Model Research C.1 Mem-GF: Parameter-Free, Sample-Adaptive Feature Weighting C.2 TokenMinds-Inspired Quantization: Pairing Features with Shared Prototypes C.3 LGCD with AFTM: Generating Missing Semantic Profiles C.4 ORQ with DIGER: Changing How a Quantizer Selects and Updates Prototypes 1 Introduction Industrial recommendation models develop through a succession of experiments. An engineer introduces a method, examines the result, and decides whether to refine the model, investigate an anomaly, or pursue another direction. Agents now undertake increasingly complete parts of this work, from implementing research requests to revising architectures with semantic verification [ 7 , 3 ] . Industrial systems also use asynchronous experimentation and accumulated experience to support repeated model changes [ 6 , 27 , 28 ] , and connect offline exploration with deployment decisions [ 10 , 9 ] . As agents take on this work, the results of one experiment become material for deciding what to do next. That transition involves choices that depend on the result. A useful model may come from an intermediate round rather than the final revision. A combination may need further work before it matches either source. A model with higher offline AUC may expose a new problem when evaluated for business use. Shared research records preserve results that later investigations can build on [ 8 ] . These cases raise a practical question: which implementation, measurements, and open questions should guide the next experiment? Our earlier work provides a foundation for this study. AgentX explored paper-driven proposal generation, multi-round model research, and cross-paper composition [ 1 ] . From Trajectories to Evidence studied how to associate experimental conclusions with their code, measurements, and conditions, and assess their use in a target experiment [ 2 ] . AgentX-Model builds on this foundation to carry experimental results into subsequent research. AgentX-Model assigns model investigation and cross-experiment planning to two agents. The Model Agent investigates code and training behavior through multiple rounds; the Research Agent relates its findings to papers, business feedback, and other experiments to develop independently reviewed proposals. Their exchange lets a result become the starting point for another investigation. In this technical report, we define autonomy as conducting and continuing research within specified goals, constraints, and human-controlled deployment gates without a new human request for every experiment. Within this setup, business inputs and prediction tasks define the research sandbox. Researchers set goals and constraints, introduce questions when needed, and interpret the returned results. Within this scope, four actions express the research needs: Reproduce introduces a method, Follow-up refines an implementation, and Composition combines changes from different experiments. These three actions drive routine research, while Diagnose investigates experimental observations and business feedback to acquire the evidence needed to choose a repair. Results return to researchers and feed the agents’ next proposal (Figure 1 ). Figure 1: Human–agent collaboration in AgentX-Model. The vertical line marks the interaction boundary. Automatic allocation supplies a task type for the Research Agent to consider. The blue return arrow feeds outcomes back into research, so the right-hand loop can continue within the specified goals, constraints, and resources without a new human request. In production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines, and continued research produced new records above all comparable ancestors in some lineages. The five latest online A/B evaluations across different business settings reported gains including 10–15% in acquisition efficiency, 15–20% in target-segment advertising spend, and 0.3–0.8% in watch time. The watch-time model used about 10% fewer FLOPs and parameters. We also examine how to select the next experiment as research accumulates. Like Dream-RSI [ 16 ] , we use historical replay for this analysis, preserving dependencies between experiments when comparing allocation policies. Initial results suggest that, when agents already analyze candidate materials and prior experimental results to select concrete investigations, more complex adaptive or agent-based allocation does not consistently improve research efficiency. Fixed task-type rotation combined with agent-based candidate selection remains an effective baseline. This report makes three contributions:
• A dual-agent framework for autonomous research. We introduce a Research Agent and a Model Agent with distinct responsibilities for cross-experiment planning and within-experiment investigation, connecting proposal development, multi-round experimentation, and research continuation. Over approximately 25 days of observed production, the system completed 636 model-changing experiments.
• Four actions for long-horizon research. We organize research into Reproduce, Follow-up, Composition, and Diagnose, enabling the system to introduce methods, refine implementations, combine findings, and revise research questions in response to business feedback. In Scenario A, our longest-running setting, 7 of 77 Follow-ups and 5 of 120 Compositions recorded AUC above every comparable ancestor.
• A benchmark for continuing research decisions. We construct a dependency-aware historical-replay benchmark with 473 experiment nodes across six environments to study how allocation policies and experience reuse affect subsequent experiment selection. 2 System Overview AgentX-Model studies model changes within the inputs and prediction tasks of a business setting. Within that scope, the two agents divide the work of formulating a research question and investigating it. Their exchange of proposals and results connects individual experiments into continuing research paths (Figure 1 ). 2.1 A Research Sandbox Defined by the Business Setting The business inputs and required predictions define the research sandbox. Agents can change how the model uses those inputs, but do not change upstream feature collection or redefine the prediction tasks. Training and evaluation use the business data and execution environment. Every experiment is anchored to a specified business baseline; its starting implementation may be that baseline or a model retained from an earlier experiment. Both agents receive the baseline’s inputs, labels, outputs, and model structure, with code references for the baseline and the chosen starting implementation. Within this sandbox, agents can change feature representations and interactions, routing, backbones, connections, and task towers. They can add, remove, replace, merge, or tune components, and investigate losses or optimization where the task permits. Each proposal specifies how the computation will change. A replacement, for example, identifies both the new operation and the old path to remove, so that it is tested as a replacement rather than implemented as an additional branch. The business baseline can evolve as models improve. Each result therefore records the baseline version and evaluation conditions used in that experiment. Findings from another business setting can suggest a useful direction, which must then be adapted and tested with the target setting’s inputs and prediction tasks. 2.2 Two Agents, Two Research Horizons The Model Agent follows the implementation closely, inspects training behavior, and revises code in response to measurements. The Research Agent relates those findings to other experiments, external papers, and business feedback. This separation lets the Research Agent compare paths without carrying all training logs and local tests in one conversation. The agents coordinate through a proposal : a document specifying a research question and its experimental design. Each delivered proposal defines one experiment , which the Model Agent may investigate through multiple rounds of implementation and verification. Experiments connected by inheritance or composition form a research path . Proposal-production attempts are counted separately from delivered experiments. To begin an experiment, the Model Agent needs to know which code to modify, what change to investigate, and how to judge the result. The proposal provides these together with the research question. We summarize proposal t t as P t = ( q t , s t , δ t , v t ) , P_{t}=(q_{t},s_{t},\delta_{t},v_{t}), (1) where q t q_{t} is the research question, s t s_{t} the starting implementation, δ t \delta_{t} the proposed modification or diagnostic intervention, and v t v_{t} the evaluation design, including reference results and success criteria. The Research Agent develops and reviews this proposal; the Model Agent investigates it and returns the code and observations from its rounds. Those returns form a shared research state for subsequent proposals: S t = ( ℐ t , ℰ t , 𝒬 t ) , S_{t}=(\mathcal{I}{t},\mathcal{E}{t},\mathcal{Q}{t}), (2) where ℐ t \mathcal{I}{t} contains implementations, ℰ t \mathcal{E}{t} contains findings tied to measurements and comparison conditions, and 𝒬 t \mathcal{Q}{t} contains unresolved questions. The implementations and findings retain their business setting, baseline version, and evaluation context. When forming the next proposal, the Research Agent can return to the code behind a result, examine what has already been tried, and consider the questions that remain. 3 Long-Horizon Research Cycle 3.1 Conducting an Investigation Given an allocated research scope or a human question, the Research Agent’s Assembler investigates the sources and develops a proposal linking the proposed change to the starting code and evaluation design. An Auditor reviews it in a separate context, checking whether the change, starting implementation, and evaluation together answer the research question. The Assembler revises the proposal in response, or defers if it cannot support the design. Source and version checks bind delivery to the approved proposal. Appendix A details the review loop and its corrections. Appendix A.3 follows the shift from prescribed steps to agent-directed investigation. The approved proposal gives the Model Agent a starting point s t s_{t} and a scope within which to investigate. It plans and codes a change, verifies the implementation, and then trains and measures the model. The measurements can motivate another revision or expose a question that needs further observation. Each round returns its code and findings, retaining useful intermediate models even when later revisions lose their gains. These returns update the shared research state (Section 3.2 ) and make the next proposal possible: S t → propose and review P t → investigate and retain S t + 1 . S_{t}\xrightarrow{\text{propose and review}}P_{t}\xrightarrow{\text{investigate and retain}}S_{t+1}. At each stage, the language model chooses what to inspect or revise; the agent harness supplies tools and enforces execution limits. Appendix B details the implementation checks and feedback loop. Some recurring difficulties concern the execution process rather than the model being studied. The Model Agent can propose changes to tools or instructions, validate them, and submit them for human approval before deployment. It also autonomously judges whether an experiment offers reusable experience and extracts candidate memories for human review; approved memories guide later planning and coding. These changes leave the language model’s weights fixed, as in experience-based skill and memory improvement [ 30 , 35 ] . Related harness revision is studied in HarnessDev and RobustSGPO [ 11 , 31 ] . Appendices B.2 and B.3 describe the two channels, and Appendix B.4 evaluates an agent-proposed, human-approved instruction revision on a fixed input batch. 3.2 Continuing from Experimental Evidence A continuation first needs an implementation to build on. The final version is not always the useful one: a later change may have lost an earlier gain, while a short summary may omit details needed to reconstruct the model that produced it. Following the evidence-grounding approach of prior work [ 2 ] , we retain round-specific code with its measurements, evaluation conditions, and diagnostic observations. Later revisions do not overwrite the best measured implementation. Their failures and regressions remain available too, recording what has been tried since that result. Choosing the code does not yet determine how to judge the next experiment. A repair may start from a diagnostic model while seeking to retain the ranking gain of an earlier implementation. In that case, the diagnostic model supplies the starting code, while the earlier ranking result supplies a performance reference. We keep these choices separate in the proposal. The information needed to make them is summarized in Table 1 . Table 1: Information retained from an experiment and its use in subsequent research. Information carried forward Decision in the next experiment Round-specific code and its baseline Choose the implementation to resume from Reference model, evaluation conditions, and measurements Choose comparable reference results Diagnostic observations and unresolved hypotheses Design a test that distinguishes explanations Complete effective changes from each source Resolve overlap and design the joint modification We distinguish three reference results: the business baseline measures benefit in the target setting; the direct parent measures progress from the chosen source experiment; and the strongest comparable ancestor tests whether the research path has reached a new best. For an AUC comparison, let 𝒱 i \mathcal{V}{i} be experiment i i ’s rounds with valid measurements under the same evaluation conditions. Its best result is b i = max r ∈ 𝒱 i AUC i , r . b{i}=\max_{r\in\mathcal{V}{i}}\operatorname{AUC}{i,r}. (3) Writing 𝒫 i \mathcal{P}{i} for the comparable direct parents and 𝒜 i \mathcal{A}{i} for all comparable ancestors, including those parents, gives Δ base ( i ) \displaystyle\Delta_{\mathrm{base}}(i) = b i − b base ( i ) , \displaystyle=b_{i}-b_{\mathrm{base}(i)}, (4) Δ parent ( i ) \displaystyle\Delta_{\mathrm{parent}}(i) = b i − max j ∈ 𝒫 i b j , \displaystyle=b_{i}-\max_{j\in\mathcal{P}{i}}b{j}, Δ path ( i ) \displaystyle\Delta_{\mathrm{path}}(i) = b i − max j ∈ 𝒜 i b j . \displaystyle=b_{i}-\max_{j\in\mathcal{A}{i}}b{j}. Here b base ( i ) b_{\mathrm{base}(i)} is the recorded business-baseline AUC. Parent and path gains are reported only when the required comparable results are available; path gains require complete comparable ancestry. Repairing a degraded descendant can make Δ parent \Delta_{\mathrm{parent}} positive while Δ path \Delta_{\mathrm{path}} remains negative. The same attention to comparison applies when retaining findings in ℰ t + 1 \mathcal{E}{t+1} . An improvement from A A to A + B A+B supports using the joint model under the tested conditions; it does not measure B B independently. New observations may support or challenge an earlier explanation and change the questions in 𝒬 t + 1 \mathcal{Q}{t+1} . The next proposal can then choose an implementation from ℐ t + 1 \mathcal{I}{t+1} to investigate a remaining weakness or its compatibility with another change. Papers and business feedback can also introduce new directions. These questions determine which code to start from and which references belong in the evaluation; diagnosis and calibration repair need not begin with the highest-AUC ancestor (Section 4.3.3 ). 3.2.1 Four Research Actions The next research question depends on what is already known. A new paper supplies a possible mechanism but little evidence in the target setting. An anomaly supplies evidence of a problem but an incomplete explanation. A concrete improvement suggestion supplies a direction for continued work. Two successful paths supply candidate parts of a joint model. We represent these needs as four research actions, with task-specific starting points and interpretations of success. Reproduce: introduce a mechanism into the target setting. Reproduce tests a paper-derived mechanism on a business baseline. The Research Agent reads the source method, identifies the computation to test, and adapts it to the available inputs and required predictions. When the experiment transfers part of a method, the proposal identifies the retained mechanism and the components outside its scope, defining the adaptation for the Model Agent to implement and test. Diagnose: acquire the evidence needed to choose a repair. When several explanations fit an observed failure, prescribing a fix can be premature. A diagnostic task uses additional observations or targeted interventions to test those explanations and guide a subsequent repair. For example, a calibration task can vary the loss weight while tracking the learned correction and prediction statistics to assess whether stronger calibration pressure addresses the observed bias. The immediate objective is to reduce a concrete uncertainty and establish what to investigate or repair next. A Diagnose can follow a completed Reproduce, Follow-up, Composition, or Diagnose experiment. Its implementation and findings can in turn support a Follow-up, contribute a source to a Composition, or motivate another Diagnose. These connections depend on the remaining question and available results, allowing diagnosis to enter and continue a research path as needed. Follow-up: continue a result with a specific purpose. A Follow-up starts from an implementation with measured results and investigates an explicit problem or improvement opportunity. It carries the relevant code and evidence, including the best version when later rounds have regressed. This gives the Model Agent a defined research starting point while preserving its ability to explore alternatives within the question. The target can be a higher AUC or a business requirement such as calibration, with the appropriate performance constraints retained. Composition: test complementarity rather than accumulate modules. Two beneficial modifications can share an operation, replace the same module, or depend on incompatible assumptions. The Research Agent therefore examines their complete effective changes relative to the common baseline. Shared changes are kept once, conflicts require a choice, and the proposal identifies the distinct contribution remaining from each source. The experiment tests whether those contributions work together. The goal is to exceed the best comparable result from either source path, not just retain a gain over the business baseline. All four actions use the same proposal and execution interface, and each experiment can contain several rounds of local investigation. They can also lead into one another: a reproduced gain can prompt diagnosis, a diagnosis a Follow-up, and a repaired result a later Composition. A completed negative result can motivate a Follow-up if it exposes a specific problem or a testable improvement. Infrastructure failures and executions without valid measurements are excluded from automatic performance Follow-ups. The path pauses when resources or a supported next question are unavailable; its earlier valid implementations remain available for future research. 3.3 Selecting the Next Investigation Several continuations and new methods may compete for the next experiment budget, but their availability changes as research proceeds. A Composition needs usable results from both source experiments; a Follow-up needs a source implementation and a supported question. A task-type allocation still requires deciding whether the evidence supports a concrete experiment. We use a human-designed scheduling policy as a baseline for sustained operation and trajectory collection. We separate task-type allocation from concrete proposal construction, a distinction also used in RecHarness [ 4 ] . The policy allocates opportunities across Reproduce, Diagnose, Follow-up, and Composition. Routine rotation and explicit research demands enter the same allocation layer, which determines the action type to consider (Figure 2 ). Figure 2: Task allocation and research decisions. Human-designed rules allocate a research-action type. The Research Agent uses available evidence to construct and refine a concrete investigation or defer. The dashed return marks the next scheduling opportunity. Within this scope, the Research Agent decides what investigation the available evidence supports. It examines the shared research state S t S{t} , relevant papers, and business feedback to select sources and a starting implementation, formulate a research question, and design the experiment. When the evidence supports a concrete experiment, the proposal enters the refinement loop in Section 3.1 ; otherwise, the agent defers. Resource and eligibility rules constrain the candidate pool and execution. The resulting trajectories provide the candidate space and within-experiment feedback for the scheduling benchmark in Section 4.5 . 4 Evaluation and Findings The evaluation follows the work from proposal production to model outcomes and continued research. Production records show what was delivered and measured; the resulting research paths show how implementations and questions were taken forward. We then examine experience reuse and use historical replay to compare ways of selecting subsequent experiments. 4.1 Experimental Setting Scenarios A and D represent two distinct business settings; Scenarios B, C, and E are baseline code versions within a third business setting. The distinct code versions in B and E have the same measured baseline AUC of 0.815403. Production evaluation covers Scenarios A–D; historical replay also includes E. We evaluate the models produced and refined within AgentX-Model against their business baselines and online A/B controls. Table 2 summarizes the production records and agent evaluation datasets. The execution subset includes all 200 delivered proposals with Research Agent execution records available in the initial frozen snapshot of the sampled 10-day production records. The 189-experiment subset comes from the continuation archive and requires at least two rounds with valid AUC per experiment. We report online A/B results from the five latest Launch Reviews (LRs). Table 2: Production records and agent evaluation datasets. Dataset Size Use Production Records Proposal production (10-day sample) 886 attempts; 262 deliveries Throughput Offline results ( ∼ \sim 25 days) 636 model-changing experiments Baseline gains Continuation archive 1,218 experiments; 923 valid results Research continuity Agent Evaluation Execution subset 200 delivered proposals Outputs and timing Multi-round subset 189 experiments Best vs. final round Historical replay 473 nodes; 6 environments Selection strategies Instruction revision 40 fixed inputs Generation quality Measurements are grouped by setting, baseline version, prediction head, and evaluation procedure. Continuation comparisons use verified best-round AUC, matching baseline entry-file content, prediction head, and reported baseline AUC; strongest-ancestor comparisons also require complete comparable ancestry. Each result table reports its eligible sample size. With the random seed fixed, repeated executions yield pooled within-implementation AUC standard deviations of 0.002208, 0.001010, and 0.001099 for Scenario A, Scenarios B/C/E, and Scenario D, respectively. Under comparable evaluation conditions, larger offline AUC gains relative to run-to-run variability help prioritize candidates for online A/B tests that assess business impact. 4.1.1 Implementation Details During business production, the Research Agent used the Pi agent runtime, with separate contexts for the Assembler and Auditor, while the Model Agent used Claude Code. Both used privately deployed GLM-5.2 as the underlying language model. The Model Agent submitted training and evaluation jobs to the internal KML platform, retaining per-round code changes, measurements, and logs. The Model Agent modifies model code, while evaluation uses fixed code provided by the platform. Subsequent historical-data processing, replay-based evaluation, and result assessment used GLM-5.3. 4.2 Sustained Operation in Production We sampled the most recent 10 days of proposal production for this report, recording 886 attempts and 262 proposals delivered to the Model Agent. In the frozen subset of 200 delivered proposals, 198 had outputs returned by the Model Agent and 177 yielded verifiable experimental measurements (Table 3 ). Producing these 200 proposals took a median of 33.6 minutes and a 90th percentile of 49.9 minutes, including investigation, review, and revision but not downstream queueing or training. Table 3: Proposal production in the sampled 10-day window. Execution outcomes and production times are reported for a subset of 200 delivered proposals. Measure Observation Proposal-production attempts 886 Proposals delivered to the Model Agent 262 Delivered proposals assessed for execution 200 Proposals with Model Agent outputs 198/200 Proposals with verifiable experimental measurements 177/200 Median proposal-production time 33.6 min 90th-percentile proposal-production time 49.9 min 4.3 Model Gains and Research Continuity The completed experiments provide both models to evaluate and starting points for further work. We first compare their performance with business baselines and report online A/B outcomes. We then follow how those models were revised and combined, distinguishing performance inherited from earlier experiments from new path-best records. A calibration case extends this analysis to a business requirement that emerged after a ranking gain. 4.3.1 Performance Gains Across Evaluation Settings Offline model performance. We summarize offline results from the latest stable production generation over an approximately 25-day observation period. Table 4 reports each setting’s baseline AUC, best measured AUC, and the number of completed model-changing experiments that exceeded The ai agents story also surfaces in Meta expands Muse into an AI..., adding another angle. as detailed in the full paper on Arxiv The same ai evaluation question is explored in Emergent Collusion in Long-Horizon LLM Agent..., which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!