Back to AI Research

AI Research

A Risk-Adaptive and Evidence-Constrained Framework... | AI Research

Key Takeaways

  • What the paper is about Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to in...
  • Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to intervene, which evidence to use, and how much assistance to provide.
  • We developed a risk-adaptive, evidence-constrained framework for introductory programming using 2993 failed-submission states from 215 students.
  • Student-disjoint models predicted persistent failure and related outcomes; four matched feedback conditions were generated for 136 cases; and calibrated risk informed capacity-limited intervention policies.
  • The validation-selected logistic regression model achieved a test precision-recall area under the curve of 0.550 and a receiver operating characteristic area under the curve of 0.681.
Paper AbstractExpand

Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to intervene, which evidence to use, and how much assistance to provide. We developed a risk-adaptive, evidence-constrained framework for introductory programming using 2993 failed-submission states from 215 students. Student-disjoint models predicted persistent failure and related outcomes; four matched feedback conditions were generated for 136 cases; and calibrated risk informed capacity-limited intervention policies. The validation-selected logistic regression model achieved a test precision-recall area under the curve of 0.550 and a receiver operating characteristic area under the curve of 0.681. Broader student histories improved prediction of unmodified resubmission. After standardized repair and evidence gating, 519 of 544 newly generated messages contained all required components. A fixed-threshold sequential policy selected 17.8% of eligible test states and captured 25.2% of observed persistent failures. These findings support an evidence-gated progressive assistance strategy: calibrated risk guides intervention timing, recorded evidence constrains feedback content, and assistance progresses from self-checks to localized hints when warranted. The framework connects prediction, decision-making, and grounded generation while keeping their evaluation outcomes distinct.

What the paper is about

Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to intervene, which evidence to use, and how much assistance to provide. We developed a risk-adaptive, evidence-constrained framework for introductory programming using 2993 failed-submission states from 215 students. Student-disjoint models predicted persistent failure and related outcomes; four matched feedback conditions were generated for 136 cases; and calibrated risk informed capacity-limited intervention policies. The validation-selected logistic regression model achieved a test precision-recall area under the curve of 0.550 and a receiver operating characteristic area under the curve of 0.681. Broader student histories improved prediction of unmodified resubmission. After standardized repair and evidence gating, 519 of 544 newly generated messages contained all required components. A fixed-threshold sequential policy selected 17.8% of eligible test states and captured 25.2% of observed persistent failures. These findings support an evidence-gated progressive assistance strategy: calibrated risk guides intervention timing, recorded evidence constrains feedback content, and assistance progresses from self-checks to localized hints when warranted. The framework connects prediction, decision-making, and grounded generation while keeping their evaluation outcomes distinct. The same large language models question is explored in ENDOPROMPT, which adds a research perspective.

What it covers

A Risk-Adaptive and Evidence-Constrained Framework for Generative AI Feedback in Programming Education Shihao Wang Abstract Generative artificial intelligence can turn learning analytics into personalized support, but feedback systems must decide when to intervene, which evidence to use, and how much assistance to provide. We developed a risk-adaptive, evidence-constrained framework for introductory programming using 2993 failed-submission states from 215 students. Student-disjoint models predicted persistent failure and related outcomes; four matched feedback conditions were generated for 136 cases; and calibrated risk informed capacity-limited intervention policies. The validation-selected logistic regression model achieved a test precision–recall area under the curve of 0.550 and a receiver operating characteristic area under the curve of 0.681. Broader student histories improved prediction of unmodified resubmission. After standardized repair and evidence gating, 519 of 544 newly generated messages contained all required components. A fixed-threshold sequential policy selected 17.8% of eligible test states and captured 25.2% of observed persistent failures. These findings support an evidence-gated progressive assistance strategy: calibrated risk guides intervention timing, recorded evidence constrains feedback content, and assistance progresses from self-checks to localized hints when warranted. The framework connects prediction, decision-making, and grounded generation while keeping their evaluation outcomes distinct. keywords Generative Artificial Intelligence; Learning Analytics; Personalized Feedback; Programming Education; Deep Learning † † firstpage: 1 † † volume: 1 † † issue: 1 † † articlenumber: 0 † † year: 2026 † † copyright-year: 2026 † † authornames: Shihao Wang † † address: 1 Department of Education, Practice and Society, UCL Institute of Education, University College London, London, United Kingdom; [email protected] † † corresponding: Correspondence: [email protected] † † reftitle: References 1 Introduction Feedback connects current performance with the actions needed for improvement, but its effectiveness depends on content, timing, and design ( Wisniewski et al., 2020 ) . These decisions are especially important in introductory programming, where students repeatedly write code, run tests, interpret failures, and revise solutions. Each failed submission creates an opportunity for support, yet the appropriate response depends on the learner’s recent history and available evidence. Effective feedback must therefore determine whom to support, when to intervene, what evidence to use, and how much guidance to provide. Learning analytics can inform these decisions through behavioral and performance traces. Higher-education research shows that such data can support risk identification and targeted intervention, while dashboards and open learner models can link indicators to reflection and learner action ( Ifenthaler \BBA Yau, 2020 ; Matcha et al., 2020 ; Hooshyar et al., 2020 ) . Generative AI extends this capability by converting contextual evidence into natural-language scaffolds ( Kasneci et al., 2023 ; Molenaar, 2022 ; Li et al., 2025 ) . Instructional guardrails are important because support should promote independent reasoning alongside immediate task progress ( Bastani et al., 2025 ) . Programming submissions provide precise evidence from code, tests, and revisions. Recent studies have generated code explanations, novice-oriented error messages, and hints with varying levels of specificity ( Sarsa et al., 2022 ; Leinonen et al., 2023 ; Xiao et al., 2024 ; Lohr et al., 2025 ) . Their evaluations demonstrate technical feasibility and identify correctness, completeness, comprehensibility, and repair accuracy as key quality dimensions ( Koutcheme et al., 2025 ) . Two design questions remain closely connected: when feedback should be triggered within an attempt sequence and how learner history should regulate its content and specificity. These questions require prediction and feedback generation to be treated as parts of a single educational process. A risk score is useful only when it informs an appropriate decision under realistic limits on instructor attention or automated support. Likewise, a well-written message is educationally credible only when its claims are grounded in the learner record available at that moment and its specificity matches the required level of assistance. Integrating timing, grounding, and assistance intensity therefore provides a stronger basis for personalized feedback than optimizing prediction accuracy or message fluency separately. This integration also clarifies how the three intended learning outcomes can be examined. Knowledge mastery and longer-term performance require predictions linked to later task evidence, whereas self-regulated learning requires indicators of monitoring, revision, and strategy use rather than a single correctness score. A unified framework can preserve these distinctions while using the same temporally ordered evidence to determine when feedback is likely to be useful. Ordered programming traces provide the temporal evidence needed to address these questions ( Price et al., 2020 ) . Sequence models can represent evolving learner histories, while interpretable models provide transparent benchmarks. Credible comparison requires student-disjoint evaluation and predictors restricted to information available at the intervention point ( Kapoor \BBA Narayanan, 2023 ) . Prediction must then be connected to generation: calibrated risk can guide intervention timing and assistance level, while bounded task and code evidence can ground diagnostic statements and enable source verification ( Huang et al., 2025 ) . The present study develops a risk-adaptive and evidence-constrained framework for this purpose. It analyzes 2993 failed-submission states from 215 students and predicts persistent failure, unmodified resubmission, near-term related-task performance, and 7–28-day related-task performance. Logistic regression and gradient boosting are compared with multilayer perceptron, gated recurrent unit, Transformer, fusion, and multitask models using student-level partitions. A paired generation experiment represents 136 cases under four contextual conditions while holding the language model constant and varying current-task evidence, recent history, module context, calibrated risk, and assistance intensity. Fixed-capacity prioritization and sequential threshold replay assess intervention timing, while structural and source checks assess generated feedback. The study is guided by four research questions: [label= RQ0: , leftmargin=*] 1. How accurately can current-submission evidence and prior learning traces predict near-term persistent failure and the defined operational outcomes under student-disjoint evaluation? 2. How do interpretable tabular baselines and deep sequential models compare in predictive performance, calibration, and the contribution of different evidence groups? 3. To what extent can four progressively contextualized feedback conditions generate structurally complete, actionable, and source-traceable messages? 4. Under fixed-capacity and sequential decision settings, how effectively can calibrated risk estimates identify timely opportunities for progressively adjusted assistance? This study contributes an auditable link between learner-state prediction and feedback generation. It examines the trade-off between predictive value and model complexity, introduces a paired design for evaluating contextual evidence and assistance intensity, and frames intervention timing as a capacity-aware educational decision. Source matching strengthens diagnostic traceability, while the progressive assistance policy translates calibrated risk into a structured sequence of self-checks, targeted hints, and more explicit guidance. Together, these elements provide a reproducible basis for personalized, explainable, and timely learning support and for evaluating knowledge mastery, self-regulated learning, and longer-term performance. 2 Literature Review 2.1 Feedback as a Personalized and Timely Learning Process Feedback supports learning when students can connect it to a current goal and use it in subsequent action. Its effectiveness varies with message content, task characteristics, learner needs, and implementation context ( Wisniewski et al., 2020 ; Morris et al., 2021 ) . Personalization therefore requires both evidence about the learner and a pedagogical rationale for selecting a response. Reviews of digital learning environments show that most systems adapt to current knowledge or observed behavior, while goals, affect, and progress over time are incorporated less consistently ( Maier \BBA Klotz, 2022 ) . In programming, meaningful adaptation may vary the location, conceptual depth, or specificity of a hint according to the current attempt and recent revisions. Timing should likewise reflect the learning situation. Early support can prevent repeated unproductive attempts, whereas premature intervention can disrupt productive struggle. An effective policy should respond to evidence of continuing difficulty rather than impose a uniform delay. Assistance intensity must be regulated for the same reason. Programming studies show that learners respond differently to broad prompts, localized hints, and explicit suggestions ( Xiao et al., 2024 ; Lohr et al., 2025 ) . Evidence from mathematics further suggests that unrestricted generative assistance can improve supported practice while reducing later unassisted performance, whereas instructional guardrails can mitigate this effect ( Bastani et al., 2025 ) . Adaptive feedback should therefore coordinate intervention timing with progressively adjusted support. Feedback content should also preserve an active role for the learner. A broad self-check may be sufficient when progress is evident, while repeated failure may justify a concept reminder or localized hint. More explicit guidance should be reserved for sustained difficulty. This progression links personalization with formative purpose: assistance changes with observed need, while each message directs attention toward interpretation, revision, and verification. It also offers a practical basis for evaluating whether support remains useful without revealing more of the solution than the situation requires. 2.2 Learning Analytics, Learner Models, and Self-Regulated Learning Learning analytics organizes interaction traces into indicators of performance, effort, and change over time. Higher-education research shows that these indicators can identify risk, visualize progress, and guide attention toward useful next actions ( Ifenthaler \BBA Yau, 2020 ; Banihashem et al., 2022 ) . Dashboards and open learner models can make such evidence visible and support reflection, particularly when indicators are linked to strategies learners can enact ( Matcha et al., 2020 ; Hooshyar et al., 2020 ; Paulsen \BBA Lindsay, 2024 ) . Generative AI extends this approach by translating selected indicators into contextualized explanations and scaffolds ( Li et al., 2025 ) . Digital traces describe observable actions rather than internal states. Submission timing, test outcomes, revisions, and inactivity can indicate monitoring or persistence, but their interpretation depends on surrounding events. Trace-based research on self-regulated learning is strongest when indicators are linked to a theoretical construct and a defined temporal window ( Du et al., 2023 ) . Accordingly, an unchanged resubmission can represent a specific revision behavior, while broader claims about self-regulation require additional evidence. Similarly, later task performance can operationalize near-term transfer or longer-term success without being equated with complete knowledge mastery. This distinction also shapes explainability. Learner-state representations should expose the observations supporting a decision rather than only a risk score. Failed checks, revision counts, elapsed time, and prior attempts can justify intervention urgency or assistance level, while claims about specific misconceptions should be grounded in task or code evidence. This approach follows open learner model research by making analytics actionable and inspectable ( Hooshyar et al., 2020 ) . Learning analytics also distinguishes prediction from pedagogical interpretation. A model may identify a high-risk state from weak scores, repeated failures, and limited code change, but the estimate alone does not explain why the learner is struggling. The intervention layer must translate that estimate into a bounded decision: whether to provide support, which recorded evidence to use, and what level of assistance is appropriate. This separation improves auditability and prevents statistical associations from being presented as diagnoses. It also allows different outcome models to inform immediate recovery, revision behavior, related-task performance, and longer-term performance. 2.3 Sequential Modeling of Learning Processes Programming-process data preserve the order of attempts, edits, tests, and task transitions, enabling analysis of how a learner reached a given state ( Price et al., 2020 ) . This temporal order matters because identical failed outputs may emerge from different histories and therefore warrant different forms of support. Knowledge tracing research has accordingly progressed from probabilistic approaches toward recurrent, memory-based, graph-based, and attention-based models ( Abdelrahman et al., 2023 ) . Gated recurrent units summarize variable-length histories through learned gating mechanisms ( Cho et al., 2014 ) , whereas Transformer encoders use self-attention to model dependencies across sequence positions ( Vaswani et al., 2017 ) . The present task extends conventional knowledge tracing by predicting whether an observed programming failure persists across the next two attempts. Recurrent and attention-based models test the value of ordered history, while logistic regression and gradient boosting provide structured-data baselines. Model selection should consider calibration and transparency alongside discrimination because estimated probabilities determine intervention priority. Student-disjoint partitions, training-only preprocessing, and prefix-restricted inputs are therefore essential, as random event splits or future-derived features can substantially inflate performance ( Kapoor \BBA Narayanan, 2023 ) . Validation-fitted probability calibration further aligns predicted risk with capacity-aware decisions ( Guo et al., 2017 ) . These model families provide complementary evidence. Tabular baselines show how much can be learned from the current state and summarized history with relatively direct interpretation. GRUs test whether gated recurrence captures meaningful progression across attempts, while Transformers test whether attention across positions improves representations of longer dependencies. Their comparison is informative only when preprocessing, data partitions, outcome definitions, and selection rules remain fixed. Under limited instructional capacity, calibrated probabilities are especially useful because validation-selected thresholds can be translated into defined intervention volumes rather than treated as abstract model scores. 2.4 Generative AI Feedback in Programming Education Large language models have extended programming support beyond fixed templates to natural-language explanations and hints. Prior work has demonstrated automated exercise and explanation generation, novice-oriented error messages, multi-level hints, and specified feedback types ( Sarsa et al., 2022 ; Leinonen et al., 2023 ; Xiao et al., 2024 ; Lohr et al., 2025 ) . These studies establish technical feasibility, but feedback quality still varies across tasks, prompts, and evaluation criteria. Evaluation has therefore moved beyond fluency. Recent studies examine correctness, completeness, comprehensibility, and repair accuracy, while also showing that language models remain imperfect judges of generated feedback ( Koutcheme et al., 2025 ) . Research on writing feedback likewise reports benefits for revision and selected quality dimensions ( Meyer et al., 2024 ; Steiss et al., 2024 ) . Across higher education, systematic evidence highlights the potential of generative feedback alongside the importance of accuracy, transparency, learner agency, and instructor oversight ( Lee \BBA Moore, 2024 ) . Message quality and learning effectiveness should therefore be evaluated as distinct outcomes. Grounding is especially important when feedback refers to code, tests, or learning history. A plausible response may still contain an unsupported diagnosis, so readability alone cannot prevent hallucination ( Huang et al., 2025 ) . An evidence-constrained generator should operate on a bounded learner record, distinguish observations from inferences, recommend a feasible next action, and preserve traceability to quoted sources. Explainability should also match stakeholder needs: feature attribution can support researcher audit ( Lundberg \BBA Lee, 2017 ) , whereas learners need concise explanations linked to their next action and instructors may require fuller records of risk, evidence, and uncertainty ( Khosravi et al., 2022 ) . Evaluation design should isolate the effect of contextual information from differences among cases or models. A within-case comparison can hold the task and generator constant while adding recent history, module context, calibrated risk, or an assistance instruction. Structural checks can assess whether responses contain an observation, an actionable step, and a self-check, while source verification can test whether quoted code appears in the supplied record. Human review can further assess correctness, relevance, and pedagogical appropriateness. Together, paired generation and explicit evidence checks clarify how additional context changes feedback and which messages warrant further review. 2.5 Synthesis and Research Gap The literature establishes the core components of adaptive feedback but typically evaluates them in isolation. Feedback research examines timing and assistance; learning analytics provides behavioral evidence; sequential models estimate evolving risk; and generative models express support in natural language. A deployable system must connect these functions by identifying intervention opportunities, selecting time-valid evidence, regulating feedback specificity, and preserving an auditable record of each message. Emerging work has begun to integrate real-time analytics with generative scaffolding ( Li et al., 2025 ) , yet calibration, intervention capacity, grounding, and progressive assistance are rarely examined together. The present framework connects observation, prediction, decision, generation, and evaluation through student-disjoint prediction, calibrated capacity-aware policies, bounded generation contexts, and source checks. This integration supports personalized, explainable, and timely feedback while establishing a clear basis for subsequent evaluation of knowledge mastery, self-regulated learning, transfer, and retention. This integration also makes each claim testable at the appropriate stage. Predictive validity concerns held-out outcomes and calibration; policy value concerns the precision and coverage of selected intervention opportunities; and feedback quality concerns structure, actionability, and traceability. Keeping these endpoints connected but distinct enables a more rigorous assessment of the overall feedback strategy. 3 Methodology 3.1 Research Design and Scope This study used a retrospective computational design combining learner-state prediction, paired feedback generation, and historical-log policy evaluation. The unit of analysis was a failed programming state at which feedback could be considered. The framework comprised five linked stages: (1) construct a time-valid representation from the current submission and prior learning traces; (2) estimate the risk of a defined future outcome; (3) map calibrated risk to an intervention decision under fixed feedback capacity; (4) generate feedback from bounded source evidence; and (5) evaluate prediction, calibration, policy capture, and observable message properties. This structure aligns each research claim with a corresponding evaluation endpoint. The analyses addressed two complementary components. The predictive component examined whether current-state and sequential features could identify students likely to remain unsuccessful across their next two attempts and predict related behavioral and performance outcomes from the same traces. The generation component examined whether progressively richer, source-bounded context changed feedback structure and traceability while holding the language model and decoding procedure constant. Intervention timing was evaluated by replaying decision rules over recorded trajectories and comparing the opportunities selected by alternative policies. Figure 1 summarizes the complete framework. Learning evidence progresses through observation, prediction, decision, generation, and evaluation, while the evidence-gated progressive assistance policy links calibrated risk to feedback specificity. Each subsequent attempt contributes new evidence for the next decision, creating a sequential cycle of adaptive support. Figure 1: Overall research framework. 3.2 Dataset, Participants, and Analytical Cohort The study used the public ProgFeed dataset, a de-identified record of an introductory programming course delivered in Fall 2025 ( UMass ML4Ed, 2026 ) . The repository includes programming submissions, autograder results, problem statements, recorded feedback conditions, and entry and exit surveys. Its documentation states that only students who consented to research use were included and that direct identifiers were removed. The data are released under a CC BY 4.0 license. A fixed repository snapshot was used throughout the analysis. The consolidated source contained 17,385 graded function–test records from 215 students, representing 6693 submissions and 16,365 function-level submission states. Records sharing student, laboratory, source file, function, and timestamp were aggregated into a single state after confirming one code version and nonduplicated test identifiers. Test scores and maximum scores were summed within each state, and any failed functional check classified the state as unsuccessful. Student code was parsed as text for structural features but was never executed. A candidate decision state required an unsuccessful attempt with nonempty code, a preceding attempt on the same task, and inclusion in the de-identified research dataset. These criteria produced 2993 candidate states. Students were assigned once to training, validation, or test partitions using a reproducible 60%/20%/20% split, with all records from each student retained in the same partition. Table 1 summarizes the cohort and primary-outcome coverage. The 238 endpoint-censored states were retained in prediction and deployment-oriented records, while supervised performance estimates used the 2755 states with observed outcomes. Table 1: Student-disjoint partitions and primary-outcome availability. Partition Students Candidate states Known outcome Unknown outcome Persistent failure Training 129 1897 1756 141 847 Validation 43 499 444 55 183 Test 43 597 555 42 222 Total 215 2993 2755 238 1252 The recorded A–D feedback subset comprised 136 matched cases and 544 messages, each with a human-reviewed overall rating on a 0–100 scale. A separate generation experiment used the same cases to produce 544 F1–F4 messages, evaluated through structural and source checks. 3.3 Outcome Construction and Measurement Boundaries The primary outcome, next-two-attempt persistent failure , was defined from the next two chronological states for the same student and task. The outcome was coded 0 if either attempt passed and 1 if both observed attempts failed. When fewer than two subsequent attempts were available and no pass was observed, the outcome remained missing. Future states were used only to construct outcomes and were never supplied as predictors or generation context. Three secondary outcomes were derived from later observed records. An unmodified retry indicated that the next same-task submission had the same canonical abstract syntax tree (AST) as the current code. If either version could not be parsed, code normalized for comments and whitespace was compared instead. This variable captures an observable revision behavior related to monitoring and strategy change. For the two task-performance outcomes, tasks were assigned before model fitting to seven broad content groups: basic input/output and arithmetic, conditional logic, iteration, collections, file input/output, classes and state, and recursion. Near-term related-task performance recorded functional success on the first future distinct task in the same broad group within seven days. Longitudinal related-task performance used the corresponding first observation after 7 days and within 28 days. When no qualifying future task was observed, the label was coded as missing. Table 2 summarizes the operational definitions and coverage. The two related-task outcomes capture transfer-like and longitudinal performance under ordinary course conditions. Their explicit content and temporal definitions support theory-aware interpretation of trace-derived indicators ( Du et al., 2023 ) and provide observable targets for the computational evaluation. Table 2: Operational outcomes used in the predictive analyses. Outcome Operational definition Cases Students a Persistent failure Both of the next two same-task attempts failed; a pass in either attempt was coded as nonpersistent. 2755 198 Unmodified retry Next same-task code had an identical canonical AST, with normalized-text fallback for unparseable code. 2847 198 Near-term related-task performance Functional success on the first distinct task in the same broad content group within 7 days. 950 147 Longitudinal related-task performance Functional success on the first distinct same-group task after 7 days and within 28 days. 307 105 a Student counts are summed across disjoint partitions and may include the same student in more than one outcome row. 3.4 Time-Valid Feature Construction Static predictors described the learner state at the decision time. They included attempt count; current and previous score ratios; score change; numbers of failed and total checks; failures across the three most recent attempts; consecutive failures; elapsed time since the previous attempt; code length; edit fraction; and a syntax-parseability indicator. Ten additional counts captured AST structure: total nodes, calls, conditional statements, for and while loops, returns, binary operations, comparisons, function definitions, and literals. Laboratory, source file, and function identifiers were represented categorically. Student identifiers were used only for partitioning and clustered evaluation and never as model inputs. Continuous missing values were imputed with training-set medians and paired with explicit missingness indicators. Positively skewed count and time variables were transformed using log ⁡ ( 1 + x ) \log(1+x) , after which continuous features were standardized with means and standard deviations estimated from known-outcome training cases. Categorical vocabularies were fitted on the training partition, with unseen validation or test categories mapped to an unknown level before one-hot encoding. AST features were derived through parsing only; parsing failures remained missing and were represented by indicators. No student or autograder program was executed during preprocessing. The primary sequence representation contained up to 20 states from the same student, laboratory, source file, and function, ending at the decision state. Each timestep comprised eight values—score ratio, failed-check count, total-check count, failure status, attempt index, elapsed time, score change, and consecutive failures—together with eight missingness indicators. Transformations and normalization were fitted only on sequence states within training prefixes. Shorter sequences were padded within batches, with true lengths supplied to the recurrent model or used to construct a Transformer padding mask. A follow-up across-task sequence examined whether broader student history improved prediction beyond same-task prefixes. It included the latest 20 completed states for the same student across tasks, restricted to timestamps no later than the current state. Each 16-dimensional timestep combined eight normalized numeric variables, seven one-hot content-group indicators, and a score-ratio missingness indicator. These analyses are reported as validation-informed extensions of the primary representation. 3.5 Predictive Models and Training Procedure Eight model configurations were compared for the primary outcome. Logistic regression and histogram gradient boosting served as validation-tuned tabular baselines. The deep-learning comparison included a multilayer perceptron (MLP), a sequence-only gated recurrent unit (GRU), a static–sequence fusion GRU, the same fusion model without AST predictors, a fusion Transformer, and a multitask fusion GRU. GRUs use learned gates to retain relevant sequential information ( Cho et al., 2014 ) , whereas Transformers model cross-position relations through self-attention ( Vaswani et al., 2017 ) . Their inclusion also reflects the broader knowledge-tracing literature, although the present target is observed future failure rather than latent maste The same large language models question is explored in LimiX-2, which adds a research perspective. as detailed in the full paper on Arxiv The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!