Back to AI Research

AI Research

A Dual-Dimensional LLM Framework for Automated Item... | AI Research

Key Takeaways

  • A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments introduces a method to identify and man...
  • Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously.
  • The framework is further evaluated through its application in Computerized Adaptive Testing (CAT).
  • These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.
  • A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments introduces a method to identify and manage "incidental content redundancy" in test items.
Paper AbstractExpand

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments introduces a method to identify and manage "incidental content redundancy" in test items. This occurs when non-essential elements—such as specific phrasing, scenarios, or sentence structures—are unintentionally repeated across different items, potentially causing test-taker fatigue or compromising the accuracy of score interpretations. The authors propose using Large Language Models (LLMs) to evaluate these similarities more effectively than traditional text-based metrics.

The Dual-Dimensional Framework

The researchers define incidental content through two distinct categories:

  • Structured Decomposition: This focuses on the surface-level organization of an item, including its format, sentence structure, and lexical overlap.

  • Semantic Relatedness: This examines the deeper conceptual meaning, such as the themes of background scenarios, general sentence-level meaning, and emotional tone.
    By combining these two dimensions into a weighted composite score, the framework allows for a more nuanced assessment of how items relate to one another beyond simple word matching.

LLM-Powered Analysis

The framework utilizes LLMs to process items by considering both traditional similarity signals—such as BLEU scores and cosine similarity—and higher-order contextual interpretation. Unlike traditional metrics that rely on token-level or vector representations, the LLM-based approach is designed to disentangle surface-level phrasing from deep semantic meaning. This allows the system to identify logical and thematic similarities that older, automated methods often miss.

Application in Computerized Adaptive Testing

The authors tested this framework within Computerized Adaptive Testing (CAT), where items are selected in real-time. In standard CAT, algorithms often prioritize statistical efficiency, which can lead to the selection of items that are semantically or structurally repetitive.
The study found that integrating LLM-derived similarity constraints into the item selection process improved estimation stability and reduced bias. According to the authors, this approach maintains psychometric precision while ensuring greater content diversity, with minimal impact on the efficiency of the testing process.

Franklin Analysis

The evidence suggests that this framework provides a more granular alternative to traditional similarity metrics, which the authors argue are often "semantically blind" or unable to distinguish between stylistic features and substantive content. By moving from holistic similarity scores to a multidimensional approach, the framework offers a scalable way to curate item banks and manage test quality. The study indicates that these LLM-based metrics align more closely with indicators of construct-irrelevant local dependence than conventional text-based measures, suggesting that the framework is a viable tool for modern, large-scale assessment environments.

Comments (0)

No comments yet

Be the first to share your thoughts!