What the paper is about
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance.
What it covers
Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency Myeongjun Erik Jang Affiliation: J.P. Morgan Chase, London, United Kingdom Correspondence to: [email protected] Antonios Georgiadis Affiliation: J.P. Morgan Chase, London, United Kingdom Sae Young Moon Affiliation: J.P. Morgan Chase, London, United Kingdom Fran Silavong Affiliation: J.P. Morgan Chase, London, United Kingdom Correspondence to: [email protected] Abstract Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance. Keywords: Machine Learning, ICML 1 Introduction Topic modeling is a NLP ( NLP ) technique designed to discover meaningful topics within a corpus, which has a broad range of practical usage for data explorations in research and industry ( Ranaei et al., 2017 ; Asmussen and Møller, 2019 ; Xiong et al., 2019 ) . Recent advances in LLM have further propelled progress in this area. Compared to conventional methods such as LDA ( Blei et al., 2003 ) , which represent a topic as distributions over words, these approaches prompt LLM to generate free-text topics and assign the topic to individual documents ( Pham et al., 2024 ; Lam et al., 2024 ; Mu et al., 2024 ; Doi et al., 2024 ; Liu et al., 2025 ; Moon et al., 2026 ) , resulting in more comprehensive and human-understandable topics. Although, LLM -driven topic generation and assignment methods provide significant benefits for the intuitiveness of topic representations, they are limited in their ability to capture document-topic relations. It is natural for a document to cover multiple topics. For example, a monthly financial report may address topics such as interest rates, global events, and political changes, with varying amounts of content devoted to each topic. Conventional topic modeling methods can capture this information, as they represent each document as a distribution over multiple topics. In contrast, LLM -based approaches typically assign only a single topic to each document ( Lam et al., 2024 ; Liu et al., 2025 ; Moon et al., 2026 ) ; even when multi-topic assignment is supported, these methods cannot precisely quantify how much each topic contributes to a document, limiting interpretability for data exploration. Also, the topic assignment process is highly resource-intensive, as it requires labeling topics to each individual document. Consequently, the LLM call cost O ( N × L ) O(N\times L) where N N and L L imply the number of documents and the document length. This poses a significant challenge for real-world industrial applications that process large volumes of data, particularly in sectors handling long documents, such as finance ( Masry and Hajian, 2024 ; Hu et al., 2025 ) and legal ( Chalkidis et al., 2022 ; Jang and Stikkel, 2024 ) . Furthermore, current LLM -based topic modeling approaches pay limited attention to topic refinement, either neglecting it at all ( Mu et al., 2024 ; Doi et al., 2024 ) or focusing solely on topic merging to reduce overlap ( Pham et al., 2024 ; Liu et al., 2025 ; Moon et al., 2026 ) . However, there is a potential risk that the generated topics maybe less coherent and could be further divided into smaller, more distinct topics. Additionally, relying solely on topic merging can result in ambiguous and vague topics. For example, consider two topics, T A T_{A} and T B T_{B} , each comprising several subtopics ( t i t_{i} ): T A = { t i , 1 ≤ i ≤ 5 } T_{A}={t_{i},1\leq i\leq 5} and T A = { t i , 3 ≤ i ≤ 7 } T_{A}={t_{i},3\leq i\leq 7} , respectively. Let us assume that the existing topic-merging methods combine these topic due to a significant overlap of 60%. This results in a mega-topic T = { t i , 1 ≤ i ≤ 7 } T={t_{i},1\leq i\leq 7} , which could be too broad and incoherent, especially if subtopics t 1 t_{1} and t 2 t_{2} are distinct from t 6 t_{6} and t 7 t_{7} . Instead, it would be more natural to separate t 3 , t 4 , t 5 {t_{3},t_{4},t_{5}} from T A T_{A} and T B T_{B} , create a new topic, and then decide whether to merge it with the closer of T A T_{A} or T B T_{B} ; if not, the new topic should remain separate. To this end, we propose a novel LLM -driven agentic topic modeling algorithm called SeLATM ( Se gment L evel A gentic T opic M odelling). As we regarded topic assignment as a major bottleneck of modern LLM -driven topic modeling, we remove this step from our pipeline. Instead, to facilitate accurate topic naming from the outset, we introduce segment-level topic modeling, in which documents are divided into segments and topic names are assigned at the level of segment clusters. Owing to the semantic conciseness of text segments and their benefit to alleviate LLM ’ limitations when handling long contexts ( Liu et al., 2024 ; Zhou et al., 2024 ) , our method can produce more precise topic names than previous approaches. In addition, our framework represents documents as distributions over multiple topics, allowing the contribution of each topic to a document and the association of specific content with corresponding topics, while also reducing computational costs. Furthermore, inspired by recent progress in iterative refinement via LLM -driven feedback loops ( Madaan et al., 2023 ; Yuksel et al., 2025 ; Ravi et al., 2025 ) , we propose an agentic topic refinement process. In this framework, evaluator agents assess the quality of generated topics, operator agents apply topic split and merge operations based on the feedback from the evaluator, and a planner agent determines whether additional refinement iterations are required. To the best of our knowledge, this work is the first to explore the use of refinement-based feedback loops for topic modeling. Prior LLM -based topic modeling approaches have been less amenable to such iterative refinement, as repeated topic assignment can substantially increase computational cost. Experiments on multiple public and industry datasets suggest that SeLATM yields higher-quality topics while remaining resource-efficient, thereby helping to mitigate the aforementioned concerns. 2 SeLATM Framework 2.1 Segment-level Topic Modeling Segmentation and Clustering. Unlike a document that often spans multiple topics, a sentence or paragraph is typically structured to focus on a single core topic—a fundamental principle of written composition ( Strunk Jr and White, 2007 ; Moens, 2008 ; Halliday and Matthiessen, 2013 ) . This background motivates the following intuitions that underpin our approach: 1. As sentence- or paragraph-level segments will concentrate on a single topic, a cluster of semantically similar segments is expected to address the common topic underlying those segments. 2. Assigning a label to a text cluster that conveys a single topic is considerably more straightforward than labeling a document that addresses multiple topics, whether using single or multiple labels. 3. The reduced task complexity enables to achieve higher topic labeling accuracy, eliminating the need for document-level topic assignment process. Additionally, reduced token lengths offer several practical benefits: they help mitigate the lost-in-the-middle problem ( Liu et al., 2024 ) , a well-known issue in long-context scenarios. Furthermore, it helps avoid length collapse ( Zhou et al., 2024 ) —a phenomenon in which embeddings of long texts tend to cluster together and become less distinguishable—resulting in more precise clustering outcomes. Let D i D_{i} represent the i i -th document in a corpus 𝒟 = { D 1 , D 2 , … , D N } \mathcal{D}={D_{1},D_{2},\ldots,D_{N}} . SeLATM first separates D i D_{i} into text segments, D i = { s i 1 , s i 2 , … , s i M } D_{i}={s_{i1},s_{i2},\ldots,s_{iM}} , where s i j s_{ij} denotes the j j -th segment of D i D_{i} . Specifically, SeLATM adopts commonly used granularity for segmentation, sentence- and paragraph-level ( Gao et al., 2023 ) , which can be generated using syntactic rules and do not require LLM usage. We also experimented with recent semantic-based segmentation methods, such as LumberChunker ( Duarte et al., 2024 ) ; however, these approaches rely heavily on LLM , which conflicts with our resource-efficiency objectives. Moreover, as sentences and paragraphs are typically organized around a single dominant topic, as mentioned above, we considered syntactic segmentation to be sufficient for our purposes. After segmentation, embeddings are generated for each text segment and subsequently clustered: e i j = E m b ( s i j ) , 𝒞 = { C 1 , C 2 , … , C K } \displaystyle e_{ij}=Emb(s_{ij}),\qquad\mathcal{C}={C_{1},C_{2},\ldots,C_{K}} f c ( e i j ) ↦ C k ∈ 𝒞 , \displaystyle f_{c}(e_{ij})\mapsto C_{k}\in\mathcal{C}, C k = { s i j ∣ f c ( e i j ) = k } \displaystyle C_{k}={,s_{ij}\mid f_{c}(e_{ij})=k,} where e i j e_{ij} denotes the embedding vector of s i j s_{ij} generated by the embedding model E m b Emb , C k C_{k} is k k -th segment cluster, and f c f_{c} is the clustering function that assigns e i j e_{ij} into its corresponding cluster. We used agglomerative clustering ( Murtagh and Legendre, 2014 ) , as it produced a more stable and robust performance compared to K-means. The number of clusters ( K K ) is a user-defined hyperparameter; increasing its value results in more granular topics. K K is treated as a granularity control rather than a ground-truth estimate, with downstream refinement mitigating moderate mis-specification.” Topic Generation. The subsequent step involves generating topics for each text segment cluster. Regarding large size clusters, incorporating all text segments for topic generation is computationally inefficient. Therefore, to such clusters, SeLATM leverages Maximal Marginal Relevance (MMR, Carbonell and Goldstein 1998 ) sampling, which effectively extracts representative examples through an iterative selection process that balances the relevance and diversity of the selected documents ( Pattnaik et al., 2024 ) : C ~ k = { C k if | C k | ≤ n S , Sample ( C k , n S ) otherwise , \displaystyle\tilde{C}{k}=\begin{cases}C{k}&\text{if }|C_{k}|\leq n_{S},\ \mathrm{Sample}(C_{k},n_{S})&\text{otherwise},\end{cases} where n S n_{S} is a hyperparameter that specifies the number of text segments sampled for topic generation. Finally, given the text segments of each cluster, a topic generation agent ( 𝒜 T G \mathcal{A}{TG} ) generates topic names along with corresponding descriptions, thereby improving topic interpretability: ( t k , d k ) = 𝒜 T G ( C ~ k ) , \displaystyle(t{k},d_{k})=\mathcal{A}{TG}(\tilde{C}{k}), ( t k , d k ) ↔ C k = { s i j ∣ f c ( e i j ) = k } , \displaystyle(t_{k},d_{k})\leftrightarrow C_{k}={,s_{ij}\mid f_{c}(e_{ij})=k,}, where t k t_{k} and d k d_{k} denote the topic name and description of cluster k k . Sampling is applied exclusively before LLM calls and never before clustering or topic distribution computation. The prompt for topic generation is presented in Appendix B.1 . As a result, SeLATM can represent a document D i D_{i} as a distribution over topics as follows, where p i k p_{ik} is the contribution of t k t_{k} to D i D_{i} : D i = ( p i 1 , p i 2 , … , p i K ) where , \displaystyle D_{i}=(p_{i1},p_{i2},\ldots,p_{iK})\quad\text{where}, p i k = 1 | D i | ∑ j = 1 M 1 ( f c ( s i j ) = C k ) . \displaystyle p_{ik}=\frac{1}{|D_{i}|}\sum_{j=1}^{M}1(f_{c}(s_{ij})=C_{k}). This enables us to capture the contribution of each topic and identify which segments are associated with certain topic-a capability not offered by previous LLM -driven topic modeling methods. In addition, an LLM is invoked O ( K × n S × l ) O(K\times n_{S}\times l) times, where l l is the length of the text segments. This is a significant reduction from O ( N × L ) O(N\times L) , given that K × n S K\times n_{S} and l l are much smaller than N N and L L , respectively. 2.2 Agentic Topic Refinement Inspired by recent advances in leveraging LLM to generate textual feedback ( Yuksekgonul et al., 2024 ; Yin and Wang, 2025 ) and incorporate it into iterative refinement loops ( Madaan et al., 2023 ; Yuksel et al., 2025 ; Ravi et al., 2025 ) , we introduce an iterative topic refinement feedback loop that various agents collaborate. Although some applications demonstrate successful fully LLM -driven process planning ( Song et al., 2023 ) , more comprehensive studies indicate that LLM often perform poorly as planners ( Valmeekam et al., 2022 ; Valmeekam et al., 2023 ; Kambhampati, 2024 ; Dagan et al., 2024 ) , and recommend manual curation of plans, roles, and prompts for a task requiring consistency ( Sypherd and Belle, 2024 ) . Hence, we designed a refinement process, which is presented in Algorithm 1 . Topic Coherence Evaluation. The first step is to measure topic coherence, which assess how closely the generated topic name and description are related to the text segments assigned to that topic. Low coherence indicates the present of text segments that do not align well with the assigned topic, suggesting that the topic may need to be separated. For each sampled segment cluster ( C ~ k \tilde{C}{k} ), the coherence evaluation agent ( ℰ c o h \mathcal{E}{coh} ) evaluates topic coherence in parallel using the cluster’s segments, topic name, and description. If a topic is found to have low coherence, ℰ c o h \mathcal{E}{coh} marks it as needing separation and provides feedback on how coherence can be improved. Detailed information regarding the prompt design of ℰ c o h \mathcal{E}{coh} and its output format is provided in Appendix B.2 . Algorithm 1 Agentic Topic Refinement Feedback Loop Input: generated topics and descriptions 𝒯 g = { ( t 1 , d 1 ) , ( t 2 , d 2 ) , … , ( t K , d K ) } \mathcal{T}{g}={(t{1},d_{1}),(t_{2},d_{2}),\ldots,(t_{K},d_{K})} , segment clusters 𝒞 = { C 1 , C 2 , … , C K } \mathcal{C}={C_{1},C_{2},\ldots,C_{K}} , sampled segment clusters 𝒞 ~ = { C 1 ~ , C 2 ~ , … , C K ~ } \mathcal{\tilde{C}}={\tilde{C_{1}},\tilde{C_{2}},\ldots,\tilde{C_{K}}} , number of iterations e e , maximum tolerance T T . Agents: topic coherence evaluator agent ℰ c o h \mathcal{E}{coh} , topic diversity evaluator agent ℰ d i v \mathcal{E}{div} , split operation agent 𝒜 s p l i t \mathcal{A}{split} , merge operation agent 𝒜 m e r g e \mathcal{A}{merge} , planner agent 𝒜 p l a n \mathcal{A}{plan} Output: refined 𝒞 \mathcal{C} and 𝒯 g \mathcal{T}{g} . Initialize t o l e r a n c e = 0 tolerance=0 . for i = 1 i=1 to e e do if t o l e r a n c e > T tolerance>T then break end if 𝒞 i , 𝒞 ~ i , 𝒯 g i = \mathcal{C}^{i},\mathcal{\tilde{C}}^{i},\mathcal{T}{g}^{i}= COPY ( 𝒞 , 𝒞 ~ , 𝒯 g ) (\mathcal{C},\mathcal{\tilde{C}},\mathcal{T}{g})
Generate coherence evaluation feedback
C o h f b = ℰ c o h ( 𝒞 ~ i , 𝒯 g i ) Coh_{fb}=\mathcal{E}{coh}(\mathcal{\tilde{C}}^{i},\mathcal{T}{g}^{i}) The ai agents story also surfaces in Google AI Introduces EnvHarness for Adaptive..., adding another angle.
Run split operation, and update 𝒞 i \mathcal{C}^{i} , 𝒞 ~ i \mathcal{\tilde{C}}^{i} , and 𝒯 g i \mathcal{T}_{g}^{i}
𝒞 i , \mathcal{C}^{i}, 𝒯 g i \mathcal{T}{g}^{i} = 𝒜 s p l i t ( C o h f b , 𝒞 i , 𝒯 g i ) =\mathcal{A}{split}(Coh_{fb},\mathcal{C}^{i},\mathcal{T}_{g}^{i}) 𝒞 ~ i \mathcal{\tilde{C}}^{i} = UPDATE( 𝒞 ~ i \mathcal{\tilde{C}}^{i} , 𝒞 i \mathcal{C}^{i} )
Generate diversity evaluation feedback
D i v f b = ℰ d i v ( 𝒯 g i ) Div_{fb}=\mathcal{E}{div}(\mathcal{T}{g}^{i}) The ai agents story also surfaces in Google launches Gemini 3.8 Flash and..., adding another angle.
Run merge operation, and update 𝒞 i \mathcal{C}^{i} , 𝒞 ~ i \mathcal{\tilde{C}}^{i} , and 𝒯 g i \mathcal{T}_{g}^{i}
𝒞 ~ i , 𝒯 g i = 𝒜 m e r g e ( D i v f b , 𝒞 ~ i , 𝒯 g i ) \mathcal{\tilde{C}}^{i},\mathcal{T}{g}^{i}=\mathcal{A}{merge}(Div_{fb},\mathcal{\tilde{C}}^{i},\mathcal{T}_{g}^{i}) 𝒞 i \mathcal{C}^{i} = UPDATE( 𝒞 ~ i \mathcal{\tilde{C}}^{i} , 𝒞 i \mathcal{C}^{i} )
Re-evaluate refined topics
C o h f b r , D i v f b r = ℰ c o h ( 𝒞 ~ i , 𝒯 g i ) , ℰ d i v ( 𝒯 g i ) Coh_{fb}^{r},Div_{fb}^{r}=\mathcal{E}{coh}(\mathcal{\tilde{C}}^{i},\mathcal{T}{g}^{i}),\mathcal{E}{div}(\mathcal{T}{g}^{i}) o u t = 𝒜 p l a n ( C o h f b , C o h f b r , D i v f b , D i v f b r ) out=\mathcal{A}{plan}(Coh{fb},Coh_{fb}^{r},Div_{fb},Div_{fb}^{r}) if o u t . i s _ i m p r o v e d out.is_improved is t r u e true then 𝒞 = 𝒞 i \mathcal{C}=\mathcal{C}^{i} and 𝒞 ~ = 𝒞 ~ i \mathcal{\tilde{C}}=\mathcal{\tilde{C}}^{i} and 𝒯 g = 𝒯 g i \mathcal{T}{g}=\mathcal{T}{g}^{i} else t o l e r a n c e tolerance += 1 1 end if end for Topic Split Operation. For topics identified as needing separation, the split operation agent ( 𝒜 split \mathcal{A}{\text{split}} ) performs the separation. Specifically, the text segments of the topic ( C k C{k} ) are re-clustered with the number of clusters K K set to 2. Please note that C k C_{k} is used instead of C ~ k \tilde{C}{k} , because segments that are not included in the sampled segment cluster must be assigned to one of the new topics. Subsequently, we measure the proportion of text segments in the larger of the two separated clusters. If this proportion exceeds certain threshold (80% in our experiments), we skip the split operation, as this indicates that, despite 𝒜 split \mathcal{A}{\text{split}} recommending separation, the majority of text segments still pertain to a single topic. Otherwise, 𝒜 T G \mathcal{A}{TG} regenerates the topic names and descriptions for each separated segment cluster, this time incorporating the feedback provided by ℰ c o h \mathcal{E}{coh} into the prompt. Once the names and descriptions are re-generated, we update 𝒞 ~ \mathcal{\tilde{C}} accordingly by removing the split topic and adding the new topics resulting from the separation. Topic Diversity Evaluation. After the split operation, topic diversity is assessed by topic diversity evaluation agent ( ℰ d i v \mathcal{E}{div} ), which evaluates how distinctive each topic is from the others. This is a holistic, one-shot evaluation that takes all topic names and descriptions ( 𝒯 g \mathcal{T}{g} ) as input and identifies which topic groups should be merged, returning the corresponding topic indices along with feedback on how merging can improve topic diversity. Detailed information about the prompt used for ℰ d i v \mathcal{E}{div} and its output format is provided in Appendix B.3 . Topic Merging Operation. Given that ℰ d i v \mathcal{E}{div} identifies the need for topic merging and specifies the groups of topics to be united, the merge operation agent ( 𝒜 m e r g e \mathcal{A}{merge} ) combines the sampled text segments ( C ~ k \tilde{C}{k} ) of each identified topic group and run 𝒜 T G \mathcal{A}{TG} to generate the name and description of the merged topic. As with the topic split operation, feedback from ℰ d i v \mathcal{E}{div} is incorporated into the topic generation prompt. Once the merging operation is completed for all identified topic groups, 𝒞 \mathcal{C} is updated accordingly by replacing the merged topics with the newly formed topics. Planning Agent. The refined topics are re-evaluated by ℰ c o h \mathcal{E}{coh} and ℰ d i v \mathcal{E}{div} , and the results are then passed to the planning agent ( 𝒜 p l a n \mathcal{A}{plan} ), which determines whether improvements have been achieved. Specifically, the planning agent first invokes LLM to generate the summary reports for both the pre- and post-refinement iteration. It then compares these reports to determine whether any improvement has been achieved. If an improvement is observed, the topic modeling results ( 𝒞 \mathcal{C} , 𝒞 ~ \mathcal{\tilde{C}} , and 𝒯 g \mathcal{T}{g} ) are updated with the refined ones and the iteration continues. Otherwise, no updates are made and only the t o l e r a n c e tolerance parameter is incremented by 1. The refinement loop proceeds until either the maximum tolerance ( T T ) or the maximum number of iterations ( e e ) is reached. Details regarding the prompts used for 𝒜 p l a n \mathcal{A}{plan} and its output format is provided in Appendix B.4 . 2.3 Multi-view Topic Parenting In addition to topics, SeLATM also provides parent topics because presenting higher-level topics along with their relationship with base topics can offer valuable insights from an analyst’s perspective. Unlike the topic generation stage, where only text segments are available, parent topics are established based on the topics themselves, allowing additional access to the topic names and descriptions. In this regard, we employ multi-view clustering method ( Pattnaik et al., 2024 ) . Specifically, the multi-view embedding of each topic ( e k e{k} ) is constructed by concatenating the centroid of the text segment embeddings ( e c k e_{c_{k}} ), the topic name embedding ( e t k e_{t_{k}} ), and the topic description embedding ( e d k e_{d_{k}} ). Additionally, the natural language representation of each topic is formed by concatenating the topic name and description. With these representations, we repeat agglomerative clustering to categorize similar topics and use 𝒜 T G \mathcal{A}_{TG} to generate the name and description of parent topics. 3 Experiment Design Table 1 : Basic statistics of datasets used for experiments. Data set Data size Token Len The ai agents story also surfaces in Arm unveils AI-native mobile platform for..., adding another angle.
of Labels
Banking77 1,947 13 77 Bills 1,000 233 101 Wiki 1,100 3950 207 CCC 603 116 - BCR 1,250 15 - CI 3,000 17 - 3.1 Datasets We selected three publicly available datasets, Bills ( Hoyle et al., 2022 ) and Wiki ( Gao et al., 2023 ) , which are widely used for topic modeling evaluation, and Banking77 ( Casanueva et al., 2020 ) for its specificity to the finance domain. As noted by Pham et al. 2024 , using the entire training corpus for topic modeling is impractical. Therefore, we applied the same sampling strategy as Pham et al., 2024 for our experiments. We also conduct experiments on three in-house business datasets. The C ustomer C hatbot C onversation (CCC) dataset consists of conversation messages from a consumer banking chatbot application. The B anking C hatbot R eview (BCR) dataset contains customer reviews of the baking chatbot application. The C ustomer I ssues (CI) dataset comprises short summaries describing issues raised by customers during their daily banking activities. All three industry datasets does not contain ground-truth labels. Table 1 shows basic statistics of the datasets. 3.2 Evaluation Metrics The public datasets include ground-truth labels, allowing the use of standard clustering-based evaluation metrics as employed in previous studies ( Hoyle et al., 2022 ; Pham et al., 2024 ) . Accordingly, we calculated the Harmonic Mean of Purity (P1), Adjusted Rand Index (ARI), and Normalized Mutual Information (NMI) metrics, with detailed explanations provided in Appendix A.1 . These metrics require a single prediction to be compared with the labels. Therefore, for SeLATM , we used the topic with the highest contribution as the prediction for evaluation. If there was a tie in top contributions, we treated it as a new topic. Clustering-based metrics are limited in that they cannot be applied to datasets without ground-truth labels and cannot assess the quality of assigned topic names. Therefore, we evaluated topic accuracy ( 𝒯 𝒜 \mathcal{TA} ) and topic completeness ( 𝒯 𝒞 \mathcal{TC} ) ( Moon et al., 2026 ) , which measure how accurately topic names are assigned to documents and whether any topics are missing from the documents, respectively. We employed LLM -as-a-Judge building on its recent success on automated evaluations ( Zheng et al., 2023 ; Li et al., 2024 ; Jang and Silavong, 2025 ) . As LLM are prone to producing inconsistent predictions ( Jang and Lukasiewicz, 2023 ; Wang et al., 2024 ) , we used self-consistency decoding for more reliable evaluations ( Wang et al., 2023 ) . Specifically, for each 𝒯 𝒜 \mathcal{TA} and 𝒯 𝒞 \mathcal{TC} , we prompted LLM to generate five predictions—each with a distinct reasoning path—given a document and its assigned topic name, following the label schema presented in Table 2 . The predictions are converted to their corresponding scores, and the final score is computed as the average of these five scores. The prompts of 𝒯 𝒜 \mathcal{TA} and 𝒯 𝒞 \mathcal{TC} evaluations are presented in Appendix B.5 . When the output consists of multiple topics, we concatenated the generated topics into a single text for measuring 𝒯 𝒞 \mathcal{TC} , considering the definition of the metric. However, this strategy can provide additional benefits when measuring 𝒯 𝒜 \mathcal{TA} . Since SeLATM provides a distribution of topics, we used the topic ratios as weights and calculated the weighted average of each individual topic’s 𝒯 𝒜 \mathcal{TA} score. For methods that do not provide topic weights, we applied the same approach as used for 𝒯 𝒞 \mathcal{TC} . Table 2 : Textual labels and corresponding scores used for LLM-as-a-Judge evaluation. 𝒯 𝒜 \mathcal{TA} 𝒯 𝒞 \mathcal{TC} Score Incorrect Not Covered 0 Partially Correct Minorly Covered 1/3 Mostly Correct Mostly Covered 2/3 Completely Correct Complete 1 3.3 Baseline Methods As baselines, we selected the following approaches from three distinct topic modeling families, based on their reported performance and code availability.
• Embedded topic models : BERTopic ( Grootendorst, 2022 ) , CTop2Vec ( Angelov and Inkpen, 2024 ) .
• LLM -driven topic generation and assignment : TopicGPT ( Pham et al., 2024 ) , LLoom ( Lam et al., 2024 ) , TIDE ( Moon et al., 2026 ) .
• LLM -assisted neural topic models : LLM-ITL ( Yang et al., 2025 ) . Further details about the baselines and their implementations are provided in Appendix A.2 . 3.4 Models and Hyperparameters For experiments, we used OpenAI models; gpt-4o-2024-08-06 for language generation and text-embedding-3-small-1 for generating embedding vectors. The same LLM were also applied to baselines unless their code implementations require specific models. Regarding LLM -as-a-Judge, we used a different model, gpt-4.1-mini-2025-04-14 . The number of topics ( K K ) is a necessary hyperparameter for SeLATM , TIDE, and LLM-ITL. Based on the number of ground-truth labels, we set K K to 100 for the Banking77 and Bills datasets, and to 250 for the Wiki dataset. For the business datasets, which do not contain ground-truth labels, we set K K to 50. SeLATM requires an additional hyperparameter, n S n_{S} , which specifies the number of sampled text segments for each cluster. For each dataset, we identified the best n S n_{S} through hyperparameter search; detailed experimental results are provided in Appendix A.3 . Table 3 : Experimental results on the public datasets. We report an average of 5 repetitions. The best performance for each evaluation metric is highlighted in bold; underlined values indicate statistically significant performance gap over the best-performing baselines (or over SeLATM when a baseline performs best), determined by t-test with a p-value < < 0.05. ‘ SeLATM Iter:0’ refers to SeLATM results without the refinement loop; italicized values indicate where ‘ SeLATM Iter:0’ outperforms the baselines. Models Banking77 Bills Wiki P1 ARI NMI 𝒯 𝒜 \mathcal{TA} 𝒯 𝒞 \mathcal{TC} P1 ARI NMI 𝒯 𝒜 \mathcal{TA} 𝒯 𝒞 \mathcal{TC} P1 ARI NMI 𝒯 𝒜 \mathcal{TA} 𝒯 𝒞 \mathcal{TC} BERTopic .705 .468 .817 .634 .667 .359 .230 .558 .099 .066 .303 .144 .652 .137 .076 C-Top2vec .557 .406 .788 .470 .704 .424 .230 .767 .350 .379 .490 .341 .685 .501 .574 TopicGPT .172 .078 .532 .734 .683 .385 .244 .709 .748 .740 .392 .183 .779 .826 .627 LLooM .111 .037 .322 .283 .259 .159 .048 .295 .501 .424 .110 .046 .424 .544 .353 TIDE .624 .486 .800 .782 .867 .440 .285 .711 .759 .719 .389 .259 .816 .641 .621 LLM-ITL .244 .001 .562 .379 .419 .395 .033 .723 .129 .110 .491 .046 .802 .168 .010 SeLATM Iter:0 .684 .542 .832 .955 .953 .440 .255 .663 .795 .769 .522 .357 .805 .832 .574 SeLATM Iter:1 .680 .539 .832 .962 .956 .440 .260 .661 .803 .782 .524 .359 .803 .807 .567 SeLATM Iter:3 .672 .533 .830 .966 .960 .443 .258 .660 .823 .792 .522 .355 .803 .842 .587 SeLATM Iter:5 .663 .526 .826 .962 as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!