During a network outage, two modules may each produce a sequence of alarms. Knowing that the sequences occurred together does not explain which underlying failure caused the other. Inferring Causal Relations between Two Sequences of Events with Language Models studies that problem when only one observed trace from each process is available.
The researchers propose comparing language-model probability estimates for two possible directions. They do not ask a chatbot to explain what named alarms mean. Instead, they encode events as symbols and use a pretrained model to score their sequence structure.
Orienting a relationship that is assumed to exist
The formal task assumes a direct causal relationship between the two hidden processes. It then asks whether the first process causes the second or the reverse. This is narrower than discovering whether two arbitrary traces are causally related at all.
The method estimates a score for each direction using negative log-likelihoods. One part scores the presumed cause sequence; another scores the presumed effect conditioned on that cause. The proposed Sequence Generation Rule chooses the direction with the lower estimated score.
The interpretation relies on the authors' model of asymmetric sequence generation. A score comparison is not an intervention on the real system, and it does not remove the need to check whether that model fits a particular application.
Reducing the influence of event names
Ordinary language-model probabilities depend on token meanings and frequencies. That can distort a task whose symbols stand for arbitrary events, rather than familiar words.
The researchers therefore map distinct event symbols to randomly selected tokens through one-to-one transformations. These encodings preserve the original pattern of repeated events while changing its labels. Scores are estimated across multiple mappings instead of depending on a single choice of vocabulary.
Exhaustively trying every possible mapping would be impractical. The procedure limits the vocabulary size and number of mappings it considers. It also introduces sequence replication to give short traces more context for pattern recognition. Replication changes the model's input presentation; it does not supply additional independent observations of the underlying system.
What the reported evaluation establishes
The authors report testing the approach on synthetic data and real IT-monitoring data, and describe better results than conventional causal-discovery baselines on several converted time-series datasets. Those are research findings reported by the paper, not independent validation of an operational incident-response system.
The method addresses a difficult data constraint: conventional approaches often require repeated observations to estimate distributions or event intensities. Its proposed alternative draws on a pretrained model's sequence-scoring ability without fine-tuning on a collection of the target traces.
For a practitioner, the direct-causation assumption and the meaning of the probability score are central. The work offers a way to investigate direction under that setup; it should not be read as proof that a language model can infer any real-world cause from two short logs.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!