Back to AI Research

AI Research

UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Pric... | AI Research

Key Takeaways

  • What the paper is about Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB f...
  • The UQ-regression variant outputs a calibrated Gaussian over the future tick displacement, while the UQ-classification variant outputs a categorical distribution over down/up/stationary.
  • Both expose a scalar confidence (predicted signal-to-noise ratio or class probability) that supports selective prediction.
  • On large, economically meaningful moves, the tightest confidence tier reaches a directional F1 of 0.88 (down) and 0.83 (up) at the 5-second horizon.
  • The same ai evaluation question is explored in [GRUET](/learn/research/1rVoJMJR2aBZYjFJeaSM), which adds a research perspective.
Paper AbstractExpand

Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recently completed windows whose outcomes are already realised. The UQ-regression variant outputs a calibrated Gaussian over the future tick displacement, while the UQ-classification variant outputs a categorical distribution over down/up/stationary. Both expose a scalar confidence (predicted signal-to-noise ratio or class probability) that supports selective prediction. On 5.2 billion LOB events across seven cryptocurrency assets and horizons of 5, 10 and 15 seconds, UQ-regression attains near-nominal 68% interval coverage, and restricting to the most confident 10% of predictions raises directional macro F1 by 0.11-0.15 for UQ-regression and 0.05-0.11 for UQ-classification, at every horizon. On large, economically meaningful moves, the tightest confidence tier reaches a directional F1 of 0.88 (down) and 0.83 (up) at the 5-second horizon.

What the paper is about

Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recently completed windows whose outcomes are already realised. The UQ-regression variant outputs a calibrated Gaussian over the future tick displacement, while the UQ-classification variant outputs a categorical distribution over down/up/stationary. Both expose a scalar confidence (predicted signal-to-noise ratio or class probability) that supports selective prediction. On 5.2 billion LOB events across seven cryptocurrency assets and horizons of 5, 10 and 15 seconds, UQ-regression attains near-nominal 68% interval coverage, and restricting to the most confident 10% of predictions raises directional macro F1 by 0.11-0.15 for UQ-regression and 0.05-0.11 for UQ-classification, at every horizon. On large, economically meaningful moves, the tightest confidence tier reaches a directional F1 of 0.88 (down) and 0.83 (up) at the 5-second horizon. The same ai evaluation question is explored in GRUET, which adds a research perspective.

What it covers

UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting Thanks: Corresponding author. Derrick Gilchrist Edward Manoharan

  • Affiliation: Data Science Research Centre Affiliation: Tampere University Affiliation: Tampere, Finland Affiliation: [email protected] Eljas Linna Affiliation: Data Science Research Centre Affiliation: Tampere University Affiliation: Tampere, Finland Email: [email protected] Kestutis Baltakys Affiliation: Data Science Research Centre Affiliation: Tampere University Affiliation: Tampere, Finland Email: [email protected] Hao Dong Affiliation: Data Science Research Centre Affiliation: Tampere University Affiliation: Tampere, Finland Email: [email protected] Juho Kanniainen Affiliation: Data Science Research Centre Affiliation: Tampere University Affiliation: Tampere, Finland Email: [email protected] Abstract Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB , a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recently completed windows whose outcomes are already realised. The UQ-regression variant outputs a calibrated Gaussian over the future tick displacement, while the UQ-classification variant outputs a categorical distribution over down/up/stationary. Both expose a scalar confidence (predicted signal-to-noise ratio or class probability) that supports selective prediction. On 5.2 billion LOB events across seven cryptocurrency assets and horizons of 5, 10 and 15 seconds, UQ-regression attains near-nominal 68% interval coverage, and restricting to the most confident 10% of predictions raises directional macro F1 by 0.11–0.15 for UQ-regression and 0.05–0.11 for UQ-classification, at every horizon. On large, economically meaningful moves, the tightest confidence tier reaches a directional F1 of 0.88 (down) and 0.83 (up) at the 5-second horizon. 1 Introduction The limit order book (LOB) is the mechanism through which price discovery occurs in most electronic markets ( Gould et al., 2013 ) . It aggregates resting buy (bid) and sell (ask) orders by price level; the best bid and best ask define the spread, and their midpoint, the mid-price, is the standard instantaneous valuation of the asset ( Glosten and Milgrom, 1985 ; Kolm et al., 2023 ) . An incoming order that crosses the spread executes immediately against resting liquidity, whereas one that does not rests in the book at its own price level ( Briola et al., 2025 ) . Because thousands of such events arrive every second, the LOB carries microstructural signal about the next few seconds of price formation that predictive models seek to exploit ( Sirignano and Cont, 2019 ; Zhang et al., 2019a ; Lucchese et al., 2024 ) . Deep LOB forecasters are overwhelmingly point predictors : they emit a direction or a displacement, but not how much that output should be trusted. For a trading decision this omission is costly. Predictability varies sharply across market regimes and even across consecutive windows, and every action incurs fees and slippage, so the practically relevant question is rarely β€œwhich way will the price move?” but β€œis this forecast reliable enough to act on?”, the setting of selective prediction ( Chow, 1970 ; El-Yaniv and Wiener, 2010 ; Geifman and El-Yaniv, 2017 ) . Answering it requires input-dependent (aleatoric) uncertainty that is calibrated : an interval claimed to cover 68% of outcomes should do so empirically. Figure 1 illustrates the output we target. Instead of a single value, the model returns a distribution 𝒩 ⁑ ( ΞΌ , Οƒ 2 ) \mathcal{N}(\mu,\sigma^{2}) over the mid-price displacement at t + h t+h , and the ratio | ΞΌ | / Οƒ |\mu|/\sigma provides a scalar confidence that can be thresholded. We obtain this with UQ-LOB , a module attached to the hidden representation of a pretrained LOB encoder. Inspired by conditional and attentive neural processes ( Garnelo et al., 2018a ; Kim et al., 2019 ) , UQ-LOB conditions each prediction on a context set of recently completed windows whose realised outcomes are already known at time t t , so that the forecast can adapt to the prevailing regime without any gradient update. Figure 1: Confidence-aware mid-price forecasting. The model sees only the input window ending at t t (shaded) and returns a Gaussian over the displacement at t + h t+h ; its 68% and 95% intervals give a thresholdable measure of how much the forecast should be trusted. To the best of our knowledge, no prior work equips mid-price forecasting with calibrated, input-dependent uncertainty of this kind. We instantiate UQ-LOB as a UQ-regression variant, which outputs a Gaussian over the tick displacement, and a UQ-classification variant, which outputs a categorical distribution over the three directional classes (down, up, stationary), and we compare the two under matched confidence gating. Our contributions are:

β€’ An encoder-agnostic, in-context UQ head. An attentive head that conditions on the realised outcomes of recent windows and produces a calibrated Gaussian 𝒩 ⁑ ( ΞΌ , Οƒ 2 ) \mathcal{N}(\mu,\sigma^{2}) over tick displacement, with confidence given by its signal-to-noise ratio | ΞΌ | / Οƒ |\mu|/\sigma . Its objective combines a heteroscedastic likelihood with magnitude weighting, a directional margin, and a differentiable magnitude-conditional calibration penalty that turns the post-hoc calibration metrics of Levi et al. (2022) into a training signal.

β€’ A classification counterpart on the same trunk , trained on the three-class label with confidence given by its maximum class probability, enabling a controlled comparison between continuous and discrete targets at identical parameter count.

β€’ An evaluation protocol for selective LOB prediction based on directional macro F1 , which removes the trivial effect of the stationary class under gating, a magnitude-restricted analysis on large moves, and prior-matched random baselines. On 5.2 billion events across seven assets and three horizons, confidence gating improves directional reliability at every horizon, and the regression head matches the dedicated classifier at the top confidence decile while additionally providing a magnitude and a calibrated interval. 2 Related Work LOB forecasting. A large share of deep learning work on LOB data targets next-message or next-event prediction ( Nagy et al., 2023 ) , an autoregressive formulation in which errors compound as the effective horizon grows. The FI-2010 benchmark ( Ntakaris et al., 2018 ) standardised a second line of work that predicts the direction of the mid-price over a horizon of k k future events as a three-class label ( Zhang et al., 2019a ; Xiao et al., 2025 ) . Architectures for this task have progressed from CNNs ( Tsantekidis et al., 2017 ) and LSTMs ( Tsantekidis et al., 2020 ) to bilinear and attention-based models ( Tran et al., 2018 ; Wallbridge, 2020 ) , pretrained Transformers ( Xiao et al., 2025 ) , and BERT-style encoders trained on Level-2/Level-3 event streams ( Linna et al., 2025 ) . All of these, including work that treats the displacement as a regression target ( Yang et al., 2025 ) , are point predictors with no accompanying measure of predictive confidence. Uncertainty in LOB forecasting. Quantile regression ( Zhang et al., 2019b ) jointly models several return quantiles and avoids quantile crossing, but targets ask- and bid-side quantiles separately and fixes the quantile levels in advance rather than yielding a single predictive distribution. BDLOB ( Zhang et al., 2018 ) applies dropout variational inference ( Gal and Ghahramani, 2016 ) to a LOB CNN and shows that the resulting posterior predictive uncertainty can be used for position sizing and for avoiding unnecessary trades; Magris et al. (2023) obtain predictive distributions over class probabilities from a Bayesian bilinear network. Both derive uncertainty from a distribution over network weights (epistemic uncertainty) and require multiple stochastic passes at inference, whereas our head predicts input-dependent (aleatoric) uncertainty in a single forward pass and conditions it on recently realised outcomes. Calibration, selective prediction, and neural processes. Heteroscedastic Gaussian regression ( Nix and Weigend, 1994 ; Kendall and Gal, 2017 ) provides aleatoric uncertainty but is prone to loss attenuation ( Seitzer et al., 2022 ; Stirn et al., 2023 ) ; post-hoc recalibration ( Guo et al., 2017 ; Kuleshov et al., 2018 ) , deep ensembles ( Lakshminarayanan et al., 2017 ) , and conformal methods ( Angelopoulos and Bates, 2023 ) improve calibration but do not by themselves make it input-dependent, and predictive uncertainty is known to degrade under distribution shift ( Ovadia et al., 2019 ) , which is endemic to financial data. Selective prediction with a reject option ( Chow, 1970 ; El-Yaniv and Wiener, 2010 ; Geifman and El-Yaniv, 2017 ) and the maximum softmax probability as a confidence score ( Hendrycks and Gimpel, 2016 ) form the evaluation lens we adopt. Finally, neural processes ( Garnelo et al., 2018b ) and their conditional and attentive variants ( Garnelo et al., 2018a ; Kim et al., 2019 ) learn predictive distributions conditioned on an observed context set; we adapt this in-context formulation to the LOB setting, where the context is the set of most recently completed windows whose labels are already realised. 3 UQ-LOB: The Uncertainty Quantification Head Point predictions of mid-price displacement carry no indication of their own reliability, which is particularly problematic in LOB markets, where volatility, liquidity and directional momentum shift rapidly across regimes. UQ-LOB addresses this by producing, for a target window with hidden representation 𝐑 t \mathbf{h}^{t} , a full predictive distribution conditioned on a context set π’ž \mathcal{C} of (representation, realised label) pairs from recently completed windows (Figure 2 ). The head operates on the encoder’s hidden representation rather than on raw market features, so any pretrained LOB encoder can serve as its starting point (Section 4.3 ). Figure 2: The UQ-LOB. Context and target representations from a pretrained encoder are projected into a shared space; each context representation is fused with its (scaled) realised label, refined by self-attention, and read out by the target through cross-attention. The decoder outputs either a Gaussian ( ΞΌ , Οƒ 2 ) (\mu,\sigma^{2}) over tick displacement (UQ-regression) or three class logits (UQ-classification). 3.1 Problem Setup and Labels The event stream of each asset is segmented into non-overlapping windows of L = 512 L=512 events (Section 4.2 ). Let t t denote the timestamp of the last event of a window, i.e. the model’s present moment. To reflect execution latency, the start price p start p_{\mathrm{start}} is the mid-price at the first event after a small fixed delay d d past t t , and the end price p end p_{\mathrm{end}} is the mean mid-price over the k = 10 k=10 events immediately preceding t start + h t_{\mathrm{start}}+h , which damps single-tick noise without washing out genuine displacement at the tick rates of our data. The regression target is the raw tick displacement y h ​ ( t ) = p end βˆ’ p start y_{h}(t)=p_{\mathrm{end}}-p_{\mathrm{start}} . The classification label thresholds the same displacement at Β± Ξ΄ h ​ ( t ) \pm\delta_{h}(t) , a volatility-scaled cutoff derived from the training-split log-return distribution of each asset and horizon and converted to ticks at the current price level (Appendix A.7 ). All labels are computed independently per asset and horizon. 3.2 In-Context Conditioning The head follows conditional neural processes (CNPs) ( Garnelo et al., 2018a ) , which cast regression as an in-context learning problem: rather than learning a fixed input–output map, a CNP summarises an observed context set π’ž = { ( 𝐱 i , y i ) } i = 1 C \mathcal{C}={(\mathbf{x}{i},y{i})}{i=1}^{C} into a representation 𝐫 ⁑ ( π’ž ) \mathbf{r}(\mathcal{C}) and predicts p ⁑ ( y βˆ— ∣ 𝐱 βˆ— , π’ž ) = p ⁑ ( y βˆ— ∣ 𝐱 βˆ— , 𝐫 ⁑ ( π’ž ) ) p(y^{}\mid\mathbf{x}^{},\mathcal{C})=p(y^{}\mid\mathbf{x}^{},\mathbf{r}(\mathcal{C})) . The original CNP aggregates context points by a mean, which weighs every point equally; attentive neural processes ( Kim et al., 2019 ) let the query attend selectively to the most relevant points, and we adopt this form. In our setting the inputs are windows. Let 𝐗 i ∈ ℝ L Γ— D \mathbf{X}{i}\in\mathbb{R}^{L\times D} be a raw window of L L events with D D features, encoded by a pretrained encoder into 𝐑 i = f ΞΈ ​ ( 𝐗 i ) ∈ ℝ d h \mathbf{h}{i}=f{\theta}(\mathbf{X}{i})\in\mathbb{R}^{d{h}} . A training instance consists of N = C + 1 N=C+1 windows: C C context pairs { ( 𝐑 i c , y i c ) } i = 1 C {(\mathbf{h}{i}^{c},y{i}^{c})}{i=1}^{C} and a single withheld target ( 𝐑 t , y t ) (\mathbf{h}^{t},y^{t}) . The context set is drawn from the windows that most recently precede the target, subject to a causal constraint : if a context window ends at t i c t{i}^{c} , its label is realised at e i c = t i c + d + h e_{i}^{c}=t_{i}^{c}+d+h , and we require e i c ≀ t βˆ€ i ∈ { 1 , … , C } , e_{i}^{c};\leq;t\qquad\forall,i\in{1,\dots,C}, (1) so that every context label is known at the moment the target prediction is made. This is strictly stronger than merely requiring the context and target input windows not to overlap: it also excludes context windows whose label horizon runs past t t , which would leak future price information into the prediction. The same construction is used at training, validation and test time, so the test-time predictor only ever conditions on information available at t t (Appendix A.5 ). 3.3 Head Architecture Figure 2 shows the head. Context and target representations are projected into a shared space by a common linear map W p W_{p} : 𝐑 ~ i c = W p ​ 𝐑 i c \tilde{\mathbf{h}}{i}^{c}=W{p}\mathbf{h}{i}^{c} and 𝐑 ~ t = W p ​ 𝐑 t \tilde{\mathbf{h}}^{t}=W{p}\mathbf{h}^{t} . Context labels are divided by a fixed, asset- and horizon-specific scale y ref y_{\mathrm{ref}} (Appendix B.4 ) and projected through a tanh \tanh -bounded linear layer, 𝐲 i proj = tanh ⁑ ( W y ​ y i c / y ref ) \mathbf{y}{i}^{\mathrm{proj}}=\tanh!\big(W{y},y_{i}^{c}/y_{\mathrm{ref}}\big) . Each projected representation is concatenated with its projected label and passed through the context encoder , a three-layer GELU MLP, 𝐫 i = MLP enc ( [ 𝐑 ~ i c ; 𝐲 i proj ] ) , i = 1 , … , C . \mathbf{r}{i}=\mathrm{MLP}{\mathrm{enc}}\big(\left[\tilde{\mathbf{h}}{i}^{c},;,\mathbf{y}{i}^{\mathrm{proj}}\right]\big),\qquad i=1,\dots,C. (2) Self-attention over { 𝐫 i } i = 1 C {\mathbf{r}{i}}{i=1}^{C} then lets each context point re-weigh itself by inter-context relevance, yielding a refined set { 𝐫 i ref } i = 1 C {\mathbf{r}{i}^{\mathrm{ref}}}{i=1}^{C} . The target queries this refined context by cross-attention: the projected target embedding is mapped to a query and the projected context embeddings to keys through learned maps W q , W k W_{q},W_{k} , while the refined representations serve directly as values , πͺ = W q ​ 𝐑 ~ t \mathbf{q}=W_{q}\tilde{\mathbf{h}}^{t} , 𝐀 i = W k ​ 𝐑 ~ i c \mathbf{k}{i}=W{k}\tilde{\mathbf{h}}{i}^{c} , 𝐯 i = 𝐫 i ref \mathbf{v}{i}=\mathbf{r}{i}^{\mathrm{ref}} , giving 𝐫 t = softmax ⁑ ( πͺ ​ 𝐊 ⊀ d r ) ​ 𝐕 , \mathbf{r}^{t}=\mathrm{softmax}!\left(\frac{\mathbf{q},\mathbf{K}^{\top}}{\sqrt{d{r}}}\right)\mathbf{V}, (3) where 𝐊 , 𝐕 \mathbf{K},\mathbf{V} stack the keys and values row-wise. Separating keys from values lets the target decide which context windows are relevant from their representation alone, while retrieving the label-informed summary of those windows: recent, similarly structured LOB states receive higher weight than dissimilar ones. The attended summary is concatenated with the target’s own projected embedding and passed to a decoder , a five-layer GELU MLP, which thus sees both the inferred market context and the target’s own features. The UQ-regression decoder outputs [ ΞΌ , Ξ½ ] = MLP dec ​ ( [ 𝐫 t ; 𝐑 ~ t ] ) , Οƒ 2 = y ref 2 β‹… softplus ⁑ ( Ξ½ ) + Ο΅ , [\mu,,\nu]=\mathrm{MLP}{\mathrm{dec}}\big(\left[\mathbf{r}^{t},;,\tilde{\mathbf{h}}^{t}\right]\big),\qquad\sigma^{2}=y{\mathrm{ref}}^{2}\cdot\mathrm{softplus}(\nu)+\epsilon, (4) giving p ⁑ ( y t ∣ 𝐑 t , π’ž ) = 𝒩 ⁑ ( ΞΌ , Οƒ 2 ) p(y^{t}\mid\mathbf{h}^{t},\mathcal{C})=\mathcal{N}(\mu,\sigma^{2}) with Ο΅ \epsilon a small constant for numerical stability. The UQ-classification decoder outputs three logits [ β„“ down , β„“ up , β„“ stat ] = MLP dec cls ​ ( [ 𝐫 t ; 𝐑 ~ t ] ) [\ell_{\mathrm{down}},\ell_{\mathrm{up}},\ell_{\mathrm{stat}}]=\mathrm{MLP}{\mathrm{dec}}^{\mathrm{cls}}([\mathbf{r}^{t};\tilde{\mathbf{h}}^{t}]) . The two variants share the encoder, projection and attention trunk and differ only in the decoder (Appendix A.2 ). 3.4 Training Objectives UQ-regression. The regression head minimises a weighted sum of four terms, each computed on target windows only: β„’ reg = Ξ» 1 ​ β„’ NLL + Ξ» 2 ​ β„’ WMAE + Ξ» 3 ​ β„’ dir + Ξ» 4 ​ β„’ calib . \mathcal{L}{\mathrm{reg}}=\lambda_{1}\mathcal{L}{\mathrm{NLL}}+\lambda{2}\mathcal{L}{\mathrm{WMAE}}+\lambda{3}\mathcal{L}{\mathrm{dir}}+\lambda{4}\mathcal{L}{\mathrm{calib}}. (5) (i) Heteroscedastic NLL ( Nix and Weigend, 1994 ; Kendall and Gal, 2017 ) , β„’ NLL = 𝔼 ⁑ [ 1 2 ​ log ⁑ ( 2 ​ Ο€ ​ Οƒ 2 ) + ( y βˆ’ ΞΌ ) 2 / ( 2 ​ Οƒ 2 ) ] \mathcal{L}{\mathrm{NLL}}=\mathbb{E}\big[\tfrac{1}{2}\log(2\pi\sigma^{2})+(y-\mu)^{2}/(2\sigma^{2})\big] . Letting Οƒ \sigma depend on the input allows the network to temper the residual of hard examples by inflating Οƒ \sigma ( loss attenuation ). Seitzer et al. (2022) show this can happen before ΞΌ \mu has converged: since βˆ‚ β„’ NLL / βˆ‚ ΞΌ ∝ 1 / Οƒ 2 \partial\mathcal{L}{\mathrm{NLL}}/\partial\mu\propto 1/\sigma^{2} , an inflated Οƒ \sigma starves the gradient reaching ΞΌ \mu and biases ΞΌ \mu toward small magnitudes. (ii) Magnitude-weighted MAE , which counteracts this bias by up-weighting large displacements, β„’ WMAE = 𝔼 ⁑ [ w ⁑ ( y ) ​ | ΞΌ βˆ’ y | y ref ] , w ⁑ ( y ) = min ⁑ ( 1 + ( | y | y ref ) p , w max ) , \mathcal{L}{\mathrm{WMAE}}=\mathbb{E}!\left[w(y),\frac{|\mu-y|}{y_{\mathrm{ref}}}\right],\qquad w(y)=\min!\left(1+\left(\frac{|y|}{y_{\mathrm{ref}}}\right)^{p},\ w_{\max}\right), (6) with the clamp w max w_{\max} bounding the influence of extreme outliers ( p = 4 p=4 , w max = 40 w_{\max}=40 throughout). (iii) Directional loss. Neither of the above targets the discrete up/down/stationary boundary Ξ΄ i \delta_{i} used at evaluation time. Writing ΞΌ ~ i = ΞΌ i / y ref \tilde{\mu}{i}=\mu{i}/y_{\mathrm{ref}} , we add β„’ dir = 𝔼 [ [ | y i | β‰₯ Ξ΄ i ] ( ΞΌ ~ i βˆ’ sign ( y i ) m ) 2 + [ | y i | < Ξ΄ i ] | ΞΌ ~ i βˆ’ sign ( y i ) m | ] , \mathcal{L}{\mathrm{dir}}=\mathbb{E}!\left[\mathbbm{1}!\left[|y{i}|\geq\delta_{i}\right]\big(\tilde{\mu}{i}-\mathrm{sign}(y{i}),m\big)^{2};+;\mathbbm{1}!\left[|y_{i}|<\delta_{i}\right]\big|\tilde{\mu}{i}-\mathrm{sign}(y{i}),m\big|\right], (7) where Ξ΄ i \delta_{i} is the asset-, horizon- and window-specific threshold of Appendix A.7.3 and m = 1 m=1 is a fixed margin. Directional examples receive a quadratic pull toward a confidently signed target Β± m \pm m ; stationary examples only a linear one, so their (weaker, bounded-gradient) sign signal is retained without a strong pull in magnitude. This class-conditional quadratic/linear switching is structurally related to the reverse-Huber (berHu) loss ( Zwald and Lambert-Lacroix, 2012 ; Laina et al., 2016 ) , which switches on residual magnitude rather than on the true class. (iv) Magnitude-conditional calibration loss. The three terms above shape ΞΌ \mu ; none directly corrects Οƒ \sigma . We therefore add β„’ calib = 1 B ​ βˆ‘ b = 1 B ( 𝔼 i ∈ ℬ b ​ [ | ΞΌ i βˆ’ y i | Οƒ i ] βˆ’ 2 Ο€ ) 2 , \mathcal{L}{\mathrm{calib}}=\frac{1}{B}\sum{b=1}^{B}\left(\mathbb{E}{i\in\mathcal{B}{b}}!\left[\frac{|\mu_{i}-y_{i}|}{\sigma_{i}}\right]-\sqrt{\tfrac{2}{\pi}}\right)^{2}, (8) where { ℬ b } b = 1 B {\mathcal{B}{b}}{b=1}^{B} partitions the batch into B B quantile bins of | ΞΌ | |\mu| and 2 / Ο€ = 𝔼 ​ | Z | \sqrt{2/\pi}=\mathbb{E}|Z| , Z ∼ 𝒩 ⁑ ( 0 , 1 ) Z\sim\mathcal{N}(0,1) , is the expected standardised residual of a calibrated Gaussian. Unlike remedies that reweight the NLL globally ( Seitzer et al., 2022 ) or decouple the mean and variance branches architecturally ( Stirn et al., 2023 ) , β„’ calib \mathcal{L}{\mathrm{calib}} turns the post-hoc calibration metrics of Levi et al. (2022) into a differentiable penalty, binning on predicted magnitude so that calibration is enforced separately for small and large predicted moves. All loss hyperparameters are listed in Appendix A.4 . UQ-classification. The classification head minimises a class-weighted cross-entropy over the three-class label, β„’ cls = βˆ’ βˆ‘ c = 0 2 w c 1 [ y cls = c ] log p c \mathcal{L}{\mathrm{cls}}=-\sum_{c=0}^{2}w_{c},\mathbbm{1}[y_{\mathrm{cls}}=c]\log p_{c} , with w c ∝ 1 / f c w_{c}\propto 1/f_{c} the inverse training-split frequency of class c c , normalised so that βˆ‘ c w c = 3 \sum_{c}w_{c}=3 , which counteracts the class imbalance documented in Appendix B . 3.5 Confidence Scores and Selective-Prediction Protocol Confidence. For UQ-regression we use the predicted signal-to-noise ratio SNR = | ΞΌ | / Οƒ \mathrm{SNR}=|\mu|/\sigma : a large predicted displacement paired with a small predicted uncertainty is both worth acting on and sharply estimated. For UQ-classification we use the softmax confidence max c ⁑ p c \max_{c}p_{c} ( Hendrycks and Gimpel, 2016 ) . Gating at percentile q q retains the ( 100 βˆ’ q ) % (100-q)% most confident predictions of each asset–horizon test set and evaluates only on those; q = 0 q=0 is the full test set. Directional macro F1. To compare the heads on a common footing, the regression mean ΞΌ \mu is mapped to a three-class label with the same rule as the ground truth, using a threshold multiplier calibrated on a held-out split so that the predicted stationary proportion matches the true one (Appendix A.7.3 ). Under gating, the stationary class behaves degenerately: confident predictions are large in magnitude and therefore predominantly directional, so stationary F1 collapses while down/up F1 rise (Table 8 ), and three-class macro F1 conflates the two effects. We therefore report directional macro F1 , the mean of the down and up F1 scores, and provide a prior-matched reference: a classifier predicting at random with the true class priors attains an expected F1 of p c p_{c} for class c c , so its directional macro F1 is ( 1 βˆ’ p stat ) / 2 (1-p_{\mathrm{stat}})/2 . Because gating changes the class composition of the retained subset, this reference is only valid for the full test set and rises under gating; the large-move analysis of Section 5.3 fixes it at 0.5 by construction. Calibration and magnitude-weighted fit. We report the empirical coverage of the nominal 68% and 95% intervals, cov 68 = Pr [ | y βˆ’ ΞΌ | ≀ Οƒ ] \mathrm{cov}{68}=\Pr[|y-\mu|\leq\sigma] and cov 95 = Pr [ | y βˆ’ ΞΌ | ≀ 1.96 Οƒ ] \mathrm{cov}{95}=\Pr[|y-\mu|\leq 1.96\sigma] , and the negative log predictive density (NLPD). Point fit is summarised by a magnitude-weighted R 2 R^{2} , R w 2 = 1 βˆ’ βˆ‘ i w i ​ ( y i βˆ’ ΞΌ i ) 2 βˆ‘ i w i ​ ( y i βˆ’ y Β― w ) 2 , y Β― w = βˆ‘ i w i ​ y i βˆ‘ i w i , R^{2}{w}=1-\frac{\sum{i}w_{i}(y_{i}-\mu_{i})^{2}}{\sum_{i}w_{i}(y_{i}-\bar{y}{w})^{2}},\qquad\bar{y}{w}=\frac{\sum_{i}w_{i}y_{i}}{\sum_{i}w_{i}}, (9) with w i = w ⁑ ( y i ) w_{i}=w(y_{i}) from Equation 6 . Large moves are those that cross the stationary threshold and, after fees, are the only ones a trader can profit from; WR 2 therefore measures fit on the moves that matter rather than on the small fluctuations that dominate standard R 2 R^{2} . 4 Experimental Setup 4.1 Dataset We use Level-2 and Level-3 order book data collected from the Kraken cryptocurrency exchange through its public market data API for seven USD-quoted assets: Bitcoin (BTC), Ethereum (ETH), Solana (SOL), Litecoin (LTC), Dogecoin (DOGE), Sui (SUI) and Bittensor (TAO). The data span November 2025 to February 2026 and contain approximately 5.2 billion events over 600 asset-days, with event rates from 39 to 156 events/s and mean inter-event times from 6.4 to 26.0 ms (Table 6 , Appendix B ). For every asset, data up to 18 January 2026 are used for training, 19–31 January for validation and 1–15 February for testing (about 68%, 15% and 17% of asset-days). The validation period is split into two equal halves for model selection and for calibrating the threshold multiplier of Appendix A.7.3 ; the test period is strictly out-of-sample and chronologically after both. Because no data-sharing agreement is in place with the exchange, the raw data cannot be redistributed; we release code, all label thresholds and scale values, and the dataset statistics needed to reconstruct the pipeline on comparable data (Reproducibility Statement). 4.2 Input Representation Following Linna et al. (2025) , each event is encoded as a discrete token from a vocabulary of | 𝒱 | = 439 |\mathcal{V}|=439 entries built from five categorical attributes (event type, side, discretised volume, discretised price distance from the opposing best quote, and a simultaneity flag); events deeper than L max = 10 L_{\max}=10 levels are discarded. Each event also carries seven continuous features. Three follow Linna et al. (2025) : log inter-event time, log tick distance from the opposing best quote and log relative volume. Four capture directional order-flow pressure at the top of the book. Building on the order flow imbalance of Cont et al. (2014) , the depth-normalised flow imbalance DNFI t = Ξ” ​ Q 1 , t b βˆ’ Ξ” ​ Q 1 , t a Q 1 , t b + Q 1 , t a ∈ [ βˆ’ 1 , + 1 ] , QI t = Q 1 , t b βˆ’ Q 1 , t a Q 1 , t b + Q 1 , t a ∈ [ βˆ’ 1 , + 1 ] , \mathrm{DNFI}{t}=\frac{\Delta Q^{b}{1,t}-\Delta Q^{a}{1,t}}{Q^{b}{1,t}+Q^{a}{1,t}}\in[-1,+1],\qquad\mathrm{QI}{t}=\frac{Q^{b}{1,t}-Q^{a}{1,t}}{Q^{b}{1,t}+Q^{a}{1,t}}\in[-1,+1], (10) tracks the net change in best-level queue sizes Q 1 , t b , Q 1 , t a Q^{b}{1,t},Q^{a}{1,t} relative to resting depth ( Ξ” \Delta denotes the change from the previous event); two cumulative sums of DNFI over the preceding 50 and 200 events capture short- and medium-term flow momentum; and the queue imbalance QI t \mathrm{QI}{t} complements the flow signal with a static snapshot of resting-liquidity asymmetry. The stream is segmented into non-overlapping windows of L = 512 L=512 events, each forming one model input (Appendix A.6 ). 4.3 Encoder, Training and Model Size As encoder we use D-TABL, a deeper variant of the temporal attention-augmented bilinear network of Tran et al. (2018) adapted to our 512 512 -event windows (Appendix A.1 ). It is first pretrained on the three-class label alone, independently for each of the 21 asset–horizon pairs. UQ-LOB is then attached to the pretrained encoder and the whole model is fine-tuned end-to-end with Equation 5 (UQ-regression) or the weighted cross-entropy (UQ-classification). Each training instance consists of N = 16 N=16 windows, C = 15 C=15 context windows and one withheld target, and each step processes a batch of 16 such instances. The encoder has 0.27M parameters; the shared projection and context/attention trunk adds 0.79M, and each decoder 0.17M, so both variants total 1.23M parameters (Table 4 ). Training all 42 models requires about 210 GPU-hours on a single NVIDIA GH200; full details and hyperparameters are given in Appendices A.3 and A.4 . 5 Results 5.1 Calibration of the Regression Head Figure 3: UQ-regression on consecutive BTC-USD test windows at the 10 s horizon: predicted ΞΌ \mu with 68% and 95% intervals against the realised displacement; crosses mark realisations outside the 95% interval. Figure 3 gives a qualitative view of the predictive distribution, and Table 3 quantifies it across horizons. Empirical 68% coverage is 68.1–68.2%, essentially nominal, and it is stable across horizons. The 95% interval, however, covers only 90.4–90.7% of outcomes: the residual distribution is heavier tailed than a Gaussian, and Equation 8 constrains only the first absolute moment of the standardised residual, so it enforces the one-sigma level but not the tails. We return to this in Section 7 . 5.2 Confidence Gating Figure 4: Magnitude-weighted R 2 R^{2} (Equation 9 ) of the UQ-regression head as a function of the SNR percentile retained, per asset and horizon. Table 1: Directional macro F1 (mean of down and up F1; stationary excluded) as a function of the confidence percentile retained, for UQ-regression (SNR-gated) and UQ-classification (softmax-gated), pooled over the seven assets. Row 0 0 is the full test set. A prior-matched random classifier attains ( 1 βˆ’ p stat ) / 2 (1-p{\mathrm{stat}})/2 on the full test set: 0.337 at 5 s ( p stat = 0.326 p_{\mathrm{stat}}=0.326 ), 0.350 at 10 s ( p stat = 0.300 p_{\mathrm{stat}}=0.300 ) and 0.362 at 15 s ( p stat = 0.276 p_{\mathrm{stat}}=0.276 ); because gating changes the class mix of the retained subset, this full-set reference does not apply to the gated rows (Section 5.2 ). H 5s 10s 15s UQ-Regression UQ-Classification UQ-Regression UQ-Classification UQ-Regression UQ-Classification Conf. F1( ↓ \downarrow ) F1( ↑ \uparrow ) Mean F1( ↓ \downarrow ) F1( ↑ \uparrow ) Mean F1( ↓ \downarrow ) F1( ↑ \uparrow ) Mean F1( ↓ \downarrow ) F1( ↑ \upar The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective. as detailed in the full paper on Arxiv The same ai evaluation question is explored in Advancing Model Research in AgentX, which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!