What the paper is about
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different velocity constraints. We find that flow matching has advantages for estimating distances involving complicated distributions and continuous variables. Furthermore, this framework enables the design of new distance metrics in a principled way. Together, flow matching provides a unified approach for understanding, estimating, and designing neural representational dissimilarity metrics. The same ai evaluation question is explored in A Risk-Adaptive and Evidence-Constrained Framework for..., which adds a research perspective.
What it covers
A Flow Matching Framework for Neural Representational Dissimilarity Thanks: Code: https://github.com/AgeYY/FlowRDM . Zeyuan Ye Affiliation: Department of Neuroscience, The University of Texas at Austin Xue-Xin Wei Affiliation: {y.zeyuan, weixx}@utexas.edu Abstract Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different velocity constraints. We find that flow matching has advantages for estimating distances involving complicated distributions and continuous variables. Furthermore, this framework enables the design of new distance metrics in a principled way. Together, flow matching provides a unified approach for understanding, estimating, and designing neural representational dissimilarity metrics. Figure 1: Overview of the proposed flow matching framework for neural representational dissimilarity. ( A ) Different stimuli evoke condition-specific neural responses, producing different population-response distributions whose dissimilarity is to be quantified. ( B ) Flow matching learns velocity fields v 1 𝒱 ( x , t ) v_{1}^{\mathcal{V}}(x,t) and v 2 𝒱 ( x , t ) v_{2}^{\mathcal{V}}(x,t) that transport a shared source distribution to the fitted distributions q 1 𝒱 q_{1}^{\mathcal{V}} and q 2 𝒱 q_{2}^{\mathcal{V}} . Here, we define representational dissimilarity as the Jeffreys divergence between q 1 𝒱 q_{1}^{\mathcal{V}} and q 2 𝒱 q_{2}^{\mathcal{V}} . We prove that different choices of 𝒱 \mathcal{V} correspond to different distance metrics. 1 Introduction A central question in neuroscience is how the brain represents information through stochastic neural population responses. Representing information means not only “storing” it but also “structuring” it in an appropriate format to support downstream computation given the biological constraints ( Kriegeskorte and Kievit, 2013 ; Kriegeskorte and Wei, 2021 ) . How can we study representational structures? Representational similarity analysis (RSA) takes a geometric perspective, probing representational structure by measuring pairwise distances between neural population response distributions across conditions, such as stimuli or task states ( Kriegeskorte et al., 2008a ; Nili et al., 2014 ) . RSA has been widely used to study how neural representations support perception and task performance ( Kriegeskorte et al., 2008a ; Cichy et al., 2014 ; Diedrichsen and Kriegeskorte, 2017 ) and to compare representational systems across species and computational models, including deep neural network models ( Kriegeskorte et al., 2008b ; Kriegeskorte, 2009 ; Kornblith et al., 2019 ) . A key step in RSA is to define and estimate representational distances. Because neural responses are stochastic, this often becomes a mathematical question of how to compare distributions. No single distance metric is best for every purpose as different metrics reveal different properties of the distributions ( Walther et al., 2016 ; Diedrichsen and Kriegeskorte, 2017 ; Kriegeskorte and Wei, 2021 ) . For example, correlation distance compares the relative pattern of mean activity across neurons. Euclidean distance measures separation between mean response vectors. Mahalanobis distance additionally accounts for response covariance. Kullback-Leibler (KL) divergence compares full response distributions. Lastly, Fisher information measures how sensitive the response distribution changes with a small change of the stimulus variable ( Kullback and Leibler, 1951 ; Amari, 2016 ) . Given the diversity of these useful distance metrics, a natural question arises: can they be unified under a single mathematical framework? Such a framework, if exists, could offer several benefits. First, it would make the statistical assumptions underlying different metrics more explicit, clarifying their interpretation. Second, it would provide a principled way to design new distance metrics. Finally, it could provide a unified estimator for these metrics, so that improvements to the estimator could benefit multiple metrics rather than requiring separate efforts for each one. Here, we provide such a framework using flow matching. Flow matching is a modern deep generative modeling method that is widely used in image, video, and language generation ( Lipman et al., 2023 ; Esser et al., 2024 ; Davtyan et al., 2023 ; Hu et al., 2024 ) . It starts by generating a data sample from a simple source distribution (e.g., a Gaussian distribution) and then evolves the sample toward a target distribution through a velocity field. In our flow-matching framework for neural dissimilarity, we define the representational distance as the Jeffreys divergence between two target neural-response distributions under a given velocity constraint. We mathematically prove that this Jeffreys divergence reduces to existing distance metrics under different velocity-field constraints (Fig. 1 ). We show that flow matching offers advantages when estimating distances involving complex distributions and continuous variables, and allows the principled design of new distance metrics. 2 Background and Related Work Representational dissimilarity analysis. RSA summarizes neural, behavioral, or model representations by a representational dissimilarity matrix (RDM), whose entries measure the pairwise separation between conditions ( Kriegeskorte et al., 2008a ; Kriegeskorte, 2009 ; Kriegeskorte and Kievit, 2013 ; Nili et al., 2014 ) . Common RDM entries include correlation, Euclidean, Mahalanobis, and cross-validated Mahalanobis distances; these choices differ in how they treat response scale, noise covariance, and finite-sample bias ( Walther et al., 2016 ; Diedrichsen and Kriegeskorte, 2017 ) . Fisher information estimation. Fisher information measures how rapidly p θ ( x ) p_{\theta}(x) changes with a continuous stimulus or behavioral variable, θ \theta ( Fisher, 1922 ; Amari, 2016 ; Seung and Sompolinsky, 1993 ; Pouget et al., 2000 ) . Its inverse lower-bounds the variance of unbiased estimators, making Fisher information a local measure of discriminability ( Seriès et al., 2009 ; Rao, 1945 ; Brunel and Nadal, 1998 ; Abbott and Dayan, 1999 ; Averbeck et al., 2006 ) . With repeated trials at nearby stimulus values, Fisher information can be estimated from tuning-curve derivatives and response covariance or from decoder-based estimator ( Kanitscheider et al., 2015 ; Kohn et al., 2016 ) . Continuous and naturalistic experiments, however, have no exact repeats. Estimation procedures need to exploit information shared across nearby values of the condition θ \theta , for example by using Gaussian processes ( Rasmussen and Williams, 2006 ; Ye and Wessel, 2025 ; Nejatbakhsh et al., 2023 ) . However, Gaussian processes scale poorly ( Bruinsma et al., 2020 ) . Our flow-matching framework unifies Fisher information estimation with the other distance metrics that have traditionally been treated separately. Continuous normalizing flows and flow matching. Normalizing flows model densities through invertible transformations ( Rezende and Mohamed, 2015 ; Papamakarios et al., 2021 ) , while continuous normalizing flows transport samples through an ODE ( Chen et al., 2018 ) . Roughly speaking, a sample begins at x 0 ∼ ρ 0 x_{0}\sim\rho_{0} , follows the velocity field v v over time, and arrives at a distribution intended to approximate the observed data. In the original continuous-flow formulation, the velocity parameters are trained by maximum likelihood, which requires solving the ODE and estimating the divergence of the velocity field during optimization, which can be slow ( Chen et al., 2018 ; Grathwohl et al., 2019 ) . Flow matching reformulates this training problem. Instead of maximizing likelihoods, one specifies a probability path between source samples and data samples, computes the target velocity along that path, and trains v v by regressing onto those velocities ( Lipman et al., 2023 ; Albergo and Vanden-Eijnden, 2023 ) . Flow matching thus provides a practical way to learn a transport map from a source distribution to an empirical response distribution. 3 A flow matching framework for computing representational distance Flow matching learns a time-dependent velocity field that transports samples from a simple source distribution to a target distribution ( Lipman et al., 2023 ) . Consequently, different constraints on the velocity field yield different approximations to the target distribution. Let x ∈ ℝ d x\in\mathbb{R}^{d} denote a neural population response vector, where d d is the number of neurons, and let p c ( x ) p_{c}(x) denote the unknown target distribution under condition c ∈ 𝒞 c\in\mathcal{C} . Starting from a source sample X 0 ∼ ρ 0 X_{0}\sim\rho_{0} , sampling follows the ODE d X t d t = v c 𝒱 ( X t , t ) , t ∈ [ 0 , 1 ] , \frac{dX_{t}}{dt}=v^{\mathcal{V}}{c}(X{t},t),\qquad t\in[0,1], (1) where the velocity field v c 𝒱 ( x , t ) v_{c}^{\mathcal{V}}(x,t) specifies the direction and speed of the sample at each flow time and 𝒱 \mathcal{V} denotes the chosen velocity class. Integrating Eq. ( 1 ) from t = 0 t=0 to t = 1 t=1 defines a transformation T c , 1 𝒱 T_{c,1}^{\mathcal{V}} and the fitted distribution q c 𝒱 = ( T c , 1 𝒱 ) # ρ 0 , q_{c}^{\mathcal{V}}=(T_{c,1}^{\mathcal{V}}){#}\rho{0}, (2) so that T c , 1 𝒱 ( X 0 ) ∼ q c 𝒱 T_{c,1}^{\mathcal{V}}(X_{0})\sim q_{c}^{\mathcal{V}} when X 0 ∼ ρ 0 X_{0}\sim\rho_{0} . The fitted distribution q c 𝒱 q_{c}^{\mathcal{V}} approximates p c p_{c} . Here, t t is an artificial flow-matching coordinate, not physical time. In general, our framework has three steps: (1) choose a source distribution and a velocity class 𝒱 \mathcal{V} ; (2) learn the velocity field by flow matching to approximate the target distributions; and (3) use the learned model to compute the Jeffreys divergence between the approximated distributions as the representational distance. Notation. We use a , b ∈ 𝒞 a,b\in\mathcal{C} for a pair of categorical conditions being compared and μ c = 𝔼 X ∼ p c [ X ] \mu_{c}=\mathbb{E}{X\sim p{c}}[X] for the mean response under condition c c . We use θ ∈ Θ \theta\in\Theta for a continuously varying condition, such as stimulus orientation, physical time, or behavior. For simplicity, we describe the framework below using the categorical condition c c ; the same framework applies to the continuous condition θ \theta . 1. Choosing a source distribution and velocity class. We first choose a source distribution ρ 0 \rho_{0} and a velocity class 𝒱 \mathcal{V} . The source specifies the initial distribution, which is typically shared across conditions. Meanwhile, the velocity class specifies the transformations allowed from the source to each target. For example, an unconstrained velocity field, parameterized by a neural network, permits flexible nonlinear transformations, whereas a constant, state-independent velocity restricts the transformation to a translation. These choices determine which properties of the target distributions can be represented (see Fig. 2 ) and, consequently, which representational distance is induced. Section 4 presents the specific choices used to recover different distance metrics. Figure 2: Velocity classes 𝒱 \mathcal{V} used in flow matching determine the fitted distributions q c 𝒱 q_{c}^{\mathcal{V}} (also see Table 1 ). (A) A two-component Gaussian-mixture target, shown as samples and dotted density contours. (B) A constant velocity translates the Gaussian source (black region) without changing its shape. (C) An affine velocity allows an affine transformation of the source distribution, yielding a Gaussian approximation to the target. (D) An unconstrained velocity captures the two modes of the target. Arrows indicate velocity direction, not magnitude. 2. Flow-matching training. Given the source distribution and velocity class, flow matching learns the velocity field with a simple regression objective ( Lipman et al., 2023 ; Albergo and Vanden-Eijnden, 2023 ) . During training, we draw source samples X 0 ∼ ρ 0 X_{0}\sim\rho_{0} and observed responses X 1 ( c ) ∼ p c X_{1}^{(c)}\sim p_{c} and connect them along a predefined path X t ( c ) = α t X 0 + β t X 1 ( c ) , ( α 0 , β 0 ) = ( 1 , 0 ) , ( α 1 , β 1 ) = ( 0 , 1 ) , X_{t}^{(c)}=\alpha_{t}X_{0}+\beta_{t}X_{1}^{(c)},\qquad(\alpha_{0},\beta_{0})=(1,0),\quad(\alpha_{1},\beta_{1})=(0,1), (3) where α t \alpha_{t} and β t \beta_{t} are prescribed interpolation schedules controlling the contributions of the source and target samples at flow time t t . The boundary conditions ensure that the path begins at the source sample and ends at the recorded response. Its target instantaneous velocity is U t ( c ) = α ˙ t X 0 + β ˙ t X 1 ( c ) , U_{t}^{(c)}=\dot{\alpha}{t}X{0}+\dot{\beta}{t}X{1}^{(c)}, (4) where the dots denote derivatives with respect to t t . A neural network receives ( X t ( c ) , t , c ) (X_{t}^{(c)},t,c) and is trained to predict this target velocity. The flow-matching objective is ℒ FM ( v ) = 𝔼 C ∼ π 𝒞 , t ∼ Unif [ 0 , 1 ] X 0 ∼ ρ 0 , X 1 ( C ) ∼ p C [ ‖ v C ( X t ( C ) , t ) − U t ( C ) ‖ 2 ] , \mathcal{L}{\mathrm{FM}}(v)=\mathbb{E}{\begin{subarray}{c}C\sim\pi_{\mathcal{C}},;t\sim\operatorname{Unif}[0,1]\ X_{0}\sim\rho_{0},;X_{1}^{(C)}\sim p_{C}\end{subarray}}\left[\left|v_{C}(X_{t}^{(C)},t)-U_{t}^{(C)}\right|^{2}\right], (5) where C ∼ π 𝒞 C\sim\pi_{\mathcal{C}} denotes a randomly sampled categorical condition. For a continuous condition, we instead sample Θ ∼ π Θ \Theta\sim\pi_{\Theta} and replace C C , p C p_{C} , X 1 ( C ) X_{1}^{(C)} , and v C v_{C} by Θ \Theta , p Θ p_{\Theta} , X 1 ( Θ ) X_{1}^{(\Theta)} , and v Θ v_{\Theta} , respectively. The population minimizer v 𝒱 = arg min v ∈ 𝒱 ℒ FM ( v ) v^{\mathcal{V}}=\arg\min_{v\in\mathcal{V}}\mathcal{L}{\mathrm{FM}}(v) is learned by backpropagation through the velocity network. Optionally, the learned velocity field can be fine-tuned by maximizing the log likelihood (Section B.1 ). We use this fine-tuning only for the new distance metric introduced in Supplementary Section C , where it is practically useful for optimization; all other results use flow-matching training alone for simplicity. 3. Jeffreys-divergence computation. We define the representational distance between conditions a a and b b as the Jeffreys divergence between their fitted distributions ( Kullback and Leibler, 1951 ; Amari, 2016 ) , D 𝒱 ( a , b ) = D KL ( q a 𝒱 ∥ q b 𝒱 ) + D KL ( q b 𝒱 ∥ q a 𝒱 ) , D{\mathcal{V}}(a,b)=D_{\mathrm{KL}}!\left(q_{a}^{\mathcal{V}}\middle|q_{b}^{\mathcal{V}}\right)+D_{\mathrm{KL}}!\left(q_{b}^{\mathcal{V}}\middle|q_{a}^{\mathcal{V}}\right), (6) where each directed KL divergence is D KL ( q a 𝒱 ∥ q b 𝒱 ) = 𝔼 x ∼ q a 𝒱 [ log q a 𝒱 ( x ) − log q b 𝒱 ( x ) ] . D_{\mathrm{KL}}!\left(q_{a}^{\mathcal{V}}\middle|q_{b}^{\mathcal{V}}\right)=\mathbb{E}{x\sim q{a}^{\mathcal{V}}}\left[\log q_{a}^{\mathcal{V}}(x)-\log q_{b}^{\mathcal{V}}(x)\right]. (7) When the source distribution is Gaussian and the velocity class 𝒱 \mathcal{V} is affine (linear in the state), the fitted distributions q c 𝒱 q_{c}^{\mathcal{V}} can be obtained by solving deterministic ODEs for their means and covariances (Supplementary Section A.3 ), from which the Jeffreys divergence can be computed. In other cases, we estimate the Jeffreys divergence using Monte Carlo. We draw samples X i ( a ) ∼ q a 𝒱 X_{i}^{(a)}\sim q_{a}^{\mathcal{V}} and X i ( b ) ∼ q b 𝒱 X_{i}^{(b)}\sim q_{b}^{\mathcal{V}} for i = 1 , … , M i=1,\ldots,M by integrating the sampling ODE in Eq. ( 1 ) and compute D ^ 𝒱 ( a , b ) = 1 M ∑ i = 1 M [ log q a 𝒱 ( X i ( a ) ) − log q b 𝒱 ( X i ( a ) ) ] + 1 M ∑ i = 1 M [ log q b 𝒱 ( X i ( b ) ) − log q a 𝒱 ( X i ( b ) ) ] . \widehat{D}{\mathcal{V}}(a,b)=\frac{1}{M}\sum{i=1}^{M}\bigl[\log q_{a}^{\mathcal{V}}(X_{i}^{(a)})-\log q_{b}^{\mathcal{V}}(X_{i}^{(a)})\bigr]+\frac{1}{M}\sum_{i=1}^{M}\bigl[\log q_{b}^{\mathcal{V}}(X_{i}^{(b)})-\log q_{a}^{\mathcal{V}}(X_{i}^{(b)})\bigr]. (8) These log densities are computed using the continuous change-of-variables identity in Eq. ( 111 ). 4 Flow matching with different velocity constraints lead to different representational metrics We now present our main theoretical results, i.e., various commonly used representational distance metrics arise from our flow matching framework as Jeffreys divergences under different velocity classes. These correspondences are summarized in Table 1 . For each velocity class, the proof follows a similar general strategy. We substitute the analytical form of the constrained velocity field into the population flow-matching loss and solve for its minimizer, which determines the induced approximate distribution. Evaluating the Jeffreys divergence between the resulting condition-specific approximations then yields the corresponding distance metric. Complete proofs for all rows of Table 1 are provided in the Supplementary Section A.4 . Table 1: Distance metrics and their corresponding velocity classes. We use c c for a categorical condition and θ \theta for a continuous condition. All rows use the shared standard Gaussian source ρ 0 = 𝒩 ( 0 , I ) \rho_{0}=\mathcal{N}(0,I) except the starred geometry-template distance, which uses a source distribution concentrated around a low-dimensional manifold (Supplementary Section C ). The cosine and correlation rows equal 2 r 2 2r^{2} times their conventional definitions. In the Mahalanobis row, A t A_{t} is shared across conditions. In the linear-Fisher row, A ¯ θ , t \bar{A}{\theta,t} is shared locally across nearby conditions (see Supplementary Section A.4 ). Velocity field Induced distance v c ( x , t ) = β ˙ t b c v{c}(x,t)=\dot{\beta}{t}{\color[rgb]{0.5703,0.1367,0.1367}b{c}} Euclidean distance v c ( x , t ) = β ˙ t b c , ‖ b c ‖ = r v_{c}(x,t)=\dot{\beta}{t}{\color[rgb]{0.5703,0.1367,0.1367}b{c}},\ |{\color[rgb]{0.5703,0.1367,0.1367}b_{c}}|=r Cosine distance v c ( x , t ) = β ˙ t b c , 1 ⊤ b c = 0 , ‖ b c ‖ = r v_{c}(x,t)=\dot{\beta}{t}{\color[rgb]{0.5703,0.1367,0.1367}b{c}},\ \mathbf{1}^{\top}{\color[rgb]{0.5703,0.1367,0.1367}b_{c}}=0,\ |{\color[rgb]{0.5703,0.1367,0.1367}b_{c}}|=r Correlation distance v c ( x , t ) = β ˙ t b c + A t ( x − β t b c ) v_{c}(x,t)=\dot{\beta}{t}{\color[rgb]{0.5703,0.1367,0.1367}b{c}}+{\color[rgb]{0.5703,0.1367,0.1367}A_{t}}(x-\beta_{t}{\color[rgb]{0.5703,0.1367,0.1367}b_{c}}) Mahalanobis distance v c ( x , t ) = β ˙ t b c + A c , t ( x − β t b c ) v_{c}(x,t)=\dot{\beta}{t}{\color[rgb]{0.5703,0.1367,0.1367}b{c}}+{\color[rgb]{0.5703,0.1367,0.1367}A_{c,t}}(x-\beta_{t}{\color[rgb]{0.5703,0.1367,0.1367}b_{c}}) Gaussian Jeffreys divergence v θ ( x , t ) = β ˙ t b θ + A ¯ θ , t ( x − β t b θ ) v_{\theta}(x,t)=\dot{\beta}{t}{\color[rgb]{0.5703,0.1367,0.1367}b{\theta}}+{\color[rgb]{0.5703,0.1367,0.1367}\bar{A}{\theta,t}}(x-\beta{t}{\color[rgb]{0.5703,0.1367,0.1367}b_{\theta}}) Linear Fisher information v c ( x , t ) , v θ ( x , t ) ∈ 𝒱 unconstrained {\color[rgb]{0.5703,0.1367,0.1367}v_{c}(x,t)},\ {\color[rgb]{0.5703,0.1367,0.1367}v_{\theta}(x,t)}\in\mathcal{V}{\mathrm{unconstrained}} Jeffreys divergence; full Fisher information v c ( x , t , τ ) , v θ ( x , t , τ ) {\color[rgb]{0.5703,0.1367,0.1367}v{c}(x,t;\tau)},\ {\color[rgb]{0.5703,0.1367,0.1367}v_{\theta}(x,t;\tau)} Auxiliary τ \tau -dependent distance v c ( x , t ) = a c , t + Ω c , t ( x − o c , t ) + λ c , t ( x − o c , t ) , Ω c , t ⊤ = − Ω c , t v_{c}(x,t)={\color[rgb]{0.5703,0.1367,0.1367}a_{c,t}}+{\color[rgb]{0.5703,0.1367,0.1367}\Omega_{c,t}}(x-{\color[rgb]{0.5703,0.1367,0.1367}o_{c,t}})+{\color[rgb]{0.5703,0.1367,0.1367}\lambda_{c,t}}(x-{\color[rgb]{0.5703,0.1367,0.1367}o_{c,t}}),\ {\color[rgb]{0.5703,0.1367,0.1367}\Omega_{c,t}}^{\top}=-{\color[rgb]{0.5703,0.1367,0.1367}\Omega_{c,t}} Geometry-template Jeffreys divergence ∗ ∗ : new distance introduced in this paper; dark red : outputs from trainable neural networks Fig. 2 illustrates this correspondence for three representative velocity classes. A velocity field that is constant in x x can only translate the Gaussian source. At the population optimum, this translation shifts the source mean to the target-distribution mean while leaving its covariance unchanged, so the Jeffreys divergence reduces to the squared Euclidean distance between target means. An affine-in- x x velocity field can both translate and linearly transform the Gaussian source. Its population optimum therefore produces the moment-matched Gaussian approximation to the target distribution, and the resulting Jeffreys divergence is the Gaussian Jeffreys divergence. Finally, in the population limit, an unconstrained velocity field recovers the target distribution itself, yielding the full Jeffreys divergence. Beyond a unification of existing distance metrics, the flow-matching framework enables principled designs of new metrics by specifying the source distribution as a distributional “template” and choosing a velocity class that constrains how the template can be transformed. The resulting distance is the Jeffreys divergence between the two fitted target distributions. As a demonstration, we design a new distance metric using a velocity class restricted to similarity transformations, thereby preserving the geometry of the source distribution when approximating the target distributions (Supplementary Section C ). This metric can be useful when the geometry of the target distributions is known a priori or when one wishes to emphasize a particular geometric property. 5 Applications We evaluate the flow-matching framework across a number of applications. First, we use a simple synthetic categorical dataset to test whether different velocity constraints recover multiple representational distance metrics (Section 5.1 ). The remaining four settings involve continuous variables: Fisher information in simulations and neural recordings (Section 5.2 ), and time-resolved RDMs in simulations and neural recordings (Section 5.3 ). The later four settings highlight an important advantage of flow matching for complex distributions that depend on continuous variables. 5.1 Flow matching recovers multiple distance metrics on a synthetic categorical dataset We first test whether practically flow matching with different velocity constraints enables the estimation of commonly used representational distance metrics, as predicted by our analytical results. We construct a five-condition Gaussian dataset in ℝ 3 \mathbb{R}^{3} , and fit flow matching with the corresponding velocity constraint for each metric in Table 1 . Accuracy, measured by the Pearson correlation between the unique off-diagonal entries of the estimated and ground-truth RDMs, improves with total sample size across all six metrics (Fig. 3 A). The estimated RDMs closely resemble the corresponding ground-truth RDMs defined by the six metrics (Fig. 3 B), supporting our main theoretical claim that flow matching can be used as a unified estimator for these metrics. Notably, most of these metrics depend only on first- and second-order moments, for which robust and efficient estimators such as Ledoit–Wolf shrinkage are available ( Ledoit and Wolf, 2004 ) . Indeed, when comparing flow matching with these specialized estimators on this simple Gaussian dataset, we did not observe advantages in estimated accuracy for flow matching (Fig. S1 ). In later applications, we will show that for datasets involving complex distributions and continuous variables, flow matching exhibits various advantages. Figure 3: Flow matching recovers distance metrics on a synthetic categorical dataset. ( A ) Pearson correlation between the unique off-diagonal entries of the estimated and ground-truth RDMs as a function of total sample size N N . Curves and error bars show means and sample standard deviations over five repetitions; the inset visualizes the five-condition dataset. ( B ) Ground-truth RDMs (top) and the corresponding flow-matching estimates using 2,000 total samples (bottom). 5.2 Flow matching estimates Fisher information in simulations and neural data Fisher information is a local measure of representational distance. Linear Fisher information quantifies how precisely small stimulus changes can be estimated by an optimal, locally unbiased linear readout of neural responses ( Brunel and Nadal, 1998 ; Abbott and Dayan, 1999 ; Averbeck et al., 2006 ; Moreno-Bote et al., 2014 ; Kanitscheider et al., 2015 ; Le and Wei, 2026 ) , whereas full Fisher information measures the sensitivity of the entire response distribution through the log-likelihood score, including information accessible only through nonlinear readouts. Estimating Fisher information can be challenging, partly because it is difficult to model how response distributions depend on continuous conditions. In contrast, our flow-matching approach can, in principle, interpolate naturally across conditions by approximating the velocity field with a deep neural network, without requiring additional data binning or modeling assumptions. Below, we evaluate the performance of this flow-matching framework in estimating both quantities. Simulations. We first estimate linear Fisher information using a toy dataset whose conditional response distribution is Gaussian given a continuous variable θ \theta . We fit a condition-dependent affine (linear-in-state) velocity field, v θ ( x , t ) = β ˙ t b θ + A ¯ θ , t ( x − β t b θ ) v_{\theta}(x,t)=\dot{\beta}{t}b{\theta}+\bar{A}{\theta,t}(x-\beta{t}b_{\theta}) (Table 1 ). As baselines, we consider several existing methods. Local linear estimator (OLE) estimates linear Fisher information by binning the data and performing linear classification on adjacent bins ( Kanitscheider et al., 2015 ) . Gaussian-process kernel regression (GKR; see Section D.1.3 ) uses Gaussian-process regression to estimate the distributional mean as a function of θ \theta and then uses kernel regression to estimate covariances ( Ye and Wessel, 2025 ) . The Wishart process jointly models the mean and covariance with Gaussian processes ( Nejatbakhsh et al., 2023 ) (Section D.1.4 ). On this Gaussian dataset, affine flow yields a smooth estimate that closely follows the ground-truth linear Fisher-information when varying the sample-size and number of neurons(Fig. 4 B, C). GKR and the Wishart process perform fairly well on low-dimensional datasets, but their errors increase substantially with response dimension (Fig. 4 C; also see Pearson correlations in Supplementary Figure S2 A). Figure 4: Flow matching performs best in estimating Fisher information on toy datasets. ( A ) Illustration of the Gaussian conditional-response dataset. ( B ) Representative linear Fisher-information estimates ( d = 300 , N = 1000 d=300,N=1000 ). ( C ) Mean relative absolute error (MRAE) versus total sample size at d = 300 d=300 (left) and versus response dimension at N = 1,000 N=1{,}000 (right). ( D ) Equal-weight Gaussian scale-mixture dataset. ( E ) Representative full Fisher-information estimates. ( F ) MRAE versus total sample size at d = 20 d=20 (left) and versus response dimension at N = 7,000 N=7{,}000 (right). Points and error bars in ( C ) and ( F ) show means and standard deviations over five repetitions. Furthermore, we evaluate full Fisher information on another Gaussian mixture dataset (Fig. 4 D). GKR and the Wishart process fail in estimation because they assume Gaussian distributions, whereas unconstrained flow successfully targets the full Fisher information (Fig. 4 E and F). Neural data. Next, we apply flow matching to neural recording data. We analyze six mouse visual-cortex recording sessions from Stringer et al. (2021) , each containing responses from 10,000 to 25,000 neurons across about 4,000 static-grating trials (Fig. 5 A). Direct estimation of Fisher information in this high-dimensional response space is statistically challenging and computationally demanding. Following common dimensionality-reduction practice ( Cunningham and Yu, 2014 ) , we performed PCA separately for each session using only the training trials and apply the fixed projection to the training, validation, and test responses (200 principal components, about 70% of the variance). All models are fitted and evaluated in this dimensionality-reduced PCA space. Figure 5: Flow matching estimates Fisher information in mouse visual-cortex recordings. ( A ) Experimental paradigm: visual-cortex responses are recorded while mice view static gratings. ( B ) Session-wise mean held-out log likelihood relative to GKR. Gray lines connect estimates from the same session. Bars and error bars show the across-session mean and standard error of the mean (SEM). ( C ) Linear Fisher information from affine flow and full Fisher information from unconstrained flow versus grating orientation. Gray lines show six individual sessions, each from a different mouse; black curves and error bars show the across-session mean and SEM. For a given orientation range, visualizations of the PC1–PC2 subspace show that, while the Gaussian estimators (Bin + LW, GKR, the Wishart process, and affine flow) generally fit the held-out data well, some discrepancies remain (Supplementary Fig. S3 ). By contrast, unconstrained flow captures a non-elliptical density contour and visually fits the data better. Quantitatively The same ai evaluation question is explored in AutoViewMem, which adds a research perspective. as detailed in the full paper on Arxiv The same ai evaluation question is explored in LLM-Generated Feature Pools for Time Series..., which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!