Inverse problems involve the recovery of unknown model parameters from noisy indirect observations. They are central to myriad applications, e.g., in imaging, geophysics, and remote sensing. Current research involves understanding computational and statistical aspects of inference with complex models, particularly partial differential equations, and the development of new data-driven methodologies for formulating and solving inverse problems given rich sources of information.
This workshop will survey recent theoretical and algorithmic advances in the field, collectively illustrating an interplay between optimization, sampling, machine learning, and statistical theory. The workshop has natural connections to other FoCM 2026 workshops on Computational Optimal Transport, Foundations of Data Science and Machine Learning, Foundations of Numerical PDEs, and Continuous Optimization, among others.
Organizers
MIT
University of Chicago
University of Cambridge
Speakers
Semi-plenary speakers
The University of Texas at Austin
University of Klagenfurt
Invited speakers
Emory University
University of Sydney
Rice University
University of Manchester
Korea Advanced Institute of Science and Technology
University of Utah
University of Cambridge
EPFL
Columbia University
University of Klagenfurt
University of Vienna
FU Berlin
Bocconi University
EPFL
University of Chicago
INRIA
Monday, 13 July
14:00-14:30
Rebecca Willett (University of Chicago)
Inverse medium scattering is an ill-posed, nonlinear wave-based imaging problem arising in medical imaging, remote sensing, and non-destructive testing. Machine learning (ML) methods offer increased inference speed and flexibility in capturing prior knowledge about imaging targets than classical optimization-based approaches; however, they perform poorly in regimes where the scattering behavior is highly nonlinear. A key limitation is that ML methods struggle to incorporate the physics governing the scattering process, which are typically inferred implicitly from the training data or loosely enforced via architectural design. In this talk, I will present a method that endows a machine learning framework with explicit knowledge of the problem’s physics, in the form of a differentiable solver that represents the forward model. The proposed method progressively refines reconstructions of the scattering potential using measurements at increasing wave frequencies, following a classical strategy to stabilize recovery. Empirically, we find that our method provides high-quality reconstructions at a fraction of the computational or sampling costs of competing approaches. This is joint work with Owen Melia, Olivia Tsang, Vasileios Charisopoulos, Yuehaw Khoo, and Jeremy Hoskins.
14:30-15:00
Kui Ren (Columbia University)
The task of simultaneously reconstructing multiple physical coefficients in PDEs from observed data is ubiquitous in applications. We propose an integrated data-driven and model-based iterative reconstruction framework for such joint inversion problems where additional data on the unknown coefficients are supplemented for better reconstructions. Our method couples the supplementary data with the PDE model to make the data-driven modeling process consistent with the model-based reconstruction procedure. This coupling strategy allows us to characterize the impact of learning uncertainty on the joint inversion results for two typical inverse problems. Numerical evidence is provided to demonstrate the feasibility of using data-driven models to improve the joint inversion of multiple coefficients in PDEs.
15:00-15:30
Maarten De Hoop (Rice University)
Transformers are now central to modern machine learning, yet a rigorous understanding of their training dynamics remains incomplete, particularly in regimes involving many layers, many attention heads, and many tokens. We introduce a mathematical framework for the gradient-based training of transformers in a simultaneous mean-field limit in which depth (number of layers) and width (number of attention heads) are treated continuously. In contrast with training residual networks, whose infinite-depth limit leads to controlling a neural ordinary differential equation, transformers give rise to controlling a neural partial differential equations: Token distributions evolve through layers as probability measures, coupled to attention mechanisms whose parameters are themselves represented by layer-dependent probability measures.
Within this framework, we establish well-posedness of the forward transformer dynamics and characterize the evolution of token distributions by flow maps solving ordinary differential equations in suitable function spaces. We then derive the training dynamics as a conditional Wasserstein gradient flow over attention-parameter measures. Using adjoint sensitivity analysis, we obtain an explicit formula for the gradient of the risk in terms of backward adjoint equations, and prove existence and uniqueness of the resulting gradient-flow curves.
A central technical contribution is an analysis of the neural tangent kernel associated with attention. We give necessary and sufficient conditions for its injectivity, showing that this property is equivalent to the linear independence of certain log-sum-exp functions modulo affine functions. These conditions hold for broad classes of token distributions, including discrete measures, uniform distributions, and Gaussian mixtures. Under this injectivity assumption and a small-initial-loss condition, we prove convergence of the gradient flow to global minimizers. The results provide a rigorous foundation for training infinitely deep and wide transformers and clarify the role of attention-induced coupling in transformer optimization.
Joint research with R. Barboni, T. Furuya and G. Peyré.
16:00-16:30 Coffee Break
16:30-17:00
Richard Nickl (University of Cambridge)
We study optimal statistical inference procedures for the states of time evolution phenomena occurring in data assimilation’ or filtering problems. There it is a common practice to assign a Gaussian process prior on the initial condition of a dynamical system and to update it to a Bayesian posterior measure in the space of possible trajectories given a discrete sample of the process. In key applications the dynamics are non-linear, such as with Navier-Stokes equations in geophysical sciences or reaction-diffusion equations in biochemistry. While Bayesian posterior distributions are widely computed by filtering or MCMC methods, little is known about the statistical behaviour of these posterior measures in non-linear settings. In this talk we will introduce a theoretical framework for such models and then present recent results, known asBernstein-von Mises theorems’, that show that the posterior measures are approximated in function space by the Gaussian laws of solutions to certain SPDEs that involve the inverse Fisher information of the underlying statistical model.
17:00-17:30
Anna Little (University of Utah)
Motivated by biological imaging problems such as cryo-electron microscopy and cryo-electron tomography, this talk develops a functional approach to hidden signal recovery in low-SNR regimes where alignment or particle picking is unreliable. I will first discuss functional multi-reference alignment, where a connection to classical deconvolution and a multivariate Kotlarski identity enables recovery of non-bandlimited signals from second-order statistics. I will then describe recent work on functional multi-target detection, where many unknown translations of a compactly supported signal appear inside one large, noisy observation. In this setting, the bispectrum of the hidden signal is first estimated from third-order autocorrelations and then inverted via functional frequency marching or a Kotlarski-type integral formula, yielding explicit recovery guarantees that depend on signal smoothness and noise properties.
17:30-18:00
Mikyoung Lim (Korea Advanced Institute of Science and Technology)
Layer potential techniques are widely used to represent solutions of transmission problems in partial differential equations. For a simply connected planar inclusion, there exists a unique exterior conformal mapping, and geometric function theory relates the properties of this mapping to the geometry of the inclusion. We derive an analytic inversion formula for the conductivity inclusion problem by combining geometric function theory with the layer potential representation. We then present a Bayesian inversion approach for noisy measurements.
Tuesday, 14 July
14:00-14:30
Tiangang Cui (University of Sydney)
Optimal experimental design for inverse problems—a task of identifying the most informative observation functionals—is notoriously difficult. This difficulty stems primarily from the evaluation and optimisation of design criteria. The challenge is further exacerbated by several factors: the high computational cost of forward model evaluations in PDE-constrained settings; the growing number of such evaluations required as the dimensionality increases; and, most critically, the demands of the sequential setting, where each new design must be computed conditional on the updated solution of the inverse problem after assimilating the accumulated data.
We propose to address these difficulties using functional inequalities, which provide (i) an ultra fast surrogate for evaluating design criteria via a closed-form solution, and (ii) a quantitative measure of the effective dimensionality of the inverse problem, thereby bypassing the curse of dimensionality. To ensure the validity of the assumptions underlying these functional-inequality-based design criteria, we couple them with measure transport. This results in an integrated framework that not only overcomes the core challenges of optimal design but also makes the methodology naturally applicable to the otherwise intractable sequential design setting.
14:30-15:00
Julianne Chung (Emory University)
We consider hyperparameter estimation for inverse problems with Gaussian priors and forward models that may have uncertain model parameters. As the resulting posterior is non-Gaussian, in general, and sampling this distribution is computationally expensive, we use an empirical Bayes framework for estimating the maximum a posteriori estimate of the hyperpameters by considering the marginalized posterior distribution. Since the optimization problem is computationally challenging due to repeated evaluation of log determinants, we propose a Majorization-Minimization with Monte Carlo approach for hyperparameter estimation. Specifically, we replace a challenging optimization problem with a sequence of simpler ones by utilizing a majorization function for the log-determinant term, and we combine it with a Monte Carlo estimator for efficient computation. We provide theoretical results, and we demonstrate that the method can be used for estimating model parameters (e.g., unknown parameters that define the forward process) by considering separable nonlinear inverse problems.
15:00-16:00 semi-plenary talk
Omar Ghattas (The University of Texas at Austin)
We address real-time Bayesian inverse problems governed by time-shift-invariant wave equations, with particular focus on tsunami inference and optimal experimental design. Efforts are underway to instrument subduction zones with ocean bottom acoustic pressure sensors to provide tsunami early warning. Our goal is to create a physics-based early-warning system that employs this pressure data, along with the 3D coupled acoustic–gravity wave equations, to infer the earthquake-induced spatiotemporal seafloor motion in real time. The Bayesian solution of this inverse problem then provides the seafloor forcing to forward propagate the tsunamis toward populated areas along coastlines and issue forecasts with quantified uncertainties.
In the context of the Cascadia Subduction Zone, a single forward wave propagation requires 1 hour on a supercomputer. The Bayesian inverse problem, with a billion uncertain parameters, formally requires hundreds of thousands of adjoint wave propagations; thus real time inference appears to be intractable. We propose a novel approach to enable exact solution of the inverse and prediction problems in real time. The key is to exploit the time-shift-invariance of the parameter-to-observable map, which permits FFT diagonalization and fast GPU implementation. We demonstrate that tsunami inverse problems with a billion parameters can be solved exactly in a fraction of a second.
This fast Bayesian inversion capability is then exploited to solve the optimal experimental design problem of placement of seafloor pressure sensors to maximize expected information gain in predictive quantities of interest. Time permitting, we discuss data-driven prior
construction, goal-oriented dimension reduction, and construction of fast surrogates for nonlinear shallow water equation-based tsunami predictions, which are more accurate in shallower waters.
This work is joint with Stefan Henneking, Sreeram Venkat, Bowen Shi, and Yuhang Li at UT Austin, and Alice Gabriel at UCSD.
16:00-16:30 Coffee Break
16:30-17:00
Olivier Zahm (INRIA)
This contribution provides guidelines to derive a posteriori error estimators for quasi-optimal nonconforming finite element discretizations of symmetric elliptic problems. The resulting estimators are defined for all source terms that are admissible to the underlying weak formulations. More importantly, they are eqIn this talk, we address the problem of efficiently sampling multimodal probability distributions, where standard Markov Chain Monte Carlo methods often suffer from poor mixing and mode trapping. To mitigate these issues, we propose GRiLS (Gradient-free Riemannian Langevin Sampler), a novel preconditioning framework that enhances exploration without requiring gradient evaluations. Our approach introduces a Riemannian metric which reshapes the local geometry in order to facilitate transitions across modes. The resulting gradient-free MCMC algorithm is particularly suitable for complex, computationally expensive targets where derivatives are unavailable or impractical. Empirical results on multimodal benchmarks demonstrate that GRiLS achieves significantly improved mixing and convergence compared to existing gradient-based (MALA) and gradient-free (RW, pCN) MCMC approaches.uivalent to the error in a strict sense. In particular, their data oscillation part is bounded by the error and, furthermore, can be designed to be bounded by classical data oscillations. The estimators are computable, except for the data oscillation part. Since even the computation of some bound of the oscillation part is not possible in general, it must be handled on a case-by-case basis. The abstract framework is illustrated by examples involving second- and fourth-order problems.
17:00-17:30
Jonas Latz (University of Manchester)
Gaussian processes (GPs) have gained popularity as flexible machine learning models for regression and function approximation with an in-built method for uncertainty quantification. However, GPs suffer when the amount of training data is large or when the underlying function contains multi-scale features that are difficult to represent by a stationary kernel. To address the former, training of GPs with large-scale data is often performed through inducing point approximations, also known as sparse GP regression (GPR), where the size of the covariance matrices in GPR is reduced considerably through a greedy search on the data set. To aid the latter, deep GPs have gained traction as hierarchical models that resolve multi-scale features by combining multiple GPs. Posterior inference in deep GPs requires a sampling or, more usual, a variational approximation. Variational approximations lead to large-scale stochastic, non-convex optimisation problems and the resulting approximation tends to represent uncertainty incorrectly. In this work, we combine variational learning with MCMC to develop a particle-based expectation-maximisation method to simultaneously find inducing points within the large-scale data (variationally) and accurately train the deep GPs (sampling-based). The result is a highly efficient and accurate methodology for deep GP training on large-scale data. We test our method on standard benchmark problems.
17:30-18:00
Claudia Schillings (FU Berlin)
We present a randomised Ensemble Kalman Optimization method for linear inverse problems. The scheme injects additive noise into the ensemble update to overcome the subspace property and ensemble collapse that affect standard ensemble Kalman inversion in the small-ensemble regime. We also discuss a projected variant that controls the ensemble spread via a moving-ball constraint in observation space. We outline theoretical results on the behaviour of both schemes and illustrate their performance on representative examples.
This is joint work with Jonas Latz (University of Manchester), Alix Leroy (University of Edinburgh) and Simon Weissmann (University of Mannheim).
Wednesday, 15 July
14:00-14:30
Victor Panaretos (Ecole Polytechnique Fédérale de Lausanne)
The kernel trick is often associated with computational and technical shortcuts in statistics and machine learning. In this talk I revisit two classical problems, where the kernel trick reveals mathematical structure that can inform methodology. The first is computerised tomography, where through the prism of reproducing kernels, the X-ray transform acquires a projection-like interpretation. This leads to a function-valued regression framework that reconciles continuum theory with discrete/noisy practice. The second concerns two-sample testing, where probability measures can be embedded into Gaussian measures on an RKHS through combined first-and-second moment kernel embeddings. In this setting, distinguishing probability distributions becomes discriminating between mutually singular Gaussian measures, revealing a striking separation phenomenon. While both examples benefit from the kernel trick, what is arguably more noteworthy is the latent structure that is elicited and can be operationalised in statistical (and) inverse problems. Based on joint work with Ho Yun (Columbia), Leonardo Santoro (EPFL), and Kartik Waghmare (ETH Zürich).
14:30-15:00
Botond Szabo (Bocconi University)
We consider the recovery of an unknown function $f$ from a noisy observation of the solution $u_f$ to a partial differential equation that can be written in the form $\L u_f=c(f,u_f)$, for a differential operator $\L$ that is rich enough to recover $f$ from $\L u_f$. Examples include the time-independent Schr\”odinger equation $\Delta u_f = 2u_ff$, the heat equation with absorption term $(\partial_t -\Delta_x/2) u_f=fu_f$, and the Darcy problem $
abla\cdot (f abla u_f) = h$. We transform this problem into the linear inverse problem of recovering $\L u_f$ under the Dirichlet boundary condition, and show that Bayesian methods with priors placed either on $u_f$ or $\L u_f$ for this problem yield optimal recovery rates not only for $u_f$, but also for $f$. We also derive frequentist coverage guarantees for the corresponding Bayesian credible sets. Adaptive priors are shown to yield adaptive contraction rates for $f$, thus eliminating the need to know the smoothness of this function. The results are illustrated by numerical experiments on synthetic data sets.
15:00-15:30
Sven Wang (EPFL)
Bayesian and related methodologies have been extremely popular for estimation, prediction and inference in complex statistical models, for instance arising from differential equations (PDEs/SDEs). We present recent statistical and computational guarantees which underpin those methodologies. We will discuss statistical convergence with growing statistical sample size, as well as recent progress in studying the (polynomial) computational complexity of the numerical algorithms required. This includes polynomial-time mixing results for high-dimensional Markov Chain Monte Carlo (MCMC) methods as well as recent “generalized M-estimators” introduced in arxiv.org/abs/2601.09007.
16:00-16:30 Coffee Break
16:30-17:00
Otmar Scherzer (University of Vienna)
The starting point of this talk is the paper of Engl, Kunisch, Neubauer ”Convergence rates for Tikhonov Regularization of Nonlinear Ill–Posed Problems”, Inverse Problems, 5.3 (1989), 523-540. The analysis of this paper is based on functional analysis. We use their fundamental techniques to analyze Tikhonov regularization which is based on data-driven approaches. In order to do so we consider discretization of Tikhonov regularization, as it has been analyzed by Neubauer and S., Num. Funct. Anal. Opt. 11, 1-2 (1990), 85-99 ”Finite–dimensional approximation of Tikhonov regularized solutions of non-linear ill-posed problems”.
17:00-17:30
Elena Resmerita (University of Klagenfurt)
Image restoration in the presence of additive noise is well understood and
can be addressed by a wide range of methods. In contrast, non-additive noise
models are significantly more challenging due to their data-dependent nature.
In this work, we focus on multiplicative noise and adopt a variational approach.
While image denoising under multiplicative noise has been extensively studied
from both theoretical and numerical perspectives, the combined setting of blur
and multiplicative noise remains insufficiently understood, particularly at the
theoretical level. We propose a unified variational framework for image deblurring under mul-
tiplicative noise, that encompasses several well-established models, including
Rudin-Osher, Aubert-Aujol, and so on. Within this framework, we es-
tablish well-posedness, convergence results, and error estimates under natural
assumptions on the blurring operator.
17:30-18:30 semi-plenary talk
Barbara Kaltenbacher (University of Klagenfurt)
Image restoration in the presence of additive noise is well understood and can be addressed by a wide range of methods. In contrast, non-additive noise models are significantly more challenging due to their data-dependent nature. In this work, we focus on multiplicative noise and adopt a variational approach. While image denoising under multiplicative noise has been extensively studied from both theoretical and numerical perspectives, the combined setting of blur and multiplicative noise remains insufficiently understood, particularly at the theoretical level.
We propose a unified variational framework for image deblurring under multiplicative noise, that encompasses several well-established models, including Rudin-Osher, Aubert-Aujol, and so on. Within this framework, we establish well-posedness, convergence results, and error estimates under natural assumptions on the blurring operator.
Posters
Melissa Adrian
(Purdue University)
Data assimilation algorithms estimate the state of a dynamical system from partial observations, where the successful performance of these algorithms hinges on costly parameter tuning and on employing an accurate model for the dynamics. This paper introduces a framework for jointly learning the state, dynamics, and parameters of filtering algorithms in data assimilation through a process we refer to as auto-differentiable filtering. The framework leverages a theoretically motivated loss function that enables learning from partial, noisy observations via gradient-based optimization using auto-differentiation. We further demonstrate how several well-known data assimilation methods can be learned or tuned within this framework. To underscore the versatility of auto-differentiable filtering, we perform experiments on dynamical systems spanning multiple scientific domains, such as the Clohessy-Wiltshire equations from aerospace engineering, the Lorenz-96 system from atmospheric science, and the generalized Lotka-Volterra equations from systems biology. Finally, we provide guidelines for practitioners to customize our framework according to their observation model, accuracy requirements, and computational budget.
Matthias Beckmann
(University of Hamburg)
Many algorithms in optimization generate a sequence that converges weakly to a solution. Under linearity assumptions, sometimes one obtains more: strong convergence and a formula for the limit.
In this talk, I will report on recent work (In recent years, the topic of high dynamic range (HDR) tomography has started to gather attention due to recent advances in hardware technology. The issue is that registering high-intensity projections that exceed the dynamic range of the detector cause sensor saturation, which, in turn, leads to a loss of information. Inspired by the multi-exposure fusion strategy in computational photography, a common approach is to acquire multiple Radon projections at different exposure levels that are algorithmically fused to facilitate HDR reconstructions.
As opposed to this, we introduce a new single-shot approach which is based on the Modulo Radon Transform, a novel generalization of the conventional Radon transform. In this case, Radon projections are folded via a modulo non-linearity, which allows HDR values to be mapped into the dynamic range of the sensor and, thus, avoids saturation or clipping. We discuss how to recover a function from given MRT samples backed by mathematical reconstruction guarantees. Our theoretical results are illustrated by numerical simulations and prototype modulo ADC based hardware experiments, where we report recovery from measurements 10 times larger than the sensor’s dynamic range while benefiting from lower quantization noise.with Sedi Bartz, Yuan Gao, and Walaa Moursi) on the classical KM method and some accelerated algorithms.
Vasileios Charisopoulos
(University of Washington)
Computed tomography (CT) is an imaging modality with widespread use in biomedical imaging, non-destructive materials testing, and remote sensing. Standard approaches for inverting CT measurements, which rely on a numerically poorly conditioned preprocessing step, often struggle to reconstruct signals with high dynamic range (e.g., X-ray imaging of tissue with embedded metal). In this talk, I will present an iterative method based on nonsmooth optimization that operates directly on the measurement space and yields faster and more accurate reconstructions.
Joint work with R. Willett.
Link to paper: https://doi.org/10.1137/24M1678982
Ivan Hasenohr
(University of Klagenfurt)
We are interested in a parameter reconstruction problem emerging from MRI applications: consider a bilinear parabolic PDE of the form $\dot z^\star + H_0(\theta^\star, z^\star) + C(q, z^\star) = H_1\theta^\star$, where the state $z^\star(t,x)$ is subject to the action of a command $q(t,x)$ and of an unknown time-independent parameter $\theta^\star(x)$. Assuming partial state observations, we aim at reconstructing the parameter $\theta^\star$.
To do so, we introduce an online reconstruction method in the spirit of the works of Baumeister et al. (1997) and Kügler et al. (2010): considering on time-dependent state and parameter estimates $z(t,x)$ and $\theta(t,x)$ evolving under the combined action of observation-based numerical controls and of a carefully designed physical control $q^\star(t,x)$, we prove their convergence towards the true state and parameters.
This research was funded in whole or in part by the Austrian Science Fund (FWF) \href{https://dx.doi.org/10.55776/F100800}{10.55776/F100800}, and is joint work with Barbara Kaltenbacher.
Duc Hoang
(HU Berlin)
We consider the problem of learning nonlinear operators G_0: X ->Y from a finite number of samples by framing it as a nonparametric regression problem between Hilbert spaces. Two measurement models corresponding to infinite-dimensional versions of the random design model and the white noise model are considered. The rates derived in recent work by Reinhardt, Wang, and Zech (2024) are shown to be minimax optimal for the class of holomorphic operators that are uniformly bounded in a stronger output norm, in particular showing that the dimension-independent algebraic statistical rate n^(-kappa/(kappa+1)) is feasible. Here, kappa describes the approximation rate and depends on certain smoothness exponents (s, t) in the input and output Hilbert spaces. In particular, the curse of sample complexity in the learning theory of Lipschitz operators (or more general spaces of operators with finite regularity) is circumvented. As a consequence, the same proof techniques yield lower bounds on the approximation-theoretic continuous nonlinear N-widths, which are matched by the expression-rate bounds in Herrmann, Schwab, and Zech (2024). The applicability of our theory is shown in the context of the parameter-to-solution operator in the Darcy flow model. The resulting operator estimators may serve as data-driven surrogates for forward operators in Bayesian inverse problems.
Herrmann, L., Schwab, C., & Zech, J. (2024). Neural and spectral operator surrogates: unified construction and expression rate bounds. Advances in Computational Mathematics
Reinhardt, N., Wang, S., & Zech, J. (2024). Statistical learning theory for neural operators. arXiv preprint arXiv:2412.17582
Nick Huang
(Simon Fraser University)
Inverse problems are often framed in a probabilistic setting known as Bayesian inverse problems. Numerically, these problems are often solved by sampling from a prior distribution updated with noisy measurements, g = A(f) + e, a technique known as posterior sampling. Posterior sampling allows us to overcome the issue of ill-posedness by assigning probabilities to multiple solutions of an inverse problem, as well as quantifying any uncertainties arising from noisy data. Much of Bayesian inverse theory is centered around a Gaussian prior. However, in many applications, such as MRI, this is infeasible. Recently, generative models that learn priors through deep learning and extensive data have been able to produce high-quality samples, yet have little in the way of theoretical guarantees. In this poster, we present a framework to characterize the number of measurements required for stable and accurate recovery for posterior sampling when using arbitrary priors. We provide explicit bounds in the setting of linear inverse problems, such as MRI, with a particular focus on the case of generative priors. Finally, we present some preliminary results for infinite dimensional inverse problems modeled by PDEs. This poster is joint work with Ben Adcock (SFU), and is partially based on the following paper (NeurIPS, 2025): https://arxiv.org/abs/2505.10630.
Erik Jansson
(University of Cambridge)
Template-based reconstruction, also known as indirect shape matching, can be formulated as a matching problem in which a template is deformed to explain indirectly observed data through a given forward model. We study this problem by formulating a gradient flow directly on a Lie group which acts on a shape space. Equipping the deformation group with a right-invariant metric produces evolution equations. In finite-dimensional groups, this results in ordinary differential equations, while for infinite-dimensional groups it leads to partial differential equations. The poster focuses on the geometric derivation of these equations in a general setting, highlighting the roles of the group action, the choice of metric, and the induced momentum map. The framework applies to both finite-dimensional groups, such as direct products of SO(3), as well as infinite-dimensional groups, such as diffeomorphism groups, with examples illustrating how the general construction specializes in concrete cases.
Charlotte Lämbgen
(FU Berlin)
Optimal control problems governed by partial differential equations are widely used to design systems that achieve a desired target while minimizing control costs. Classical formulations typically assume that the governing model dynamics are known exactly. In practice, however, these dynamics often depend on uncertain parameters or incomplete knowledge of the underlying processes, which can significantly affect optimal decisions.
This project develops a Bayesian perspective on optimal control under uncertainty, combining ideas from optimal control theory and Bayesian inverse problems. The central idea is to reinterpret tracking-type optimal control problems as inference problems by introducing a probabilistic observation model that links the desired target with the system output through a potentially uncertain forward operator and observational noise. Within this framework, prior information about the control is encoded probabilistically, and the link to the original problem is established via the maximum a posteriori (MAP) estimator of the resulting posterior distribution.
As a first step, we analyze the linear Gaussian setting, where the forward operator is fixed. In this case, the posterior distribution can be derived explicitly, and the MAP estimator coincides with the minimizer of the classical quadratic optimal control problem. This establishes a precise link between deterministic optimal control formulations and Bayesian inference in both finite- and infinite-dimensional settings.
Building on this foundation, we investigate how the framework can be extended to uncertain forward models, where the system operator depends on unknown parameters. This leads to models in which the control and model parameters are treated jointly, raising new theoretical questions about marginalization, risk measures, and the resulting optimization objectives.
In the long term, the developed framework will be applied to epidemiological models of virus dynamics, where uncertainty in transmission rates, immunity, and behavioral responses plays an important role in designing robust intervention strategies such as vaccination policies.
David Mis
(Rice University)
Function graph transformers have been recently introduced as a measure-theoretic framework for learning nonlinear operators between function spaces. The central idea is to represent a function as a probability measure supported on the function’s graph. In this formulation, finite token sets are empirical graph measures, and refinement of the discretization corresponds to convergence of measures.
This viewpoint leads to a graph-preserving class of measure-theoretic transformers: transformer layers act by pushforward maps on measures while preserving the property that outputs remain single-valued functions. We show that this structural constraint is compatible with standard transformer architectures: graph-preserving measure maps can be approximated by finite compositions of softmax self-attention layers and pointwise multilayer perceptrons, with a particular volume-preserving condition on the positional encoding of the input tokens. The resulting universal approximation theory applies to broad classes of nonlinear operators, accommodates inputs of low or negative Sobolev regularity, and naturally supports query points on output domains different from the input domain.
As a computational demonstration, we apply this framework to acoustic wavespeed reconstruction from boundary measurements, treating the inverse Helmholtz problem as an operator from measured boundary responses to an interior coefficient field. Function graph transformers provides a natural way to encode irregular source-receiver-frequency measurements and query the reconstructed wavespeed throughout the spatial domain.
This poster is based on joint work with Takashi Furuya, Ivan Dokmanić, Maarten V. de Hoop, and Matti Lassas.
Pablo Muñoz
(University of Klagenfurt)
We present a computational framework for model-based quantitative imaging through the identification of spatially varying relaxation parameters in the Bloch-Torrey equation, a time dependent PDE governing magnetization dynamics in magnetic resonance imaging. The equation couples RF induced precession, longitudinal and transverse relaxation, and diffusion advection transport, a structure that poses particular challenges for both discretization and adjoint derivation due to the interplay of rotational, dissipative, and elliptic operators acting on different time scales.
The forward problem is discretized using a first order Lie operator splitting that preserves the physical structure of each sub process: precession is handled using the exact Rodrigues formula, relaxation using pointwise exponentials, and diffusion advection using an implicit P1 finite element scheme. The order of accuracy is rigorously verified using particular known solutions, confirming the expected temporal and spatial convergence rates.
A key contribution is the derivation of a discrete adjoint consistent with the splitting scheme. In particular, the adjoint of the Rodrigues sub-step is its exact inverse rotation, a nontrivial requirement that ensures gradient consistency without resorting to continuous adjoints or automatic differentiation. Gradients of the least-squares cost functional are used within an L-BFGS method formulated in the L²(Ω) inner product, enabling distributed parameter identification for three spatially varying fields: the longitudinal rate R₁, the transverse rate R₂, and the equilibrium magnetization M_eq.
Numerical experiments on a synthetic phantom demonstrate accurate simultaneous reconstruction of all three parameters from CPMG multi echo measurements, across a hierarchy of mesh refinements with warm-started interpolation between levels.
Marco Pauleti
(Karlsruhe Institute of Technology)
In this work, we present a unified convergence analysis of inexact Newton regularizations with general uniformly convex penalty terms for nonlinear ill-posed problems in Banach spaces. These schemes consist of an outer (Newton) iteration and an inner iteration, which provides the update of the current outer iterate. To this end, the nonlinear problem is linearized about the current iterate and the resulting linear system is inexactly solved by an inner regularization method. In our analysis, we only rely on generic assumptions of the inner methods, and we show that a variety of regularization techniques satisfies these assumptions. For instance, gradient-type and iterated-Tikhonov methods are covered. Numerical experiments based on the inverse problem of Electrical Impedance Tomography (EIT) illustrate the impact of different uniformly convex penalty terms.
Fabian Schneider
(TU Wien)
Many problems in science and engineering are challenging to model accurately due to incomplete knowledge of underlying physical mechanisms, poorly characterized measurement uncertainty, or the high computational cost of detailed simulations. As experimental data become increasingly available, machine learning methods provide a flexible alternative to explicit parametric modelling. We focus on the recently developed framework of neural likelihood approximation based on minimizing the Kullback–Leibler divergence between the true posterior and an approximate posterior.
Our approach is fully non-parametric and avoids restrictive parametric assumptions, such as Gaussian distributions or normalizing flows, by including the normalization into the objective.
We first analyze the properties of the objective function and characterize classes of functions that are equivalent with respect to it. We then investigate the implementation and convergence of data-driven methods as a function of the available data.
Finally, we validate our findings through numerical experiments on a PDE-based semiconductor model.
Vaibhav Silmana
(Indian Institute of Technology)
is a fundamental challenge in modern machine learning. A particularly important instance is covariate shift, where the input distributions change across domains while the conditional relationship between inputs and outputs remains invariant. While this setting has been extensively studied for scalar outputs and finite-dimensional problems, significantly less attention has been given to functional data, where outputs are functions or high-dimensional objects. Such scenarios arise naturally in applications including inverse problems, multi-task learning, and functional regression.
In this work, we develop a general regularization framework for unsupervised domain adaptation in vector-valued regression under the covariate shift assumption. Our approach is formulated within the framework of vector-valued reproducing kernel Hilbert space (vRKHS), which provides a principle way to model regression problems with functional outputs. By restricting the class of hypothesis spaces, we derive a practical learning algorithm based on importance-weighted regularized least squares, which is capable of handling distribution discrepancy under covariate shift.
We provide a comprehensive theoretical analysis of the proposed framework. In particular, under general source conditions and assumptions on the kernel and data distribution, we establish optimal convergence rates for the learned estimator both in the Bochner space L2 norm and in the vRKHS norm. These results extend classical learning theory for scalar-valued problems to the setting of vector-valued regression under distributional shift, providing a theoretical foundation for regularized operator learning with functional data.
To address the practical challenge of selecting regularization parameters and kernels, we introduce an aggregation strategy that combines multiple candidate estimators constructed with different regularization levels and kernel functions. This approach simultaneously resolves parameter tuning and kernel selection, and we provide a theoretical justification for its effectiveness.
Finally, we demonstrate the effectiveness of the proposed approach through numerical experiments on a real-world face image dataset, where the task is to reconstruct images from sinogram measurements under distributional distortions such as Gaussian and motion blur. The results show that the method remains robust under perturbations.
Overall, this work provides a first unified theoretical and algorithmic framework for learning from functional data under covariate shift, bridging operator learning, regularization theory, and domain adaptation.
Allard van Belois
(University of Vienna)
This work addresses the inverse problem of identifying unknown parameters in Wasserstein gradient flows using time continuous measurements at a limited number of spatial points. The core challenge is to select an appropriate subspace that enables accurate approximation of the underlying PDE while mitigating the curse of dimensionality.
By integrating adaptive approximation techniques, such as reduced order modeling and dynamical subspace selection, with the geometric structure of Wasserstein gradient flows, we aim to develop a framework for stable and efficient parameter reconstruction. The approach emphasizes the preservation of physical properties, such as entropy dissipation, and seeks to provide a robust methodology for inverse problems in transport dominated systems.
Xiaoye Wu
(University of Vienna)
In this poster we present a novel data assimilation scheme for approximating the solution and unknown physical parameter of time evolution PDEs from sparse observations.
Our approach approximates the solution in an infinite dimensional Hilbert space via a neural network with time-evolving weights. We propose a variational principle that updates continually both the network weights and the unknown physical parameter of the physical model at each time by minimising a joint functional of the PDE residual and observation error. Our variational principle reduces to a linear least-squares problem at each discrete time step when the PDE residual depends linearly on the physical parameter.
Lastly, to handle advection-dominated dynamics, where part of function that is of interest may drift out of observation range, we implement a dynamical sensor placement strategy that evolves sensor locations to preserve stability and accuracy of the reconstructed dynamics.
Ho Yun
(EPFL)
Quantum State Tomography (QST) is the gold standard for characterizing quantum systems, yet existing statistical frameworks often rely on rigid algebraic structures—such as Pauli observables—that assume sparsity. We introduce a unified statistical framework based on the Quantum Covariance Embedding (QCE). By mapping Positive Operator-Valued Measures (POVMs) into a Reproducing Kernel Hilbert Space (RKHS), we generalize classical kernel mean embeddings to the operator-valued setting and introduce the Quantum Maximum Discrepancy (QMD) to metrize the space of measurements.
We apply this geometry to QST, reformulating the reconstruction task as a tensorized linear regression problem. This perspective yields a unified design theory: we introduce the concept of Unitary Designs and identify a structural constant which analytically establishes the statistical superiority of entangled measurements (e.g., Mutually Unbiased Bases) over local measurements (e.g., Pauli Observables). Furthermore, we derive minimax lower bounds for general density estimation without sparsity assumptions. Finally, we propose the QUAntum Regression with Kernels (QUARK) estimator. Universal across quantum systems, QUARK achieves these optimal bounds and provides a flexible methodology that unifies standard linear inversion with smooth kernel regression.
Štěpán Zapadlo
(University of Graz)
Magnetic Resonance Imaging (MRI) is a key non-invasive imaging modality offering large versatility in creating high-resolution images. Reconstructing MR images from measurements requires knowledge of the MR physics involved in the measurement process, which are highly complex (e.g., involving quantum-mechanical effects) and are commonly modeled via the Bloch equations. However, the Bloch equations are only an approximation and neglect relevant effects such as magnetization transfer, which has been shown to play a significant role in certain imaging sequences. While multi-pool extensions like the Bloch–McConnell equations can account for these phenomena, they remain approximate and introduce additional modeling choices, including the number of pools. In this work, we take a more general perspective and address the challenge of uncovering hidden physics in MRI through a structured model learning approach. Specifically, we build upon the fundamental Bloch model and propose its extension with a potentially non-linear source term that captures the unknown dynamics associated with the multi-pool structure of the Bloch-McConnell model. The source term is formulated as a neural network with a highly interpretable architecture, strongly inspired by principles derived from the Bloch-McConnell model. We demonstrate the versatility and generalizability of the proposed framework and evaluate the learning process using artificially generated data in a multi-pool setting.
