PDE Inverse Problems#
A PDE inverse problem uses indirect observations of a state field to infer an unknown coefficient, source, boundary condition, or initial condition. The unknown may itself be a function, so the statistical model should be defined before a mesh or basis is chosen. This subsection separates that function-space model from the finite representations used to compute with it.
Parameter-to-observable map#
Let \(D\subset\mathbb{R}^d\) be the physical domain, let \(m\) denote the unknown parameter field, and let \(u\) denote the state. Write the governing equations schematically as
Let \(\mathcal{X}\) be a function space for \(m\), let \(\mathcal{V}\) be a state space for \(u\), and let \(\mathcal{X}_{\mathrm{ad}}\subseteq\mathcal{X}\) contain the physically admissible parameters. Assume that the forward problem has a unique solution for every \(m\in\mathcal{X}_{\mathrm{ad}}\). It then defines the solution map
Measurements rarely contain the entire state. An observation operator \(\mathcal{O}:\mathcal{V}\to\mathbb{R}^q\) may select sensor values, spatial averages, boundary fluxes, or other measured quantities. The parameter-to-observable map is
For a true parameter \(m^\dagger\), consider the additive Gaussian observation model
where \(\boldsymbol{\Gamma}\in\mathbb{R}^{q\times q}\) is a positive-definite measurement-noise covariance matrix. The corresponding data-misfit potential is
This model treats the PDE and its inputs as exact. Unknown forcing, boundary data, initial conditions, or model discrepancy must be included explicitly if they are scientifically relevant. Even when the forward PDE is well posed, sparse observations and smoothing dynamics can make the inverse map nonunique or unstable.
Function-space posterior#
Suppose that \(\mathcal{X}\) is a separable Hilbert space. A Gaussian prior on the unknown field is a probability measure
where \(m_0\in\mathcal{X}\) is the prior mean and \(\mathcal{C}:\mathcal{X}\to\mathcal{X}\) is a self-adjoint, positive, trace-class covariance operator. Trace class means that the covariance eigenvalues have a finite sum, which ensures that prior draws have finite expected squared norm. Choose the prior so that \(\mu_0(\mathcal{X}_{\mathrm{ad}})=1\). It should encode physical length scales, amplitudes, smoothness, and boundary behavior before the computational mesh is selected. If a coefficient must be positive, one may instead assign a Gaussian prior to an unconstrained field \(z\) and set \(m=\exp(z)\), provided that this transformation is well defined and scientifically appropriate.
Under standard conditions ensuring that the forward map and posterior are well defined, the posterior measure \(\mu^{\mathbf{y}}\) is defined relative to the prior by
with
Here \(0<Z(\mathbf{y})<\infty\). The Radon–Nikodym derivative is the posterior density relative to the prior: the observations reweight the probabilities already assigned by the prior. This formulation avoids introducing a nonexistent infinite-dimensional Lebesgue density. It also makes clear that the finite-dimensional computations below approximate one fixed function-space Bayesian inverse problem (Stuart, 2010).
Karhunen–Loève truncation#
Let \((\lambda_j,\varphi_j)\) be the eigenpairs of \(\mathcal{C}\), ordered so that \(\lambda_1\geq\lambda_2\geq\cdots\geq 0\). A draw from the Gaussian prior has the Karhunen–Loève representation
where the series converges in mean square in \(\mathcal{X}\). Retaining the first \(J\) modes gives
The mean-square prior truncation error is
This identity quantifies the error under the prior (Lord et al., 2014). The coefficients \(\boldsymbol{\xi}\) therefore provide a finite-dimensional parameterization with a standard normal prior. The corresponding posterior density is proportional to
The truncation order \(J\) is an approximation choice, not a physical parameter. A retained-prior-variance criterion can guide its initial selection, but it is not sufficient: a low-variance mode may still affect the observations. Posterior expectations and scientifically relevant predictions should be checked as \(J\) increases (Stuart, 2010).
Numerical approximation#
The exact map \(\mathcal{G}\) is replaced in computation by a numerical map \(\mathcal{G}_{h,\tau}\), where \(h\) denotes the spatial or temporal discretization scale and \(\tau\) denotes the algebraic-solver tolerance. Define the observable-space numerical error
At the true parameter,
Consequently, a likelihood built with \(\mathcal{G}_{h,\tau}\) can mistake numerical error for measurement noise or parameter information. The mesh should be refined and the solver tolerance tightened until posterior summaries and predictions are stable. If this is not affordable, the approximation error must be modeled and validated; it should not be absorbed into \(\boldsymbol{\Gamma}\) without a defensible stochastic model (Kaipio and Somersalo, 2005; Kaipio and Somersalo, 2007).
Mesh refinement and KLE refinement address different approximations. Refining \(h\) improves the PDE solve for a fixed field, while increasing \(J\) expands the field representation. A coherent calculation keeps the physical prior fixed as both are refined and checks convergence of posterior expectations or predictions, rather than comparing mesh-dependent parameter vectors alone (Cotter et al., 2010; Stuart, 2010). The prior regularizes weakly informed directions, but it does not create information that the experiment did not collect.
Coefficient and source inversion#
The thermal-conductivity example begins with a two-parameter coefficient field and shows how sparse temperature observations constrain the conductivity through a steady heat equation. The contaminant-location example infers a low-dimensional source location and shows how sensor symmetry can create a multimodal posterior. Together, they isolate the two principal mechanisms developed here: an unknown that changes the PDE operator and an unknown that enters its forcing.