# Confusion about MCMC and INLA

**URL:** <https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147>\
**Category:** version agnostic\
**Tags:** modeling\
**Created:** [November 20, 2024, 11:15am UTC](https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147 "2024-11-20T11:15:41Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![majf](https://avatars.discourse-cdn.com/v4/letter/m/ba9def/32.png) [@majf](https://discourse.pymc.io/u/majf)\
**Post date:** [November 20, 2024, 11:15am UTC](https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147/1 "2024-11-20T11:15:41Z")

</div>

Hi,

I’m having some confusion about when I can use the Laplace approximation instead of MCMC for the posterior distribution. I understand that the Laplace approximation requires the assumption that the posterior distribution is normal, whereas MCMC does not. For a problem where the likelihood probabilities are normally distributed and the prior probabilities are uniformly distributed, must the posterior probabilities be of the form of (or approximate to be) normally distributed? I’m rather confused in which case it would be skewed.

best regards

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [November 21, 2024, 3:50am UTC](https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147/2 "2024-11-21T03:50:01Z")

</div>

CC @theo

---

<div class="post-metadata">

**Author:** ![theo](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/theo/32/4945_2.png) [@theo](https://discourse.pymc.io/u/theo)\
**Post date:** [November 21, 2024, 11:35am UTC](https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147/3 "2024-11-21T11:35:02Z")

</div>

Hey, so there’s a couple of things at play here. Your question has INLA in the title which is a bit different to the vanilla Laplace approximation. I’ll answer the question based on the Laplace approximation and then I’ll explain the differences with INLA at the end.

## Laplace approximation

The Laplace approximation approximates the posterior density by fitting a multivariate Gaussian (normal) to _all_ the parameters with a mean vector equal to the MAP (mode, minimise the negative log-posterior density) estimate and a precision matrix (inverse covariance, inverted Hessian matrix of log posterior wrt params) evaluated at the mode.

Don’t worry too much about the details but as you are fitting a Gaussian, this becomes _exact_ when the posterior is Gaussian and a good approximation unless you have multimodal (many MCMC methods also struggle with this) or highly skewed distributions.

Richard McElreath uses the Laplace approximation in his early parts of Statistical Rethinking and I’m sure he will explain it clearer than me. If you are worried about whether it is a good approximation for your model and whether it is worth using an approximation for gains in speed, **try fitting both MCMC and the Laplace approximation**. If the Laplace approximation produces a similar posterior to MCMC then happy days.

## Integrated nested Laplace approximation (INLA)

With INLA, we approximate the marginal posterior distribution of _some subset_ of the parameters, referred to as the marginal Laplace approximation.

Then, integrate out the remaining parameters using another method.

- **Integrated**. Using numerical integration
- **Nested**. Because we need p(θ∣y) to get p(u∣y)
- **Laplace approximations**. The methods used to obtain parameters for the Gaussian approximation.

So the difference is we only use Laplace approximation on _some_ parameters, which we hope have a Gaussian posterior, and use some other, more flexible inference method (like numerical integration in the case of R-INLA) for the other parameters.

As in the [R-INLA docs](https://www.r-inla.org/faq#h.yjpp8vc0btg8), INLA works well when the “full conditional density for the latent field to be _near_ Gaussian”, i.e. the posterior density for the subset of the parameters we want to approximate with the Laplace is Gaussian.

PyMC work on INLA is currently paused (we need a pytensor optimiser and I need to sort my life out) but we are very, very open for contributions on this if people are interested in having this in PyMC. It’s not that far off and the [issue tracker is here](https://github.com/pymc-devs/pymc-experimental/issues/340).

I have [some notes on INLA](https://theorashid.github.io/notes/INLA) with further reading at the bottom, I also recommend [Adam Howes’ thesis chapter](https://athowes.github.io/thesis/naomi-aghq.html) on this which builds up from these ideas from scratch and compares to MCMC.

---

<div class="post-metadata">

**Author:** ![bob-carpenter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bob-carpenter/32/8124_2.png) [@bob-carpenter](https://discourse.pymc.io/u/bob-carpenter)\
**Post date:** [November 25, 2024, 11:28pm UTC](https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147/4 "2024-11-25T23:28:28Z")

</div>

The Laplace approximation is just a second-order Taylor expansion. You approximate a distribution with a multivariate normal distribution centered at the posterior mode with covariance given by the negative inverse Hessian. The gradient term drops out because you’re at a mode.

Many Bayesian models, particularly hierarchical models, do not have modes. In that case, you can’t use Laplace approximation for the whole distribution because there’s no mode. But in these situations, we often can used integrated nested Laplace approximation (aka INLA).

INLA is being carried out to marginalize a density p(\alpha, \beta) to p(\alpha) = \int p(\alpha, \beta) \, \text{d}\beta. Typically, this is used for hierarchical models where \alpha are the high-level population parameters (e.g., population mean and variance) and have a known prior p(\alpha) whereas the \beta are low-level coefficients such as regression coefficients or random effects, so \alpha is low-dimensional and \beta is high-dimensional. Then we factor p(\alpha, \beta) = p(\alpha) \cdot p(\beta \mid \alpha) and the nested Laplace approximation is of p(\beta \mid \alpha). Even though the distribution of p(\beta \mid \alpha) is usually simple, everything’s conditioned on the observed data y and we’re looking at doing all of this in the posterior p(\alpha, \beta \mid y). All the detail’s on the Wikipedia page (they use \theta, x were I used \alpha, \beta): [Integrated nested Laplace approximations - Wikipedia](https://en.wikipedia.org/wiki/Integrated_nested_Laplace_approximations)

---

<div class="post-metadata">

**Author:** ![majf](https://avatars.discourse-cdn.com/v4/letter/m/ba9def/32.png) [@majf](https://discourse.pymc.io/u/majf)\
**Post date:** [December 17, 2024, 4:51pm UTC](https://discourse.pymc.io/t/confusion-about-mcmc-and-inla/16147/5 "2024-12-17T16:51:21Z")

</div>

Thanks for the answer, I’m asking this question mainly because I’m confused as to how well the Laplace approximation approximates the posterior distribution, and I’m wondering if it would be helpful in this regard to evaluate the degree of approximation by resorting to solving for the radius of the domain of trust in the domain of trust algorithm, both of which perform a second-order Taylor expansion on the function? It feels like my thinking is perhaps a bit amateurish
