# Sampling from a NormalMixture with wide spread of parameters

**URL:** <https://discourse.pymc.io/t/sampling-from-a-normalmixture-with-wide-spread-of-parameters/9003>\
**Category:** version agnostic\
**Created:** [March 9, 2022, 3:59pm UTC](https://discourse.pymc.io/t/sampling-from-a-normalmixture-with-wide-spread-of-parameters/9003 "2022-03-09T15:59:32Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dhajnes](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/dhajnes/32/4434_2.png) [@dhajnes](https://discourse.pymc.io/u/dhajnes)\
**Post date:** [March 9, 2022, 3:59pm UTC](https://discourse.pymc.io/t/sampling-from-a-normalmixture-with-wide-spread-of-parameters/9003/1 "2022-03-09T15:59:32Z")

</div>

Hello,

I am trying to sample from a Gaussian Mixture (`pm.NormalMixture()`) that has two (as an illustratory example, in practice there’s more, e.g. 8 - 36) Normals mixed, although their means are too far away from each other.

Here: N\_1(300, 100^2) and N\_2(70000, 5000^2)  
The mixture is mixed with weights 0.7 \cdot N\_1 + 0.3 \cdot N\_2, where the weights are based on a Dirichlet distribution, which is above the NormalMixture in the hierarchy.

Although when I sample the posterior of the NormalMixture, there is mainly the N\_2, which should not happen. It should be more of the N\_1.

The code used to generate the following graphs:

```auto
with pm.Model() as miniBN:
    mat = pm.Dirichlet('mat', a=mat_probs)
    elasticity = pm.NormalMixture('elasticity', w=mat, mu=ELAS_MUS, sigma=ELAS_SIGS)
    burn_in = 500
    trace = pm.sample(10000, tune=5000, target_accept=0.9)
    chain = trace[burn_in:]
    pm.plot_trace(chain)
    pm.plot_posterior(chain)
    plt.show()

```

## Question:

Am I just undersampling? Should I sample more? It does not seem to help if I ramp up the sampling to say `pm.sample(50000, tune=1000, target_accept=0.9)`. What is the proper way of handling such situations?

# Graphs:

### Trace:

 ![Screenshot from 2022-03-09 16-42-33](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/0/067a058b8f0943b3dcb87a9a93d10e2870582e4f.png)

### Posterior:

 ![Screenshot from 2022-03-09 16-45-24](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/1/18204bd5434ebc8b233c46882d6ef4a418fe248b.png)

---

<div class="post-metadata">

**Author:** ![chartl](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/chartl/32/1515_2.png) [@chartl](https://discourse.pymc.io/u/chartl)\
**Post date:** [March 10, 2022, 6:04pm UTC](https://discourse.pymc.io/t/sampling-from-a-normalmixture-with-wide-spread-of-parameters/9003/2 "2022-03-10T18:04:34Z")

</div>

Try using SMC instead of NUTS.

> Sampling from distributions with multiple peaks with standard MCMC methods can be difficult, if not impossible, as the Markov chain often gets stuck in either of the minima. A Sequential Monte Carlo sampler (SMC) is a way to ameliorate this problem.

[https://docs.pymc.io/en/v3/pymc-examples/examples/samplers/SMC2\_gaussians.html](https://docs.pymc.io/en/v3/pymc-examples/examples/samplers/SMC2_gaussians.html)

---

<div class="post-metadata">

**Author:** ![dhajnes](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/dhajnes/32/4434_2.png) [@dhajnes](https://discourse.pymc.io/u/dhajnes)\
**Post date:** [March 10, 2022, 6:25pm UTC](https://discourse.pymc.io/t/sampling-from-a-normalmixture-with-wide-spread-of-parameters/9003/3 "2022-03-10T18:25:41Z")

</div>

Hi @chartl,

thanks for you reply. I have tried the Sequential sampling, although it looks like there is a problem of practically zero values between the two peaks. It’s basically a zero plane between (\mu=300) and (\mu=70000). That may be a reason why it retrusn `LinAlgError: singular matrix`.

Do you have any proposal how to go around this problem? Perhaps adding some really small value? That would probably end up in the use of general `Mixture()` which I would like to avoid.

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [March 10, 2022, 7:45pm UTC](https://discourse.pymc.io/t/sampling-from-a-normalmixture-with-wide-spread-of-parameters/9003/4 "2022-03-10T19:45:46Z")

</div>

I guess it’s just impossible to jump from one mode to the other. You might be better of using a non marginal mixture here, with a latent indicator variable and two latent normals for the modes
