# About sampling start with find\_MAP

**URL:** <https://discourse.pymc.io/t/about-sampling-start-with-find-map/936>\
**Category:** Questions\
**Created:** [March 13, 2018, 9:30am UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936 "2018-03-13T09:30:36Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![square](https://avatars.discourse-cdn.com/v4/letter/s/49beb7/32.png) [@square](https://discourse.pymc.io/u/square)\
**Post date:** [March 13, 2018, 9:30am UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/1 "2018-03-13T09:30:36Z")

</div>

I have been looking for solutions to my model because didn’t give good convergence. When I was searching for information I saw some people using

> start = pm.find\_MAP()  
> train\_trace = pm.sample( train\_samp, tune= train\_tune, start=start, njobs=4 )

and on the documentation it states

> map : Use the MAP as starting point. This is discouraged.

My questions are

1. Is _pm.sample( …,start=start, …)_ equivalent to pm.sample( …,init= ‘map’, …)?
2. Why is it discouraged?  
I tried with and without _start= pm.find\_MAP()_, and the convergence with the later one is improved a lot. But then I saw the documentation that it’s not suggested… I am wondering if there was just luck or something I missed or problems I didn’t know…

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [March 13, 2018, 9:55am UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/2 "2018-03-13T09:55:47Z")

</div>

> [@square](#):
>
> Is pm.sample( …,start=start, …) equivalent to pm.sample( …,init= ‘map’, …)?

Yes

> [@square](#):
>
> Why is it discouraged?

The short answer is that in high dimension, the MAP is usually not in the typical set (where most of the volume of the posterior parameter space). There are quite a lot of resource explaining this (e.g., concentration of measure in high dimension etc).

> [@square](#):
>
> the convergence with the later one is improved a lot

It depends on how you measure convergence. In some low dimensional cases, it might be faster as the MAP is indeed in the typical set. However, in most (high dimension) cases, if using find\_MAP() you get a nice trace it usually means there are latent problems with your model (One scenario could be that there are multimode, and using MAP as start the trace is stucked in just one of them).  
In general, the default works better in a sense that, if it is slow or not converging you can be pretty sure that there are either problem of your model, or places could be optimized.

---

<div class="post-metadata">

**Author:** ![square](https://avatars.discourse-cdn.com/v4/letter/s/49beb7/32.png) [@square](https://discourse.pymc.io/u/square)\
**Post date:** [March 13, 2018, 10:18am UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/3 "2018-03-13T10:18:12Z")

</div>

Thanks a lot for the explanation. 😃  
I would like to ask about the documentation

> For almost all continuous models, `NUTS` should be preferred. There are hard-to-sample models for which NUTS will be very slow causing many users to use Metropolis instead. This practice, however, is rarely successful. NUTS is fast on simple models but can be slow if the model is very complex or it is badly initialized. In the case of a complex model that is hard for NUTS, Metropolis, while faster, will have a very low effective sample size or not converge properly at all. A better approach is to instead try to improve initialization of NUTS, or reparameterize the model.

Does _initialization of NUTS_ mean something like

> start = pm.sampling.init\_nuts( init=‘nuts’,…)  
> pm.sample( …, start=start , …)

But wouldn’t it also be a problem if the model has multimode?

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [March 13, 2018, 11:42am UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/4 "2018-03-13T11:42:12Z")

</div>

The initialization of NUTS means something like:

```python
step = pm.NUTS(... #initial parameters)

```

There are many parameters in NUTS, for example, the step size, the mass matrix, the parameter for dual averaging for step size… Some of these parameters will be set if you are calling the step method directly as above, which is what we refer to as initialization. As a result, in general you _should not_ initialized the step method as above.

Instead, what we recommend is to use the `init=...` kwarg in `pm.sample()` to do the initialization. The main thing to point out is that, although internally it is calling the function `pm.sampling.init_nuts( init=‘jitter+adapt_diag’,…)`, it is returning not only the start point but more importantly a scheme for tunning the mass matrix of NUTS during some tunning steps.

> [@square](#):
>
> But wouldn’t it also be a problem if the model has multimode?

There is always a problem if the modes are well separated. But I think in practice, usually the modes are not as well separate and a good initialization could give appropriate scaling to the NUTS kernel - the random velocity is then have enough energy to go from one mode to the other.

---

<div class="post-metadata">

**Author:** ![square](https://avatars.discourse-cdn.com/v4/letter/s/49beb7/32.png) [@square](https://discourse.pymc.io/u/square)\
**Post date:** [March 13, 2018, 1:57pm UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/5 "2018-03-13T13:57:48Z")

</div>

Thank you for the reply!! 🙂

---

<div class="post-metadata">

**Author:** ![square](https://avatars.discourse-cdn.com/v4/letter/s/49beb7/32.png) [@square](https://discourse.pymc.io/u/square)\
**Post date:** [March 13, 2018, 4:06pm UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/6 "2018-03-13T16:06:08Z")

</div>

In the documentation

> pymc3.sampling.sample(draws=500, step=None, init=‘auto’, n\_init=200000, …)

Is _n\_init_ suggested to be a great number because of some reason? Because it takes quite a long time when I have 800 datapoints 5 nodes in input layer , 5 nodes hidden layer… I was wondering if I can freely choose a number or if I should take something into consideration when I change it.  
Thank a lot. 🙂

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [March 13, 2018, 4:41pm UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/7 "2018-03-13T16:41:29Z")

</div>

So I take it you are using advi for initialization? Usually it stop well before 200000 (we have some simple heuristic to check if it converges and do an early stop if so)

---

<div class="post-metadata">

**Author:** ![square](https://avatars.discourse-cdn.com/v4/letter/s/49beb7/32.png) [@square](https://discourse.pymc.io/u/square)\
**Post date:** [March 13, 2018, 7:23pm UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/8 "2018-03-13T19:23:21Z")

</div>

No I was using nuts, I though it should take the value of draw, but it seems to take 200000 if no value is passed into the function.

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [March 13, 2018, 7:56pm UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/9 "2018-03-13T19:56:06Z")

</div>

Oh you are using `init=nuts`? in that case yes you can reduce the number of `n_init`. But I dont think it usually works too well. The default `init=jitter+adapt_diag` is usually recommended

---

<div class="post-metadata">

**Author:** ![bob-carpenter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bob-carpenter/32/8124_2.png) [@bob-carpenter](https://discourse.pymc.io/u/bob-carpenter)\
**Post date:** [May 24, 2024, 8:44pm UTC](https://discourse.pymc.io/t/about-sampling-start-with-find-map/936/10 "2024-05-24T20:44:09Z")

</div>

I came here to see what PyMC’s current suggestion was, because I just wrote [a blog post about initializing at the mode](https://statmodeling.stat.columbia.edu/2024/05/24/hmc-fails-when-you-initialize-at-the-mode/) on Andrew Gelman’s blog where I tried initialization at the mode for a 10K-dimensional standard normal. Summary: it’s a bad idea, but NUTS can compensate with its default warmup.

> [@square](#):
>
> Why is it discouraged?

First, there may not be a mode. In hierarchical models, you don’t get modes unless you have specialized zero-avoiding priors.

To add a little more detail to what @junpenglao said, the problem isn’t just that the mode is atypical. We start with random inits in the tails all the time. The difference is that the mode has the maximum density and so any proposal is by definition of a lower density state. The density usually drops off very fast moving away from the mode in high dimensions, making the acceptance rate basically zero for any reasonably large step size. This is a problem for HMC, for the original NUTS, and for the multinomial form of NUTS introduced by Michael Betancourt and currently used in Stan.

Luckily, the default step size tuning in NUTS will compensate by starting with a small step size in Phase I of warmup and then finishing with a reasonable step size in Phase III of warmup before sampling. (This is assuming you do the warmup similarly to how we do it in Stan or have another means of ongoing stepsize adaptation).
