# Hierarchical model - NUTS - warnings after sampling - reparametrization?

**URL:** <https://discourse.pymc.io/t/hierarchical-model-nuts-warnings-after-sampling-reparametrization/194>\
**Category:** Questions\
**Created:** [July 26, 2017, 10:14pm UTC](https://discourse.pymc.io/t/hierarchical-model-nuts-warnings-after-sampling-reparametrization/194 "2017-07-26T22:14:38Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![JWarmenhoven](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jwarmenhoven/32/51_2.png) [@JWarmenhoven](https://discourse.pymc.io/u/JWarmenhoven)\
**Post date:** [July 26, 2017, 10:14pm UTC](https://discourse.pymc.io/t/hierarchical-model-nuts-warnings-after-sampling-reparametrization/194/1 "2017-07-26T22:14:38Z")

</div>

I am trying to get rid of some divergence & max\_tree\_depth warnings in the models from Chapter 20 of Kruschke (2015). The models contain multiple nominal predictors. (I basically translated JAGS code into PyMC3)

I tried reparametrization of the models, which helped a lot in speeding up NUTS. But there are still issues. Since the original model was defined for JAGS, which uses Slice sampling if I am not mistaken, could the model/priors be suboptimal for NUTS?

> **[Jupyter Notebook Viewer](https://nbviewer.jupyter.org/github/JWarmenhoven/DBDA-python/blob/master/Notebooks/Chapter%2020.ipynb)**
>
> Check out this Jupyter notebook!

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [July 27, 2017, 9:08am UTC](https://discourse.pymc.io/t/hierarchical-model-nuts-warnings-after-sampling-reparametrization/194/2 "2017-07-27T09:08:19Z")

</div>

Did you try the new initialization for NUTS (if you upgrade to master it use `advi+adapt_diag` as initialization)? I had a quick try it seems works better.

Having looked at your models, I will comment on one point of potential improvement:  
The sigma (`ySigma` in the first two models and `sigma` in the last model) might have the best parameterization. In general, a Uniform prior for sigma is not ideal, and a `tt.max()` would likely break the gradient, maybe try a HalfCachy prior?
