# Non-centered parameterization of a beta distribution?

**URL:** <https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872>\
**Category:** Questions\
**Created:** [March 2, 2021, 12:09am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872 "2021-03-02T00:09:38Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![bridgeland](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bridgeland/32/2773_2.png) [@bridgeland](https://discourse.pymc.io/u/bridgeland)\
**Post date:** [March 2, 2021, 12:09am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/1 "2021-03-02T00:09:38Z")

</div>

@twiecki wrote a [great description](https://twiecki.io/blog/2017/02/08/bayesian-hierchical-non-centered/) of non-centering a hierarchical model, both why the sampler sometimes has difficulty exploring the space and what to do about it. [Diagnosing Biased Inference with Divergences](https://docs.pymc.io/notebooks/Diagnosing_biased_Inference_with_Divergences.html) is also a good description, showing tools for diagnosing divergences, and prescribing non-centering to avoid them. And Richard McElreath [describes](https://xcelab.net/rm/statistical-rethinking/) an amazingly simple model—which he dubs the _Devil’s Funnel_—that exhibits divergent transitions, fixable by reparameterizing to a non-centered representation (2e: page 421-423).

But all three of dsecriptions of non-centering share common feature: the variable that is reparameterized takes a normal distribution.

What if the problem variable is a beta instead? Is there a way to non-center a beta distribution?

For example, consider the following variation of McElreath’s _Devil’s Funnel_:

```
import pymc3 as pm
import numpy as np
import arviz as az

max_sigma = 0.25
with pm.Model() as lucifers_funnel:
    v = pm.HalfNormal('v', 3.0)
    x = pm.Beta('x', mu=0.5, sigma=max_sigma/pm.math.exp(v))

with lucifers_funnel:
    trace = pm.sample(random_seed=19950526)

```

This _Lucifer’s Funnel_ is a lot like McElreath’s _Devil’s Funnel_. And as with _Devil’s Funnel_, _Lucifer’s Funnel_ exhibits many divergences.

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/c/cda5fa55ea377699c2cef91fcdb65f7652f0715f.png)

_Lucifer’s Funnel_ desperately needs non-centering, to somehow move **v** out of **pm.Beta()**, and to redefine **x** with a **pm.Deterministic()**. But how? Is there a transformation that accomplishes non-centering for this beta?

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [March 2, 2021, 6:59am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/2 "2021-03-02T06:59:52Z")

</div>

Your model is not hierarchical right? ~~That’s what the non centered parametrization is usually needed for.~~

Isn’t your issue simply that sigma can easily get too close to zero? That would indicate you needed a more constrained prior on v, or model sigma directly.

Also did you try changing target accept or tune iterations? That can be a reasonable approach when you have a tough posterior.

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [March 2, 2021, 7:09am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/3 "2021-03-02T07:09:43Z")

</div>

To answer your original question though. I am not familiar with a way to parametrize the beta in a non centered way. You can’t simply multiply by the sigma and add a mean because it must be bound between 0 and 1. You can parametrize in terms of a Dirichlet times a concentration, but that is not really the same. The closest would be perhaps to drop the beta entirely and model it with a logitnormal: [Logit-normal distribution - Wikipedia](https://en.m.wikipedia.org/wiki/Logit-normal_distribution)

---

<div class="post-metadata">

**Author:** ![bridgeland](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bridgeland/32/2773_2.png) [@bridgeland](https://discourse.pymc.io/u/bridgeland)\
**Post date:** [March 2, 2021, 9:01pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/4 "2021-03-02T21:01:29Z")

</div>

My understanding is _Lucifer’s Funnel_ is in fact a hierarchical model. At least it is a variation of the McElreath’s _Devil’s Funnel_, described in a chapter devoted to hierarchical models.

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/7/715faeecb3a7427167ad85d001435c5359a67761.jpeg)

(If this is a copyright violation, many apologies. [Buy the book.](https://www.amazon.com/Statistical-Rethinking-Bayesian-Examples-Chapman/dp/036713991X/) It’s useful, well-written, and funny. I literally laughed out loud while reading it, frightening those around me.)

To answer your question, I can reduce the divergences by increasing target accept, both in this toy model and in my actual model. But I really would like to change the difficult geometry that causes the divergences, to remove the divergences entirely.

Approximate the beta with a logitnormal? I’ll look into that. But it seems a heavy-handed approach if there is some way to reparameterize the beta into some form that is mathematically identical but numerically different. Assuming such a form exists.

---

<div class="post-metadata">

**Author:** ![chartl](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/chartl/32/1515_2.png) [@chartl](https://discourse.pymc.io/u/chartl)\
**Post date:** [March 5, 2021, 5:18am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/7 "2021-03-05T05:18:35Z")

</div>

I don’t think this is so much a funnel as a problem with the boundary condition that \sigma^2 \< \mu(1-\mu). Even though you have the max sigma, your HalfNormal can still attempt to sample values near 0; and thus near the boundary.

Edit: It’s probably not even that. My guess is that the gradient of 1/\exp(v) blows up somewhere and generates the divergence; but with the following parameterization (which is logistic in the transformed variable) the gradients are better behaved and don’t cause the same kinds of problems.

Using the fully independent parameterization of:

\alpha = \frac{\mu}{\gamma}  
\beta = \frac{1-\mu}{\gamma}

whereby \sigma^2 = \frac{\gamma}{1 + \gamma}\mu(1-\mu)

does not generate such divergences:

```auto
import pymc3 as pm
import numpy as np
import arviz as az

max_sigma = 0.25
with pm.Model() as lucifers_funnel:
    loggamma = pm.Normal('lgamma', 3.0)
    gamma = pm.Deterministic('gamma', pm.math.exp(loggamma)/(1 + pm.math.exp(loggamma)))
    x = pm.Beta('x', mu=0.5, sigma=gamma * 0.5 * (1-0.5))

with lucifers_funnel:
    trace = pm.sample(random_seed=19950526)NUTS: [x, lgamma]
Sampling 4 chains, 0 divergences: 100%|██████████| 4000/4000 [00:00<00:00, 7567.42draws/s]

```

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [March 5, 2021, 6:04am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/9 "2021-03-05T06:04:11Z")

</div>

@chartl that’s quite interesting. Out of curiosity, do you see any advantage in an approach like this compared to just using a logitnormal for controlling both mu and sigma (i.e. discarding the beta distribution altogether)?

---

<div class="post-metadata">

**Author:** ![chartl](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/chartl/32/1515_2.png) [@chartl](https://discourse.pymc.io/u/chartl)\
**Post date:** [March 5, 2021, 9:40pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/10 "2021-03-05T21:40:13Z")

</div>

Depends on the setting. You can recognize the factor \gamma/(1+\gamma) as just an inflation factor over the expected Binomial variance; and in genetics the parameter \gamma (or sometimes 1-\gamma or sometimes the whole ratio) is known as F\_{\mathrm{ST}} and reflects the variability of allele frequencies across sub-populations. It is a common target for Bayesian inference.

I suppose that in a logitnormal model, \sigma^2 could serve the same purpose; but in this case it would reflect some latent-scale variance unlinked to an “excess over the binomial variance”; so it may be less clear to interpret.

Also if one observes Binomial counts, it’s convenient to stick with a conjugate distribution.

---

<div class="post-metadata">

**Author:** ![jonsedar](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jonsedar/32/2590_2.png) [@jonsedar](https://discourse.pymc.io/u/jonsedar)\
**Post date:** [March 8, 2021, 5:23am UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/11 "2021-03-08T05:23:46Z")

</div>

Nice one, thanks @chartl - I might try this parameterisation in future

---

<div class="post-metadata">

**Author:** ![bridgeland](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bridgeland/32/2773_2.png) [@bridgeland](https://discourse.pymc.io/u/bridgeland)\
**Post date:** [March 12, 2021, 3:47pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/12 "2021-03-12T15:47:19Z")

</div>

@chartl’s gamma-sigma parameterization also samples well in my actual model, not just the toy example worked here.

---

<div class="post-metadata">

**Author:** ![chartl](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/chartl/32/1515_2.png) [@chartl](https://discourse.pymc.io/u/chartl)\
**Post date:** [March 12, 2021, 6:01pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/13 "2021-03-12T18:01:24Z")

</div>

You can thank Sewall Wright for this one, I think. 🙂

---

<div class="post-metadata">

**Author:** ![bridgeland](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bridgeland/32/2773_2.png) [@bridgeland](https://discourse.pymc.io/u/bridgeland)\
**Post date:** [March 12, 2021, 10:03pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/14 "2021-03-12T22:03:25Z")

</div>

When I see him, I will thank him. Hopefully not soon.

---

<div class="post-metadata">

**Author:** ![chartl](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/chartl/32/1515_2.png) [@chartl](https://discourse.pymc.io/u/chartl)\
**Post date:** [March 18, 2021, 2:46pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/15 "2021-03-18T14:46:55Z")

</div>

While I was searching for something completely unrelated, I came across this old thread about beta parametrization, and it links through to “implicit normalization of gradients”:

> [@Matt's trick / central formulation of beta distribution?](https://discourse.pymc.io/t/matts-trick-central-formulation-of-beta-distribution/1728/2):
>
> You can have a look at the recent paper [Implicit Reparameterization Gradients](https://arxiv.org/pdf/1805.08498.pdf), many of the tricks should also apply. Also see [https://en.wikipedia.org/wiki/Beta\_distribution#Related\_distributions](https://en.wikipedia.org/wiki/Beta_distribution#Related_distributions)

---

<div class="post-metadata">

**Author:** ![bridgeland](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bridgeland/32/2773_2.png) [@bridgeland](https://discourse.pymc.io/u/bridgeland)\
**Post date:** [March 22, 2021, 7:43pm UTC](https://discourse.pymc.io/t/non-centered-parameterization-of-a-beta-distribution/6872/16 "2021-03-22T19:43:28Z")

</div>

I’m still puzzling my way through the math in that paper. But implementing their approach would seem to require making changes to PyMC3, rather than just reparameterizing the model. (Or perhaps I am misunderstanding.)
