# Difficulties fitting a mixture of dirichlet distributions

**URL:** <https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810>\
**Category:** Questions\
**Created:** [September 18, 2020, 4:37pm UTC](https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810 "2020-09-18T16:37:46Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![72nd](https://avatars.discourse-cdn.com/v4/letter/7/f1d935/32.png) [@72nd](https://discourse.pymc.io/u/72nd)\
**Post date:** [September 18, 2020, 4:37pm UTC](https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810/1 "2020-09-18T16:37:46Z")

</div>

I’m trying to fit a mixture of dirichlet distributions but am running into difficulties. While I don’t get any errors, the fits that I get are obviously incorrect. Here is my code:

```auto
%pylab inline
import numpy as np
import pymc3 as pm
import theano.tensor as tt
import pandas as pd
import random
import math

def create_data():
    dir1 = pm.distributions.multivariate.Dirichlet.dist(np.array([1, 5, 2]))
    dir2 = pm.distributions.multivariate.Dirichlet.dist(np.array([7, .5, 1]))
    dir3 = pm.distributions.multivariate.Dirichlet.dist(np.array([2, 3, 3]))
    data = np.concatenate((dir1.random(size=700), dir1.random(size=200), dir1.random(size=100)), axis=0)
    return data

def dirichlet(n_dim, suffix=""):
    if not isinstance(suffix, str):
        suffix = str(suffix)
    b = pm.HalfNormal("b" + suffix, sigma=10)
    a = pm.Dirichlet("a" + suffix, np.ones(n_dim))
    c = pm.Deterministic("c" + suffix, a * b)
    return pm.Dirichlet.dist(c, shape=3)

def stick_breaking(beta):
    portion_remaining = tt.concatenate([[1], tt.extra_ops.cumprod(1 - beta)[:-1]])
    return beta * portion_remaining
    
def estimate_model(data, n_clusters, n_features):
    with pm.Model() as model:
        alpha = pm.Gamma('alpha', 1., 1.)
        beta = pm.Beta('beta', 1, alpha, shape=n_clusters)
        w = pm.Dirichlet('w', stick_breaking(beta), shape=n_clusters)
        obs = pm.Mixture('obs', w, [dirichlet(3, k) for k in range(n_clusters)], observed=data)

        trace = pm.sample(50000, tune=10000)
        pm.traceplot(trace, ["w", "a0", "b0", "c0", "a1", "b1", "c1"])
        return pm.summary(trace)

data = create_data()
summ = estimate_model(data, 10, 3)
print(summ)

```

To match the data generating process, I should end up with 3 w’s with significant weight, at about .7, .2, and .1, and the c’s should reflect the vectors in the create\_data method. I never get anything close to this result.

Any suggestions?

---

<div class="post-metadata">

**Author:** ![72nd](https://avatars.discourse-cdn.com/v4/letter/7/f1d935/32.png) [@72nd](https://discourse.pymc.io/u/72nd)\
**Post date:** [September 22, 2020, 3:10pm UTC](https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810/2 "2020-09-22T15:10:40Z")

</div>

If it helps, here are the warnings I get when running:

```auto
Sampling 4 chains for 2_000 tune and 10_000 draw iterations (8_000 + 40_000 draws total) took 250 seconds.
There were 7001 divergences after tuning. Increase `target_accept` or reparameterize.
The acceptance probability does not match the target. It is 0.7014387282091336, but should be close to 0.8. Try to increase the number of tuning steps.
There were 7056 divergences after tuning. Increase `target_accept` or reparameterize.
There were 6769 divergences after tuning. Increase `target_accept` or reparameterize.
There were 6957 divergences after tuning. Increase `target_accept` or reparameterize.
The acceptance probability does not match the target. It is 0.6761237144949118, but should be close to 0.8. Try to increase the number of tuning steps.
The rhat statistic is larger than 1.4 for some parameters. The sampler did not converge.
The estimated number of effective samples is smaller than 200 for some parameters.

```

My traceplot looks like:

 ![discourse_trace](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/5/591765c0d0b3239b2cf39068073e47afac0a5af1.jpeg)

I have tried looking in previous discourse threads for answers, but have been unable to find any that address my situation. Is there any additional information I need to provide?

---

<div class="post-metadata">

**Author:** ![AlexAndorra](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/alexandorra/32/9142_2.png) [@AlexAndorra](https://discourse.pymc.io/u/AlexAndorra)\
**Post date:** [September 23, 2020, 8:44am UTC](https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810/3 "2020-09-23T08:44:41Z")

</div>

Hi @72nd,  
Dirichlet mixtures are quite hard to fit. I’m not well versed in it, but I’m guessing you have a “label-switching” problem: your `c`s seem to be multimodal. A usual fix is to constrain the `c`s to be ordered, IIRC.  
You can find a good introduction to these models in @aloctavodia’s book, [Bayesian Analysis with Python](https://github.com/aloctavodia/BAP).  
Hope this helps 🖖

---

<div class="post-metadata">

**Author:** ![72nd](https://avatars.discourse-cdn.com/v4/letter/7/f1d935/32.png) [@72nd](https://discourse.pymc.io/u/72nd)\
**Post date:** [September 23, 2020, 3:35pm UTC](https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810/4 "2020-09-23T15:35:29Z")

</div>

@AlexAndorra, thank you so much for replying.  
I’d run across the ordering solution in the past, but I don’t think I can apply it to multivariate data because there isn’t an ordering on tuples. Per your suggestion, I tried ordering the b’s, since they can be ordered. It did not meaningfully change the fit behavior, probably because the a’s kept label switching. It strikes me as plausible that ordering the a’s by their first component might fix this problem, but I don’t know how to do that. Is there some way to use an ordering transform on multidimensional data that I’m not aware of?

---

<div class="post-metadata">

**Author:** ![AlexAndorra](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/alexandorra/32/9142_2.png) [@AlexAndorra](https://discourse.pymc.io/u/AlexAndorra)\
**Post date:** [September 24, 2020, 11:04am UTC](https://discourse.pymc.io/t/difficulties-fitting-a-mixture-of-dirichlet-distributions/5810/5 "2020-09-24T11:04:16Z")

</div>

I’m sorry, I don’t know much about this type of models yet 😕  
Maybe other people here will be able to help you. I also hope you’ll be able to find something useful in Osvaldo’s book.  
Good luck with this, and sorry I couldn’t help more!
