# Index for hierarchical model

**URL:** <https://discourse.pymc.io/t/index-for-hierarchical-model/5738>\
**Category:** Questions\
**Created:** [September 4, 2020, 8:01pm UTC](https://discourse.pymc.io/t/index-for-hierarchical-model/5738 "2020-09-04T20:01:25Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![newerjazz](https://avatars.discourse-cdn.com/v4/letter/n/a87d85/32.png) [@newerjazz](https://discourse.pymc.io/u/newerjazz)\
**Post date:** [September 4, 2020, 8:01pm UTC](https://discourse.pymc.io/t/index-for-hierarchical-model/5738/1 "2020-09-04T20:01:25Z")

</div>

Hi I am having problems with indexing a hierarchical model and have read the examples and searched for previous questions.

Below is working code that reproduces problem. There are 3 expts; each expt contains 1000 data points. How do I index them all? Thanks in advance!

import numpy as np  
import pymc3 as pm  
import theano.tensor as tt

n\_drops = 1000

s1 = np.random.normal(35,7,n\_drops)  
s2 = np.random.normal(40,6,n\_drops)  
s3 = np.random.normal(45,4,n\_drops)  
s = np.stack((s1,s2,s3),axis=0)

mu1 = np.mean(s1)\*np.ones(n\_drops)  
mu2 = np.mean(s2)\*np.ones(n\_drops)  
mu3 = np.mean(s3)\*np.ones(n\_drops)  
mu = np.stack((mu1,mu2,mu3),axis=0)

sd1 = np.std(s1)\*np.ones(n\_drops)  
sd2 = np.std(s2)\*np.ones(n\_drops)  
sd3 = np.std(s3)\*np.ones(n\_drops)  
sd = np.stack((sd1,sd2,sd3),axis=0)

######### Non-hierarchical works ##############################  
tau\_true = 1.5  
active\_true = s \> mu + tau\_true \* sd  
Y\_true = np.multiply(s, active\_true)  
Y\_true = np.sum(Y\_true,axis=1)

with pm.Model() as model\_c:  
ϵ = pm.HalfCauchy(‘ϵ’, 50)  
tau\_ = pm.Normal('tau\_ ', mu = 0, sd = 2)

```
active = s > mu + tau_ * sd

μ = pm.Deterministic('μ', ((s*active)).sum(axis=1))
Y = pm.Normal('Y', mu=μ, sd=ϵ, observed=Y_true)

trace_c = pm.sample(cores=1)

```

pm.traceplot(trace\_c);  
print(pm.summary(trace\_c))

######### Hierarchical indexing problem ##############################

## The 3 datasets belong to 2 diff categories

type\_idx = np.array((0,0,1))

tau\_true = np.array([1.5, 1.5, 2.0]).reshape(3,1)  
active\_true = s \> mu + tau\_true \* sd  
Y\_true = np.multiply(s, active\_true)  
Y\_true = np.sum(Y\_true,axis=1)

with pm.Model() as model\_d:  
ϵ = pm.HalfCauchy(‘ϵ’, 50)  
tau\_ = pm.Normal('tau\_ ', mu = 0, sd = 2, shape=2)

```
active = s > mu + tau_[type_idx] * sd
## get error here 
## ValueError: Input dimension mis-match. (input[0].shape[1] = 3, input[1].shape[1] = 1000)

μ = pm.Deterministic('μ', ((s*active)).sum(axis=1))
Y = pm.Normal('Y', mu=μ, sd=ϵ, observed=Y_true)

trace_d = pm.sample(cores=1)

```

pm.traceplot(trace\_d);  
print(pm.summary(trace\_d))

---

<div class="post-metadata">

**Author:** ![newerjazz](https://avatars.discourse-cdn.com/v4/letter/n/a87d85/32.png) [@newerjazz](https://discourse.pymc.io/u/newerjazz)\
**Post date:** [September 8, 2020, 2:45am UTC](https://discourse.pymc.io/t/index-for-hierarchical-model/5738/2 "2020-09-08T02:45:10Z")

</div>

Hi hope someone can chime in here. In the Radon example simply index each number with its category (county) works. However I have an array of 1000 numbers for each experiment. How do I index each of these 1000 numbers for their category?

Thank you in advance!

---

<div class="post-metadata">

**Author:** ![BioGoertz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/biogoertz/32/4489_2.png) [@BioGoertz](https://discourse.pymc.io/u/BioGoertz)\
**Post date:** [September 9, 2020, 10:08am UTC](https://discourse.pymc.io/t/index-for-hierarchical-model/5738/3 "2020-09-09T10:08:31Z")

</div>

You just need to reshape your `type_idx` into shape (3,1) instead of (3,):

```python
type_idx = np.array((0,0,1)).reshape(3,1)

```

That being said, in both your models you get a ton of divergences, potentially leading to biased estimates (even though it does capture your `tau_true` well). You might want to look at reformulating your model a bit. This is probably down to the fact that your noise term is infinitesmal, on the order of 10^-11, so your observations aren’t really Normal, though this could be a quirk of your toy data here.

---

<div class="post-metadata">

**Author:** ![newerjazz](https://avatars.discourse-cdn.com/v4/letter/n/a87d85/32.png) [@newerjazz](https://discourse.pymc.io/u/newerjazz)\
**Post date:** [September 10, 2020, 12:35am UTC](https://discourse.pymc.io/t/index-for-hierarchical-model/5738/4 "2020-09-10T00:35:57Z")

</div>

Thank you BioGoertz,

It works now! After adding simulated noise to Y\_true there are almost no more divergences. Thanks!

---

<div class="post-metadata">

**Author:** ![BioGoertz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/biogoertz/32/4489_2.png) [@BioGoertz](https://discourse.pymc.io/u/BioGoertz)\
**Post date:** [September 11, 2020, 11:25am UTC](https://discourse.pymc.io/t/index-for-hierarchical-model/5738/5 "2020-09-11T11:25:25Z")

</div>

Great!
