# Constructing the random effects model matrix Z

**URL:** <https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066>\
**Category:** Questions\
**Created:** [April 8, 2018, 7:31pm UTC](https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066 "2018-04-08T19:31:27Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jack\_Caster](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jack_caster/32/644_2.png) [@Jack\_Caster](https://discourse.pymc.io/u/Jack_Caster)\
**Post date:** [April 8, 2018, 7:31pm UTC](https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066/1 "2018-04-08T19:31:27Z")

</div>

I am trying to replicate the sleep deprivation study found in the R package LME4 (for fitting linear mixed effects models) and described in this [paper](https://arxiv.org/pdf/1406.5823.pdf). The dataset `sleepstudy` contains

> the average reaction time per day for subjects in a sleep deprivation study. On day 0 the subjects had their normal amount of sleep. Starting that night they were restricted to 3 hours of sleep per night. The observations represent the average reaction time on a series of tests given each day to each subject.

What I want to to is to fit a random effects model on intercept and slope for each subject. And I would like to use the matrix parametrization. That is, `X` is the design matrix for the fixed effects given by the formula `Reaction ~ 1 + Days`, and `Z` is the design matrix for the random effects. However, I have some doubts on how to design such matrix `Z` to get intercept’s and slope’s coefficients.

What I did is to simply create a design matrix from the formula `Reaction ~ 0 + Subject + Subject:Days`. Is this the right way? On one hand the results I get seems reasonable (see [notebook](https://github.com/JackCaster/GLM_with_PyMC3/blob/master/notebooks/LME%20-%20Sleep%20study.ipynb)). On the other hand, in the paper describing LME4 ([link](https://arxiv.org/pdf/1406.5823.pdf)) at page 9 they use a more complex method based on Khatri-Rao and Kronecker product. Am I missing something?

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [April 8, 2018, 9:20pm UTC](https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066/2 "2018-04-08T21:20:26Z")

</div>

Nice notebook. What you are doing is correct. Sometimes (e.g. in `brms` which based on Stan), ppl use an indexing method for the random effects. Khatri-Rao with Kronecker product is more for when you have a random intercept _and_ slope model, but you can construct by hand similarly - the aim is to create a sparse matrix Z where there are blocks of values on the diagonal.

~~In case you have not seen it, I have this repository doing mixed effect model with varies libraries and parameterization: [https://github.com/junpenglao/GLMM-in-Python](https://github.com/junpenglao/GLMM-in-Python)~~ [EDIT] Just realized I already reply to you about this repo on a similar topic before 😅

---

<div class="post-metadata">

**Author:** ![Jack\_Caster](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jack_caster/32/644_2.png) [@Jack\_Caster](https://discourse.pymc.io/u/Jack_Caster)\
**Post date:** [April 9, 2018, 8:48am UTC](https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066/3 "2018-04-09T08:48:28Z")

</div>

Just to get it straight. As I want to have a random effect on intercept **and** slope, either I construct the matrix `Z` as `Reaction ~ 0 + Subject + Subject:Days` (as I did), or I construct the same matrix with the fancy Khatri-Rao method. The two approaches should be equivalent right? (I will verify this myself later when I am back from work).

> [@junpenglao](#):
>
> [EDIT] Just realized I already reply to you about this repo on a similar topic before 😅

No problem 😉 But it seems you do not have an example with random effects on slopes, or do you?

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [April 9, 2018, 9:14am UTC](https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066/4 "2018-04-09T09:14:23Z")

</div>

> [@Jack\_Caster](#):
>
> But it seems you do not have an example with random effects on slopes, or do you?

No I do not.

What you are doing in your notebook is correct. Also, I agree with your approach separating the random intercept and slope instead of multiplying a big matrix - you have much more controls on the prior.

---

<div class="post-metadata">

**Author:** ![Jack\_Caster](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jack_caster/32/644_2.png) [@Jack\_Caster](https://discourse.pymc.io/u/Jack_Caster)\
**Post date:** [April 9, 2018, 3:41pm UTC](https://discourse.pymc.io/t/constructing-the-random-effects-model-matrix-z/1066/5 "2018-04-09T15:41:11Z")

</div>

> [@junpenglao](#):
>
> you have much more controls on the prior

Speaking on which. I made a small mistake and I set the same `sigma_Z` for both intercept and slope of the random effect. Now, instead, I declared `sigma_Z_intercept` and `sigma_Z_slope`. The predictions are way better. Before the slope was not allow to vary, hence the prediction for subject 309 (for example) was off. I updated the notebook. I should clean it up and post it somewhere (sometime 🙂 )

 ![download](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/1X/cb8b1bac91ebd79676c414760e7f76b5a2f41ff0.png)
