# Modelling multiple correlated variables

**URL:** <https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470>\
**Category:** Questions\
**Created:** [December 17, 2021, 11:21am UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470 "2021-12-17T11:21:17Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![fredzett](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/fredzett/32/4665_2.png) [@fredzett](https://discourse.pymc.io/u/fredzett)\
**Post date:** [December 17, 2021, 11:21am UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/1 "2021-12-17T11:21:17Z")

</div>

I am completley new to pymc3 and really like the api and the flexibility that comes along.

However, I am stuck at modelling something a cash flow model. For sake of simplicity let’s say I want to model the following:

y = \sum\_{t=0}^T \dfrac{CF\_t}{(1+r)^t}

where all CF and r are normally distributed and CFs are correlated.

How would I best model this assuming the following priors

1. all CFs follow a normal distribution with different mus and stds
2. all CFs are correlated
3. r follows a normal distribution

My main challenges are:

1. how do I best create T parameters with different mus and std. I know that I can pass a shape parameter, but I cannot pass different mus and stds? Is the only way to create individual parameters?
2. how do I model that y is the sum of different parameters?
3. how do I model the correlation part? (this is likely a prob programming question rather than a pymc3 question

My below attempt works but has two pitfalls:

1. if T = 100 I had to model 100 separate variables
2. it does not account for correlations between cf1, cf2 and cf3

```auto
T = 3
model = pm.Model()
with model:
    # Define priors
    cf1 = pm.Normal("cf1",100,10)
    cf2 = pm.Normal("cf2",100,15)
    cf3 = pm.Normal("cf3",100,20)

    i = pm.Normal("i", 0.1,0.04)
    t = range(0,T)
    
    # Model output and treat as deterministic to keep track
    y = pm.Deterministic("y",cf1/(1+i)**t[0] + cf2/(1+i)**t[1] + cf3/(1+i)**t[2])
    
    trace = pm.sample()

```

Can anyone point me to the right direction?

Thanks for your help!

(Note that this is a very simplified version of what I really want to achieve)

---

<div class="post-metadata">

**Author:** ![mjedrz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/mjedrz/32/4433_2.png) [@mjedrz](https://discourse.pymc.io/u/mjedrz)\
**Post date:** [December 17, 2021, 12:48pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/2 "2021-12-17T12:48:55Z")

</div>

I think I can answer 3: if you want to include correlations in your prior, then you need to create a covariance matrix, which will contain the information about correlations between variables (covariance is a measure of how the variables vary together). Then you need to load that covariance matrix into a multivariate distribution, which in your case will be:  
`cfs = pm.MvNormal("cfs", mu=[100, 100, 100], cov=your_covariance_matrix, shape=3)`

---

<div class="post-metadata">

**Author:** ![fredzett](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/fredzett/32/4665_2.png) [@fredzett](https://discourse.pymc.io/u/fredzett)\
**Post date:** [December 17, 2021, 1:03pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/3 "2021-12-17T13:03:49Z")

</div>

Great thanks for the swift help. I tried this multiple times and it didn’t work given I forgot to set the shape parameter…THANKS!

---

<div class="post-metadata">

**Author:** ![Dirk\_Nachbar1](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/dirk_nachbar1/32/4668_2.png) [@Dirk\_Nachbar1](https://discourse.pymc.io/u/Dirk_Nachbar1)\
**Post date:** [December 21, 2021, 5:12pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/4 "2021-12-21T17:12:03Z")

</div>

Random remark, are you sure you want to model i/r as normal, and allow it to be negative? Maybe you want a half-normal.

---

<div class="post-metadata">

**Author:** ![fredzett](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/fredzett/32/4665_2.png) [@fredzett](https://discourse.pymc.io/u/fredzett)\
**Post date:** [December 22, 2021, 1:46pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/5 "2021-12-22T13:46:22Z")

</div>

Yes, you are right. I will not model it as normal. This was just a toy example and the normal prior does not make sense here. Should have used a different prior to avoid causing confusion.

---

<div class="post-metadata">

**Author:** ![Kenneth\_Ottosen](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/kenneth_ottosen/32/7337_2.png) [@Kenneth\_Ottosen](https://discourse.pymc.io/u/Kenneth_Ottosen)\
**Post date:** [October 31, 2023, 1:04pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/6 "2023-10-31T13:04:45Z")

</div>

I have a follow-up PYMC3 question. What would I have to do in case I want to set the correlation between variables when variables are not all normal but a mixture of different kinds of distributions? In that case, it would not be appropriate to use a multivariate normal distribution (pm.MvNormal), as it only generates normal distributions. I am wondering if there is something like a multivariate distribution that can output different kind of distributions? I would like to do something similar to what is possible with Crystal Ball where it is possible to correlate different input variables having different distributions → [https://www.oracle.com/docs/tech/middleware/correlated-assumptions.pdf](https://www.oracle.com/docs/tech/middleware/correlated-assumptions.pdf)

---

<div class="post-metadata">

**Author:** ![jessegrabowski](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jessegrabowski/32/5010_2.png) [@jessegrabowski](https://discourse.pymc.io/u/jessegrabowski)\
**Post date:** [October 31, 2023, 2:14pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/7 "2023-10-31T14:14:22Z")

</div>

I think the easiest option is to move the correlation out of the output distributions and into the structure of the model. For example, [this paper](https://www.tandfonline.com/doi/pdf/10.1080/02664760802684177?casa_token=AMsSvWVooaUAAAAA:Zbm5G9WxQ2Z4ZlvyPePHPBBnOHamYdCjb7lH28Y5baLuUzWQj4inJp0KZ-JY0a78pBvCOQdl-hKNoyI) (implemented in PyMC [here](https://www.pymc.io/projects/examples/en/latest/case_studies/rugby_analytics.html)) uses latent structure to model **conditionally** independent Poisson variables, which end up being correlated due to the hierarchical model structure .

I think the other alternative would be to work with copula models. @jonsedar has been doing a lot of work on these lately, I think he might be able to chime in with some general remarks on the subject?

---

<div class="post-metadata">

**Author:** ![Kenneth\_Ottosen](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/kenneth_ottosen/32/7337_2.png) [@Kenneth\_Ottosen](https://discourse.pymc.io/u/Kenneth_Ottosen)\
**Post date:** [November 1, 2023, 3:22pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/8 "2023-11-01T15:22:04Z")

</div>

@jessegrabowski - Thank you very much for the links, were really helpful 👍

---

<div class="post-metadata">

**Author:** ![jonsedar](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jonsedar/32/2590_2.png) [@jonsedar](https://discourse.pymc.io/u/jonsedar)\
**Post date:** [November 6, 2023, 8:05pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/9 "2023-11-06T20:05:34Z")

</div>

You rang? Yes I’ve been doing a few things with copulas on marginals recently - lots of bashing head against the wall, but I could try to help if you want

---

<div class="post-metadata">

**Author:** ![jessegrabowski](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jessegrabowski/32/5010_2.png) [@jessegrabowski](https://discourse.pymc.io/u/jessegrabowski)\
**Post date:** [November 6, 2023, 8:08pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/10 "2023-11-06T20:08:28Z")

</div>

Was just curious if you had any comments on the usefulness of directly modeling correlation between RVs by using copulas vs injecting conditional correlation via model structure.

Disclaimer: my knowledge of copulas begins and ends with how to spell the word.

---

<div class="post-metadata">

**Author:** ![jonsedar](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jonsedar/32/2590_2.png) [@jonsedar](https://discourse.pymc.io/u/jonsedar)\
**Post date:** [November 6, 2023, 8:26pm UTC](https://discourse.pymc.io/t/modelling-multiple-correlated-variables/8470/11 "2023-11-06T20:26:04Z")

</div>

I’m probably so far down the copula rabbit hole that (potentially better) alternatives are a distant hope… sunk costs and all that 😃

I think copulas are quite nice in the case that you want to allow the marginals to correlate in a pooled fashion regardless of the sub-models on each marginal - my intention is that this gains stability and flexibility in the design of the sub-models. Also to more easily use non-Gaussian copulas.

The alternative would be to require several features to form the sub-models of both marginals, and to correlate the sub-model coefficients, I think this would be a good example: [McElreath, 2014](http://xcelab.net/rmpubs/Mcelreath%20Koster%202014.pdf) where he uses an MvN to correlate hierarchical hyperparams. I’m not sure if/how one could achieve this with a non-Gaussian copula…
