# Separation between model and data/parameters

**URL:** <https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835>\
**Category:** General\
**Created:** [July 24, 2024, 3:13pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835 "2024-07-24T15:13:53Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![bob-carpenter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bob-carpenter/32/8124_2.png) [@bob-carpenter](https://discourse.pymc.io/u/bob-carpenter)\
**Post date:** [July 23, 2024, 3:19pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/1 "2024-07-23T15:19:47Z")

</div>

> [@How to generate an increasing sequence?](https://discourse.pymc.io/t/how-to-generate-an-increasing-sequence/14794/6):
>
> the following is perfectly valid and works out to similar (but not identical!

If you generate `x ~ normal(mu, sigma)` and sort, you get the same distribution as if you have an ordering constraint and apply the same distribution to the ordered variables.

> [@How to generate an increasing sequence?](https://discourse.pymc.io/t/how-to-generate-an-increasing-sequence/14794/6):
>
> In PyMC you are declaring a generative _forward_ model

We think the same way in Stan. It’s just a matter of how you’re defining the forward model. In Stan, it looks like this:

```stan
parameters {
  ordered[K] y;
}
model {
  y ~ normal(mu, sigma);
}

```

It’d be the same with the ordered transformation that @jessgrabowski mentioned.

> [@How to generate an increasing sequence?](https://discourse.pymc.io/t/how-to-generate-an-increasing-sequence/14794/6):
>
> The disadvantage of doing this is that you lose access to prior and posterior predictive sampling, because the forward model is not aware of the transformations

This can be a problem with constraints. The sorting trick is actually how to do Monte Carlo sampling most easily in this setting. You see the same thing in Stan with something like sum-to-zero constraints, which we don’t have built in. Or an ICAR model on spatial effects. It’s not easy to generate a random sum-to-zero or ICAR variate with Monte Carlo, so we have to resort to MCMC for prior and posterior checks (but then we can’t interleave discrete sampling).

P.S. Hope it’s OK to jump on your forums. I’m starting to work with normalizing flows in JAX and need a way to generate JAX models for testing, so I’m checking out PyMC and Pyro.

P.P.S. The original design for Stan assumed a graphical modeling framework like BUGS, but then I realized it was easy to add imperative features to make coding within Stan easier. We’ve come to realize that we gave up a lot of automation through that decision. I really like how data/parameter neutral and graphically-oriented Turing.jl is, but I don’t use Julia.

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [July 23, 2024, 8:54pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/2 "2024-07-23T20:54:22Z")

</div>

> [@bob-carpenter](#):
>
> I really like how data/parameter neutral and graphically-oriented Turing.jl is, but I don’t use Julia

What do you mean? Something about model being defined separately form the observations/ constraining transformations? Or something else?

> [@bob-carpenter](#):
>
> .S. Hope it’s OK to jump on your forums

Haha definitely OK

---

<div class="post-metadata">

**Author:** ![bob-carpenter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bob-carpenter/32/8124_2.png) [@bob-carpenter](https://discourse.pymc.io/u/bob-carpenter)\
**Post date:** [July 24, 2024, 3:13pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/3 "2024-07-24T15:13:53Z")

</div>

> [@bob-carpenter](#):
>
> I really like how data/parameter neutral and graphically-oriented Turing.jl is

> [@How to generate an increasing sequence?](https://discourse.pymc.io/t/how-to-generate-an-increasing-sequence/14794/11):
>
> What do you mean? Something about model being defined separately form the observations …

Exactly. It’s like BUGS that way. Whereas in PyMC, you have things like `Normal("y", mu=intercept + slope * x, sigma=sigma, observed=y)` where `observed=y` is part of the model code for observations whereas something like `y = Normal(...)` would be used for parameters.

I just did a search through your examples and see you have an example

```python
data = pm.MutableData("data", observed_data[0])
...
pm.Normal("y", mu=mu, sigma=1, observed=data)

```

which gives you some of what I was looking for in terms of allowing you to plug-and-play different data in the same model (I think). But it still builds `observed=data` in. So what I was wondering if there’s a way you could write a single PyMC model that you could use for data simulation to do something like simulation based calibration.

The other thing BUGS does which is super-convenient is allow partially instantiated variables (R directly supports this, too, which is costly in compute but useful). So you can do posterior prediction in a regression of `y` on `x` by having a partially known `y` and fully known `x`. In Stan this all has to get declared up front. We precompile the C++ class for the model with all of the data, parameters and generated quantities declared as such, but we don’t construct an instance until we have data. New data means an entirely new model instance because our models are immutable once constructed.

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [July 24, 2024, 3:59pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/4 "2024-07-24T15:59:17Z")

</div>

@bob-carpenter I moved the thread to its own topic.

You can use `pm.observe` to define observed nodes in pre-existing model thse days. You can also do a lot of model mutation with `pm.do`, including setting parameters to specific values, taking prior predictive draws and then conditioning on those for calibration workflows.

Our README shows an example of such workflow: [pymc/README.rst at main · pymc-devs/pymc · GitHub](https://github.com/pymc-devs/pymc/blob/main/README.rst#linear-regression-example)

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [July 24, 2024, 4:38pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/5 "2024-07-24T16:38:14Z")

</div>

Colab that was used as the basis for that readme section: [Google Colab](https://colab.research.google.com/drive/1JyIeu2NMl_z6Y8ZniZn3jDrzzUL3i81q?usp=sharing)

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [July 24, 2024, 4:41pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/6 "2024-07-24T16:41:57Z")

</div>

> [@bob-carpenter](#):
>
> The other thing BUGS does which is super-convenient is allow partially instantiated variables (R directly supports this, too, which is costly in compute but useful). So you can do posterior prediction in a regression of `y` on `x` by having a partially known `y` and fully known `x`.

I didn’t follow this part. Could you give it another try?

---

<div class="post-metadata">

**Author:** ![bob-carpenter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bob-carpenter/32/8124_2.png) [@bob-carpenter](https://discourse.pymc.io/u/bob-carpenter)\
**Post date:** [July 24, 2024, 5:42pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/7 "2024-07-24T17:42:52Z")

</div>

BUGS lets you have missing data, for example,

```R
y = c(3.2, 1.9, NA, NA)
x = c(1.5, 2.8, 3.1, -1.3)

```

A BUGS model like

```auto
y ~ dnorm(alpha + beta * x, sigma)
alpha ~ dnorm(0, 1) 
beta ~ dnorm(0, 1);
sigma ~ dlnorm(0, 1)

```

and use the above `x, y` as data, then you get posterior predictive inference for `y[3]` and `y[4]`. If you write a probabilistic model for `x`, then you can impute missing covariates. It’s super convenient.

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [July 24, 2024, 5:51pm UTC](https://discourse.pymc.io/t/separation-between-model-and-data-parameters/14835/8 "2024-07-24T17:51:24Z")

</div>

Ahh, I see. PyMC does exactly the same when it finds nan in the observation (but not yet when observations are set via the pm.observe command).
