# Calling \`pm.set\_data\` multiple times

**URL:** <https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306>\
**Category:** v5\
**Created:** [February 1, 2023, 10:52pm UTC](https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306 "2023-02-01T22:52:41Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![schwarls37](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/schwarls37/32/4823_2.png) [@schwarls37](https://discourse.pymc.io/u/schwarls37)\
**Post date:** [February 1, 2023, 10:52pm UTC](https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306/1 "2023-02-01T22:52:41Z")

</div>

In the [predicting on hold out data](https://www.pymc.io/projects/examples/en/latest/howto/api_quickstart.html#predicting-on-hold-out-data) section (4.1) in the quickstart, how come I can’t change the data again, like this (if I’ve already changed the data once as in the tutorial with `"x_obs": [-1, 0, 1.0]` and `coords={"idx": [1001, 1002, 1003]}`)?

```auto
with model:
    # change the value and shape of the data
    pm.set_data(
        {
            "x_obs": [.0, 0, .0],
            # use dummy values with the same shape:
            "y_obs": [.0, 0, .0]
        },
        coords={"idx": [1004, 1005, 1006]},
    )
    idata.extend(pm.sample_posterior_predictive(idata))
idata.posterior_predictive["obs"].mean(dim=["draw", "chain"])

```

It just returns the same thing as the original code and hasn’t updated the out of sample data at all.

---

<div class="post-metadata">

**Author:** ![cluhmann](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/cluhmann/32/3083_2.png) [@cluhmann](https://discourse.pymc.io/u/cluhmann)\
**Post date:** [February 2, 2023, 12:07am UTC](https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306/2 "2023-02-02T00:07:39Z")

</div>

Are you creating the data via `pm.MutableData()`?

---

<div class="post-metadata">

**Author:** ![schwarls37](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/schwarls37/32/4823_2.png) [@schwarls37](https://discourse.pymc.io/u/schwarls37)\
**Post date:** [February 2, 2023, 4:55pm UTC](https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306/3 "2023-02-02T16:55:26Z")

</div>

Yes, I believe so. When the following code is run the second attempt to adjust the mutable data has no effect.

```auto
x = rng.standard_normal(100)
y = x > 0

coords = {"idx": np.arange(100)}
with pm.Model() as model:
    # create shared variables that can be changed later on
    x_obs = pm.MutableData("x_obs", x, dims="idx")
    y_obs = pm.MutableData("y_obs", y, dims="idx")

    coeff = pm.Normal("x", mu=0, sigma=1)
    logistic = pm.math.sigmoid(coeff * x_obs)
    pm.Bernoulli("obs", p=logistic, observed=y_obs, dims="idx")
    idata = pm.sample()

with model:
    # change the value and shape of the data
    pm.set_data(
        {
            "x_obs": [-1, 0, 1.0],
            # use dummy values with the same shape:
            "y_obs": [0, 0, 0],
        },
        coords={"idx": [1001, 1002, 1003]},
    )

    idata.extend(pm.sample_posterior_predictive(idata))

print(idata.posterior_predictive["obs"].mean(dim=["draw", "chain"]))

with model:
    # change the value and shape of the data
    pm.set_data(
        {
            "x_obs": [.0, 0, .0],
            # use dummy values with the same shape:
            "y_obs": [.0, 0, .0]
        },
        coords={"idx": [1004, 1005, 1006]},
    )
    idata.extend(pm.sample_posterior_predictive(idata))
idata.posterior_predictive["obs"].mean(dim=["draw", "chain"])

```

---

<div class="post-metadata">

**Author:** ![cluhmann](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/cluhmann/32/3083_2.png) [@cluhmann](https://discourse.pymc.io/u/cluhmann)\
**Post date:** [February 2, 2023, 5:20pm UTC](https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306/4 "2023-02-02T17:20:08Z")

</div>

I’m not sure what happens when you try to extend an inferenceData objects with multiple groups, each of which shares a name (e.g., `posterior_predictive`).

This seems to work for me?

```python
import pymc as pm
import numpy as np

rng = np.random.default_rng(12345)

x = rng.standard_normal(100)
y = x > 0

coords = {"idx": np.arange(100)}
with pm.Model() as model:
    # create shared variables that can be changed later on
    x_obs = pm.MutableData("x_obs", x, dims="idx")
    y_obs = pm.MutableData("y_obs", y, dims="idx")

    coeff = pm.Normal("x", mu=0, sigma=1)
    logistic = pm.math.sigmoid(coeff * x_obs)
    pm.Bernoulli("obs", p=logistic, observed=y_obs, dims="idx")
    idata = pm.sample()

with model:
    # change the value and shape of the data
    pm.set_data(
        {
            "x_obs": [-1, 0, 1.0],
            # use dummy values with the same shape:
            "y_obs": [0, 0, 0],
        },
        coords={"idx": [1001, 1002, 1003]},
    )

    pp1 = pm.sample_posterior_predictive(idata)

with model:
    # change the value and shape of the data
    pm.set_data(
        {
            "x_obs": [.0, 0, .0],
            # use dummy values with the same shape:
            "y_obs": [.0, 0, .0]
        },
        coords={"idx": [1004, 1005, 1006]},
    )
    pp2 = pm.sample_posterior_predictive(idata)

print(pp1.posterior_predictive["obs"].mean(dim=["draw", "chain"]))

#<xarray.DataArray 'obs' (idx: 3)>
#array([0.027 , 0.5025, 0.977])
#Coordinates:
# * idx (idx) int64 1001 1002 1003

print(pp2.posterior_predictive["obs"].mean(dim=["draw", "chain"]))

#<xarray.DataArray 'obs' (idx: 3)>
#array([0.51625, 0.5025 , 0.501])
#Coordinates:
# * idx (idx) int64 1004 1005 1006

```

---

<div class="post-metadata">

**Author:** ![OriolAbril](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/oriolabril/32/2497_2.png) [@OriolAbril](https://discourse.pymc.io/u/OriolAbril)\
**Post date:** [February 2, 2023, 7:24pm UTC](https://discourse.pymc.io/t/calling-pm-set-data-multiple-times/11306/5 "2023-02-02T19:24:51Z")

</div>

The issue is (was) with the usage of extend, it is not a merge or concatenation function.

As explained in [its docstring](https://python.arviz.org/en/stable/api/generated/arviz.InferenceData.extend.html), extend has by default `join="left"` which means that groups that are both present in `idata` and in `other` are kept from `idata` without modifying it, `join="right"` would replace the repeated groups in `idata` in order to keep the ones from `other`.
