# Out of sample predict issue

**URL:** <https://discourse.pymc.io/t/out-of-sample-predict-issue/12344>\
**Category:** General\
**Created:** [June 20, 2023, 4:55am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344 "2023-06-20T04:55:55Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Number\_Huang](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/number_huang/32/5877_2.png) [@Number\_Huang](https://discourse.pymc.io/u/Number_Huang)\
**Post date:** [June 20, 2023, 4:55am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/1 "2023-06-20T04:55:55Z")

</div>

the following code is from the bart-bikling example,I changed it to a classifier.but it reports the shape mismatch issue,seems that the set\_data(X\_test) not work.

```auto
from pathlib import Path
import arviz as az
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import pymc as pm
import pymc_bart as pmb
from sklearn.model_selection import train_test_split
bikes = pd.read_csv(pm.get_data("bikes.csv"))
features = ["hour", "temperature", "humidity", "workingday"]
X = bikes[features]
Y = bikes["count"]
Y2 = Y.apply(lambda x:1 if x>180 else 0)
RANDOM_SEED=100
X_train, X_test, Y_train, Y_test = train_test_split(X, Y2, test_size=0.2, random_state=RANDOM_SEED)
with pm.Model() as model_oos_regression:
    X1 = pm.MutableData("X", X_train.values)
    Y1 = Y_train.values.flatten()
    #α = pm.Exponential("α", 1)
    μ = pmb.BART("μ", X1, Y1)
    #y = pm.NegativeBinomial("y", mu=pm.math.exp(μ), alpha=α, observed=Y, shape=μ.shape)
    #y = pm.Deterministic("y", pm.invlogit(μ))
    pm.Bernoulli("y",observed=Y1,p=pm.Deterministic("p1", pm.invlogit(μ)))
    idata = pm.sample(random_seed=RANDOM_SEED)
    #idata_oos_regression = pm.fit(method=pm.ADVI()).sample()

    #predict out sample
    pm.set_data({"X":X_test.values})
    # posterior_predictive_oos_regression_test = pm.sample_posterior_predictive(
    # trace=idata_oos_regression, random_seed=RANDOM_SEED,
    # var_names=['y'],
    # return_inferencedata=True,
    # predictions=True
    # )
    idata.extend(pm.sample_posterior_predictive(idata))
    #pred = posterior_predictive_oos_regression_test.predictions
    yHat = idata.posterior_predictive['y'].mean(("chain", "draw")).to_numpy()
    print(f"yHat-len={len(yHat)},X_test-len={len(X_test)}")
    assert len(yHat)==len(X_test)

```

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [June 20, 2023, 5:49am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/2 "2023-06-20T05:49:31Z")

</div>

You have to specify how the shape of y depends on its parameters. It’s illustrated in the examples here: [pymc.set\_data — PyMC 5.5.0 documentation](https://www.pymc.io/projects/docs/en/stable/api/generated/pymc.set_data.html)

Otherwise you need to provide dummy values for y with the correct shape

---

<div class="post-metadata">

**Author:** ![Number\_Huang](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/number_huang/32/5877_2.png) [@Number\_Huang](https://discourse.pymc.io/u/Number_Huang)\
**Post date:** [June 20, 2023, 6:11am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/3 "2023-06-20T06:11:41Z")

</div>

thank u Richard, not clear still.  
pm.Bernoulli(“y”,observed=Y1,p=pm.Deterministic(“p1”, pm.invlogit(μ)), **shape=Y1.shape** )  
that is what I changed,failed still.  
in the document of set\_data,the x and y has the same shape,but it is not my case.

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [June 20, 2023, 6:42am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/4 "2023-06-20T06:42:47Z")

</div>

The shape should depend on `μ` somehow? Otherwise `shape=Y1.shape` is the default anyway. If you have no other source of shape information other than the observations, you will need to use dummy variables when doing posterior predictive, to force the right shape

---

<div class="post-metadata">

**Author:** ![Number\_Huang](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/number_huang/32/5877_2.png) [@Number\_Huang](https://discourse.pymc.io/u/Number_Huang)\
**Post date:** [June 20, 2023, 7:05am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/5 "2023-06-20T07:05:22Z")

</div>

yes,μ.shape works!! thank u ricardo.  
Another question is  
pm.Bernoulli(“y”,observed=Y1,p=pm.Deterministic(“p1”, pm.invlogit(μ)), **shape=μ.shape** )  
for prediction, which variable shall I predict? “y” or “p1”?  
my understanding is p1 is invlogit(μ), that is the sigmoid value,which should be a classifier’s predict\_proba, but what’s the ‘y’?

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [June 20, 2023, 7:07am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/6 "2023-06-20T07:07:25Z")

</div>

`y` is a bernoulli draw from `p1`. You can predict whichever is more useful for you (or both)

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [June 20, 2023, 7:08am UTC](https://discourse.pymc.io/t/out-of-sample-predict-issue/12344/7 "2023-06-20T07:08:25Z")

</div>

Linking to the new entry in the FAQ for future readers: [Frequently Asked Questions - #18 by ricardoV94](https://discourse.pymc.io/t/frequently-asked-questions/74/18)
