# Posterior predictive on testing data in State Space Models

**URL:** <https://discourse.pymc.io/t/posterior-predictive-on-testing-data-in-state-space-models/12103>\
**Category:** v5\
**Tags:** modeling\
**Created:** [May 10, 2023, 11:04pm UTC](https://discourse.pymc.io/t/posterior-predictive-on-testing-data-in-state-space-models/12103 "2023-05-10T23:04:05Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![dushyant](https://avatars.discourse-cdn.com/v4/letter/d/e480ec/32.png) [@dushyant](https://discourse.pymc.io/u/dushyant)\
**Post date:** [May 10, 2023, 11:04pm UTC](https://discourse.pymc.io/t/posterior-predictive-on-testing-data-in-state-space-models/12103/1 "2023-05-10T23:04:05Z")

</div>

I have been working on modeling the below State Space Model

X\_{i,t+1} \sim N(\mu\_t, Q)\\ \mu\_{i,t} = Ax\_{i,t} + BU\_{i}\\ y\_{i,t} \sim Poisson(\lambda\_{i,t}) \\\lambda\_{i,t} = C \exp(X\_{i,t})

where X\_{i,t+1} is the hidden state for ith sample at time t, U\_i is known exogenous variables, \lambda\_{i,t} is the prediction for ith sample at time t. Q, A, B and C are parameters of the model. Below is the code for the model

```auto
with pm.Model() as mod:

    sd_dist = pm.Exponential.dist(1.0, shape=latent_variables)
    mu1 = np.zeros(latent_variables)
    
    chol1, corr, stds = pm.LKJCholeskyCov('chol_cov1', n=latent_variables, eta=2,
    sd_dist=sd_dist, compute_corr=True)
    pred_Z = pm.MvNormal('pred_Z', mu=mu1, chol=chol1, shape=(latent_variables))
    
    mu2 = np.zeros(num_features)
    chol2, corr, stds = pm.LKJCholeskyCov('chol_cov2', n=num_features, eta=2,
    sd_dist=sd_dist, compute_corr=True)
    pred_B = pm.MvNormal('B', mu=mu2, chol=chol2, shape=(num_features))
    
    A_pt = pt.as_tensor_variable(A_block)

    sigmas_Q = pm.HalfNormal('sigmas_Q', sigma=1, shape= (latent_variables))
    Q = pt.diag(sigmas_Q)
    u = pm.MutableData("input", u_scale)
    #x0 = pm.MutableData("x_init", x_0)
    data = pm.MutableData("data", data_train)
    samples = u.shape[0]
    x0 = pt.zeros((num_samples , latent_variables))
    def step(x, A, Q, samples):
        innov = pm.MvNormal.dist(mu=0, tau=Q, size = (samples))
        
        next_x = pt.nlinalg.matrix_dot(x,A) + innov
        
        return next_x, collect_default_updates([x, A, Q, samples], [next_x])
    
    hidden_states_pt, updates = pytensor.scan(step, 
                                              outputs_info=[x0], 
                                              non_sequences=[A_pt, Q, samples],
                                              n_steps=T, 
                                              strict=True)
    
    mod.register_rv(hidden_states_pt, name='hidden_states', initval=pt.zeros((T, samples, latent_variables)))
    
    
    #pred_Z = pm.HalfNormal('pred_Z', sigma=1, size= latent_variables)
    temp1 = pt.transpose(pt.nlinalg.matrix_dot(u,pred_B))
    
    #hidden_states_pt = pt.reshape(hidden_states_pt,(T,num_samples,latent_variables))
    #lambdas = pt.exp(hidden_states_pt[:, np.arange(0, num_samples * latent_variables, 2)]+temp1)
    lambdas = pt.exp(pt.squeeze(pt.nlinalg.matrix_dot(hidden_states_pt,pred_Z))+temp1)

    obs = pm.Poisson('obs', lambdas, observed=data)

```

I estimate the model parameters using the training data \lambda and U. Now I want to do some prediction on test data U\_{test}. I was thinking of using posterior predictive samples but that won’t work because the hidden states X need to be estimated again for the new sample. What will be the best way to get predictions for the new samples and also estimate the hidden states?

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [May 11, 2023, 4:48am UTC](https://discourse.pymc.io/t/posterior-predictive-on-testing-data-in-state-space-models/12103/2 "2023-05-11T04:48:38Z")

</div>

Do these Utest happen after the U in your original model? In that case you can create a new model which starts where U “left off” and call posterior predictive there.

I have an example in [this WIP notebook](https://colab.research.google.com/drive/11Ik3PgOzfemI5uEz-E7d__jZvd264lAZ#scrollTo=WieZcy4nLNTy) with a single timeseries, but the same concept applies. You can sample new unobserved variables in posterior predictive.

---

<div class="post-metadata">

**Author:** ![dushyant](https://avatars.discourse-cdn.com/v4/letter/d/e480ec/32.png) [@dushyant](https://discourse.pymc.io/u/dushyant)\
**Post date:** [May 12, 2023, 12:06am UTC](https://discourse.pymc.io/t/posterior-predictive-on-testing-data-in-state-space-models/12103/3 "2023-05-12T00:06:38Z")

</div>

Thanks! This is really helpful!
