# Sampling a gaussian using ADVI

**URL:** <https://discourse.pymc.io/t/sampling-a-gaussian-using-advi/1100>\
**Category:** Questions\
**Created:** [April 19, 2018, 2:11am UTC](https://discourse.pymc.io/t/sampling-a-gaussian-using-advi/1100 "2018-04-19T02:11:13Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![rahuldave](https://avatars.discourse-cdn.com/v4/letter/r/6f9a4e/32.png) [@rahuldave](https://discourse.pymc.io/u/rahuldave)\
**Post date:** [April 19, 2018, 2:11am UTC](https://discourse.pymc.io/t/sampling-a-gaussian-using-advi/1100/1 "2018-04-19T02:11:13Z")

</div>

I’m trying to do a 'hello world" for the new ADVI interface.

This was my old code, which produced an almost exactly matching variational posterior

```python
data = np.random.randn(100)
with pm.Model() as model: 
    mu = pm.Normal('mu', mu=0, sd=1, testval=0)
    sd = pm.HalfNormal('sd', sd=1)
    n = pm.Normal('n', mu=mu, sd=sd, observed=data)
advifit = pm.variational.advi( model=model, n=100000)
means, sds, elbo = advifit

```

In the new way, i create a ADVI object

```python
advifit = pm.ADVI( model=model)
advifit.fit(n=10000)
advifit.approx.mean.eval(), advifit.approx.std.eval()

```

this gives me:  
`(array([-0.06538046, -0.03497334]), array([0.11616796, 0.08825936]))`

is the first array the means of the variational approximations for mu and sd for my model? Or is there something going on with the parametrization (sd has a negative mean). And in general, given the advifit object, what is the officially sanctioned way of getting samples from it? I looked at the quickstart but landed up getting more confused, and the API docs dont seem to go into the Approximation objects.

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [April 19, 2018, 5:57am UTC](https://discourse.pymc.io/t/sampling-a-gaussian-using-advi/1100/2 "2018-04-19T05:57:25Z")

</div>

> [@rahuldave](#):
>
> is the first array the means of the variational approximations for mu and sd for my model?

Yes, but they are the approximation of the free parameters in the model. PyMC3 automatically transform the bounded parameters to the real line. In this case, the `sd` is only positive as it is halfnormal distributed, but for sampling and VI PyMC3 operates on the unbounded version of it.  
You can check what are the parameters actually being sample/approximate by doing:

```python
model.free_RVs
Out[4]: [mu, sd_log__]

```

> [@rahuldave](#):
>
> And in general, given the advifit object, what is the officially sanctioned way of getting samples from it?

You can do `advifit.approx.sample(1000)` which gives you a MCMC trace of 1000 iteration just like a trace returned from sampling.

---

<div class="post-metadata">

**Author:** ![rahuldave](https://avatars.discourse-cdn.com/v4/letter/r/6f9a4e/32.png) [@rahuldave](https://discourse.pymc.io/u/rahuldave)\
**Post date:** [April 19, 2018, 2:31pm UTC](https://discourse.pymc.io/t/sampling-a-gaussian-using-advi/1100/3 "2018-04-19T14:31:47Z")

</div>

Thus we are actually doing a normal on log(sd)? That would make the negative mean then a mean of log(sd), correct? That would make sense. I am confused by the rho=log(1+exp(s) parameter…is that part of this transformation?

Thanks for the “official way to get the samples”!!

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [April 19, 2018, 2:47pm UTC](https://discourse.pymc.io/t/sampling-a-gaussian-using-advi/1100/4 "2018-04-19T14:47:55Z")

</div>

> [@rahuldave](#):
>
> I am confused by the rho=log(1+exp(s) parameter…is that part of this transformation?

You can find more information in the original paper [https://arxiv.org/pdf/1603.00788.pdf](https://arxiv.org/pdf/1603.00788.pdf) basically you want the approximation parameter also on the real line so that you wont have a problem of a too large learning rate will push the sd invalid
