# Adaptive Minibatch size

**URL:** <https://discourse.pymc.io/t/adaptive-minibatch-size/1935>\
**Category:** Questions\
**Created:** [September 18, 2018, 12:27pm UTC](https://discourse.pymc.io/t/adaptive-minibatch-size/1935 "2018-09-18T12:27:04Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![gokl](https://avatars.discourse-cdn.com/v4/letter/g/3da27b/32.png) [@gokl](https://discourse.pymc.io/u/gokl)\
**Post date:** [September 18, 2018, 12:27pm UTC](https://discourse.pymc.io/t/adaptive-minibatch-size/1935/1 "2018-09-18T12:27:04Z")

</div>

Hi,

I’m using ADVI with Minibatches and I observed two things.

One is that using small batch\_size values I get to a relatively good convergence quickly, whereas with large batch sizes it takes long.

The other thing is that the variance of the elbo plot using small batch sizes is relatively large compared to bigger batch sizes. This is actually creating convergence problems for me, as the parameters fluctuate too much so that CheckParametersConvergence does not detect convergence.

So my idea is to use an adaptive batch size for Minibatches which starts with a small batch size to get to a good estimate for the parameters quickly and increase the batch size to get stable estimates later on.  
Do you think this approach makes sense?

I tried to figure out how to change the batch\_size of a pm.Minibatch, and what I came up with makes me wonder if there is a better way of doing this

```
def change_minibatch_size(minibatch, size):
    minibatch.minibatch.owner.inputs[1].owner.inputs[0].owner.inputs[0] \
        .owner.inputs[0].owner.inputs[0].owner.inputs[1].data[0] = size

```

I got to this solution after looking at the graph:

`theano.printing.debugprint(minibatch)`

```
ViewOp [id A] 'Minibatch'   
 |AdvancedSubtensor1 [id B] ''   
   |<TensorType(int64, vector)> [id C]
   |Reshape{1} [id D] ''   
     |Elemwise{Cast{int64}} [id E] ''   
     | |Elemwise{add,no_inplace} [id F] ''   
     | |Elemwise{mul,no_inplace} [id G] ''   
     | | |mrg_uniform{TensorType(float64, vector),no_inplace}.1 [id H] ''   
     | | | |<TensorType(int32, matrix)> [id I]
     | | | |TensorConstant{(1,) of 10} [id J]
     | | |InplaceDimShuffle{x} [id K] ''   
     | | |Elemwise{sub,no_inplace} [id L] ''   
     | | |UndefinedGrad [id M] ''   
     | | | |Elemwise{sub,no_inplace} [id N] ''   
     | | | |Elemwise{Cast{float64}} [id O] ''   
     | | | | |Subtensor{int64} [id P] ''   
     | | | | |Shape [id Q] ''   
     | | | | | |<TensorType(int64, vector)> [id C]
     | | | | |Constant{0} [id R]
     | | | |TensorConstant{1e-16} [id S]
     | | |UndefinedGrad [id T] ''   
     | | |Elemwise{Cast{float64}} [id U] ''   
     | | |TensorConstant{0.0} [id V]
     | |InplaceDimShuffle{x} [id W] ''   
     | |UndefinedGrad [id T] ''   
     |MakeVector{dtype='int64'} [id X] ''   
       |Subtensor{int64} [id Y] ''   
         |Shape [id Z] ''   
         | |Elemwise{Cast{int64}} [id E] ''   
         |Constant{0} [id BA]
```

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [September 23, 2018, 6:29am UTC](https://discourse.pymc.io/t/adaptive-minibatch-size/1935/2 "2018-09-23T06:29:34Z")

</div>

I am wondering if you can set the minibatch size as a theano.shared variable, and decay it during training, similar to the learning rate decay in [https://docs.pymc.io/notebooks/lda-advi-aevb.html?highlight=lda#AEVB-with-ADVI](https://docs.pymc.io/notebooks/lda-advi-aevb.html?highlight=lda#AEVB-with-ADVI)

As for the correctness of doing so, I really have no intuition. Sound plausible so I will be very interested to heard back about your result.
