# Exploring Streaming / Minibatch Inference in PyMC

**URL:** <https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646>\
**Category:** v5\
**Created:** [March 20, 2026, 11:44am UTC](https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646 "2026-03-20T11:44:52Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jhanani](https://avatars.discourse-cdn.com/v4/letter/j/f17d59/32.png) [@jhanani](https://discourse.pymc.io/u/jhanani)\
**Post date:** [March 20, 2026, 11:44am UTC](https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646/1 "2026-03-20T11:44:52Z")

</div>

Hello! I am Jhanani. I’ve been recently following the Streaming/Online Inference for GSOC 2026.

Recently, I have been working with PyMC and implemented a small prototype to understand streaming and minibatch-style inference.

This demo illustrates how posterior distributions for `theta_A` and `theta_B` evolve as new data chunks arrive, highlighting the concept of streaming or minibatch inference. We observe that the posterior means gradually converge to the true parameter values, demonstrating that knowledge accumulates effectively over successive chunks. While this example manually updates priors using previous posterior counts, true PyMC minibatch inference would use `pm.Minibatch` with variational inference (ADVI) to handle large datasets efficiently. This conceptual demonstration provides a foundation for exploring real minibatch streaming inference with larger datasets and integration with Dask.

Here is the code snippet and output:

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/8/89e638c6bd26b469b80909fbb2563f5a1fc5a3eb.png)

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/pymc3/original/2X/e/eab28658adff1c87cb391948b7397dcbe02c8c58.png)

I would appreciate your guidance on the following:

1. Should I focus on implementing `pm.Minibatch` with ADVI first, or continue refining conceptual streaming approaches like this demo?

2. When using minibatches in PyMC, how should the likelihood be scaled to correctly approximate full-data inference?

3. For this project, should I primarily focus on variational inference (ADVI), or also explore extending MCMC methods such as NUTS?

4. For the proposal, would a notebook demonstrating minibatch inference with simulated data be sufficient, or should I aim for a more advanced prototype?

---

<div class="post-metadata">

**Author:** ![zaxtax](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/zaxtax/32/3082_2.png) [@zaxtax](https://discourse.pymc.io/u/zaxtax)\
**Post date:** [March 23, 2026, 5:43pm UTC](https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646/2 "2026-03-23T17:43:13Z")

</div>

I think a simple notebook is a great start! No need for an advanced prototype just yet

---

<div class="post-metadata">

**Author:** ![jhanani](https://avatars.discourse-cdn.com/v4/letter/j/f17d59/32.png) [@jhanani](https://discourse.pymc.io/u/jhanani)\
**Post date:** [March 24, 2026, 7:09pm UTC](https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646/3 "2026-03-24T19:09:54Z")

</div>

Thank you for the guidance. I will focus on developing a clear notebook to explore streaming-style inference in PyMC.

I wanted to clarify the direction to ensure I’m aligning well with the project goals. I am considering a few approaches, such as focusing on minibatch-based inference, simulating streaming data with incremental updates, or incorporating tools like Dask for handling chunked data.

Which of these directions would you recommend prioritizing at this stage?

---

<div class="post-metadata">

**Author:** ![zaxtax](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/zaxtax/32/3082_2.png) [@zaxtax](https://discourse.pymc.io/u/zaxtax)\
**Post date:** [March 25, 2026, 5:32pm UTC](https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646/4 "2026-03-25T17:32:22Z")

</div>

Focus on the base functionality in the library itself. This is a good chance to get to know how pytensor works, how we do minibatching now, and how we might do it better. The dask integration will come later.

---

<div class="post-metadata">

**Author:** ![jhanani](https://avatars.discourse-cdn.com/v4/letter/j/f17d59/32.png) [@jhanani](https://discourse.pymc.io/u/jhanani)\
**Post date:** [March 26, 2026, 5:37pm UTC](https://discourse.pymc.io/t/exploring-streaming-minibatch-inference-in-pymc/17646/5 "2026-03-26T17:37:52Z")

</div>

Thank you so much for the guidance. I will try to focus on these areas.
