# Production grade prediction

**URL:** <https://discourse.pymc.io/t/production-grade-prediction/270>\
**Category:** Questions\
**Created:** [August 20, 2017, 4:53pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270 "2017-08-20T16:53:04Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![gyarnykh](https://avatars.discourse-cdn.com/v4/letter/g/4491bb/32.png) [@gyarnykh](https://discourse.pymc.io/u/gyarnykh)\
**Post date:** [August 20, 2017, 4:53pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/1 "2017-08-20T16:53:04Z")

</div>

How people generally transfer their models/make production live predictions after fitting Pymc3 models? Do they use different package for that? Plus to that, how would you make online updates to your model in that case (i.e. online learning)?

I know that main value of bayesian tasks is inference and better insights into models on small/medium datasets, however, results where integration over parameter distribution gives superior prediction results [http://twiecki.github.io/blog/2017/03/14/random-walk-deep-net/](http://twiecki.github.io/blog/2017/03/14/random-walk-deep-net/), actually makes you willing to try this instead of neural nets packages that are by default suited for this online updating tasks

---

<div class="post-metadata">

**Author:** ![twiecki](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/twiecki/32/6930_2.png) [@twiecki](https://discourse.pymc.io/u/twiecki)\
**Post date:** [August 21, 2017, 11:13am UTC](https://discourse.pymc.io/t/production-grade-prediction/270/2 "2017-08-21T11:13:28Z")

</div>

There isn’t too much info available on that, alas. Nicole has done some nice work in creating a sklearn wrapper for PyMC3 models, that’s probably a good starting point for productizing: [https://github.com/parsing-science/ps-toolkit/blob/master/ps\_toolkit/pymc3\_models/HLM.py](https://github.com/parsing-science/ps-toolkit/blob/master/ps_toolkit/pymc3_models/HLM.py). Then there is also `sampled` by @colcarroll [https://github.com/ColCarroll/sampled](https://github.com/ColCarroll/sampled)

---

<div class="post-metadata">

**Author:** ![AustinRochford](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/austinrochford/32/92_2.png) [@AustinRochford](https://discourse.pymc.io/u/AustinRochford)\
**Post date:** [August 21, 2017, 1:42pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/3 "2017-08-21T13:42:58Z")

</div>

FWIW, I haven’t really found a good, convenient way to do online learning with PyMC3.

---

<div class="post-metadata">

**Author:** ![colcarroll](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/colcarroll/32/9_2.png) [@colcarroll](https://discourse.pymc.io/u/colcarroll)\
**Post date:** [August 24, 2017, 4:33pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/4 "2017-08-24T16:33:25Z")

</div>

I keep meaning to try doing something with the [Interpolated](https://github.com/pymc-devs/pymc3/blob/master/pymc3/distributions/continuous.py#L1934) distribution to get something working online, though it feels like that might still be too computationally intensive to do (correctly) on the fly.

I keep meaning to build a sample PyMC3 web app, just to work out how much work needs to be done to make that sort of thing easy-ish. Perhaps that would be followed up with a more comprehensive [dash](https://github.com/plotly/dash) project.

Current plan for `sampled` is to use it to prebuild models and distribute them (note: you can do this with base PyMC3 as well). This lets you isolate the modeling step from the training step.  
Something like

```auto
from models import naive_bayes
import pymc3 as pm

corpus, labels = load_training_data()

with naive_bayes(corpus=corpus, labels=labels):
    trace = pm.sample()

```

---

<div class="post-metadata">

**Author:** ![ferrine](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ferrine/32/10_2.png) [@ferrine](https://discourse.pymc.io/u/ferrine)\
**Post date:** [August 26, 2017, 4:41pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/5 "2017-08-26T16:41:03Z")

</div>

I suggest using variational approx for that task if model is large or use pm.Empirical to wrap a trace. This way you can build pure theano graph for interested expression wrt inputs and save it properly

---

<div class="post-metadata">

**Author:** ![springcoil](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/springcoil/32/36_2.png) [@springcoil](https://discourse.pymc.io/u/springcoil)\
**Post date:** [September 16, 2017, 9:59pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/6 "2017-09-16T21:59:14Z")

</div>

I keep meaning to dig into this too with a simple dash app or something just to see how it works.

Maybe at work I’ll hack together something.

---

<div class="post-metadata">

**Author:** ![springcoil](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/springcoil/32/36_2.png) [@springcoil](https://discourse.pymc.io/u/springcoil)\
**Post date:** [September 23, 2017, 2:56pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/7 "2017-09-23T14:56:12Z")

</div>

I tried to build an API - [https://gist.github.com/springcoil/b8ef6d073349b9102ec49845b757f6b4](https://gist.github.com/springcoil/b8ef6d073349b9102ec49845b757f6b4) with Sampled from @colcarroll Unfortunately I couldn’t quite get this to work. Does anyone else have any insight into making this work? My web API knowledge isn’t great 🙂

---

<div class="post-metadata">

**Author:** ![colcarroll](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/colcarroll/32/9_2.png) [@colcarroll](https://discourse.pymc.io/u/colcarroll)\
**Post date:** [September 24, 2017, 3:27pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/8 "2017-09-24T15:27:58Z")

</div>

I just pushed a small Flask app I’ve been working on – first time trying the `cookiecutter` app to set something up, but it seems pretty plug and play – see the `README` for two lines to get it running locally. It exposes a bit of an API, and also creates a UI to interact with it and visualize the results using `d3js`. I’ll try to add more bells and whistles later 😃

> **[ColCarroll/flymc3](https://github.com/ColCarroll/flymc3)**
>
> Flask + PyMC3. Contribute to ColCarroll/flymc3 development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![dunovank](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/dunovank/32/3860_2.png) [@dunovank](https://discourse.pymc.io/u/dunovank)\
**Post date:** [March 29, 2021, 4:16pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/9 "2021-03-29T16:16:52Z")

</div>

Hey @twiecki , I’m getting a 404 from this link: [https://github.com/parsing-science/ps-toolkit/blob/master/ps\_toolkit/pymc3\_models/HLM.py](https://github.com/parsing-science/ps-toolkit/blob/master/ps_toolkit/pymc3_models/HLM.py)

Anyone have an updated link?

P.S. Also, if anyone has additional examples on putting a pymc3 model into production - ideally with the ability to periodically (scheduled or event-triggered) update posteriors with new observations - I’d be very grateful. Thanks!

---

<div class="post-metadata">

**Author:** ![twiecki](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/twiecki/32/6930_2.png) [@twiecki](https://discourse.pymc.io/u/twiecki)\
**Post date:** [March 30, 2021, 8:50pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/10 "2021-03-30T20:50:43Z")

</div>

Hey @dunovank, good to see you here!

We’re working with several clients who run models in production but there isn’t really a great resource. A replacement for that toolbox (which is also abandonware) can be found here: [GitHub - pymc-learn/pymc-learn: pymc-learn: Practical probabilistic machine learning in Python](https://github.com/pymc-learn/pymc-learn)

Then there is also [Automating daily runs for rt.live’s COVID-19 data using Airflow & ECS | by Mike Krieger | Medium](https://medium.com/@mikekrieger/automating-daily-runs-for-rt-lives-covid-19-data-dcda26ed2e2e)

Feel free to contact me directly if you want to discuss more.

---

<div class="post-metadata">

**Author:** ![dunovank](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/dunovank/32/3860_2.png) [@dunovank](https://discourse.pymc.io/u/dunovank)\
**Post date:** [March 30, 2021, 10:35pm UTC](https://discourse.pymc.io/t/production-grade-prediction/270/11 "2021-03-30T22:35:28Z")

</div>

Great, thanks @twiecki !

Basically, I have a docker pod that includes some representation learning and vector indexing services as well as a really simple binary classifier (like super simple, single predictor, logistic regression). Currently I’m just using sklearn for the classifier but the data I’m working with has a natural hierarchical structure that would benefit from partial pooling approach.

Anyway, I will most likely reach out directly at some point in the coming weeks - would be great to catch up and get any advice on deploying pymc3.
