# Nested Parallel in PyMC3

**URL:** <https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371>\
**Category:** Questions\
**Created:** [December 4, 2020, 9:01pm UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371 "2020-12-04T21:01:46Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![HONGWEI\_LIU](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/hongwei_liu/32/3394_2.png) [@HONGWEI\_LIU](https://discourse.pymc.io/u/HONGWEI_LIU)\
**Post date:** [December 4, 2020, 9:01pm UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371/1 "2020-12-04T21:01:46Z")

</div>

Dear PyMC community,

My forward model (written by Fortran) is quite complex and time-consuming to calculate. Therefore, I have to use parallel in my forward model. However, these conflicts with PyMC3 parallel sampling. I am wondering if there is a way to allow the multiprocessing in PyMC3 sampling while my forward model can still be done through parallel computation.

Also, do you have any advice on speeding up PyMC3 for the time-consuming forward model? Many thanks for your help.

Best regards,

Hongwei

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [December 5, 2020, 8:26am UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371/2 "2020-12-05T08:26:22Z")

</div>

Have you try sampling with `core=1`? This will disable the pymc3 parallel sampling so it does not conflict your custom forward computation (or at least surface the bug if there is any)

---

<div class="post-metadata">

**Author:** ![HONGWEI\_LIU](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/hongwei_liu/32/3394_2.png) [@HONGWEI\_LIU](https://discourse.pymc.io/u/HONGWEI_LIU)\
**Post date:** [December 6, 2020, 1:22am UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371/3 "2020-12-06T01:22:42Z")

</div>

Hi Junpeng,

Thanks for the reply. In the case where I use core =1, it works well. But the entire process becomes very slow since there are many cores that are not used during sampling. I am wondering if there is a way to sample using multicores while my forward model is also using multiprocessing.

The multiprocessing in my forward model was implemented using Python Ray library. When I set pm.sample with core number larger than 2, the sampling stuck at the initial stage and not using any cores.

Thanks for your time again.

Hongwei

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [December 6, 2020, 8:36am UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371/4 "2020-12-06T08:36:33Z")

</div>

Maybe try controlling the number of cores `forward model` use? I am not sure how well it will works. We sometimes observe similar multiprocess race issue using GP, where some linear algebra operation is multi-threaded

---

<div class="post-metadata">

**Author:** ![junpenglao](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/junpenglao/32/8_2.png) [@junpenglao](https://discourse.pymc.io/u/junpenglao)\
**Post date:** [December 6, 2020, 8:38am UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371/5 "2020-12-06T08:38:47Z")

</div>

You can take a look of the following to see if there is something useful:

> [@PYMC3 Multiprocessing issue on a kubernetes cloud](https://discourse.pymc.io/t/pymc3-multiprocessing-issue-on-a-kubernetes-cloud/4824/2):
>
> The source of the other threads will not be multiprocessing I think, but either openmp through theano if you have big datasets somewhere, or blas if you are using eg matrix-vector products. You can configure the theano parallelization using as described here: [http://deeplearning.net/software/theano/library/config.html](http://deeplearning.net/software/theano/library/config.html) For BLAS it depends a bit on which implementation you are using (one of MKL, openblas, blis probably). On an intel cpu usually MKL is the goto implementation, you should get tha…

---

<div class="post-metadata">

**Author:** ![HONGWEI\_LIU](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/hongwei_liu/32/3394_2.png) [@HONGWEI\_LIU](https://discourse.pymc.io/u/HONGWEI_LIU)\
**Post date:** [December 6, 2020, 6:06pm UTC](https://discourse.pymc.io/t/nested-parallel-in-pymc3/6371/6 "2020-12-06T18:06:19Z")

</div>

Hi Junpeng,

Thank you very much for your help. I will see if a different implementation of parallelization can solve my problem.

Thanks again,

Hongwei
