# What is the suitable regression model for proportions data with 0 and 1 values?

**URL:** <https://discourse.pymc.io/t/what-is-the-suitable-regression-model-for-proportions-data-with-0-and-1-values/14598>\
**Category:** General\
**Tags:** modeling, model-checking\
**Created:** [June 13, 2024, 5:37am UTC](https://discourse.pymc.io/t/what-is-the-suitable-regression-model-for-proportions-data-with-0-and-1-values/14598 "2024-06-13T05:37:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mustapha\_Momoh](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/mustapha_momoh/32/8242_2.png) [@Mustapha\_Momoh](https://discourse.pymc.io/u/Mustapha_Momoh)\
**Post date:** [June 13, 2024, 5:37am UTC](https://discourse.pymc.io/t/what-is-the-suitable-regression-model-for-proportions-data-with-0-and-1-values/14598/1 "2024-06-13T05:37:46Z")

</div>

Hello PyMC Community,

I am analyzing a dataset for Regression Discontinuity Design where the dependent variable, _y_, represents the fraction of students at each unique score who rated their experience of taking an exam remotely above 4, i.e., above ‘somewhat satisfactory’. The independent variable, _x_, corresponds to these unique scores. However, my data includes exact 0s and 1s, for instances where all students at a performance level either rated below or entirely above 4. I have 2 questions:

1. Given that _y_ can be exactly 0 and 1 as well as fractional values in between, it is suitable to consider this as a straightforward probability distribution? If so, are there specific transformations you would recommend for y so I can model this using a Beta regression model? Currently, my beta regression model does not converge due to presence of these boundary values.

2. If you think it is a straightforward probability distribution, what regression model would best suit this data, especially given the presence of 0 and 1 values? Perhaps a Zero-and-One inflated Beta distribution? If so, is there an existing PyMC tutorial on such model that might help?

I appreciate any insights or suggestions.

I have attached a simulated data that illustrates the issue below. The independent variable i.e. _test\_performance_ values have been centered at 0 based on an observed discontinuity at 4. The _threshold_ variable represent treatment assignment. _Fractions_ is the dependent variable  
[simulated.csv](https://discourse.pymc.io/uploads/short-url/3SiWn6IIEZkrAm5et2RRFKGXSqL.csv) (1.2 KB)

---

<div class="post-metadata">

**Author:** ![jessegrabowski](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/jessegrabowski/32/5010_2.png) [@jessegrabowski](https://discourse.pymc.io/u/jessegrabowski)\
**Post date:** [June 13, 2024, 7:27am UTC](https://discourse.pymc.io/t/what-is-the-suitable-regression-model-for-proportions-data-with-0-and-1-values/14598/2 "2024-06-13T07:27:52Z")

</div>

You could try a LogisticNormal, which is `expit(Normal)`. It lies between zero and one, and admits boundary values. It’s also parameterized by `mu` and `sigma`, which might be more familiar to work with than the beta distribution parameters (the scale doesn’t influence the location).

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [June 13, 2024, 10:05am UTC](https://discourse.pymc.io/t/what-is-the-suitable-regression-model-for-proportions-data-with-0-and-1-values/14598/3 "2024-06-13T10:05:22Z")

</div>

For the zeros-ones you can either model them separately, or create a Mixture similar to how Hurdle Mixtures are implement in PyMC with components = `[pm.DiracDelta.dist(0), pm.Truncated.dist(..., lower=0 + eps, upper=1-eps), pm.DiracDelta.dist(1)]`, where `0/1 +- eps` are the smallest value you can register above 0 or below 1.

Either way, you will have to decide how to parametrize the weights, which inform the probability of observing 0, 1 or something in between.

---

<div class="post-metadata">

**Author:** ![Mustapha\_Momoh](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/mustapha_momoh/32/8242_2.png) [@Mustapha\_Momoh](https://discourse.pymc.io/u/Mustapha_Momoh)\
**Post date:** [June 13, 2024, 11:06am UTC](https://discourse.pymc.io/t/what-is-the-suitable-regression-model-for-proportions-data-with-0-and-1-values/14598/4 "2024-06-13T11:06:20Z")

</div>

Thanks @jessegrabowski and @ricardoV94 for your suggestions. I will try them and report back.
