# Ordinal regression with non-normally distributed response variable

**URL:** <https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356>\
**Category:** version agnostic\
**Tags:** bambi, modeling\
**Created:** [October 1, 2025, 9:47am UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356 "2025-10-01T09:47:28Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![leomein](https://avatars.discourse-cdn.com/v4/letter/l/87869e/32.png) [@leomein](https://discourse.pymc.io/u/leomein)\
**Post date:** [October 1, 2025, 9:47am UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/1 "2025-10-01T09:47:28Z")

</div>

Hi!

I just read through [Regression Models with Ordered Categorical Outcomes — PyMC example gallery](https://www.pymc.io/projects/examples/en/latest/generalized_linear_models/GLM-ordinal-regression.html) and am thinking about applying such a model to my data (currently using logistic regression because the response can be interpreted in a binary way).

Is my understanding right that by default a logistic distribution of the latent continuous response variable _Z_ is assumed? In my data there are four categories and categories 1 and 4 have more than twice as much data as categories 2 and 3. I’m wondering how to best deal with this situation and if a logit or probit link is appropriate here?

Thanks!

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [October 1, 2025, 12:13pm UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/2 "2025-10-01T12:13:22Z")

</div>

The “amount of data” shouldn’t matter?

The shape of the latent distributions has pretty subtle/minor effects, it’s usually just chosen for computational convenience.

I like this resource: [Ordinal Regression](https://betanalpha.github.io/assets/case_studies/ordinal_regression.html)

---

<div class="post-metadata">

**Author:** ![tcapretto](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/tcapretto/32/4902_2.png) [@tcapretto](https://discourse.pymc.io/u/tcapretto)\
**Post date:** [October 1, 2025, 12:32pm UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/3 "2025-10-01T12:32:27Z")

</div>

I want to add another resource I also like: [Ordinal Regression Models in Psychology: A Tutorial - Paul-Christian Bürkner, Matti Vuorre, 2019](https://journals.sagepub.com/doi/full/10.1177/2515245918823199)

---

<div class="post-metadata">

**Author:** ![leomein](https://avatars.discourse-cdn.com/v4/letter/l/87869e/32.png) [@leomein](https://discourse.pymc.io/u/leomein)\
**Post date:** [October 1, 2025, 4:57pm UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/4 "2025-10-01T16:57:07Z")

</div>

Thanks both! Hopefully these will clear up any misconceptions I seem to have.

---

<div class="post-metadata">

**Author:** ![leomein](https://avatars.discourse-cdn.com/v4/letter/l/87869e/32.png) [@leomein](https://discourse.pymc.io/u/leomein)\
**Post date:** [October 2, 2025, 7:20am UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/5 "2025-10-02T07:20:41Z")

</div>

I think this section from [Ordinal Regression](https://betanalpha.github.io/assets/case_studies/ordinal_regression.html) explains it:

> This construction holds for _any_ probability distribution over _X_. Differences in the shape of the distribution can be compensated by reconfiguring the interior cut points to achieve the desired ordinal probabilities. Consequently we have the luxury of selecting a probability distribution based on computational convenience, in particular the expense of the cumulative distribution function and the ultimate cost of computing the interval probabilities.

Thanks again!

---

<div class="post-metadata">

**Author:** ![ricardoV94](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/ricardov94/32/5775_2.png) [@ricardoV94](https://discourse.pymc.io/u/ricardoV94)\
**Post date:** [October 2, 2025, 11:39am UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/6 "2025-10-02T11:39:26Z")

</div>

I guess the shape matter a bit in relation to the cutpoints prior, but would be weird to try to think of an exotic distribution to fit with a specific prior on the cutpoints than the other way around.

---

<div class="post-metadata">

**Author:** ![bob-carpenter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.pymc.io/bob-carpenter/32/8124_2.png) [@bob-carpenter](https://discourse.pymc.io/u/bob-carpenter)\
**Post date:** [October 2, 2025, 6:01pm UTC](https://discourse.pymc.io/t/ordinal-regression-with-non-normally-distributed-response-variable/17356/7 "2025-10-02T18:01:51Z")

</div>

> [@leomein](#):
>
> Is my understanding right that by default a logistic distribution of the latent continuous response variable _Z_ is assumed?

The continuous response variable is typically marginalized out rather than realized explicitly. The ordinal logit distribution then takes on the character of differences of logistic cdfs based on the cutpoints. The reason this is done for Hamiltonian Monte Carlo is that the result remains continuously differentiable. If you include the latent variable, then you get discontinuities in derivatives that are not good for HMC sampling (the feedback essentially gets “cut”).

You can replace the logistic cdf with another cdf to get a different latent distribution assumption. The reason logit is popular is that the cdf has a simple explicit form—the inverse logit function 1 / (1 + \exp(-x)). If you move to ordinal probit, this gets replaced with \Phi(x), the standard normal  
cdf, which is much harder to compute because there’s no explicit form.
