Extremely Small Probabilities
Join the DZone community and get the full member experience.
Join For Freeone objection to modeling adult heights with a normal distribution is that the former is obviously positive but the latter can be negative. however, by this model negative heights are astronomically unlikely. i’ll explain below how one can take “astronomically” literally in this context.
a common model says that men’s and women’s heights are normally distributed with means of 70 and 64 inches respectively, both with a standard deviation of 3 inches. a woman with negative height would be 21.33 standard deviations below the mean, and a man with negative height would be 23.33 standard deviations below the mean. these events have probability 3 × 10 ^{ 101 } and 10 ^{ 120 } respectively. or to write them out in full
0.00000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000003
and
0.000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000001.
as i mentioned on twitter yesterday, if you’re worried about probabilities that require scientific notation to write down, you’ve probably exceeded the resolution of your model. i imagine most probability models are good to two or three decimal places at most. when model probabilities are extremely small, factors outside the model become more important than ones inside.
according to wolfram alpha, there are around 10 ^{ 80 } atoms in the universe. so picking one particular atom at random from all atoms in the universe would be on the order of a billion trillion times more likely than running into a woman with negative height. of course negative heights are not just unlikely, they’re impossible. as you travel from the mean out into the tails, the first problem you encounter with the normal approximation is not that the probability of negative heights is overestimated, but that the probability of extremely short and extremely tall people is under estimated. there exist people whose heights would be impossibly unlikely according to this normal approximation. see examples here .
probabilities such as those above have no practical value, but it’s interesting to see how you’d compute them anyway. you could find the probability of a man having negative height by typing
pnorm(23.33)
into r or
scipy.stats.norm.cdf(23.33)
into python. without relying on such software, you could use the bounds
with x equal to 21.33 and 23.33. for a proof of these bounds and tighter bounds see these notes .
Published at DZone with permission of John Cook, DZone MVB. See the original article here.
Opinions expressed by DZone contributors are their own.
Trending

Send Email Using Spring Boot (SMTP Integration)

How To Manage Vulnerabilities in Modern CloudNative Applications

Cypress Tutorial: A Comprehensive Guide With Examples and Best Practices

Using OpenAI Embeddings Search With SingleStoreDB
Comments