Another common mistake is ignoring the Prior Probability of the Hypothesis or in other words, ignoring the ratio of the Hypothesis. In our case, it would mean that while thinking about the probability that someone has Hypoillness, given that the person is experiencing some of its symptoms, we ignore the fact that the ratio of having Hypoillness is extremely low as it is a very rare illness, so the probability of a person having it lessens a lot right away. This is an extremely important aspect and ignoring it can give us drastically incorrect results.
Example 1.31. Walk or watch?
You want to go for a walk this morning, you look outside your window and see the weather. It's a bright sunny morning but as you are about to step out of the house, you get an alert from your weather app that there is a forecast for rain in the next few minutes, should you give up on your walk?
The morning looks too pleasant to pass up a walk, so you decide to do quick research to find out the accuracy of the app's forecasts. You go to the forecaster's website and find that the app has a rain forecast accuracy of 90% i.e., out of the 100 days on which it rained, the app predicted the rain on 90 of those days. This is pretty good! Along with this piece of information, you also find that out of the 100 dry days, the app predicted the dry days on 80 of those days. An accuracy of 80%, is not bad either!
Based on the accuracy, it looks like the app forecasts pretty reliably so you decide to give up on your walk and continue to watch Netflix. The whole day passed by but it did not rain and you missed your walk! Why? Because you did not use Bayes' Theorem!
Let's take a step back and ask the same questions that we asked in the example of Hypoillness
For the answer to the first question, say, it turns out that it rains only 10% of the time in your region. So, in a total of 100 days, it rains on 10 of those days. This means that with an accuracy of 90%, out of the 10 wet days, the app forecasts correctly 9 times.
For the second question, we know that out of a total of 100 days, there are 90 dry days and 10 wet days. For dry days, it has an accuracy of 80% i.e., it forecasts rain 20% of the time. So, out of the 90 dry days, the app forecasts rain on 18 days. Thus, the app forecasts rain on 27 days (18 incorrect + 9 correct).
Now, let's put all of this information in the context of the Bayes' Theorem equation:
Substituting the probabilities into the Bayes’ formula yields:
Observe that our accuracy of 90% reduced drastically to a mere 33%. But why did it happen?
It happened because the 90% accuracy answers the question “Given that it rains, what is the probability that the app forecasts rain, denoted by ( | )”? The reason why this accuracy is higher is that the app performs much better in forecasting wet days. Maybe because wet days have some specific aspects like temperature changes, wind etc. that help in making accurate forecasts.
However, the question we are interested in is “Given the app forecasts rain, what is the probability that it will actually rain”? The accuracy here is just 33%, which is pretty unreliable!
What happened here?
Although the app often correctly predicts rain when it actually rains, it doesn't rain very often, so the number of days on which it rains and on which rain is forecasted is quite small (9 days). Although the app rarely forecasts rain when it doesn't rain, there are many days on which it doesn't rain, so there are many opportunities for an incorrect forecast (18 days out of 100). Thus a prediction of rain is more often associated with a dry day than with a wet day. And that's what happened to you today!
Through our exploration of this example, the practical utility of Bayes' Theorem in deciphering complex probability scenarios becomes clear. By incorporating prior probabilities and updating them with new evidence, Bayes' Theorem is particularly invaluable in domains like financial forecasting, clinical research, and machine learning, where the ability to adapt to new data and adjust predictions accordingly is paramount.