The dictionary meaning of “Empirical” is based on observation or experience. Empirical probability also referred to as experimental probability, centers on real-world experiments and observations. Empirical probability is rooted in the notion that the probability of an event is based on the frequency of that event's occurrence in past experiments.
|
Empirical Probability is also called “A posteriori” or "Frequentist" or “Relative Frequency” probability |
Empirical probability is determined by counting the number of times an event has occurred () and then dividing it by the total number of trials () of an experiment. For an event , it can be denoted by the notation:
The empirical probability becomes a more accurate approximation of the true probability as the number of trials in an experiment increases.
Example 1.4. Let's examine the “coin-flipping” experiment through the lens of empirical probability. Consider a fair coin that was flipped three times (), with “heads” appearing zero times (), as illustrated in Fig 1.2.
Since none of the previous trials yielded “heads”, the probability that the next coin flip will result in “heads” is:
Alternatively, in a second example, if the coin was flipped 100 times (), with “heads” appearing 30 times (), the probability changes to:
Finally, in a third example, if the coin was flipped 100,000 times (), with “heads” appearing 48,000 times (), the probability changes to:
So, as evident, the empirical probability can change based on the number of trials. Also, as the number of trials increases, the probability tends to stabilize or converge towards a specific value, increasing its accuracy.
Example 1.5. Let’s look at a risk disclosure taken from one of the stock exchanges (Fig 1.3). The exchange looked at the historical performance of its individual traders and gathered the below data:
|
The ratio "9 out of 10" does not represent an actual count of traders analyzed. Instead, it's a way of expressing that among a large group of traders whose data was analysed, approximately 90% faced net losses. In mathematical terms, this could be stated as a proportion of 9/10 or 0.9. |
What is the probability that a randomly selected individual trader (in equity Futures and Options segment) has incurred losses?
In this context, the experiment can be defined as “the process of selecting individual traders for the analysis of their performance in equity futures trading by the stock exchange company”. Considering that the proportion of traders whose performance has been analysed is 10, the number of trials () of the experiment is 10.
The event refers to an “individual trader incurring a net loss”. Given that the proportion of traders who have faced losses is 9, the number of occurrences () of this event is 9. As a result, the probability that a randomly selected individual trader has incurred losses is:
Example 1.6. Consider another example where a blogger has launched a new newsletter and wants to determine the chances of their subscribers opening the newsletter email. To do this, they examine the historical data from their previous newsletter and discover that out of a total of 1,000 subscribers, 200 opened the newsletter emails. So, what would be the probability of the subscribers opening the new newsletter?
In this case, the experiment can be defined as “analysing a subscriber’s email behavior” and since there are 1000 subscribers, the number of trials () of the experiment is 1000. The event refers to “a subscriber opening the newsletter email”, thus the number of occurrences of the event () is 200. The empirical probability of a subscriber opening the email can be calculated as:
What if the blogger reduced the sample size and examined the historical data of just 10 subscribers instead of 1000? Now, there would have been only one subscriber who opened the email (Fig 1.4), so the event “a subscriber opening the newsletter email” would have occurred just once.
Consequently, our empirical probability would have been 1/10 = 0.1 or 10%, which is significantly different from the 20% probability observed earlier with 1000 subscribers, demonstrating the idea that as the number of trials increases, the empirical probability converges to the true or actual probability.
While empirical probability can be applied to situations where the outcomes are not equally likely, it also has a few limitations. Firstly, as shown earlier, it leaves the question of how many experiments (e.g., 10, 100, 1000, 10,000) are needed to obtain a reliable approximation of the true probability of the event. Secondly, there may be situations in the real world where no observed data is available and it's not feasible to repeat the experiments several times. This is where our next probability interpretation becomes relevant.