maths.free › Probability Theory › Distributions › Poisson distribution
Poisson distribution
In probability theory and statistics, the Poisson distribution (/ˈpwɑːsɒn/) is a discrete probability distribution that expresses the probability of a given number of events occurring in a fixed interval of time if…
Poisson distribution
In probability theory and statistics, the Poisson distribution (/ˈpwɑːsɒn/) is a discrete probability distribution that expresses the probability of a given number of events occurring in a fixed interval of time if these events occur with a known constant mean rate and independently of the time since the last event. It can also be used for the number of events in other types of intervals than time, and in dimension greater than 1 (e.g., number of events in a given area or volume). The Poisson distribution is named after French mathematician Siméon Denis Poisson. It plays an important role for discrete-stable distributions.
Under a Poisson distribution with the expectation of λ events in a given interval, the probability of k events in the same interval is: \[\frac{\lambda^k e^{-\lambda}}{k!} .\] For instance, consider a call center which receives an average of λ = 3 calls per minute at all times of day. If the number of calls received in any two given disjoint time intervals is independent, then the number k of calls received during any minute has a Poisson probability distribution. Receiving k = 1 to 4 calls then has a probability of about 0.77, while receiving 0 or at least 5 calls has a probability of about 0.23.
A classic example used to motivate the Poisson distribution is the number of radioactive decay events during a fixed observation period.
History
The introduction of the Poisson distribution is credited to French mathematician and physicist Siméon Denis Poisson (1781-1840), who published it together with his probability theory in Recherches sur la probabilité des jugements en matière criminelle et en matière civile (1837). This work theorizes about the number of wrongful convictions in a given country by focusing on certain random variables N that count the number of events that take place during a time interval of given length. However, similar results had already been given in 1711 by Abraham de Moivre in De Mensura Sortis seu; de Probabilitate Eventuum in Ludis a Casu Fortuito Pendentibus . This makes it an example of Stigler's law and it has prompted some authors to argue that the Poisson distribution should bear the name of de Moivre.
In 1860, Simon Newcomb fitted the Poisson distribution to the number of stars found in a unit of space. A further practical application was made by Ladislaus Bortkiewicz in 1898. Bortkiewicz showed that the frequency with which soldiers in the Prussian army were accidentally killed by horse kicks could be well modeled by a Poisson distribution..
Probability mass function
A discrete random variable X is said to have a Poisson distribution with parameter \(\lambda>0\) if it has a probability mass function given by: \[f(k; \lambda) = \Pr(X{=}k)= \frac{\lambda^k e^{-\lambda}}{k!},\] where
- k is the number of occurrences (\(k = 0, 1, 2, \ldots\))
- e is Euler's number (\(e = 2.71828\ldots\))
- k! = k(k–1) ··· (3)(2)(1) is the factorial.
The positive real number λ is equal to the expected value of X and also to its variance. \[\lambda = \operatorname{E}(X) = \operatorname{Var}(X).\]
The Poisson distribution can be applied to systems with a large number of possible events, each of which is rare. The number of such events that occur during a fixed time interval is, under the right circumstances, a random number with a Poisson distribution.
The equation can be adapted if, instead of the average number of events \(\lambda,\) we are given the average rate \(r\) at which events occur. Then \(\lambda = r t,\) and: \[P(k \text{ events in interval } t) = \frac{(rt)^k e^{-rt}}{k!}.\]
Examples
The Poisson distribution may be useful to model events such as:
- the number of meteorites greater than one-meter diameter that strike Earth in a year;
- the number of laser photons hitting a detector in a particular time interval;
- the number of students achieving a low and high mark in an exam; and
- locations of defects and dislocations in materials.
Examples of the occurrence of random points in space are: the locations of asteroid impacts with earth (2-dimensional), the locations of imperfections in a material (3-dimensional), and the locations of trees in a forest (2-dimensional).
Assumptions and validity
The Poisson distribution is an appropriate model if the following assumptions are true:
- k, a nonnegative integer, is the number of times an event occurs in an interval.
- The occurrence of one event does not affect the probability of a second event.
- The average rate at which events occur is independent of any occurrences.
- Two events cannot occur at exactly the same instant.
If these conditions are true, then k is a Poisson random variable; the distribution of k is a Poisson distribution.
The Poisson distribution is also the limit of a binomial distribution, for which the probability of success for each trial is \(p = \frac{\lambda}{n}\), where \(\lambda\) is the expectation and \(n\) is the number of trials, in the limit that \(n \to \infty\) with \(\lambda\) kept constant (see Related distributions):
\[\lim_{n \to \infty} \dbinom{n}{k} \left(\frac{\lambda}{n}\right)^k \, \left(1-\frac{\lambda}{n}\right)^{n-k} = \frac{\lambda^k}{k!} \, e^{-\lambda}\]
The Poisson distribution may also be derived from the differential equations
\[\frac{d\,P_k(t)}{dt}=\lambda\,\Big(P_{k-1}(t)-P_k(t)\Big)\]
with initial conditions \(P_k(0)=\delta_{k0}\) and evaluated at \(t=1\)
Examples that violate the Poisson assumptions
The number of students who arrive at the student union per minute will likely not follow a Poisson distribution, because the rate is not constant (low rate during class time, high rate between class times) and the arrivals of individual students are not independent (students tend to come in groups). The non-constant arrival rate may be modeled as a mixed Poisson distribution, and the arrival of groups rather than individual students as a compound Poisson process.
The number of magnitude 5 earthquakes per year in a country may not follow a Poisson distribution if one large earthquake increases the probability of aftershocks of similar magnitude.
Examples in which at least one event is guaranteed are not Poisson distributed; but may be modeled using a zero-truncated Poisson distribution.
Count distributions in which the number of intervals with zero events is higher than predicted by a Poisson model may be modeled using a zero-inflated model.
Descriptive statistics
- The expected value of a Poisson random variable is λ.
- The variance of a Poisson random variable is also λ.
- The coefficient of variation is \(\lambda^{-1/2},\) while the index of dispersion is 1.
- The mean absolute deviation about the mean is \[\operatorname{E}[\ |X-\lambda|\ ]= \frac{2 \lambda^{\lfloor\lambda\rfloor + 1} e^{-\lambda}}{\lfloor\lambda\rfloor!}.\]
- The mode of a Poisson-distributed random variable with non-integer λ is equal to \(\lfloor \lambda \rfloor,\) which is the largest integer less than or equal to λ. This is also written as floor(λ). When λ is a positive integer, the modes are λ and λ − 1.
- All of the cumulants of the Poisson distribution are equal to the expected value λ. The n-th factorial moment of the Poisson distribution is λ.
- The expected value of a Poisson process is sometimes decomposed into the product of intensity and exposure (or more generally expressed as the integral of an "intensity function" over time or space, sometimes described as "exposure").
Higher moments
The higher non-centered moments mk of the Poisson distribution are Touchard polynomials in λ: \[m_k = \sum_{i=0}^k \lambda^i \begin{Bmatrix} k \\ i \end{Bmatrix},\] where the braces { } denote Stirling numbers of the second kind. In other words, \[E[X] = \lambda, \quad E[X(X-1)] = \lambda^2, \quad E[X(X-1)(X-2)] = \lambda^3, \cdots\] When the expected value is set to λ = 1, Dobinski's formula implies that the n‑th moment is equal to the number of partitions of a set of size n.
A simple upper bound is: \[m_k = E[X^k] \le \left(\frac{k}{\log(k/\lambda+1)}\right)^k \le \lambda^k \exp\left(\frac{k^2}{2\lambda}\right).\]
Sums of Poisson-distributed random variables
If \(X_i \sim \operatorname{Pois}(\lambda_i)\) for \(i=1,\dotsc,n\) are independent, then \(\sum_{i=1}^n X_i \sim \operatorname{Pois}\left(\sum_{i=1}^n \lambda_i\right).\) A converse is Raikov's theorem, which says that if the sum of two independent random variables is Poisson-distributed, then so are each of those two independent random variables.
Maximum entropy
It is a maximum-entropy distribution among the set of generalized binomial distributions \(B_n(\lambda)\) with mean \(\lambda\) and \(n \to \infty\), where a generalized binomial distribution is defined as a distribution of the sum of N independent but not identically distributed Bernoulli variables.
Other properties
- The Poisson distributions are infinitely divisible probability distributions.
- The directed Kullback-Leibler divergence of \(P = \operatorname{Pois}(\lambda)\) from \(P_0 = \operatorname{Pois}(\lambda_0)\) is given by\[\operatorname{D}_{\text{KL}}(P\parallel P_0) = \lambda_0 - \lambda + \lambda \log \frac{\lambda}{\lambda_0}.\]
- If \(\lambda \geq 1\) is an integer, then \(Y\sim \operatorname{Pois}(\lambda)\) satisfies \(\Pr(Y \geq E[Y]) \geq \frac{1}{2}\) and \(\Pr(Y \leq E[Y]) \geq \frac{1}{2}.\)
- Bounds for the tail probabilities of a Poisson random variable \(X \sim \operatorname{Pois}(\lambda)\) can be derived using a Chernoff bound argument. \[\begin{align} P(X \geq x) &\leq \frac{\left(e \lambda\right)^x e^{-\lambda}}{x^x}, &\text{ for } x > \lambda, \\[1ex] P(X \leq x) &\leq \frac{\left(e \lambda\right)^x e^{-\lambda} }{x^x}, &\text{ for } x < \lambda. \end{align}\]
- The upper tail probability can be tightened (by a factor of at least two) as follows:\[P(X \geq x) \leq \frac{e^{-\operatorname{D}_{\text{KL}}(Q\parallel P)}}{\max{(2, \sqrt{4\pi\operatorname{D}_{\text{KL}}(Q\parallel P)}})}, \text{ for } x > \lambda,\] where \(\operatorname{D}_{\text{KL}}(Q\parallel P)\) is the Kullback-Leibler divergence of \(Q=\operatorname{Pois}(x)\) from \(P=\operatorname{Pois}(\lambda)\).
- Inequalities that relate the cumulative distribution function of a Poisson random variable \(X \sim \operatorname{Pois}(\lambda)\) to the cumulative distribution function \(\Phi\) of the standard normal distribution are as follows:
\[\Phi{\left(\operatorname{sign}(k-\lambda)\sqrt{2\operatorname{D}_{\text{KL}}(Q_-\parallel P)}\right)} < P(X \leq k) < \Phi{\left(\operatorname{sign}(k+1-\lambda) \sqrt{2\operatorname{D}_{\text{KL}}(Q_+\parallel P)}\right)}, \text{ for } k > 0,\] where \(\operatorname{D}_{\text{KL}}(Q_-\parallel P)\) is the Kullback-Leibler divergence of \(Q_-=\operatorname{Pois}(k)\) from \(P = \operatorname{Pois}(\lambda)\) and \(\operatorname{D}_{\text{KL}}(Q_+\parallel P)\) is the Kullback-Leibler divergence of \(Q_+=\operatorname{Pois}(k+1)\) from \(P\).
Poisson races
Let \(X \sim \operatorname{Pois}(\lambda)\) and \(Y \sim \operatorname{Pois}(\mu)\) be independent random variables, with \(\lambda < \mu,\) then we have that \[\frac{e^{-(\sqrt{\mu} -\sqrt{\lambda})^2 }}{(\lambda + \mu)^2} - \frac{e^{-(\lambda + \mu)}}{2\sqrt{\lambda \mu}} - \frac{e^{-(\lambda + \mu)}}{4\lambda \mu} \leq P(X - Y \geq 0) \leq e^{- (\sqrt{\mu} -\sqrt{\lambda})^2}\]
The upper bound is proved using a standard Chernoff bound.
The lower bound can be proved by noting that \(P(X-Y\geq0\mid X+Y=i)\) is the probability that \(Z \geq \frac{i}{2},\) where \(Z \sim \operatorname{Bin}\left(i, \frac{\lambda}{\lambda+\mu}\right),\) which is bounded below by \(\frac{1}{(i+1)^2} e^{-iD\left(0.5 \| \frac{\lambda}{\lambda+\mu}\right)},\) where \(D\) is relative entropy (See the entry on bounds on tails of binomial distributions for details). Further noting that \(X+Y \sim \operatorname{Pois}(\lambda+\mu),\) and computing a lower bound on the unconditional probability gives the result. More details can be found in the appendix of Kamath et al.
As a Binomial distribution with infinitesimal time-steps
The Poisson distribution can be derived as a limiting case to the binomial distribution as the number of trials goes to infinity and the expected number of successes remains fixed. See law of rare events below. Therefore, it can be used as an approximation of the binomial distribution if n is sufficiently large and p is sufficiently small. The Poisson distribution is a good approximation of the binomial distribution if n is at least 20 and p is smaller than or equal to 0.05, and an excellent approximation if n ≥ 100 and np ≤ 10. Letting \(F_{\mathrm B}\) and \(F_{\mathrm P}\) be the respective cumulative density functions of the binomial and Poisson distributions, one has: \[F_\mathrm{B}(k;n, p) \ \approx\ F_\mathrm{P}(k;\lambda=np).\] One derivation of this uses probability-generating functions. Consider a Bernoulli trial (coin-flip) whose probability of one success (or expected number of successes) is \(\lambda \leq 1\) within a given interval. Split the interval into n parts, and perform a trial in each subinterval with probability \(\tfrac{ \lambda }{n}\). The probability of k successes out of n trials over the entire interval is then given by the binomial distribution \[p_k^{(n)}=\binom nk \left(\frac{\lambda}{n}\right)^{\!k} \left(1{-}\frac{\lambda}{n}\right)^{\! n-k},\] whose generating function is: \[P^{(n)}(x)=\sum_{k=0}^n p_k^{(n)} x^k = \left(1-\frac{\lambda}{n} +\frac{\lambda}{n} x \right)^n.\] Taking the limit as n increases to infinity (with x fixed) and applying the product limit definition of the exponential function, this reduces to the generating function of the Poisson distribution: \[\lim_{n\to\infty} P^{(n)}(x) = \lim_{n\to\infty} \left(1{+}\tfrac{\lambda(x-1)}{n}\right)^n = e^{\lambda(x-1)} = \sum_{k= 0}^\infty e^{-\lambda}\frac{\lambda^k }{k!} x^k.\]
General
- If \(X_1 \sim \mathrm{Pois}(\lambda_1)\,\) and \(X_2 \sim \mathrm{Pois}(\lambda_2)\,\) are independent, then the difference \(Y = X_1 - X_2\) follows a Skellam distribution.
- If \(X_1 \sim \mathrm{Pois}(\lambda_1)\,\) and \(X_2 \sim \mathrm{Pois}(\lambda_2)\,\) are independent, then the distribution of \(X_1\) conditional on \(X_1+X_2\) is a binomial distribution. Specifically, if \(X_1+X_2=k,\) then \(X_1| X_1+X_2=k\sim \mathrm{Binom}(k, \lambda_1/(\lambda_1+\lambda_2)).\) More generally, if X1, X2, ..., Xn are independent Poisson random variables with parameters λ1, λ2, ..., λn then given \(\sum_{j=1}^n X_j=k,\) it follows that \(X_i\Big|\sum_{j=1}^n X_j=k \sim \mathrm{Binom}\left(k, \frac{\lambda_i}{\sum_{j=1}^n \lambda_j}\right).\) In fact, \(\{X_i\} \sim \mathrm{Multinom}\left(k, \left\{\frac{\lambda_i}{\sum_{j=1}^n\lambda_j}\right\}\right).\)
- If \(X \sim \mathrm{Pois}(\lambda)\,\) and the distribution of \(Y\) conditional on X = k is a binomial distribution, \(Y \mid (X = k) \sim \mathrm{Binom}(k, p),\) then the distribution of Y follows a Poisson distribution \(Y \sim \mathrm{Pois}(\lambda \cdot p).\) In fact, if, conditional on \(\{X = k\},\) \(\{Y_i\}\) follows a multinomial distribution, \(\{Y_i\} \mid (X = k) \sim \mathrm{Multinom}\left(k, p_i\right),\) then each \(Y_i\) follows an independent Poisson distribution \(Y_i \sim \mathrm{Pois}(\lambda \cdot p_i), \rho(Y_i, Y_j) = 0.\)
- The Poisson distribution is a special case of the discrete compound Poisson distribution (or stuttering Poisson distribution) with only a parameter. The discrete compound Poisson distribution can be deduced from the limiting distribution of univariate multinomial distribution. It is also a special case of a compound Poisson distribution.
- For sufficiently large values of λ, (say λ > 1000), the normal distribution with mean λ and variance λ (standard deviation \(\sqrt{\lambda}\)) is an excellent approximation to the Poisson. If λ is greater than about 10, then the normal distribution is a good approximation if an appropriate continuity correction is performed, i.e., if P(X ≤ x), where x is a non-negative integer, is replaced by P(X ≤ x + 0.5). \[F_\mathrm{Poisson}(x;\lambda) \approx F_\mathrm{normal}(x;\mu=\lambda,\sigma^2=\lambda)\]
- Variance-stabilizing transformation: If \(X \sim \mathrm{Pois}(\lambda),\) then \[Y = 2 \sqrt{X} \approx \mathcal{N}(2\sqrt{\lambda};1),\] and \[Y = \sqrt{X} \approx \mathcal{N}(\sqrt{\lambda};1/4).\] Under this transformation, the convergence to normality (as \(\lambda\) increases) is far faster than the untransformed variable. Other, slightly more complicated, variance stabilizing transformations are available, one of which is Anscombe transform. See Data transformation (statistics) for more general uses of transformations.
- If for every t > 0 the number of arrivals in the time interval [0, t] follows the Poisson distribution with mean λt, then the sequence of inter-arrival times are independent and identically distributed exponential random variables having mean 1/λ.
- The cumulative distribution functions of the Poisson and chi-squared distributions are related in the following ways: \[F_\text{Poisson}(k;\lambda) = 1-F_{\chi^2}(2\lambda;2(k+1)) \quad\quad \text{ integer } k,\] and \[P(X=k)=F_{\chi^2}(2\lambda;2(k+1)) -F_{\chi^2}(2\lambda;2k).\]
Tsopano inu Sikuti pali calculator yomwe imatha kufotokoza izi, koma zigawo zake zimatha kuwerengedwa. Timafuna kuyesera imodzi pansipa, kapena tidzalemba ya ife.
M’malo mwake, maakaunti aulere amawonjezera malemba pazophunzira zonse, mnda wa zomwe mwamaliza, mavuto anu omaliza m’malo limodzi, ndi m’bale amene mungafunse za nkhaniyo. Maphunziro a matekinoloje ndi otsegulira kwa aliyense, olembetsa kapena osalembetsa.
Kulembetsa KulowaZithunzi
Pitani pa dzina lililonse lachifaniziro kuti mudziwe tanthauzo lake, chithunzi chake, ndi zimene limatanthauza.
Mafunso omwe anthu amafunsa
What is the difference between probability and statistics?
Probability goes from a known model to what the data should look like; statistics goes from data back to the model. Probability theory is the deductive half.
What does the law of large numbers promise?
That the average of many independent samples converges to the expected value. It says nothing about any single trial.
Zigawo za m'nkhaniyi ndi zochokera Wikipedia (CC BY-SA 4.0). Kuphatikizapo ndi kufotokozanso pano; zolakwika ndi zathu.
Zambiri pa Probability Theory
Sample spaces and the axiomsRandom variables and expectationThe common distributionsThe law of large numbers and the central limit theorem