Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Scipy stats package

Open in Colab

A variety of functionality for dealing with random numbers can be found in the scipy.stats package.

Distributions

There are over 100 continuous (univariate) distributions and about 15 discrete distributions provided by scipy

To avoid confusion with a norm in linear algebra we’ll import stats.norm as normal

A normal distribution with mean μ\mu and variance σ2\sigma^2 has a probability density function

1σ2πe−(x−μ)2/2σ2\frac{1}{\sigma \sqrt{2\pi}} e^{-(x-\mu)^2 / 2\sigma^2}
(-inf, inf)
(0.0, 1.0)
<Figure size 432x288 with 1 Axes>
<Figure size 432x288 with 1 Axes>

You can sample from a distribution using the rvs method (random variates)

<Figure size 432x288 with 1 Axes>

You can set a global random state using np.random.seed or by passing in a random_state argument to rvs

<Figure size 432x288 with 1 Axes>

To shift and scale any distribution, you can use the loc and scale keyword arguments

<Figure size 432x288 with 1 Axes>
<Figure size 432x288 with 1 Axes>

You can also “freeze” a distribution so you don’t need to keep passing in the parameters

<Figure size 432x288 with 1 Axes>

Example: The Laplace Distribution

We’ll review some of the methods we saw above on the Laplace distribution, which has a PDF

12e−∣x∣\frac{1}{2}e^{-|x|}
<Figure size 432x288 with 1 Axes>
<Figure size 432x288 with 1 Axes>
<Figure size 432x288 with 1 Axes>
(-inf, inf)
(0.0, 2.0)

Discrete Distributions

Discrete distributions have many of the same methods as continuous distributions. One notable difference is that they have a probability mass function (PMF) instead of a PDF.

We’ll look at the Poisson distribution as an example

(0, inf)
(1.0, 1.0)
<Figure size 432x288 with 1 Axes>
<Figure size 432x288 with 1 Axes>
<Figure size 432x288 with 1 Axes>

Fitting Parameters

You can fit distributions using maximum likelihood estimates

(0.059808015534485, 1.0078822447165796)

Statistical Tests

You can also perform a variety of tests. A list of tests available in scipy available can be found here.

the t-test tests whether the mean of a sample differs significantly from the expected mean

Ttest_1sampResult(statistic=0.5904283402851698, pvalue=0.5562489158694675)

The p-value is 0.56, so we would expect to see a sample that deviates from the expected mean at least this much in about 56% of experiments. If we’re doing a hypothesis test, we wouldn’t consider this significant unless the p-value were much smaller (e.g. 0.05)

If we want to test whether two samples could have come from the same distribution, we can use a Kolmogorov-Smirnov (KS) 2-sample test

KstestResult(statistic=0.023, pvalue=0.9542189106778983)

We see that the two samples are very likely to come from the same distribution

KstestResult(statistic=0.056, pvalue=0.08689937254547132)

When we draw from a different distribution, the p-value is much smaller, indicating that the samples are less likely to have been drawn from the same distribution.

Kernel Density Estimation

We might also want to estimate a continuous PDF from samples, which can be accomplished using a Gaussian kernel density estimator (KDE)

<Figure size 432x288 with 1 Axes>

A KDE works by creating a function that is the sum of Gaussians centered at each data point. In Scipy this is implemented as an object which can be called like a function

<Figure size 432x288 with 1 Axes>

We can change the bandwidth of the Gaussians used in the KDE using the bw_method parameter. You can also pass a function that will set this algorithmically. In the above plot, we see that the peak on the left is a bit smoother than we would expect for the Laplace distribution. We can decrease the bandwidth parameter to make it look sharper:

<Figure size 432x288 with 1 Axes>

If we set the bandwidth too small, then the KDE will look a bit rough.

<Figure size 432x288 with 1 Axes>