SciPy
In short: A scientific Python library built on NumPy — offers advanced functions for optimisation, statistics, signal processing and linear algebra.
In more detail: While NumPy provides the basic data structure (arrays) and basic operations, SciPy adds more specialised scientific algorithms (e.g. numerical integration, Fourier transforms) that you’d otherwise have to implement yourself. Frequently used together with NumPy and Pandas in the same data science project.
In Depth
Themed submodules
from scipy import stats, optimize
# Statistical test: do two groups differ significantly?
t_stat, p_value = stats.ttest_ind(group_a, group_b)
# Optimisation: find the minimum of a function
result = optimize.minimize(lambda x: (x - 3) ** 2, x0=0)
print(result.x) # close to 3.0SciPy is organised into themed submodules, each for a scientific field: scipy.optimize (optimisation problems, e.g. minimising a function or solving a system of equations), scipy.stats (statistical tests and probability distributions), scipy.signal (signal processing, filters), scipy.linalg (advanced linear algebra beyond NumPy’s basic functions), scipy.integrate (numerical integration, differential equations), scipy.interpolate (interpolation between data points), and several others. This division reflects SciPy’s role as a collection of mature, well-tested implementations of classic numerical/scientific methods, rather than a monolithic tool — you import exactly the submodule you need.
Historical background
SciPy emerged in the late 1990s/early 2000s as a community project aiming to make many functions from commercial scientific software (like MATLAB) freely available. Together with NumPy (which was originally part of the same project before splitting off as an independent base library), SciPy formed the foundation of today’s Python scientific ecosystem — practically every later library in this space (scikit-learn for machine learning, Matplotlib for visualisation) builds directly or indirectly on the data structures and conventions established by SciPy/NumPy.
Division of labour in the Python ecosystem
The division of labour in the Python data science ecosystem is usually clear: NumPy provides the basic array data structure and simple operations (addition, matrix multiplication), Pandas builds on that with tabular data management with named columns and time-series functionality, and SciPy supplements both with specialised scientific algorithms you’d otherwise have to laboriously implement yourself. A typical project combines all three: Pandas for reading in/cleaning raw data, NumPy for numerical processing, SciPy for the actual statistical analysis or optimisation.
Limits
SciPy is designed for classic numerical/statistical methods, not modern machine learning workflows — for that, more specialised libraries exist, like scikit-learn (classic ML) or PyTorch/TensorFlow (deep learning), which partly build on SciPy concepts but bring their own, optimised data structures (e.g. GPU-capable tensors).
See also: NumPy, Data Science, Pandas