EMZETT.
Login

Pandas

In short: The standard Python library for tabular data — loads, filters, groups and transforms data similar to a spreadsheet, just programmable.

In more detail: Pandas’s central data structure is the DataFrame (a table with named columns), built on NumPy. It lets you read in, clean, aggregate and export CSV/Excel/database data — the de-facto standard for data analysis in Python.

In Depth

Basic structure: DataFrame and Series

import pandas as pd
 
df = pd.read_csv("sales.csv")
revenue_per_category = df.groupby("category")["revenue"].sum()
 
# Filtering
big_sales = df[df["revenue"] > 1000]
 
# Computing a new column
df["revenue_with_tax"] = df["revenue"] * 1.19
 
# Handling missing values
df = df.dropna(subset=["customer_id"])

A DataFrame consists of named columns (each column itself a Series object, technically a NumPy array of a uniform type) and an index for the rows — conceptually similar to a spreadsheet, but fully controllable programmatically and without the row limits of an Excel file (Pandas easily handles millions of rows, as long as memory is sufficient). A single column on its own (df["revenue"]) is a Series — one-dimensional, with the same index as the DataFrame it came from.

Typical operations

  • Filtering: boolean indexing like df[df["revenue"] > 1000] — reads almost like a SQL WHERE clause.
  • Grouping and aggregating: groupby corresponds to SQL GROUP BY, followed by an aggregation function (sum, mean, count).
  • Merging: merge corresponds to a SQL JOIN between two DataFrames based on shared columns.
  • Reshaping: pivot tables (pivot_table) turn rows into columns and vice versa, useful for cross-tabulation analyses.
  • Time series: built-in functions for date/time columns (resampling to daily/monthly level, timezone conversion) — one reason Pandas was originally developed for financial data analysis (the name stands for “panel data”).

Role in the data science workflow

In a typical data science workflow, Pandas usually handles the cleaning and preprocessing phase (“data wrangling”): handling missing values (dropping or filling them), fixing data types (e.g. converting text dates into real date objects), renaming or recomputing columns, removing duplicates — before the prepared data is passed on to more specialised libraries like SciPy (statistical tests) or machine learning tools (scikit-learn). In practice, data science teams often spend most of a project’s time exactly in this Pandas-heavy cleaning phase, not in the actual model training.

See also: NumPy, Data Science, SciPy