Pandas GroupBy: The Analyst's Most Useful Tool
SkillVeris Team
Data Science Team

You will understand the split-apply-combine pattern that makes GroupBy intuitive rather than magic.
In this guide, you'll learn:
- You will translate real business questions directly into GroupBy operations.
- You will use aggregation functions like sum, mean, and count to summarize groups.
- You will group by multiple columns and apply several aggregations at once with the agg method.
- You will avoid common pitfalls like the confusing multi-level index and forgotten NaN handling.
1Why GroupBy Is So Powerful
Pandas GroupBy is the tool that answers almost every 'per category' business question — revenue per region, average order per customer, count of sign-ups per month — by splitting your data into groups, applying a calculation to each, and combining the results into a tidy summary. Learn it well and a huge share of everyday analysis becomes a single line of code.
Almost every real analytics question is a grouping question in disguise. The moment you hear 'by', 'per', or 'for each', you are describing a GroupBy. That is why experienced analysts reach for it constantly and why it is worth understanding deeply rather than copying snippets.
This guide explains the pattern behind GroupBy, walks through real examples, and points out the traps that trip up beginners.
2The Split-Apply-Combine Pattern
GroupBy feels like magic until you see the three-step pattern underneath it, called split-apply-combine. First, pandas splits the DataFrame into groups based on the column you choose — all the West rows together, all the East rows together. Second, it applies a function to each group independently, such as summing revenue. Third, it combines those per-group results into a single output table.
Once you picture those three steps, GroupBy stops being mysterious. Every GroupBy you ever write is just choosing what to split on, what to apply, and letting pandas combine the results for you. The syntax df.groupby('region')['revenue'].sum() maps exactly onto split by region, apply sum to revenue, combine into a series.
🔑Split, apply, combine
Split the data into groups, apply a calculation to each group, combine the answers into one table. Every GroupBy is these three steps — nothing more.
3Your First GroupBy
Suppose you have a sales table with columns for region, product, and revenue, and a stakeholder asks for total revenue by region. The answer is df.groupby('region')['revenue'].sum(). Pandas groups the rows by region, sums the revenue within each, and returns a neat series of one number per region.
Swap the function to change the question. Use .mean() for average revenue per region, .count() for how many sales each region made, .max() for the biggest single sale, or .min() for the smallest. The grouping stays the same; only the applied calculation changes. That flexibility is why one pattern answers so many questions.
4Turning Questions Into GroupBy
The real skill is hearing a business question and immediately seeing the GroupBy. The word after 'by' or 'per' is what you group on; the metric they want is what you aggregate. Practice this translation and analysis becomes fast and almost automatic.
- Average order value per customer: df.groupby('customer_id')['order_total'].mean().
- Number of orders per month: df.groupby('month')['order_id'].count().
- Total revenue per product category: df.groupby('category')['revenue'].sum().
- Highest-spending customer in each region: group by region, then find the max.
5Grouping by Multiple Columns
Real questions often need more than one grouping level. 'Revenue by region and by product' means grouping on both columns at once: df.groupby(['region', 'product'])['revenue'].sum(). Pandas creates a group for every combination that exists in the data and sums within each, giving you a detailed breakdown.
This returns a result with a multi-level index, which is powerful but can confuse beginners because it does not look like a flat table. Calling .reset_index() at the end flattens it back into ordinary columns, which is almost always what you want for further analysis or exporting to a report.
6Multiple Aggregations at Once With agg
Often you want several summaries in one shot — total, average, and count of revenue per region together. The agg method does this: df.groupby('region')['revenue'].agg(['sum', 'mean', 'count']) returns a table with a column for each statistic. This is far tidier than running three separate GroupBy calls.
You can go further and apply different functions to different columns, for example summing revenue while averaging quantity, by passing a dictionary to agg. This flexibility lets you build a rich summary table — essentially a custom report — in a single, readable line of code.
A quick agg recipe
A common pattern is producing a per-group summary table with named metrics.
Pass a list like ['sum', 'mean', 'count'] to get several stats on one column.
Pass a dict like {'revenue': 'sum', 'quantity': 'mean'} for per-column functions.
Follow with .reset_index() to get a clean, flat report table.7Common Pitfalls
The first trap is the multi-level index confusion just mentioned — remember .reset_index() when the output looks unusually nested. The second is forgetting how missing values behave: by default GroupBy excludes rows where the grouping key is NaN, so if a region column has blanks, those sales silently vanish from your totals.
A third subtlety is that most aggregation functions skip NaN values in the data being summarized, which is usually what you want but can surprise you if you expected them counted. When totals do not reconcile with a figure you trust, missing values in either the key or the value column are the usual culprit.
⚠️Watch the NaNs
GroupBy drops rows with a NaN grouping key by default, so blank categories disappear from your summary. If your grouped totals do not match the overall total, check for missing values in the column you grouped on.
8GroupBy and SQL Are the Same Idea
If you know SQL, GroupBy will feel familiar because it is the direct equivalent of GROUP BY. df.groupby('region')['revenue'].sum() is the pandas version of SELECT region, SUM(revenue) FROM sales GROUP BY region. The same split-apply-combine logic underlies both.
This connection is handy: reasoning about a query in SQL terms often clarifies the pandas version and vice versa. Analysts fluent in both simply pick whichever tool fits the data's location — SQL when it lives in a database, pandas when it is already in a DataFrame.
9Frequently Asked Questions
What does GroupBy actually do in pandas? It follows the split-apply-combine pattern: it splits your DataFrame into groups based on a column, applies a calculation like sum or mean to each group, and combines the results into a summary table.
How do I group by more than one column? Pass a list of column names, such as df.groupby(['region', 'product']). Pandas creates a group for each existing combination and aggregates within it, returning a result with a multi-level index you can flatten with reset_index.
How do I get multiple statistics at once? Use the agg method with a list of functions, like agg(['sum', 'mean', 'count']), or a dictionary mapping columns to functions for different aggregations per column. This produces a tidy summary table in one call.
Why are some rows missing from my grouped results? By default GroupBy drops rows where the grouping key is NaN, so blank categories disappear. Check for missing values in your grouping column if the totals do not reconcile with the overall figure.
Is pandas GroupBy the same as SQL GROUP BY? Conceptually yes. df.groupby('region')['revenue'].sum() is equivalent to SELECT region, SUM(revenue) FROM sales GROUP BY region. Both use split-apply-combine logic, just in different environments.
Where can I practice pandas GroupBy for free? SkillVeris offers free courses and study notes on Python and pandas, with hands-on examples of GroupBy and the split-apply-combine pattern using realistic datasets.
10Next Steps
Pandas GroupBy earns its reputation as the analyst's most useful tool because almost every 'per category' question maps onto its split-apply-combine pattern. Once you can hear a business question and instantly see what to group on and what to aggregate, a large slice of daily analysis becomes a single clear line of code — with agg and multi-column grouping ready when questions get richer.
You can practice GroupBy and the rest of pandas for free on SkillVeris, where the Python and data analysis courses and study notes walk you through realistic examples step by step. Take a dataset you know, ask three 'per' questions of it, and answer each with a GroupBy — the pattern will stick fast.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Data Science Team
Our data team shares real-world analytics, ML, and SQL insights grounded in industry practice.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.