COVID Data Analysis: A Guided Pandas Project
SkillVeris Team
Engineering Team

You will load a real time-series COVID dataset into a pandas DataFrame and parse dates correctly so every later step works.
In this guide, you'll learn:
- You will compute 7-day rolling averages to smooth reporting noise and reveal the underlying trend rather than daily spikes.
- You will calculate daily new cases from cumulative totals using the diff() method, a common source of beginner bugs.
- You will compare regions fairly by normalizing per 100,000 population instead of using raw counts.
- You will build clear line charts with labeled axes, sources, and honest y-axis scaling that does not exaggerate change.
1What This COVID Pandas Project Teaches You
This guided COVID data analysis project teaches you the core pandas workflow used in every real data job: load a messy time series, clean it, compute rolling averages, and visualize trends honestly. By the end you will have a working notebook that turns raw daily case counts into a clear, defensible story about how an outbreak moved over time.
COVID data is a perfect teaching dataset because it is real, freely available, and full of the exact quirks you meet in professional work: cumulative versus daily counts, reporting gaps on weekends, revised numbers, and populations of wildly different sizes. Wrestling with those quirks is the point, not a distraction.
You do not need prior epidemiology knowledge. You need basic Python and a willingness to check your own numbers. We will treat the data with respect, because careless COVID charts genuinely misled people, and learning to avoid that is part of becoming a trustworthy analyst.
2Getting a Clean Dataset
Use a reputable public source such as the Our World in Data COVID dataset or the Johns Hopkins CSSE archive. Both publish tidy CSV files with one row per country per day, containing columns like date, location, total_cases, new_cases, and population. Download a single CSV so your project is reproducible offline.
Load it with pandas using read_csv and immediately parse the date column: pass parse_dates=['date'] so pandas treats it as real timestamps rather than strings. If dates are strings, sorting and rolling windows silently break, which is the most common beginner mistake in any time-series project.
💡Always inspect first
Before any analysis, run df.head(), df.info(), and df['date'].min(), df['date'].max(). Thirty seconds of looking prevents an hour of debugging a wrong assumption about the data's shape.
3Cumulative Versus Daily: The diff() Trick
Many COVID datasets report total_cases as a cumulative running total that only ever goes up. To get daily new cases you subtract each day from the one before. In pandas this is one line: after sorting by date within a location, new = df['total_cases'].diff(). The first value becomes NaN because there is no prior day, which is correct.
Watch out when you have multiple countries in one DataFrame. A plain diff() will subtract the last day of one country from the first day of the next, producing nonsense. Group first: df.groupby('location')['total_cases'].diff(). Grouping before differencing is the single most important habit in this project.
4Smoothing With 7-Day Rolling Averages
Raw daily counts are jagged because many places did not report on weekends, creating artificial dips and Monday spikes. A rolling average smooths this so you see the trend, not the reporting artifact. Use df['new_cases'].rolling(window=7).mean() to average each day with the six before it.
A 7-day window is standard for COVID precisely because it cancels the weekly reporting cycle: every window contains exactly one of each weekday. Plot the smoothed line on top of a faint raw line so readers can see both the noise and the trend, which is more honest than hiding the raw data.
- window=7 averages a full week to remove weekday reporting effects.
- Use min_periods to decide whether early days with incomplete windows show or stay blank.
- center=False (the default) means each point looks backward, matching how news dashboards report.
- Never present a rolling average without saying which window you used.
5Comparing Regions Fairly
Comparing raw case counts between a country of 5 million and one of 300 million is meaningless. Normalize by population to get cases per 100,000 people: df['cases_per_100k'] = df['new_cases'] / df['population'] * 100000. This is the metric public health officials actually compare.
Once normalized, a small country with an intense outbreak can correctly appear worse than a large country with more total cases. That reversal is not an error; it is the whole reason we normalize. Always state your denominator in the chart so readers know they are seeing rates, not counts.
6Visualizing Honestly
Use matplotlib or seaborn to draw line charts of the 7-day average over time. Label both axes, add a title that states the metric and region, and cite your data source and download date directly on the figure. An unlabeled chart is not analysis; it is decoration.
Resist tempting distortions. Do not truncate the y-axis to make a small rise look like a cliff, and do not mix cumulative and daily series on one axis. If you compare countries, put them on the same normalized scale. Honest defaults protect both your readers and your credibility.
⚠️Charts persuade even when they lie
A truncated axis or a raw-count comparison can imply a story the data does not support. When in doubt, show the full range and the normalized rate, and let the numbers speak plainly.
7Reading Trends and Turning Points
With a smoothed curve you can describe the outbreak in plain language: a rising slope means acceleration, a flattening top means a peak, and a sustained decline means the wave is receding. Compute the week-over-week percent change to quantify this instead of eyeballing it.
Be careful attributing causes. A drop after a policy change might be caused by that policy, by seasonality, by behavior, or by reduced testing. As an analyst you describe what the data shows and flag what it cannot prove. Correlation in a single time series is rarely enough to claim a cause.
8Handling Missing and Revised Data
Real COVID data has holidays with zero reports, negative daily values from corrections, and countries that changed how they counted mid-pandemic. Decide a policy for each: you might clip negative daily cases to zero, forward-fill short gaps, or simply leave NaN and let the rolling average handle it.
Document every decision in a markdown cell. Reproducibility means someone else can rerun your notebook and understand why a number looks the way it does. The goal is not perfect data, which does not exist, but transparent handling of imperfect data.
9Ways to Extend the Project
Once the core analysis works, stretch yourself. Add a small function that takes a country name and returns its peak date and peak 7-day average. Build a faceted chart comparing several countries at once. Join in a vaccinations file and explore whether case trends shifted after rollout.
- Write a reusable function that loads, cleans, and returns a country-level time series.
- Add a second dataset (deaths or vaccinations) and align it by date with a merge.
- Export a clean tidy CSV of your processed data as a portfolio artifact.
- Turn the final chart into a short written summary an editor could publish.
10Frequently Asked Questions
Do I need epidemiology knowledge for this project? No. You need basic Python and pandas. The epidemiology concepts you use, like per-capita rates and smoothing, are explained as you go and are really just careful data handling.
Why use a 7-day rolling average instead of another window? Because COVID reporting follows a weekly cycle, a 7-day window contains exactly one of each weekday and cancels that artificial pattern, revealing the true trend underneath the weekend dips.
How do I turn cumulative totals into daily counts? Sort by date within each location, then use groupby('location')['total_cases'].diff(). Grouping first prevents pandas from subtracting one country's last day from the next country's first day.
Where can I get reliable COVID data now? Archived datasets from Our World in Data and Johns Hopkins CSSE remain publicly available and are ideal for practice because they are tidy, documented, and reproducible offline.
Is it okay that some daily values are negative? Negative daily counts come from official corrections to earlier totals. Decide a documented policy, such as clipping them to zero or leaving them, and note your choice so the analysis stays transparent.
Can I learn all of this for free? Yes. SkillVeris offers free pandas and data analysis courses with hands-on projects, so you can build this notebook without paying for anything.
11Next Steps
You now have a repeatable pandas workflow: load and parse dates, difference cumulative totals within groups, smooth with a 7-day rolling average, normalize per capita, and visualize honestly. Those five moves reappear in almost every time-series analysis you will ever do, from web traffic to sales to sensor data.
To go further, explore the free pandas, Python, and data visualization courses on SkillVeris, and pair this project with study notes on time series and statistics. Rebuild the notebook from scratch once without looking, and you will genuinely own the skill.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Engineering Team
Our engineering team documents real build journeys so you can learn by doing, not just reading.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.