Working With CSV and Excel Files in Python
SkillVeris Team
Engineering Team

You will read CSV and Excel files into pandas with the options that handle real-world messiness.
In this guide, you'll learn:
- You will write results back out to CSV and multi-sheet Excel workbooks cleanly.
- You will handle encoding, delimiter, and header problems that trip up beginners.
- You will clean common data issues like missing values, bad types, and stray whitespace.
- You will work with multiple sheets and combine several files into one dataset.
1Reading Tabular Files in Python
To work with CSV and Excel files in Python, you load them into pandas with pd.read_csv('file.csv') or pd.read_excel('file.xlsx'), which turns the file into a DataFrame you can filter, clean, and analyse in code. This one skill replaces hours of manual spreadsheet clicking with a few repeatable lines.
CSV (comma-separated values) is a plain-text format that any tool can open, while Excel files are richer binary workbooks with multiple sheets, formulas, and formatting. pandas reads both, and knowing the quirks of each saves you from the errors that stall most beginners.
Everything here uses pandas, so make sure it is installed with pip install pandas, plus openpyxl for Excel support. With those in place you can automate almost any tabular chore.
2Reading CSV Files Well
The basic call is df = pd.read_csv('data.csv'), but real files rarely behave. read_csv accepts dozens of arguments to cope. Use sep=';' when the delimiter is a semicolon, common in European exports. Use encoding='latin-1' or encoding='utf-8-sig' when strange characters appear. Use header=None when the file has no column names, then supply your own.
Other frequent lifesavers include skiprows to jump over metadata lines at the top, usecols to load only the columns you need, and parse_dates to turn date columns into real datetime objects on the way in. Reading a file correctly the first time avoids a cascade of cleaning later.
💡Peek before you parse
Open the file in a text editor first to see its delimiter, whether it has a header, and how dates and decimals are formatted. Thirty seconds of looking prevents most read errors.
3Reading Excel Workbooks
Excel files add the complication of multiple sheets. By default pd.read_excel loads the first sheet, but you choose one with sheet_name='Sales' or by index with sheet_name=1. Pass sheet_name=None to load every sheet at once into a dictionary of DataFrames keyed by sheet name, which is perfect for processing a whole workbook.
Watch for merged cells, header rows that start partway down, and formatting that looks like data but is not. The skiprows, header, and usecols arguments work here too. Because Excel stores true dates and numbers, type handling is often cleaner than CSV, but formulas are read as their last calculated value, not the formula itself.
4Cleaning What You Load
Freshly loaded data almost always needs a wash. Start by inspecting with df.info() and df.head() to spot wrong types and missing values. Numbers stored as text are a classic problem; convert them with pd.to_numeric(df['col'], errors='coerce'), which turns unparseable entries into NaN so you can decide what to do with them.
From there, strip stray whitespace with df['name'] = df['name'].str.strip(), standardise casing, drop duplicate rows with drop_duplicates(), and rename awkward columns. Handle missing values deliberately, either dropping them with dropna() or filling with fillna() and a documented choice.
- Fix data types, especially numbers and dates read as text.
- Trim whitespace and normalise inconsistent casing.
- Remove duplicate rows that would distort any total.
- Rename columns to short, code-friendly names.
- Decide and document how each missing value is handled.
5Writing Results Back Out
Once you have cleaned or summarised data, save it. Write CSV with df.to_csv('output.csv', index=False), where index=False stops pandas from adding an unwanted row-number column. For Excel, df.to_excel('output.xlsx', index=False) produces a workbook, and you can control the sheet name with sheet_name.
To write several DataFrames into one workbook on separate sheets, use an ExcelWriter as a context manager and call to_excel on each DataFrame with a different sheet_name. This is how you turn a script's output into a tidy, shareable report a colleague can open in Excel.
⚠️Mind the index
Forgetting index=False is the most common export mistake, leaving a stray unnamed column in every file. Add it by default unless you specifically want the index saved.
6Combining Many Files
A frequent real task is merging a folder of monthly CSVs into one dataset. Use Python's glob module to list matching files with glob.glob('data/*.csv'), read each into a DataFrame, collect them in a list, and combine with pd.concat(list_of_dfs, ignore_index=True). In a few lines you have unified twelve files that would take an hour to copy-paste.
If the files share a key column rather than the same layout, use merge instead of concat to join them side by side, much like a database join or a spreadsheet VLOOKUP. Choosing concat for stacking and merge for joining is a distinction worth learning early.
7Automating the Boring Parts
The real reward is automation. Any spreadsheet routine you do by hand every week, opening a file, filtering rows, adding a column, saving a report, can become a script that runs in seconds and never makes a typo. Wrap your steps in a function, point it at a folder, and let it process everything.
This repeatability is also more trustworthy than manual work. A script does the same thing every time, and you can read it to see exactly what happened, which matters when someone questions a number in your report.
8Common Pitfalls
A handful of issues account for most beginner frustration with tabular files, and recognising them turns cryptic errors into quick fixes.
- Encoding errors from special characters, fixed with the right encoding argument.
- Wrong delimiters producing one giant column, fixed with sep.
- Numbers with currency symbols or thousands separators read as text.
- Excel dates shifting because of format ambiguity, avoided by parsing explicitly.
- Silent data loss when a filter or merge drops rows you did not expect, so always check row counts.
9Frequently Asked Questions
Do I need Excel installed to read xlsx files in Python? No. pandas reads Excel files through the openpyxl library, which does not require Microsoft Excel. Install it with pip install openpyxl and you can read and write workbooks anywhere Python runs.
What is the difference between CSV and Excel files? A CSV is a plain-text file where values are separated by a delimiter, with no formatting or multiple sheets. An Excel file is a richer binary workbook that supports several sheets, formulas, and formatting. pandas handles both.
Why do I get encoding errors when reading a CSV? The file was saved with a character encoding different from the default pandas expects. Try encoding='utf-8-sig' for files exported from Excel, or encoding='latin-1' as a common fallback until the special characters read correctly.
How do I read only one sheet from an Excel file? Pass the sheet_name argument to pd.read_excel, either as the sheet's name like sheet_name='Q1' or its position like sheet_name=0. Use sheet_name=None to load all sheets into a dictionary at once.
How can I combine multiple CSV files into one? List the files with the glob module, read each into a DataFrame, and stack them with pd.concat while passing ignore_index=True. If the files share a key column instead of a layout, use merge to join them.
Is it faster than doing this in Excel by hand? Dramatically, especially for repetitive work on many files or large datasets. A script processes thousands of rows in seconds, runs identically every time, and leaves a readable record of exactly what it did.
10Next Steps
Reading, cleaning, writing, combining, and automating tabular files is one of the highest-return skills in practical Python. With pandas you replace hours of manual spreadsheet work with a handful of dependable lines, and the same patterns scale from one small CSV to a folder of large workbooks.
You can learn all of this free on SkillVeris, where the Python and data analysis courses and study notes cover file handling, pandas, and cleaning in short, hands-on lessons. Pick a repetitive spreadsheet task you already dread, automate it, and you will feel the payoff immediately.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Engineering Team
Our engineering writers turn abstract code concepts into hands-on, project-driven learning experiences.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.