Pandas often reads data with wrong types — numbers as strings, booleans as objects, categories as plain strings. Wrong types cause silent errors in math, filtering, and groupby operations.
# Always check types first
df.dtypes
# Quick summary of types + non-null counts
df.info()import pandas as pd
df = pd.DataFrame({
'student_id': ['101', '102', '103'], # should be int
'gpa': ['3.5', '2.8', '3.9'], # should be float
'age': [20, 21, 19], # fine as int
'us_citizen': ['Y', 'N', 'Y'], # could be category
'enrolled': [1, 0, 1] # could be bool
})
# String → int
df['student_id'] = df['student_id'].astype(int)
# String → float
df['gpa'] = df['gpa'].astype(float)
# int → bool
df['enrolled'] = df['enrolled'].astype(bool)
# string → category (saves memory for low-cardinality columns)
df['us_citizen'] = df['us_citizen'].astype('category')# Pass a dict of {column: dtype}
df = df.astype({
'student_id': int,
'gpa': float,
'enrolled': bool
})# → Integer
df['col'].astype(int)
df['col'].astype('int32') # smaller memory
df['col'].astype('int64') # default int size
# → Float
df['col'].astype(float)
df['col'].astype('float32') # smaller memory
# → String
df['col'].astype(str)
# → Boolean
df['col'].astype(bool)
# → Category (like enum — good for repeated string values)
df['col'].astype('category')# Silently skips conversion if it fails (keeps original type)
df['col'].astype(int, errors='ignore')
⚠️ Preferpd.to_numeric()witherrors='coerce'instead — it gives you more control (see below).
# errors='coerce' turns bad values into NaN instead of crashing
df['gpa'] = pd.to_numeric(df['gpa'], errors='coerce')
# Then fill or drop the NaNs
df['gpa'].fillna(0)
df.dropna(subset=['gpa'])When to use over astype(): When the column has messy data like
'3.5', 'N/A', '', None mixed together.
df['enrollment_date'] = pd.to_datetime(df['enrollment_date'])
# Custom format
df['enrollment_date'] = pd.to_datetime(df['enrollment_date'], format='%d/%m/%Y')
# Coerce bad dates to NaT (Not a Time)
df['enrollment_date'] = pd.to_datetime(df['enrollment_date'], errors='coerce')
# Extract parts after converting
df['enroll_year'] = df['enrollment_date'].dt.year
df['enroll_month'] = df['enrollment_date'].dt.month# Map string flags to actual booleans
df['is_citizen'] = df['us_citizen'].map({'Y': True, 'N': False})
# Or use == comparison
df['is_citizen'] = df['us_citizen'] == 'Y'
# Then cast to int (0/1) for math
df['is_citizen_int'] = df['is_citizen'].astype(int)# Downcast integers to smallest fitting type
df['age'] = pd.to_numeric(df['age'], downcast='integer')
# Downcast floats
df['gpa'] = pd.to_numeric(df['gpa'], downcast='float')
# Convert all object columns that look numeric
for col in df.select_dtypes('object').columns:
df[col] = pd.to_numeric(df[col], errors='ignore')print(df.dtypes) # see all types
print(df['gpa'].dtype) # single column type
# Memory usage before and after
df.memory_usage(deep=True)| Goal | Method |
|---|---|
| Check all types | df.dtypes or df.info() |
| Convert one column | df['col'].astype(int) |
| Convert multiple columns | df.astype({'col1': int, 'col2': float}) |
| Safe numeric conversion | pd.to_numeric(df['col'], errors='coerce') |
| Convert to datetime | pd.to_datetime(df['col']) |
| String flag → bool | df['col'].map({'Y': True, 'N': False}) |
| Reduce memory usage | pd.to_numeric(..., downcast='integer') |
| Low-cardinality strings | .astype('category') |
Clean numeric strings ('3.5', '101') → astype(float) / astype(int)
Messy data with NaN/blanks/errors → pd.to_numeric(..., errors='coerce')
Date strings → pd.to_datetime()
'Y'/'N' or '1'/'0' flags → .map() or == comparison
Repeated string values → astype('category')