Skip to content

Instantly share code, notes, and snippets.

@sreekarun
Created April 5, 2026 13:58
Show Gist options
  • Select an option

  • Save sreekarun/cd6f682467c38d48a61157ad801a11a4 to your computer and use it in GitHub Desktop.

Select an option

Save sreekarun/cd6f682467c38d48a61157ad801a11a4 to your computer and use it in GitHub Desktop.
Change data types

Changing Data Types with astype() in Pandas


WHY IT MATTERS

Pandas often reads data with wrong types — numbers as strings, booleans as objects, categories as plain strings. Wrong types cause silent errors in math, filtering, and groupby operations.

# Always check types first
df.dtypes

# Quick summary of types + non-null counts
df.info()

1. Basic astype() — convert a single column

import pandas as pd

df = pd.DataFrame({
    'student_id': ['101', '102', '103'],   # should be int
    'gpa':        ['3.5', '2.8', '3.9'],   # should be float
    'age':        [20, 21, 19],             # fine as int
    'us_citizen': ['Y', 'N', 'Y'],          # could be category
    'enrolled':   [1, 0, 1]                 # could be bool
})

# String → int
df['student_id'] = df['student_id'].astype(int)

# String → float
df['gpa'] = df['gpa'].astype(float)

# int → bool
df['enrolled'] = df['enrolled'].astype(bool)

# string → category (saves memory for low-cardinality columns)
df['us_citizen'] = df['us_citizen'].astype('category')

2. Convert Multiple Columns at Once

# Pass a dict of {column: dtype}
df = df.astype({
    'student_id': int,
    'gpa':        float,
    'enrolled':   bool
})

3. Common Type Conversions Cheat Sheet

# → Integer
df['col'].astype(int)
df['col'].astype('int32')    # smaller memory
df['col'].astype('int64')    # default int size

# → Float
df['col'].astype(float)
df['col'].astype('float32')  # smaller memory

# → String
df['col'].astype(str)

# → Boolean
df['col'].astype(bool)

# → Category (like enum — good for repeated string values)
df['col'].astype('category')

4. Safe Conversion with errors='ignore'

# Silently skips conversion if it fails (keeps original type)
df['col'].astype(int, errors='ignore')

⚠️ Prefer pd.to_numeric() with errors='coerce' instead — it gives you more control (see below).


5. Safer Numeric Conversion — pd.to_numeric()

# errors='coerce' turns bad values into NaN instead of crashing
df['gpa'] = pd.to_numeric(df['gpa'], errors='coerce')

# Then fill or drop the NaNs
df['gpa'].fillna(0)
df.dropna(subset=['gpa'])

When to use over astype(): When the column has messy data like '3.5', 'N/A', '', None mixed together.


6. Date Conversion — pd.to_datetime()

df['enrollment_date'] = pd.to_datetime(df['enrollment_date'])

# Custom format
df['enrollment_date'] = pd.to_datetime(df['enrollment_date'], format='%d/%m/%Y')

# Coerce bad dates to NaT (Not a Time)
df['enrollment_date'] = pd.to_datetime(df['enrollment_date'], errors='coerce')

# Extract parts after converting
df['enroll_year']  = df['enrollment_date'].dt.year
df['enroll_month'] = df['enrollment_date'].dt.month

7. Boolean Conversion from 'Y'/'N' strings

# Map string flags to actual booleans
df['is_citizen'] = df['us_citizen'].map({'Y': True, 'N': False})

# Or use == comparison
df['is_citizen'] = df['us_citizen'] == 'Y'

# Then cast to int (0/1) for math
df['is_citizen_int'] = df['is_citizen'].astype(int)

8. Downcast to Save Memory — large DataFrames

# Downcast integers to smallest fitting type
df['age'] = pd.to_numeric(df['age'], downcast='integer')

# Downcast floats
df['gpa'] = pd.to_numeric(df['gpa'], downcast='float')

# Convert all object columns that look numeric
for col in df.select_dtypes('object').columns:
    df[col] = pd.to_numeric(df[col], errors='ignore')

9. Check Before and After

print(df.dtypes)          # see all types
print(df['gpa'].dtype)    # single column type

# Memory usage before and after
df.memory_usage(deep=True)

🧠 Quick Reference

Goal Method
Check all types df.dtypes or df.info()
Convert one column df['col'].astype(int)
Convert multiple columns df.astype({'col1': int, 'col2': float})
Safe numeric conversion pd.to_numeric(df['col'], errors='coerce')
Convert to datetime pd.to_datetime(df['col'])
String flag → bool df['col'].map({'Y': True, 'N': False})
Reduce memory usage pd.to_numeric(..., downcast='integer')
Low-cardinality strings .astype('category')

🧠 Mental Model — Which method to use?

Clean numeric strings ('3.5', '101')  →  astype(float) / astype(int)
Messy data with NaN/blanks/errors     →  pd.to_numeric(..., errors='coerce')
Date strings                          →  pd.to_datetime()
'Y'/'N' or '1'/'0' flags             →  .map() or == comparison
Repeated string values                →  astype('category')
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment