Pandas dataframe.drop_duplicates()
Pandas drop_duplicates() method helps in removing duplicates from the Pandas Dataframe allows to remove duplicate rows from a DataFrame, either based on all columns or specific ones in python.
By default, drop_duplicates() scans the entire DataFrame for duplicate rows and removes all subsequent occurrences, retaining only the first instance being the simple and efficient method. Let’s see a quick example:
import pandas as pd
data = {
"Name": ["Alice", "Bob", "Alice", "David"],
"Age": [25, 30, 25, 40],
"City": ["NY", "LA", "NY", "Chicago"]
}
df = pd.DataFrame(data)
display(df)
# Removing duplicates
unique_df = df.drop_duplicates()
display(unique_df)
Output:

Pandas dataframe.drop_duplicates()
This example demonstrates how duplicate rows are removed while retaining the first occurrence using pandas.DataFrame.drop_duplicates() since it’s commonly used and recommended.
dataframe.drop_duplicates() Syntax in Python :
Syntax: DataFrame.drop_duplicates(subset=None, keep=’first’, inplace=False)
Parameters:
- subset: Subset takes a column or list of column label. It’s default value is none. After passing columns, it will consider them only for duplicates. ( Optional)
- keep: keep is to control how to consider duplicate value. It has only three distinct value and default is ‘first’.
- If ‘first‘, it considers first value as unique and rest of the same values as duplicate.
- If ‘last‘, it considers last value as unique and rest of the same values as duplicate.
- If False, it consider all of the same values as duplicates
- inplace: Boolean values, removes rows with duplicates if True.
Return type: DataFrame with removed duplicate rows depending on Arguments passed.
Python dataframe.drop_duplicates() : Examples
Duplicate rows can arise due to merging datasets, incorrect data entry, or other reasons. The drop_duplicates() works by identifying duplicates based on all columns (default) or specified columns and removing them as per your requirements. Below, we are discussing examples of dataframe.drop_duplicates() method:
1. Dropping Duplicates Based on Specific Columns
You can target duplicates in specific columns using the subset parameter. This helps when certain fields are more relevant for identifying duplicates.
import pandas as pd
df = pd.DataFrame({
'Name': ['Alice', 'Bob', 'Alice', 'David'],
'Age': [25, 30, 25, 40],
'City': ['NY', 'LA', 'SF', 'Chicago']
})
# Drop duplicates based on the 'Name' column
result = df.drop_duplicates(subset=['Name'])
print(result)
Output
Name Age City 0 Alice 25 NY 1 Bob 30 LA 3 David 40 Chicago
Here, duplicates are removed based solely on the Name column, ignoring the other fields. This is helpful when specific columns uniquely identify rows.
2. Keeping the Last Occurrence
By default, drop_duplicates() retains the first occurrence of duplicates. However, you can retain the last duplicate instead using keep='last'.
import pandas as pd
df = pd.DataFrame({
'Name': ['Alice', 'Bob', 'Alice', 'David'],
'Age': [25, 30, 25, 40],
'City': ['NY', 'LA', 'NY', 'Chicago']
})
# Keep the last occurrence of duplicates
result = df.drop_duplicates(keep='last')
print(result)
Output
Name Age City 1 Bob 30 LA 2 Alice 25 NY 3 David 40 Chicago
The keep='last' parameter ensures the last occurrence of each duplicate is retained instead of the first.
3. Dropping All Duplicates
To remove all rows with duplicates, use keep=False. This keeps only rows that are entirely unique.
import pandas as pd
df = pd.DataFrame({
'Name': ['Alice', 'Bob', 'Alice', 'David'],
'Age': [25, 30, 25, 40],
'City': ['NY', 'LA', 'NY', 'Chicago']
})
# Drop all duplicates
result = df.drop_duplicates(keep=False)
print(result)
Output
Name Age City 1 Bob 30 LA 3 David 40 Chicago
With keep=False, all occurrences of duplicate rows are removed, leaving only rows that are entirely unique across all columns.
4. Modifying the Original DataFrame Directly
To modify the original DataFrame directly without creating a new one, use inplace=True.
import pandas as pd
df = pd.DataFrame({
'Name': ['Alice', 'Bob', 'Alice', 'David'],
'Age': [25, 30, 25, 40],
'City': ['NY', 'LA', 'NY', 'Chicago']
})
# Modify the DataFrame in place
df.drop_duplicates(inplace=True)
print(df)
Output
Name Age City 0 Alice 25 NY 1 Bob 30 LA 3 David 40 Chicago
Using inplace=True modifies the original DataFrame directly, saving memory and avoiding the need to assign the result to a new variable.
Python | Pandas dataframe.drop_duplicates() – FAQs
What does drop_duplicates do in Python?
The
drop_duplicates()method in pandas is used to remove duplicate rows from a DataFrame. By default, this method considers all columns to identify duplicates, but you can specify particular columns to look for duplicates. When duplicates are found, only the first occurrence is kept, and the rest are dropped unless specified otherwise.
What does duplicated () do in pandas?
To remove duplicates from a DataFrame, you can use the
drop_duplicates()method. This method allows you to specify whether to drop duplicates across all columns or just a subset.Example:
# Remove duplicates considering only column 'A'
clean_df = df.drop_duplicates(subset=['A'])
print(clean_df)
What is the difference between duplicated and Drop_duplicates?
duplicated(): This method returns a boolean Series indicating whether each row is a duplicate (True) or not (False). This is useful when you need to identify duplicates without actually removing them from the DataFrame.drop_duplicates(): This method removes the duplicate rows from the DataFrame and returns a new DataFrame with only unique rows or modifies the original DataFrame in place ifinplace=Trueis specified.Both methods can operate on the entire DataFrame or on specific columns specified by the
subsetparameter and can be controlled further by parameters likekeep, which can specify whether to keep the first occurrence, the last, or none of the duplicates.
How to check if pandas DataFrame has duplicates?
To check if a Pandas DataFrame has duplicates, use the
duplicated()method combined withany(). Example:has_duplicates = df.duplicated().any()
print(has_duplicates) # True if duplicates exist
How Do I Merge Two DataFrames and Delete Duplicates?
To merge two DataFrames and remove duplicates, you can use
merge()followed bydrop_duplicates(). This ensures that the merged DataFrame does not contain duplicate rows.Example:
df1 = pd.DataFrame({
'A': [1, 2, 3],
'B': ['a', 'b', 'c']
})
df2 = pd.DataFrame({
'A': [4, 2, 3],
'B': ['d', 'b', 'c']
})
# Merge DataFrames and remove duplicates
merged_df = pd.merge(df1, df2, how='outer').drop_duplicates()
print(merged_df)


