How to get gender from a name in Python: gender-guesser, name data and APIs compared
There are three practical ways to get gender from a first name in Python: an offline library such as gender-guesser, your own lookup table built from official baby-name counts, or a name-gender API. The library is free and needs no network, but it only knows Latin-script names: on our 4,610-name test set it answered 83% of Latin-script names and none of the names written in Arabic, Chinese, Korean, Cyrillic, Greek, Hebrew or Thai script. A table built from birth records gives you real counts for one country. An API covers more countries and scripts and returns the evidence behind each answer, at a price per name.
Every code sample below was run before publishing. Pick by the names you actually have, not by which approach looks simplest.
Option 1: gender-guesser (offline, free)
gender-guesser is a Python package wrapping a dictionary of about 40,000 first names compiled by Jörg Michael in 2007–2008. It runs locally and needs no key.
pip install gender-guesser
import gender_guesser.detector as gender
detector = gender.Detector(case_sensitive=False)
print(detector.get_gender("Andrea")) # female
print(detector.get_gender("Andrea", "italy")) # male
print(detector.get_gender("Leslie")) # mostly_female
print(detector.get_gender("Avery")) # andy
print(detector.get_gender("Somchai")) # unknown
It returns one of six strings: male, female, mostly_male, mostly_female, andy (androgynous) and unknown. Three things to know before relying on it:
- Case matters by default. Without
case_sensitive=False,"andrea"returnsunknown. Data exported from forms is often lower-case. - Country is a name, not a code. The second argument takes the package's own country names (
"italy","great_britain","usa"), so map your ISO codes first. Without a country,Andreacomes back female; with"italy", male. - The data is from 2008 and Latin-script only. Names that became common since then, and any name written in another script, come back
unknown. The package is licensed under the GPLv3, which matters if you ship it inside proprietary software.
How it did on 4,610 names
We ran gender-guesser 0.4.0 on the same labelled fixture we use for our own benchmark: 4,610 first names from 25 countries. mostly_male and mostly_female were counted as answers; andy and unknown as no answer. No country was passed.
| Names | gender-guesser coverage | Correct when answered | Correct end to end |
|---|---|---|---|
| Latin script (3,195) | 83.0% | 98.2% | 81.4% |
| Arabic, Chinese, Korean, Cyrillic, Greek, Hebrew, Thai (1,412) | 0% | — | 0% |
| All 4,610 | 57.5% | 98.2% | 56.4% |
When it answers, it is very accurate: 98.2% on Latin-script names. The limit is coverage. If your names are mostly Western and you can treat a missing answer as unknown, it is a reasonable free starting point. If your data includes customers from Asia, the Middle East or Eastern Europe in their own scripts, about half your rows will come back empty. The measurement script is reproducible and listed at the end of this page.
Option 2: build a lookup from official name counts (pandas)
Several countries publish how many babies were registered with each first name and sex. The US Social Security Administration's baby names data is the best known: one file per year from 1880, with every name given to at least five babies of one sex. Download names.zip from that page in a browser, unzip it into a folder called names, and build a table:
import glob
import pandas as pd
births = pd.concat(
pd.read_csv(path, names=["name", "sex", "count"])
for path in glob.glob("names/yob*.txt")
)
totals = births.pivot_table(
index="name", columns="sex", values="count", aggfunc="sum", fill_value=0
)
totals["records"] = totals["F"] + totals["M"]
totals["female_share"] = totals["F"] / totals["records"]
def gender_from_ssa(name, threshold=0.9, min_records=100):
"""Return (gender, female_share, records); gender is None when unsure."""
key = name.strip().capitalize()
if key not in totals.index:
return None, None, 0
row = totals.loc[key]
records = int(row["records"])
share = float(row["female_share"])
if records < min_records:
return None, share, records
if share >= threshold:
return "female", share, records
if share <= 1 - threshold:
return "male", share, records
return None, share, records
The function returns the evidence, not just a label, and refuses to answer below a threshold or a minimum count. Keep both: a split name such as Leslie should come back unknown, not as whichever gender is slightly ahead.
This approach gives you real counts and full control, and it is free. Its limits are the limits of the data: one country, names with at least five births in a year, and no accents or non-Latin characters (the SSA files are ASCII). Summing every year since 1880 also weights old usage heavily; for a list of current customers you may want to restrict to recent decades. Other countries publish comparable files, including England and Wales (ONS), France (INSEE) and Canada; combining them means reconciling different spellings, periods and licences, which is most of the work.
Option 3: a name-gender API
An API is the choice when your names span countries and scripts, when you need the country to change the answer, or when you want the evidence returned with each result. NameGender has a Python client:
pip install namegender-client
import os
from namegender import NameGender
client = NameGender(os.environ["NAMEGENDER_API_KEY"])
result = client.name("Andrea", country="IT")
print(result["gender"], result["probability"], result["sample_size"], result["source"])
probability is the observed share of the dominant gender (0 to 100), sample_size is the number of counted records behind it, and source says where the answer came from. Italy publishes no name counts, so Andrea in Italy comes back male with a sample_size of 0 and a confidence of unverified; the same name with country="US" comes back female, backed by counted US birth registrations. The guide to probability, confidence and sample_size explains how to combine the fields.
Enrich a DataFrame in bulk
For a column of names, send unique name and country pairs in batches of up to 100, then merge the results back. Repeated names are looked up once.
import pandas as pd
df = pd.DataFrame({
"first_name": ["Andrea", "Andrea", "Jordan", "Ayşe", "Andrea"],
"country": ["IT", "US", "US", "TR", "IT"],
})
rows = []
keys = df[["first_name", "country"]].drop_duplicates()
for country, group in keys.groupby("country"):
names = group["first_name"].tolist()
for start in range(0, len(names), 100):
chunk = names[start:start + 100]
response = client.bulk(chunk, country=country)
for name, item in zip(chunk, response["results"]):
rows.append({
"first_name": name,
"country": country,
"gender": item["gender"],
"probability": item["probability"],
"sample_size": item["sample_size"],
})
df = df.merge(pd.DataFrame(rows), on=["first_name", "country"], how="left")
client.bulk() returns the whole response: a summary and a results list in the same order as the names you sent. An unknown result has gender set to None; keep it as a missing value rather than filling it with a default. For a one-off file, the dashboard also takes a CSV or XLSX upload, which may need no code at all (adding a gender column to a spreadsheet).
On the same 4,610-name fixture, with the country passed, NameGender answered 94.3% of names and 91.1% were correct end to end; on Latin-script names alone, 95.5%. It is weaker on some scripts than others: Thai names were correct end to end 49% of the time and Hebrew 64%. How accurate is gender prediction from a name? has the full breakdown. Accounts get 500 free credits on signup and 100 a day after that, enough to test on your own data first.
Other APIs work the same way from Python. Genderize, for example, takes a plain HTTP request:
import requests
response = requests.get(
"https://api.genderize.io",
params={"name": "andrea", "country_id": "IT"},
timeout=10,
)
print(response.json()) # name, gender, probability (0 to 1), count
Its probability is a fraction rather than a percentage, and count plays the role of sample_size. Gender API comparison measures several services on the same 991 names.
Which one to use
| Your situation | Use |
|---|---|
| Mostly Western first names, no budget, offline | gender-guesser, with case_sensitive=False and unknowns kept |
| One country with published counts, need full control | Your own table from that country's birth records |
| Names from many countries or scripts, or the country should change the answer | An API that takes a country and returns its evidence |
| A research study | Any of the above, validated on a labelled sample of your own data |
Whichever you choose, check it on 200 rows where you already know the answer, report how many names it left unknown, and do not use the result to decide anything about an individual. For studies, the guidelines for gender inference in research cover thresholds and what to report.
Common mistakes
- Treating
unknownas an error. It is the correct answer for a name without enough evidence. Store it as missing, not as"male"or an empty string. - Sending full names to a first-name lookup.
"Maria Garcia"is not a first name. Split the name first, or use an endpoint that parses full names; splitting full names covers the edge cases. - Ignoring the country.
Andrea,JeanandKimchange gender between countries. Names that change gender by country lists the common ones. - Looping one request per row. Deduplicate first and use bulk calls; a customer list of 200,000 rows usually contains far fewer unique first names.
The gender-guesser figures on this page were measured on 7 October 2026 with gender-guesser 0.4.0 and scripts/benchmark-gender-guesser.py against tests/Fixtures/international_accuracy.csv, the fixture published with our benchmark methodology.
Every claim on this page is measurable against your own list. The free tier is enough to check it.
Related
-
Gender inference from names in research: guidelines for accuracy and reporting
Inferring gender from names in a study: validate on your data, fix a threshold, report unknowns, check bias by script and country, and what to write in methods.
-
Is this name male or female? How to check, and when the answer is a coin flip
Check whether a first name is male or female in three steps. Real birth-record numbers for Riley, Robin, Kim, Jean and Andrea show when to trust the answer.
-
How accurate is gender prediction from a name? Measured by script and by probability
Measured on 4,610 labelled names from 25 countries: 94% answered, 97% of answers correct, 91% end to end. Results by script, and where the errors come from.