Gender inference from names in research: guidelines for accuracy and reporting
Inferring gender from first names is a defensible research instrument for aggregate questions, such as the share of women among first authors in a field, if five conditions hold: you measure its accuracy on a labelled sample of your own corpus, you fix a probability threshold before looking at the results, you report how many names were left unknown, you check whether errors cluster in particular countries or scripts, and you never use the result to decide anything about an individual. Most criticism of name-based studies is criticism of a study that skipped one of these steps.
This page sets out each step, with the numbers to report and a template for the methods section. We run a name-to-gender API, so some examples use our published measurements; the guidelines apply whichever tool or dataset you use.
What a name-based estimate measures
A name-based estimate answers one question: in the population whose records the tool draws on, what share of people with this first name were recorded as male or as female? It is a statement about how a name has been used, not about how the person in your dataset identifies.
Three consequences follow, and they belong in your limitations section:
- The categories are binary. Birth registers and most name lists record two sexes. Non-binary people are not represented, and a name-based method cannot represent them.
- The estimate is about a population, not a person. A name that is 98% female in national records is still male for two people in a hundred. Individual errors are expected; the method is valid only when they average out or can be bounded.
- The population matters. The same name can lean opposite ways in different countries.
Andreais mostly female in the United States and male in Italy;Jeanis male in France and mostly female in the United States. A tool that ignores country answers a different question from the one you are asking. Names that change gender by country lists the common cases.
Use it for aggregate questions only
Appropriate uses compare groups: gender shares among authors, reviewers, grant applicants, inventors or speakers, and how those shares change over time or between fields. The research question is about a distribution, and the analysis can carry the uncertainty of the instrument.
Inappropriate uses act on individuals: screening, selection, targeting, or any decision about a named person. A study design that requires knowing a specific person's gender should ask that person. When the stakes for individuals are real, self-report is the method; inference is not a substitute.
Guideline 1: know where the evidence comes from
Tools differ more in their evidence than in their algorithms. Before choosing one, find out:
- Source of the gender association. Official birth or population registers with counts, social-media profiles, scraped name lists, or a statistical model. Registers are the most transparent; model outputs are the hardest to audit.
- Whether the evidence is counted. "1,900 of 2,000 people with this name were recorded as female" is a different claim from "this name appears on a list of female names". Ask whether the tool returns the number of records behind each answer.
- Coverage by country and script. A tool built from European and North American registers will answer Chinese, Thai or Hebrew names less often or less accurately. Look for published results by script, not a single accuracy figure.
NameGender returns, with every answer, the observed share (probability), the number of counted records (sample_size), a confidence tier and the source layer. Counted data comes from official name statistics in eleven countries; elsewhere the evidence is a list of attested names without counts, and those answers are marked unverified. The guide to probability, confidence and sample_size explains each field.
Guideline 2: validate on your own corpus
A vendor's accuracy figure describes the vendor's test set. Your corpus has its own mix of countries, eras and name formats, so measure on it.
- Draw a random sample of 200 to 500 records from the corpus you will analyse.
- Establish the gender of each by an independent route: self-identification on a profile page, pronouns in a biography, or a published list. Record the route and any records you could not label.
- Run the sample through the tool with exactly the settings you will use in the full analysis: the same country hints, the same threshold.
- Report three numbers: coverage (names answered / names submitted), accuracy when answered (correct / answered) and end-to-end accuracy (correct / submitted).
Sample size sets how precisely you know the accuracy. At 90% accuracy, the 95% confidence interval is roughly ±4.2 points with 200 labelled names and ±2.9 points with 400 (1.96 × √(0.9 × 0.1 / n)). If your comparison between groups hinges on a difference of a few points, label more.
All three numbers are needed. A tool can reach 99% accuracy when answered by refusing every difficult name, and 100% coverage by guessing on every name. How accurate is gender prediction from a name? shows the three measures side by side on 4,610 names: 94.3% coverage, 96.6% accuracy when answered, 91.1% end to end.
Guideline 3: fix a threshold before you see the results
Every tool returns some answers with weak evidence. Decide in advance which answers you will accept, write the rule down, and apply it unchanged. Choosing the threshold after seeing which one gives the clearest result is a form of p-hacking.
A threshold only works if the tool's probability means what it says. Check that answers at higher probability are right more often. In our latest benchmark run, on 7 October 2026:
| Returned probability | Answers | Actually correct |
|---|---|---|
| 50 to 69 | 79 | 60.8% |
| 70 to 79 | 54 | 74.1% |
| 80 to 89 | 96 | 80.2% |
| 90 to 99 | 2,629 | 97.2% |
| 100 | 1,489 | 99.2% |
On those figures, accepting only answers at 90 or above keeps 95% of the answers and lowers the error rate among them to 2.1%. Common choices in published work are 0.8, 0.9 and 0.95. Higher thresholds trade coverage for accuracy, and the names that drop out are not random: they are disproportionately unisex names and names from countries with weaker data.
If the tool reports how many records stand behind an answer, add a minimum count as well. A probability of 100 resting on three records is weaker evidence than a probability of 96 resting on 40,000.
Guideline 4: report unknowns and bound their effect
Names below the threshold, names the tool does not know and records with only an initial all end up unclassified. Do not drop them silently. Report how many there are, and show how much they could move your result.
A simple bound: suppose 1,000 authors, of whom 850 are classified and 340 of those as women. The share of women among classified authors is 40%. If all 150 unclassified authors were men, the true share would be 34%; if all were women, 49%. If your conclusion survives the whole range, say so. If it does not, the unknowns are the finding, and more labelling is the next step.
Report the bound, or a sensitivity analysis that reassigns unknowns in proportion to the classified share within the same country, alongside the headline estimate.
Guideline 5: check whether errors cluster
Misclassification is rarely evenly spread. It concentrates where the tool's data is thin, and that is often correlated with the groups being compared. If one field has more authors with East Asian names than another, a tool that is weaker on those names will bias the comparison between fields, not just add noise.
Our own results show how large the spread can be. On the same benchmark, Latin-script names were correct end to end 96% of the time, Chinese names in Han characters 91%, Hebrew 64% and Thai 49%. The full table is in coverage by writing system. Other tools have different weak spots, which is why the check has to be run on yours.
In practice: stratify the validation sample by country or script, report coverage and accuracy within each stratum, and say which strata are too small to support a conclusion.
Guideline 6: handle name formats deliberately
- Initials. "J. Smith" carries no gender signal. Report the share of records with initials only; in older bibliographic data it can be large and is not random across fields or eras.
- Name order. Chinese, Japanese, Korean and Hungarian names are often written surname first. Send the given name, not the first token. If your data mixes orders, parse explicitly; splitting full names covers the common cases.
- Transliteration. A romanised Chinese given name carries much less gender information than the original characters, because many characters share one romanisation. Use the original script when your source has it.
- Country context. If your records include affiliation or nationality, pass it as a country hint. If they do not, say that results are worldwide estimates and name the countries where that is likely to matter.
Guideline 7: make it reproducible
Name data changes. Registers publish new years, tools add sources, and answers change with them. Record enough for someone else to rerun the classification:
- the tool or dataset, and its version: for NameGender, the
data_versionreturned with each response; - the date of the run;
- the input field used (given name, full name), and how names were parsed;
- country hints, and where they came from;
- the threshold and any minimum record count;
- the validation sample, its labelling route and its results.
Keep the raw responses, not only the final labels, so the threshold can be changed later without rerunning.
Guideline 8: treat names as personal data
Author lists, survey responses and administrative records link names to identifiable people. Send only the field the lookup needs, check your ethics approval and data-protection basis cover sending it to a third-party service, and prefer tools that do not keep what you send. NameGender API keys can be set to not store the names they send (how that works).
Publish aggregates, not individual inferred labels. A released dataset with an inferred gender attached to each named author invites exactly the individual-level use that the method does not support.
Reporting checklist
| Item | What to report |
|---|---|
| Construct | Name-associated gender in a reference population; binary; not gender identity |
| Tool | Name, version or data version, date of run |
| Input | Field used, parsing of name order, handling of initials |
| Context | Country hints and their source, or a statement that none were used |
| Threshold | Probability cut-off and minimum record count, fixed before analysis |
| Validation | Sample size, labelling route, coverage, accuracy when answered, end-to-end accuracy |
| Unknowns | Number and share unclassified, with a bound or sensitivity analysis |
| Bias | Coverage and accuracy by country or script, and strata too small to interpret |
| Ethics | Data-protection basis, retention by the tool, aggregate-only publication |
A template for the methods section
We inferred gender from first names using NameGender (data version 2026.10, accessed [date]). For each author we submitted the given name, parsed from the full name with surname-first order for [countries], and the country of the first listed affiliation as a hint. We accepted classifications with a reported probability of at least 0.90 and treated all others, and authors listed with initials only, as unclassified. On a random validation sample of [n] authors whose gender we established from [route], coverage was [x]%, accuracy among classified names [y]% and end-to-end accuracy [z]%; results by world region are in Table S1. [k]% of authors in the full corpus were unclassified. Assigning all unclassified authors to either gender moves the estimated share of women from [a]% to [b]%. Name-based inference reflects how names are used in the reference population and does not capture gender identity or non-binary genders.
Replace the bracketed parts, and the tool details if you use another service.
Further reading
The literature on name-based gender inference is mostly from bibliometrics and computational social science. Useful starting points:
- Santamaría and Mihaljević (2018), Comparison and benchmark of name-to-gender inference services, PeerJ Computer Science. Compares several services on one labelled dataset and discusses how to tune thresholds.
- Karimi, Wagner, Lemmerich, Jadidi and Strohmaier (2016), Inferring gender from names on the web: a comparative evaluation of gender detection methods. Shows that accuracy varies considerably by country of origin.
- Wais (2016), Gender prediction methods based on first names with genderizeR, The R Journal. A worked approach to choosing a threshold against a labelled sample.
- Mihaljević, Tullney, Santamaría and Steinfeldt (2019), Reflections on gender analyses of bibliographic corpora, Frontiers in Big Data. Discusses the conceptual limits of the method.
- Lockhart, King and Munsch (2023), Name-based demographic inference and the unequal distribution of misrecognition, Nature Human Behaviour. On who is misclassified, and why that matters for conclusions.
Running this with NameGender
Researchers at universities and non-profit institutes can apply for 100,000 free credits for non-commercial studies. Our benchmark and methodology pages publish the fixtures behind the figures on this page, including the scripts where we do badly, so you can cite them or rerun them.
Every claim on this page is measurable against your own list. The free tier is enough to check it.
Related
-
How to get gender from a name in Python: gender-guesser, name data and APIs compared
Three working ways to get gender from a first name in Python: gender-guesser, baby-name counts with pandas, and an API, measured on the same 4,610 names.
-
Is this name male or female? How to check, and when the answer is a coin flip
Check whether a first name is male or female in three steps. Real birth-record numbers for Riley, Robin, Kim, Jean and Andrea show when to trust the answer.
-
How accurate is gender prediction from a name? Measured by script and by probability
Measured on 4,610 labelled names from 25 countries: 94% answered, 97% of answers correct, 91% end to end. Results by script, and where the errors come from.