Continue → Overview

Choosing a gender dataset for names

A row count is not enough to evaluate a dataset. Provenance, country granularity, observation counts, licensing and update behaviour decide whether the data can support your use case.

100 free credits every day. No card.

Define the delivery model first

A downloadable dataset gives you local control, predictable latency and the ability to run offline. It also makes you responsible for normalisation, matching, updates, provenance and deletion policy. An API keeps those parts managed but adds a network dependency and usage accounting.

There is a practical middle path for one-time enrichment: upload a CSV or XLSX file, map its columns and download the result. That avoids building an integration while preserving row-level probability, evidence and source fields in the output.

NameGender currently exposes the maintained data through lookup and batch products rather than promising an unrestricted raw database download. If raw redistribution rights are essential, establish that licensing requirement before evaluating accuracy or price.

Inspect provenance at row level

A useful dataset identifies where a result came from. Counted national registration sources can support sample-size claims; a global source without counts can extend geographic coverage but cannot justify the same evidence label. Blending both into one unexplained confidence score hides that difference.

Ask whether the license permits your exact use: internal enrichment, model training, resale, publication and redistribution are different rights. Open availability on the web does not automatically allow a vendor to repackage the underlying data or allow you to export it onward.

Version information matters as much as source names. Record a data version or processing date with every enrichment so a later correction can be scoped and rerun rather than debated from memory.

Benchmark the populations you hold

Large Western registration datasets can make an aggregate benchmark look strong while masking poor coverage in Arabic, Indic or East Asian scripts. Build a fixture that reflects your own customer distribution and report each country or script separately.

Measure coverage, accuracy when answered and end-to-end accuracy. A conservative dataset can appear accurate by returning unknown on every difficult row; an aggressive one can appear complete by guessing. Neither single percentage describes the tradeoff.

Include spelling variants, diacritics, ambiguous names, empty cells and deliberate non-names. Then repeat the same fixture after a data update. That is the minimum needed to know whether a dataset improved, regressed or merely grew in row count.

Questions

Can I download the complete NameGender database?

The maintained product is currently delivered through the API and managed batch processing. Contact support if your use requires a separately licensed data delivery.

What fields should a gender dataset include?

At minimum: normalized name, country or market, result, probability, evidence count, source and version.

Is a larger dataset always more accurate?

No. Duplicate records, uneven country coverage and unknown provenance can increase row count without improving decisions.

How should I compare two datasets?

Run the same labelled fixture against both and report coverage, accuracy when answered and end-to-end accuracy by market.

Related pages

The name gender database

What is actually in the name-to-gender dataset: seven counted national registries, WGND 2.0 for the rest, versioned snapshots and a per-answer source field.

Customer and lead gender enrichment

Append a gender column to a customer list, CRM export or lead database, with a confidence field you can filter on and no monthly credit expiry.

Bulk gender detection from CSV and Excel

Upload a CSV or XLSX name list and get a gender column back, with confidence fields, automatic deduplication and the cost shown before charging.

Our coverage, script by script — including where we fail

A reproducible benchmark across 4,610 names, 25 countries and nine writing systems. Two scripts return almost nothing, and we publish those numbers too.

How to compare name-to-gender APIs without trusting anyone's marketing

What to measure before you pick a vendor, why a single accuracy percentage tells you nothing, and how to run the same test on all of them — including us.

Check it against your own list

Every number on this page is reproducible with a free key. If your data breaks it, that is the more interesting result.