Indian name gender
The strongest non-Latin market we measure. Romanised and Devanagari input both resolve, and end-to-end accuracy sits close to the European figure.
100 free credits every day. No card.
Measured 28 August 2026 on a fixed, reproducible fixture. Not a population sample — a fixed input anyone can rerun against us or against a competitor.
There is no single Indian naming system
India spans many languages, scripts and naming conventions. Some regions place the family name first, some use a patronymic, some use initials expanded from a village or father name, and many people are recorded in Latin script with spellings that vary between documents.
What makes this market work despite that is the given-name stock itself: it is large, distinctive, and strongly gendered. Priya, Ananya and Lakshmi are unambiguous in a way that Andrea and Jean are not.
Real output, country hint IN
Devanagari input is bridged to the Latin entry and reports source: script, so you can tell a direct match from a transliterated one. India publishes no registration counts, so these carry confidence unverified — accuracy is what the measurement above speaks to.
| Input | Gender | Probability |
|---|---|---|
| Priya | female | 95 |
| Rahul | male | 95 |
| Ananya | female | 95 |
| Arjun | male | 95 |
| प्रिया (Priya) | female | 95 |
| राहुल (Rahul) | male | 95 |
Initials and expanded names
South Indian records frequently carry initials — R. Kumar, K. S. Priya. Single letters are treated as initials and skipped, so the first full token is what gets classified. When the only full token is a surname or a place name, the answer is meaningless regardless of its confidence tier; the name field tells you when that happened.
Spelling variance is the main source of misses
Romanised Indian names vary widely between documents — Lakshmi, Laxmi, Lakshmy. Near-spellings resolve through fuzzy matching, and matched_as names the entry that was actually used, so you can audit whether the substitution was reasonable rather than trusting it silently.
That field is worth keeping in your output. A fuzzy match is a judgement the system made on your behalf, and it is usually right; when it is wrong it is wrong in a specific, checkable way, which is far better than a silent miss. If matched_as is populated on more rows than you expected, the fix is normalising your source spellings rather than raising a threshold.
Everything here returns confidence: unverified, because India publishes no name registration counts. Since that describes almost every market outside seven Western countries, a filter of "reject anything unverified" is not a quality bar — it is a filter that removes most of the world. Use the measured accuracy above and a probability threshold instead.
Questions
Which scripts are supported?
Latin and Devanagari resolve today, the latter through a script bridge reported as source: script. Other Indic scripts are not separately measured and should be tested on your own data first.
Should I pass country=IN?
Yes when you have it. It narrows the pool, though answers will carry confidence unverified because no registration counts are published.
How are surnames handled?
Titles and initials are stripped and the first full token is classified. Where the family name comes first, send the given name alone.
What about unisex Indian names?
They return a low probability rather than a forced pick. Filter on it.
Related pages
Predict gender from an Arabic name written in Arabic script or transliterated. How the consonant skeleton index works and why it beats guessing a spelling.
Predict gender from a Turkish name. Dotted and dotless i handling, the genuinely unisex names, and why correct Turkish answers are still marked unverified.
Every name lookup returns a probability, a sample size and a confidence tier. What each one measures, how the tiers are set, and where to put your threshold.
A reproducible benchmark across 4,610 names, 25 countries and nine writing systems. Two scripts return almost nothing, and we publish those numbers too.
Some ordinary names flip gender across borders, and the error is systematic rather than random. What the country parameter does, and where the data comes from.
Check it against your own list
Every number on this page is reproducible with a free key. If your data breaks it, that is the more interesting result.