Continue → Overview

Japanese name gender

This is the weakest market we publish numbers for, and the sample behind those numbers is small. Both facts are on this page because you need them before an integration, not three weeks into one.

100 free credits every day. No card.

Names measured
51
Coverage
98.0%
Accuracy when answered
78.0%
End-to-end
76.5%

Measured 28 August 2026 on a fixed, reproducible fixture. Not a population sample — a fixed input anyone can rerun against us or against a competitor.

Read the measurement before anything else

The figures above come from 51 Japanese names in a fixed fixture. That is a small sample and a wide confidence interval, so treat it as a signal of the right order of magnitude rather than a precise rate. Kana is measured on three names, which is too few to report at all.

We are publishing the page anyway because the alternative — a page implying parity with our European coverage — would be worse. If you process Japanese names at volume, measure on your own list before committing.

Why kanji is genuinely hard

A Japanese given name is written in kanji, but the reading is not determined by the characters. The same characters take multiple readings, and parents choose readings freely. Gender signal attaches to the reading at least as much as to the characters, and the reading is exactly what the written form does not tell you.

This is a different problem from Chinese. There, the character carries connotation reliably; here, an identical string can be two different names.

Real output, country hint JP — including a wrong one

The last row is the failure mode this market produces: a confident, wrong answer from a bridged reading. It is on the page because it is representative, not because it is rare.

Input Gender Probability Verdict
太郎 (Taro) male 95 correct
花子 (Hanako) female 95 correct
さくら (Sakura) female 95 correct
翔太 (Shota) male 95 correct
Yuki female 75 genuinely unisex; the 75 is the useful part
陽菜 (Hina) male 99 wrong — a female name returned as male at high confidence

What we would actually recommend

For segmentation and aggregate reporting, the current accuracy is usable if you keep the unresolved rows as their own bucket. For addressing individual people in Japanese, it is not — set a high probability threshold and route the rest to review, or use a name field the person filled in themselves.

Romaji input (Haruto, Sakura, Akiko) avoids the reading ambiguity and is often the more reliable path when your system stores it.

Questions

Why publish a page for a market you are weak in?

Because the alternative is a buyer discovering it mid-integration. A page that says 78% on 51 names is more useful than one that says nothing.

Is kana supported?

Kana input resolves, but only three kana names sit in the fixture, so we have no honest rate to publish for it.

Will this improve?

It needs real Japanese name data with counts, not a cleverer bridge. That is a sourcing problem and we have not solved it yet.

What about surname order?

Japanese names are also written surname first in Japanese order. Send the given name alone, or read the name field to see what was classified.

Related pages

Korean name gender

Predict gender from a Korean name in Hangul or romanised form. Measured coverage and accuracy, competing romanisation systems, and surname-first handling.

Chinese name gender

Predict gender from a Chinese name in Han characters or pinyin. Measured coverage and accuracy, the surname-first trap, and why characters beat romanisation.

Gender probability by name

Every name lookup returns a probability, a sample size and a confidence tier. What each one measures, how the tiers are set, and where to put your threshold.

Our coverage, script by script — including where we fail

A reproducible benchmark across 4,610 names, 25 countries and nine writing systems. Two scripts return almost nothing, and we publish those numbers too.

Chinese name gender: what a character can and cannot tell you

Coverage on Han names is 99.5% but accuracy is 86.9% — the inverse of the usual pattern. Why pinyin is weaker evidence, and the name-order trap.

Check it against your own list

Every number on this page is reproducible with a free key. If your data breaks it, that is the more interesting result.