How Typographical Errors in Japanese Standards Created Unicode's 'Ghost Characters'
A legacy Japanese character set established in 1978 introduced phantom kanji that still inhabit global software systems.

In 1978, Japan's Ministry of Economy, Trade and Industry formulated a national character encoding standard that later became designated as JIS X 0208. The specification served as the foundational blueprint for modern Japanese digital text handling. However, following its release, software engineers and linguists identified several kanji characters embedded within the system that possessed no known definitions, historical records, or established pronunciations. These unexplained additions came to be referred to as ghost characters, or 幽霊文字.
For nearly two decades, the anomalous characters persisted in regional software standards without clear explanation. As reported by Hacker News, a formal inquiry was eventually launched in 1997 to determine the origin of the mysterious entries. While standard rules mandated that each added character maintain a recorded source, the archived documentation often provided minimal detail, frequently citing entire reference works without specific page citations.
A primary reference work cited in the standard's documentation was the "Overview of National Administrative Districts" (国土行政区画総覧), an exhaustive catalog of Japanese geographical place names. Researchers attempting to verify character entries faced significant logistical challenges, as the publication comprised a seven-volume set totaling approximately 6,300 pages. Finding individual characters across thousands of unindexed entries required exhaustive manual verification.
The 1997 investigation successfully identified the origins of most ghost characters by conducting direct interviews with the catalogers who drafted the 1978 benchmark. The inquiry established that a substantial portion of the unknown characters had been created accidentally through physical document processing mistakes and typographical misinterpretations during the standard's assembly.
One notable error involved the character 妛, which resulted from a physical layout defect while attempting to record a place name containing the character component "山 over 女." Because system operators were unable to print the combined glyph as a single entity, the two sub-components were printed separately, manually cut out, and pasted onto a submission page before being photocopied. During the copying process, the physical seam where the paper edges met created a dark line that was misidentified as an intentional stroke, creating an entirely new character. The legitimate historic character, 𡚴, was omitted from the JIS standard and was not integrated into digital character tables until years later.
Of all the ghost characters examined during the investigation, exactly one entry lacked both a documented administrative source and any identifiable historical antecedent: 彁. Investigators concluded that the character most likely originated from a misreading of the existing character 彊, though no specific document tracing the precise point of error could be located.
When global computing bodies subsequently developed the Unicode standard, existing national character sets—including JIS X 0208—were incorporated to preserve backward compatibility. As a result of this integration, as well as separate character unification processes across Chinese, Japanese, and Korean text standards, these typographical artifacts remain permanently coded into modern operating systems and digital character sets worldwide.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



