katakana
Enable an AI agent to recognise katakana, record how its signs function in context, assess reading and representation uncertainty, and choose transformations that preserve intended meaning.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
Researched by: Codex + Grok
Purpose and description
Enable an AI agent to recognise katakana, record how its signs function in context, assess reading and representation uncertainty, and choose transformations that preserve intended meaning.
Katakana is a Japanese mora syllabary of predominantly angular graphemes, historically isolated from fragments of Chinese characters, that in modern Japanese orthography functions as the marked phonetic script for loanwords, foreign and scientific names, onomatopoeia, and emphasis, in complementary distribution with hiragana and kanji.
It can be Locate katakana spans and flag uncertain boundaries in mixed-script material; Transcribe katakana while preserving small forms, voicing marks, and uncertain characters; Propose reading candidates and identify the context needed to distinguish them; Compare spelling variants under an explicitly chosen equivalence rule; Convert width or script representation while retaining the original and reporting losses; Check a spelling against a stated orthographic convention without treating unfamiliar forms as errors.
Distinguishing features
Identify characters against a katakana repertoire rather than angular appearance alone; hiragana and unrelated symbols may share visual features.
Distinguish lookalikes such as katakana ロ, kanji 口, and Latin or numeric shapes using character identity when available and contextual evidence when it is not.
Check small versus full-size forms, such as ャ versus ヤ and ッ versus ツ, because size can change how a sequence is read.
Treat the prolonged sound mark ー as a contextual sign rather than sufficient evidence that surrounding text is katakana.
Do not infer katakana solely from a word being foreign: establish script identity independently from the reason that script was chosen.
Scope
+ Identification of katakana characters and sequences within mixed-script text
+ Full-size and small kana, voicing marks, and prolonged-sound notation
+ Relationships between written katakana sequences and intended readings
+ Contextual uses such as loanword spelling, sound symbolism, names, and emphasis
+ Width variants, encoding, rendering, and transformations affecting katakana identity
- Japanese grammar and the general interpretation of sentences
- Hiragana and kanji repertoires except where needed to distinguish or relate scripts
- The identity and attributes of people, places, products, or organisms named in katakana
- Loanword etymology and borrowing history beyond evidence needed to explain a spelling
- Complete models of languages that use katakana-based orthographies
- General font design, document layout, and text-encoding infrastructure
Characteristics
- Script membership
- katakana character; contextual shared sign; other script; unresolved Separates katakana identity from adjacent punctuation and visually similar characters.
- Observed written form
- Exact text sequence or image-backed transcription, retaining uncertainty Provides a recoverable original before interpretation or normalization.
- Kana size
- full-size; small; ambiguous; not applicable Distinguishes sequences whose readings differ despite similar constituent shapes.
- Voicing-mark attachment
- Base kana linked to dakuten or handakuten; absent; unresolved attachment Prevents detached or misread marks from silently changing the transcription.
- Width representation
- fullwidth; halfwidth; mixed; unresolved Supports comparison and display adaptation while retaining the source form.
- Intended reading
- Written span linked to one or more reading candidates with contextual evidence Avoids assuming that character recognition alone settles pronunciation.
- Script-choice function
- loanword notation; name notation; sound symbolism; emphasis; specialist convention; other; unknown Determines whether changing scripts would remove a communicative distinction.
- Interpretation status
- unambiguous; multiple supported readings; transcription uncertain; convention unresolved Controls when an agent may transform text automatically and when it should preserve alternatives.
Also called
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 5 bundles · 9 layers · 16 findings · 27 questions.
Katakana recognition Establish which observed signs and spans belong to katakana.
An agent must distinguish script identity from visual resemblance and surrounding language before interpreting or editing text.
Character identity
Resolve the identity of individual signs using available textual or visual evidence.
Kana versus lookalike
Record whether an apparent katakana sign is identified by encoded text, inferred from an image, or still ambiguous.
- What evidence identifies this sign as katakana rather than a similar kanji, Latin character, numeral, or symbol? definition
- Which source text or image supports the identification, and what ambiguity remains? provenance
Katakana span boundaries
Determine which adjoining signs participate in a katakana expression.
Shared-sign membership
Record how prolonged sound marks, separators, punctuation, and neighbouring scripts relate to an identified katakana span.
- Does each adjacent mark participate in the katakana spelling, separate its components, or belong to surrounding text? boundary
- Where can this span be selected or replaced without detaching a reading-relevant mark or splitting a mixed-script name? action
Kana composition and reading Represent how katakana signs combine and what readings those combinations support.
Small kana, voicing marks, and sound-length notation prevent reliable interpretation through isolated character substitution.
Size and voicing
Identify distinctions carried by kana size and attached marks.
Reading-relevant form
Preserve full-size versus small kana and the presence and attachment of dakuten or handakuten.
- Which signs are small kana, and which base kana carry dakuten or handakuten? definition
- Would the proposed transcription or replacement change a small form, drop a mark, or attach a mark to the wrong kana? action
Sequence reading
Relate complete written combinations to contextually supported readings.
Combination and length
Record how combinations such as キャ or ティ, small ッ, and prolonged sound marks contribute to a candidate reading.
- How does the applicable orthographic convention interpret each small-kana combination, small ッ, and prolonged sound mark? definition
- What linguistic, lexical, or speaker-provided evidence supports the intended reading of the complete span? provenance
- Can a reading be produced reliably, or must the agent retain alternatives and request context? action
Katakana use and convention Explain the local function of katakana and the convention against which its spelling should be assessed.
Katakana does more than spell foreign words; its use may carry emphasis, naming choices, or specialist distinctions that transformations must preserve.
Script-choice function
Identify why a particular expression appears in katakana.
Contextual role
Record supported interpretations of script choice without inferring a word's origin or status from katakana alone.
- Is katakana being used for loanword notation, a name, sound symbolism, emphasis, a specialist convention, or another supported purpose? definition
- What distinction would disappear if this expression were rewritten in hiragana, kanji, or Latin letters? action
Spelling authority
Identify the relevant language, period, community, and source of spelling authority.
Variant or error
Assess an observed spelling relative to an identified convention, including deliberate or authoritative name spellings.
- Which language, historical period, specialist practice, or house style governs this katakana occurrence? boundary
- Which source establishes the preferred spelling, particularly for a person's name, product, or quoted expression? provenance
- Does the evidence justify correction, presentation as a variant, or preservation with an unresolved note? action
Katakana representation and transformation Control changes to encoded or displayed katakana while preserving relevant distinctions.
Width conversion, mark handling, search normalization, and transliteration can alter text identity or conceal uncertainty.
Encoded and visible form
Relate the underlying text representation to the katakana a reader sees.
Width and mark integrity
Record fullwidth or halfwidth forms, separate or composed marks, and rendering problems that affect recognition.
- What width forms and voicing-mark representations occur in the original sequence? definition
- Does the displayed result preserve readable small kana and correctly associated marks in the target font and layout? measurement
Controlled equivalence
Define which distinctions a particular comparison or conversion may discard.
Conversion contract
Specify whether an operation seeks exact identity, normalized matching, reading equivalence, or a different script representation.
- For this operation, must width, kana size, sound-length spelling, and script-choice distinctions remain distinguishable? boundary
- Which conversion rule or romanization convention applies, and where does it require a reading decision? definition
- Can the original be recovered after conversion, and what lost distinctions or unresolved readings must accompany the result? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Kinds and varieties
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Modern fullwidth gojūon katakana (the everyday printed syllabary, including ン)
- Dakuon and handakuon forms (voiced and semi-voiced signs marked with dakuten or handakuten)
- Yōon combinations (palatalized two-sign spellings such as キャ, シュ, チョ)
- Small kana / sutegana (ァィゥェォッャュョ and related size-reduced letters, including the sokuon ッ)
- Chōon usage with the long-vowel mark ー, characteristic of katakana spelling
- Extended katakana for non-Japanese phonemes (ファ, ティ, トゥ, ドゥ, シェ, チェ, ヴ, ツァ, and similar)
- Halfwidth (hankaku) katakana as a legacy computing and encoding repertoire
- Ainu and other phonetic-extension katakana (small combining letters in the Unicode Katakana Phonetic Extensions set)
- Which of these kinds and varieties hold for the sense of katakana this model covers, and on what evidence? provenance
Identifiers and schemes
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Wikidata - Q82946 - Item for katakana as a writing system / syllabary.
- ISO 15924 - Kana (411) - Four-letter code Kana, numeric 411; related codes are Hira (hiragana), Hrkt (Japanese syllabaries), and Jpan (Han + hiragana + katakana).
- Unicode script - Katakana (ISO alias Kana) - Unicode Script property value matching ISO 15924 Kana.
- Unicode block - U+30A0-U+30FF - Main Katakana block, including the double-hyphen, prolonged-sound mark, and small letters.
- Unicode block - U+31F0-U+31FF - Katakana Phonetic Extensions, chiefly small letters for Ainu and similar phonetic use.
- Unicode halfwidth katakana - U+FF66-U+FF9D (plus related halfwidth voiced marks in the Halfwidth and Fullwidth Forms block) - Legacy hankaku katakana corresponding to the JIS X 0201 katakana set.
- IETF BCP 47 script subtag - Kana - Used as in ja-Kana for Japanese text written in katakana.
- JIS X 0208 row - ku 05 (row 5) for fullwidth katakana in the 1978/1983/1990 coded character set - Traditional Japanese industrial coded-character addressing, distinct from Unicode scalars.
- Which of these identifiers and schemes hold for the sense of katakana this model covers, and on what evidence? provenance
Standards and regulation
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- ISO 15924 (ISO) - script code Kana / 411 for katakana.
- ISO/IEC 10646 and the Unicode Standard (ISO/IEC JTC 1/SC 2 and the Unicode Consortium) - coded character identity and rendering classes.
- JIS X 0201 (Japanese Industrial Standards Committee) - 8-bit set that includes halfwidth katakana.
- JIS X 0208 and JIS X 0213 (Japanese Industrial Standards Committee) - fullwidth katakana in the Japanese coded kanji sets.
- 現代仮名遣い, Cabinet Notification of 1986 (Government of Japan / Agency for Cultural Affairs) - modern kana orthography, which katakana spelling follows where applicable.
- 外来語の表記, Cabinet Notification No. 1 of 1991 (Government of Japan / Agency for Cultural Affairs) - official loanword spelling in katakana.
- MEXT Course of Study for Japanese language (Ministry of Education, Culture, Sports, Science and Technology) - school order and inventory of katakana taught in compulsory education.
- Which of these standards and regulation hold for the sense of katakana this model covers, and on what evidence? provenance
Real-world use
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Japanese newspapers, packaging, menus, and advertising write gairaigo, foreign personal and place names, and many product names in katakana.
- Japanese biological and technical prose often sets species names, chemical names, and similar labels in katakana rather than hiragana.
- Manga, games, and signage use katakana for onomatopoeia, robot or foreigner speech, and emphasis analogous to italics.
- Ainu publications and signage in Hokkaidō use standard plus extended katakana as a practical orthography.
- Legacy Japanese computing, POS terminals, and some forms still store or display hankaku katakana from the JIS X 0201 repertoire.
- Historical telegrams and early data systems in Japan were commonly katakana-only.
- Learners encounter katakana as the second syllabary after hiragana, then as the default script for decoding loanword vocabulary.
- Which of these real-world use hold for the sense of katakana this model covers, and on what evidence? provenance
Typical measurements
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Basic gojūon inventory (including ン) - 46 - letters
- Stroke count of a single modern katakana letter - 1-6 (most letters 2-4) - strokes
- Typographic advance of a fullwidth katakana letter - 1 - em
- Typographic advance of a halfwidth katakana letter - 0.5 - em
- Main Unicode Katakana block size - 96 (U+30A0-U+30FF) - code points
- Morae encoded by one basic letter (excluding ッ and ー, which modify a neighbouring mora) - 1 - mora
- Which of these typical measurements hold for the sense of katakana this model covers, and on what evidence? provenance
Failure modes and hazards
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Homoglyph confusion with kanji or hiragana of similar shape (カ/力, ロ/口, ニ/二, タ/夕, ヘ/へ, ツ/ッ), causing OCR, identity, and proofreading errors.
- Halfwidth-katakana mojibake when JIS X 0201, Shift_JIS, and UTF-8 are mixed, a long-standing Japanese data-interchange failure.
- Lossy transcription of foreign phonemes (notably English /l/ vs /r/, /v/, interdental fricatives), producing unstable or competing spellings.
- Omission or mis-sizing of ッ versus ツ, which changes gemination and can change the word.
- Confusion of the chōonpu ー with a hyphen, minus, or dash in horizontal and especially vertical text.
- The so-called katakana wall: dense loanword text that is legally Japanese but opaque to many readers.
- Security and accessibility issues when visually similar mixed-script strings (katakana, kanji, Latin, Cyrillic) are treated as identical.
- Which of these failure modes and hazards hold for the sense of katakana this model covers, and on what evidence? provenance
Regional variation
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Standard Japanese school katakana versus Hokkaidō Ainu practice, which needs the phonetic-extension small letters.
- Ryukyuan writing sometimes uses katakana as a phonetic overlay or as a full orthography, not only for loanwords.
- Loanword spelling drift inside Japan (コンピュータ vs コンピューター, バイオリン vs ヴァイオリン), with older and newer Cabinet-aligned forms coexisting.
- Hankaku katakana is a Japanese information-systems habit, not a general East Asian print practice.
- Historical colonial and occupation use of katakana in Taiwan, Korea, and the Pacific does not match present Japanese gairaigo conventions.
- Outside Japan, katakana appears mainly in Japanese-language teaching and in Unicode/font support, not as a living orthography.
- Which of these regional variation hold for the sense of katakana this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Hiragana - Same mora inventory and largely the same phonology; katakana is angular and historically from character fragments, hiragana from cursive whole characters. Functional test: native grammatical endings and default native lexicon are hiragana; gairaigo, many names, and marked diction are katakana.
- Kanji - Kanji are morphographic; each katakana letter encodes a mora, not a morpheme. Test: dakuten/handakuten and yōon combination behaviour apply to kana, not to a kanji's meaning.
- Man'yōgana - Man'yōgana uses full Chinese characters for sound; katakana are abbreviated derived signs. Test: a man'yōgana graph is still a kanji with a lexical identity; a katakana letter is not.
- Hangul - Hangul is a featural alphabet grouped into syllable blocks for Korean; katakana is a linear mora syllabary for Japanese (and some Ainu) phonology. Test: language, syllable-block layout, and Unicode script Hangul vs Katakana.
- Bopomofo (Zhuyin) - Both are phonetic East Asian letter sets, but bopomofo transcribes Mandarin initials/finals/tones and is a different Unicode script. Test: inventory (tones and zh/ch/sh series) and Script property Bopomofo vs Katakana.
- Lookalike CJK ideographs (力, 口, 二, 夕, 卜) - Same-looking kanji do not take dakuten, do not form yōon with small ャュョ, and have different Unicode code points and Script=Han.
- Latin capitals and fullwidth Latin - Angular square letters in Japanese fonts can resemble katakana. Test: Unicode Script Latin vs Katakana, and whether the string participates in kana orthography (ー, ッ, dakuten).
- Which of these neighbouring kinds and how to tell them apart hold for the sense of katakana this model covers, and on what evidence? provenance
Sources
- Katakana Range: 30A0-30FF - Unicode Consortium - Official coded repertoire, code-point range, character names, and the distinction between ordinary katakana, small letters, the prolonged-sound mark, and related signs.
- Katakana Phonetic Extensions Range: 31F0-31FF - Unicode Consortium - The extra small katakana used especially for Ainu and other phonetic extensions beyond the modern Japanese school inventory.
- 外来語の表記 (Notation of Loanwords), Cabinet Notification No. 1 of 28 June 1991 - Agency for Cultural Affairs / Government of Japan - The official modern Japanese conventions that make katakana the default script for gairaigo and that list accepted spellings, including how far extended katakana may be used for foreign sounds.
- ISO 15924, Codes for the representation of names of scripts - International Organization for Standardization (registration authority: Unicode Consortium) - The script identifier Kana (411) for katakana, and its neighbours Hira, Hrkt, and Jpan in the same code set.
What the second pass must settle
- Which authoritative repertoire and orthographic references should establish coverage of modern, historical, and specialist katakana forms?
- How should the model represent katakana used for languages other than Japanese while leaving each language's phonology and grammar to neighbouring models?
- Which spelling authorities should take precedence when general conventions conflict with personal names, brand spellings, or faithful quotation?
- Which equivalence rules are appropriate for search, deduplication, accessibility output, and archival preservation, especially for width and sound-length variants?
- What evidence threshold should permit automatic resolution of visually ambiguous kana or uncertain readings?