← Back to catalogue
Research draft

hiragana

vr.tr.hiragana · INF.MED

Enable an AI agent to recognise hiragana, assess the identity and usability of its written forms, and choose appropriate reading, annotation, conversion or preservation actions.

Thing Registry Information and virtual systems

Research draft, second pass

A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.

Researched by: Codex

Purpose and description

Enable an AI agent to recognise hiragana, assess the identity and usability of its written forms, and choose appropriate reading, annotation, conversion or preservation actions.

It can be Identify hiragana spans and flag uncertain characters, small-kana distinctions or voicing marks for review; Segment hiragana into reading units and propose context-supported readings; Check spelling against an explicitly selected contemporary or historical convention; Convert eligible spans to katakana or romanisation while retaining the original and recording ambiguities; Normalise equivalent digital representations after checking which distinctions the task must preserve; Annotate grammatical roles, reading relationships and uncertain historical forms without rewriting the source.

Distinguishing features

Determine script identity from the character or interpreted glyph rather than roundness alone: hiragana あ and katakana ア can represent corresponding readings while remaining different script characters.

Distinguish hiragana from kanji by its kana identity and contextual phonographic use; a cursive appearance alone does not establish that a handwritten form is hiragana.

Test small versus full-sized character identity: ゃ and や, or っ and つ, cannot be treated as interchangeable visual variants.

Separate script identity from encoded block membership: a combining voicing mark or shared punctuation may participate in hiragana text without independently identifying that text as hiragana.

Distinguish hiragana spelling from a phonetic transcription: particle uses of は, へ and を require contextual reading rather than a single invariant sound assignment.

Scope

+ Identification of hiragana characters and sequences, including small kana and voiced or semi-voiced forms

+ Relationships between written hiragana, mora structure and context-dependent readings

+ Hiragana spelling conventions, grammatical uses and relationships to neighbouring scripts

+ Distinctions among contemporary, historical and nonstandard forms or usages

+ Integrity of hiragana across handwritten, printed and digitally encoded representations

- Japanese vocabulary, syntax and discourse beyond what is needed to interpret hiragana usage

- Kanji identity, meanings and internal character structure

- Katakana inventory and usage except where comparison or conversion involves hiragana

- Japanese phonetics and speech performance beyond the reading distinctions represented by hiragana

- Document authorship, copyright and access permissions

- General font engineering, keyboard design and text storage infrastructure

Characteristics

Recognition granularity
individual character; orthographic unit; text span; inventory or convention Prevents a conclusion about one glyph from being applied automatically to an entire word, passage or writing system.
Character and modifier identity
identified hiragana character; small or full-sized form; unmarked, dakuten or handakuten form; unresolved identity Size and voicing distinctions can change reading and must survive recognition and transformation.
Reading relationship
written unit linked to one or more context-supported readings, with alternatives retained Supports reading assistance without assuming that every character has one context-free pronunciation.
Mora count
morae under a stated interpretation; unresolved when context is insufficient Separates written character count from rhythmic units, including small-kana combinations, small っ and ん.
Orthographic function
whole-word spelling; particle; okurigana; reading annotation; expressive spelling; other or unresolved The function determines which reading and editing rules apply.
Convention alignment
conforms to named convention; intentional departure; suspected error; unresolved Separates mistakes from historical spelling, specialist usage and deliberate stylistic choices.
Historical or variant relationship
observed form linked to a documented contemporary, historical or variant identification Allows recognition of older material without silently replacing its distinctions with contemporary spelling.
Representation integrity
verified against source; equivalent encoding variation; ambiguous glyph; suspected recognition error; lossy transformation Determines whether the agent can safely search, copy, normalise or convert the text.

Also called

okurigana

Where this came from

wikidata · CC0 1.0

Drafted structure

Bundle to layer to finding to question, as the second pass will find it: 5 bundles · 10 layers · 10 findings · 20 questions.

Hiragana character identity Establishes which hiragana characters are present and which visible marks participate in their identity.

An agent must distinguish character identity from glyph appearance before reading or editing hiragana.

Script membership

Identifies hiragana within mixed Japanese text and ambiguous visual material.

Hiragana identification basis

Record the evidence for identifying a character or span as hiragana, preserving alternatives when the source is unclear.

  1. Which characters are identifiable as hiragana, and which marks are shared with other Japanese scripts? definition
  2. Does identification come from encoded text, a visible glyph or an interpretation of handwriting? provenance

Size and voicing distinctions

Tracks small kana and voicing distinctions that affect interpretation.

Contrastive kana features

Record whether size and dakuten or handakuten belong to the intended character rather than resulting from rendering or recognition noise.

  1. Is the observed kana small or full-sized, and is that distinction supported by the source? definition
  2. Is a voicing mark present, absent or ambiguous, and would changing it alter the intended reading? boundary
Hiragana reading units Relates hiragana sequences to mora structure and context-sensitive readings.

Character-by-character substitution cannot reliably describe how hiragana sequences are read.

Mora segmentation

Distinguishes character boundaries from the units relevant to reading and mora counting.

Sequence-to-mora analysis

Record how combinations such as きゃ, small っ and ん contribute to the sequence's interpreted mora structure.

  1. Which characters form a combined reading unit, and which contribute a separate mora? definition
  2. How many morae does the sequence have under the selected reading, and where is the analysis uncertain? measurement

Contextual reading

Captures readings that depend on grammatical or lexical context.

Spelling-reading correspondence

Record the contextual interpretation of particle spellings and vowel sequences rather than treating spelling as a complete phonetic transcription.

  1. Are は, へ or を functioning as particles here, and what reading does that support? definition
  2. What lexical or grammatical evidence supports interpreting a vowel sequence as a long vowel or as separate vowels? provenance
Hiragana orthographic role Explains what hiragana is doing within the surrounding Japanese writing.

Appropriate interpretation and editing depend on whether hiragana spells a word, marks grammar or annotates another expression.

Word and grammar functions

Identifies whole-word spelling, particles and kana attached to kanji.

Hiragana functional span

Record the role and extent of each relevant hiragana span, including its attachment to neighbouring kanji when used as okurigana.

  1. Does this span spell an entire word, function as a particle or supply okurigana? definition
  2. Which surrounding word or kanji does the span belong to, and where does that relationship end? boundary

Reading annotation and script choice

Distinguishes reading annotations from body text and records evidence for deliberate hiragana spelling.

Annotation target and writing choice

Record whether hiragana annotates another expression or is the expression's chosen spelling, without inferring authorial intent from script alone.

  1. Is this hiragana a reading annotation, and which exact base-text span does it annotate? boundary
  2. What evidence, if any, establishes that hiragana was deliberately selected in place of a kanji or katakana spelling? provenance
Hiragana conventions and history Assesses spelling and character forms against the conventions relevant to their context.

An agent must distinguish errors from historical, specialist and intentional uses before proposing correction.

Applicable spelling convention

Names the orthographic reference used to evaluate the text.

Convention-relative assessment

Record the applicable spelling convention and classify departures only when there is enough contextual evidence.

  1. Which documented convention governs this text, including relevant treatment of じ versus ぢ and ず versus づ? provenance
  2. Does a departure warrant correction, preservation as intentional usage or an unresolved annotation? action

Historical and variant forms

Handles older spellings and forms without assuming contemporary equivalence.

Historical form interpretation

Record source-supported identification of forms such as ゐ and ゑ, and assess how hentaigana or other historical variants relate to this model's scope.

  1. Which dated source or specialist reference supports the identification and interpretation of this historical form? provenance
  2. Would a contemporary transcription discard a spelling or form distinction that the task requires preserving? boundary
Hiragana representation and transformation Assesses digital and visual fidelity and constrains transformations of hiragana text.

Normalisation, OCR and script conversion can preserve appearance while changing identity, reading or recoverability.

Encoding and source fidelity

Checks correspondence among source glyphs, recognised characters and encoded sequences.

Kana representation equivalence

Record whether differences reflect equivalent encodings, distinct characters or errors, especially around combining voicing marks and small kana.

  1. Do differing encodings represent the same intended kana, or do they preserve a meaningful character distinction? boundary
  2. Which source comparison verifies that OCR or transcription retained small kana and voicing marks correctly? provenance

Controlled kana conversion

Defines task-appropriate normalisation, katakana conversion and romanisation.

Transformation contract

Record the conversion's purpose, convention, contextual dependencies and recoverability before changing hiragana.

  1. Should the transformation preserve written spelling or represent contextual reading, and which conversion convention applies? action
  2. Which distinctions or ambiguities would be lost, and what original text or annotation must accompany the result? boundary

What the second pass must settle

  • Does the registry intend hiragana to include hentaigana and other historical forms directly, or should this model link to separately owned historical-script models?
  • Which authoritative orthographic references and editions should govern contemporary spelling, historical spelling and okurigana assessment?
  • Which specialist small-kana usages, dialect spellings and expressive spellings require explicit coverage rather than an unresolved-use category?
  • Which shared marks, including iteration and prolonged-sound marks where encountered, should this model interpret directly and which should belong to a neighbouring Japanese writing model?
  • Which romanisation and modernisation conventions should be supported, and what evidence is required before resolving an ambiguous reading during conversion?