character
a written symbol that is used to represent speech
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
recalled by Codex without web access - no source was read
Researched by: Codex
Purpose and description
Enable an agent to recognise a written character, assess its identity and interpretation in context, and choose operations that preserve its intended linguistic distinctions.
A character is a conventional written symbol used within a writing system to represent a unit of language, such as a speech sound, syllable or morpheme.
It can be Identify a character from an occurrence using graphic, script and linguistic context.; Record competing identifications and the evidence needed to distinguish them.; Associate a character with supported contextual readings and orthographic functions.; Transcribe or encode a character while preserving its source form and unresolved distinctions.; Compare variants and propose substitutions under explicit orthographic rules.; Flag transformations that could merge distinct characters or lose evidence..
Distinguishing features
The symbol must have an attested linguistic role in a writing convention; visual resemblance to writing alone does not establish character status.
A change of handwriting or typeface need not change character identity; determine whether the relevant orthography treats the forms as equivalent.
A character identity and a digital code point are separate descriptions; establish their correspondence rather than assuming a one-to-one mapping.
Identical-looking forms may have different identities in different scripts or conventions; appearance alone is insufficient.
The character is a repeatable symbolic identity; a particular printed or handwritten occurrence is an instance carrying contextual evidence.
Scope
+ Character identity within a specified script or orthographic tradition
+ Relations to represented speech units and context-dependent readings
+ Graphic variants and combinations that affect character recognition
+ Digital representations and their correspondence to character identity
+ Recognition uncertainty and conditions for safe transcription or substitution
- Fictional persons, personality traits and other senses of character
- Complete languages, speech systems and writing systems
- Words, sentences and documents as composed linguistic works
- Fonts, typefaces and glyph designs as independently managed artifacts
- Physical inscriptions, pages and display devices as material carriers
- Encoding standards and text-processing software as complete systems
Characteristics
- Writing-system membership
- Script and orthographic tradition, with language and period where relevant Establishes the conventions under which a form has a linguistic identity.
- Speech-representation function
- Segmental, syllabic, morphographic, auxiliary, mixed or unresolved Guides interpretation without assuming that every character denotes one sound.
- Contextual reading
- Supported reading linked to language, surrounding characters and usage conditions Prevents assigning a single pronunciation where interpretation depends on context.
- Graphic equivalence
- Form-to-identity relation: accepted variant, distinct character, disputed or unresolved Separates harmless visual variation from distinctions that transcription must preserve.
- Combination behaviour
- Independent, combining, contextually joining, convention-dependent or unresolved Helps determine whether a visible component can be interpreted or manipulated independently.
- Digital representation
- Encoding standard and version, code point or sequence, and mapping evidence Supports storage and interchange without equating encoding units with linguistic units.
- Usage status
- Current, historical, specialised, obsolete, disputed or unknown within a named community Constrains substitution and interpretation to the appropriate usage tradition.
- Identification status
- Resolved, provisional, ambiguous or illegible for a particular occurrence Determines whether an agent can commit to a transcription or must retain alternatives.
Analytical facets
- substance
- information
- origin
- conceptual
- agency
- inert
- mobility
- not-applicable
- scale
- not-applicable
- affordances
- readable
Also called
+72
Where this came from
oewn:2024 · CC BY 4.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 19 findings · 29 questions.
Identity and boundaries Determine what counts as this character and which neighbouring symbolic units remain distinct.
A visible mark, character identity and orthographic unit cannot safely be treated as interchangeable.
Writing convention
Locate character identity within the conventions that establish its linguistic use.
Character membership
Record the script, orthography and evidence supporting recognition as a character.
- Which script and orthographic tradition recognise this symbol, and under what name or identifier? definition
- Which attested usage or authoritative description supports its membership in that tradition? provenance
Unit boundaries
Separate character identity from glyphs, components and larger orthographic units.
Segmentation and identity
Make explicit whether a form is one character, a component or a combination under the relevant convention.
- Does the orthography treat this form as one character, part of a character or a sequence of characters? boundary
- What evidence distinguishes this character from a similar glyph or a larger graphemic unit? boundary
Linguistic function Describe how the character participates in representing speech.
Interpretation requires the relevant language and orthographic rules, not an assumed fixed sound value.
Represented unit
Identify the linguistic contribution attributed to the character.
Speech correspondence
Record whether and how the character represents a speech unit or modifies its representation.
- What speech unit or speech-related distinction does this character represent in the specified orthography? definition
- Does its contribution require other characters, and which combinations establish that contribution? boundary
Reading conditions
Capture the conditions that select among possible readings.
Context-dependent readings
Associate readings with their linguistic context and supporting evidence.
- Which readings are supported for each relevant language, period and surrounding-character context? provenance
- When context leaves multiple readings possible, what alternatives must an agent retain? action
Graphic realisation Describe the visible evidence used to recognise the character across forms and combinations.
Recognition must tolerate permitted variation while preserving graphic distinctions that carry identity.
Diagnostic form
Identify graphic features that distinguish the character from plausible alternatives.
Recognition features
Record distinguishing components and ambiguities in observed forms.
- Which components, marks or spatial relations distinguish this character from its nearest visual confusables? definition
- Which distinguishing features remain observable in the occurrence being identified? measurement
Variants and combinations
Handle alternative forms, joining behaviour and attached marks without losing identity.
Form equivalence
Record the conventions under which graphic differences preserve or change character identity.
- Which handwritten, typographic or positional forms count as variants of this character in the relevant tradition? boundary
- When a mark is added, removed or joined, does the result retain this identity, form a combination or become another character? boundary
Encoding and transcription Connect character identity to digital representations and controlled transformations.
A transcription can display plausibly while encoding a different identity or discarding a meaningful distinction.
Representation mapping
Document how the character is represented in a specified encoding environment.
Encoding correspondence
Record supported code points or sequences and limits of their correspondence to the character.
- Which code point or sequence represents this character in the named encoding standard and version? provenance
- Does that representation uniquely identify this character, or does interpretation require additional script or orthographic context? boundary
Transformation policy
Specify when transcription and text transformations preserve the distinctions required by the task.
Permitted transformations
Assess changes against character identity, source fidelity and the intended use.
- Which proposed normalisation, case conversion or substitution preserves the distinctions required by this task? action
- What source form and mapping record must be retained when transliteration or fallback representation changes or loses distinctions? action
Attestation and interpretive state Track where character claims apply and whether an observed occurrence can be interpreted reliably.
Historical usage and ambiguous occurrences require qualified identifications rather than unconditional assignments.
Usage context
Locate the character's use in a community, period and orthographic practice.
Attested use
Separate documented use from proposed, reconstructed or disputed character assignments.
- Which dated examples or reference works attest this character in the specified writing community? provenance
- Is its status current, historical, specialised or disputed within that community, and what supports that assessment? provenance
Occurrence assessment
Assess uncertainty when mapping a particular inscription or rendering to a character identity.
Identification decision
Record candidate identities, missing evidence and the permissible next step.
- Which candidate characters remain compatible with the visible form and linguistic context? boundary
- Should the agent commit a transcription, preserve alternatives or request a clearer image or specialist review? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Check these first
Recalled without web access and unsourced; every item is a lead to verify.
- This covers the linguistic written-symbol sense; the broader computing sense also includes punctuation, controls and other encoded elements.
- Logographic characters generally represent linguistic units rather than directly specifying pronunciation; representing speech should not imply a purely phonetic system.
- Character, grapheme and user-perceived character are not interchangeable; a researcher should establish the intended unit for any inventory or measurement.
- Which of these check these first hold for the sense of character this model covers, and on what evidence? provenance
Kinds and varieties
Recalled without web access and unsourced; every item is a lead to verify.
- Alphabetic letters
- Consonant letters in abjads
- Syllabic characters
- Logographic characters
- Vowel signs
- Diacritical marks
- Which of these kinds and varieties hold for the sense of character this model covers, and on what evidence? provenance
Identifiers and schemes
Recalled without web access and unsourced; every item is a lead to verify.
- Unicode code point notation - U+0041 - Identifies LATIN CAPITAL LETTER A; a written character may require multiple Unicode code points.
- Unicode character name - LATIN CAPITAL LETTER A - A standardized name for an encoded character, distinct from its pronunciation or language-specific letter name.
- ISO 15924 - Latn - Identifies the Latin script, not an individual character; script codes provide classification context.
- Which of these identifiers and schemes hold for the sense of character this model covers, and on what evidence? provenance
Standards and regulation
Recalled without web access and unsourced; every item is a lead to verify.
- The Unicode Standard - Unicode Consortium: character encoding, properties and text-processing specifications.
- ISO/IEC 10646 - ISO and IEC: the Universal Coded Character Set.
- Unicode Standard Annex #15, Unicode Normalization Forms - Unicode Consortium: normalization of equivalent encoded sequences.
- Unicode Standard Annex #29, Unicode Text Segmentation - Unicode Consortium: default boundaries for grapheme clusters and other text units.
- ISO 15924 - ISO: codes for the representation of names of scripts.
- Which of these standards and regulation hold for the sense of character this model covers, and on what evidence? provenance
Real-world use
Recalled without web access and unsourced; every item is a lead to verify.
- Recording language in handwriting, print and digital text.
- Teaching reading, spelling and writing.
- Transcribing speech with phonetic symbols.
- Encoding, searching and exchanging textual information.
- Recognizing written language through optical character recognition.
- Which of these real-world use hold for the sense of character this model covers, and on what evidence? provenance
Typical measurements
Recalled without web access and unsourced; every item is a lead to verify.
- Unicode code points per user-perceived character - One or more; no single fixed count applies across writing systems and combining sequences. - code points
- Which of these typical measurements hold for the sense of character this model covers, and on what evidence? provenance
Failure modes and hazards
Recalled without web access and unsourced; every item is a lead to verify.
- Confusing characters with glyphs or code points causes incorrect counting, selection and editing.
- Incorrect encoding interpretation produces garbled text or replacement characters.
- Missing fonts or inadequate shaping support makes characters unreadable or renders them incorrectly.
- Visually confusable characters can cause transcription errors and facilitate identifier spoofing.
- Naive comparison or diacritic removal can miss equivalent sequences or erase meaningful distinctions.
- Which of these failure modes and hazards hold for the sense of character this model covers, and on what evidence? provenance
Regional variation
Recalled without web access and unsourced; every item is a lead to verify.
- The same character can represent different sounds in different languages and orthographies.
- Chinese character use includes simplified and traditional conventions, with additional regional differences in preferred forms.
- Letter inventories, diacritic use and alphabetic ordering vary among languages sharing a script.
- Which of these regional variation hold for the sense of character this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Recalled without web access and unsourced; every item is a lead to verify.
- glyph - A glyph is a rendered visual form; one character can have several glyphs, and one glyph can render several characters.
- grapheme - A grapheme is a contrastive unit of a particular writing system and can comprise more than one character.
- code point - A code point is a numerical position in a coded character set; it is an encoding concept rather than a linguistic symbol.
- phoneme - A phoneme is a contrastive sound unit in a language; a character is written, and their correspondence need not be one-to-one.
- letter - A letter is a character belonging to an alphabet or comparable letter inventory; syllabic and logographic characters need not be letters.
- script - A script is a system or repertoire of written signs; a character is an individual symbol used within it.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of character this model covers, and on what evidence? provenance
What the second pass must settle
- Does the registry definition include punctuation, numerals and other auxiliary writing signs whose relationship to speech is indirect?
- Which sources should govern character identity when orthographic practice, historical scholarship and encoding standards use different boundaries?
- When should visually identical symbols used across scripts or languages share a character identity, and when should they remain distinct?
- How should this model align character, grapheme, combining component and user-perceived text unit without assuming universal equivalence?
- What evidence threshold should permit automatic identification or substitution for historical, damaged or poorly documented character forms?