← Back to catalogue
Research draft

Tibetan script

vr.tr.tibetan-script · INF.MED

Enable an AI agent to recognise Tibetan script, assess the integrity and usability of its representations, and choose appropriate reading, encoding, transcription and preservation actions.

Thing Registry Information and virtual systems

Research draft, second pass

A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.

recalled by Codex without web access - no source was read

Researched by: Codex

Purpose and description

Enable an AI agent to recognise Tibetan script, assess the integrity and usability of its representations, and choose appropriate reading, encoding, transcription and preservation actions.

Tibetan script is a Brahmi-derived abugida in which consonant letters carry an inherent vowel modified by vowel signs, with consonant combinations often written as vertical stacks, used to write Tibetan, Dzongkha and other languages.

It can be Identify Tibetan-script spans and propose a graphic style and orthographic profile with explicit uncertainty.; Inspect syllable construction, stack composition and punctuation against a selected profile.; Diagnose missing glyphs, suspect character sequences and legacy encoding dependencies.; Transliterate or transcribe a passage while preserving source alignment and marking unresolved readings.; Prepare search or segmentation derivatives while retaining the original representation.; Route ambiguous handwriting, specialised signs and potentially destructive conversions for expert review..

Distinguishing features

Identify characteristic Tibetan letterforms and their structural relationships; do not infer the script solely from a document's Tibetan-language label.

Distinguish base letters, subjoined forms and vowel signs within a stack rather than treating every visible component as an independent alphabetic letter.

Recognise tsheg and shad signs and inspect their contextual functions rather than applying whitespace-based word segmentation.

Distinguish headed and headless graphic forms from changes of language or changes to a text's underlying character sequence.

Separate Tibetan-script writing from Latin transliteration, including Wylie, even when both represent the same passage.

Scope

+ Letters, vowel signs, subjoined forms, punctuation and other script signs

+ Construction and segmentation of written syllables and letter stacks

+ Relationships among Tibetan script styles, including headed and headless forms

+ Language-specific and historical orthographic profiles using Tibetan script

+ Character encoding, rendering, transliteration and transcription fidelity

- Tibetan and other languages as systems of grammar, speech and meaning

- Authorship, interpretation, editions and rights of individual works

- Conservation and ownership of manuscripts, inscriptions and printed objects

- Fonts and input applications as independently maintained software products

- Religious doctrines and cultural practices expressed through Tibetan-script texts

Characteristics

Orthographic profile
Named language, historical period and orthographic convention; mixed or undetermined Recognition and validation require the relevant convention rather than a single universal account of acceptable Tibetan-script writing.
Graphic style
Headed, headless, named subtype, mixed or undetermined Style affects recognition, transcription and the suitability of reference examples or OCR.
Repertoire coverage
Covered, partly covered or unassessed against an identified repertoire An agent must know whether uncommon signs and specialised combinations can be represented without substitution.
Stack integrity
Verified, suspect, malformed or unresolved under a stated orthographic profile A plausible-looking stack may contain an incorrect component or digital sequence.
Segmentation policy
Graphic component, stack, written syllable or lexical word; policy and exceptions identified Counting, searching and annotation produce different results at different boundaries.
Digital representation
Unicode text, legacy encoding, glyph-mapped text, image only, mixed or unknown Appearance alone does not establish that text is searchable, interoperable or safely convertible.
Transliteration correspondence
Linked source and target spans, named scheme and version, with unresolved or lossy mappings A transliteration must remain traceable to its source and distinguish spelling preservation from pronunciation guidance.
Recognition error rate
Errors per 100 reference units, with unit, sample and reference transcription specified OCR quality cannot be judged consistently without defining whether errors concern characters, stacks or written syllables.

Where this came from

wikidata · CC0 1.0

Drafted structure

Bundle to layer to finding to question, as the second pass will find it: 5 bundles · 9 layers · 16 findings · 26 questions.

Script identity and orthographic profiles Establish which Tibetan-script tradition and writing convention an agent is handling.

The script must be recognised independently of language, while interpretation of its combinations requires a specific orthographic context.

Script and language boundaries

Separate graphic-system identification from identification of the language expressed.

Independent script identification

Record graphic evidence for identifying Tibetan script and separately qualify any language attribution.

  1. Which observed letterforms, stacks and punctuation support identification as Tibetan script? definition
  2. Does the evidence identify a language, or only the script used to write it? boundary

Orthographic context

Identify the conventions governing spelling and specialised combinations.

Profile selection

Associate validation with an evidenced orthographic profile, allowing historical variation and Sanskrit-related usage to remain explicit.

  1. Which reference or documented practice establishes the orthographic profile for this material? provenance
  2. Which apparent exceptions require a historical or Sanskrit-related profile rather than correction as errors? boundary
Syllables, stacks and boundaries Describe how graphic components form readable units and how those units are separated.

Tibetan-script processing depends on internal syllable structure and distinctions between graphic, syllabic and lexical boundaries.

Letter and stack composition

Track components and their positions without collapsing visual arrangement into linguistic function.

Syllable component analysis

Record proposed root, prefix, suffix, superscribed, subjoined and vowel components where the selected profile supports those roles.

  1. Which components are present, and what evidence supports their proposed roles within the written syllable? definition
  2. Does the combination violate the selected profile, or does its analysis remain unresolved? boundary

Tsheg, shad and segmentation

Interpret separators, punctuation and layout without assuming that written syllables are lexical words.

Explicit boundary policy

Keep sign recognition, written-syllable segmentation and word segmentation separately assessable.

  1. How are tsheg, shad, spaces and line endings interpreted in this passage and its orthographic profile? definition
  2. What additional linguistic evidence supports any proposed word boundaries beyond the visible separators? measurement
Graphic variation and reading Handle Tibetan-script styles and the uncertainty involved in reading particular graphic forms.

A shared script identity does not make headed print, headless handwriting and unusual signs equally recognisable.

Headed and headless forms

Relate graphic styles to the underlying repertoire while retaining uncertain identifications.

Style-sensitive recognition

Record the evidence for a headed or headless classification and any finer attribution.

  1. Which visible features justify the proposed style classification? definition
  2. Which documented examples support mapping an unfamiliar form to a particular letter or sign? provenance

Ambiguous and specialised signs

Preserve evidence where damage, handwriting or specialised notation prevents a secure reading.

Qualified readings

Distinguish observed marks, proposed readings and editorial reconstruction.

  1. Which readings remain possible, and which visible details distinguish them? measurement
  2. Should the agent preserve an image-linked unresolved sign, supply a qualified transcription or request specialist review? action
Encoding, rendering and conversion Assess whether Tibetan-script representations preserve character identity and remain usable across digital tools.

Correct appearance, correct character encoding and reversible conversion are separate requirements for Tibetan text.

Character sequences and display

Examine the relationship between stored characters and rendered stacks or signs.

Encoding and display integrity

Record encoding dependencies and distinguish sequence defects from font or rendering defects.

  1. Is the passage Unicode text, a legacy representation or an image, and how was that established? provenance
  2. When a stack displays incorrectly, does the evidence locate the problem in the stored sequence, font coverage or rendering system? measurement
  3. Which conversion or normalisation is justified, and how will the original sequence and appearance be retained for comparison? action

Transliteration and search

Manage transformations used for transcription, interchange and retrieval.

Traceable transformations

Specify how a named transliteration or search transformation treats stacks, punctuation, unusual signs and uncertain readings.

  1. Which scheme and version are used, and does the output preserve spelling or approximate pronunciation? definition
  2. Which distinctions disappear in the transformed output, and can the original Tibetan sequence be recovered? measurement
  3. Which transformations may be used for retrieval while keeping a faithful source transcription available? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.

A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.

Reported evidence

Findings from the breadth pass, kept separate from the structural claims.

Check these first

Recalled without web access and unsourced; every item is a lead to verify.

  • This describes the writing system, not an individual text, edition, font or physical manuscript.
  • Uchen and Ume are broad distinctions; a fuller classification of regional and calligraphic styles needs specialist verification.
  • Claims are recalled knowledge, not findings checked against sources in this response.
  1. Which of these check these first hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Kinds and varieties

Recalled without web access and unsourced; every item is a lead to verify.

  • Uchen: headed letterforms commonly used in printing and formal writing
  • Ume: headless letterforms encompassing several handwriting and calligraphic styles
  1. Which of these kinds and varieties hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Identifiers and schemes

Recalled without web access and unsourced; every item is a lead to verify.

  • ISO 15924 - Tibt - Identifies the Tibetan script, not a particular language.
  • Unicode block - U+0F00-U+0FFF - The Tibetan block contains encoded letters, vowel signs, subjoined consonants, digits and other signs; the range does not imply that every position is assigned.
  1. Which of these identifiers and schemes hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Standards and regulation

Recalled without web access and unsourced; every item is a lead to verify.

  • ISO 15924, issued by ISO, provides codes for script names.
  • The Unicode Standard, maintained by the Unicode Consortium, specifies Tibetan character encoding and character properties.
  1. Which of these standards and regulation hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Real-world use

Recalled without web access and unsourced; every item is a lead to verify.

  • Writing and publishing Tibetan-language texts.
  • Writing Dzongkha in Bhutan.
  • Preserving and transmitting Tibetan Buddhist scriptures, commentaries and ritual texts.
  • Producing educational materials, signage and digital communications.
  • Calligraphy and manuscript production.
  1. Which of these real-world use hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Failure modes and hazards

Recalled without web access and unsourced; every item is a lead to verify.

  • Fonts or shaping systems with inadequate Tibetan support can display consonant stacks and vowel signs incorrectly.
  • Incorrect character ordering can disrupt rendering, searching and text comparison.
  • Treating the tsheg syllable delimiter as a word-space produces unreliable word segmentation.
  • OCR can confuse similar letterforms or misread stacked and handwritten characters.
  • Confusing script identification with language identification can misclassify Dzongkha or other texts as Tibetan.
  1. Which of these failure modes and hazards hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Regional variation

Recalled without web access and unsourced; every item is a lead to verify.

  • Tibetan-speaking communities use differing manuscript and calligraphic traditions within the script.
  • Bhutan uses Tibetan script for Dzongkha, with distinctive orthographic and calligraphic conventions.
  1. Which of these regional variation hold for the sense of Tibetan script this model covers, and on what evidence? provenance

Neighbouring kinds and how to tell them apart

Recalled without web access and unsourced; every item is a lead to verify.

  • Tibetan language - A language has vocabulary and grammar; Tibetan script is a writing system that can represent more than one language.
  • Dzongkha - Dzongkha is a language written using Tibetan script, not a separate name for the script.
  • Wylie transliteration - Wylie represents Tibetan spelling with Latin characters rather than Tibetan letterforms.
  • Tibetan typeface - A typeface is a particular visual design of the script's glyphs; the script encompasses the underlying writing system.
  1. Which of these neighbouring kinds and how to tell them apart hold for the sense of Tibetan script this model covers, and on what evidence? provenance

What the second pass must settle

  • Which authoritative references should define the initial repertoire and orthographic profiles, including historical spellings and Sanskrit-related combinations?
  • What level of classification among headed and headless styles is both operationally useful and reliably supported by reference specimens?
  • Where should the model draw the boundary between Tibetan-script signs, specialised notation and ornamental marks?
  • Which encoding, normalisation and transliteration specifications cover the intended material, and which cases prevent reversible conversion?
  • Which representative samples and expert transcriptions should establish recognition and segmentation quality across print, handwriting and historical material?