Arabic alphabet
Enable an agent to recognise the Arabic alphabet, record its repertoire and writing conventions, assess a representation against an identified convention, and choose transformations that preserve intended distinctions.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
recalled by Codex without web access - no source was read
Researched by: Codex
Purpose and description
Enable an agent to recognise the Arabic alphabet, record its repertoire and writing conventions, assess a representation against an identified convention, and choose transformations that preserve intended distinctions.
The Arabic alphabet is the conventional set of 28 basic letters used to write Arabic in a predominantly right-to-left, joining script, structurally classified as an abjad because consonants and long vowels are normally represented while short vowels are usually omitted or indicated by optional marks.
It can be Classify a sign as an Arabic letter, orthographic mark, contextual form or neighbouring script extension.; Validate a proposed repertoire and order against an identified Arabic orthographic profile.; Explain contextual joining and distinguish a shaping failure from a different letter sequence.; Assess whether removing dots, vowel marks or hamza-related distinctions would lose information.; Compare encoded and displayed text while preserving logical order and letter identity.; Choose a documented normalisation or search policy and identify transformations requiring human review..
Distinguishing features
The subject is a conventional repertoire used to write Arabic, rather than every character belonging to the broader Arabic script.
Consonantal letters and context-dependent vowel representation make the term abjad relevant; the registered name alphabet must not imply obligatory letters for every vowel.
A letter can have contextual joining forms without those forms becoming separate alphabet members.
Distinguishing dots can determine letter identity, whereas optional vocalisation marks serve a different orthographic role.
Arabic text has a right-to-left base direction, but mixed text cannot be represented correctly by simply reversing a character sequence.
Scope
+ Arabic letter identities, names, repertoire membership and ordering conventions
+ Relationships between letters, contextual forms, ligatures and distinguishing dots
+ Arabic-specific use of vowel signs, hamza, shadda and other orthographic marks
+ Right-to-left writing, joining constraints and interactions with mixed-direction text
+ Digital representations and transformations that affect Arabic letter identity or orthographic distinctions
+ Explicitly identified historical, regional and specialised orthographic profiles
- General properties of alphabets and writing systems owned by the wider alphabet model
- Arabic grammar, vocabulary, pronunciation and dialect systems beyond their relationship to written signs
- Additional repertoires for other languages using Arabic script
- Individual manuscripts, books, editions and inscriptions as authored or material objects
- Typeface design, calligraphic artworks and font licensing as independently managed things
- Translation and transliteration systems except where their mappings clarify alphabet boundaries
Characteristics
- Orthographic profile
- Named convention with language variety, period, region and specialised use where applicable Repertoire and acceptable spellings must be assessed against a stated convention.
- Repertoire membership
- Sign related to profile as core letter, orthographic variant, auxiliary sign or excluded extension Prevents broader Arabic-script characters from being silently treated as Arabic alphabet members.
- Letter count
- Number of members under an explicit counting rule A count is interpretable only when treatment of hamza and other disputed or dependent forms is recorded.
- Joining behaviour
- Applicable joining class, named standard and exceptional context Supports recognition of valid connections and diagnosis of broken rendering.
- Vocalisation coverage
- Unvocalised, partially vocalised or fully vocalised relative to a declared profile Missing marks can be conventional rather than evidence of corruption.
- Ordering convention
- Named sequence, such as hijai or abjad, with its actual member order Alphabet teaching order, historical order and software collation need not be interchangeable.
- Representation fidelity
- Unchecked, verified for a stated purpose, lossy or ambiguous A visually plausible rendering may conceal character substitutions or lost distinctions.
- Representational equivalence
- Two forms related as the same letter, encoding equivalent, orthographic alternative or distinct signs under a named rule Constrains normalisation, search expansion and transcription.
Also called
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 18 findings · 28 questions.
Repertoire and boundaries Establishes which signs belong to the Arabic alphabet under the selected convention.
Arabic alphabet membership must remain distinct from Arabic script coverage and from the generic alphabet concept.
Alphabet specialisation
Locates this registered thing within writing-system concepts.
Arabic alphabet sense
Record the Arabic-language repertoire intended by the entry and its relationship to alphabet, abjad and Arabic script.
- Which Arabic writing convention defines the repertoire being modelled? definition
- Which properties are inherited from the wider alphabet model, and which require Arabic-specific treatment? boundary
Membership and counting
Separates alphabet members from related signs and extensions.
Explicit repertoire rule
Record the member list and the rule used to classify hamza, alif variants, ta marbuta, alif maqsura and lam-alif.
- Which of these signs are counted as independent letters, variants, auxiliary signs or ligatures in the selected authority? definition
- What letter count follows from that rule, and which Arabic-script extensions are excluded? measurement
Letter identity and shaping Connects abstract letters with visible forms without equating appearance with identity.
Arabic recognition depends on distinguishing identity-bearing features from contextual shaping and stylistic variation.
Dots and letter identity
Records how dots and base shapes distinguish letters.
Identity-bearing features
Identify the features needed to distinguish letters within a specified writing tradition.
- Which letters share similar base shapes, and which dot patterns distinguish them in this profile? definition
- When dots are absent or unclear, what contextual evidence is required before assigning a letter identity? action
Joining and ligatures
Records contextual forms, connection constraints and multi-letter glyphs.
Contextual form resolution
Relate visible forms to underlying letters, including letters that interrupt joining and ligatures such as lam-alif.
- Which joining contexts are available for each letter, and which apparent connections are stylistic? definition
- How can an agent recover the underlying letter sequence from a ligature or contextual glyph without treating it as a new alphabet member? action
Orthographic signs and vocalisation Models the roles of written vowel information and supplementary orthographic signs.
The presence, absence and placement of Arabic marks affect interpretation, while their status differs from that of core letters.
Vowel representation
Distinguishes vowel-related letter uses from added vocalisation marks.
Vocalisation policy
Record how the profile uses alif, waw, ya and short-vowel signs, and what omission of signs permits.
- How does the profile distinguish consonantal and vowel-related uses of these letters? definition
- Which marks are expected for the intended use, and how is vocalisation completeness assessed? measurement
Hamza and additional marks
Captures sign-specific distinctions that generic diacritic handling would obscure.
Mark-specific interpretation
Record the profile's treatment of hamza and its carriers, madda, shadda, sukun and specialised signs.
- Which distinctions involving hamza, its carriers and other marks are required by the selected orthography? definition
- Which marks can be omitted for a specified task, and which omissions change spelling or remove needed reading information? action
Direction and digital representation Preserves Arabic text identity across storage, shaping, display and comparison.
Arabic display order, joining and character sequences interact in ways that can produce plausible-looking but incorrect text.
Logical and visual order
Distinguishes stored sequence from displayed arrangement.
Mixed-direction integrity
Record expected reading order and the effect of embedded numbers, Latin text, punctuation and directional controls.
- What logical sequence should the text preserve when Arabic appears beside numbers or left-to-right text? definition
- Which display and copy-paste checks would expose reordered characters or misplaced punctuation? action
Encoding and equivalence
Separates standard encoding equivalence from task-specific substitutions.
Controlled text transformation
Record encoding, presentation forms, combining marks, tatweel and joining controls before choosing transformations.
- Which character sequences are equivalent under the identified encoding standard, and which merely look similar? boundary
- For this task, may an agent fold alif variants, remove vocalisation or tatweel, or alter joining controls, and what information would each operation lose? action
Ordering and attested variation Makes order and convention-dependent differences explicit and traceable.
A single unqualified modern repertoire cannot adequately represent historical ordering or regional and specialised practice.
Alphabetic order
Records named member sequences and their intended applications.
Order versus collation
Distinguish hijai and abjad sequences from rules used to sort complete words.
- Which authority supplies the precise hijai or abjad sequence used by this profile? provenance
- Which additional rules govern sorting words containing hamza forms, vowel marks or other orthographic variants? boundary
Historical and specialised profiles
Bounds differences by period, region and use.
Attested profile differences
Record sourced differences in dotting, sign inventory and orthographic practice, including specialised Quranic usage where applicable.
- Which dated or institutionally identified sources establish each claimed regional, historical or specialised difference? provenance
- Does the difference change repertoire or orthography, or does it belong to calligraphic style or an individual document? boundary
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Check these first
Recalled without web access and unsourced; every item is a lead to verify.
- This is recalled knowledge, not source-verified research.
- The conventional count of 28 needs qualification when discussing hamza, derived forms and pedagogical counting practices.
- The sense assumed is the basic alphabet used for Arabic, not every character used by Arabic-script languages.
- Which of these check these first hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Identifiers and schemes
Recalled without web access and unsourced; every item is a lead to verify.
- ISO 15924 - Arab; 160 - Identifies the Arabic script, including its use beyond Arabic, rather than this alphabet alone.
- Unicode - Arabic block: U+0600-U+06FF - Contains core Arabic letters, marks and other characters; additional blocks support the wider script. The block is not an alphabet inventory.
- Which of these identifiers and schemes hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Standards and regulation
Recalled without web access and unsourced; every item is a lead to verify.
- The Unicode Standard, issued by the Unicode Consortium, specifies character encoding and properties for Arabic text.
- Unicode Standard Annex #9, Unicode Bidirectional Algorithm, issued by the Unicode Consortium, specifies bidirectional text ordering.
- ISO 15924, issued by ISO, provides codes for identifying scripts, including Arabic.
- Which of these standards and regulation hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Real-world use
Recalled without web access and unsourced; every item is a lead to verify.
- Writing Arabic in education, administration, publishing and everyday communication.
- Writing Qur'anic text with specialized orthographic and recitation marks.
- Representing Arabic in digital documents, messaging, search and typesetting.
- Providing the letter-system foundation for adaptations used to write languages such as Persian and Urdu.
- Which of these real-world use hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Typical measurements
Recalled without web access and unsourced; every item is a lead to verify.
- Conventional basic-letter inventory - 28 - letters
- Which of these typical measurements hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Failure modes and hazards
Recalled without web access and unsourced; every item is a lead to verify.
- Omitted short-vowel marks can leave words ambiguous without linguistic context.
- Missing or misrecognized dots can change letter identity, particularly in handwriting and optical character recognition.
- Incorrect joining or contextual shaping can produce broken or misleading text.
- Mixed Arabic, Latin text and numerals can display in an unintended order when bidirectional processing is mishandled.
- Inappropriate normalization of hamza-bearing letters, alif variants or other distinctions can impair spelling fidelity and search.
- Which of these failure modes and hazards hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Regional variation
Recalled without web access and unsourced; every item is a lead to verify.
- Maghrebi manuscript and calligraphic traditions use some letter forms and dotting conventions that differ from common eastern practices, notably for fa and qaf.
- Qur'anic orthography and marking conventions differ in some respects from ordinary modern Arabic spelling and typography.
- Arabic-script alphabets for other languages add or modify letters; these are adaptations of the script rather than regional inventories of the Arabic alphabet.
- Which of these regional variation hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Recalled without web access and unsourced; every item is a lead to verify.
- alphabet - The broader registry concept concerns letter-based writing systems generally; this entry specifies the Arabic inventory and conventions, while noting its stricter classification as an abjad.
- abjad - An abjad is a type of writing system organized primarily around consonant letters; the Arabic alphabet is a particular instance.
- Arabic script - The script encompasses graphic conventions and characters used across multiple languages; the Arabic alphabet is the basic letter inventory for Arabic.
- Arabic language - Arabic is a language with spoken and written varieties; the alphabet is a means of representing it in writing.
- Arabic calligraphy - Calligraphy concerns the artistic formation and arrangement of writing; styles such as Naskh and Kufic do not constitute separate basic alphabets.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of Arabic alphabet this model covers, and on what evidence? provenance
What the second pass must settle
- Which authoritative orthographic profile should serve as the default when the registry supplies no definition?
- How should the publication count and classify hamza, alif variants, ta marbuta and alif maqsura while preserving alternative documented conventions?
- Which historical and regional practices need explicit profiles, and which should remain references to neighbouring models?
- How much specialised Quranic marking belongs in this alphabet model rather than a linked orthographic or recitation model?
- Which encoding-standard version and task-specific comparison policies should govern digital equivalence and permitted information loss?