question mark
Enable an AI agent to recognise a question mark, assess whether its form and use fit its context, and decide whether to preserve, insert, replace, explain or flag it.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
Researched by: Codex + Grok
Purpose and description
Enable an AI agent to recognise a question mark, assess whether its form and use fit its context, and decide whether to preserve, insert, replace, explain or flag it.
A punctuation grapheme that marks interrogative illocution (and, by extension, uncertainty or a missing value) in writing and encoded media, whose canonical Latin-script form is Unicode U+003F QUESTION MARK, inherited from ASCII 0x3F and historically the medieval punctus interrogativus.
It can be Locate and classify an occurrence while retaining its original representation.; Associate the sign with the text, token or component whose interpretation it affects.; Check its placement and function against an identified language or notation convention.; Insert, remove or replace it when editing authority and contextual rules support the change.; Explain its evidenced role or flag ambiguity for review.; Compare source and output representations to detect recognition errors or unintended substitutions..
Distinguishing features
Distinguish the sign from a question: a question-mark occurrence can be recorded even when no interrogative content is present.
Distinguish identity from resemblance: a curved stroke and dot are visual evidence, while the encoded character, source or declared convention may support a different identity.
Distinguish a literal question mark from a symbol with a formally assigned role by checking the enclosing notation and whether its interpreter gives it operational meaning.
Distinguish a question mark used to signal uncertainty from an interface control by recording whether the sign itself is merely displayed or belongs to an actionable component.
Distinguish a standalone sign from a sequence or composite by checking character boundaries, visual grouping and the governing convention before splitting it.
Scope
+ Identification of a question-mark sign independently of its font or visual styling
+ Individual occurrences and their attachment to surrounding text or interface elements
+ Interrogative, uncertainty-signalling, literal and formally assigned uses
+ Encoded, drawn and rendered representations and their correspondence
+ Context-sensitive validation and permitted transformations
- The content, answerability and resolution of a question
- Complete grammar and punctuation systems of individual languages
- Whole typefaces, font production and general text layout
- Programming languages, query languages and URI syntax beyond the role assigned to a question-mark occurrence
- Complete help controls, status indicators or other interface components containing the sign
Characteristics
- Sign identity
- Identified question mark; candidate variant; lookalike; unresolved Prevents visual resemblance from being treated as sufficient proof of identity.
- Recorded representation
- Encoded text; vector outline; raster image; handwriting; mixed representation Determines how identity can be checked and which edits preserve the source.
- Language or notation context
- Reference to the applicable language, writing convention, formal notation or interface convention; unknown Placement and interpretation depend on the convention governing the occurrence.
- Contextual function
- Interrogative punctuation; uncertainty annotation; literal mention; formal syntax; interface symbol; other evidenced function; unresolved The same visible sign can require different interpretation and handling.
- Attachment
- Reference to the affected text span, token, annotation target or interface component; unattached; unresolved Identifies what the occurrence qualifies or participates in.
- Sequence position
- Start offset and length with the indexing unit declared; unavailable for unsegmented visual sources Supports exact selection without confusing character boundaries with visual boundaries.
- Contextual validity
- Unchecked; supported by the governing convention; disputed; inconsistent with the governing convention Separates an observed occurrence from a judgement that its use is appropriate.
- Representation integrity
- Verified against source; suspected recognition error; suspected substitution; illegible; unresolved Helps distinguish an intentional question mark from a damaged or incorrectly transcribed sign.
- Editing constraint
- Source-preserving only; editorial changes permitted; changes require syntax validation; changes require review; unspecified Prevents punctuation correction from changing quoted evidence or executable meaning.
Also called
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 18 findings · 28 questions.
Sign identity Establish what counts as this sign and what evidence identifies an occurrence.
An agent must distinguish a question mark from a resembling shape before interpreting or changing it.
Family boundary
Separate the registered sign from candidate variants and neighbouring punctuation.
Variant membership
Record whether an observed form belongs to this entry, with unresolved membership retained explicitly.
- What evidence establishes that this form belongs to the registered question mark rather than a neighbouring sign? boundary
- Which visual or encoded differences are treated as variants by the convention governing this occurrence? definition
Occurrence evidence
Identify the particular sign present in a source.
Identity support
Keep the source representation and the evidence used to identify it, including disagreements between appearance and encoding.
- Was this occurrence identified from encoded text, an image, handwriting or a source description? provenance
- Do its visible form and recorded encoding support the same identity, and what remains unresolved if they disagree? boundary
Interrogative attachment Record when and how an occurrence functions as punctuation for interrogative content.
Recognising the sign alone does not establish what passage it punctuates or whether its placement is appropriate.
Question scope
Determine the span to which interrogative punctuation applies.
Punctuated span
Associate an interrogative occurrence with its relevant clause, sentence, fragment or quoted span.
- Which exact span does this question mark punctuate? boundary
- Does the sign belong to quoted material, enclosing text or another explicitly identified span? boundary
Placement convention
Assess position and punctuation relationships using the applicable writing convention.
Placement assessment
Record the rule used to judge position, spacing and relationships to nearby punctuation without assuming one universal convention.
- Which language and editorial convention govern this question mark's position and spacing? provenance
- Under that convention, should this occurrence be preserved, moved, supplemented or removed? action
Non-interrogative roles Distinguish uncertainty, literal display and formally assigned meanings from interrogative punctuation.
Treating every question mark as a sentence-level question can corrupt annotations, interface interpretation or formal expressions.
Annotation and display
Identify occurrences used as annotations, examples of the sign or interface symbols.
Displayed role
Record an evidenced non-interrogative role and the target or component supplying its context.
- What evidence shows that this occurrence signals uncertainty, names the sign itself or serves an interface role? definition
- Which assertion, value or interface component supplies the meaning that the sign alone cannot establish? boundary
Formal interpretation
Identify whether a surrounding formal system interprets the occurrence.
Notation-assigned role
Link an operational meaning to the specific notation and parsing context that assign it.
- Which formal notation and version, if applicable, assign a role to this question mark in this position? provenance
- What validation is needed before changing it to establish whether parsing or behaviour would change? action
Representation and legibility Assess whether the recorded and presented forms preserve the intended sign.
An apparent question mark may reflect transcription or rendering problems, while a valid occurrence may be difficult to perceive.
Source fidelity
Compare the current occurrence with its source across recognition and conversion steps.
Transcription integrity
Track whether the question mark was present in the source or appeared during transcription, conversion or substitution.
- Can this occurrence be verified against the original source as an intentional question mark? provenance
- Which recorded transformation first introduced a discrepancy between the source sign and the current representation? provenance
Perceptible form
Assess recognition in the actual presentation context, including relevant nonvisual output.
Recognition in presentation
Record whether intended readers can distinguish the sign and recover its contextual role from the presented output.
- At the intended display or print size, is this question mark distinguishable from nearby marks and plausible lookalikes? measurement
- Where nonvisual access is required, does the tested output convey the sign's relevant role or require an accompanying label? action
Controlled question-mark edits Determine when an occurrence may be changed and what must be checked afterward.
Changing a question mark can alter interrogative force, expressed uncertainty, source fidelity or formal behaviour.
Edit justification
Connect proposed insertion, deletion or replacement to a specific contextual need.
Permitted transformation
Record the intended change, its authority and the interpretation it is expected to preserve or alter.
- Is the proposed change correcting representation, changing punctuation or deliberately changing the meaning supplied by this question mark? action
- Do the source-preservation and editing constraints permit that specific change? boundary
Edit verification
Check the consequences of a question-mark transformation in its surrounding context.
Context preservation check
Verify the resulting attachment, presentation and interpretation against the stated purpose of the edit.
- After the edit, does the surrounding passage or expression retain the intended interrogative, uncertainty-signalling or formal meaning? action
- What comparison, rendering inspection or notation-specific validation demonstrates that the change achieved its intended result? measurement
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Kinds and varieties
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Latin terminal question mark (U+003F ?), the default closer in most Latin-script orthographies
- Inverted opening question mark (U+00BF ¿), paired with ? to bracket Spanish and related interrogatives
- Arabic/RTL question mark (U+061F ؟), a mirrored form used with Arabic, Persian, Urdu, N'Ko, Syriac and Adlam
- Fullwidth East Asian question mark (U+FF1F ?), the official CJK layout form occupying one em
- Greek question mark (U+037E ;), a semicolon-shaped mark canonically equivalent to U+003B
- Armenian question mark (U+055E ՞), a diacritic placed over the stressed vowel of the question word
- Interrobang (U+203D ‽) and inverted interrobang (U+2E18 ⸘), single glyphs combining question and exclamation
- Reversed/percontation mark (U+2E2E ⸮), a historic rhetorical-question form, distinct from the Arabic mirrored mark
- Which of these kinds and varieties hold for the sense of question mark this model covers, and on what evidence? provenance
Identifiers and schemes
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Unicode code point - U+003F - Character name QUESTION MARK; general category Po; script Common; inherited from ASCII.
- ASCII / ISO/IEC 646 - 0x3F (decimal 63) - Single-byte identity mapping in UTF-8.
- HTML named character reference - ? or ? or ? - Inverted form is ¿ (U+00BF).
- URI percent-encoding - %3F - Unencoded ? is a reserved delimiter, not a data character, in RFC 3986.
- Adobe Glyph List - question - PostScript/glyph-name mapping for U+003F.
- Which of these identifiers and schemes hold for the sense of question mark this model covers, and on what evidence? provenance
Standards and regulation
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Unicode Standard (Unicode Consortium) - encodes U+003F and the related interrogative repertoire (¿ ؟ ? ; ՞ ‽ ⁇ ⁈ ⁉ ⸮ ⹔).
- ISO/IEC 10646 Universal Coded Character Set (ISO/IEC JTC 1/SC 2) - aligned with Unicode for these code points.
- ISO/IEC 8859-1 Latin-1 (ISO) - includes inverted question mark 0xBF.
- RFC 3986 Uniform Resource Identifier: Generic Syntax (IETF) - '?' delimits the query component of a URI.
- Ortografía de la lengua española (Real Academia Española, from the 1754 Ortografía) - obligatory paired ¿…? for Spanish interrogatives.
- GB/T 15834 标点符号用法 (Standardization Administration of China / Ministry of Education) - Chinese punctuation, including the fullwidth question mark.
- Lexique des règles typographiques en usage à l'Imprimerie nationale (Imprimerie nationale, France) - thin non-breaking space before ? in French.
- Requirements for Chinese Text Layout (W3C) - fullwidth ? metrics and stacking of repeated marks.
- POSIX.1 / IEEE Std 1003.1 (IEEE and The Open Group) - glob '?' matches exactly one character.
- Which of these standards and regulation hold for the sense of question mark this model covers, and on what evidence? provenance
Real-world use
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Terminates (or, in Spanish, brackets) written interrogative sentences and clauses in print, captions, UI copy and messaging.
- Opens the query string of a URL, e.g. /search?q=term&lang=en, so that following pairs are passed to a script.
- Acts as a one-character wildcard in Unix, DOS and Access filename globbing (contrast * for a string).
- In POSIX/Perl-style regular expressions, marks the preceding atom optional and can make a quantifier lazy.
- Placeholder for bind parameters in JDBC/ODBC-style prepared SQL.
- Prefixed to a diagnosis in English medical notes to mean 'query/possible' (e.g. ?diverticulitis).
- Chess annotation: '?' a bad move, '??' a blunder, '?!' dubious, '!?' interesting.
- Marks uncertain dates or data in running text, often in brackets: Genghis Khan (1162?-1227).
- Shown as a fallback glyph when a character cannot be mapped into the target encoding, or as Apple's flashing-folder '?' when no boot volume is found.
- Which of these real-world use hold for the sense of question mark this model covers, and on what evidence? provenance
Typical measurements
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Unicode scalar value - 63 (U+003F) for the Latin mark; related marks occupy U+00BF, U+061F, U+FF1F and others - code point
- UTF-8 encoded length - 1 octet (0x3F) for U+003F; 2 octets for ¿ and ؟; 3 octets for fullwidth ? - byte
- Typographic advance width - about 0.4-0.6 em for Latin U+003F (East Asian Width: Narrow); 1 em for U+FF1F - em
- French preceding space - narrow/thin non-breaking space before ? in France and Belgium; none in English and usually none in Canadian French - space character
- Which of these typical measurements hold for the sense of question mark this model covers, and on what evidence? provenance
Failure modes and hazards
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Encoding substitution: unmappable characters silently become '?' or '�', destroying the original data.
- Unescaped '?' in a URI path is parsed as the start of the query string (seen as %3F when wrongly encoded in the path).
- Literal '?' treated as a metacharacter in glob, regex, SPARQL variables, or JDBC parameter binding.
- Omitted Spanish opening ¿, which the RAE treats as an error in formal writing and which hides interrogative scope.
- Greek U+037E normalised to semicolon U+003B, so a question mark and a Latin semicolon become identical in processing.
- Visual confusion with glottal-stop letter ʔ (U+0294) in phonetic text and with the replacement character � (U+FFFD).
- Stacked ??? or ?! vs dedicated ⁇/‽ producing inconsistent line-breaking and East Asian width (clreq: three marks should occupy two ems).
- Which of these failure modes and hazards hold for the sense of question mark this model covers, and on what evidence? provenance
Regional variation
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Spanish, Asturian, Galician and some Spanish-influenced languages require inverted ¿ at the start of the interrogative span; informal writing often drops it.
- France and Belgium put a thin non-breaking space before ?; Canadian French generally does not (except typically before the colon).
- Greek and Church Slavonic use a semicolon-shaped question mark; Hebrew is RTL but keeps the Latin-oriented ?.
- Arabic-script and several other RTL orthographies use mirrored ؟; Adlam also has a distinct initial question mark.
- Chinese official usage is fullwidth ?; Korean practice is typically halfwidth; Japanese historically relies on the particle か and still does not sanction ? in government or school texts, though informal use is common.
- Solomon Islands Pidgin may wrap a yes/no question in a pair of question marks because intonation is the only spoken cue.
- Which of these regional variation hold for the sense of question mark this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Exclamation mark (!, U+0021) - Marks exclamatory, not interrogative, illocution; both are terminal punctuation and are combined as ?! or as the interrobang.
- Interrobang (‽, U+203D) - A single invented glyph for question-plus-exclamation; ordinary practice still writes ?! or !? as two characters.
- Semicolon (;, U+003B) - In Latin script it is a pause mark, not a question mark; in Greek the question mark is this same shape and U+037E normalises to it.
- Full stop / period (., U+002E) - Declarative sentence terminator; a question mark occupies the same syntactic slot at the end of a sentence.
- Latin letter glottal stop (ʔ, U+0294) - A phonetic letter resembling '?' without the dot; it is not punctuation and does not terminate a clause.
- Replacement character (�, U+FFFD) - A decoding-failure symbol often drawn as a question mark in a diamond; it is not an interrogative mark and should not be stored as U+003F.
- Inverted question mark (¿, U+00BF) - The opening member of a Spanish pair, not a substitute for the closer; it lives in Latin-1, not in ASCII.
- Asterisk wildcard (*) - In globbing, '*' matches a string of any length while '?' matches exactly one character; they are not interchangeable.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of question mark this model covers, and on what evidence? provenance
Sources
- Question mark - Definition, history (punctus interrogativus), script variants, Unicode inventory, computing/URI/regex/wildcard uses, chess and medical notation, neighbours (glottal stop, replacement character, Greek semicolon).
- Upside-down question and exclamation marks - Spanish paired ¿…? rule from the Real Academia Española (1754), encodings U+00BF/U+00A1, nested ¿¡…!? practice.
- Wildcard character - Question mark as a single-character glob in Unix, MS-DOS and Access, contrasted with * / % / _.
- Requirements for Chinese Text Layout - 中文排版需求 - Fullwidth ? (U+FF1F) as 問號/问号; sequential ??? compressed to two ems; ⁇ (U+2047) taking one em.
- General Punctuation - Related punctuation block: interrobang, double question mark, question-exclamation combinations, emoji presentation of U+2049.
- U+003F QUESTION MARK - Code point properties: Basic Latin, general category Other Punctuation, UTF-8 0x3F, HTML ?, URL %3F, East Asian Width narrow, not mirrored.
What the second pass must settle
- Does this registry entry own only the ordinary question-mark sign, or also inverted, script-specific and other conventional variants?
- Which repeated or combined question-mark forms should be represented as sequences of occurrences, and which belong to separate registered signs?
- Which authoritative language and editorial references should support placement and spacing judgements in the first supported contexts?
- Where should ownership divide between this sign model and neighbouring models for uncertainty annotations and question-mark interface symbols?
- What evidence threshold should establish intentional use when only a rendered image or potentially lossy transcription survives?