speech
(language) communication by word of mouth
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
recalled by Codex without web access - no source was read
Researched by: Codex
Purpose and description
Enable an agent to recognise spoken language, assess how reliably it communicates in context, and decide how it may be interpreted, relayed or transformed.
Speech is the production and use of articulated vocal signals to express linguistic meaning in spoken communication.
It can be Segment spoken exchanges into attributable utterances while retaining overlap and uncertain boundaries.; Transcribe speech with timestamps, alternatives and markers for inaudible or uncertain wording.; Identify candidate languages and interpret utterances against their conversational context.; Request repetition or clarification when wording, reference or intended action remains unclear.; Translate, caption or summarise speech while preserving attribution and marking lost or uncertain vocal cues.; Relay or reuse speech within established permissions while distinguishing original delivery from transformed output..
Distinguishing features
The communication uses spoken linguistic expression; a written transcript represents speech but is not itself a spoken occurrence.
Laughter, coughing and other vocal sounds do not become speech solely because they communicate something; their relationship to spoken language must be established.
The entry covers spoken communication generally, including brief conversational turns, rather than only a formal address called 'a speech'.
A repeated sentence can have the same wording while constituting a different utterance with a different speaker, time, addressee or communicative force.
Human, synthetic and replayed voices require separate origin labels; an audible voice alone does not establish a live human speaker.
Scope
+ Spoken utterances and their boundaries within an exchange
+ Linguistic content, language varieties and changes between languages
+ Vocal delivery, timing and prosodic cues relevant to interpretation
+ Speaker and addressee roles, turn-taking and conversational repair
+ Live, recorded and mediated delivery of spoken language
+ Interpretation uncertainty and conditions governing capture, attribution and reuse
- Language as an abstract grammatical and semantic system
- Writing and signed languages as distinct modes of linguistic expression
- Nonlinguistic vocal sounds considered independently of accompanying speech
- Speech anatomy, disorders, diagnosis and treatment
- The recording file, microphone or transmission infrastructure as an independently managed object
- A prepared address as an authored work, and its publication or edition history
Characteristics
- Language and variety
- Named or coded languages and varieties by segment; mixed, uncertain or unidentified Determines which linguistic resources and interpretive assumptions are appropriate.
- Utterance extent
- Start and end offsets in seconds relative to a stated recording or interaction clock, with boundary uncertainty Makes claims about wording, timing and speaker changes traceable to a particular stretch of speech.
- Speaking rate
- Words per minute or syllables per second, with segmentation and pause-inclusion method stated Supports pacing and accessibility decisions without treating different measurement methods as interchangeable.
- Intelligibility
- Proportion of words correctly recovered under a stated listener task and listening condition; unknown when untested Separates evidence of listener understanding from signal quality or automated transcription confidence.
- Prosodic delivery
- Observed stress, intonation, rhythm and pause patterns, with annotation method and uncertainty Preserves vocal cues that wording alone may fail to convey.
- Participant roles
- Attributed speaker or voice source, intended addressee and other participants; unknown roles allowed Supports interpretation of references and turns while separating voice attribution from verified identity.
- Delivery mode
- In-person live, remotely transmitted live, recorded playback, synthetic rendering or mixed Determines whether immediate response is possible and what origin evidence is needed.
- Interpretation status
- Uninterpreted, provisional, participant-confirmed, disputed or unresolved, attached to a specific interpretation Prevents a plausible reading from being treated as established speaker intent.
- Reuse conditions
- Links to applicable consent, permissions and restrictions, specifying action, audience and duration; unknown where unestablished Lets an agent distinguish permission to listen from permission to record, share, quote or generate a derivative.
Analytical facets
- substance
- information
- origin
- conceptual
- agency
- inert
- mobility
- not-applicable
- scale
- not-applicable
- affordances
- readable
Also called
+105
Where this came from
oewn:2024 · CC BY 4.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 19 findings · 29 questions.
Utterance and linguistic expression Establishes which spoken occurrence is being modelled and what linguistic material it contains.
Speech unfolds in time, and its boundaries and wording cannot always be recovered as clean sentences.
Spoken occurrence
Locates an utterance within an interaction or recording and distinguishes it from repetitions and representations.
Utterance boundaries
Record the evidence used to delimit speech, retaining incomplete turns, interruptions and uncertain endpoints.
- What counts as the modelled unit here: a turn, an intonation unit, an utterance or another explicitly defined segment? definition
- Where does the segment begin and end, and how are overlap and uncertain boundaries represented? measurement
Spoken wording
Captures language choice and recoverable expression without silently regularising the speaker's delivery.
Language and wording evidence
Separate observed speech, transcription hypotheses and edited wording, including code-switching and disfluencies where relevant.
- Which languages or varieties occur in each segment, and what evidence supports their identification? definition
- Which words, repetitions, fillers or unfinished constructions are recoverable, and which remain uncertain or were editorially normalised? provenance
Vocal delivery and intelligibility Represents the audible delivery and the conditions under which listeners can recover it.
Speech communicates through timing and vocal patterns as well as words, and its recoverability depends on listening conditions.
Prosody and timing
Records vocal emphasis and temporal organisation relevant to the particular utterance.
Audible delivery cues
Describe stress, intonation, pauses and rate while distinguishing acoustic observations from interpretations of attitude or emotion.
- Which stress, pitch movement or pause patterns distinguish possible readings of this utterance? measurement
- Which proposed meanings of these vocal cues depend on language, context or listener judgement rather than observation alone? boundary
Listener recoverability
Assesses whether wording can be understood by the intended listeners under specified conditions.
Intelligibility and obstruction
Record listener evidence and obstacles such as masking noise, overlap or transmission loss without conflating them with speaker competence.
- How much wording can the relevant listeners recover, under what conditions and using what assessment method? measurement
- Would repetition, slower delivery, reduced noise or accompanying text address the observed difficulty? action
Meaning in conversation Connects spoken wording to references, communicative acts and the evolving exchange.
The same words can function differently depending on who says them, what preceded them and how participants respond.
Reference and communicative force
Tracks contextual dependencies and candidate interpretations of what a speaker is doing by speaking.
Contextual interpretation
Separate recoverable wording from proposed referents and communicative force, retaining alternatives when context is insufficient.
- What context resolves expressions such as 'I', 'here', 'that one' or 'tomorrow', and which references remain unresolved? definition
- What evidence supports interpreting the utterance as a request, assertion, question, quotation, joke or another communicative act? provenance
Turns and repair
Represents participation, uptake and attempts to correct problems in speaking or understanding.
Interactional response
Track turn relations, backchannels, interruptions and clarification sequences without assuming that silence or acknowledgement establishes agreement.
- How does this turn relate to the preceding speech, and what participant response provides evidence of understanding or misunderstanding? provenance
- Should an agent answer, wait, request clarification or confirm an interpretation before acting on this utterance? action
Voice origin and transformation Traces whose expression is represented and how speech reaches its present audible or textual form.
A voice may be live, replayed, dubbed or synthesised, and a transcript or translation can introduce changes that affect attribution.
Speaker and voice origin
Distinguishes audible voice sources from verified people and from authors of the wording.
Attribution evidence
Record the basis and limits of speaker attribution, including cases where a person reads another author's words or a system renders them.
- Who or what produced the audible voice, who supplied the wording, and which of these attributions are verified? provenance
- Is this a live utterance, playback, synthetic rendering or composite, and what evidence establishes that status? boundary
Representations and derivatives
Relates the spoken occurrence to recordings, transcripts, captions, translations and edits.
Transformation trace
Make omissions, wording changes and alignment limits visible across representations of the same speech.
- Which source segment supports this transcript, caption, translation or edited audio, and what transformations were applied? provenance
- Which timing, emphasis, ambiguity or participant cues were lost or altered enough to affect interpretation? boundary
Audience and permitted use Connects speech to its intended audience and the established conditions for agent handling.
Hearing an utterance does not by itself establish its intended audience or authorise every subsequent use.
Address and exposure
Distinguishes intended recipients from incidental listeners and later audiences.
Audience context
Record whether speech is addressed to the agent, another participant or a wider audience, and whether capture changes its exposure.
- To whom was this speech addressed, and was the agent an intended recipient or an incidental listener? boundary
- Would the proposed relay, quotation or storage expose the speech to an audience beyond the established context? action
Capture and reuse conditions
Records action-specific authority and restrictions relevant to processing spoken communication.
Authorised speech handling
Keep the basis for recording, transcribing, sharing and deriving new output explicit, with unresolved conditions marked as unknown.
- What established consent, permission or applicable rule governs recording, transcription, retention and sharing in this context? provenance
- Does the proposed use, including generated voice output where relevant, fall within those conditions or require additional authority? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Check these first
Recalled without web access and unsourced; every item is a lead to verify.
- This is recall-based information; no sources or current standard editions were checked.
- Speaking-rate figures are illustrative, not universal norms; word counts are not directly comparable across all languages.
- The listed kinds overlap rather than forming an exclusive taxonomy; sung speech occupies a boundary with singing.
- Which of these check these first hold for the sense of speech this model covers, and on what evidence? provenance
Kinds and varieties
Recalled without web access and unsourced; every item is a lead to verify.
- Spontaneous conversational speech
- Prepared or read-aloud speech
- Monologic speech
- Whispered speech
- Sung speech
- Which of these kinds and varieties hold for the sense of speech this model covers, and on what evidence? provenance
Identifiers and schemes
Recalled without web access and unsourced; every item is a lead to verify.
- ISO 639-3 - Three-letter language codes - Identifies the language used, not an individual speech event or speech as a medium.
- BCP 47 language tags - Language subtags optionally combined with script, region and variant subtags - Labels the language of speech content; does not uniquely identify a speaker, recording or utterance.
- Which of these identifiers and schemes hold for the sense of speech this model covers, and on what evidence? provenance
Standards and regulation
Recalled without web access and unsourced; every item is a lead to verify.
- International Phonetic Alphabet - International Phonetic Association; provides notation for transcribing speech sounds.
- ITU-T Recommendation P.800 - International Telecommunication Union; specifies methods for subjective assessment of transmission quality, including speech quality.
- IEC 60268-16 - International Electrotechnical Commission; specifies objective rating of speech intelligibility using the Speech Transmission Index.
- Which of these standards and regulation hold for the sense of speech this model covers, and on what evidence? provenance
Real-world use
Recalled without web access and unsourced; every item is a lead to verify.
- Face-to-face conversation and social coordination
- Teaching, explanation and oral transmission of knowledge
- Broadcasting, public address and oral performance
- Telephone and voice communication services
- Voice interfaces, dictation and speech recognition
- Which of these real-world use hold for the sense of speech this model covers, and on what evidence? provenance
Typical measurements
Recalled without web access and unsourced; every item is a lead to verify.
- Speaking rate - Approximately 100-200 in many ordinary adult speaking contexts; strongly dependent on language, task and treatment of pauses - words per minute
- Word intelligibility score - 0-100; the observed result depends on listeners, material and listening conditions - percent correctly identified words
- Speech Transmission Index - 0-1, the scale bounds rather than a typical conversational range - dimensionless
- Which of these typical measurements hold for the sense of speech this model covers, and on what evidence? provenance
Failure modes and hazards
Recalled without web access and unsourced; every item is a lead to verify.
- Noise, reverberation and competing voices can mask speech and reduce intelligibility.
- Ambiguity, unfamiliar vocabulary and differences in pronunciation can produce misunderstanding.
- Speech production or hearing difficulties can impede communication, with effects depending on available accommodations.
- Transmission dropouts, clipping and audio processing artifacts can remove or distort meaningful cues.
- Overheard, recorded or intercepted speech can disclose private information.
- Which of these failure modes and hazards hold for the sense of speech this model covers, and on what evidence? provenance
Regional variation
Recalled without web access and unsourced; every item is a lead to verify.
- Languages and regional varieties differ in speech sounds, phonological contrasts, rhythm, stress and intonation.
- Conventions for turn-taking, politeness, address and acceptable silence vary across communities.
- Multilingual communities vary in patterns of code-switching and the spoken varieties used in different settings.
- Which of these regional variation hold for the sense of speech this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Recalled without web access and unsourced; every item is a lead to verify.
- language - Language is a system of linguistic expression; speech is its vocal realization, while language can also be signed or written.
- voice - Voice concerns vocal sound production and its qualities; speech requires linguistic articulation and can occur without vocal-fold vibration, as in whispering.
- utterance - An utterance is a particular communicative production; speech is the broader activity or medium.
- public speech - A public speech is a particular address or genre; the requested sense includes everyday spoken communication.
- speech recording - A recording preserves a particular realization of speech on a carrier; it is distinct from the communicative activity and its linguistic content.
- nonlinguistic vocalization - Sounds such as coughing or unarticulated laughter need not encode words or linguistic structure, although they may accompany speech.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of speech this model covers, and on what evidence? provenance
What the second pass must settle
- Does the authoritative registry treatment include sung linguistic expression, private self-talk and synthetic spoken output under this sense of speech, or require explicit links to neighbouring entries?
- Which existing Vercy world model, if any, already owns spoken communication and should be reused rather than duplicated?
- What segmentation conventions best accommodate overlap, unfinished utterances and differences across languages without imposing written-sentence boundaries?
- Which intelligibility measures are suitable for the intended listeners and languages, and how should their results remain distinct from automated recognition confidence?
- Which jurisdiction-specific and context-specific sources must establish capture and reuse conditions, especially for incidental speech and generated voice derivatives?