transcription
Enable an agent to recognise a transcription task, record its source and representation rules, assess fidelity, and determine whether the resulting transcript is fit for its intended use.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
recalled by Codex without web access - no source was read
Researched by: Codex
Purpose and description
Enable an agent to recognise a transcription task, record its source and representation rules, assess fidelity, and determine whether the resulting transcript is fit for its intended use.
In its linguistic and documentary sense, transcription is the production of a written representation of speech or an existing text according to conventions governing what is preserved, interpreted, or normalized.
It can be Select and partition source material into traceable transcription assignments; Apply an explicit transcription convention to create or revise text; Align transcript segments with source times or page regions; Flag ambiguous readings and route them for targeted review; Compare a transcript with its source or an adjudicated reference; Accept a version for a stated use or return specific passages for correction.
Distinguishing features
The output represents identifiable linguistic content in a source; newly composed text without such a source is not transcription.
The output preserves source language under declared notation rules; rendering meaning in another language introduces translation.
Omissions and normalisations must be accountable to a declared convention; selecting only main ideas produces a summary instead.
The object modelled is the source-to-text activity; its resulting transcript is a related information artefact.
Speech recognition, optical character recognition, and manual entry can perform parts of the activity; the method alone does not establish transcript fidelity.
Scope
+ Identifying source material, its modality, and the portions selected for transcription
+ Choosing orthographic, phonetic, diplomatic, or edited transcription conventions
+ Producing text with speaker attribution, segmentation, and source alignment where required
+ Representing uncertainty, unintelligible passages, and editorial intervention
+ Reviewing and correcting a transcript against its source
- Biological transcription from a nucleic acid template
- Musical transcription into notation or for another instrumentation
- Translation into a different language
- Summarisation, interpretation, and substantive rewriting of source content
- Recording equipment and the capture of source media as independent activities
- Publication, archival custody, and downstream analysis of completed transcripts
Characteristics
- Source modality
- recorded speech | live speech | signed discourse | manuscript | printed text | mixed Determines what evidence can be inspected and which transcription conventions are applicable.
- Source and selected extent
- Source identifier with time intervals, page regions, or other stable locators Makes coverage and individual readings traceable to the material being transcribed.
- Representation convention
- Named and versioned convention specifying notation, omissions, normalisation, and editorial additions Defines which differences from the source are intentional and which count as errors.
- Language and notation
- Source languages or varieties, output script, and any phonetic or sign notation Supports assignment to suitable transcribers and interpretation of the output.
- Production method
- manual | automated | assisted; associated tool or model version where applicable Records how the initial reading was produced and where review may be needed.
- Review state
- unstarted | in progress | awaiting review | reviewed with unresolved passages | accepted for stated use Prevents an unreviewed draft from being treated as an accepted transcript.
- Coverage
- Percentage of selected source extent processed, with denominator and excluded spans recorded Separates completion from correctness and exposes unprocessed material.
- Unresolved readings
- Count of marked spans, with source locations and uncertainty types Shows where the transcript still requires evidence or expert judgement.
- Evaluated transcription error
- Metric and value, with reference transcript, comparison unit, normalisation rules, and evaluated extent Makes quality claims interpretable without assuming one metric suits every transcription convention.
Also called
Where this came from
wikidata · CC0 1.0
Also registered as vr.tr.transcription-process
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 18 findings · 28 questions.
Source and task boundary Establishes what linguistic material is being transcribed and what completion covers.
A transcript cannot be assessed for fidelity or completeness without a bounded source.
Source identification
Records the material from which readings must be recoverable.
Identifiable source evidence
Record the source version, modality, languages, and stable locators available for checking the transcript.
- Which recording, live event, image, or physical text supplies the linguistic content? provenance
- Which source version and locators allow a reviewer to recover the evidence for a passage? provenance
Selected coverage
Distinguishes assigned material from intentional exclusions and unfinished work.
Bounded transcription assignment
Record selected intervals or regions and identify gaps without silently treating them as absent source content.
- Which speakers, time intervals, pages, marginal notes, or other regions belong to this assignment? boundary
- How much of the selected extent has been processed, and which spans remain excluded or unfinished? measurement
Representation conventions Defines what the written representation preserves, normalises, annotates, or omits.
Two transcripts may differ legitimately because they serve different purposes and follow different conventions.
Notation and detail
Specifies the transcription system and its required level of linguistic detail.
Declared representation policy
Identify the notation and rules for phenomena such as pronunciation, disfluency, spelling, layout, and signed expression.
- Which convention defines the required orthographic, phonetic, diplomatic, or other representation? definition
- Which source features must be retained, and which may be normalised or omitted? boundary
Editorial intervention
Makes additions and transformations distinguishable from source content.
Visible editorial decisions
Record how supplied punctuation, expanded abbreviations, corrected spelling, and explanatory insertions are represented.
- How are supplied words, expansions, corrections, and explanatory annotations distinguished from directly represented content? definition
- Which proposed edits would cross this task's boundary into translation, summary, or rewriting? boundary
Segmentation and attribution Connects passages to source locations and, where relevant, to speakers or textual hands.
Correct words alone may be insufficient when their order, source position, or attribution is wrong.
Source alignment
Defines transcript units and their correspondence to source spans.
Traceable transcript segments
Record segmentation and alignment at the resolution required for checking or downstream use.
- Are segments defined by turns, utterances, words, lines, pages, or another explicit rule? definition
- What timestamp or region precision is required, and how is alignment error assessed? measurement
Speaker and hand attribution
Handles attribution without converting an uncertain distinction into an asserted identity.
Supported content attribution
Separate local speaker or hand labels from verified identities and account for overlap or uncertain attribution.
- What evidence supports each speaker or hand label and any association with a named person? provenance
- How should simultaneous speech, changing hands, and uncertain attribution be represented? action
Uncertainty and recovery Preserves the distinction between readable content, competing readings, and unavailable evidence.
Transcription must expose uncertainty so plausible completions do not become unsupported source claims.
Uncertain passages
Locates and classifies passages whose reading is not settled.
Explicit unresolved readings
Mark uncertainty at the affected span and distinguish illegibility, unintelligibility, missing source, and competing interpretations.
- Which exact spans are unresolved, and what prevents a reliable reading of each? measurement
- How does the convention distinguish an uncertain candidate reading from wholly unrecoverable content? definition
Evidence-assisted recovery
Governs attempts to resolve difficult passages while retaining their evidential basis.
Reviewable reading recovery
Record the source inspection, contextual evidence, or specialist judgement used to revise an uncertain reading.
- Which replay, image enhancement, parallel witness, vocabulary resource, or specialist review could resolve this passage? action
- What evidence supports the revised reading, and does any uncertainty remain? provenance
Production and fidelity review Records how the transcript was produced, checked, corrected, and accepted.
A completed text and an automated confidence score do not by themselves establish fitness for a particular use.
Transcription production
Tracks the method and revisions that produced the current transcript.
Recoverable production history
Associate transcript versions with their human or automated production steps and subsequent corrections.
- Who or what produced the initial transcription, using which tools, versions, and relevant settings? provenance
- Which passages were subsequently corrected, by whom, and on what source evidence? provenance
Fidelity and acceptance
Evaluates the transcript against its declared convention and intended use.
Use-specific quality decision
Record review coverage, error criteria, unresolved defects, and the basis for accepting or returning the transcript.
- What review method and comparison rules assess wording, omissions, attribution, and alignment for this use? measurement
- Which remaining defects require correction before acceptance, and who can make that decision? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Check these first
Recalled without web access and unsourced; every item is a lead to verify.
- No sense is recorded, and ACT.PRC alone does not distinguish linguistic, documentary, biological, or musical transcription; this description provisionally covers linguistic and documentary uses.
- The varieties overlap: a phonetic transcript can also be verbatim, and conventions determine what fidelity means.
- Standards are recalled rather than checked; their editions, precise scope, and applicability should be verified.
- Which of these check these first hold for the sense of transcription this model covers, and on what evidence? provenance
Kinds and varieties
Recalled without web access and unsourced; every item is a lead to verify.
- Verbatim transcription of speech
- Edited or normalized transcription of speech
- Phonetic transcription
- Phonemic transcription
- Conversation-analytic transcription
- Diplomatic transcription of written documents
- Which of these kinds and varieties hold for the sense of transcription this model covers, and on what evidence? provenance
Standards and regulation
Recalled without web access and unsourced; every item is a lead to verify.
- ISO 24624, issued by the International Organization for Standardization, addresses transcription of spoken language.
- The International Phonetic Alphabet, maintained by the International Phonetic Association, provides symbols and conventions for representing speech sounds.
- The TEI Guidelines, maintained by the Text Encoding Initiative Consortium, provide conventions for encoding transcribed speech and documentary texts.
- Which of these standards and regulation hold for the sense of transcription this model covers, and on what evidence? provenance
Real-world use
Recalled without web access and unsourced; every item is a lead to verify.
- Creating searchable records of interviews, meetings, and oral histories
- Documenting speech for linguistic analysis
- Preparing records of court proceedings
- Producing reference text for training and evaluating speech-recognition systems
- Making manuscripts and archival documents accessible as machine-readable text
- Which of these real-world use hold for the sense of transcription this model covers, and on what evidence? provenance
Typical measurements
Recalled without web access and unsourced; every item is a lead to verify.
- Word error rate against a reference transcript - Task-dependent; zero indicates agreement, and insertion errors can make the value exceed 100 - %
- Transcription effort relative to recording duration - Depends on recording quality, notation detail, language, and workflow - labour hours per recording hour
- Which of these typical measurements hold for the sense of transcription this model covers, and on what evidence? provenance
Failure modes and hazards
Recalled without web access and unsourced; every item is a lead to verify.
- Omitting, substituting, or inserting words, especially under noise or overlapping speech
- Assigning speech to the wrong speaker
- Normalizing speech in ways that erase analytically important features or change apparent meaning
- Exposing confidential or personally identifying information
- Treating an uncertain reading or an automatically generated passage as verified source content
- Which of these failure modes and hazards hold for the sense of transcription this model covers, and on what evidence? provenance
Regional variation
Recalled without web access and unsourced; every item is a lead to verify.
- Orthography, punctuation, and dialect representation follow language-specific and regional conventions.
- Court-reporting requirements and transcript certification practices differ by jurisdiction.
- Privacy and consent requirements for recordings and transcripts differ by jurisdiction and use.
- Which of these regional variation hold for the sense of transcription this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Recalled without web access and unsourced; every item is a lead to verify.
- translation - Translation expresses content in another language; transcription ordinarily represents the source language.
- transliteration - Transliteration maps writing between scripts; speech transcription maps spoken material into writing or notation.
- automatic speech recognition - Automatic speech recognition is a computational method for recognizing speech; transcription is the broader activity and can be performed manually.
- captioning - Captioning adds presentation requirements such as synchronization and often speaker or non-speech sound information.
- biological transcription - Biological transcription synthesizes RNA from a DNA template rather than representing speech or documents.
- musical transcription - Musical transcription represents music in notation or adapts it for another performing medium.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of transcription this model covers, and on what evidence? provenance
What the second pass must settle
- Does vr.tr.transcription denote linguistic transcription, biological transcription, musical transcription, or a broader process concept?
- Does an existing Vercy world model already cover the intended sense and therefore need to be linked instead of duplicated?
- Should signed-language notation, manuscript transcription, and speech transcription share this registry entry, or are any already owned by neighbouring entries?
- Which recognised transcription conventions and practitioner sources should ground the preservation, normalisation, and uncertainty rules?
- Which intended uses require particular review coverage, error measures, attribution evidence, or acceptance thresholds?