← Back to catalogue
Research draft

Turing test

vr.tr.turing-test · ACT.ACT

Enable an AI agent to recognise a Turing-test proposal or trial, assess what its conversational evidence supports, and determine which conclusions or further evaluations are warranted.

Thing Registry Activities and processes

Research draft, second pass

A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.

recalled by Codex without web access - no source was read

Researched by: Codex

Purpose and description

Enable an AI agent to recognise a Turing-test proposal or trial, assess what its conversational evidence supports, and determine which conclusions or further evaluations are warranted.

The Turing test is a family of behavioural tests derived from Alan Turing's imitation game, in which a machine's conversational performance is assessed by whether human judges can distinguish it from a human interlocutor under specified conditions.

It can be Classify a proposed evaluation as a core conversational Turing test, a documented extension or a neighbouring benchmark.; Compare two implementations and identify protocol differences that prevent direct comparison of their results.; Check that judge instructions, communication restrictions and passing rules are specified before interpreting outcomes.; Assess a claimed pass against trial records, human controls and the declared decision rule.; Identify additional trials or controls needed to resolve uncertainty about conversational indistinguishability.; Constrain reports to the tested conditions and distinguish observed performance from philosophical interpretation..

Distinguishing features

A core conversational instance asks human judges to assess whether an interlocutor is human or machine; scoring answer correctness alone does not establish a Turing test.

The protocol restricts identity-revealing channels so that judgments depend on the permitted interaction rather than direct inspection of the participant.

A claimed pass must refer to a specified protocol, judgment rule and population; the label alone supplies no universal passing threshold.

Human conversational performance supplies the comparison target, unlike benchmarks defined solely by task accuracy or expert solutions.

The model distinguishes evidence of successful conversational imitation from additional claims about intelligence, understanding or consciousness.

Scope

+ Definitions and interpretations of the Turing test, with their proponents and disagreements

+ Imitation-game roles, communication restrictions and human comparison conditions

+ Operational protocols, judgment rules and protocol-specific passing criteria

+ Trial records, judge decisions, uncertainty and reproducibility

+ Claims about intelligence or understanding made from test outcomes

- General theories of intelligence, consciousness or personhood beyond their relationship to interpreting this test

- Machine architectures, training procedures and datasets except where they affect a trial's validity

- CAPTCHAs and automated bot-detection systems as independently modelled mechanisms

- General AI capability benchmarks without human-versus-machine conversational judgment

- Embodied or perceptual intelligence evaluations beyond documenting their relationship to Turing-test variants

Characteristics

Interpretive position
Attributed position: operational criterion of intelligence, evidence relevant to intelligence, conversational imitation benchmark, or other explicitly defined interpretation Different interpretations license different conclusions from the same trial outcome.
Protocol lineage
Relationship to an identified formulation, publication or implemented variant The name Turing test can conceal substantive differences in roles, channels and decision rules.
Participant arrangement
Paired human-machine comparison, single-interlocutor classification, or another explicitly documented arrangement Participant arrangement changes the judgment task and its comparison baseline.
Interaction budget
Minutes per encounter and/or permitted conversational turns Brief exchanges and sustained questioning expose different opportunities for discrimination.
Permitted communication
Text-only or specified additional modalities, including timing and metadata visibility Available cues determine whether judges assess conversational behaviour or other identifying signals.
Judge population
Defined recruitment population with language, relevant expertise and prior exposure Results depend on who judges and what those judges know about the task.
Machine-as-human judgment rate
Proportion of eligible machine encounters classified as human, with denominator and uncertainty This records a common outcome without treating it as sufficient evidence of passing every variant.
Human-as-human judgment rate
Proportion of eligible human encounters classified as human, with denominator and uncertainty Human controls reveal judge behaviour and conditions that make even human participants appear machine-like.
Passing rule
Explicit metric, threshold, aggregation method and treatment of abstentions, or no passing rule specified A verdict cannot be assessed independently of the rule that produced it.
Trial assessment state
Proposed, running, completed but unassessed, assessable, inconclusive, or invalid under the stated protocol Completion of an interaction does not establish that its evidence supports a verdict.

Where this came from

wikidata · CC0 1.0

Drafted structure

Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 17 findings · 35 questions.

Meaning and interpretation Records which Turing-test concept is intended and what significance its proponents assign to performance.

An agent must not silently substitute one interpretation of this contested concept for another.

Formulation and lineage

Identifies the formulation being modelled and its relationship to other imitation-game formulations.

Attributed test definition

Record an attributed definition and the roles and judgment task it specifies; historical and later formulations require explicit comparison.

  1. Which source and author define the version of the Turing test used here? provenance
  2. What must the interrogator distinguish, and which participants occupy each role? definition

Contested significance

Separates agreement about a procedure from disagreement about what its success means.

Intelligence claim

Record whether a position treats passing as a definition, sufficient condition, necessary condition or limited evidence of intelligence, together with attributed objections.

  1. Who holds the stated relationship between passing and intelligence, and how do they justify it? provenance
  2. Would this position allow an intelligent system to fail or a system without understanding to pass? boundary
Imitation-game design Defines the conversational encounter that gives a trial its identity.

Roles, channels and conversational constraints determine which discrimination task was actually tested.

Roles and comparison

Specifies judges, machine interlocutors, human controls and the information given to each.

Judge task and human control

Record whether judges compare participants or classify separate encounters, including disclosed machine prevalence and human-control instructions.

  1. Do judges choose between paired participants or classify each interlocutor independently? definition
  2. What are judges told about participant identities, machine prevalence and the purpose of the interaction? measurement
  3. What are human controls instructed to do, and how closely do their conditions match the machine's? measurement

Conversation envelope

Records the channels, duration, topics and assistance available during an encounter.

Permitted behaviour and cues

Make the interaction limits explicit, including personas, topic restrictions, external tools, response timing and visible metadata.

  1. Which modalities, topics, time limits and turn limits define the permitted conversation? definition
  2. Which tools, human assistance, personas or response transformations are permitted for each participant? boundary
  3. Could formatting, latency or interface metadata reveal identity independently of conversational content? measurement
Judgment and verdict Connects individual judge responses to a defensible protocol-specific outcome.

A headline passing claim is uninterpretable without its scoring rule, human comparison and uncertainty.

Decision rule

Defines how classification, confidence and abstention responses become reported outcomes.

Declared passing criterion

Record the chosen success criterion and when it was fixed, without assuming every Turing-test variant shares it.

  1. What exact metric, threshold and aggregation rule determine a pass? definition
  2. Was the rule fixed before outcomes were inspected, and where is that version recorded? provenance
  3. How are uncertain judgments, abstentions and incomplete encounters counted? measurement

Outcome and uncertainty

Records machine and human outcomes with the statistical context needed to interpret them.

Comparative judgment evidence

Preserve outcome counts, denominators, judge-level dependence and uncertainty rather than only a passing percentage.

  1. How often were machine and human interlocutors each classified as human, and over how many eligible encounters? measurement
  2. How do uncertainty estimates account for repeated judges, repeated systems and participant assignment? measurement
  3. Does the evidence support the stated criterion, or is the verdict inconclusive? action
Trial integrity and replication Assesses whether recorded conversations and judgments support the claimed trial conditions.

Identity leakage, selective reporting and an unrecorded system configuration can undermine a claimed result.

Blinding and trial record

Checks identity concealment, assignment records and completeness of retained evidence.

Auditable encounters

Link judgments to transcripts and assignment records while identifying exclusions, protocol deviations and possible identity disclosures.

  1. Can every included judgment be traced to an encounter, assignment and applicable instruction set? provenance
  2. Which encounters were excluded or interrupted, and were those decisions governed by a declared rule? measurement
  3. What evidence shows whether judges received unintended identity cues? measurement

System and population stability

Determines which system configuration and judge population a result describes.

Replicable test conditions

Record system version, prompting, memory, assistance and judge recruitment sufficiently to assess replication and limits of transfer.

  1. Which system version, instructions, memory state and external assistance produced the tested conversations? provenance
  2. How were judges recruited, and what language skills, expertise or prior system exposure characterised them? measurement
  3. Which conditions must a replication preserve, and which should it vary to test robustness? action
Inference and test boundaries Controls how trial outcomes are interpreted and how neighbouring evaluations are classified.

Conversational indistinguishability does not by itself settle every claim associated with machine intelligence.

Limits of passing

Separates observed judge behaviour from broader theoretical and practical conclusions.

Licensed conclusions

Identify the narrow result established under the protocol and the additional premises required for claims about understanding, general competence or consciousness.

  1. What conclusion follows directly from the observed judgments under these particular conditions? boundary
  2. Which additional premises connect that outcome to a claim about intelligence or understanding, and who defends them? provenance
  3. Which complementary evaluation would address an important capability claim left unresolved by this trial? action

Variants and neighbours

Distinguishes conversational test variants from extensions and evaluations with a different target.

Variant membership

Classify an evaluation by its actual judgment task and channels, documenting departures from the selected core formulation.

  1. Does the evaluation retain human judgment of human-versus-machine conversational identity? boundary
  2. If it adds embodiment, perception or another task, which attributed formulation makes it an extension of the Turing test? provenance
  3. Should its outcome be compared with the core conversational test, reported separately or linked to a neighbouring model? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.

A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.

Reported evidence

Findings from the breadth pass, kept separate from the structural claims.

Check these first

Recalled without web access and unsourced; every item is a lead to verify.

  • This describes the conversational machine-intelligence sense; there is no single universally accepted protocol or passing threshold.
  • Behavioural and functionalist interpretations give successful performance evidential weight, while critics dispute whether it establishes understanding; neither interpretation should be presented as settled.
  • The relationship between Turing's original imitation game and later test formulations is interpretively disputed; historical details and named variants should be checked against their original texts.
  1. Which of these check these first hold for the sense of Turing test this model covers, and on what evidence? provenance

Kinds and varieties

Recalled without web access and unsourced; every item is a lead to verify.

  • Text-only conversational Turing test
  • Restricted-domain conversational Turing test
  • Total Turing test, extending evaluation to perception and physical interaction
  1. Which of these kinds and varieties hold for the sense of Turing test this model covers, and on what evidence? provenance

Real-world use

Recalled without web access and unsourced; every item is a lead to verify.

  • Framing philosophical debate about machine intelligence and behavioural evidence.
  • Evaluating whether conversational systems can produce apparently human dialogue.
  • Organising human-versus-machine judging competitions, including the historical Loebner Prize.
  • Studying how judges attribute intelligence and humanness to conversational partners.
  1. Which of these real-world use hold for the sense of Turing test this model covers, and on what evidence? provenance

Typical measurements

Recalled without web access and unsourced; every item is a lead to verify.

  • Machine conversations classified as human - 0-100; interpretation depends on the protocol and human comparison condition - %
  • Judge accuracy in distinguishing humans from machines - 0-100; chance performance depends on the judging task - %
  • Conversation duration - Protocol-dependent; no universally required duration - minutes
  1. Which of these typical measurements hold for the sense of Turing test this model covers, and on what evidence? provenance

Failure modes and hazards

Recalled without web access and unsourced; every item is a lead to verify.

  • Mistaking successful human imitation for proof of understanding, consciousness, or general intelligence.
  • Results varying with judge expertise, language, conversation length, topic restrictions, and participant instructions.
  • Evasion, deliberate mistakes, or fabricated personas exploiting judges' expectations without demonstrating broad competence.
  • Reporting that a system passed without specifying the protocol, comparison group, threshold, or uncertainty.
  • Treating failure to imitate human conversational habits as evidence that a system lacks useful intelligence.
  1. Which of these failure modes and hazards hold for the sense of Turing test this model covers, and on what evidence? provenance

Regional variation

Recalled without web access and unsourced; every item is a lead to verify.

  • Language, dialect, and cultural expectations affect judgments of conversational humanness and limit comparisons across populations.
  1. Which of these regional variation hold for the sense of Turing test this model covers, and on what evidence? provenance

Neighbouring kinds and how to tell them apart

Recalled without web access and unsourced; every item is a lead to verify.

  • Imitation game - Turing's original presentation begins with a game involving concealed male and female participants; the later label Turing test usually denotes human-machine conversational discrimination.
  • CAPTCHA - A CAPTCHA automatically screens users for access control, whereas a conventional Turing test uses human judges to assess machine imitation of human conversation.
  • Chinese room argument - This is a philosophical argument about whether symbol manipulation entails understanding, rather than an empirical conversational test.
  • Artificial general intelligence - General intelligence concerns competence across a broad range of tasks; conversational indistinguishability is a narrower behavioural criterion.
  1. Which of these neighbouring kinds and how to tell them apart hold for the sense of Turing test this model covers, and on what evidence? provenance

What the second pass must settle

  • Which primary and scholarly sources best resolve differences between Turing's imitation-game formulations and later standard conversational interpretations?
  • Which attributed positions treat passing as necessary, sufficient or merely evidential for intelligence, and where do their arguments diverge?
  • Which implemented protocols justify their passing thresholds through human controls, statistical reasoning or historical convention?
  • How strongly do judge expertise, disclosed machine prevalence, personas and conversation duration change outcomes across replicated studies?
  • Which independently auditable trials support generalisation beyond a particular system configuration, judge population and interaction setting?