← Back to catalogue
Research draft

hate speech

vr.tr.hate-speech · ACT.ACT

Enable an AI agent to recognise possible hate speech, record the evidence and interpretive framework behind that assessment, evaluate contextual harm, and select proportionate responses.

Thing Registry Activities and processes

Research draft, second pass

A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.

recalled by Codex without web access - no source was read

Researched by: Codex

Purpose and description

Enable an AI agent to recognise possible hate speech, record the evidence and interpretive framework behind that assessment, evaluate contextual harm, and select proportionate responses.

Hate speech is a contested category of expression that attacks, demeans or promotes hostility, discrimination or violence against people on grounds of group identity, with definitions differing over covered characteristics, context and thresholds of harm.

It can be Assess an expression against a named definition and explain which criteria are supported or unresolved.; Compare outcomes across definitions without collapsing their differences.; Request missing conversational, linguistic, cultural, or dissemination context.; Preserve necessary evidence with controlled access and minimal repetition of harmful content.; Recommend proportionate responses within the agent's authority, including contextualisation, review, or escalation.; Reassess a decision when an appeal, corrected translation, or new contextual evidence arrives..

Distinguishing features

Identify whether hostility targets people through actual or attributed group identity; disagreement with an idea alone does not establish this connection.

Test whether the expression performs an attack, degradation, exclusion, threat, or encouragement of hostility under a named definition, rather than treating offensiveness as sufficient.

Determine whether the speaker endorses the hostile message or quotes, rejects, documents, satirises, or reclaims it.

Resolve coded meanings through contextual evidence rather than classifying an isolated word or symbol automatically.

Assess descriptive classification, platform-rule violation, and legal restriction separately because their criteria may differ.

Scope

+ Communicative acts that attack, demean, exclude, threaten, or encourage hostility toward people through attributed group identity

+ Competing definitions and their treatment of targets, expression, intent, context, and harm

+ Explicit, coded, visual, audiovisual, and multimodal expressions

+ Distinguishing endorsement from reporting, quotation, counterspeech, satire, and reclamation

+ Evidence, uncertainty, and proportionate handling of suspected instances

- Prejudice or hatred without an identifiable communicative act

- Harassment lacking the identity connection required by the applicable definition

- Discrimination in employment, housing, or services as conduct beyond the expression itself

- Violent acts and operational threat management beyond the speech assessment

- General misinformation assessment and platform governance as whole systems

Characteristics

Assessment framework
Named definition, research codebook, community standard, or legal instrument with issuer, version, and applicability Makes the criteria and limits of a classification explicit.
Target and identity connection
Targeted people or group; attributed identity; explicit, contextual, disputed, or unresolved connection Separates identity-directed hostility from neighbouring forms of abuse or criticism.
Communicative function
Derogation, dehumanisation, collective blame, exclusion, threat, encouragement of hostility, other, or unresolved; multiple values permitted Records what the expression does rather than relying on a keyword match.
Speaker stance toward the hostile content
Endorsement, rejection, reporting, quotation with unclear stance, satire, reclamation, mixed, or unresolved Helps prevent documentation and resistance from being classified as endorsement.
Expression and coding
Text, speech, image, gesture, audio, video, or mixed; explicit, coded, or ambiguous Identifies which linguistic, cultural, and multimodal evidence interpretation requires.
Observed dissemination
Views, recipients, shares, or repetitions within a stated observation period; unknown where unavailable Supports assessment of exposure while keeping measured reach distinct from assumed influence.
Contextual harm assessment
Evidence-supported description of potential harm, affected people, urgency, and uncertainty Supports response decisions without treating offensive language as a complete measure of risk.
Classification outcome
Meets, does not meet, or unresolved under each named framework, with reasons and evidence gaps Preserves disagreement and avoids presenting one assessment as universally authoritative.

Also called

Hate speech in Japanhate speech in the United Statesanti-LGBTQ rhetoriconline hate speechincitement to genocide

Where this came from

wikidata · CC0 1.0

Drafted structure

Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 18 findings · 28 questions.

Definitions and boundaries Records what counts as hate speech under particular frameworks and where those frameworks disagree.

The registry has no definition, and a useful model must expose competing criteria rather than silently select one.

Framework criteria

Identifies the authority, scope, and decision criteria of each definition.

Definition-dependent classification

Record each assessment against an identified framework, including which identities and communicative functions it covers.

  1. Whose definition is being applied, and what target, expression, intent, or harm criteria does it require? definition
  2. Which original document and version establish those criteria? provenance

Neighbouring expression

Tests boundaries with insult, criticism, harassment, threats, and discriminatory advocacy.

Identity and attack boundary

Require an explained connection between the target, attributed identity, and hostile communicative function.

  1. What evidence connects the attack to people's identity rather than solely to their actions, ideas, or institutional role? boundary
  2. Would the expression meet the selected definition if its offensive vocabulary were replaced by neutral wording? boundary
Targets and hostile functions Represents who is targeted and what the communication seeks or appears to do to them.

Hate-speech assessment needs more precision than a label of negative sentiment.

Target identification

Identifies individual, collective, indirect, and intersecting targets.

Attributed group membership

Record the identity attributed by the expression without asserting that the target actually holds it.

  1. Who is targeted, and which actual or attributed group membership makes them the target? definition
  2. Is the target named directly, inferred through a proxy, or unresolved, and what supports that interpretation? provenance

Hostile communicative function

Distinguishes forms of identity-directed hostility and their combinations.

Attack and exclusion mechanisms

Describe the expression's function, including degradation, dehumanisation, collective blame, exclusion, threats, or encouragement of hostile action.

  1. Which hostile function is evidenced by the expression, and which specific content supports it? definition
  2. Does the message express contempt, advocate unequal treatment, threaten harm, or encourage others to act, and where is the distinction uncertain? boundary
Meaning and context Determines how the expression functions within its conversation, language, medium, and audience.

Identical words or symbols can support different assessments depending on stance and context.

Quotation and stance

Separates hostile content appearing in a message from the communicator's position toward it.

Endorsement versus reference

Record evidence of endorsement, rejection, documentation, satire, or reclamation without treating a claimed purpose as conclusive.

  1. Does the surrounding communication endorse the hostile content or use it for reporting, counterspeech, satire, or reclamation? boundary
  2. Which conversational turns, captions, or other contextual evidence establish the speaker's stance? provenance

Coded and multimodal meaning

Examines indirect references, local usage, translation, and interactions between modalities.

Context-supported decoding

Treat proposed coded meanings as interpretations requiring evidence and consideration of plausible alternatives.

  1. What evidence supports the alleged hostile meaning of a euphemism, symbol, meme, or combination of text and imagery in this setting? provenance
  2. Could dialect, translation, local usage, or missing audiovisual context materially change the assessment? boundary
Exposure and harm Assesses circulation, audience context, and plausible pathways from expression to harm.

Classification alone does not establish the severity, urgency, or consequences of an instance.

Circulation and audience

Records observed dissemination and relevant speaker-audience relationships.

Exposure without assumed impact

Separate measured circulation from estimates of influence or audience response.

  1. What reach, repetition, or amplification was observed, over what period, and through which measurement source? measurement
  2. What evidence shows the speaker's authority or influence with this audience, rather than merely their follower count? provenance

Harm pathways and urgency

Examines potential intimidation, exclusion, discrimination, and violence in the specific setting.

Evidence-based harm assessment

Record plausible harm pathways, observed effects, and uncertainty separately; identify concrete signals requiring urgent review.

  1. Which harms are observed, which are anticipated, and what evidence connects them to this expression? measurement
  2. Do specific targets, calls to act, timing, or surrounding events warrant urgent threat review under the applicable procedure? action
Decisions and contestability Connects supported assessments to authorised responses and revisable decisions.

An agent must distinguish interpreting speech from having authority to restrict it and must preserve routes for correction.

Response selection

Evaluates responses against the applicable rules, evidence, urgency, and potential adverse effects.

Proportionate authorised response

Record the basis for contextualisation, review, distribution limits, removal, or escalation as distinct decisions.

  1. What action is authorised by the applicable rule and supported by the evidence and assessed urgency? action
  2. Could the proposed response suppress reporting, counterspeech, or a targeted community's own expression, and how should that affect the decision? action

Review and evidence

Preserves the grounds for a decision and enables correction without unnecessary redistribution.

Revisable assessment record

Retain necessary evidence, framework versions, reasons, and unresolved disagreements with appropriate access controls.

  1. What minimum evidence and context must be retained so another reviewer can reconstruct the assessment? provenance
  2. What new evidence, appeal, or framework change would require reconsideration, and who can revise the outcome? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.

A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.

Reported evidence

Findings from the breadth pass, kept separate from the structural claims.

Check these first

Recalled without web access and unsourced; every item is a lead to verify.

  • This is a recall-based description; no sources were consulted.
  • There is no single universally accepted definition; a researcher should specify whether the intended category is legal, institutional or analytical.
  • The listed kinds overlap, and current jurisdiction-specific law and platform policies require verification before application.
  1. Which of these check these first hold for the sense of hate speech this model covers, and on what evidence? provenance

Kinds and varieties

Recalled without web access and unsourced; every item is a lead to verify.

  • Identity-based insults and derogation
  • Dehumanizing or demonizing expression
  • Advocacy of exclusion or discrimination
  • Incitement to hatred or violence
  • Identity-based threats
  • Coded or implicit expressions of group hatred
  1. Which of these kinds and varieties hold for the sense of hate speech this model covers, and on what evidence? provenance

Standards and regulation

Recalled without web access and unsourced; every item is a lead to verify.

  • United Nations: International Covenant on Civil and Political Rights, Article 19 protects freedom of expression subject to specified restrictions; Article 20(2) requires prohibition by law of advocacy of national, racial or religious hatred constituting incitement to discrimination, hostility or violence.
  • United Nations: International Convention on the Elimination of All Forms of Racial Discrimination, Article 4 addresses racist propaganda, incitement and organizations.
  • United Nations Office of the High Commissioner for Human Rights: Rabat Plan of Action provides a six-part threshold test for assessing incitement, covering context, speaker, intent, content and form, extent, and likelihood including imminence.
  • Council of Europe Committee of Ministers: Recommendation CM/Rec(2022)16 on combating hate speech distinguishes levels of severity and corresponding responses.
  • Council of the European Union: Framework Decision 2008/913/JHA addresses certain forms and expressions of racism and xenophobia through criminal law.
  1. Which of these standards and regulation hold for the sense of hate speech this model covers, and on what evidence? provenance

Real-world use

Recalled without web access and unsourced; every item is a lead to verify.

  • A category in platform content policies and moderation decisions.
  • A concept used in evaluating potentially unlawful incitement and discriminatory expression.
  • A research category for studying prejudice, intergroup hostility and online abuse.
  • A monitoring category in atrocity prevention and assessments of threats to targeted communities.
  • A focus of education, counterspeech and support for people targeted by identity-based abuse.
  1. Which of these real-world use hold for the sense of hate speech this model covers, and on what evidence? provenance

Typical measurements

Recalled without web access and unsourced; every item is a lead to verify.

  • Proportion of sampled content classified as hate speech under a specified definition - No universal typical range; depends on sampling, platform, language and classification rules - Percent of sampled items
  • Precision and recall of a hate-speech classifier against an annotated reference set - No universal typical range; depends on the dataset and annotation standard - Dimensionless proportion
  1. Which of these typical measurements hold for the sense of hate speech this model covers, and on what evidence? provenance

Failure modes and hazards

Recalled without web access and unsourced; every item is a lead to verify.

  • Expression can intimidate targeted people, normalize discrimination and contribute to conditions supporting violence; effects depend on context.
  • Overbroad definitions can suppress legitimate criticism, dissent, satire or discussion of religion and politics.
  • Quotation, counterspeech and reclaimed slurs can be misclassified when context is ignored.
  • Automated systems can miss coded hostility and disproportionately flag language used by marginalized communities.
  • Inconsistent definitions, sampling and annotation can make prevalence estimates misleading or incomparable.
  1. Which of these failure modes and hazards hold for the sense of hate speech this model covers, and on what evidence? provenance

Regional variation

Recalled without web access and unsourced; every item is a lead to verify.

  • United States constitutional doctrine generally does not recognize hate speech as a standalone exception to free-speech protection; particular expression may fall within categories such as true threats or incitement.
  • European jurisdictions commonly restrict some racist or other identity-based expression, but protected characteristics, thresholds and penalties vary.
  • Across jurisdictions, treatment of religious hatred, criticism of religion and blasphemy differs; these categories should not be treated as interchangeable.
  1. Which of these regional variation hold for the sense of hate speech this model covers, and on what evidence? provenance

Neighbouring kinds and how to tell them apart

Recalled without web access and unsourced; every item is a lead to verify.

  • Offensive speech - Offensiveness alone does not establish an identity-based attack or satisfy a particular hate-speech definition.
  • Incitement - Incitement concerns prompting others toward prohibited conduct or outcomes; it need not involve group identity, and many definitions of hate speech include expression below an incitement threshold.
  • Harassment - Harassment concerns abusive or intimidating conduct toward people and need not be based on group identity.
  • Hate crime - A hate crime requires a criminal offense with the relevant bias element; expression classified as hate speech is not necessarily criminal.
  • Discrimination - Discrimination concerns unequal treatment or disadvantage on specified grounds; hate speech concerns expression, though an act can involve both.
  • Criticism of beliefs or institutions - Criticizing ideas, practices or institutions does not by itself constitute an attack on people because of their identity; wording and context determine possible overlap.
  1. Which of these neighbouring kinds and how to tell them apart hold for the sense of hate speech this model covers, and on what evidence? provenance

What the second pass must settle

  • Which reference definitions should anchor this registry entry, and how do their identity, intent, and harm criteria differ?
  • How should the model represent disagreement over attacks on dominant groups, institutions, religions, and ideologies without conflating criticism with attacks on people?
  • What evidence is sufficient to classify emerging coded expressions across languages and communities?
  • Which methods can assess cumulative harm from repeated identity-directed expression without attributing collective effects to an isolated message without support?
  • How should practical response thresholds balance uncertainty, urgency, and the risk of suppressing documentation, counterspeech, or reclaimed language?