hate speech
Enable an AI agent to recognise possible hate speech, record the evidence and interpretive framework behind that assessment, evaluate contextual harm, and select proportionate responses.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
recalled by Codex without web access - no source was read
Researched by: Codex
Purpose and description
Enable an AI agent to recognise possible hate speech, record the evidence and interpretive framework behind that assessment, evaluate contextual harm, and select proportionate responses.
Hate speech is a contested category of expression that attacks, demeans or promotes hostility, discrimination or violence against people on grounds of group identity, with definitions differing over covered characteristics, context and thresholds of harm.
It can be Assess an expression against a named definition and explain which criteria are supported or unresolved.; Compare outcomes across definitions without collapsing their differences.; Request missing conversational, linguistic, cultural, or dissemination context.; Preserve necessary evidence with controlled access and minimal repetition of harmful content.; Recommend proportionate responses within the agent's authority, including contextualisation, review, or escalation.; Reassess a decision when an appeal, corrected translation, or new contextual evidence arrives..
Distinguishing features
Identify whether hostility targets people through actual or attributed group identity; disagreement with an idea alone does not establish this connection.
Test whether the expression performs an attack, degradation, exclusion, threat, or encouragement of hostility under a named definition, rather than treating offensiveness as sufficient.
Determine whether the speaker endorses the hostile message or quotes, rejects, documents, satirises, or reclaims it.
Resolve coded meanings through contextual evidence rather than classifying an isolated word or symbol automatically.
Assess descriptive classification, platform-rule violation, and legal restriction separately because their criteria may differ.
Scope
+ Communicative acts that attack, demean, exclude, threaten, or encourage hostility toward people through attributed group identity
+ Competing definitions and their treatment of targets, expression, intent, context, and harm
+ Explicit, coded, visual, audiovisual, and multimodal expressions
+ Distinguishing endorsement from reporting, quotation, counterspeech, satire, and reclamation
+ Evidence, uncertainty, and proportionate handling of suspected instances
- Prejudice or hatred without an identifiable communicative act
- Harassment lacking the identity connection required by the applicable definition
- Discrimination in employment, housing, or services as conduct beyond the expression itself
- Violent acts and operational threat management beyond the speech assessment
- General misinformation assessment and platform governance as whole systems
Characteristics
- Assessment framework
- Named definition, research codebook, community standard, or legal instrument with issuer, version, and applicability Makes the criteria and limits of a classification explicit.
- Target and identity connection
- Targeted people or group; attributed identity; explicit, contextual, disputed, or unresolved connection Separates identity-directed hostility from neighbouring forms of abuse or criticism.
- Communicative function
- Derogation, dehumanisation, collective blame, exclusion, threat, encouragement of hostility, other, or unresolved; multiple values permitted Records what the expression does rather than relying on a keyword match.
- Speaker stance toward the hostile content
- Endorsement, rejection, reporting, quotation with unclear stance, satire, reclamation, mixed, or unresolved Helps prevent documentation and resistance from being classified as endorsement.
- Expression and coding
- Text, speech, image, gesture, audio, video, or mixed; explicit, coded, or ambiguous Identifies which linguistic, cultural, and multimodal evidence interpretation requires.
- Observed dissemination
- Views, recipients, shares, or repetitions within a stated observation period; unknown where unavailable Supports assessment of exposure while keeping measured reach distinct from assumed influence.
- Contextual harm assessment
- Evidence-supported description of potential harm, affected people, urgency, and uncertainty Supports response decisions without treating offensive language as a complete measure of risk.
- Classification outcome
- Meets, does not meet, or unresolved under each named framework, with reasons and evidence gaps Preserves disagreement and avoids presenting one assessment as universally authoritative.
Also called
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 18 findings · 28 questions.
Definitions and boundaries Records what counts as hate speech under particular frameworks and where those frameworks disagree.
The registry has no definition, and a useful model must expose competing criteria rather than silently select one.
Framework criteria
Identifies the authority, scope, and decision criteria of each definition.
Definition-dependent classification
Record each assessment against an identified framework, including which identities and communicative functions it covers.
- Whose definition is being applied, and what target, expression, intent, or harm criteria does it require? definition
- Which original document and version establish those criteria? provenance
Neighbouring expression
Tests boundaries with insult, criticism, harassment, threats, and discriminatory advocacy.
Identity and attack boundary
Require an explained connection between the target, attributed identity, and hostile communicative function.
- What evidence connects the attack to people's identity rather than solely to their actions, ideas, or institutional role? boundary
- Would the expression meet the selected definition if its offensive vocabulary were replaced by neutral wording? boundary
Targets and hostile functions Represents who is targeted and what the communication seeks or appears to do to them.
Hate-speech assessment needs more precision than a label of negative sentiment.
Target identification
Identifies individual, collective, indirect, and intersecting targets.
Attributed group membership
Record the identity attributed by the expression without asserting that the target actually holds it.
- Who is targeted, and which actual or attributed group membership makes them the target? definition
- Is the target named directly, inferred through a proxy, or unresolved, and what supports that interpretation? provenance
Hostile communicative function
Distinguishes forms of identity-directed hostility and their combinations.
Attack and exclusion mechanisms
Describe the expression's function, including degradation, dehumanisation, collective blame, exclusion, threats, or encouragement of hostile action.
- Which hostile function is evidenced by the expression, and which specific content supports it? definition
- Does the message express contempt, advocate unequal treatment, threaten harm, or encourage others to act, and where is the distinction uncertain? boundary
Meaning and context Determines how the expression functions within its conversation, language, medium, and audience.
Identical words or symbols can support different assessments depending on stance and context.
Quotation and stance
Separates hostile content appearing in a message from the communicator's position toward it.
Endorsement versus reference
Record evidence of endorsement, rejection, documentation, satire, or reclamation without treating a claimed purpose as conclusive.
- Does the surrounding communication endorse the hostile content or use it for reporting, counterspeech, satire, or reclamation? boundary
- Which conversational turns, captions, or other contextual evidence establish the speaker's stance? provenance
Coded and multimodal meaning
Examines indirect references, local usage, translation, and interactions between modalities.
Context-supported decoding
Treat proposed coded meanings as interpretations requiring evidence and consideration of plausible alternatives.
- What evidence supports the alleged hostile meaning of a euphemism, symbol, meme, or combination of text and imagery in this setting? provenance
- Could dialect, translation, local usage, or missing audiovisual context materially change the assessment? boundary
Exposure and harm Assesses circulation, audience context, and plausible pathways from expression to harm.
Classification alone does not establish the severity, urgency, or consequences of an instance.
Circulation and audience
Records observed dissemination and relevant speaker-audience relationships.
Exposure without assumed impact
Separate measured circulation from estimates of influence or audience response.
- What reach, repetition, or amplification was observed, over what period, and through which measurement source? measurement
- What evidence shows the speaker's authority or influence with this audience, rather than merely their follower count? provenance
Harm pathways and urgency
Examines potential intimidation, exclusion, discrimination, and violence in the specific setting.
Evidence-based harm assessment
Record plausible harm pathways, observed effects, and uncertainty separately; identify concrete signals requiring urgent review.
- Which harms are observed, which are anticipated, and what evidence connects them to this expression? measurement
- Do specific targets, calls to act, timing, or surrounding events warrant urgent threat review under the applicable procedure? action
Decisions and contestability Connects supported assessments to authorised responses and revisable decisions.
An agent must distinguish interpreting speech from having authority to restrict it and must preserve routes for correction.
Response selection
Evaluates responses against the applicable rules, evidence, urgency, and potential adverse effects.
Proportionate authorised response
Record the basis for contextualisation, review, distribution limits, removal, or escalation as distinct decisions.
- What action is authorised by the applicable rule and supported by the evidence and assessed urgency? action
- Could the proposed response suppress reporting, counterspeech, or a targeted community's own expression, and how should that affect the decision? action
Review and evidence
Preserves the grounds for a decision and enables correction without unnecessary redistribution.
Revisable assessment record
Retain necessary evidence, framework versions, reasons, and unresolved disagreements with appropriate access controls.
- What minimum evidence and context must be retained so another reviewer can reconstruct the assessment? provenance
- What new evidence, appeal, or framework change would require reconsideration, and who can revise the outcome? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Check these first
Recalled without web access and unsourced; every item is a lead to verify.
- This is a recall-based description; no sources were consulted.
- There is no single universally accepted definition; a researcher should specify whether the intended category is legal, institutional or analytical.
- The listed kinds overlap, and current jurisdiction-specific law and platform policies require verification before application.
- Which of these check these first hold for the sense of hate speech this model covers, and on what evidence? provenance
Kinds and varieties
Recalled without web access and unsourced; every item is a lead to verify.
- Identity-based insults and derogation
- Dehumanizing or demonizing expression
- Advocacy of exclusion or discrimination
- Incitement to hatred or violence
- Identity-based threats
- Coded or implicit expressions of group hatred
- Which of these kinds and varieties hold for the sense of hate speech this model covers, and on what evidence? provenance
Standards and regulation
Recalled without web access and unsourced; every item is a lead to verify.
- United Nations: International Covenant on Civil and Political Rights, Article 19 protects freedom of expression subject to specified restrictions; Article 20(2) requires prohibition by law of advocacy of national, racial or religious hatred constituting incitement to discrimination, hostility or violence.
- United Nations: International Convention on the Elimination of All Forms of Racial Discrimination, Article 4 addresses racist propaganda, incitement and organizations.
- United Nations Office of the High Commissioner for Human Rights: Rabat Plan of Action provides a six-part threshold test for assessing incitement, covering context, speaker, intent, content and form, extent, and likelihood including imminence.
- Council of Europe Committee of Ministers: Recommendation CM/Rec(2022)16 on combating hate speech distinguishes levels of severity and corresponding responses.
- Council of the European Union: Framework Decision 2008/913/JHA addresses certain forms and expressions of racism and xenophobia through criminal law.
- Which of these standards and regulation hold for the sense of hate speech this model covers, and on what evidence? provenance
Real-world use
Recalled without web access and unsourced; every item is a lead to verify.
- A category in platform content policies and moderation decisions.
- A concept used in evaluating potentially unlawful incitement and discriminatory expression.
- A research category for studying prejudice, intergroup hostility and online abuse.
- A monitoring category in atrocity prevention and assessments of threats to targeted communities.
- A focus of education, counterspeech and support for people targeted by identity-based abuse.
- Which of these real-world use hold for the sense of hate speech this model covers, and on what evidence? provenance
Typical measurements
Recalled without web access and unsourced; every item is a lead to verify.
- Proportion of sampled content classified as hate speech under a specified definition - No universal typical range; depends on sampling, platform, language and classification rules - Percent of sampled items
- Precision and recall of a hate-speech classifier against an annotated reference set - No universal typical range; depends on the dataset and annotation standard - Dimensionless proportion
- Which of these typical measurements hold for the sense of hate speech this model covers, and on what evidence? provenance
Failure modes and hazards
Recalled without web access and unsourced; every item is a lead to verify.
- Expression can intimidate targeted people, normalize discrimination and contribute to conditions supporting violence; effects depend on context.
- Overbroad definitions can suppress legitimate criticism, dissent, satire or discussion of religion and politics.
- Quotation, counterspeech and reclaimed slurs can be misclassified when context is ignored.
- Automated systems can miss coded hostility and disproportionately flag language used by marginalized communities.
- Inconsistent definitions, sampling and annotation can make prevalence estimates misleading or incomparable.
- Which of these failure modes and hazards hold for the sense of hate speech this model covers, and on what evidence? provenance
Regional variation
Recalled without web access and unsourced; every item is a lead to verify.
- United States constitutional doctrine generally does not recognize hate speech as a standalone exception to free-speech protection; particular expression may fall within categories such as true threats or incitement.
- European jurisdictions commonly restrict some racist or other identity-based expression, but protected characteristics, thresholds and penalties vary.
- Across jurisdictions, treatment of religious hatred, criticism of religion and blasphemy differs; these categories should not be treated as interchangeable.
- Which of these regional variation hold for the sense of hate speech this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Recalled without web access and unsourced; every item is a lead to verify.
- Offensive speech - Offensiveness alone does not establish an identity-based attack or satisfy a particular hate-speech definition.
- Incitement - Incitement concerns prompting others toward prohibited conduct or outcomes; it need not involve group identity, and many definitions of hate speech include expression below an incitement threshold.
- Harassment - Harassment concerns abusive or intimidating conduct toward people and need not be based on group identity.
- Hate crime - A hate crime requires a criminal offense with the relevant bias element; expression classified as hate speech is not necessarily criminal.
- Discrimination - Discrimination concerns unequal treatment or disadvantage on specified grounds; hate speech concerns expression, though an act can involve both.
- Criticism of beliefs or institutions - Criticizing ideas, practices or institutions does not by itself constitute an attack on people because of their identity; wording and context determine possible overlap.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of hate speech this model covers, and on what evidence? provenance
What the second pass must settle
- Which reference definitions should anchor this registry entry, and how do their identity, intent, and harm criteria differ?
- How should the model represent disagreement over attacks on dominant groups, institutions, religions, and ideologies without conflating criticism with attacks on people?
- What evidence is sufficient to classify emerging coded expressions across languages and communities?
- Which methods can assess cumulative harm from repeated identity-directed expression without attributing collective effects to an isolated message without support?
- How should practical response thresholds balance uncertainty, urgency, and the risk of suppressing documentation, counterspeech, or reclaimed language?