AI Model Evaluation
Provide the format-neutral context an agent needs to plan, execute, record, interpret, approve and disclose an evaluation of a specified AI model or system version, and to retain the resulting evidence.
Bundle → Layer → Finding → Questions Filled
5 bundles · 13 layers · 28 findings · 112 questions
Evaluation Scope and Mandate What is being evaluated, for what claim, under whose authority, and by whom.
Subject and Claim Scope
The exact object under test and the claim the evaluation is designed to support or refute.
Evaluation subject binding
The immutable binding between the evaluation and the model or system version, weights, configuration and inference settings under test.
- Which authoritative identifier and version designate the exact model or system under evaluation? identity
- Is the evaluated unit the bare model, a guarded endpoint, or the full deployed system including scaffolding and tools? classification
- Which inference configuration, decoding parameters, system prompt, tooling and safety filters form part of the evaluated subject? composition
- How is a hosted or continuously updated endpoint pinned so results remain attributable to a fixed subject? temporal
Intended purpose and claim scope
The intended purpose, deployment context and the specific claim the evaluation supports, with explicit out-of-scope uses and generalisability limits.
- What precise claim about capability, safety or quality is this evaluation designed to support or refute? definition
- Which documented intended purpose, deployment context and affected populations bound the validity of the claim? relationship
- Which uses, populations, languages or environments are explicitly outside the evaluated claim? constraint
- What documented limits on generalisability apply outside the tested distribution? quality
Mandate, Actors and Independence
Why the evaluation exists, what triggered it, and who is competent and independent enough to perform it.
Evaluation mandate and trigger
The legal, contractual, policy or research basis obliging the evaluation, the event that triggered this instance, and the decision that consumes it.
- Under which legal, regulatory, contractual or internal-policy basis is this evaluation required or commissioned? authority
- Which event triggered this evaluation instance — release gate, capability threshold, incident, drift alarm or scheduled review? event
- Which decision or gate consumes the result, and who is accountable for that decision? decision
Assessor identity, competence and independence
Who designed, executed, scored and reviewed the evaluation, their competence evidence, degree of independence from the development team, and declared conflicts.
- Which named organisations and roles designed, executed, scored and reviewed this evaluation? ownership
- What degree of independence separates the evaluators from the development team, and how is it evidenced? classification
- What evidence establishes that assessors and human raters are competent for the domain and risk being measured? evidence
- Which conflicts of interest were declared and how were they managed? constraint
Evaluation Design The protocol: method class, task and dataset binding, metrics, disaggregation, acceptance criteria and human or adversarial procedures, all fixed before results are observed.
Method Class and Task Binding
Which method class is used and how tasks, datasets, splits and ground truth are bound to the evaluation.
Evaluation method class and baseline
Classification of the evaluation method — statistical, formal, empirical, human-subject or adversarial — with justification, declared blind spots and comparison baselines.
- Which method class is applied, and why is it appropriate to the property being assessed? classification
- Which properties can this method class not evidence, and what complementary method covers that gap? constraint
- Against which baseline, reference implementation or comparator system are results interpreted? measurement
- How is the method shown to reflect the state of the art at the time of execution? requirement
Task and dataset binding with contamination control
Version-pinned reference to evaluation tasks, datasets and splits, the ground-truth definition, and controls against train/test contamination and held-out set erosion.
- Which dataset version, snapshot or split identifiers, and which checksums, are bound to this evaluation? relationship
- How is ground truth or the reference answer defined, produced and quality-assured for each item? definition
- What controls establish that the evaluation items were not present in training or tuning data? validation
- How is a held-out or canary split protected from repeated exposure that would invalidate future results? access
Scenario, task and adaptation protocol
A HELM scenario is a task plus domain plus language. Inspect Task composes dataset, solver or agent, and scorer, optionally with a sandbox and tools. Standardized few-shot or other adaptation is required so comparisons attribute performance to the subject. UK AISI autonomous-system submissions must be Inspect tasks and should expose a binary success scorer.
- What task, domain, language or modality, and user-facing use case define this scenario, and which HELM or Inspect eval identifier does it correspond to? identity
- What adaptation protocol (zero-shot, few-shot, chain-of-thought, agent loop, tool use) is frozen, including prompt templates and in-context example policy? process
- If the task is agentic, which tools, MCP servers and sandbox (Docker, Kubernetes or other) isolate untrusted model actions, and what approval gates exist for tool calls? security
- Does the scenario depend on a geographic, legal or language locale, and if so which locale constraints apply to items and raters? spatial
Metrics, Disaggregation and Acceptance
How results are computed, how they are broken down, and what counts as pass or fail.
Metric specification
Formal specification of each metric: construct measured, computation, aggregation, reporting precision and the scorer implementation and version producing it.
- What is the exact computational definition, unit and aggregation rule for each reported metric? definition
- Which construct does the metric claim to measure, and what evidence links the metric to that construct? quality
- Which scorer implementation, version and configuration produced the metric values? provenance
- What rounding, significant-figure and tie-breaking conventions govern reported values? constraint
Disaggregation factors and subgroup analysis
Factors along which results must be broken down, the sufficiency rules for subgroup reporting, and the disparity comparisons performed.
- Which factors must results be disaggregated by, and how was that factor list justified? composition
- What minimum sample size per subgroup is required before a disaggregated result is reportable? measurement
- Which disparity comparison is applied across subgroups, and what gap is treated as a finding? validation
- How are protected attributes obtained or inferred for disaggregation, and under which lawful basis? privacy
Acceptance criteria and thresholds
Pass/fail criteria, probabilistic thresholds, risk tiers and safety margins fixed and recorded before results are observed, with change control for later alteration.
- Which acceptance thresholds and probabilistic criteria were fixed before results were observed, and where is that pre-registration recorded? requirement
- Who set each threshold, on what basis, and with what safety margin? authority
- How is a change to a threshold after results are known authorised, justified and disclosed? lifecycle
- What consequence follows automatically from breaching each threshold? decision
Human and Adversarial Protocols
Procedures for evaluations whose validity depends on people: raters and red teams.
Human evaluation protocol
Recruitment, representativeness, instructions, blinding, agreement measurement and participant protection for human-rated evaluation.
- How were raters recruited and how does the rater population relate to the affected population? composition
- Which inter-rater agreement statistic is computed and what value is required for the result to be usable? measurement
- Which human-subject protections, consent arrangements and ethics approvals govern the evaluation? privacy
- How are raters blinded to system identity and condition to prevent expectancy effects? process
Adversarial and red-team protocol
Threat model, attacker capability and knowledge assumptions, attack-class coverage, elicitation effort and stopping rules for adversarial testing.
- Which threat model, attacker goals, capabilities and knowledge assumptions define this adversarial test? security
- Which attack classes — evasion, poisoning, privacy, prompt injection, misaligned output — are in and out of scope? classification
- How much elicitation effort, tooling and time was expended, and how is that effort evidenced as sufficient? measurement
- What stopping rule ends the exercise, and how are unresolved attack avenues recorded? constraint
Execution, Results and Evidence Run records, environment capture, result and item-level evidence, uncertainty, integrity and provenance.
Run Execution and Reproducibility
The record of an individual execution and what is needed for an independent party to repeat it.
Run record, configuration and resource accounting
Identity of a single run, its complete configuration, timing, and the compute, token, latency and energy quantities attributable to it.
- Which identifier uniquely designates this run, and how does it relate to the campaign and to repeated epochs? identity
- Which configuration values — seeds, sampling parameters, batch, concurrency, epochs, prompt template — are captured with the run? composition
- What are the run start and end times, and how are they distinguished from result publication and record ingestion times? temporal
- What compute, token, latency, monetary and energy quantities are recorded for the run? measurement
Environment capture and replication
Software and hardware environment, dependency versions, bounded non-determinism, replication tolerance and declared irreproducible components.
- Which code revision, dependency versions, container image and hardware description are recorded for the run? provenance
- Which sources of non-determinism are permitted, and how are they bounded or seeded? constraint
- What exact procedure and inputs allow an independent party to replicate the run, and what tolerance defines success? process
- Which parts of the run are irreproducible, such as a retired endpoint or a one-off rater panel, and how is that recorded? exception
Results, Item Evidence and Uncertainty
What was measured, the item-level evidence behind it, and how confident the numbers are.
Result record, aggregation and item-level evidence
Reported metric values keyed by subject, task, split, metric and factor, the traceable aggregation path from per-item scores, and the retained per-item inputs, outputs, traces and scores.
- How is a single reported result value uniquely identified across subject, task, split, metric and factor? identity
- What is the traceable path from per-item scores through reducers to each headline value? composition
- Which fields are retained per evaluated item, and which are required to re-score without re-running inference? evidence
- When a result is recomputed or corrected, how are the prior value, the scorer version and the reason retained? state
Uncertainty and statistical validity
Sample size adequacy, variance, interval estimates, significance rules and multiple-comparison control governing what the numbers may be said to show.
- Which uncertainty estimate accompanies each headline value, and by what procedure was it computed? measurement
- How was the number of items and repetitions determined to be adequate for the claimed resolution? validation
- Under which rule is a difference between two systems or two versions declared meaningful rather than noise? constraint
- How are multiple comparisons across metrics, factors and checkpoints controlled? quality
Scoring and judging
Inspect scoring maps output to a Score against a Target, with built-in match, multiple-choice, math, model-graded and perplexity scorers, custom scorers, multiple scorers and epoch reductions. Model grading uses another model as judge. Human raters and CJE-style calibration of judges against oracles are first-class. ISO/IEC TR 29119-11 highlights the oracle problem when expected results are hard to specify. Refusals may be logged separately from task success.
- Which scorer or judge configuration produced each sample score, including grader model, rubric, and whether scoring was live or deferred and later re-scored? process
- For each sample, what were the input, output, target, score value, explanation and refusal or scanner flags? measurement
- If a model or crowd judge was used, what agreement with gold oracles, inter-rater reliability and bias controls support treating those judgments as evidence? quality
- Do sample inputs, outputs or rater notes contain personal data or prohibited content, and what minimization or redaction applies? privacy
Evidence Integrity and Provenance
Packaging, sealing and lineage that make the evidence transferable and auditable.
Evidence package and integrity
Bundling, manifesting, digesting and sealing of evaluation evidence so it can be transferred to auditors or authorities without loss of integrity.
- Which items constitute the sealed evidence package for this evaluation, and how is completeness asserted? composition
- Which digest, signature or timestamping mechanism protects the package, and who holds the signing authority? security
- How is a later alteration of any packaged artefact detected and reported? validation
- In what neutral form is the package transferred to an auditor or authority without loss of meaning? interoperability
Provenance and lineage graph
Graph linking datasets, subject artefacts, runs, scorers, results, findings and decisions to responsible agents and times, using entity/activity/agent semantics.
- Which generation, derivation, usage and attribution edges must be recorded between evaluation entities? relationship
- Which agent — person, organisation or software — is recorded as responsible for each activity and artefact? ownership
- How are activity start and end times recorded and reconciled with the time the record reached the evaluation store? temporal
- How is lineage preserved when datasets, models or scorers are governed in external systems? interoperability
Interpretation, Findings and Decision Turning results into findings with declared validity limits, and into authorised decisions with lifecycle control.
Findings and Validity
Atomic findings and the declared threats to their validity.
Evaluation finding record
An atomic, separately citable finding derived from results: what was observed, its severity and likelihood, the evidence behind it, and its state.
- How is a single finding identified and cited independently of the report that first presented it? identity
- Which specific results, items and traces substantiate the finding? evidence
- Which severity and likelihood scale classifies the finding, and who assigns it? classification
- What states can a finding occupy — open, accepted, mitigated, disputed, withdrawn — and what moves it between them? state
Validity threats and declared limitations
Threats to construct, internal and external validity, benchmark saturation or gaming, disclosure of inconclusive results, and periodic review of the measurement approach itself.
- Which threats to construct, internal and external validity are declared for this evaluation? quality
- How is benchmark saturation, gaming or optimisation against the test set assessed and disclosed? validation
- How are inconclusive or negative results recorded so they cannot be silently dropped? evidence
- How and when is the effectiveness of the measurement approach itself reviewed? process
Decision, Waivers and Lifecycle
The authorised decision that consumes the evaluation, any conditional release, and the states and triggers that govern repetition.
Approval decision, conditions and waivers
The authorised approve, approve-with-conditions, reject or defer decision, its signatories, the evidence version reviewed, the authorised scope, and any time-bounded waiver with compensating controls.
- Who holds the authority to approve or reject on the basis of this evaluation, and under which delegation? authority
- What decision was taken, on which sealed evidence version, and with what conditions attached? decision
- Which deployment scope, markets, populations and time window does the authorisation cover? constraint
- On what documented basis may release proceed when an acceptance criterion is not met, and when does that waiver expire? exception
Evaluation lifecycle and re-evaluation triggers
States an evaluation instance passes through, the conditions requiring repetition or invalidation, and the cadence of periodic re-evaluation and in-production measurement.
- Which states does an evaluation instance pass through from planned to archived, and which transitions are irreversible? lifecycle
- Which changes to model, data, threat landscape, scaffolding or deployment context oblige re-evaluation? event
- At what cadence is periodic re-evaluation and in-production measurement required? temporal
- Under what conditions is a completed evaluation declared invalid, and how are dependent approvals affected? state
Assurance, Records Governance and Interoperability Disclosure to differentiated audiences, external assurance, access and retention of records, and alignment and exchange with external frameworks.
Reporting and External Assurance
How results are disclosed to each audience and how external evaluators participate.
Evaluation report and audience tiering
Report instances derived from the evaluation, tiered by audience, with rules preventing public claims from exceeding internal evidence, and versioning for corrections and retractions.
- Which audience tiers exist and what content is withheld or summarised in each? classification
- How is it ensured that public claims never exceed what the full internal evidence supports? validation
- How are report versions, corrections and retractions identified and communicated? lifecycle
- Which report content is mandated by regulation or code, and by what deadline? requirement
External and third-party evaluation
Arrangements for independent evaluators: access tier, scope, safe harbour, publication rights, exemption basis and handling of dissent.
- What level of access — black-box, grey-box, weights, fine-tuning or internal tooling — is granted to external evaluators? access
- Which contractual terms bound scope, timing, safe harbour and publication rights? constraint
- On what basis may an external evaluation be omitted, and how is that basis evidenced? exception
- How is disagreement between internal and external assessments recorded and resolved? process
Records Access and Retention
Who may see evaluation records, how sensitive detail is handled, and how long records live.
Access control and hazardous detail handling
Default access rules per record scope, redaction of uplift-relevant or exploit-enabling detail, handling of personal data in traces, and access auditing.
- What is the default access rule for each record scope, and which roles hold read or write rights? access
- Which uplift-relevant or exploit-enabling details must be redacted, held separately, or shared only with authorities? security
- How are personal data in prompts, traces and rater records minimised, lawfully based and protected? privacy
- Which access events are logged, and for how long are those logs themselves retained? evidence
Retention, deletion and legal hold
Retention periods and legal bases per record class, resolution of deletion-versus-evidence conflicts, legal-hold handling, and the tombstone that keeps prior claims auditable after deletion.
- What minimum and maximum retention period applies to each evaluation record class, and on what legal basis? retention
- How are conflicts resolved between data-minimisation deletion duties and evidence-retention duties? exception
- How is a legal hold applied, recorded and released across evaluation records? process
- What summary or digest must survive after primary records are deleted so that prior claims remain auditable? requirement
Alignment and Exchange Interoperability
Declared alignments to external frameworks and the identifier and comparability conventions enabling exchange.
Standards alignment, conformance limits and exchange interoperability
Which external framework clauses each record claims to address, how alignment is kept distinct from certified conformance, which identifier schemes are used, and when results from different organisations or rounds may be compared.
- Which external framework clauses does each evaluation record claim to address, and what evidence supports the claim? interoperability
- How are alignment statements prevented from being read as certified conformance? constraint
- Which identifier schemes and namespaces are used for subjects, datasets, metrics and reports across organisations? identity
- Under what conditions may results from two organisations or two rounds be compared without misleading readers? measurement
- Where two frameworks impose conflicting requirements, how is the conflict recorded and resolved? decision
Classifiers Filled
- Family
- World Models
- Category
- Information and virtual systems
- Entry kind
- aggregate
- Navigation path
- NAV.INF.AI.EVL
- Domain
- INF.AI.EVL
- Industry
- Cross-industry
- Tags
- aimodelevaluationinf.ai.evl
What it is Filled
One evaluation instance: a bounded, protocol-governed measurement applied to a version-pinned AI model or system under a declared configuration, producing runs, results, findings, decisions and retained evidence. The model owns the evaluation record graph and its governance; it references but does not own the dataset, the model artefact, the organisational management system, the risk register or the incident process.
In scope
- Evaluation mandate, trigger and consuming decision or gate
- Binding of the evaluation to a model or system version, configuration and inference settings
- Method-class selection (statistical, formal, empirical, human-subject, adversarial) and protocol design
- Task and dataset version binding, split discipline, ground-truth definition and contamination control
- Metric specification, aggregation, disaggregation factors and pre-registered acceptance thresholds
- Human-rater and red-team protocols including threat models and elicitation effort
- Run execution records, environment capture, determinism and replication
- Result records, per-item evidence, uncertainty and statistical validity
- Findings, severity, validity threats and declared limitations
- Approval, conditional approval, waiver, rejection and release gating
- Re-evaluation triggers, evaluation lifecycle states and cadence
- Evidence packaging, integrity, provenance and lineage
- Audience-tiered reporting, external and third-party evaluation arrangements
- Access control, sensitivity handling, retention, deletion and legal hold
- Declared alignments to external frameworks and exchange interoperability
Out of scope
- Construction, licensing, annotation and internal structure of evaluation datasets (belongs to the benchmark dataset model)
- Model training, fine-tuning, packaging and build provenance of the evaluated artefact (belongs to the model artefact model)
- Organisation-wide AI management system policy, competence programme and internal audit programme
- Risk identification, treatment selection and residual-risk ownership in a risk register
- Serious-incident capture, statutory reporting deadlines and corrective-action workflow
- Notified-body conformity assessment, declarations of conformity and certification records
- Continuous production telemetry and observability outside a bounded evaluation protocol
- Procurement scoring, vendor selection and commercial cost-benefit analysis
- Human-subject research ethics approval procedures (referenced, not defined)
- Compute infrastructure inventory and capacity management
Why it exists Filled
Provide the format-neutral context an agent needs to plan, execute, record, interpret, approve and disclose an evaluation of a specified AI model or system version, and to retain the resulting evidence.
Distinguishing features Filled
- Records a planned, pre-registered evaluation of a specific model or system version, not a generic benchmark.
- Unlike a benchmark dataset record, it binds a subject, protocol, run and sealed evidence package.
- Unlike a governance assessment, it supplies evidence; it does not decide on residual risk or approval.
- Registers negative and inconclusive results with the same discipline as positive ones.
What robots and AI may and may not do Filled
Must not
- Make a claim about capability, safety or quality that exceeds the sealed evidence.
- Change thresholds or analysis plan after seeing results.
- Act as sole approver of a decision based on runs it executed.
- Publish hazardous-capability detail outside its restricted sensitivity.
- Drop negative or inconclusive results from the record.
Only with a human decision
- Approving a release or deployment decision based on evaluation results.
- Authorising disclosure of hazardous-capability findings.
May
- Register an evaluation protocol with thresholds and disaggregation factors before runs.
- Execute runs and compute results against the bound subject version.
- Check evaluation data for contamination with training data.
- Re-score a result through supersession while keeping the original.
Moral aspects Filled
- Overstated evaluation claims mislead users and regulators about real risk.
- Results not disaggregated by group can hide unfair performance gaps.
- Red-team findings can enable misuse if published carelessly.
Who is affected
- People affected by the evaluated model
- Model developers and deployers
- Regulators and the public relying on claims
Owners Filled
Steward
Name an accountable owner for the evaluation record set and a separate approver role; the approver may not be the sole executor of the runs relied on.
Roles
- Evaluation Owner
- Registers the protocol and its pre-registration time and maintains the mandate basis and scope statement; Ensures the retention schedule, sensitivity scheme and code lists are applied to the instance; Accountable for completeness of the sealed evidence package
- Evaluator or Test Engineer
- Designs and executes runs, capturing environment, configuration, timing and item-level traces; Computes metrics with a recorded scorer implementation and version; Records replication tolerance and any irreproducible components; may not be the sole approver of the consuming decision
- Independent Assessor
- Reviews method-class fit, validity threats, uncertainty treatment and evidence sufficiency; Records dissent where internal and independent readings differ; Confirms that reported claims do not exceed sealed evidence
- Approver or Accountable Deployer
- Takes approve, approve-with-conditions, reject or defer decisions within a recorded delegation; Signs decision records bound to a specific evidence version and authorised scope; Grants, extends and revokes waivers with compensating controls and explicit expiry
- Red Team Lead
- Owns the threat model, attack-class scope, granted affordances and elicitation effort record; Applies restricted sensitivity labels to uplift-relevant detail before distribution; Records unresolved attack avenues and follow-up commitments
- Records and Access Steward
- Applies sensitivity labels, access grants, expiries and revocations across all four scopes; Operates the retention schedule, legal holds, deletion tombstones and access audit logs; Verifies evidence-package digests before submission or publication and raises findings on mismatch
Links to other meta-models Filled
references
- WM-AI-009 — Benchmark dataset - Version-pinned binding of evaluation tasks, splits, ground truth and checksums. This model stores the reference, version, split designation and contamination check only; dataset construction, licensing and internal structure remain there.
- WM-SFT-004 — AI or software model artefact - Identifies the evaluated subject artefact, its version and digest. This model never mints artefact identifiers and defers packaging and build provenance to the artefact model.
- Candidate sibling — AI incident and post-market monitoring model - Supplies re-evaluation triggers and receives findings. Not registered in the relations file, so this link is a proposal marked as a gap rather than a canonical relation.
- Candidate sibling — AI risk register and treatment model - Receives evaluation findings as measurement evidence against identified risks and owns treatment and residual-risk acceptance. Not yet registered; proposed only.
aligned
- NIST AI RMF 1.0 MEASURE function (NIST AI 100-1) - Structuring alignment for documented TEVV test sets, metrics and tools, metric selection, independent assessment, deployment-condition measurement, fairness evaluation and measurement-effectiveness review. Voluntary and non-certifiable.
- Regulation (EU) 2024/1689 Articles 9, 15, 19, 55, 60, 72 and Annex IV - Legal alignment for prior-defined metrics and probabilistic thresholds, documented adversarial testing, dated and signed test logs and reports, minimum log retention, real-world testing and post-market re-evaluation.
- General-Purpose AI Code of Practice — Safety and Security chapter - Voluntary compliance route supplying evaluation triggers, risk tiers with safety margins, acceptance determination before proceeding, external evaluation expectations and the model report before release.
- ISO/IEC 42001:2023 clause 9 and Annex A - Alignment of monitoring, measurement, internal audit and management review duties to the evaluation record set. Asserted at clause-title level only because the normative text is paywalled.
- MLCommons Croissant 1.0 - Interoperable description of referenced evaluation data, including semantic versioning, sha256 checksums and RecordSet or Field structure used for integrity verification of bound datasets.
- MLPerf Inference Rules - Reference profile for comparability: divisions and categories, scenario-specific metrics, quality targets, reporting precision, bounded non-determinism, mandatory replicability and an independent audit process.
- Hugging Face model card evaluation section - Projection target for downstream-deployer disclosure of testing data, factors, metrics, results, limitations and environmental impact.
- NIST AI 100-2 E2025 adversarial machine learning taxonomy - Vocabulary for attacker goals, capabilities and knowledge, lifecycle stages of attack, and attack classes including evasion, poisoning, privacy, indirect prompt injection and misaligned outputs.
- ISO/IEC 25059:2023 quality model for AI systems - Vocabulary alignment for the quality characteristics a metric claims to measure. Treated as provisional because the edition is under revision and the text is paywalled.
composes
- W3C PROV-O - Supplies entity, activity and agent classes with generation, derivation, attribution, association and timing properties for the evaluation lineage graph rather than redefining provenance semantics locally.
neighbor
- WM-AI-009 benchmark dataset - Dataset creation, licensing, annotation workforce and internal record structure are owned there. This model retains only a version-pinned reference, checksum, split designation and contamination check — the minimum needed for reproducibility and comparability.
- WM-SFT-004 AI or software model artefact - Artefact identity, versioning, packaging and build provenance are owned there. This model records a subject binding (identifier, version, digest, configuration) and never mints artefact identifiers.
- AI management system (ISO/IEC 42001 AIMS) - Clause 9 duties to determine what is monitored, run internal audits and hold management reviews are organisational and belong to an AIMS model. This model covers a single evaluation instance and the evidence it yields to that system.
- AI risk register and treatment model - Risk identification, treatment selection and residual-risk acceptance live in a risk model. Evaluation supplies measurement evidence and findings that reference risks; it does not own risk objects.
- Serious incident and post-market monitoring model - Incident capture, reporting deadlines and corrective action belong to an incident model. Evaluation consumes incidents as re-evaluation triggers and emits findings, not incident reports.
- Conformity assessment and certification records - Notified-body assessment and declarations of conformity are separate records. Evaluation evidence is an input; this model forbids presenting alignment as certified conformance.
- Production observability and telemetry - Continuous telemetry is a monitoring concern. An evaluation instance is bounded by a registered protocol with thresholds fixed before results are observed, even when executed on production traffic.
parent
- WM-SFT-004
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- Authoritative master-system identifier issued by the system of record for the entity — the model artefact registry for the evaluation subject, the dataset registry for evaluation data, and the quality or records management system for the evaluation instance, decision and finding.
- Governed global identifier or IRI where no master-system identifier exists — for example a PROV-O or JSON-LD IRI, a DOI, or a registry-governed URI under a declared namespace authority.
- UUID or ULID minted by the adopting Dimension, recorded together with the minting agent and the minting event time.
- A date, a run timestamp, a benchmark name, a file name or a metric score is never an identifier; such values may only be recorded as attributes.
Direct properties not applicable Not applicable
Not applicable
Institutional or informational subject: no invented physical properties.
Recognition optional Filled
- An evaluation names a subject version, a registered protocol, datasets, runs, metrics and a sealed result.
- Often confused with a benchmark dataset, a model card, a governance assessment and production monitoring.
Capabilities and actions required Filled
- Register evaluation protocol: Fix the method, task and dataset bindings, metrics, disaggregation factors and acceptance thresholds as an immutable pre-registered protocol version before any result is observed.
- Bind evaluation subject: Pin the model or system version, configuration and digest that the evaluation will measure, including snapshotting a hosted endpoint.
- Bind evaluation dataset and verify contamination: Pin dataset versions, splits and checksums by reference and record the train/test contamination check.
- Execute evaluation run: Run the registered protocol against the bound subject and dataset, capturing configuration, environment, timing, usage and item-level traces.
- Compute and aggregate results: Apply scorers and reducers to item scores to produce disaggregated and headline values with uncertainty estimates.
- Re-score and supersede a result: Recompute a score or aggregate while preserving the original value, the scorer version and the reason for change.
- Raise evaluation finding: Create a citable finding from results, with evidence links, severity, likelihood and initial state.
- Seal evidence package: Assemble, digest and sign the set of artefacts constituting the evidence for a decision or a submission.
- Decide approval or rejection: Take and sign the authorised decision that consumes the evaluation, bound to a specific sealed evidence version and authorised scope.
- Issue time-bounded waiver: Authorise release despite an unmet acceptance criterion, with compensating controls and an explicit expiry.
- Publish tiered evaluation report: Derive an audience-specific report from sealed evidence, applying redaction rules and a consistency check against the full internal record.
- Grant scoped evaluation access: Issue time-bounded access to evaluation records at a declared scope and tier, with purpose recorded and access logged.
- Trigger re-evaluation: Open a new evaluation instance in response to a qualifying change, incident, drift alarm or scheduled review, linked to the prior instance.
- Apply retention, hold or deletion action: Enforce the retention schedule for a record class, apply or release a legal hold, or delete with a tombstone.
- Scan evaluation integrity: Review transcripts for refusals, evaluation awareness, leakage, harness failure and other threats to validity.
- Package reproducibility evidence: Assemble pins, hashes and unreleased-component list so another party can attempt a re-run or audit.
Hazards and failure modes required Filled
- False assurance from contaminated or unrepresentative test data.
- Selective reporting that hides failures.
- Leakage of dangerous capability detail.
- Results applied to a different model version than tested.
Standards and interfaces required Filled
- ISO/IEC 42001 AI management system.
- ISO/IEC TS 4213 assessment of machine learning classification performance.
- NIST AI Risk Management Framework.
- Model cards and datasheets practice.
Context of use required Filled
- Regulatory duties modelled here are EU-centric under Regulation (EU) 2024/1689; the GPAI obligations in Articles 53 to 55 applied from 2 August 2025 and high-risk system obligations phase in on a later schedule.
- US material (NIST AI 100-1, AI 100-2 E2025) is voluntary guidance; no US federal evaluation mandate is assumed.
- UK AI Security Institute Inspect is used as an engineering reference for run and log structure, not as a regulatory requirement.
- Personal-data handling assumes a GDPR-like regime for lawful basis, minimisation and retention; other regimes may differ materially, particularly on collecting protected attributes for fairness disaggregation.
- No assumption is made about Chinese, Indian, Brazilian, Japanese, Korean or Canadian AI evaluation regimes; adopters in those jurisdictions must add local requirements to the mandate and retention findings.
- Digest algorithm choice (SHA-256 in the cited Croissant profile) is illustrative; jurisdictions or sectors with mandated cryptographic suites must substitute their own.
- EU Articles 10, 15, 43, 55, 72 and 92 apply to providers and deployers in the Union for in-scope systems and GPAI models; they are alignments, not global duties.
- NIST AI RMF and TEVV-Athlon are voluntary US measurement frameworks, not binding law.
- UK AISI Inspect and the Autonomous Systems Evaluation Standard are de facto frontier-evaluation engineering practice, not a statute.
- Croissant adoption is concentrated in Hugging Face, Kaggle and OpenML ecosystems and is not universal.
- Classification-performance controls in ISO/IEC TS 4213 are assumed transferable as hygiene controls (leakage, environment, baselines) even when the metric set does not apply.
Sources Filled
- Artificial Intelligence Risk Management Framework (AI RMF 1.0) — AI RMF Core - National Institute of Standards and Technology (NIST)
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) - European Union — Publications Office (EUR-Lex)
- EU Artificial Intelligence Act — article and annex reference texts (Articles 9, 19, 55; Annex IV) - Future of Life Institute
- The General-Purpose AI Code of Practice - European Commission — DG CNECT / AI Office
- ISO/IEC TS 4213:2022 — Information technology — Artificial intelligence — Assessment of machine learning classification performance - ISO/IEC JTC 1/SC 42 (via IEC Webstore)
- ISO/IEC 42001:2023 — Information technology — Artificial intelligence — Management system - ISO/IEC JTC 1/SC 42
- PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
- Croissant Format Specification 1.0 - MLCommons
- MLPerf Inference Rules - MLCommons
- Annotated Model Card Template - Hugging Face
- Inspect AI — Eval Logs - UK AI Security Institute
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2 E2025) - National Institute of Standards and Technology (NIST)
- ISO/IEC 25059:2023 — Software engineering — Systems and software Quality Requirements and Evaluation (SQuaRE) — Quality model for AI systems - ISO/IEC JTC 1/SC 7 and SC 42
- HELM code architecture documentation - Stanford Center for Research on Foundation Models (CRFM)
- ISO/IEC 24029-2:2023 — Artificial intelligence (AI) — Assessment of the robustness of neural networks — Part 2: Methodology for the use of formal methods - ISO/IEC JTC 1/SC 42
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 - National Institute of Standards and Technology
- The TEVV-Athlon Framework for Evaluating AI Systems, NIST AI 200-2 ipd - National Institute of Standards and Technology
- ISO/IEC TS 4213:2022 Information technology — Artificial intelligence — Assessment of machine learning classification performance - ISO/IEC JTC 1/SC 42
- Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act) - European Union
- Croissant Format Specification 1.1 - MLCommons Association
- Inspect AI: Framework for Large Language Model Evaluations - UK AI Security Institute
- Holistic Evaluation of Language Models (HELM) - Stanford Center for Research on Foundation Models
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 - National Institute of Standards and Technology
- ISO/IEC TR 29119-11:2020 Software and systems engineering — Software testing — Part 11: Guidelines on the testing of AI-based systems - ISO/IEC JTC 1/SC 7 with ISO/IEC JTC 1/SC 42
- Autonomous Systems Evaluation Standard - UK AI Security Institute
Open questions
- Evaluation-instance identity: the base never asks which authoritative master-system identifier designates the evaluation instance itself, only the subject, run, result, finding and report. Draft an instance-scoped identity node rather than importing grok's campaign-scoped wording.
- Campaign or programme grouping: decide whether a body of work spanning several evaluation instances is a sibling aggregate, a parent programme object or a service layer, and where the identifier for it lives. Grok's campaign node is held pending this decision.
- Test-oracle definition where expected results are unavailable or non-deterministic (ISO/IEC TR 29119-11 oracle problem). The base's ground-truth question presumes a reference answer exists; primary text must be retrieved before a node is drafted.
- Continual-learning and post-deployment feedback-loop evaluation (EU AI Act Article 15(4) and Article 72 monitoring plans), including where the boundary sits against the post-market monitoring and incident neighbour that the base excludes.
- Cross-organisation exchange of evaluation results: the base's identifier-scheme and comparability guidance is a design position with no cited normative registry behind it, and grok independently notes MLPerf, Evals and lm-evaluation-harness schemas are projections rather than canon.
- Jurisdictional coverage beyond the EU and voluntary US frameworks: the base marks spatial/jurisdictional scope as a gap and neither provider searched Chinese, Indian, Japanese, Korean, Canadian or Brazilian evaluation regimes, nor sector regimes such as medical device, aviation or SR 11-7 model validation.
- Full normative text of ISO/IEC 42001:2023, ISO/IEC 25059:2023, ISO/IEC TS 4213:2022 and ISO/IEC 24029-2:2023 is paywalled. Alignment is asserted at scope or clause-title level only and must be re-verified against purchased copies before any conformance statement.
- Sector-specific evaluation regimes — medical device software, automotive functional safety, aviation, and financial model risk management such as SR 11-7 — are not modelled and may impose additional validation, independent review and documentation duties.
- Agentic and long-horizon tool-use evaluation (environment sandboxing, task success over many steps, autonomous replication and self-proliferation testing) is only partially covered by the run and trace findings; no cited primary source supplies a settled structure for it.
- Continual-learning and online evaluation, where the subject changes during the evaluation window, is acknowledged through endpoint pinning and lifecycle invalidation but not fully elaborated.
- Environmental accounting appears only as optional run-level data elements; no authoritative measurement methodology for AI evaluation energy or emissions is specified here.
- Human-rater labour conditions, compensation and psychological safeguards for exposure to harmful content are referenced through participant protection but not modelled in detail.
- Cross-organisation evaluation result exchange has no settled normative registry; the identifier-scheme guidance is a design position, not a cited requirement.
- ISO/IEC 24029 formal and statistical robustness methods are neighboring techniques not modelled in full; primary text was not retrieved beyond catalogue scope lines.
- ISO/IEC TS 42119-2:2025 testing-of-AI overview and the in-progress ISO/IEC CD 4213 expansion to regression, clustering and recommendation were not retrieved as full text.
- ISO/IEC 25059 AI quality model, ISO/IEC 42001 AIMS audit evidence and ISO/IEC DIS 24970 AI logging are not encoded.
- MLPerf, Hugging Face Evaluate, OpenAI Evals and EleutherAI lm-evaluation-harness schemas are widely used projections not treated as canonical.
- Model Cards (Mitchell et al.) and data cards are reporting formats; only finding and results artifacts are modelled here.
- Medical-device, aviation and SR 11-7-style financial model-validation regimes are out of scope and unmapped.
- China, sectoral and other non-EU, non-US evaluation statutes were not searched in this pass.
- Inter-laboratory reproducibility protocols and energy-measurement metrology beyond ISO/IEC TS 4213 mentions remain thin.
- Human-subjects statistical-power design for ARIA-style contextual evaluations is only pointed at via TEVV-Athlon field-testing tools.
Machine files
Provenance
world-models research · reviewable-draft
Built from: models/wm-ai-003-ai-model-evaluation/spec.yaml, ver-cy/world-models/card-supplements/wm-ai-003-ai-model-evaluation.json