EM-AI-03 · Subject model · W3
AI evaluation and safety evidence
Evaluation plan, benchmark, run, results, risk assessment and decision on permissible use. Metric, dataset and evidence are reused.
Queued for research
Claude: not-started; Grok: not-started.
Research note, in Russian: Entire research brief pending
Subject boundary and candidate types
- EvaluationPlan
- EvaluationRun
- BenchmarkSpecification
- SafetyAssessment
- DeploymentDecision
Deep research questions
- How to detect benchmark leakage into training?
- How to separate capability measurement from permissibility of use?
- How to account for uncertainty and a context change?
Verifiable invariants
- A result pins the model and dataset versions
- An evaluation is not universal outside its context
- A decision has an accountable owner and conditions
End-to-end acceptance scenario
A change of dataset and application context leads to a separate evaluation and an explicit decision, not to carrying over the previous approval.
Negative case
One high benchmark score declares the system safe for all tasks.
Approaches to compare
- NIST AI RMF: context, evaluation and risk management
- PROV/SPDX: provenance of data, artifacts and actions
- Model registry/evaluation tooling practice; map system, artifact, run and endpoint
Candidates in the live catalogue
- WM-AI-003 · AI Model Evaluation · 0.3.0-research.1 · installable
Semantic fit requires boundary research; a published model does not by itself complete this card. - WM-AI-009 · Evaluation Dataset / Benchmark · 0.3.0-research.1 · installable
Semantic fit requires boundary research; a published model does not by itself complete this card. - WM-AI-008 · AI Safety / Governance Assessment · 0.3.0-research.1 · installable
Semantic fit requires boundary research; a published model does not by itself complete this card. - WM-AI-010 · AI Incident Report · 0.3.0-research.1 · installable
Semantic fit requires boundary research; a published model does not by itself complete this card.
Result requirements
Every card is executed together with the full research contract: definitions, fields and cardinalities, lifecycle, sources, data mastership, rights, the five object facets, at least eight invariants, positive and negative examples, dependencies, migration and applicability limits.