EM-AI-03 · Subject model · W3

AI evaluation and safety evidence

Evaluation plan, benchmark, run, results, risk assessment and decision on permissible use. Metric, dataset and evidence are reused.

Queued for research

Claude: not-started; Grok: not-started.

Research note, in Russian: Entire research brief pending

Subject boundary and candidate types

Deep research questions

Verifiable invariants

End-to-end acceptance scenario

A change of dataset and application context leads to a separate evaluation and an explicit decision, not to carrying over the previous approval.

Negative case

One high benchmark score declares the system safe for all tasks.

Approaches to compare

Candidates in the live catalogue

Result requirements

Every card is executed together with the full research contract: definitions, fields and cardinalities, lifecycle, sources, data mastership, rights, the five object facets, at least eight invariants, positive and negative examples, dependencies, migration and applicability limits.

Machine-readable assignment JSON