Assessment / Evaluation
Give an agent the context needed to understand, create, inspect and operate an assessment or evaluation record: what was assessed, against which criteria at which version, by which method, on what evidence, yielding which criterion outcomes, scores, composite result and stated conclusion, with declared validity, assurance, lifecycle, disclosure and interoperability.
Bundle → Layer → Finding → Questions Filled
6 bundles · 14 layers · 28 findings · 106 questions
Assessment Frame, Subject and Scope What this assessment is, who mandated it, what it is about, and where its boundary lies.
Identity and Mandate
How the assessment record is identified and versioned, and the authority and purpose under which it exists.
Assessment record identity, granularity and versioning
The identifier that authoritatively designates this assessment, the unit of record it covers, and how successive versions and supersessions are distinguished from independent re-assessments.
- Which identifier authoritatively designates this assessment record, and which system issued it? identity
- What is the atomic unit of the record: one subject, one criteria set, one occasion, or a combination? definition
- How are successive versions of the same assessment distinguished from a new, independent assessment? lifecycle
- Which prior assessment, plan or request does this record derive from or fulfil? provenance
Assessment type, stakes and mandate
The class of assessment, the party relationship between assessor and subject, the scheme or legal basis authorising it, and the consequence class its result carries.
- Which class of assessment is this: conformity, diagnostic, formative, summative, risk, quality, appraisal or peer review? classification
- Is this a first-party, second-party or third-party assessment relative to the subject? relationship
- Under what mandate, scheme or legal basis is the assessment carried out? authority
- What consequences attach to the result, and who commissioned it? ownership
Subject and Scope
What was assessed, in what state, and how the assessed boundary and sample were drawn.
Subject identification and assessed state
Which object of assessment was examined, at which pinned version, configuration or state, whether it is an individual, batch, population or system, and where it was located when that matters.
- Which identifier and identifier scheme designate the object of assessment? identity
- What version, configuration or state of the subject was assessed, and how is that state pinned? state
- Is the subject an individual item, a batch, a population or a system, and how is membership defined? composition
- Where was the subject located or deployed at the time of assessment, when location affects the result? spatial
Scope boundary, sampling and generalisation limits
What was explicitly inside and outside the assessed scope, how a sample was selected and sized if the whole subject was not assessed, and how far the result may legitimately be generalised.
- What is explicitly inside and outside the assessed scope? definition
- If a sample was assessed rather than the whole subject, how was the sample selected and sized? process
- To what extent may the result be generalised beyond the assessed sample or scope? constraint
- What happens to the assessment if the scope changes while it is in progress? exception
Criteria, Method and Scale Which criteria apply at which version, how they were examined, and the scale and decision rule by which results are expressed and decided.
Criteria Binding
Reference and version binding to the criteria source, and which criteria were selected, excluded or added.
Criterion reference and version binding
The catalogue, standard or rubric supplying the criteria, the version pinned for this assessment, the locator for each individual criterion, and any interpretation applied to ambiguous criterion text.
- Which criteria catalogue, standard or rubric supplies the criteria, at which version? relationship
- How is an individual criterion addressed unambiguously within that source? identity
- Where the criterion text is ambiguous, which interpretation or application note was applied? definition
- How is the exact criteria version preserved so the result stays reproducible after the catalogue changes? temporal
Criterion selection, applicability and tailoring
Which criteria from the bound source were selected, which were declared not applicable and why, and whether any criteria were added beyond the source and under whose authority.
- Which criteria from the source were selected for this assessment, and on what basis? decision
- Which criteria were declared not applicable, and what justifies each exclusion? exception
- Were any criteria added beyond the referenced source, and under whose authority? authority
Method and Execution
How each criterion was examined and what must be preserved for the result to be reproducible.
Assessment method, instrument, mode and coverage
Which methods were used to examine each criterion, whether results were produced automatically, manually or semi-automatically, the depth and coverage applied, and the instrument, rubric or item set that operationalises the criterion.
- Which assessment methods were used, such as examine, interview, test, inspect or elicited judgement? process
- Was each result produced automatically, manually or semi-automatically, and by which agent? classification
- What depth and coverage were applied, and what do they imply about assurance? measurement
- Which instrument, rubric, item set or procedure defines how each criterion is examined? composition
Execution parameters and reproducibility
The tool, model or engine references, versions, configuration profiles and input snapshots that were in force when machine-generated results were produced, captured so an independent party can reproduce them.
- Which tool, model or engine produced each machine-generated result, at which version? provenance
- What configuration, profile, thresholds and input data snapshot were in force during execution? constraint
- What must be preserved for an independent party to reproduce the same results? quality
Scale and Decision Rule
The vocabulary and scale in which results are expressed, and the rule that converts them into a decision.
Outcome scale and result vocabulary
The measurement level, bounds and favourable direction of any numeric scale, the controlled vocabulary for non-numeric outcomes, and how indeterminate, inapplicable, untested and error states are kept distinct from a negative result.
- Is the outcome nominal, ordinal, interval or ratio, and what are its bounds? measurement
- Which controlled vocabulary expresses non-numeric outcomes for this assessment? classification
- Which direction of the scale is favourable, and what are the best and worst attainable values? definition
- How are untested, not-checked, error and not-selected states distinguished from a negative result? exception
Decision rule, thresholds and uncertainty handling
The rule that converts scored or determined values into a conformity or grade decision, the cut scores, acceptance limits and guard bands it uses, how uncertainty is accounted for beforehand, and the conditions under which the assessment must abstain.
- What decision rule converts measured or scored values into a conformity or grade decision? decision
- What cut scores, acceptance limits or guard bands apply, and who set them? constraint
- How is measurement or rater uncertainty accounted for before the decision is taken? measurement
- Under what conditions must the assessment abstain rather than decide? exception
Evidence and Observation What was relied on, how it links to criteria, whether it is sound, and what was actually observed.
Evidence Corpus
The evidence items relied on, their linkage to criteria, and their quality and admissibility.
Evidence items and criterion linkage
The discrete evidence items relied on, their types, the criterion or determination statement each supports or contradicts, and the times at which each was created, collected and received.
- What discrete evidence items were relied on, and of what type is each? composition
- Which criterion or determination statement does each evidence item support or contradict? relationship
- What makes the evidence set sufficient to support the determination for each criterion? evidence
- When was each evidence item created, collected and received? temporal
Evidence quality, admissibility and reuse limits
Whether each item is authentic and unaltered since collection, whether it is still current for the assessed state, which offered evidence was rejected or down-weighted, and which items carry confidentiality or personal-data constraints limiting reuse.
- How is each evidence item shown to be authentic and unaltered since collection? security
- Is the evidence still current for the assessed state, and when does it cease to be? validation
- Which offered evidence was rejected or down-weighted, and why? quality
- Which evidence carries confidentiality, personal-data or privilege constraints that limit reuse? privacy
Observation and Measurement
Raw observed results with distinct event and recording times, and the uncertainty and traceability that qualify quantitative values.
Observation records, actors and timing
The raw result observed or returned for each check before judgement is applied, the actor or device that produced it, the separation of the time the observed condition held from the time the observation was recorded, and the place of observation where it matters.
- What raw result was observed or returned for each check before judgement was applied? measurement
- Which actor or device produced the observation, and under which role? provenance
- How are the time the observed condition held and the time the observation was recorded kept distinct? temporal
- Where was the observation taken when site or environment materially affects the result? spatial
Measurement uncertainty and traceability
The stated uncertainty of each quantitative result and how it was evaluated, the quantity, unit and reference scale each value belongs to, and the calibration or reference standard the measurement is traceable to.
- What is the stated uncertainty of each quantitative result and how was it evaluated? measurement
- Which quantity, unit and reference scale does each numeric value belong to? definition
- To what reference standard or calibration is the measurement traceable? provenance
Determination, Scoring and Conclusion Criterion-level verdicts and scores, their aggregation into a composite result, and the stated conclusion with its validity.
Criterion-Level Outcomes
Per-criterion verdicts with rationale and locators, and the raw and normalised scores that accompany them.
Criterion verdict, rationale and locator
The outcome value assigned to each criterion from the bound vocabulary, the reasoning connecting cited evidence to that outcome, the severity or materiality of negative outcomes, and where in the subject each negative outcome occurs.
- What outcome value was assigned to each criterion, from which vocabulary? state
- What reasoning connects the cited evidence to each criterion outcome? evidence
- How are negative outcomes graded by severity or materiality? classification
- Where in the subject does each negative outcome occur? spatial
Criterion scores and normalisation
The raw score recorded for each criterion against its maximum, how raw scores are normalised or rescaled before aggregation, and which scale definition version each recorded score refers to.
- What raw score, if any, was recorded for each criterion and against which maximum? measurement
- How are raw scores normalised or rescaled before aggregation? process
- Which scale definition version does each recorded score refer to? interoperability
Composite Result, Conclusion and Validity
How criterion results roll up, what conclusion is stated and by whom, and for how long and under what conditions it holds.
Aggregation model and composite result
The scoring or aggregation model that combines criterion results, the weights and precedence rules it applies, the form in which the composite is expressed, and how it is computed when some criteria are untested or indeterminate.
- Which scoring or aggregation model combines criterion results into a composite result? process
- What weights or precedence rules apply, and how are unweighted and inapplicable items handled? constraint
- How is the composite result expressed, and must its inputs be published alongside it? measurement
- How is a composite computed when some criteria are untested or indeterminate? exception
Conclusion statement, authority and qualification
The stated conclusion and its form, the actor accountable for it and whether the decision was separated from the determination work, the qualifications, adverse opinions or dissent attached, and what the conclusion entitles or obliges.
- What is the stated conclusion, and is it binary, graded or narrative? definition
- Who is accountable for the conclusion, and was the decision separated from the determination work? authority
- What qualifications, adverse opinions, dissent or minority positions attach to the conclusion? quality
- What does the conclusion entitle or oblige, and which downstream act, if any, must follow? relationship
Validity window, conditions and confidence
From when until when the conclusion is asserted to hold, the conditions whose breach voids it early, the surveillance or re-assessment needed to keep it current, and the confidence or assurance level claimed.
- From when until when is the conclusion asserted to hold? temporal
- Which conditions, if breached, void the conclusion before its end date? constraint
- What ongoing surveillance or re-assessment is required to keep the conclusion current? requirement
- What confidence or assurance level is claimed, and on what basis? quality
Assurance of the Assessment Itself Whether the people, tools and instrument behind the result can be relied on.
Assessor and Independence
Who performed the assessment, in what roles, with what competence and authorisation, and under what declared conflicts.
Assessor identity, roles and contribution
Which persons, teams, bodies or automated agents performed the assessment, how the roles of collecting, determining, reviewing and deciding were distributed, and which parts of the result are attributable to which agent.
- Which persons, teams, bodies or automated agents performed the assessment? identity
- How are the roles of collecting, determining, reviewing and deciding distributed? ownership
- Which parts of the result are attributable to which assessor or agent? provenance
Competence, authorisation and impartiality
The competence, qualification or accreditation supporting each assessor's participation, the designated scope within which the assessing body may act, declared relationships that could impair impartiality, and what the assessment explicitly does not entitle the assessor to assert.
- What competence, qualification or accreditation supports each assessor's participation? requirement
- Within what designated or accredited scope is the assessing body entitled to act here? authority
- What relationships, interests or prior work could impair impartiality, and how were they treated? constraint
- What does this assessment explicitly not entitle the assessor to assert? exception
Quality Control and Instrument Quality
Review and moderation of this assessment, and the standing evidence that the instrument measures what the conclusion claims, reliably and fairly.
Review, moderation and rater agreement
The independent review, second marking or moderation performed before finalisation, the inter-rater or inter-tool agreement measured against a threshold, and how disagreements between assessors or tools were resolved.
- What independent review, second marking or moderation was performed before finalisation? process
- What inter-rater or inter-tool agreement was measured, and what threshold applied? measurement
- How were disagreements between assessors or tools resolved? decision
Instrument validity, reliability and fairness evidence
The standing evidence that the instrument measures what the conclusion claims, that its results are precise and repeatable, that bias and accessibility barriers across subject groups have been checked, and the populations or contexts for which it is not validated.
- What evidence supports that the instrument measures what the conclusion claims? validation
- What evidence supports the precision and repeatability of the results? quality
- What checks were made for bias, differential functioning or accessibility barriers across subject groups? validation
- For which populations, contexts or subject types is the instrument not validated? constraint
Lifecycle, Disclosure and Interoperability How the record moves through its states, how it is corrected and challenged, who may see what, how long it is kept, and how it projects into external formats.
Record Lifecycle
States and transitions of the assessment record, and how it is corrected, challenged and superseded.
Assessment state and transitions
The states an assessment record may occupy and the vocabulary defining them, the permitted transitions with their actors and preconditions, what becomes immutable at finalisation, and how a record that should never have existed is marked without erasing its history.
- What states may an assessment record occupy, and which vocabulary defines them? state
- Which transitions are permitted, who may trigger them, and what preconditions apply? lifecycle
- What becomes immutable at finalisation, and what may still change afterwards? constraint
- How is a record that should never have existed marked, without deleting its history? exception
Correction, appeal and supersession
When an amendment is required rather than a new assessment, how the subject or a relying party may challenge the conclusion and within what period, how the outcome of a challenge is reflected, and how the chain from the original to the currently effective conclusion is navigated.
- What circumstances require an amendment rather than a wholly new assessment? decision
- How may the subject or a relying party challenge the conclusion, and within what period? process
- How is the outcome of an appeal or complaint reflected in the record? event
- How is the chain from the original to the currently effective conclusion navigated? relationship
Disclosure and Retention
Who may see which part of the record, what is redacted before release, and how long each part is kept before disposition is referred onward.
Disclosure classification and redaction
The disclosure class applying separately to the conclusion, the scores and the underlying evidence, the terms under which the assessment may be published or cited, what must be redacted before external release and who approves it, and whether a conclusion may be released without its evidence.
- What disclosure class applies to the conclusion, the scores and the underlying evidence separately? access
- Under what terms, embargo or licence may the assessment be published or cited? ownership
- What must be redacted before external release, and who approves the redaction? privacy
- Can a relying party receive the conclusion without the evidence, and what integrity guarantee accompanies it? security
Retention class and disposition referral
The retention class and trigger applying to each part of the record, which policy and system execute destruction or transfer when retention ends, and how a legal hold or open appeal suspends disposition.
- What retention class and minimum retention period apply to each part of the record? retention
- What event starts the retention clock for the assessment and for its evidence? temporal
- Which policy and which system execute destruction or transfer when retention ends? ownership
- How does a legal hold or open appeal suspend disposition? exception
Interoperability
Projection into external assessment reporting formats and binding of the vocabularies, units and languages used.
External alignment and vocabulary binding
Which external assessment reporting formats this record must project into, how local outcome values map into each target vocabulary and where that mapping loses information, which code systems, units and language tags are bound at which versions, and what is claimed as conformance to an external standard.
- Which external assessment reporting formats must this record project into? interoperability
- How do local outcome values map to each target vocabulary, and where is the mapping lossy? interoperability
- Which code systems, units and language tags are bound, and at which versions? classification
- What is claimed as conformance to an external standard, and what evidence supports that claim? validation
Classifiers Filled
- Family
- World Models
- Category
- Activities and processes
- Entry kind
- aggregate
- Navigation path
- NAV.ACT.ASM
- Domain
- ACT.ASM
- Industry
- Cross-industry
- Tags
- assessmentevaluationact.asm
What it is Filled
Format-neutral structure of a single assessment/evaluation instance and the bindings it carries to reusable definitional components (criteria source, instrument, scale, decision rule). It spans framing and scoping, criteria binding and tailoring, method and execution parameters, evidence and observation capture, criterion-level determination, aggregation, conclusion and validity, assurance of assessor competence and impartiality, and record lifecycle, disclosure, retention referral and external projection. It is subject-domain agnostic and applies equally to conformity assessment, security control assessment, accessibility evaluation, educational measurement, clinical risk assessment and peer review. It is modelled as an aggregate because the assessment act, its inputs and its determinations share one identity boundary and one finalisation event.
In scope
- Identity, granularity and versioning of an assessment record and its supersession chain
- Assessment type, stakes, mandate, scheme reference and first/second/third-party relationship
- Subject reference, pinned assessed state, scope boundary, sampling and generalisation limits
- Binding of criteria to a named catalogue at a pinned version, plus selection, tailoring and interpretation notes
- Declared method, mode, depth and coverage, and the instrument or rubric that operationalises each criterion
- Execution parameters needed for reproducibility: tool and model references, versions, configuration, input snapshots
- Evidence item references, criterion linkage, sufficiency, admissibility and confidentiality qualifiers
- Observation records with distinct event and recording times, plus declared uncertainty and traceability pointers
- Criterion-level outcomes, rationale, severity and locators; raw, normalised and maximum scores
- Aggregation model, weighting, composite score or level, and result-vector expression
- Conclusion statement, decision authority, qualifications, dissent, validity window and surveillance obligations
- Assessor identity and roles, competence and authorisation references, impartiality and conflict declarations
- Independent review, moderation, rater agreement, and validity/reliability/fairness evidence for the instrument
- Record state machine, correction, appeal and supersession; disclosure classification, redaction and retention referral
- Crosswalks and vocabulary bindings for projection into external assessment reporting formats
Out of scope
- Internal structure, lifecycle and state management of the assessed subject itself
- Authoring, versioning, publication and normative interpretation of the criteria catalogue or standard
- Runtime execution and orchestration of testing tools, scanners, graders or evaluator engines and their operational telemetry
- Issuance, signing, suspension, withdrawal and verification of certificates, credentials or attestations
- Enforcement of consequences: certification decisions with legal effect, licensing, sanction, market withdrawal or remediation workflow
- Identity management, employment and qualification records of persons and organisations
- Physical or digital custody, storage and chain-of-custody execution for evidence objects
- Calibration programmes, metrological traceability management and instrument maintenance
- Training, validation and deployment lifecycle of predictive or generative scoring engines
- Platform audit-trail and access-log semantics, storage and evaluation
- Execution of retention, destruction, transfer or erasure actions
- Risk register lifecycle, risk treatment and plan-of-action-and-milestones case management
- Population-level aggregation of many assessments into reputation scores or league tables
- Commercial terms, scheduling and cost accounting of assessment services
Why it exists Filled
Give an agent the context needed to understand, create, inspect and operate an assessment or evaluation record: what was assessed, against which criteria at which version, by which method, on what evidence, yielding which criterion outcomes, scores, composite result and stated conclusion, with declared validity, assurance, lifecycle, disclosure and interoperability.
Distinguishing features Filled
- Records a determination against version-pinned criteria with a declared method, scale and decision rule.
- Separates determination from decision: the assessor and the actor who records the conclusion can differ.
- Unlike a measurement record, it interprets observations against criteria rather than storing raw values.
- Unlike an attestation or certificate, it does not issue a credential with legal effect.
What robots and AI may and may not do Filled
Must not
- Record a conclusion without a version-pinned criterion and supporting evidence.
- Publish a score without its scale, aggregation model and decision rule.
- Treat an untested or indeterminate outcome as a pass.
- Infer regulatory or contractual effect from the quality of the method.
- Delete an assessment to settle a dispute or hide an unfavourable result.
Only with a human decision
- Recording a conclusion that has contractual, regulatory or employment consequences.
- Changing the decision rule or weights after evidence has been seen.
- Resolving a dispute about an assessment outcome.
May
- Frame an assessment and bind criteria at an explicit version.
- Register evidence and link each item to the criteria it supports.
- Compute a composite result with the declared aggregation model and publish it with its scale reference.
- Record indeterminate, not-applicable and untested outcomes explicitly.
Moral aspects Filled
- Assessments of people affect careers and access to opportunities; criteria and evidence must be open to the person assessed.
- Aggregate scores can hide failures on single critical criteria.
- Bias in criteria or method can disadvantage groups systematically.
Who is affected
- Persons or organisations being assessed
- Decision makers relying on the result
- Assessors whose judgement is recorded
Owners Filled
Steward
Name a single accountable assessment owner and, separately, a decision authority entitled to record conclusions; the two roles must be distinguishable in every record so determination and decision remain separable.
Roles
- Assessment Owner
- Accountable for the record's scope, completeness and timeliness; Approves scope and sampling changes and records their effect on results already captured; Names the assessors, the reviewer and the decision authority for the record
- Assessor / Evaluator
- Performs selection and determination against the bound criteria; Registers evidence references, linkages, sufficiency judgements and observations with dual timestamps; Records criterion outcomes, rationale and scores, including indeterminate and not-applicable values; Where the assessor is an automated agent, is identified as such with its tool reference, version and configuration
- Technical Reviewer / Moderator
- Reviews determinations independently of the assessor who made them; Records agreement statistics and adjudicates disagreements between assessors or tools; May return a record to draft with reasons, and may not decide a conclusion on determinations they reviewed
- Decision Authority
- Applies the declared decision rule to the composite result and its uncertainty; Records the conclusion, its type, qualifications, dissent and validity window; Accountable for the referral to any downstream attestation and for stating what the conclusion does not entitle
- Criteria and Vocabulary Steward
- Maintains bindings to external criteria catalogues, scales, outcome vocabularies and decision rules; Publishes crosswalks, successor mappings and deprecations for target reporting formats; Assesses the impact of catalogue and vocabulary version changes on records that are still in draft
- Disclosure and Records Officer
- Sets disclosure class and publication terms per record part and approves redactions; Issues and withdraws public renditions and records the redaction rule version applied; Assigns the retention class and trigger and refers disposition to the owning records-retention policy model
Links to other meta-models Filled
child
- WM-ACT-009 (parent model in NAV.ACT) - Inherit generic activity semantics — actor, occurrence timing, status and provenance — and specialise them for criterion-referenced judgement producing a determination and a conclusion. Generic activity identity and timing machinery is not restated here.
references
- Criteria / requirement catalogue model (control catalogue, standard clause set, rubric registry) - Carry the criteria source reference, its pinned version, criterion locators, local interpretation notes and tailoring decisions. Criterion text, structure, normative status, versioning and publication remain with the catalogue.
- Subject-domain model of the assessed object - Carry the subject reference and a pinned state descriptor of the object as assessed. The subject's own attributes, lifecycle and state transitions remain with the domain model.
- Party, organisation and role model - Resolve assessor, requester, relying party and decision-maker references and their roles on this record. Party identity, contact data and organisational structure remain with the party model.
- Competence, accreditation and designation registry - Carry accreditation, notification or designation references and the scope they cover for this assessment. Granting, surveilling, suspending and withdrawing accreditation belongs to the accreditation authority.
- Evidence custody and records repository model - Carry evidence pointers, integrity digests, collection times and sensitivity classes. Storage, chain-of-custody execution and retrieval of the evidence bytes remain with the repository.
- Measurement, observation and metrology model - Carry measured values, units, stated uncertainty and calibration references used as evidence. Uncertainty evaluation methodology, calibration programmes and traceability chains remain with the metrology model.
- Attestation, certificate and credential model - Carry a referral marker and a pointer where the conclusion is attested downstream. Issuance, signing, suspension, withdrawal and verification of certificates or credentials belong entirely to that model.
- Risk register and treatment model - Carry pointers where a determination seeds a risk entry. Risk characterisation lifecycle, mitigating factors, remediation tracking and deadlines belong to the risk model.
- Records retention and disposition policy model - Carry the retention class, retention trigger, legal-hold flag and the disposition referral record. Setting periods and executing destruction, transfer or erasure belong to that policy model.
- Access control and information classification model - Carry the disclosure class and publication terms per record part. Evaluating access requests, enforcing decisions and storing access logs belong to that model.
- Assessment scheme, programme or examination framework model - Carry the scheme reference and the parameters it fixed for this occasion, such as required criteria profile, surveillance interval and eligibility. Scheme rules, cycles and governance belong to the programme model.
- Dispute, appeal and complaint case model - Carry the appeal reference, its outcome and the hold it places on disposition. Conducting proceedings, managing timelines and recording hearings belong to the case model.
composes
- Provenance model aligned to W3C PROV - Attach Entity, Activity and Agent provenance to assessment content: attribution of determinations to agents, derivation of conclusions from evidence, and start and end times of assessment activities.
aligned
- Code list, vocabulary and unit registry - Bind outcome vocabularies, scale definitions, severity schemes, unit code systems and language tags at stated versions, and publish crosswalks between them. Vocabulary authoring and deprecation belong to the registry.
neighbor
- Measurement / observation model - An assessment consumes measured or observed values as evidence and adds criterion-referenced judgement; a bare measured value with an uncertainty statement is a measurement result, not an assessment. Units, uncertainty evaluation and calibration traceability are referenced from the metrology model.
- Criteria / requirement catalogue model - This model carries a criterion reference, its pinned catalogue version, local applicability and interpretation notes only. Criterion text, structure, normative status and catalogue lifecycle belong to the catalogue, exactly as OSCAL separates a control catalogue from assessment results and XCCDF separates a Benchmark from a TestResult.
- Attestation / certificate / credential model - ISO/IEC 17000 separates determination and review from decision and from attestation. This model ends at the recorded conclusion; issuing, signing, suspending or revoking a certificate or credential is a downstream act, and in EU product law only a designated notified body may give that act legal effect.
- Evaluator or test-execution engine - EARL records the assertor and the test mode as attributes of an assertion, and XCCDF records the checking system as a reference. Naming a tool, its version and its configuration never transfers ownership of running that tool, its scheduling, or its operational logs.
- Risk register and treatment model - OSCAL keeps observations, findings and risks distinct and hands risks onward to a POA&M. This model may produce findings that seed risks, but risk characterisation lifecycle, mitigating factors, remediation tracking and deadlines belong to the risk model.
- Party, competence and accreditation registry - Assessor identity, competence evidence, accreditation and designation scope are referenced by identifier and version. Whether a body is accredited or notified, and for what scope, is asserted and maintained by the accreditation authority, not by an assessment record.
- Provenance and audit-trail models - Content provenance (which agent produced which determination, from which evidence, when) is mixed in from a PROV-aligned model and retained. Platform-level record access and change logging is a separate audit model whose storage and evaluation this model does not own.
- Records retention and disposition model - This model declares a retention class, a retention trigger event and a legal-hold flag, and records the referral. Setting periods and executing destruction, transfer or erasure is owned by the adopting Dimension's records policy.
- Aggregate rating / reputation model - schema.org AggregateRating rolls many individual ratings into counts and averages across a population. That population-level statistic is a sibling concern; this model bounds one assessment occasion on one declared subject and scope.
- Assessment programme / scheme model - A conformity assessment scheme, examination programme or continuous-monitoring programme defines recurring rules, cycles and eligibility. This model references the scheme and carries the parameters it fixed for this occasion; the scheme's own rules and cycle belong to the programme model.
parent
- WM-ACT-009
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- Authoritative master-system identifier issued by the system of record for the assessment — for example the assessing body's report or case number, the accreditation scheme's application number, or the assessment platform's record identifier.
- Governed global identifier or IRI where the assessment, its criteria binding or its published conclusion exists under a governed namespace — for example a DOI, a URN or a scheme-issued IRI.
- UUID or ULID assigned by the adopting Dimension where neither of the above exists; assigned once at record creation and never reissued or reused.
- A date, a score, a composite level, a subject name, a criteria version or a file name is never an identifier and may not be used as one, alone or in combination.
Direct properties not applicable Not applicable
Not applicable
Institutional or informational subject: no invented physical properties.
Recognition optional Filled
- An assessment names its subject, the criteria version, the method, the scale and a recorded conclusion.
- Often confused with a raw measurement, an audit engagement, a certificate and a risk assessment.
Capabilities and actions required Filled
- Frame the assessment: Open an assessment record by fixing its identity, type, mandate, subject reference, pinned subject state, scope boundary and sampling design.
- Bind criteria at version: Resolve the criteria source, pin its version, select the applicable criteria, record not-applicable justifications and any additions, and retain the binding so results stay reproducible.
- Declare method, scale and decision rule: Record the methods, modes, depth and coverage per criterion, the instrument used, the outcome scale and vocabulary, and the decision rule with its thresholds, guard bands and abstention conditions, before any result is produced.
- Register evidence and link it to criteria: Register a reference to an evidence object with its type, integrity digest, collection and receipt times, the criteria it bears on, its sufficiency judgement and any confidentiality constraints.
- Record an observation: Record a raw observed result with its method, producing actor or device, distinct event and recording times, place where relevant, and any uncertainty and unit for quantitative values.
- Record a criterion outcome: Assign an outcome value from the bound vocabulary to a criterion, with rationale, cited evidence, severity where negative, locator, and any raw and normalised score.
- Compute the composite result: Apply the aggregation model already declared on this record to the criterion outcomes and scores recorded on it, producing a composite score, level or vector together with the inputs needed to reproduce it.
- Derive and record the conclusion: Apply the declared decision rule to the composite result and its uncertainty, record the resulting conclusion with its type, qualifications, dissent and validity window, and name the accountable decision maker.
- Review and finalise: Record independent review or moderation, resolve disagreements, and transition the record to final, freezing the content that must not change afterwards.
- Amend or supersede: Issue a new version of a finalised assessment with a reason code, link it to the version it replaces, and move the effective-record flag, or mark a record as entered in error without deleting it.
- Register an appeal or complaint outcome: Record a lodged challenge against the conclusion, the rights exercised, its outcome and the resulting amendment or reaffirmation, and suspend disposition while it remains open.
- Classify disclosure and project to external formats: Set the disclosure class per record part, produce a redacted public rendition under a versioned redaction rule, and emit crosswalked projections into required external reporting formats with explicit loss statements.
- Declare retention and refer disposition: Assign a retention class and trigger to each part of the record, register any legal hold, and refer disposition to the owning records-retention policy model when the period expires.
Hazards and failure modes required Filled
- Wrong decisions from scores read without their scale or decision rule.
- Gaps hidden when untested criteria are counted as satisfied.
- Unfair outcomes from biased criteria or unvalidated methods.
Standards and interfaces required Filled
- ISO/IEC 17000 conformity assessment vocabulary and principles.
- ISO/IEC 33001 series for process assessment.
- W3C Verifiable Credentials for carrying results as attestations.
- Open Badges for learning and skill assessments.
Context of use required Filled
- EU conformity-assessment terminology — notified body, CE marking, declaration of conformity, modules — is specific to the EU New Legislative Framework; other jurisdictions use different designation regimes and the party-relationship and authority findings must be re-bound accordingly.
- The ISO/IEC 17000-series functional model (selection, determination, review, decision, attestation) and the ILAC accreditation ecosystem are assumed as the default vocabulary; jurisdictions outside that ecosystem may separate these functions differently.
- The AERA/APA/NCME Standards describe primarily United States practice; fairness, accommodation and test-taker rights expectations differ materially in other jurisdictions and the fairness finding is deliberately mechanism-only rather than threshold-setting.
- Retention periods, disclosure duties, appeal windows and erasure rights are jurisdiction-specific; all are referenced as classes and pointers, and none is fixed by this model.
- Language and locale of criteria, rationale and conclusions are not assumed to be English; BCP 47 tagging is required on every narrative field.
- Measurement traceability assumes an SI-based metrological infrastructure; assessments using non-metrological or purely judgemental scales bind a scale definition instead and leave the uncertainty elements empty.
Sources Filled
- ISO/IEC 17000:2020 Conformity assessment — Vocabulary and general principles - ISO/IEC (Joint Technical Committee)
- OSCAL Assessment Results Model (Assessment Layer) - National Institute of Standards and Technology (NIST)
- NIST SP 800-53A Rev. 5 — Assessing Security and Privacy Controls in Information Systems and Organizations - National Institute of Standards and Technology (NIST)
- NISTIR 7275 Rev. 4 — Specification for the Extensible Configuration Checklist Description Format (XCCDF) Version 1.2 - National Institute of Standards and Technology (NIST)
- Evaluation and Report Language (EARL) 1.0 Schema - World Wide Web Consortium (W3C)
- PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
- FHIR R5 RiskAssessment resource - Health Level Seven International (HL7)
- Question and Test Interoperability (QTI) 3.0 — Implementation and Best Practices - 1EdTech Consortium (formerly IMS Global)
- Common Vulnerability Scoring System v4.0 Specification Document - Forum of Incident Response and Security Teams (FIRST)
- Common Criteria for Information Technology Security Evaluation, CC:2022 Release 1, and CEM:2022 - Common Criteria Recognition Arrangement (CCRA)
- ILAC-G8:09/2019 Guidelines on Decision Rules and Statements of Conformity (ILAC Guidance Series) - International Laboratory Accreditation Cooperation (ILAC)
- Conformity assessment (Single Market — Goods, building blocks) - European Commission
- Website Accessibility Conformance Evaluation Methodology (WCAG-EM) 2.0 - World Wide Web Consortium (W3C)
- schema.org Rating, Review and AggregateRating types - Schema.org Community Group (W3C)
- JCGM publications, including JCGM 106:2012 The role of measurement uncertainty in conformity assessment and JCGM 100:2008 (GUM) - Joint Committee for Guides in Metrology (BIPM)
- Standards for Educational and Psychological Testing (open-access edition) - AERA, APA and NCME
Open questions
- Ratify the composition links and the parent link to WM-ACT-009 once the relationship contract is populated, obtaining WM-ACT-009's own rationale so the parent claim can be compared rather than inherited from a registry field.
- Resolve the supersession-chain question: decide whether an assessment series is a distinct identity that owns the effective-record flag, or whether the flag must be derivable from version metadata on each record.
- Establish the source basis and a boundary note for the access-control model that the checklist refers enforcement to, then re-derive the access dimension from actual findings and questions rather than prose.
- Source jurisdiction-specific retention periods, appeal windows, disclosure duties and erasure rights from records-management and legal instruments, none of which were consulted in this research round.
- Re-source the EU legal-effect claims to primary instruments — Decision 768/2008/EC, Regulation (EU) 2019/1020 and the Blue Guide (2022) — instead of the Commission portal summary currently marked primary_source.
- Obtain independent second-provider or qualified human expert review of the aggregate boundary, artifact-ownership rules and outcome-vocabulary crosswalks when the Grok waiver is lifted or an alternative reviewer is authorised.
- Specify the artifact-provenance semantics of the serial flag, then re-test whether af-instrument-definition, af-technical-manual, af-assessment-procedure and af-crosswalk-table are owned artifacts or referenced instance-invariant assets.
- Domain-specific criteria semantics — clinical, educational, financial, environmental, safety — are referenced but not modelled; a Dimension adopting this model in a regulated domain will need a sibling criteria model.
- Sampling statistics beyond declaration: no power analysis, confidence-interval construction or acceptance-sampling plan is modelled.
- Adaptive and computerised-adaptive testing item-selection algorithms, item exposure control and equating are outside the instrument finding.
- Benchmark harness specifics for evaluating machine-learning systems — dataset splits, contamination and leakage controls, held-out set governance — are only partly reached by the reproducibility finding.
- Panel deliberation and consensus protocols (Delphi, expert elicitation, jury scoring) are summarised as roles and agreement statistics rather than modelled as procedures.
- Cross-border mutual recognition of assessment results, and the conditions under which one jurisdiction accepts another's conclusion, are not modelled.
- Effort, cost and scheduling of assessment engagements are excluded.
- Continuous and streaming assessment is only partly covered: surveillance obligations are declared, but the monitoring cadence and its state machine belong to the programme model.
- Aggregation of many assessments into population statistics, league tables or reputation scores is excluded and left to a sibling model.
- Digital signature profiles and long-term validation of signed assessment reports are referenced through integrity digests but not specified.
Machine files
Provenance
world-models research · reviewable-draft
Built from: models/wm-act-034-assessment-evaluation/spec.yaml, ver-cy/world-models/card-supplements/wm-act-034-assessment-evaluation.json