← Back to catalogue
Published

Quality / Confidence

vr.wm-xct-026 · wm-xct-026-quality-confidence

Provide a format-neutral, attachable structure for asserting, computing, evidencing and governing quality dimensions, measures, scores, assessment methods, uncertainty and confidence about any subject, so that an agent can decide whether the subject is fit for a stated purpose and can defend that decision with evidence.

World Models Cross-cutting context XCT.QLT

Bundle → Layer → Finding → Questions Filled

6 bundles · 14 layers · 27 findings · 112 questions

Quality assertion foundations What a quality/confidence record IS: a separately identified assertion about some other subject, bound to a declared dimension and a registered, dereferenceable measure.

Assertion identity and attachment

Identity of the assertion itself and the precise unit of the subject it covers.

Identity of the quality assertion as a distinct resource

The quality/confidence record is a resource in its own right, distinct from the subject, from the measure and from the assessing activity. It must be independently identifiable so that it can be cited, superseded, disputed, embargoed and retained on its own schedule. DQV expresses this as a typed resource (QualityMeasurement, QualityAnnotation, QualityCertificate, UserQualityFeedback, QualityPolicy) that is linked to, rather than embedded in, the assessed resource; PROV then attaches generation and attribution to that resource.

  1. Which identifier identifies this quality assertion itself, as distinct from the subject it describes and the measure it applies? identity
  2. Is this record a computed measurement, a human quality annotation, a certificate, a policy statement, or user feedback? classification
  3. What natural key determines whether two records are the same assertion rather than two independent assessments of the same subject? identity
  4. Who is authorised to mint and retire identifiers for quality assertions within the adopting Dimension? authority

Subject attachment and assessment scope

An assertion is meaningless unless the assessed unit is explicit: the whole resource, one distribution, a subset, a feature type, a single instance, one attribute, or a population represented by a sample. ISO 19157-1 makes scope an explicit component of a data quality evaluation, and DQV's computedOn names the exact resource assessed. Ambiguous scope is the most common way an honest score becomes a misleading one.

  1. What exactly was assessed: the whole resource, a distribution, a subset, a field, or a single instance? composition
  2. Does the assertion apply only within a restricted extent, population segment, or operating condition, and how is that restriction expressed? constraint
  3. Was the whole scope inspected or was a sample used to infer a property of a population? measurement
  4. May one assertion cover several subjects at once, and if so how is per-subject traceability preserved? relationship

Dimension and measure vocabulary

The graded property being assessed (dimension/category) and the concrete, registered, dereferenceable procedure that produces a value (measure/metric).

Dimension and category vocabulary binding

DQV separates Dimension (a quality-related characteristic that matters to consumers) from Metric (an abstract procedure for assessing it) and groups dimensions into Categories, but deliberately declines to prescribe one authoritative list. Multiple valid vocabularies exist in parallel — ISO/IEC 25012's fifteen characteristics under inherent and system-dependent viewpoints, ISO/IEC 5259-2 for analytics and ML data, ISO 19157-1's geographic quality components, and the European Statistics Code of Practice output principles. The mixin therefore binds a vocabulary per subject class rather than embedding one, and records that binding as a governed profile.

  1. Which dimension vocabulary is bound for this subject class, at which version, and by whose authority? interoperability
  2. Is the dimension inherent to the data itself or dependent on the surrounding system and context of use? classification
  3. Which dimensions are mandatory for this subject class before it may be published or relied upon, and which are optional? requirement
  4. Where two bound vocabularies define overlapping or contradictory dimensions, which one governs? exception

Measure definition and registration

A score is uninterpretable unless the measure that produced it is itself a registered, versioned, dereferenceable object with a definition, parameters, a value type and a value structure. ISO 19157-1 specifies the components and content structure of data quality measures; ISO/TC 211 publishes these as normative machine-readable schemas; ISO/FDIS 19157-3 specifies a register of measures maintained under ISO 19135-1 registration procedures. That register standard is still at FDIS stage, so registration governance is adopted here as a defensible pattern rather than a settled normative requirement.

  1. Which registered measure produced this value, at which version, and where is its definition dereferenceable? identity
  2. What parameters, tolerances or thresholds were bound when the measure was applied, and what are their values? constraint
  3. What value type and value structure does the measure emit — a single value, a ratio, a count, a matrix, or a coverage? measurement
  4. How is a measure superseded or retired, and what happens to results already produced by the retired version? lifecycle

Domain-specific dimension extension

Additional dimensions (for example ML-specific characteristics in ISO/IEC 5259-2, linked-data interlinking, or community uniqueness) are first-class local concepts with provenance of who added them. The full list of extra ISO/IEC 5259-2 characteristics is not fully extractable from freely published text and remains a documented gap.

  1. What identifier and definition distinguish this extension dimension from standard catalogue codes? identity
  2. Who authorised the extension and against which standard extension clause? authority
  3. If an ISO/IEC 5259-2 additional ML characteristic is claimed, is the official characteristic name cited from the standard text rather than inferred? evidence
Assessment method and result How the assessment was carried out and how its numeric or categorical outcome is expressed, including uncertainty.

Assessment method

Procedure, environment, sampling design and the reference against which the subject was compared.

Evaluation procedure, sampling and environment

The executed procedure must be recorded with enough fidelity to be repeated: full inspection versus sampling, the sampling design, the data splits used, the executing agent (human, automated or hybrid), the tool and its version, and the evaluation environment. ISO/IEC TS 4213 makes evaluation setup, preprocessing, training/test/validation splits, cross-validation, documented algorithm and hyperparameters, environment and baselines explicit control criteria; ISO 19157-1 makes the evaluation method a reported component.

  1. Which documented procedure and which tool version executed this assessment? process
  2. If sampling was used, what was the sampling design, the achieved sample size, and the resulting representativeness claim? measurement
  3. Was the assessment automated, human-judged, or hybrid, and where did human judgement enter? provenance
  4. What preprocessing, exclusion or filtering was applied before measurement, and could it bias the result? quality

Reference basis, ground truth and traceability

Every quality claim is a comparison against something. That reference must be named: a specification, a reference dataset, an adjudicated ground truth, a calibrated standard, or an expert panel. VIM defines accuracy, trueness and precision strictly relative to a true or reference value, and defines metrological traceability as the property that relates a result to a reference through a documented unbroken chain of calibrations. Where no defensible reference exists, the assertion must say so rather than imply one.

  1. Against which reference, specification or ground truth was the subject compared, and who adjudicated it? evidence
  2. Is there an unbroken documented chain relating this result to a stated reference, and where is it recorded? provenance
  3. What is known about the quality and error rate of the reference itself? quality
  4. If no independent reference exists, how is that stated so the result is not read as accuracy? exception

Result expression

The value produced, its scale and structure, and the uncertainty attached to it.

Result value, scale and value structure

The result carries a value plus everything needed to interpret it: expected data type, scale type (nominal, ordinal, interval, ratio), value structure (single value, ratio, count, matrix, coverage) and, where the value is a quantity, a dereferenceable unit reference. DQV recommends dereferenceable unit identifiers rather than unit strings. Ordinal grades must be marked as ordinal so that consumers do not average them.

  1. What is the result value and what data type and value structure does it use? measurement
  2. What scale type applies, and which arithmetic operations are therefore invalid on this value? constraint
  3. If the value is a quantity, which dereferenceable unit identifier applies? interoperability
  4. What is the permitted value domain, and how are null, not-applicable and not-assessed distinguished from a poor score? validation

Uncertainty and error model

A value without an uncertainty statement cannot support a defensible conformity decision. JCGM GUM defines the evaluation of measurement uncertainty and its supplements cover distribution propagation and multiple output quantities; VIM defines standard, combined and expanded uncertainty, coverage interval, coverage probability and coverage factor. This model requires that any quantitative result declare either an uncertainty statement or an explicit, reasoned statement that none was evaluated. Uncertainty about the value is distinct from confidence in the assertion, and the two must not be merged.

  1. What uncertainty is attached to the result, expressed as a standard uncertainty, an expanded uncertainty, or an interval? measurement
  2. How was the uncertainty evaluated, and which components dominate the budget? process
  3. If no uncertainty was evaluated, why not, and what interpretation is thereby forbidden? exception
  4. How is the uncertainty of the value kept distinct from the confidence in the assertion as a whole? definition
Confidence and evidence The explicit, scale-declared statement of how much weight the assertion itself can bear, and the evidence that backs it.

Confidence expression

What the confidence statement is about, which scale it uses, and what reasons produced it.

Object and semantics of a confidence statement

Confidence in this model is a statement about the assertion — how much the assessment can be relied upon — not a property of the subject and not a probability that a proposition about the world is true. GRADE frames certainty as certainty that the true value lies on one side of a threshold or within a range, which is explicitly about the evidence base rather than about the estimate alone. Conflating confidence with a model's output probability, with coverage probability, or with a risk score is the principal failure mode this finding exists to prevent.

  1. What is this confidence statement about: the assertion, the underlying value, the method, or the subject? definition
  2. Is the confidence value a probability, and if not, what is it? classification
  3. What action is an agent permitted or forbidden to take at each confidence level? decision
  4. Where the subject is a model output that carries its own score, is that score treated as confidence, and is it calibrated? quality

Confidence scale declaration and calibration

A confidence value is uninterpretable without its scale. The scale must be named, versioned and dereferenceable, with its ordered terms and any numeric mapping stated. GRADE's four-level certainty scale is one such governed scale; a model's softmax output is not. Where a numeric confidence is claimed to be calibrated, the calibration evidence — reference set, calibration method, date and observed calibration error — must be recorded, otherwise the value is reported as uncalibrated.

  1. Which confidence scale is used, at which version, and where are its ordered terms defined? interoperability
  2. Is there an authorised numeric mapping for the scale terms, and is cross-scale comparison permitted? constraint
  3. What evidence supports a claim that numeric confidence values are calibrated, and when was it last established? evidence
  4. How is an uncalibrated confidence value marked so downstream agents do not treat it as a probability? quality

Confidence basis: evidence and agreement

Confidence must be decomposed into named reasons rather than asserted. GRADE names domains that lower certainty (risk of bias, imprecision, inconsistency, indirectness, publication bias) and domains that raise it (large effect, dose-response gradient, plausible confounding that would only strengthen the finding). This model generalises that pattern: record the type, amount and consistency of evidence, the degree of agreement among independent assessments, and each rating-down or rating-up decision with its reason, so the grade is reconstructible and challengeable.

  1. What type and amount of evidence underpins this assertion, and how consistent is it? evidence
  2. How much do independent assessors, methods or sources agree, and how was agreement measured? quality
  3. Which factors lowered or raised the confidence rating, and what is the recorded reason for each adjustment? decision
  4. What evidence that would materially change the rating is known to be missing or inaccessible? exception

Metaquality

A quality report can itself be wrong, biased or unrepresentative. Metaquality elements record confidence in the quality evaluation, whether the evaluated sample represents the quality unit, and whether quality is homogeneous across the unit. This is quality-of-quality, not a second copy of the host score.

  1. What metaquality confidence is assigned to this quality result, and on what evidence? quality
  2. How representative of the declared quality unit is the evaluated sample or census? evidence
  3. Is quality homogeneous across the unit, or do partitions differ enough that a single score is misleading? quality
  4. Who assessed metaquality, and is that party independent of the original assessor? ownership

Evidence and reproducibility

The retrievable material that substantiates the assertion and allows an independent party to re-run it.

Evidence artifacts, integrity and re-execution

An assertion is only as defensible as the material a challenger can retrieve. W3C VC Data Model 2.0 provides an evidence property, cryptographic securing mechanisms and credential status for tamper-evidence and revocation; the EU AI Act requires technical documentation and retained logs that substantiate declared performance. This finding requires that each assertion either point to retrievable, integrity-protected evidence with re-execution instructions, or declare explicitly that it is an unsubstantiated judgement.

  1. Which retrievable evidence items substantiate this assertion, and what does each one show? evidence
  2. How is the integrity of each evidence item protected and verified at retrieval time? security
  3. What is needed to re-execute this assessment and reproduce the result within stated tolerance? process
  4. If no evidence is retained, is the assertion explicitly marked as an unsubstantiated judgement? quality
  5. Is any evidence item access-restricted, and what may a consumer be told when they cannot retrieve it? access
Requirement, conformance and fitness Turning measured quality into a decision: thresholds, verdicts and a declared judgement of fitness for a stated purpose.

Requirements and conformance

Stated acceptance criteria and the verdict derived from measurement against them.

Quality requirement and acceptance threshold

A threshold is a governed object separate from both the measure and the result: it names the measure, the acceptance limit, the direction of the comparison, its owner and its effective period. Because measurement uncertainty affects conformity decisions, the decision rule must state how uncertainty is handled — for example acceptance limits with guard bands — rather than comparing a point estimate to a limit and calling it conformity.

  1. What acceptance limit applies to this measure, in which direction, and who set it? requirement
  2. How does the decision rule account for measurement uncertainty when comparing the result to the limit? decision
  3. On what basis was the threshold chosen — regulation, contract, product specification, or internal convention? authority
  4. What happens to existing verdicts when a threshold changes? lifecycle

Conformance verdict and severity

A verdict is a derived assertion that cites both a result and a requirement. Constraint validation supplies a genuinely binary conformance signal with organisational severities (violation, warning, informational); graded measurement supplies a score. The model keeps them distinct: a verdict must name the requirement it was evaluated against, and a score may never be presented as a verdict. Severity classifies consequence, not degree of failure.

  1. What is the verdict, against which requirement, and using which result? validation
  2. What severity is attached to each non-conformance, and what does that severity oblige the consumer to do? constraint
  3. Is the verdict reproducible from the recorded result, requirement and decision rule alone? provenance
  4. How are partial, conditional or waived conformance states represented without collapsing them into pass? exception

Fitness for purpose and limitations

The purpose-relative judgement and the explicit statement of what the assertion does not support.

Intended use, fitness for purpose and declared limitations

Quality is only decidable relative to a purpose: the European Statistics Code of Practice defines relevance by user need, ISO 19157-1 frames quality information as helping users decide whether data suffice for their particular application, and the EU AI Act requires providers to declare accuracy levels and metrics in the instructions for use. A fitness statement must therefore name the purpose it was judged against and, symmetrically, state the known defects, exclusions and uses the assertion does not support — otherwise absence of a caveat is read as absence of a limitation.

  1. For which stated purpose was fitness judged, and by whom? decision
  2. Is the subject judged fit, fit with conditions, or unfit for that purpose, and on which assertions does the judgement rest? validation
  3. What known defects, exclusions or systematic errors are present that a consumer must be told about? quality
  4. Which uses does this assessment explicitly not support, and where is that recorded so it travels with the data? constraint
  5. What happens to the fitness judgement when the intended purpose changes? lifecycle
Time, lifecycle and provenance When the assertion refers to, when it was produced, how long it remains valid, how it is superseded, and who stands behind it.

Time and validity

Distinct timestamps and the period over which an assertion may be relied on.

Assessment time versus subject state time

At least three times must be distinguishable: the time of the subject state that was assessed, the time the assessment produced its result, and the time the assertion was recorded or ingested. SOSA's separation of phenomenonTime from resultTime is the canonical pattern. All times use RFC 3339 with seconds and an explicit offset or Z; a date alone is never sufficient and is never an identifier.

  1. Which state of the subject, at which instant or interval, does this assertion describe? temporal
  2. When was the assessment executed and when did its result become available? temporal
  3. When was the assertion recorded or ingested into the adopting Dimension, and how does that differ from the result time? provenance
  4. What time precision and clock source are used, and is the precision adequate for the decisions taken on this assertion? quality

Validity period, currency and re-assessment

Assertions have a shelf life. Credential models express this with validFrom/validUntil; official statistics treat timeliness and punctuality as governed dimensions. This model requires an explicit validity policy per measure: a validity window or a decay statement, a re-assessment cadence, and defined consumer behaviour when an assertion is stale — never silent reuse of an expired score.

  1. From when until when may this assertion be relied upon? temporal
  2. What must a consuming agent do when the assertion is past its validity window? state
  3. How often must this measure be re-assessed, and what events force early re-assessment? process
  4. Does quality on this dimension decay predictably over time, and is that decay modelled or merely assumed? measurement

Lifecycle and supersession

States an assertion moves through and the rules that keep the historical record honest.

Assertion lifecycle, supersession and immutability

Once issued, a quality assertion is a historical fact about what was believed at a time and must not be silently edited. Register practice models item status (valid, superseded, retired) with explicit supersession pointers; credential practice adds revocation through status checking. This model requires a declared state set, an append-only supersession chain, and a distinction between correcting a defective assertion (retract plus reissue with reason) and producing a new assessment (new assertion, prior one remains valid for its period).

  1. What are the permitted lifecycle states of an assertion and which transitions are legal? lifecycle
  2. Is this record a correction of a defective assertion or an independent new assessment? state
  3. What guarantees that an issued assertion is not altered after the fact, and how is that verifiable? security
  4. When an assertion is retracted, what must happen to decisions and derived assertions that relied on it? event

Attribution and provenance

Who made the assertion, on whose authority, with what independence, and from which inputs.

Assessor attribution, competence, authority and independence

A quality assertion inherits credibility from its author. PROV supplies attribution and qualified attribution with roles; DQV distinguishes producer measurements, third-party certificates and user feedback precisely because their standing differs. This model additionally records competence or accreditation where claimed, the authority under which the assertion is issued, and independence or conflict of interest relative to the subject's owner — a self-assessment and an independent audit are not interchangeable and must not be presented as equivalent.

  1. Which agent produced this assertion, in which role, and on whose behalf? ownership
  2. Under what authority, accreditation or delegated mandate is this assertion issued? authority
  3. What is the assessor's relationship to the subject's owner, and is any conflict of interest declared? provenance
  4. What evidence supports the assessor's competence for this specific measure? evidence

Provenance lineage and derived assertions

Assertions are frequently derived: rolled up across parts, weighted into composite indices, propagated to downstream products, or recomputed from other assertions. PROV's wasDerivedFrom, used and wasGeneratedBy make that lineage explicit, and DQV's quality-metadata container lets certificates, policies, measurements and annotations carry joint provenance as a group. Aggregation is a lineage operation with its own risks: composite scores hide dimension-level failures and are invalid across incompatible scales, so the aggregation function, weights and their justification must be recorded with the derived assertion.

  1. From which assertions, datasets or activities was this assertion derived? provenance
  2. If this is a composite or rolled-up score, what function and weights produced it and who approved them? composition
  3. Are the inputs on compatible scales and scopes such that aggregation is meaningful? constraint
  4. Can a consumer recover the dimension-level results behind a composite score? interoperability
  5. How does an assertion about an input propagate to products derived from that input? relationship
Interoperability, disagreement and governance Exchanging quality statements without false equivalence, handling competing and disputed assessments, and governing access and retention.

Interoperability and comparability

Alignment to external vocabularies, and the rules that stop incomparable scores from being compared.

External alignment, mappings and recorded conflicts

External standards are alignments, not conformance claims. This model maps its slots to DQV classes and properties, to ISO/IEC 25012 and ISO/IEC 5259-2 dimension vocabularies, to ISO 19157-1 measure components, to JCGM uncertainty terms and to PROV attribution, recording for each mapping whether it is exact, broader, narrower or merely related. Conformance to any of these is asserted only where tested evidence exists; recorded conflicts include the fact that constraint-validation severities are not quality grades and that credential formats carry no confidence property.

  1. To which external term does each slot of this model map, and at what mapping strength? interoperability
  2. Is conformance to a cited standard claimed, and what test evidence supports the claim? evidence
  3. Which known conflicts or semantic mismatches exist between aligned standards, and how are they resolved? exception
  4. How are mappings maintained when an aligned standard is revised? lifecycle

Comparability and non-comparability rules

The most damaging misuse of quality data is comparing values that are not comparable. Coherence and comparability are treated as a governed output-quality principle in official statistics precisely because comparability must be established, not assumed. Two results are comparable only if the measure and its version, the parameter binding, the scope and population, the scale and unit, the reference basis and the validity period are compatible. Confidence grades on different scales are never comparable and never averageable.

  1. Under what stated conditions may two results of this measure be compared or trended? constraint
  2. Which combinations of scales, units or confidence vocabularies are explicitly forbidden from being combined? validation
  3. How is a break in comparability caused by a method or threshold change signalled to consumers? event
  4. When results are normalised for comparison, what normalisation was applied and is it reversible? process

Disagreement, feedback and rectification

Multiple assessors, contested assessments and the enforceable right to have inaccurate assertions corrected.

Competing assessments, disagreement and rectification

Several parties may legitimately assess the same subject and disagree — DQV explicitly accommodates producer measurements, third-party certificates and user feedback side by side, since fitness for purpose varies by consumer. Disagreement must be represented, not silently resolved by last-writer-wins. Where assertions concern personal data, disagreement acquires legal force: the accuracy principle obliges rectification or erasure of inaccurate data, and data subjects hold rights to rectification and erasure, so a dispute channel with defined outcomes and deadlines is mandatory rather than optional.

  1. Which other assertions cover the same subject, scope and dimension, and do they conflict? relationship
  2. What precedence rule applies when assertions conflict, and who owns that rule? decision
  3. How does an affected party challenge an assertion, and what outcomes and deadlines apply? exception
  4. Where the assertion concerns personal data, what rectification or erasure duty applies and how is it evidenced? privacy
  5. While a dispute is open, how is the assertion's disputed status made visible to consumers? state

Governance: access and retention

Who may see quality statements and their evidence, and how long they are kept.

Access sensitivity and disclosure of quality statements

A quality statement can be more sensitive than the subject it describes: a low accuracy score, an open dispute or an unresolved defect can be commercially or legally damaging, and evidence packages may contain personal or confidential data even when the subject does not. Access must therefore be classified independently at bundle, layer, finding and artifact level, with a defined redacted view for consumers who are entitled to know that a limitation exists but not to see the underlying evidence, and with a bias toward disclosing limitations that affect safety.

  1. What access classification applies to this assertion, and does it differ from that of the subject? access
  2. What may a consumer without full access be told, and what must never be inferable from the redacted view? security
  3. Which limitations must be disclosed regardless of classification because they affect safety, rights or legal duties? requirement
  4. Is the assertion under embargo before a scheduled release, and who may lift it? authority

Retention, deletion and legal hold of assertions and evidence

Retention of quality records is contested between opposing duties. Regulated AI providers must retain technical documentation, logs and performance evidence that substantiate declared accuracy; data protection law imposes storage limitation and grants erasure rights over personal data, including within evidence packages. This finding requires an explicit retention schedule per assertion class and per evidence artifact, a deletion mechanism that preserves the integrity of the historical record where the assertion itself must survive, and a legal-hold override that suspends scheduled deletion.

  1. How long must this assertion and each of its evidence artifacts be retained, and on what basis? retention
  2. When evidence must be deleted but the assertion must survive, what is retained and what is destroyed? process
  3. How is an erasure request affecting quality evidence executed and evidenced without falsifying the audit trail? privacy
  4. What suspends scheduled deletion, who may impose and release a hold, and how is the hold recorded? authority
  5. How is a conflict between a mandatory retention duty and an erasure duty resolved and documented? exception

Classifiers Filled

Family
World Models
Category
Cross-cutting context
Entry kind
mixin
Navigation path
NAV.XCT.QLT
Domain
XCT.QLT
Industry
Cross-industry
Tags
qualityconfidencexct.qlt

What it is Filled

WM-XCT-026 models the quality/confidence ASSERTION as a first-class, separately identified, attributable and time-stamped record about some other subject. It covers: what is assessed and at what granularity; which dimension and which registered measure are used; the method, sampling and reference basis; the result value, its scale and its uncertainty; a confidence statement about the assertion itself, its declared scale and its evidential basis; conformance against a stated requirement; declared fitness for purpose and limitations; validity in time; lifecycle and supersession; attribution and provenance; comparability, disagreement, access and retention. It does NOT model the subject's own semantics, nor units, provenance graphs, identifier minting, constraint languages, risk of harm or agent trust, all of which belong to composable sibling models. Storage and interface (JSON, YAML, Markdown, RDF, Git, MCP, MongoDB) are projections of this structure, never its semantics. W3C DQV is used as the primary structural anchor precisely because it deliberately refuses to prescribe a single definition of quality and instead standardises the shape of comparable assessments.

In scope

  • Quality assertions attached to any host entity, event, dataset, model output, artifact or process step
  • Dimension/category vocabularies and registered quality measures with parameters, value type and value structure
  • Assessment method: procedure, evaluation environment, full inspection versus sampling, reference/ground-truth basis
  • Result expression: value, scale type, value structure, unit reference, aggregation and composite scores
  • Measurement uncertainty and error models, including coverage interval, coverage probability and coverage factor
  • Confidence as an explicit statement about the assertion, with a declared scale, evidential basis and calibration claim
  • Requirements, acceptance thresholds, decision rules, conformance verdicts and severity
  • Declared intended use, fitness-for-purpose judgement and explicit limitations, caveats and known defects
  • Temporal validity, currency, re-assessment cadence, lifecycle states, supersession and retraction
  • Attribution, competence, independence of the assessor, and provenance/lineage of derived assertions
  • Comparability rules, competing assessments, disagreement and rectification
  • Access classification, disclosure, retention and deletion of quality statements and their evidence

Out of scope

  • The subject's own domain semantics, payload or business rules
  • Definition of units, quantity kinds and unit conversion (Measurement & Units sibling model)
  • General provenance graph semantics and activity modelling (Provenance & Attribution sibling model)
  • Identifier minting policy and identifier scheme registration (Identifier sibling model)
  • Constraint/shape languages and the authoring of validation rules (Validation & Constraint sibling model)
  • Risk, hazard and likelihood of harm; confidence in an assertion is not a risk score (Risk sibling model)
  • Trust, reputation and accreditation of agents as standalone subjects (Agent/Party and Certification sibling models)
  • Statistical and machine-learning method definitions themselves (Method/Analysis sibling model)
  • Licensing, rights and terms of use of the subject
  • Orchestration of monitoring pipelines and job scheduling (Process sibling model)
  • Sensitive-data classification schemes as such; this model only carries the classification assigned to a quality statement

Why it exists Filled

Provide a format-neutral, attachable structure for asserting, computing, evidencing and governing quality dimensions, measures, scores, assessment methods, uncertainty and confidence about any subject, so that an agent can decide whether the subject is fit for a stated purpose and can defend that decision with evidence.

Distinguishing features Filled

  • Treats a quality or confidence statement as its own attributable, time-stamped assertion about another subject.
  • Binds every result to a registered measure, method and scale, so results from different schemes are not mixed.
  • Differs from validation, which checks constraints, and from risk, which concerns harm.
  • Separates confidence in the assertion from the quality result it states.

What robots and AI may and may not do Filled

Must not

  • Aggregate scores from different scales into one number.
  • Issue a quality assertion without method and sampling basis.
  • Edit an issued assertion instead of superseding it.
  • Declare a subject fit for purpose without a declared purpose.
  • Present self-assessment as independent assessment.

Only with a human decision

  • Deciding fitness for a safety-critical or legal purpose.
  • Retracting a published assertion.

May

  • Execute an assessment under a registered measure and record the result.
  • Attach uncertainty and a scale-declared confidence with reasons.
  • Reconcile competing assertions and report disagreement.
  • Schedule re-assessment when assertions expire.

Moral aspects Filled

  • Quality labels on data about people can lead to unfair decisions if the data is poor for some groups.
  • Overstated confidence misleads users who cannot check the basis.
  • Assessors must be identifiable so their judgements can be challenged.

Who is affected

  • Users who rely on quality claims
  • Producers of the assessed subject
  • People described by the assessed data

Owners Filled

Steward

Designate a named owner for the model and a separate register manager for the quality measure register and the confidence scale definitions, with documented delegation and a successor rule.

Roles

Model owner
Own the scope, boundary notes and out-of-scope list, and resolve boundary disputes with sibling models.; Approve major versions and breaking changes, and publish migration notes.; Maintain the registry entry, review state and the record of unresolved boundaries.
Register manager
Operate the quality measure register and the confidence scale register: accept submissions, assign identifiers, set status and supersession.; Ensure every registered item is dereferenceable, versioned and accompanied by a definition, parameters, value type and value structure.; Publish register change records and retain superseded items so historical assertions remain interpretable.
Assessor
Execute assessments according to the planned scope, sampling design and documented procedure, and record the execution environment.; Record results with scale, unit reference and status codes, and evaluate uncertainty or record an explicit reasoned omission.; Assign confidence against a named scale with decomposed rating-down and rating-up reasons, and declare independence and any conflict of interest.
Quality steward
Maintain quality profiles, thresholds, decision rules, validity windows and re-assessment cadence for subject classes in scope.; Monitor staleness, trigger re-assessment on subject, method or threshold change, and manage series-break markers.; Reconcile competing assertions under the governed precedence rule and keep disagreement visible.
Evidence and records custodian
Retain evidence packages with digests, enforce the retention and disposition schedule, and apply and release legal holds.; Execute erasure requests without falsifying the audit trail, producing tombstones and deletion records.; Document and escalate conflicts between retention duties and erasure duties.
Access authority
Assign and review access classifications for assertions and evidence independently of the subject's classification.; Define and maintain redaction profiles and mandatory-disclosure lists for safety- and rights-affecting limitations.; Authorise embargo release and audit access to restricted quality statements.
Dispute and rectification handler
Receive and triage challenges to assertions, set and enforce response deadlines, and record outcomes.; Execute rectification, amendment, annotation or retraction and notify recipients of corrected assertions.; Maintain the disputed-state visibility of assertions while a challenge is open.

Links to other meta-models Filled

composes

  • Any host world-model entry that requires quality or confidence qualification - Attach quality assertions to any subject through a quality-metadata container and a subject reference, without modifying or duplicating the host's own semantics. DQV's pattern of linking quality statements to an assessed resource is the template.
  • Validation & Constraint results model (sibling) - Consume binary conformance results with focus node, source shape, constraint component and severity as evidence inputs to graded quality assertions, keeping validation authoring out of this model.
  • Evidence & Attestation model (sibling) - Carry issued quality certificates and attestations with issuer, securing mechanism, status and evidence, providing tamper-evidence and revocation that this model does not define.

references

  • Provenance & Attribution mixin (sibling; identifier to be assigned by the registry) - Resolve who generated a quality assertion, from what, and when. DQV itself delegates this to PROV rather than restating it, so this model references generation, attribution and derivation instead of redefining them.
  • Measurement, Units & Quantity Kinds model (sibling) - Resolve dereferenceable unit identifiers and quantity kinds for quantitative results; DQV recommends a dereferenceable unit reference rather than a unit string, and VIM supplies the metrological vocabulary.
  • Agent / Party & Role model (sibling) - Resolve assessor identity, role, accreditation and organisational affiliation used for attribution, competence and independence classification.
  • Sampling & Statistical Population model (sibling) - Resolve frames, sampling designs, achieved samples and representativeness claims that determine whether a measured value generalises to the declared scope.
  • Access & Disclosure Policy model (sibling) - Resolve classification schemes, redaction profiles and authorisation decisions applied to quality statements and evidence packages, which this model classifies but does not define.
  • Records Retention & Disposition model (sibling) - Resolve retention classes, schedules, legal holds and disposition actions covering assertions and evidence under conflicting retention and erasure duties.

child

  • Quality measure register (nested registry within this model's governance scope) - Hold registered measure items with identifier, version, status and supersession, maintained under registration procedures; nested rather than sibling because measure semantics are inseparable from result semantics.

aligned

  • Risk & Harm model (sibling) - Align confidence and quality inputs to risk assessment without merging them: measurement and evidence assessment is a distinct governed activity from risk treatment, and confidence grades must not be arithmetically combined with risk scores.
  • ISO/IEC 25012 and ISO/IEC 5259-2 data quality dimension vocabularies - Bind dimension terms from published vocabularies rather than embedding a proprietary list; alignment strength is recorded per term and conformance is not claimed without test evidence.
  • ISO 19157-1 data quality measure component structure - Align the measure definition slots (identifier, name, definition, parameters, value type, value structure, source reference) and the scope/evaluation/reporting separation to a published component structure.
  • Sensing System Capability model (sibling, SSN/SOSA-aligned) - Reference declared instrument capabilities qualified by operating conditions as method context, while keeping capability declarations distinct from assessments of produced results.

extends

  • Dataset / Catalogue record model (sibling) - Extend catalogue records with attached quality metadata containers so that quality carries its own provenance, validity and access rules instead of becoming inert descriptive fields.

neighbor

  • Validation & Conformance (SHACL-style constraint checking) - SHACL yields binary conformance: a focus node either conforms or produces results, and sh:resultSeverity (Violation/Warning/Info) is organisational rather than a graded score. WM-XCT-026 carries graded, scaled and uncertain assessments. A validation report is therefore an INPUT artifact to a quality assertion, not a substitute for it; conversely a quality score must never be presented as conformance without an explicit requirement and decision rule.
  • Provenance & Attribution - DQV itself reuses PROV for who produced a quality statement, when, and from what. This model references prov:wasAttributedTo / prov:wasGeneratedBy / prov:wasDerivedFrom rather than redefining provenance; it owns only the quality-specific slots (assessor competence, independence, method, evidence).
  • Measurement & Units - Uncertainty semantics (standard/expanded uncertainty, coverage interval, coverage probability, coverage factor, metrological traceability) are consumed from JCGM GUM/VIM; unit identifiers are referenced (DQV recommends sdmx-attribute:unitMeasure with dereferenceable unit definitions) and never minted here.
  • Risk & Harm - Confidence is a statement about the strength of an assessment, not a probability of harm and not an expected loss. Frameworks that grade certainty of evidence (GRADE) keep the certainty rating strictly separate from the effect estimate; this model preserves that separation and forbids arithmetic between confidence grades and risk scores.
  • Evidence & Attestation / Verifiable Credentials - A signed attestation carries issuer, proof, status and an evidence property, but W3C VC Data Model 2.0 defines no confidence property; assurance is derived by the verifier from issuer, proof, evidence and status. This model supplies the confidence semantics and treats the credential as a transport and integrity wrapper.
  • Sensing systems and system capability (SSN/SOSA) - ssn-system states device capabilities (Accuracy, Precision, Resolution, Drift, Latency) qualified by Condition, i.e. quality of an instrument under stated conditions. That is a capability declaration about a system, not an assessment of a produced result. WM-XCT-026 may cite such a capability as method context but keeps assertion-level results separate.
  • Catalogue/Dataset description (DCAT-style records) - A catalogue record describes availability and access; quality metadata is attached to it via a quality-metadata container rather than embedded as descriptive fields, so that quality statements can carry their own provenance, validity and access rules.

What else AI and robots need to interact with it Filled

Identity and identifiers required Filled

  • Authoritative master-system identifier assigned by the system of record for the subject, the measure, the certificate or the register item.
  • Governed global identifier or IRI issued by a recognised registry or standards body, including external measure, dimension, unit and scale identifiers.
  • UUID or ULID minted by the adopting Dimension, used only where neither of the above exists, and recorded together with the minting namespace.
  • A date, a version label, a filename, a storage path or a content hash alone is never an identifier; a content digest is an integrity control that may accompany, but never replace, an identifier.

Direct properties not applicable Not applicable

Not applicable

Institutional or informational subject: no invented physical properties.

Recognition optional Filled

  • A quality assertion names a subject, dimension, measure, method, result, scale and assessor.
  • Often confused with metadata about the dataset itself, a validation report and a rating.

Capabilities and actions required Filled

  • Register a quality measure: Create or revise a registered, versioned, dereferenceable measure definition with its dimension, definition, parameters, expected data type and value structure, following register-style registration and maintenance procedures.
  • Bind a quality profile to a subject class: Declare which dimensions and measures are mandatory or optional for a subject class, with thresholds, cadence and required confidence handling — the operation that makes this mixin attachable in a specific Dimension.
  • Plan an assessment scope and sampling design: Fix the assessed unit, extent restrictions, inspection mode, sampling design and reference basis before execution, so the resulting score is interpretable and not retrofitted.
  • Execute an assessment and record a result: Run the measure over the planned scope in a recorded environment and emit a result with value, scale, value structure, unit reference and status code.
  • Evaluate and attach measurement uncertainty: Evaluate uncertainty for a quantitative result and attach a standard or expanded uncertainty with coverage factor and coverage probability, or record an explicit reasoned omission.
  • Assign a scale-declared confidence with reasons: Attach a confidence value bound to a named, versioned scale, decomposed into recorded rating-down and rating-up reasons, evidence characterisation and agreement, with calibration status stated.
  • Evaluate conformance against a requirement: Apply a documented decision rule that accounts for measurement uncertainty to produce a verdict against a named requirement, with severity and any waiver reference.
  • Issue an assertion immutably: Transition an assertion from draft to issued, fix its content digest, set validity, apply access classification and retention class, and optionally wrap it in a secured credential.
  • Supersede or retract an assertion: Create a superseding assertion or a retraction with a recorded reason, preserving the append-only chain and notifying dependants.
  • Reconcile competing assertions: Detect assertions covering the same subject, scope and dimension that materially disagree, apply the governed precedence rule and record the disagreement instead of merging silently.
  • Decide fitness for a declared purpose: Combine cited assertions, verdicts, confidence and limitations into a purpose-relative fitness judgement with conditions and non-supported uses, recorded as its own assertion.
  • Expire assertions and schedule re-assessment: Detect assertions past their validity window or invalidated by subject, method or threshold change, mark them stale and enqueue re-assessment according to the profile cadence.
  • Export quality metadata in an alignment profile: Project assertions into an external target vocabulary using the maintained crosswalk, emitting only mappings of declared strength and attaching non-comparability notes.
  • Apply retention and disposition: Execute the retention schedule over assertions and evidence artifacts, honouring legal holds, producing tombstones where content is destroyed but the record must survive, and documenting retention/erasure conflicts.
  • Attach quality assertion: Bind a new quality measurement, annotation, certificate reference or policy binding to a host using the host's authoritative identifier and a new assertion identifier.
  • Ingest quality feedback: Accept a user or aggregator annotation, including conflicting assessments.
  • Derive composite score: Compute a derived measurement from other measurements with disclosed operators and weights.
  • Crosswalk dimension: Record or update a mapping from a local dimension or field to external catalogues without implying conformance.

Hazards and failure modes required Filled

  • False confidence leads to use of unfit data.
  • Mixed scales produce meaningless scores.
  • Stale assertions are trusted after the subject changed.

Standards and interfaces required Filled

  • ISO/IEC 25012 data quality model.
  • ISO 19157-1 geographic data quality.
  • W3C Data Quality Vocabulary (DQV).
  • W3C PROV-O.
  • JCGM 100:2008 Guide to the expression of uncertainty in measurement.

Context of use required Filled

  • Legal grounding for accuracy duties, rectification, erasure, storage limitation and declared accuracy metrics is drawn from EU instruments (Regulation (EU) 2016/679 and Regulation (EU) 2024/1689). Adopters in other jurisdictions must substitute equivalent instruments; the structural slots (retention class, legal basis, rectification outcome) are jurisdiction-neutral but their contents are not.
  • Output-quality principles are drawn from the European Statistical System's Code of Practice; other statistical systems use comparable but differently named principle sets, so principle names are treated as vocabulary bindings rather than universals.
  • NIST AI RMF is a voluntary United States framework, not a legal requirement; it is used only to support treating measurement as a separately governed function.
  • GRADE originates in health evidence assessment; its certainty domains are used as an evidence-decomposition pattern, not as a claim that clinical certainty grading transfers unmodified to other domains.
  • Time handling assumes RFC 3339 with explicit offsets throughout; adopters operating on local civil time or on calendar systems without a fixed offset must record the offset explicitly rather than relying on a local default.
  • ISO, W3C, BIPM/JCGM and NIST sources are treated as globally citable technical authorities; they are not a substitute for jurisdictional quality-of-data law.
  • GUM k approximately 2 for about 95 percent coverage assumes a near-normal case; other distributions need explicit justification.
  • Positional accuracy, topological consistency and gridded accuracy apply when the host is geographic or otherwise spatially referenced.
  • NIST AI RMF is a voluntary US public-authority framework; ISO/IEC 5259 is the international AI data-quality measures standard.

Sources Filled

  1. Data on the Web Best Practices: Data Quality Vocabulary (DQV) - World Wide Web Consortium (W3C)
  2. ISO/IEC 25012:2008 Software engineering — Software product Quality Requirements and Evaluation (SQuaRE) — Data quality model - ISO/IEC
  3. ISO 19157-1:2023 Geographic information — Data quality — Part 1: General requirements - ISO/TC 211
  4. JCGM publications: Guide to the expression of uncertainty in measurement (JCGM 100:2008 and GUM-series), Supplements 1–2, and JCGM 106:2012 The role of measurement uncertainty in conformity assessment - Joint Committee for Guides in Metrology (JCGM) / BIPM
  5. International Vocabulary of Metrology — Basic and general concepts and associated terms (VIM3), annotated online edition (JCGM 200:2012) - Joint Committee for Guides in Metrology (JCGM) / BIPM
  6. PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
  7. Shapes Constraint Language (SHACL) - World Wide Web Consortium (W3C)
  8. Semantic Sensor Network Ontology (SSN/SOSA) - W3C and Open Geospatial Consortium (Spatial Data on the Web Working Group)
  9. ISO 8000-8:2015 Data quality — Part 8: Information and data quality: Concepts and measuring - ISO
  10. ISO/IEC 5259-2:2024 Artificial intelligence — Data quality for analytics and machine learning (ML) — Part 2: Data quality measures - ISO/IEC JTC 1/SC 42
  11. ISO/IEC TS 4213:2022 Information technology — Artificial intelligence — Assessment of machine learning classification performance - ISO/IEC JTC 1/SC 42
  12. ISO/FDIS 19157-3 Geographic information — Data quality — Part 3: Data quality measures register - ISO/TC 211
  13. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) - European Parliament and Council of the European Union
  14. Regulation (EU) 2016/679 (General Data Protection Regulation) - European Parliament and Council of the European Union
  15. GRADE Working Group — the GRADE approach to rating certainty of evidence - GRADE Working Group
  16. European Statistics Code of Practice - Eurostat / European Statistical System Committee
  17. AI Risk Management Framework (AI RMF 1.0) - National Institute of Standards and Technology (NIST)
  18. Verifiable Credentials Data Model v2.0 - World Wide Web Consortium (W3C)
  19. FAIR Principles - GO FAIR International Support and Coordination Office
  20. ISO 19157 Data Quality Measures (DQM) XML schema, version 1.2.0 - ISO/TC 211 XML schema repository
  21. ISO/IEC 25012:2008 Software engineering — Software product Quality Requirements and Evaluation (SQuaRE) — Data quality model - International Organization for Standardization / International Electrotechnical Commission
  22. ISO 19157-1:2023 Geographic information — Data quality — Part 1: General requirements - International Organization for Standardization
  23. ISO 8000-1:2022 Data quality — Part 1: Overview - International Organization for Standardization
  24. JCGM 100:2008 Evaluation of measurement data — Guide to the expression of uncertainty in measurement (GUM 1995 with minor corrections) - Joint Committee for Guides in Metrology (BIPM, IEC, ILAC, ISO, IUPAC, IUPAP, IFCC, OIML)
  25. NIST AI Risk Management Framework 1.0 — AI Risks and Trustworthiness and AI RMF Core (Measure) - National Institute of Standards and Technology (United States)
  26. ISO/IEC 5259-2:2024 Artificial intelligence — Data quality for analytics and machine learning (ML) — Part 2: Data quality measures - ISO/IEC JTC 1/SC 42

Open questions

  • Verify the ISO/IEC 25012 fifteen characteristics and their inherent versus system-dependent assignment from normative text, then publish as a bound vocabulary profile artifact rather than as model structure.
  • Verify the ISO 19157-1:2023 element set from normative text and confirm that usability, a core element in ISO 19157:2013, is not retained in the 2023 edition; record the outcome as a version-mapping conflict in the alignment crosswalk.
  • Obtain the official ISO/IEC 5259-2 ML data-quality characteristic names from licensed text; both providers flag them as not extractable from freely published sources, so no closed ML quality catalogue may be claimed until then.
  • Verify the ISO 8000 part selection and terminology: the base cites Part 8 (concepts and measuring) with its internal syntactic/semantic/pragmatic terminology unverified, while the source cites Part 1 (overview). Determine which part governs the fitness-for-purpose grounding.
  • Resolve result polarity as question-level enrichment of the base result finding: whether a larger numeric value means better quality for a given measure, the theoretical best and worst values, and an exception code for an empty unit so that an empty sample is never reported as perfect completeness.
  • Define currentness evaluation for streaming and continuously updated hosts, specifically the sliding window or watermark that bounds the assessed state; neither the base nor the accepted additions cover this.
  • Find a primary source governing qualitative and nominal uncertainty comparable to GUM for quantitative results; both providers independently record this as an unresolved gap, and the base marks non-numeric quality semantics as its only declared coverage gap.
  • Settle the terminology guardrail between GUM coverage probability, GUM level of confidence and statistical confidence interval, and record the permitted labelling rules in the crosswalk so a coverage figure is never published under a statistical label it does not support.
  • Develop the spatial extent of the quality unit within a geospatial domain profile rather than the format-neutral mixin, including the spatial reference in which the extent is stated.
  • The specific fifteen ISO/IEC 25012 characteristic names and their inherent/system-dependent assignment were not verified from normative text; only the two-viewpoint structure and the count were confirmed from the ISO catalogue record. The model therefore binds the vocabulary by reference and does not enumerate it.
  • The internal terminology of ISO 8000-8 (syntactic, semantic and pragmatic quality; verification versus validation) is widely reported in secondary sources but was not confirmed from the normative text; only the standard's scope was verified. The fitness-for-purpose finding is grounded instead on ISO 19157-1 and the European Statistics Code of Practice.
  • A calibrated confidence-plus-likelihood exemplar from climate assessment practice was sought as a third independent grounding for ordinal confidence, but the authoritative documents were not retrievable at access time and were therefore excluded rather than cited from memory or from mirrors.
  • The full text of JCGM 100 and JCGM 106 was not read; the publication index and the annotated VIM were. Specific clause-level requirements on decision rules and guard bands are represented at the level of the concepts named in those publications' titles and in VIM definitions.
  • No domain-specific bindings are supplied for images, audio, code, physical goods or robotics sensing beyond the SSN capability alignment; the registry records a robotics factor of zero, and no robotics-specific quality node is asserted.
  • Aggregation and composite-index methodology is modelled structurally (function, weights, validity preconditions) but no primary source is cited for any particular aggregation function; specific composite indices must be justified locally.
  • Human inter-rater reliability statistics are referenced as an agreement measure without binding a specific coefficient, since no single normative choice was identified.
  • Full clause-level text of ISO/IEC 25012, ISO 19157-1, ISO 8000 parts beyond Part 1 overview, ISO/IEC 25024 measure tables and ISO/IEC 5259-2 additional ML characteristic names was not available as free complete primary HTML; those names must be copied from licensed text before claiming a closed catalogue.
  • JCGM GUM-1:2023 Bayesian/coverage-interval restatement is only partially cited via secondary catalogue notices; this model remains aligned to JCGM 100:2008/2010 as fetched.
  • UNECE NQAF, SDMX quality reporting, DDI quality statements, HL7 FHIR/clinical quality measures and FAIR metrics are not modelled; they are likely specialised siblings.
  • ISO 5725 accuracy of measurement methods, JCGM 200 VIM vocabulary and ISO/IEC 17025 calibration certificates are not fully incorporated.
  • Wang and Strong 1996 academic dimension set and DAMA-DMBOK uniqueness/validity are used only as conflict notes, not as primary catalogues.
  • Qualitative uncertainty for nominal classifications lacks a single primary standard comparable to GUM.
  • Region-specific legal accuracy duties (for example GDPR accuracy principle) are not encoded as legislation objects.

Machine files

Provenance

world-models research · reviewable-draft

Built from: models/wm-xct-026-quality-confidence/spec.yaml, ver-cy/world-models/card-supplements/wm-xct-026-quality-confidence.json