← Back to catalogue
Published

Test Case / Test Result

vr.wm-sft-015 · wm-sft-015-test-case-test-result

Provide a governed, format-neutral structure for repeatable software and system test cases together with the executions, results and evidence they produce, so an agent can define a reusable test, resolve it for execution, record what happened, adjudicate a governed verdict, retain defensible evidence, and expose that evidence to assurance and release consumers.

World Models Information and virtual systems INF.SFT.TST

Bundle → Layer → Finding → Questions Filled

9 bundles · 19 layers · 40 findings · 177 questions

Test Case Identity and Intent Everything needed to name a reusable test case unambiguously, to pin the exact revision an execution used, to state what the case exists to verify, and to bind it to the normative basis it was derived from.

Identity and Revision Control

The authoritative identifier of a test case, the identity of each revision of its definition, and the immutability guarantees that let past evidence stay interpretable.

Authoritative test case identity

How a reusable test case is uniquely and stably identified across authoring tools, repositories, management systems and execution systems, and how that identity survives renaming, relocation and tool migration. Identity is the anchor for every later reference from executions, evidence and release assurance.

  1. Which system holds the authoritative identifier for this test case, and what identifier value does it issue? identity
  2. Which alternate or legacy identifiers must resolve to this test case, and in which external systems do they live? interoperability
  3. What identity guarantees hold when the test case is renamed, moved between repositories or migrated between tools? constraint
  4. What distinguishes one test case from a variant, a duplicate and a differently-named copy of the same test? definition

Revision identity and definition immutability

How successive revisions of a test case definition are identified, superseded and frozen, so that any execution can name exactly the definition text it ran and a consumer can later prove the cited definition was not altered. Includes authorship and change-reason provenance of each revision.

  1. How is a new revision of the test case definition created, numbered and superseded? lifecycle
  2. Who authored or changed this revision, on what authority, and from which prior revision was it derived? provenance
  3. Which timestamps distinguish when a revision became effective from when the register recorded it? temporal
  4. Which changes to a definition require a new revision rather than an in-place editorial correction? constraint
  5. How can a consumer prove that the definition referenced by a past execution has not been altered? evidence

Objective, Classification and Test Basis Traceability

What the test case exists to verify, how it is classified for selection and reporting, and the typed links and derivation record connecting it to requirements, acceptance criteria, risks and coverage items.

Test objective, test item and classification

The single stated purpose of the test case expressed independently of how it is executed, the test item or component under test, and the controlled classification of test level, test type and quality characteristic that drives selection and reporting.

  1. What single objective does this test case exist to verify, stated independently of how it is executed? definition
  2. Which test level, test type and quality characteristic classify this test case, and from which controlled vocabulary? classification
  3. Which test item, component or interface is under test, and how is it referenced? relationship
  4. How is the objective checked for being single-purpose, observable and falsifiable? quality

Test basis traceability and derivation record

Typed links from a test case revision to the requirement, acceptance criterion, risk, specification statement or test assertion it was derived from, together with the design technique, coverage items and generation parameters that justify the derivation. Traceability is many-to-many by nature and does not claim one-to-one requirement coverage.

  1. Which requirement, acceptance-criterion, risk or specification statements does this test case revision trace to, and with what link type? relationship
  2. Which coverage items does this test case cover, and which test design technique produced them? composition
  3. Which normative statement gives the expected outcome its authority, and what is its prescription level? authority
  4. How is traceability completeness assessed without asserting one-to-one coverage of a requirement? validation
  5. Was the case written manually, derived from a model, or generated combinatorially, and with which generation parameters? provenance
Executable Test Specification The verifiable substance of the test case: the entry state it needs, the fixtures and data it references, the stimulus it applies, the outcome it expects with the rule used to compare against it, and the end state it leaves behind.

Entry State, Fixtures and Test Data

Conditions that must hold before execution begins and the versioned fixtures and data sets the case depends on, referenced rather than owned.

Preconditions and required entry state

The conditions that must hold before the case may be executed meaningfully, expressed so they can be checked rather than assumed, together with the disposition the definition declares when a precondition cannot be established. The definition declares intent; what actually happened in a given run is recorded by the execution structures of this model.

  1. What state must the test item, its data and the referenced environment be in before execution begins? state
  2. Which preconditions are mandatory gates and which are advisory assumptions that merely make the case inapplicable? constraint
  3. How is each precondition established: by the case itself, by a referenced fixture, or by an external set-up step? process
  4. What disposition does the definition declare when a precondition cannot be established? exception

Fixture and test data references

How the case names the fixtures, seed state and data sets it requires, with version, provisioning contract, determinism expectations and sensitivity classification, while the master lifecycle of that data stays with the owning test data system.

  1. Which fixtures and test data sets does this case require, and how are they referenced and versioned? relationship
  2. What sensitivity classification applies to the referenced test data, and is production-derived data permitted for this case? privacy
  3. Which parts of the referenced data are inputs, which are seeded state and which are expected baselines? composition
  4. What determinism and isolation guarantees must the data satisfy for the expected outcome to hold? constraint
  5. How long must a referenced data set remain resolvable for past evidence to stay interpretable? retention

Inputs, Parameters and Ordered Actions

The declared parameter surface of the case, the value sets that produce data-driven variants, and the ordered actions that apply the stimulus.

Inputs, parameters and parameterized variants

Declared inputs and parameters with types, defaults and value constraints, the basis on which values or ranges are chosen, the exclusion of invalid combinations, and how a parameter set produces data-driven variants without fragmenting the identity of the parent case.

  1. Which parameters does the test case declare, with what types, defaults and value constraints? composition
  2. When a parameter set produces many invocations, what is the identity of each variant relative to the parent test case? identity
  3. On what sampling, boundary or interaction basis are the input values or ranges chosen? measurement
  4. Which parameter combinations are excluded as invalid, unreachable or unsafe? constraint
  5. How are parameter values bound at resolution time from external configuration or credential stores? interoperability

Ordered actions and step structure

The ordered actions the test case prescribes, the granularity and identity of a step so expectations and evidence can attach to it, shared set-up versus discriminating actions, ordering significance, and delegation to reusable keywords or shared step definitions. The boundary against a test procedure that orders several cases is stated explicitly.

  1. What is the ordered sequence of actions the test case prescribes, and at what granularity is a step defined? process
  2. How is an individual step identified so that expected outcomes and captured evidence can attach to it? identity
  3. Which steps are shared set-up or teardown and which are the discriminating actions of this case? composition
  4. Is the step order significant, and which steps may be reordered, repeated or run concurrently? constraint
  5. Where a step delegates to a reusable keyword or shared step definition, how is that reference resolved? relationship

Expected Outcome, Oracle and End State

What a conforming test item must produce, the rule and tolerance by which agreement is decided, the baselines the comparison uses, and the state the case must leave behind.

Expected outcomes, oracle and tolerances

What the case declares as correct behaviour per step and overall, the oracle and comparison rule used to decide agreement, tolerances and units for approximate comparisons, the baselines the comparison references, and the normalization or masking applied before comparison. The definition declares the comparison rule; adjudicating an actual observation against it and recording the resulting verdict are performed by the execution structures of this model.

  1. What outcome must a conforming test item produce, stated per step and for the case overall? requirement
  2. Which comparison rule, tolerance and units apply when the expected outcome is numeric or approximate? measurement
  3. Which reference baseline or golden artifact does the comparison use, and how is it versioned? evidence
  4. Which parts of an observed outcome are normalized, masked or ignored before comparison? constraint
  5. What agreement criterion does the definition fix, and what is explicitly left for result adjudication to decide? decision

Postconditions, end state and side-effect limits

The state the case declares must hold after execution regardless of outcome, the cleanup or restoration obligations and who performs them, the state deliberately left for dependent cases, and the side effects that must never escape the test boundary.

  1. What state must hold after the case completes, independently of the outcome it produced? state
  2. Which cleanup or restoration obligations does the case declare, and who performs them? process
  3. Which downstream cases depend on state that this case deliberately leaves behind? relationship
  4. Which side effects are permitted, and which must never escape the test boundary? constraint
Binding, Applicability and Readiness What turns a reusable definition into something an agent may legitimately select and resolve: its binding to an executable implementation, the conditions and attributes that make it applicable and selectable, its dependencies, and the ownership and readiness state that authorises use.

Automation Binding, Applicability and Dependencies

How a definition points at the implementation that realises it, and the declared conditions, attributes and dependencies that determine when it is applicable and selectable.

Automation binding to an executable implementation

How a test case revision names the executable implementation, keyword decomposition or generated harness that realises it, the framework contract that binding must satisfy, how broken or ambiguous bindings are detected, and the automation status of the case. The implementation's own repository lifecycle stays outside this model.

  1. Which executable implementation realises this test case revision, and how is it addressed? relationship
  2. Which automation framework contract must the implementation satisfy for the binding to be valid? interoperability
  3. How is a binding detected as broken, ambiguous or pointing at a different revision than intended? validation
  4. Is this case manual, automatable, automated or generated, and who decides that status? classification
  5. What must remain true of the definition when only the implementation changes? constraint

Applicability, selection attributes and dependencies

The conditions under which a case is applicable, the tags, attributes and priority used to select it into a scope, time-bounded applicability, and declared dependencies on other cases. Evaluating a selection expression to build and dispatch an actual run belongs to the referenced suite, campaign and pipeline models.

  1. Which tags, attributes and priority values are declared for selection, and from which vocabulary? classification
  2. Under which product configurations, platforms, locales or feature flags is this case applicable? constraint
  3. Which other test cases must have succeeded, or must not run concurrently, for this case to be meaningful? relationship
  4. What selection expression would an agent evaluate to include this case in a given scope, and which model owns evaluating and scheduling it? decision
  5. Is applicability time-bounded, for example pending a fix, a deprecation date or a seasonal window? temporal

Stewardship and Readiness

Who is accountable for a definition, how it is reviewed, who may read or change it, and the readiness state that declares a revision fit to be resolved for execution.

Ownership, review and readiness for execution

The accountable owner and day-to-day maintainer of the definition, the review criteria and approval authority, the permitted readiness states and transitions of a definition revision, and the access distinctions that apply to reading and changing it. This concerns the definition's own readiness, not the state of any execution.

  1. Which role or team owns this test case definition, and who maintains it day to day? ownership
  2. Which readiness states can a definition revision hold, and which transitions between them are permitted? lifecycle
  3. Who may approve a revision as ready for execution, and on what basis? authority
  4. Who may read, propose changes to and publish this definition, and which parts are restricted? access
  5. Which review criteria must a revision satisfy before it is marked ready? quality
Execution Record and Run Context The identified execution event itself: which frozen test-case revision was exercised, with which parameter values, inside which run, against which build, in which environment, by which actor or tool, at which times, and as which attempt in a retained attempt series.

Execution identity and definitional binding

How an execution record is identified in its own right, how it is pinned to an exact test-case revision and criterion version, and which external records give it context and reproducibility.

Execution record identity and correlation

Every execution/result record is an independently identified entity, distinct from the test case it exercises, from the run that contains it and from any report that renders it. Identity is never derived from a date, a verdict value or a test-case identifier; a separate correlation key groups logically equivalent executions across runs.

  1. Which system of record issues the authoritative identifier for this execution record, and what is that identifier? identity
  2. When no master-system identifier exists, which governed IRI or Dimension-assigned UUID/ULID stands in, and how is that substitution marked? identity
  3. Which stable correlation key groups executions that are logically the same check across different runs? relationship
  4. What constraint prevents a date, a status value, a display name or a test-case identifier from acting as the record key? constraint

Binding to the exact test-case revision and parameter values

The execution binds to an immutable revision of the test case, its expected-result or oracle definition and the criterion or rule being evaluated, together with the concrete parameter values, data-set references and seeds actually used. A later edit of the test case must not silently change the meaning of a past result.

  1. Which immutable test-case revision, and which version of its expected-result or oracle definition, does this execution bind to? composition
  2. Which criterion or rule was evaluated, in which rule version and by which tool component version? identity
  3. What concrete parameter values, data-set references and randomisation seeds were bound for this execution? measurement
  4. What happens to this record if the referenced test case is subsequently modified, deprecated or withdrawn? constraint

Execution context and reproducibility references

The references that make an execution interpretable and repeatable: containing run and suite or campaign, the build or source revision under test identified by digest, the environment or configuration item, the executing actor, harness and tool version, and the observed invocation with its command line and exit status. All targets are externally owned records.

  1. Which run, suite or campaign contains this execution, and what identifies that container authoritatively? composition
  2. Which build, artifact or source revision was under test, identified by digest rather than by label? relationship
  3. Which environment or configuration item hosted the execution, and which model owns that record? relationship
  4. Which actor, harness or tool executed the case, in which version, and under what observed invocation? provenance
  5. What minimum context set must be present before this record may be treated as reproducible? quality

Execution time and attempt series

Time semantics for an execution attempt, and the append-only series of attempts produced by retries under one logical binding.

Execution time points, duration and clock provenance

Four time semantics are kept distinct: when the attempt started and ended (event time), when each outcome was observed by the executing agent, and when the record was ingested into the store. All are RFC 3339 with seconds and an explicit offset or Z, and none may be collapsed into another.

  1. What are the start and end times of this attempt, expressed with seconds and an explicit offset or Z? temporal
  2. When was each outcome observed by the executing agent, and when was the record ingested, and do these differ from the event time? temporal
  3. What duration is reported, measured from which clock and at what resolution? measurement
  4. How are unsynchronised, unknown-offset or reconstructed timestamps represented without fabricating precision? quality

Attempts, retries and append-only accumulation

A retry produces a new attempt record; it never re-opens or overwrites an earlier one. The finding covers attempt ordinal, predecessor linkage, retry trigger and the referenced retry policy, and how a disagreeing attempt series is summarised without destroying any member.

  1. What is this attempt's ordinal within the retry group, and which earlier attempt does it directly follow? event
  2. What triggered this retry — harness policy, orchestration rule or human decision — and which policy record applies? process
  3. Which rule guarantees that an earlier attempt's observations, timestamps and attachments cannot be mutated by a later attempt? constraint
  4. When attempts in one group disagree, how is the group-level outcome reported and which attempt is authoritative? state
Observed Outcome and Adjudicated Verdict What was actually observed and compared, step by step, and separately what verdict an assertor reached over those observations, including non-determinism evidence and the correction history of adjudications.

Observed outcome and comparison

The observable facts of an execution: the actual value or system state, how it was compared with the expectation, under what tolerance, and where in the subject the observation applies, including per-step granularity.

Actual outcome and comparison against expectation

The observed actual output or system state, the expectation compared against, the comparator and its version, the tolerance, normalisation or masking applied before comparison, the resulting delta, and the pointer identifying where in the subject the observation applies.

  1. What actual value, output or system state was observed, and in what representation was it captured? measurement
  2. Against which expected value or oracle was the observation compared, and by which comparator? validation
  3. What tolerance, precision, normalisation or masking was applied before the comparison was judged? constraint
  4. Where in the subject does this observation apply, expressed as a pointer, location or region? evidence
  5. How is an observation that could not be evaluated recorded without asserting either pass or fail? exception

Per-step and per-assertion outcomes

Ordered step, assertion or sub-test outcomes inside one execution, each with its own status, message and position, plus the point of first deviation and the disposition of steps after it. A container outcome is derived from members and never replaces the members' own records.

  1. Which ordered steps or assertions were executed, and what outcome does each carry in its own right? composition
  2. At which step did execution first deviate, and were the following steps executed, skipped or blocked? state
  3. How is the container or suite outcome derived from member outcomes, and is that derivation itself recorded? relationship
  4. What distinguishes a step from a nested sub-test, and how deep may the hierarchy legitimately go? definition

Verdict, non-determinism and correction

The asserted verdict over the observations, its authority and mode, the classification of non-deterministic behaviour across attempts, and the append-only correction history of adjudications.

Adjudicated verdict, assertor and status vocabulary

A verdict is an assertion made by a named assertor in a stated mode over recorded observations, expressed in a governed vocabulary covering passed, failed, error, blocked, skipped, inconclusive, not-applicable and untested, with explicit mappings to external vocabularies. The absence of a defect, of a failure element or of a report is never evidence of pass.

  1. Which governed verdict value applies to this execution, and what is that value's precise definition in the adopting Dimension? classification
  2. Who or what asserted the verdict, in which mode, and on whose authority? authority
  3. Which recorded observations and which oracle support this verdict, and was it derived or directly reported? provenance
  4. Which verdict and evidence combinations are prohibited, such as a pass with no evaluated expectation or a pass inferred from an absent defect? constraint
  5. How does this verdict map onto external vocabularies, and which mappings lose information? interoperability

Non-determinism and flaky-test evidence

Evidence that an identical binding produced differing outcomes across an attempt set, recorded as a property of that set rather than as a mutation of any attempt. Covers determinism classification, the measure and window used, and whether the classification is permitted to affect the reported verdict.

  1. Across which attempt set, under which proven-identical binding, did the observed outcomes disagree? quality
  2. How is this execution classified — deterministic, flaky, environment-induced or unknown — and on what evidence? classification
  3. May a non-determinism classification change the reported verdict, and which role is permitted to decide that? decision
  4. What instability measure is reported, over which observation window and with which denominator? measurement

Correction, supersession and result revisions

Execution records are append-only: an adjudication is corrected by adding a superseding revision that cites the record it supersedes, the reason, the correcting actor and the effective time. A rerun is a new attempt, never a correction, and superseded revisions remain retrievable.

  1. Which states may an execution record occupy, and which transitions between them are legal? lifecycle
  2. Which revision supersedes which, for what reason, by whom, and effective from when? provenance
  3. What distinguishes a correction of an adjudication from a new execution attempt, and which must be used in each case? constraint
  4. For how long must a superseded revision remain retrievable, and which policy owns that decision? retention
Evidence, Integrity and Interchange The byte-level evidence produced by an execution, the digests and attestations that seal it, the outbound references to defects and consumers, and the format projections through which results leave this model.

Evidence capture and integrity

Which byte-level evidence was captured for an attempt, how it is bound and classified, and how digests and signed attestations make the evidence set tamper-evident.

Logs, captures, traces and attachments

Byte-level evidence attached to an attempt or observation: harness standard output and error, screenshots, recordings, distributed traces, coverage data and crash dumps. Each attachment declares its role, media type, size, durable location, capture time, sensitivity classification and the attempt, step or observation it is bound to.

  1. Which evidence artifacts were captured for this attempt, and what role does each one play? evidence
  2. To which attempt, step or observation is each attachment bound, and may one attachment serve several of them? composition
  3. Do captures contain personal data, secrets or customer content, and what redaction was applied before storage? privacy
  4. Where is each artifact stored, and which reference stays meaningful if the store is rotated or purged? access

Digests, sealing and result attestation

Evidence integrity rests on named cryptographic digests over each artifact and result document, and on a signed statement binding the result to the subject under test by digest with the configuration used. Signing keys, trust roots and verification policy remain with the referenced supply-chain and identity models; this model records digests, attestation references and any verification outcome reported back to it.

  1. Which digest algorithm and value seals each evidence artifact and each emitted result document? evidence
  2. Which signed attestation, if any, binds this result to the subject under test, and under which predicate type? security
  3. What integrity-check outcome has a verifying system reported back, and how is it recorded here without this model performing the check? validation
  4. What minimum integrity metadata must be present before this result may be exported as release evidence? requirement

Outbound references and result interchange

How execution records point outward to defects, waivers and consuming builds, and how the format-neutral record is projected into interchange formats without those projections becoming authoritative.

Defect, waiver and consumer references

Typed outbound references from an execution record: defect or issue reports raised, reproduced or verified against it, waiver and known-issue references, and the builds or releases that consume the evidence. Only the reference, its role and its assertion time live here; every target keeps its own lifecycle.

  1. Which defect or issue records does this execution reference, and in which role — raised, reproduced, verified-fixed or known-issue? relationship
  2. Which model or system owns each referenced record's lifecycle, and which of its fields must never be copied into this record? ownership
  3. What rule prevents a pass verdict from being inferred from the absence of a linked defect or an empty failure list? constraint
  4. Which build or release records consume this evidence, and is that reference maintained here, at the consumer, or in both places? relationship

Result documents and format projections

The canonical execution record is format-neutral; JUnit-style XML, TAP 14, SARIF, Open Test Reporting documents and telemetry attribute sets are projections of it. This finding records which projection was emitted, its schema version, the semantics lost in the mapping, how identifiers survive projection and which side is authoritative on disagreement.

  1. Which interchange format and schema version was emitted, and which semantics does that projection lose? interoperability
  2. How do record identifiers survive projection into formats that identify tests only by name, class or ordinal? identity
  3. Which representation is authoritative when the emitted projection and the internal record disagree? quality
  4. How is an emitted document checked against its declared schema, and what is recorded when it does not validate? validation
Stewardship and controlled lifecycle Who is accountable for a test case or test result record, which roles must act before it becomes usable, and how it moves through defined states, revisions and supersession without losing its earlier meaning.

Accountability and role assignment

Assignment of ownership, stewardship and custody for records, and the role model that separates authoring, reviewing, approving and executing.

Record ownership, stewardship and custody

Identifies the accountable owner of a test case or test result record and its assurance claims, the steward performing day-to-day maintenance under delegated authority, and the custodian holding stored evidence when custody is separated from ownership. Ownership is a reference into the adopting Dimension's party registry, held with the effective moment from which it applies so that historical accountability remains reconstructable after a transfer.

  1. Which accountable party owns this test case or test result record and the assurance claims derived from it? ownership
  2. Which steward maintains the record day to day, and under what delegated authority do they act? authority
  3. Which custodian holds the stored evidence when custody is separated from record ownership? relationship
  4. What evidence records a transfer of ownership, and from which moment did it take effect? provenance

Role separation and independence constraints

Declares the author, reviewer, approver and executor roles required for a record to progress, the role combinations prohibited on the same revision, and how independence between authoring and execution is demonstrated at a given integrity level. Includes the documented exception path when a role conflict must be tolerated, so that waivers are visible rather than implicit.

  1. Which roles must have acted before a test case revision may reach an approved state? authority
  2. Which role combinations are prohibited on the same record revision? constraint
  3. How is the executor's independence from the author demonstrated at the applicable integrity level? evidence
  4. Which documented exception permits a role conflict on this record, and who granted it? exception

Lifecycle, approval and supersession

Defined states and transitions for test case definitions and result records, the approval act that gates them, and the rules that govern revision, correction and supersession.

Lifecycle states and approval gates

Defines the distinct state models for a test case definition and for a test result record, the transitions that require review and approval before a record may be cited as evidence, and the content of the approval act. Approval is captured with the signer, the moment of signing and the meaning of the signing, and is bound to the exact revision approved so it cannot be transferred to a later revision.

  1. Which lifecycle states are defined for a test case definition, and which for a test result record? state
  2. Which transitions require review and approval before the record becomes usable as evidence? lifecycle
  3. What signature manifestation is captured when a record is approved? authority
  4. How is an approval bound to the exact revision it approved rather than to the record as a whole? evidence
  5. Under what conditions may an approved record be returned to an earlier state, and what is recorded? decision

Revision identity, correction and supersession

Distinguishes a new revision of an existing test case from a new test case, and fixes the rule that an approved or finalised record is corrected by a superseding revision rather than by in-place edit. Records the validity window of each revision and defines what happens to results that reference a revision once superseded, so historical results remain interpretable against the definition actually executed.

  1. How is a revision of a test case identified, and what change makes it a new test case instead? identity
  2. Under what rule is an approved or finalised record corrected, by edit or by superseding revision? process
  3. Which timestamps establish the validity window of a revision that has been superseded? temporal
  4. What happens to test results that reference a test case revision after that revision is superseded? relationship
Provenance, evidence integrity and assurance claims How a result's origin and producing tools are recorded, how evidence artifacts are bound and proven unaltered, and how coverage, traceability and confidence claims are stated with their limits explicit.

Execution provenance and evidence integrity

Origin, agency and tool trust for a result, and the digest, attestation and immutability rules that keep its evidence verifiable.

Execution provenance and source or tool trust

Records which agent acted, under which plan, in which activity, and with which tool, harness or generator a result was produced, together with the environment and resolved dependencies relied on. Separates event time from observation and ingestion time so a delayed or replayed import cannot be mistaken for a fresh execution, and attaches a trust assessment to the producing tool so results from unqualified or self-reporting sources are distinguishable from results from qualified ones.

  1. Which agent, acting under which plan and activity, produced this test result? provenance
  2. How are execution event time, observation time and ingestion time recorded separately for one run? temporal
  3. How is the tool, harness or generator that produced the result identified and trust-assessed? quality
  4. Which environment and resolved dependency facts form part of the execution record? composition
  5. What does this provenance record explicitly not attest about the run? constraint

Evidence binding, integrity and immutability

Binds each evidence artifact to the result that cites it by cryptographic digest, so that the artifact is matched by content rather than by filename or location, and defines when evidence becomes immutable. Records the outcome of every integrity verification, including failures, so an unverifiable artifact is visibly degraded rather than silently dropped or silently trusted.

  1. What digest and algorithm bind an evidence artifact to the result record that cites it? security
  2. How is an evidence artifact shown to be unaltered after ingestion? validation
  3. At what point does evidence become immutable, and what may still change afterwards? constraint
  4. How is a failed or unverifiable integrity check represented rather than discarded? exception

Validation, coverage and traceability claims

Rules that make a record acceptable, and the disciplined statement of coverage, trace and confidence claims with their scope and caveats attached.

Record validation and confidence grading

Defines which fields and referential links must resolve before a result record is accepted, how outcomes that are neither pass nor fail are represented, and what confidence grade attaches to the record. Explicitly preserves inconclusive, blocked, aborted and error outcomes as first-class verdicts, and marks known-unreliable tests without discarding their history, so that aggregate figures are not silently improved by dropping inconvenient runs.

  1. Which mandatory fields and referential links must resolve before a result record is accepted? validation
  2. How is an inconclusive, blocked, aborted or errored execution outcome represented? exception
  3. What confidence grade is attached to a result, and on what basis is it assigned? quality
  4. How is a known-unreliable or intermittently failing test flagged without deleting its records? measurement
  5. What limit does this model state on what a passing result establishes? constraint

Coverage measurement as a bounded claim

Treats a coverage figure as a claim with a declared coverage item type, measurement method, numerator, denominator, exclusions and aggregation basis, not as a bare percentage. Coverage is technique-relative, so the same test set yields different figures under different item definitions. The claim carries a mandatory caveat that a coverage figure measures exercised items and does not by itself establish test adequacy, correctness or fitness for release.

  1. Which coverage item type, measurement method and denominator does this figure refer to? measurement
  2. Which scope boundaries and exclusions limit this coverage figure? constraint
  3. Which executions and definition revisions were aggregated to produce the figure? composition
  4. What caveat is carried so the figure is not read as proof of adequacy? quality
  5. How is a coverage claim referenced by an external gate without this model owning that gate? interoperability

Requirement, risk and control trace assertions

Holds the assertion that a test case exercises a stated requirement, risk or control, together with who asserted it, when it was last reconfirmed and whether the referenced target has since changed. Trace links are references outward: this model owns the assertion and its staleness state, while the requirement, risk or control model owns the target's definition, revision and approval.

  1. Which requirement, risk or control does this test case assert that it exercises? relationship
  2. Who asserted this trace link, and when was it last reconfirmed? provenance
  3. What happens to the trace link when the referenced target revision changes? state
  4. How is a trace link expressed so an external tool can resolve it without local knowledge? interoperability
Sensitivity, retention and controlled disclosure Classification and protection of test records that may carry personal or production data, the retention, hold and disposition regime for this model's own records, and the redaction, export and interoperability rules governing their release.

Sensitivity, personal data and retention

Classification at record, field and artifact level, treatment of personal or production-derived data, and the retention, hold and disposition regime.

Sensitivity classification and personal or production data

Assigns a sensitivity label at record level with overrides at field and artifact level, because a benign test case can cite a log containing production personal data. Records whether personal data is present, the origin class of the test data, and the minimisation or pseudonymisation measure applied before the data entered the record, so that later access, export and retention decisions rest on declared facts rather than on inspection of payloads.

  1. What sensitivity label applies to this record, and does it differ for individual fields or artifacts? classification
  2. Do the test inputs, logs or attachments contain personal or production-derived data? privacy
  3. Which minimisation or pseudonymisation measure was applied before the data entered the record? security
  4. Which fields must be masked for a reader holding only baseline access? access

Retention, legal hold and disposition

Binds each record class to a retention rule and disposition authority owned by the records-management model, records the trigger that starts the retention clock, and captures hold state that suspends disposition. Distinguishes the outcomes of disposition: full erasure, redacted stub, or tombstone that preserves the identifier and enough metadata for referential integrity. Retention clocks for a result and for its evidence artifacts may differ and are recorded separately.

  1. Which retention rule and disposition authority governs this record class? retention
  2. When does the retention clock start for the result, and when for its evidence artifacts? temporal
  3. Which legal hold suspends disposition, and who may place or release it? authority
  4. What remains after disposition: full erasure, redacted stub or tombstone? decision

Export, redaction and standards alignment

Controlled release of test records beyond their home boundary, and the disciplined mapping of this model onto external standards without unearned conformance claims.

Export packaging and redaction

Governs release of a test record and its evidence beyond its home boundary: which redaction profile applies, how integrity is re-established once redaction has changed artifact content, and what the recipient is told about what was withheld. Because artifacts are bound by digest, redaction necessarily produces new artifacts with new digests; the export manifest must therefore state the relationship to the originals rather than presenting redacted content as the original evidence.

  1. Which redaction profile applies when a test record leaves its home boundary? access
  2. How is integrity re-established for an export whose content has been redacted? security
  3. What does the export manifest tell the recipient about what was withheld? evidence
  4. Which approval is required before an export crosses an organisational or jurisdictional boundary? authority

Standards alignment, conformance limits and regional variation

Records which external standard or schema each field set aligns with, the status of each mapping, and what evidence would be needed before a conformance claim could be made. Alignment is a mapping assertion, not conformance: an implementation may align with a quality-management vocabulary and still fail its conformance requirements. Unresolved conflicts between aligned standards are held open as publication holds rather than resolved by silent preference, and regional or sector profiles that change governance obligations are recorded explicitly.

  1. Which external standard, schema or vocabulary does a given field set claim to align with? interoperability
  2. What evidence would be required before a conformance claim to that standard could be made? evidence
  3. Where do two aligned standards conflict, and how is that conflict held open? decision
  4. Which regional or sector-specific rule changes the governance obligations for these records? requirement
  5. How is a tool-specific report format bound without becoming this model's semantics? definition

Classifiers Filled

Family
World Models
Category
Information and virtual systems
Entry kind
aggregate
Navigation path
NAV.INF.SFT.TST
Domain
INF.SFT.TST
Industry
Cross-industry
Tags
testcaseresultinf.sft.tst

What it is Filled

WM-SFT-015 models a versioned test-evidence aggregate. The reusable test case definition is the aggregate root: authoritative identity and revision, title and objective, classification, requirement, risk and acceptance-criterion traceability, preconditions, fixture and test-data references, inputs and parameters, ordered actions, expected outcomes, oracle and comparison rule with tolerances, postconditions, automation binding, applicability, dependencies, variants, ownership and readiness. Governed members are the executions of a specific frozen definition revision, their observed outcomes, adjudicated verdicts and result status, the evidence captured during execution, the assurance signals derived over that evidence, and the governance record of ownership, approval, access and retention that binds them. Test suites and test sets, test plans and campaigns, requirements and acceptance criteria, risks, builds and releases, defects and test environments are typed external references: the model carries the reference, the binding and subject-specific parameters, and never reproduces those neighbours' lifecycles or operational functions.

In scope

  • Authoritative identity and revision history of a reusable test case, including the immutable definition snapshot that any execution cites.
  • Test intent: title, objective, test level and test type classification, and the design technique and coverage items from which a case is derived.
  • Typed traceability bindings from a test case revision to requirement, acceptance-criterion, risk, specification-statement and test-assertion references held by external models.
  • Preconditions and required entry state, expressed so they can be checked rather than assumed.
  • Fixture and test-data references with version, provisioning contract and sensitivity classification.
  • Declared inputs, parameters, value constraints and parameterized variants, including combinatorial or boundary-derived value sets.
  • Ordered actions and step structure, including delegation to reusable keywords or shared step definitions.
  • Expected outcomes, the oracle and comparison rule, tolerances, units and normalization rules, and postconditions with cleanup obligations.
  • Automation binding that names the executable implementation or keyword decomposition realising a test case revision.
  • Applicability conditions, selection tags and attributes, priority, and declared inter-case dependencies.
  • Governed execution records stating which definition revision was run, when, against which referenced test item and environment, by which actor and with which resolved parameters.
  • Observed outcomes, adjudicated verdicts, result status and reason codes, including reruns, retries and quarantined or unreliable results.
  • Evidence captured by an execution such as logs, measurements, screenshots, baselines and attachments, with provenance and integrity.
  • Assurance signals derived over the aggregate: traceability completeness, coverage claims, result reliability and readiness statements consumed by release assurance.
  • Governance of the aggregate: ownership and stewardship, review and approval state, access classification, and retention and disposition of test records.

Out of scope

  • Test suite and test set membership, ordering and suite-level lifecycle.
  • Test plan, campaign and cycle scheduling, resourcing and exit-criteria management.
  • Requirement, user story, acceptance-criterion and risk authoring, approval and change lifecycle.
  • Build, release and deployment identity and release-gating decisions, which WM-SFT-008 owns.
  • Defect and incident triage, assignment and resolution lifecycle.
  • Test environment provisioning, configuration baselining and infrastructure state.
  • Source code and executable test script implementation together with its repository lifecycle.
  • Pipeline orchestration, runner scheduling and job dispatch mechanics.
  • Runtime policy evaluation, enforcement engines and organisation-wide audit-trail semantics owned by referenced governance models.
  • Test tool and framework product lifecycle, licensing and vendor management.
  • Lawfulness assessment and erasure execution for production-derived personal data used as test data, which the adopting Dimension's privacy and storage models own.

Why it exists Filled

Provide a governed, format-neutral structure for repeatable software and system test cases together with the executions, results and evidence they produce, so an agent can define a reusable test, resolve it for execution, record what happened, adjudicate a governed verdict, retain defensible evidence, and expose that evidence to assurance and release consumers.

Distinguishing features Filled

  • Uses the versioned test case definition as the root, with every result tied to the exact revision it ran.
  • Separates observations from the adjudicated verdict, so evidence and judgement stay distinct.
  • Differs from a test script, which is one implementation, and from a test plan or suite, which group cases.
  • Records retries and non-determinism instead of hiding flaky results.

What robots and AI may and may not do Filled

Must not

  • Mark a test as passed without recorded observations.
  • Change an expected outcome after seeing the result to make it pass.
  • Discard failed attempts and keep only the passing retry.
  • Run tests against production data or systems without authorization.
  • Supersede an adjudicated result without recording why.

Only with a human decision

  • Accepting a waiver for a failed test before release.
  • Approving changes to acceptance criteria.
  • Running tests that touch real personal data or live safety systems.

May

  • Define and version test cases with traceability to requirements.
  • Execute a bound attempt and record observations with evidence.
  • Compare observed outcomes against expectations and tolerances.
  • Classify flakiness across attempts.

Moral aspects Filled

  • Test evidence often supports safety and regulatory claims; falsified results can put people at risk.
  • Test data drawn from real users must be protected or anonymised.
  • Testing should cover users with disabilities and diverse conditions, not only the typical case.

Who is affected

  • Users of the tested system
  • Testers and developers
  • Certifiers and regulators who rely on evidence

Owners Filled

Steward

The adopting Dimension must nominate a single accountable owner for the WM-SFT-015 package and publish the party registry against which owner, steward, custodian, author, reviewer, approver and executor references resolve.

Roles

Model owner
Hold accountability for the WM-SFT-015 package, its scope boundaries and its published governance artifacts.; Approve model revisions, breaking changes and the resolution or continued holding of alignment conflicts.; Nominate stewards and approve delegations of authority.
Record steward
Maintain test case and test result records day to day within a delegated authority.; Keep sensitivity labels, retention bindings and trace assertions current, and act on reconciliation findings.; Raise waivers where a separation-of-duties constraint cannot be met and record the compensating control.
Test author
Create and revise test case definitions, coverage item declarations and trace assertions.; Record the evidential scope note bounding what a definition can evidence.; Submit revisions for review and respond to review findings.
Reviewer
Examine a revision against the acceptance rules and the applicable integrity level before approval.; Record review findings and refusal reasons in a form bound to the revision reviewed.; Confirm that trace assertions and coverage declarations are supportable, not merely present.
Approver
Execute the approval act with printed name, moment of signing and declared meaning, bound to the revision digest.; Verify that the required role set and independence conditions for the transition are satisfied before signing.; Refuse and record the refusal where prohibited role pairs apply without a valid waiver.
Executor
Run the test procedure and produce result and provenance records with event, observation and ingestion times separated.; Record non-binary outcomes truthfully rather than coercing them to pass or fail.; Declare the tool, harness and environment used and any deviation from the approved procedure.
Evidence custodian
Hold stored evidence artifacts, maintain their digests and respond to integrity verification requests.; Report unverifiable or missing artifacts rather than substituting equivalents.; Apply storage-side protection consistent with the artifacts' sensitivity labels.
Records and privacy authority
Own retention schedules, disposition authorities and legal holds, and execute or authorise disposition.; Determine handling obligations for personal and production-derived data appearing in test records.; Adjudicate conflicts between erasure obligations and sector retention floors.
Release consumer
Cite assurance summaries and coverage claims by their stable identifiers when applying release gates owned elsewhere.; Apply their own thresholds and gate verdicts without writing back into this model.; Report citations that resolve to disposed or stale evidence.

Links to other meta-models Filled

child

  • WM-ACT-038 (registered parent model) - Positions the test-evidence aggregate beneath its registered parent so generic activity framing is inherited rather than restated; this model contributes only test-specific definition, execution, evidence, assurance and governance semantics.
  • WM-ACT-038 - Registered parent of WM-SFT-015. A test execution is an instance of a performed activity; this model specialises that activity for test execution and result recording and inherits the parent's generic activity identity and time framing rather than restating it.

references

  • WM-SFT-008 Build / Release - Release assurance references test case revisions, executions and evidence held here. This model publishes stable, resolvable evidence references and carries no release identity, no release lifecycle, no release-gating decision and no release audit trail.
  • Requirement / acceptance criterion model - Carries typed traceability links and the prescription level of the normative statement behind an expected outcome, while requirement authoring, approval and change lifecycle remain with the target.
  • Risk register / risk item model - Carries risk references that justify technique choice and priority for a case; risk identification, assessment and treatment lifecycle remain with the target.
  • Test suite / test set and test campaign model - Supplies selectable definitions and evidence to suites and campaigns; membership, ordering, scheduling, resourcing and exit criteria remain with the target.
  • Test environment / configuration model - Execution records cite an environment reference so results stay interpretable; provisioning, configuration baselining and infrastructure state remain with the target.
  • Defect / incident model - A failing result may cite a defect reference; defect triage, assignment, resolution and closure remain with the target and never become result state here.
  • Test data management model - Names data-set identifiers, versions and sensitivity classification required by a case; data generation, masking, stewardship and erasure execution remain with the target.
  • Test environment / configuration-item model (registry identifier not yet assigned) - Supplies the environment or configuration item in which an execution ran, so results are interpretable and comparable. Provisioning, drift and decommissioning lifecycles remain in the target; only the reference plus observed runtime facts are carried here.
  • Defect / issue report model (registry identifier not yet assigned) - Provides defect and issue records that an execution raised, reproduced or verified. This model records the reference, its role and its assertion time; triage, severity, state and closure belong to the target and are never mirrored here.
  • Actor, identity and access model (registry identifier not yet assigned) - Resolves the executing agent and the verdict assertor. Principal registration, credential lifecycle and entitlement evaluation stay in the target; this model stores only the reference and the assertion mode.
  • Audit log / event log model (registry identifier not yet assigned) - Receives actor, timestamp and reason for adjudications, supersessions and dispositions so they are durably recorded. Referencing an audit record grants this model no audit-trail semantics: ordering guarantees, tamper evidence and audit retention are the target's.
  • WM-SFT-008 (build or release) - Carry release and build reference identifiers cited alongside an assurance summary, and expose stable, citable assurance-claim identifiers. Release lifecycle, gate criteria, gate verdicts and deployment approval are owned by WM-SFT-008; this model never records a gate outcome or infers release fitness from its own claims.
  • Requirement, risk and control models of the adopting Dimension - Resolve trace assertion targets and their pinned revisions. This model holds the assertion, its author, its reconfirmation moment and its staleness state; requirement and risk definition, revision, approval and closure remain in the target models.
  • Defect and incident model of the adopting Dimension - Cite defect identifiers raised from failing results and blocking defects that prevent execution, aligned with the affectedByChangeRequest and blockedByChangeRequest relationship pattern. Defect triage, severity and resolution lifecycle are not mirrored here.
  • Party, person and organisation registry of the adopting Dimension - Resolve owner, steward, custodian, author, reviewer, approver, executor and export-approver references. Party mastering, identity proofing and organisational structure remain external.
  • Records management and retention schedule model - Bind each record class to a retention rule and disposition authority, and receive disposition confirmations. Schedule authorship, disposition authority and the execution of destruction, sanitisation and transfer remain in the records model.
  • Access control policy and enforcement model of the adopting Dimension - Supply scopes and sensitivity labels that an external policy decision point consumes. Policy evaluation, enforcement and session authorisation are executed externally; this model records only admissibility determinations for operations on its own records.
  • Audit and event log model of the adopting Dimension - Carry references to audit entries generated by operations on these records so an investigator can resolve them. Audit capture, sequencing, tamper-evidence and retrieval semantics are owned by the audit model and are not reproduced here.

aligned

  • ISO/IEC/IEEE 29119-3 test documentation templates - Aligns field naming and document boundaries with an international test documentation standard without claiming conformance; recorded as an alignment because clause-level template content was not verified in this research.
  • W3C Test Metadata - Aligns identity, title, purpose, status, specification reference, preconditions, inputs, expected results, version, contributor, rights and grouping to a published metadata vocabulary for tests.
  • OASIS Test Assertions Model Version 1.0 - Aligns normative source, prerequisite, predicate, prescription level, variable and tag to the traceability, precondition, oracle and parameter structures here, and adopts its many-to-many relation between normative statements and test cases.
  • ISO/IEC/IEEE 29119-5 keyword-driven testing data exchange - Aligns keyword decomposition and the interchange of test cases, test data and results with a standardised exchange format so definitions remain portable between frameworks.
  • OASIS SARIF Version 2.1.0 (Errata 01) - Alignment for run, invocation, result kind and level, artifact hashes, attachments, baseline state and correlation identifiers. Alignment only: conformance is not claimed, and SARIF's severity levels are mapped rather than adopted as this model's verdict vocabulary.
  • W3C EARL 1.0 Schema and ACT Rules Format 1.1 - Alignment for the assertion structure (assertor, subject, criterion, mode) and the outcome vocabulary including cantTell, inapplicable and untested, which supply the observation/adjudication separation and the inconclusive semantics used here.
  • in-toto Attestation Framework test-result predicate v0.1 - Alignment for binding a result to a digest-identified subject with the configuration used, and for expressing evidence integrity. Signing, key custody, trust policy and verification remain outside this model's boundary.
  • TAP 14, Open Test Reporting and JUnit-style XML (Maven Surefire schema) - Alignment for ecosystem projections of execution results, including streamed test points with SKIP/TODO directives, started/finished events with SUCCESSFUL/FAILED/ABORTED outcomes, and rerun/flaky elements. These are projections with lossy mappings, not the semantics of this model.
  • OpenTelemetry semantic conventions for test attributes - Alignment for the telemetry projection of test execution and for correlating traces to attempts. Its status vocabulary is deliberately narrower than this model's verdict list, so the mapping is recorded as lossy in both directions.
  • OASIS OSLC Quality Management Version 2.1 - Map governance-relevant fields onto TestCase, TestExecutionRecord and TestResult resource shapes and their validatesRequirement, executesTestScript, runsOnTestEnvironment and producedByTestExecutionRecord relationships, for cross-tool exchange. Alignment is asserted; conformance is not claimed without assessment evidence.
  • ISO/IEC/IEEE 29119-3:2021 test documentation templates - Map this model's lifecycle, approval and completeness elements onto the standard's documentation items so records can be rendered as recognisable test documentation. Clause text is paywalled, so the mapping is held as partial pending clause-level verification.
  • W3C PROV-O provenance ontology - Express execution provenance using Entity, Activity and Agent with wasGeneratedBy, wasAssociatedWith, actedOnBehalfOf and qualified Association carrying hadRole and hadPlan, so provenance is portable beyond this model's own projections.
  • in-toto Attestation Framework Statement v1 and SLSA Provenance - Bind evidence artifacts by digest-matched subject and express producing-platform facts using the builder, external parameters, resolved dependencies and start and finish metadata pattern. Attestation issuance, key management and verification policy remain with the attestation infrastructure.

composes

  • Sector and jurisdiction compliance profiles (data protection, AI, life sciences) - Overlay retention floors, log-keeping minima, signature-manifestation requirements and personal-data handling obligations onto the base governance rules. The profile supplies the obligation; the legal interpretation and the obligation's own lifecycle stay with the compliance model.

neighbor

  • Test procedure / test procedure specification - A test case states what is verified, under which entry state, with which expected outcome. A test procedure sequences one or more test cases with set-up and wrap-up for a specific execution occasion and carries its own identity and lifecycle; this model references a procedure but does not own its ordering or approval.
  • Executable test script / automated implementation - The script is an artifact of a source repository with its own versioning, review and build lifecycle. This model holds only a resolvable binding reference plus the contract the implementation must satisfy, so implementation change does not fork the definition.
  • Test suite / test set - A suite is a grouping and ordering construct over cases. Membership, ordering and suite-level status live with the suite model; this model declares only the tags, attributes and traceability that make a case selectable.
  • Test plan / test campaign - Planning covers scope, schedule, resourcing, entry and exit criteria for an effort. This model supplies definitions and evidence that a plan consumes and never holds schedule, staffing or exit-criteria decisions.
  • Scenario / use case - A scenario describes intended usage for stakeholders and may be executable prose. A test case adds a verdict-bearing expectation with an oracle and comparison rule; an executable scenario is treated as one authoring form of a test case specification, not as a separate lifecycle.
  • Requirement / acceptance criterion / test assertion - Normative statements and the testable assertions derived from them sit in the test basis. One assertion may require several test cases and coverage is often imperfect, so traceability is a many-to-many typed reference, not ownership of the requirement lifecycle.
  • Defect / incident report - A defect is a reported anomaly with independent triage, assignment and closure. A failing result may cite a defect reference, but defect state never becomes result state inside this model.
  • Build / release (WM-SFT-008) - Release assurance references test evidence held here. Release identity, release lifecycle, gating decisions and the release audit trail remain with WM-SFT-008; this model publishes resolvable evidence references only.
  • Test environment / configuration - Environment identity, configuration baseline and provisioning are external. An execution record cites the environment reference so results stay interpretable, without importing environment lifecycle.
  • Test data management system - Master data sets have their own stewardship, versioning and privacy controls. This model references a data-set identifier and version and classifies sensitivity for the binding, but does not own generation, masking or erasure execution.

What else AI and robots need to interact with it Filled

Identity and identifiers required Filled

  • Authoritative master-system identifier issued by the system of record that masters the subject, such as the quality or test management master system's test case identifier, the executing platform's execution-run identifier, or the records system's disposition identifier.
  • Governed global identifier or IRI from a namespace the adopting Dimension or an external registry controls, used when no master-system identifier exists.
  • UUID or ULID minted by the adopting Dimension, used only as a last resort and recorded together with the reason no higher-priority identifier was available.
  • A date, a timestamp, a filename, a storage path, a sequence position or a content digest alone is never an identifier. A digest identifies a byte sequence, not a governed record, and a filename is convenience metadata that may change without changing identity.

Direct properties not applicable Not applicable

Not applicable

Institutional or informational subject: no invented physical properties.

Recognition optional Filled

  • A test case has an identifier, revision, objective, preconditions, steps and expected outcomes; a result adds verdict and evidence.
  • Often confused with a test script, a test run log and a defect report.

Capabilities and actions required Filled

  • Define a reusable test case: Author a new test case definition from a stated objective and test basis references, capturing preconditions, fixture and data references, parameters, ordered actions, expected outcomes with a comparison rule, postconditions, applicability and ownership.
  • Version and freeze a test case revision: Create the next revision of a definition when a semantic change occurs, freeze the superseded revision, and publish the canonical snapshot and digest that executions and evidence will cite.
  • Validate a test case definition: Statically assess one definition revision for completeness, internal consistency and traceability sufficiency against a declared validation profile, and report findings to authors and reviewers.
  • Select candidate test case revisions: Return the set of test case revisions matching a selection expression over declared tags, attributes, traceability links and applicability context at a stated point in time, with the reason each case was included or excluded.
  • Resolve a definition for execution: Bind one test case revision, one parameter set, the referenced fixture and data-set versions and the automation binding into a fully-resolved definition handle that the execution structures cite when they record what was run.
  • Execute a bound test-case attempt: Run one attempt of an immutably bound test-case revision with declared parameter values in a referenced environment, emitting an observation stream and creating the attempt record.
  • Record an observation and its evidence: Append an observed outcome, per-step result, diagnostic message or evidence attachment to an open attempt, with its observation time and capturing agent.
  • Compare observed outcome against expectation: Apply the declared comparator, normalisation and tolerance to an observation and its bound expectation, producing a comparison record and delta without asserting a verdict.
  • Adjudicate a verdict over recorded observations: Assert a governed verdict for an execution, naming the assertor, the assertion mode and the observations relied upon, and rejecting prohibited verdict-and-evidence combinations.
  • Create a retry attempt: Open a new attempt under the same binding in response to a retry trigger, linking it to its predecessor and preserving all earlier attempt evidence unchanged.
  • Classify non-determinism across an attempt set: Adjudicate whether an attempt set with a proven-identical binding shows deterministic or flaky behaviour, record the instability measure and window, and state whether the classification affects the reported verdict.
  • Supersede an adjudicated result: Issue a superseding revision that corrects a previous adjudication, citing the superseded revision, the reason, the correcting actor and the effective time, while retaining the original.
  • Seal the evidence set: Close an attempt, compute named digests over every evidence artifact and the canonical result document, and assemble the evidence manifest optionally bound into a signed result attestation.
  • Export a result projection: Derive a non-authoritative interchange document for a consuming toolchain from the canonical record, applying the governed mapping table and recording the lossy mappings and schema version used.
  • Determine admissibility of a record operation: Determines whether a requested operation on a test case or test result record is admissible under this model's role assignments, separation-of-duties matrix and sensitivity profile, and records the determination with its basis. Authentication, external policy evaluation and enforcement at a runtime policy enforcement point are outside this model; the determination is a governance statement about the record operation, and the resulting audit entry is written to the adopting Dimension's audit model, which this model only references.
  • Capture review and approval of a revision: Captures a review outcome or approval act against one specific record revision, recording the signer, the moment of signing and the declared meaning of the signing, bound to the digest of the revision so the act cannot be transferred to another revision. Advances the record state only when the required role set and independence conditions for the target transition are satisfied.
  • Correct a record by superseding revision: Produces a successor revision that supersedes an approved or finalised revision, closing the predecessor's validity window and recording the correction reason, rather than editing the predecessor in place. Preserves every existing result's pointer to the revision actually executed, so historical results remain interpretable.
  • Compute a bounded assurance summary: Aggregates validated result records and coverage measurements into an assurance summary carrying the item type, method, numerator, denominator, exclusions, aggregated execution references, computation moment and mandatory caveats. The summary is a claim about exercised items and observed outcomes; it states no adequacy conclusion, defines no threshold and issues no gate verdict, which remain owned by the referencing build or release model.
  • Reconcile results against definitions and trace targets: Recomputes referential state between result records, the definition revisions they executed and the requirement, risk or control targets they assert, flagging orphaned pointers, superseded definitions and stale trace assertions. Reports divergence; it does not repair references or alter target models.
  • Verify evidence integrity: Recomputes digests for cited evidence artifacts, checks any attestations whose subject digest should match, and records the verification outcome and moment against the evidence manifest. Reports verified, unverifiable, mismatched or missing outcomes; it neither quarantines nor deletes artifacts, and it does not act as an enforcement point.
  • Produce a redacted export package: Assembles a release package for a stated recipient class by applying a redaction profile, issuing new digests for redacted artifacts with derivation links to their originals, and producing a manifest that discloses the counts and classes of withheld elements. Requires a recorded export approval before a package crosses an organisational or jurisdictional boundary.
  • Evaluate retention and propose disposition: Evaluates the bound retention rule against the recorded trigger event and hold state to determine whether disposition is due for a record and, separately, for its evidence artifacts, and emits a disposition proposal with its authority reference. Execution of destruction, sanitisation or archival transfer is performed by the records-management model and the adopting Dimension, not here.
  • Create a tombstone for a disposed record: On confirmation from the executing records system that disposition has occurred, replaces the record with a tombstone carrying the stable identifier, disposition outcome, authority reference and confirmation reference, so that inbound references from results, assurance summaries and external gates resolve to an explicit disposed state rather than to a dangling pointer.

Hazards and failure modes required Filled

  • False passes let defects reach users.
  • Flaky tests erode trust and hide real failures.
  • Tests against live systems cause damage or data exposure.

Standards and interfaces required Filled

  • ISO/IEC/IEEE 29119-3 test documentation.
  • ISO/IEC/IEEE 29119-4 test techniques.
  • ISO/IEC/IEEE 29119-5 keyword-driven testing.
  • JUnit XML report format (de facto).
  • ETSI TDL test description language.

Context of use required Filled

  • FDA General Principles of Software Validation is United States medical-device guidance; comparable duties elsewhere, including EU medical device regulation with IEC 62304, aerospace DO-178C, automotive ISO 26262 and rail EN 50128, impose analogous but differently worded documentation and retention obligations.
  • Retention periods for test evidence are jurisdiction- and sector-specific and are declared by the adopting Dimension rather than by this model.
  • Constraints on production-derived personal data used as test data assume a GDPR-like regime where one applies; other regimes permit or forbid different practices, and the privacy model reference absorbs the difference.
  • Authoring language and locale are assumed to be declared by the adopting Dimension; localized keywords in executable scenario formats affect parsing and must be recorded with the specification document.
  • Retention and access rules assume a data-protection regime of the EU GDPR type, where captured screenshots and logs may constitute personal data. Jurisdictions with different regimes will alter the retention schedule and redaction obligations without changing the model structure.
  • No sector-specific evidence regime is assumed. Regulated domains such as medical device software (IEC 62304), airborne software (DO-178C), automotive functional safety (ISO 26262) and clinical or laboratory settings impose additional mandatory evidence, independence-of-reviewer and record-retention requirements that the adopting Dimension must add as constraints.
  • Timestamps assume a network-synchronised clock. Air-gapped, embedded or offline execution environments may produce unsynchronised or unknown-offset times, which the model represents with -00:00 and a confidence note rather than by silently normalising.
  • English-language governed code lists are assumed; localisation of verdict and role labels is a presentation concern and must not create parallel code values.
  • Retention floors, log-keeping minima and technical documentation retention periods drawn from European Union regulation apply only where those regulations bind; other jurisdictions are assumed to supply their own profile through the jurisdiction profile mix-in.
  • Electronic signature and audit-trail expectations drawn from United States electronic-records regulation are treated as a sector profile for regulated life sciences, not as a universal default; organisations outside that scope may use a lighter approval manifestation.
  • Personal-data obligations assume a controller or processor model with purpose limitation and storage limitation. Jurisdictions with materially different data-protection architectures will require a different profile, and the model does not attempt to generalise across them.
  • United States national archives records requirements are used as a structural template separating retention binding, disposition authority and destruction execution. This separation is treated as good practice rather than as a legal requirement outside its jurisdiction.
  • Aviation, medical device, automotive and rail assurance regimes impose structural coverage and independence requirements that are only accommodated through the integrity level and coverage item type; no sector-specific coverage vocabulary is supplied, and one would be needed before use in those sectors.
  • Language, calendar and time-zone assumptions default to RFC 3339 with an explicit offset; organisations operating a non-Gregorian business calendar must map their retention triggers onto that representation themselves.

Sources Filled

  1. IEEE/ISO/IEC International Standard for Software and systems engineering--Software testing--Part 3: Test documentation - IEEE Standards Association (IEEE/ISO/IEC)
  2. IEEE/ISO/IEC International Standard - Software and systems engineering--Software testing--Part 4: Test techniques - IEEE Standards Association (IEEE/ISO/IEC)
  3. ISO/IEC/IEEE International Standard Software and systems engineering -- Software testing -- Part 5: Keyword-Driven Testing - IEEE Standards Association (ISO/IEC/IEEE)
  4. ISO/IEC/IEEE International Standard - Software and systems engineering -- Software testing -- Part 1: General concepts - IEEE Standards Association (ISO/IEC/IEEE)
  5. Test Metadata - World Wide Web Consortium (W3C)
  6. Test Assertions Model Version 1.0 - OASIS
  7. ITU-T Recommendation Z.161: Testing and Test Control Notation version 3: TTCN-3 core language - International Telecommunication Union (ITU-T)
  8. NIST Special Publication 800-142, Practical Combinatorial Testing - National Institute of Standards and Technology (NIST), U.S. Department of Commerce
  9. General Principles of Software Validation: Final Guidance for Industry and FDA Staff - U.S. Food and Drug Administration (CDRH and CBER)
  10. Gherkin Reference - Cucumber Open Source Project
  11. JUnit User Guide - JUnit Team
  12. Static Analysis Results Interchange Format (SARIF) Version 2.1.0 Plus Errata 01 - OASIS Open
  13. SARIF JSON Schema 2.1.0 (sarif-schema-2.1.0.json) - OASIS Static Analysis Results Interchange Format TC
  14. Evaluation and Report Language (EARL) 1.0 Schema - World Wide Web Consortium (W3C)
  15. Accessibility Conformance Testing (ACT) Rules Format 1.1 - World Wide Web Consortium (W3C)
  16. RFC 3339: Date and Time on the Internet: Timestamps - Internet Engineering Task Force (IETF)
  17. IEEE/ISO/IEC 29119-3-2021 — Software and systems engineering — Software testing — Part 3: Test documentation - IEEE Standards Association / ISO / IEC
  18. TAP Version 14 Specification - Test Anything Protocol
  19. in-toto Attestation Framework — Test Result predicate - in-toto project
  20. OpenTelemetry Semantic Conventions — Test attributes registry - OpenTelemetry (Cloud Native Computing Foundation)
  21. Open Test Reporting - JUnit team (ota4j-team)
  22. Maven Surefire test report XSD (surefire-test-report.xsd) - Apache Software Foundation
  23. Maven Surefire Plugin — Rerun failing tests - Apache Software Foundation
  24. OSLC Quality Management Version 2.1. Part 1: Specification - OASIS
  25. PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
  26. 21 CFR Part 11 - Electronic Records; Electronic Signatures - U.S. Food and Drug Administration / U.S. Government Publishing Office
  27. Regulation (EU) 2016/679 (General Data Protection Regulation) - European Union
  28. Regulation (EU) 2024/1689 (Artificial Intelligence Act) - European Union
  29. Universal Electronic Records Management (ERM) Requirements - U.S. National Archives and Records Administration (NARA)
  30. in-toto Attestation Framework: Statement layer, version 1 - in-toto project (Open Source Security Foundation)
  31. SLSA Provenance - SLSA / Open Source Security Foundation
  32. Secure Software Development Framework (SSDF), NIST SP 800-218 - U.S. National Institute of Standards and Technology (NIST)
  33. IEEE Std 1012-2024, IEEE Standard for System, Software, and Hardware Verification and Validation - IEEE

Open questions

  • Acquire clause-level text of ISO/IEC/IEEE 29119-3:2021 and 29119-4:2021 and IEEE Std 1012-2024 through licensed access, then re-map documentation-template field lists and V&V independence gradings; until then those citations support process separation only and no field-level mapping may be asserted from them.
  • Retrieve the ISTQB glossary through a non-client-rendered route (printable export or archived snapshot) so that a preferred terminology source can be used as normative evidence rather than excluded, and re-check the test case, test procedure, test script and test item definitions against the model's terminology choices.
  • Locate a normative basis for coverage adequacy thresholds — DO-178C objective tables, ISO 26262 structural coverage tables, IEC 62304 class-dependent requirements — before the coverage-adequacy gap row can be closed; the model currently and correctly delegates the threshold to the adopting Dimension.
  • Establish whether any specification-grade source defines flaky-test classification and a flake-rate denominator; the current structure rests on Maven Surefire's flakyFailure and rerunFailure harness convention, which the package itself declares is not a specification.
  • Design and source an ingestion function for externally produced result documents, covering identifier carriage, schema validation failure recording and lossy-mapping capture, since ingestion is referenced by canonicalization rules, crud.create and known_omissions but is modelled by no function.
  • Research long-horizon evidence verifiability for decade-scale retention: cryptographic algorithm agility, key rotation, trusted timestamping and long-term signature validation, which the package currently delegates and names as a real unaddressed risk.
  • Define a supplier and multi-party test-evidence acceptance policy for evidence originating outside the adopting Dimension's trust boundary, which is presently handled only through tool trust level and integrity status.
  • Investigate performance, load and reliability result semantics — percentile aggregation, warm-up exclusion, statistical significance of a regression — and decide whether they belong in a dedicated measurement model rather than as parameters and tolerances here.
  • Clause-level text of ISO/IEC/IEEE 29119-3 templates and 29119-4 technique definitions was not machine-readable in this research; catalogue scope and abstract text were used and no field-level structure rests on those citations alone.
  • The official ISTQB glossary is a client-rendered application that could not be retrieved, so ISTQB terminology is not used as normative evidence anywhere in this result despite being a preferred source class.
  • Non-functional test design parameters such as load profiles, performance workload models and security abuse cases are represented only generically through parameters, tolerances and classification.
  • Model-based testing artefacts such as behaviour models and generation grammars, and structural constructs of standardised test notations such as components, ports and defaults, are referenced but not modelled.
  • Accessibility and localisation-specific expected-outcome rules, and property-based or metamorphic oracle families, are not specialised beyond the generic comparison-rule structure.
  • Cost, effort and maintenance-burden attributes of a test case, which some management systems treat as first-class, are omitted for want of primary support.
  • Test-case definition and design: preconditions, procedure steps, expected-result authoring, oracle design, test-data design and requirement traceability are left to the sibling definition pass and are only bound to here.
  • Assurance aggregation: suite and campaign roll-ups, coverage claims, exit criteria and test completion reporting are left to the sibling assurance pass; only per-container derivation of an outcome from members is modelled here.
  • Canonical governance service layers: the merger takes these from the governance pass, so the service_layers block supplied here is deliberately concise and execution-scoped.
  • Performance, load and reliability result semantics: percentile aggregation, warm-up exclusion and statistical significance of a regression are not modelled; they would need a dedicated measurement model.
  • Hardware-in-the-loop and real-time testing: sub-millisecond timing, clock domains, sampling jitter and physical instrumentation traceability are out of reach of RFC 3339 alone and are not addressed.
  • Manual exploratory session semantics: session charters, time-boxing and note-taking conventions are referenced only through the manual assertion mode.
  • Clause-level field lists from ISO/IEC/IEEE 29119-3:2021 could not be read directly (the ISO page returned HTTP 403 and the IEEE abstract does not enumerate template fields), so that standard supports process separation rather than specific field names here.
  • Clause-level text of ISO/IEC/IEEE 29119-3:2021, ISO/IEC/IEEE 29119-4:2021 and IEEE Std 1012-2024 was not machine-extracted; only publisher metadata, scope and abstract were verified live. Field-level mappings to those standards' templates are therefore held as partial rather than asserted.
  • NIST SP 800-218 task-level text for PS.3, PW.7, PW.8 and the RV practices sits in a PDF that did not convert; only the project page's practice-group statements and the note that version 1.1 added a provenance-data task were verified. Task-level citations are avoided.
  • FDA General Principles of Software Validation is cited from its landing-page metadata only; specific statements about testing limitations, tester independence and record retention were not extracted and are not attributed to it.
  • No test management tool schema is modelled. Common execution report formats are treated as import mappings with recorded gaps, not as semantics, so format-specific fields with no model counterpart are currently undocumented.
  • Cryptographic algorithm agility, key rotation and long-term signature validity for evidence retained across decade-scale retention periods are delegated rather than modelled; long-horizon verifiability is a real unaddressed risk.
  • Manual and exploratory testing evidence, where the executing agent is a person and the tool trust concept applies weakly, is accommodated but not specialised; the provenance structure fits automated execution more naturally.
  • Cost, effort and scheduling attributes of test execution are absent; they belong to a work-management model that was not identified in the relation ledger.
  • Multi-party and supplier-supplied test evidence, where the evidence originates outside the adopting Dimension's trust boundary, is handled only through tool trust level and integrity status; no supplier attestation acceptance policy is modelled.

Machine files

Provenance

world-models research · reviewable-draft

Built from: models/wm-sft-015-test-case-test-result/spec.yaml, ver-cy/world-models/card-supplements/wm-sft-015-test-case-test-result.json