Data Lineage
Provide a format-neutral, composable pattern for asserting, evidencing and governing how datasets and their fields were derived from other datasets and fields through processing activities, so that an agent can reconstruct derivation paths, perform impact analysis and evidence provenance obligations.
Bundle → Layer → Finding → Questions Filled
7 bundles · 14 layers · 28 findings · 109 questions
Lineage Subject Identity and Perimeter What a lineage assertion is about: the identity and resolution of the dataset, field, activity and agent nodes it connects, and the declared granularity and coverage perimeter of the lineage graph.
Node Identity and Resolution
Identity rules for the four node kinds a lineage edge can touch: dataset or distribution, field, processing activity, and responsible agent.
Dataset and distribution node identity
How a lineage node that stands for a dataset or one of its physical distributions is identified, bound to a location-derived name, and resolved to the authoritative dataset record held by WM-DAT-001.
- Which authoritative master-system identifier identifies the dataset this node stands for, and which system of record issued it? identity
- What location-derived binding (namespace and name) is recorded for this node, and under which source-type convention was it constructed? interoperability
- Does this node denote the abstract dataset or one specific distribution or physical materialisation of it? classification
- How is a node re-resolved when the dataset is renamed, relocated or its namespace convention changes? exception
Field and column node identity
How a field-level lineage node is identified so that column lineage survives renames and schema evolution, given that column lineage facets key on field name while table formats assign immutable numeric field IDs.
- What stable field identifier does the source system assign to this column, and is it invariant under rename? identity
- Which field name was observed at capture time, and to which parent dataset node and schema version does that name belong? provenance
- How are nested, repeated or struct-typed fields addressed within a single field node? composition
- What happens to existing field-level edges when a column is renamed, retyped or dropped? constraint
Processing activity and agent references
How the activity that performed a derivation and the agent responsible for it are referenced, keeping job and run definition with WM-DAT-005 and agent records with the adopting Dimension's party model.
- Which processing activity is asserted to have generated the output node, and how is that activity identified? identity
- Which agent kind bears responsibility for the activity: a person, an organization or a software agent? classification
- Is the activity node internal to the observed perimeter or an opaque external activity known only by reference? relationship
Granularity and Perimeter Frame
The declared resolution of the lineage graph and the explicit statement of what lies inside, outside and opaque within the observed perimeter.
Declared granularity and coverage perimeter
The explicit declaration of the level at which lineage is asserted and the boundary of what was observed, so that absence of an edge can be distinguished from absence of observation.
- At which granularity levels is lineage asserted for this subject, and which levels are explicitly not asserted? definition
- Which systems, zones or hops lie inside the observed perimeter, and which are recorded as opaque or black-box segments? composition
- Does the absence of an upstream edge for a node mean no derivation exists, or that no observation was made? quality
- Which manual steps, spreadsheets or user-side workarounds are declared as part of the flow rather than omitted? process
Derivation and Transformation Semantics The meaning of the edge itself: what kind of dependency it asserts, how the transformation is characterised, and what structural constraints keep the resulting graph coherent.
Derivation Edge Semantics
The typed dependency relations between nodes and the characterisation of the transformation that realises them.
Edge relation typing and identity
Which relation an edge asserts, how the edge is itself identified so it can carry attributes, times and roles, and how directionality and multiplicity are constrained.
- Which relation does this edge assert between its endpoints, and what is the source of that typing? relationship
- Is the dependency direct, meaning output values are derived from input values, or indirect, meaning the input influenced the output without contributing values? classification
- Does this edge carry its own identifier and attributes, or is it an unqualified binary link? identity
- What role did the input node play in the activity that produced the output? authority
Transformation characterisation and logic evidence
How the transformation behind an edge is described: its typed subtype, whether it obfuscates values, and the reference to the expression or mapping that realises it.
- Which transformation subtype best characterises this edge, and from which controlled vocabulary is it drawn? classification
- Does the transformation mask, hash, pseudonymise or otherwise obfuscate the input values? privacy
- Where is the executable logic that realises this transformation, and at which version was it observed? evidence
- Which data-preparation operation class does this transformation fall into for documentation obligations? requirement
Graph Structure and Integrity
Constraints that make a lineage instance logically coherent, and the treatment of collections, partitions and aggregate nodes.
Structural consistency constraints
The uniqueness, typing, ordering and impossibility constraints a lineage instance must satisfy, and the normalisation rule that makes two differently expressed instances comparable.
- Which uniqueness constraints apply, so that the same generation or usage is not asserted twice with conflicting attributes? constraint
- Which event-ordering constraints must hold between generation, usage, start, end and invalidation on this instance? temporal
- Which node and edge combinations are declared impossible, including type disjointness and reflexive specialization? validation
- Are cycles permitted in the derivation graph, and how are recursive or self-referencing derivations represented? constraint
Collection, series and partition lineage
How lineage is asserted for aggregates whose members change over time: dataset series, partitions, file sets and logical collections, and how member-level edges roll up to collection-level edges.
- Is this node a collection, and if so which membership relation defines its members at the asserted time? composition
- How do member-level derivation edges aggregate into an edge asserted at collection level? relationship
- How are partition-scoped derivations expressed when only part of an output was regenerated? state
- When does a new member make a series a new node rather than an update to the existing collection node? lifecycle
Processing Context and Responsibility The context that makes a derivation reproducible and attributable: the run under which it was observed, the plan and configuration that governed it, and the chain of responsibility for both the processing and the lineage record.
Execution Reference Binding
How a lineage assertion is bound to the execution occurrence and configuration under which it was observed, without importing run lifecycle ownership.
Run reference and observation binding
The carried reference to the execution occurrence under which the derivation was observed, including the run identifier, parent run reference and terminal outcome used only to interpret completeness of the assertion.
- Under which run identifier was this derivation observed, and which model owns that run record? provenance
- What terminal outcome did the referenced run report, and how does that outcome qualify the completeness of this assertion? state
- Is this assertion part of a nested run hierarchy, and which parent run does it belong to? composition
- For a continuous or streaming activity with no terminal state, what bounded observation window does this assertion cover? temporal
Plan, code and configuration provenance
The version-pinned description of the logic and configuration that governed the derivation, sufficient to explain and re-derive it, expressed as references to code, parameters and engine identity.
- Which plan or program governed the activity, and at which immutable version was it observed? provenance
- Which externally supplied parameters and which platform-internal settings influenced the derivation? process
- Which processing engine and resolved dependencies were in effect, and how are they identified? interoperability
- Is the recorded plan sufficient to re-derive the output, and what is explicitly not captured? quality
Responsibility and Accountability
Attribution and delegation for the processing, and named accountability for the correctness of the lineage record itself.
Attribution, association and delegation
Who is responsible for the derived entity, who was associated with the activity, and on whose behalf a software agent acted, keeping party records in the identity model.
- To which agent is the derived entity attributed, and on what basis was that attribution made? ownership
- Which agents were associated with the activity, and in which roles? authority
- When a service account or automated agent acted, on whose behalf did it act and under which authorisation? authority
- Does attribution transfer to downstream derived entities, or must it be re-asserted at each hop? relationship
Lineage stewardship and accuracy attestation
Named accountability for the lineage record itself: who maintains it, who attests that it reflects the real flow, and on what cadence that attestation is refreshed.
- Which named role is accountable for the accuracy of this lineage record, distinct from the owner of the underlying data? ownership
- On what cadence, and against which evidence, is the lineage record attested as reflecting the actual flow? process
- What escalation applies when the accountable role rejects or cannot confirm the recorded lineage? exception
Time, Version and Validity Making a lineage assertion interpretable in time: separating when the derivation happened from when it was recorded, pinning the dataset states involved, and defining when an assertion stops being true.
Time Semantics
The distinct time axes a lineage assertion carries and the validity window over which it holds.
Event time versus observation time
The separation of when the derivation occurred from when the assertion about it was captured and recorded, plus nominal or scheduled time where it differs from both.
- At what event time did the derivation occur, expressed with seconds and an explicit offset? temporal
- At what observation or ingestion time was this assertion captured and recorded, and by which producer? provenance
- What nominal or logical time does the processed data represent, where that differs from wall-clock event time? temporal
- How are late-arriving or out-of-order assertions ordered and merged against already-recorded ones? process
Assertion validity, supersession and retraction
The lifecycle of the lineage assertion as a record: when it becomes effective, when it is superseded by a corrected assertion, and how a retraction is expressed without erasing history.
- Over which validity window is this assertion held to be true, and what ends that window? lifecycle
- Which prior assertion does this one supersede, and what was the reason for the correction? provenance
- How is a retracted assertion represented so that consumers who already read it can detect the retraction? exception
- Which state may a lineage record occupy, and which transitions are permitted? state
Version Binding and Schema Change
Pinning the dataset states an edge connects and preserving lineage continuity when structures change.
Version and snapshot pinning
How an edge names the specific state of the input and output datasets it relates, using the version or snapshot identity issued by the owning system, plus an integrity digest where available.
- Which specific dataset state does each endpoint of this edge refer to, and which identifier scheme expresses it? identity
- What ordering key allows two states of the same dataset to be compared? temporal
- What content digest or checksum evidences the integrity of the referenced state? evidence
- If the endpoint state is not addressable, what weaker binding is recorded and how is that weakness declared? quality
Schema and naming change continuity
Preserving lineage continuity when fields are renamed, retyped or dropped, or when a naming convention changes and location-derived node names would otherwise disconnect.
- Which schema change occurred, and which fields and downstream edges does it affect? event
- Which identifier remained stable across the change, and which binding had to be re-established? identity
- What is the risk that a namespace or naming-convention change silently disconnects existing lineage nodes? validation
- Which downstream consumers must be informed of the change, and by which route? decision
Capture, Evidence and Assurance Why a lineage assertion should be believed: how it was obtained, what supports it, how the instance is validated, and how coverage is measured and reconciled.
Capture Method and Provenance of the Lineage Record
How each assertion entered the record and how much weight it carries.
Capture method and producing agent
The method by which the assertion was obtained (runtime observation, static parsing, declaration, inference or manual entry) and the identity and schema version of the producer that emitted it.
- By which method was this assertion obtained, and what does that method structurally fail to see? provenance
- Which producer emitted the assertion, and against which immutable schema version? interoperability
- Is provenance recorded about the lineage record itself, and where is that meta-provenance held? provenance
- When two producers assert conflicting lineage for the same endpoints, which one prevails? exception
Confidence grading and supporting evidence
The graded confidence attached to an assertion and the evidence references that justify the grade, including explicit flagging of manual or uncontrolled steps.
- What confidence grade is assigned to this assertion, and against which published grading scale? measurement
- Which concrete evidence items support the assertion, and are they retrievable? evidence
- Is any part of the asserted path dependent on a manual workaround or uncontrolled step? quality
- What would falsify this assertion, and has that test been run? validation
Validation and Reconciliation
Evaluating the lineage instance against constraints and measuring how much of the real flow it actually covers.
Lineage instance validation
Structural evaluation of a lineage instance against uniqueness, typing, ordering and impossibility constraints, plus referential checks that every node reference resolves.
- Which constraint profile and schema version was the instance evaluated against? validation
- Which violations were found, at which severity, and on which nodes or edges? quality
- Which node references failed to resolve, leaving dangling or orphan nodes? validation
- Does a failing validation block acceptance of the assertions, and who decides? decision
Coverage measurement and reconciliation
Quantifying how much of the declared perimeter is actually covered by asserted lineage and reconciling asserted paths against independent evidence of the real flow.
- What proportion of the declared perimeter has asserted lineage at each granularity level? measurement
- Which asserted paths were reconciled against independent evidence, and which diverged? evidence
- Which known gaps remain unclosed, and what is the accepted remediation horizon? quality
- How is coverage reported to the accountable role and to supervisory or assurance consumers? process
Lineage Record Governance Governance of the lineage record as an object in its own right: how sensitive it is, who may see which parts, how long it is kept, and which external obligations bind it.
Sensitivity and Disclosure Scoping
Classification of the lineage record and the scoping requirements it places on disclosure, with evaluation and enforcement left to the access model.
Sensitivity of the lineage record
Classification of a lineage record in its own right, recognising that node names, transformation expressions and masking flags can disclose confidential logic or reveal where personal data flows.
- What sensitivity classification applies to this lineage record, and is it higher or lower than that of the data it describes? classification
- Which elements of the record could disclose personal data footprints or confidential business logic if released? privacy
- Which elements must be redacted or generalised in a lower-trust projection of the record? security
- Does propagating a masking flag along a derivation path change the classification of downstream nodes? constraint
Disclosure scoping and exceptions
The scope levels at which lineage may be disclosed, the default denial posture, and the named exceptions such as regulatory or incident access, with the evaluation decision left to the access model.
- At which scope may a requester read lineage: whole bundle, layer, individual finding or a specific artifact? access
- Which named exceptions permit broader disclosure than the default rule, and for how long? exception
- Which model evaluates and enforces the disclosure decision, and what does this model contribute to it? authority
Retention and External Obligation Binding
How long lineage records are kept, how deletion is represented without destroying traceability, and which external obligations bind the lineage held for a subject.
Retention, tombstoning and disposition
The retention rule for lineage records, the tombstone representation used when a referenced node or record must be removed, and the propagation duty toward downstream copies, with execution owned elsewhere.
- What retention period applies to this lineage record, and which obligation sets it? retention
- When a referenced node must be erased, is the lineage record deleted or tombstoned, and what remains readable? lifecycle
- Which downstream holders of copies or replications must be informed of an erasure, and how are they identified? relationship
- Which party actually executes the deletion in the storage substrate, and where is that execution recorded? authority
External obligation binding
Which external obligations require lineage for a given subject, what each requires, and which lineage elements satisfy the requirement, recorded as binding parameters rather than as a reproduction of the obligation.
- Which external obligations require documented provenance for this subject, and under which jurisdiction? requirement
- Which specific lineage elements are relied on to satisfy each obligation? evidence
- Does the obligation require a narrative lineage statement, a machine-readable graph, or both? interoperability
- Does the obligation require recording the original purpose of collection for personal data upstream of a derivation? privacy
Interoperability and Exchange Making lineage portable: alignment crosswalks to external provenance standards, governance of namespaces and extensions, resolution of nodes across system boundaries, and the integrity of exchanged payloads.
Standard Alignment and Extension Governance
Recorded crosswalks to external provenance vocabularies and the rules governing namespaces and custom extensions.
Alignment crosswalk and conformance claims
Element-by-element mapping between this model and external provenance vocabularies, with explicit lossiness notes and evidenced conformance claims rather than assumed compatibility.
- Which external vocabulary elements does each local element map to, and in which direction is the mapping lossless? interoperability
- Which conformance claims are made against which standard version, and what evidence supports each claim? validation
- Where do the aligned standards conflict, and which one governs in this Dimension? exception
- How is a narrative lineage statement reconciled with the machine-readable graph for the same subject? quality
Namespace and extension governance
Rules for allocating namespaces used in node identity and for defining custom extensions, so that identifiers stay stable and extensions do not collide with governed terms.
- Who allocates a namespace used in node identity, and what makes an allocation immutable once used? authority
- What prefix and naming rule prevents a custom extension from colliding with governed terms? constraint
- How is an extension schema version referenced so that a consumer can resolve exactly the version emitted? interoperability
- What is the promotion path from a local extension to a governed term, and what changes on promotion? lifecycle
Cross-system Resolution and Exchange
Stitching lineage graphs produced by different systems and exchanging them with verifiable integrity.
Cross-system node resolution and stitching
How nodes asserted independently by different producers are recognised as the same thing or as related specialisations, and what evidence is required before two graphs are joined.
- On what evidence are two independently asserted nodes judged to denote the same thing? identity
- Is the relationship between two nodes identity, alternate presentation, or specialisation of a more general node? relationship
- How are symlink or mirror identifiers for the same physical dataset recorded without creating duplicate nodes? composition
- What is done when a stitching decision is later found to be wrong? exception
Exchange payload contract and integrity
The serialization-neutral contract for a unit of lineage handed to another system, including merge semantics for partial or accumulative delivery and the integrity evidence that binds payload to producer.
- What is the minimum element set a payload must carry to be independently interpretable? requirement
- Is the payload a complete self-contained snapshot or an accumulative increment to be merged? process
- What integrity evidence binds the payload to its producer and its declared schema version? security
- How does a consumer detect that a payload it already merged has since been superseded or retracted? state
Classifiers Filled
- Family
- World Models
- Category
- Information and virtual systems
- Entry kind
- pattern
- Navigation path
- NAV.INF.DAT.LIN
- Domain
- INF.DAT.LIN
- Industry
- Cross-industry
- Tags
- datalineageinf.dat.lin
What it is Filled
WM-DAT-006 models the lineage assertion itself: the identity of the nodes it connects (dataset, distribution, field, activity, run, agent), the derivation edge and its transformation characterisation, the time and version bindings that make the assertion interpretable, the evidence and capture method that make it falsifiable, and the governance of the lineage record. Per the known-relation ledger it is a pattern composed into WM-DAT-001 (Dataset) and WM-DAT-005 (Pipeline); it carries references and bindings to datasets, jobs and runs but never reproduces their definition, schema, orchestration, execution or run lifecycle. It is storage- and interface-neutral: property graphs, event streams, catalog tables, RDF, JSON documents and narrative statements are projections of the same semantics.
In scope
- Identity and resolution of lineage nodes: dataset, distribution, field or column, processing activity, run reference and responsible agent
- Derivation edges between nodes, including direct and indirect field-level dependency and characterisation of the transformation applied
- Granularity declaration (dataset, field, partition or collection level) and the declared perimeter separating observed from opaque segments
- Time semantics: event time of derivation versus observation or ingestion time of the assertion, nominal time, and the validity window of an assertion
- Version and snapshot pinning of referenced dataset states, and continuity of lineage across schema change
- Capture method, producer, confidence grading and supporting evidence for each assertion, including provenance of the lineage record itself
- Structural validation of the lineage instance and measurement of lineage coverage and reconciliation against declared sources
- Sensitivity classification, disclosure scoping requirements, retention and disposition of lineage records
- Alignment crosswalks and exchange payload contracts for interchange with PROV, OpenLineage and DCAT consumers
Out of scope
- Dataset definition, schema authoring, distribution publication and catalog lifecycle, which are owned by WM-DAT-001
- Pipeline orchestration, scheduling, task dependency resolution, retry and failure handling, and run state transitions, which are owned by WM-DAT-005
- Execution of processing itself; this model records assertions about processing, it does not run or trigger it
- Audit-trail and log-record semantics, including tamper-evidence and log retention for regulatory logging obligations, which belong to a logging or audit-record model
- Access-control policy evaluation and enforcement; this model declares required scoping and defers evaluation to the access model
- Data quality rule definition, profiling and fitness scoring of the data content itself, which belongs to a data quality sibling
- Software build provenance and artifact attestation for code and container images
- Master data matching, entity resolution of business records, and semantic mapping of business terms
- Lineage visualisation, graph query languages and storage engines, which are interface and storage projections
Why it exists Filled
Provide a format-neutral, composable pattern for asserting, evidencing and governing how datasets and their fields were derived from other datasets and fields through processing activities, so that an agent can reconstruct derivation paths, perform impact analysis and evidence provenance obligations.
Distinguishing features Filled
- Asserts how datasets and fields derive from others through processing; it records but never runs processing.
- Each edge carries its capture method, producer and confidence, unlike a static diagram.
- Unlike an audit log, it describes derivation, not who accessed what.
- Unlike build provenance, its subject is data rather than software artefacts.
What robots and AI may and may not do Filled
Must not
- Assume lineage that was not observed or declared.
- Trigger or approve processing from lineage records.
- Cache referenced schemas as authoritative copies.
- Expose lineage that reveals classified field names or expressions to unauthorised readers.
- Delete superseded assertions instead of retracting them.
Only with a human decision
- Accepting impact analysis results as the basis for a breaking change.
- Approving retraction of lineage used in a regulatory submission.
May
- Assert lineage edges with capture method and confidence.
- Traverse lineage upstream or downstream for impact analysis.
- Measure lineage coverage and reconcile with observed runs.
- Export lineage in an exchange format.
Moral aspects Filled
- Lineage supports the right to know how data about a person was derived.
- Lineage records can themselves expose sensitive structure and must be classified.
Who is affected
- Data subjects
- Data owners and stewards
- Auditors and regulators
Owners Filled
Steward
The adopting Dimension must name a data product owner or data steward accountable for lineage accuracy per subject, distinct from the owner of the underlying data, and record the escalation path when an attestation is rejected.
Roles
- Data product owner or data steward
- Hold accountability for the accuracy of the lineage recorded for the subject; Approve the perimeter and granularity declaration and its revisions; Sign or reject the periodic lineage accuracy attestation and accept residual gaps with a stated remediation horizon; Authorise the disposition mode when retention expires or an erasure trigger applies
- Lineage producer operator
- Register the producer and maintain its immutable schema version references; Emit assertions with event time, observation time, capture method and confidence grade populated; Declare the known blind spots of the capture method and report degradation in capture coverage; Re-emit corrected assertions as supersessions rather than editing prior emissions
- Lineage registrar
- Allocate and record namespaces and extension prefixes and enforce their immutability once used; Maintain the alignment crosswalk register and the evidence supporting each conformance claim; Adjudicate producer precedence when producers assert conflicting lineage for the same endpoints; Operate the promotion and deprecation path for local extension terms
- Lineage assurance reviewer
- Run instance validation against the declared constraint profile and publish the validation report; Measure coverage per granularity level and reconcile asserted paths against independent evidence; Maintain the open gap register and report unresolved references and orphan nodes; Recommend blocking or waiver decisions on failing validations to the accountable role
- Privacy and disclosure controller
- Classify lineage records and identify disclosive elements independently of the classification of the described data; Specify redaction and generalisation requirements for lower-trust projections; Define named, time-bounded disclosure exceptions and their triggering conditions; Refer every access request to the access model for evaluation and enforcement rather than deciding it here
Links to other meta-models Filled
child
- WM-XCT-012 - WM-DAT-006 is registered as a child of the cross-cutting traceability parent and inherits its generic provenance vocabulary, agent typing and evidence grading conventions. This model specialises them for dataset and field derivation only; generic identity, authority and conflict-resolution machinery is not duplicated here.
composes
- WM-DAT-001 - Per the known-relation ledger, WM-DAT-001 composes this lineage pattern so that a dataset carries its lineage context. In that direction WM-DAT-006 supplies derivation assertions keyed on dataset and field references plus version pins; it does not supply, override or restate dataset definition, schema, distribution or catalog lifecycle, which remain owned by WM-DAT-001.
- WM-DAT-005 - Per the known-relation ledger, WM-DAT-005 composes this pattern so that a pipeline emits lineage. WM-DAT-006 receives job and run references and a read-only carried run outcome used to qualify assertion completeness; run state transitions, scheduling, retries, failure handling and operational monitoring remain owned by WM-DAT-005.
aligned
- W3C PROV-DM and PROV-O (http://www.w3.org/ns/prov#) - Alignment to the provenance conceptual core and its IRI namespace for Entity, Activity, Agent, the derivation relations and the qualified-influence pattern. Recorded as a crosswalk with declared lossiness; no conformance is claimed without evidence recorded in the alignment crosswalk artifact.
- OpenLineage specification (Job, Run, Dataset, facets) - Alignment to the runtime lineage exchange model for dataset and job naming, run identifiers, column lineage and facet extension rules. This model adopts the facet extension discipline and naming conventions as bindings; it does not adopt OpenLineage run-state semantics as its own lifecycle.
- W3C DCAT 3 catalog vocabulary - Alignment for interchange of dataset, distribution, dataset series, versioning, qualified relation, provenance statement and checksum terms with catalog consumers. Catalog records, publication and discovery remain outside this model.
references
- Logging and audit-record model (registry identifier unresolved) - Lineage assertions may be cited by regulatory logging and audit records, and disposition actions may need to be evidenced there. This model carries the reference and the binding parameters only; log capture, tamper-evidence, log retention and audit-trail semantics are owned entirely by the target and are listed in out_of_scope.
- Access control and data classification model (registry identifier unresolved) - This model declares the sensitivity classification of a lineage record, the disclosive elements and the required redaction and scoping parameters. Evaluation of an access request, enforcement of the decision and any enforcement audit trail are owned by the target model.
neighbor
- WM-DAT-001 Dataset - WM-DAT-001 owns what a dataset is: its definition, schema, distributions, classification and catalog lifecycle. WM-DAT-006 carries only a resolvable reference plus a version or snapshot pin to a dataset state, and asserts derivation between such references. A schema description reproduced inside a lineage record is a cached binding, not an authoritative schema.
- WM-DAT-005 Pipeline - WM-DAT-005 owns job definition, scheduling, orchestration, run execution and run state transitions such as START, RUNNING, COMPLETE, ABORT and FAIL. WM-DAT-006 records the run identifier and job reference under which a derivation was observed, plus the terminal outcome as a carried attribute for interpreting completeness; it does not model, drive or validate the run lifecycle.
- Logging and audit-record model - Regulatory logging obligations require automatic recording of events over the lifetime of a system. Lineage may be an input to such records but does not own log capture, tamper-evidence, log retention or audit-trail semantics; a reference to an audit record never transfers those semantics to this model.
- Software build provenance (SLSA) - SLSA provenance attests how a software artifact was built, covering build definition, resolved dependencies and builder identity. It does not describe dataset or field derivation. WM-DAT-006 may reference a build attestation as evidence for the code version of a processing activity, but the two provenance graphs are distinct and must not be merged into one node space.
- Data quality management - Quality of the data content (accuracy, completeness of values, conformance to rules) is a sibling concern. WM-DAT-006 owns only quality of the lineage assertion: capture method, confidence, coverage and reconciliation status. Quality metric facets attached to a dataset in an exchange payload are carried references to the quality model.
- Access control and data classification model - WM-DAT-006 declares that a lineage record has a sensitivity classification and which disclosure scopes apply, because lineage exposes schema names, business logic and personal-data footprints. Evaluation and enforcement of the resulting access decision, and any enforcement audit trail, remain with the access-control model.
- Catalog and discovery (DCAT) - DCAT owns catalog records, publication and discovery of datasets and dataset series. WM-DAT-006 aligns to DCAT provenance and versioning terms for interchange but does not own catalog entries, publisher metadata or distribution access URLs.
parent
- WM-XCT-012
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- Authoritative master-system identifier issued by the system of record for the referenced object, such as the catalog table identifier for a dataset, the table-format field identifier for a column, or the orchestrator run identifier for a run; this is the first choice whenever it exists.
- Governed global identifier or IRI from a normative registry or namespace, such as a provenance IRI, a catalog dataset IRI, or a namespace and name pair minted under a registered namespace.
- UUID or ULID minted by the adopting Dimension, recorded together with the resolution evidence that binds it to a master-system identifier once one becomes available.
- Never an identifier: a date, a load or emission timestamp, a file path fragment, a display label or a human-readable title. Location-derived names are bindings and must remain resolvable to one of the three identifier classes above.
Direct properties not applicable Not applicable
Not applicable
Institutional or informational subject: no invented physical properties.
Recognition optional Filled
- A lineage edge links input and output nodes through a process with capture method, time and confidence.
- Often confused with a pipeline definition, an audit log, a data catalogue entry and software build provenance.
Capabilities and actions required Filled
- Assert a lineage edge: Record a new derivation assertion between resolved nodes, carrying its relation type, dependency mode, transformation characterisation, time axes, version pins, capture method, producer and confidence grade.
- Resolve a lineage node reference: Resolve an identifier, location-derived name or symlink to a canonical lineage node, applying the identity priority order and recording the resolution evidence.
- Traverse lineage upstream or downstream: Return the set of nodes and edges reachable from a starting node in a given direction, at a given granularity, as of a specified time and version pin, with per-edge confidence and perimeter status.
- Analyse impact of a proposed change: Produce the downstream consumer and field set affected by a proposed schema, semantic or retention change, as decision support for the change process owned elsewhere.
- Validate a lineage instance: Evaluate a lineage instance against its declared constraint profile and schema version, checking uniqueness, typing, event ordering, impossibility and referential resolution, and emit a validation report.
- Measure coverage and reconcile: Compute coverage ratios per granularity level against the declared perimeter, reconcile asserted paths against independent evidence, and record divergences and open gaps.
- Supersede or retract an assertion: Replace an incorrect assertion with a corrected one, or retract it, preserving the superseded record and making the change detectable by prior readers.
- Export a lineage exchange payload: Project a scoped set of assertions into a serialization-neutral exchange payload under a named alignment profile, with producer identity, declared schema version and content digest.
- Apply retention disposition to lineage records: Mark lineage records for deletion, tombstoning or field-level redaction on expiry of their retention period or on a lawful erasure trigger, and hand the derived downstream holder set to the notifying model.
Hazards and failure modes required Filled
- Wrong impact analysis from incomplete lineage.
- False provenance claims in regulatory reporting.
- Disclosure of sensitive logic through lineage details.
Standards and interfaces required Filled
- OpenLineage.
- W3C PROV-O.
- W3C DCAT.
- SLSA provenance for build artefacts.
Context of use required Filled
- The legal anchors used are EU instruments (the AI Act, GDPR and the INSPIRE metadata implementing regulation). Equivalent obligations in other jurisdictions were not enumerated, and the obligation binding register is the intended extension point for them.
- BCBS 239 is a supervisory standard whose scope and implementation timetable are set by national supervisors and which primarily binds globally and domestically systemically important banks. Its use here as evidence for governance, reconciliation and manual-workaround control is a generalisation of supervisory expectation, not a claim of universal applicability.
- The INSPIRE lineage requirement is cited from a retained-EU-law republication by a national authority; the authentic EU text should be substituted before the citation is relied on in an EU regulatory context.
- Time handling assumes RFC 3339 offsets are available from source systems. Systems that record only local civil time, or that change offset across a daylight-saving boundary during a processing window, require an explicit offset resolution rule that this model requires but does not supply.
- The identity priority assumes a system of record exists that issues stable identifiers. In estates dominated by file drops and ad hoc extracts, only the third priority is achievable and the resulting weakness must be recorded on each affected node.
Sources Filled
- PROV-DM: The PROV Data Model - World Wide Web Consortium (W3C)
- PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
- Constraints of the PROV Data Model - World Wide Web Consortium (W3C)
- OpenLineage Object Model - OpenLineage project (LF AI & Data Foundation)
- OpenLineage Naming Conventions - OpenLineage project (LF AI & Data Foundation)
- ColumnLineageDatasetFacet - OpenLineage project (LF AI & Data Foundation)
- OpenLineage Run Cycle - OpenLineage project (LF AI & Data Foundation)
- OpenLineage Facets and custom facet rules - OpenLineage project (LF AI & Data Foundation)
- Data Catalog Vocabulary (DCAT) - Version 3 - World Wide Web Consortium (W3C)
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (AI Act) - European Parliament and Council of the European Union (EUR-Lex)
- Regulation (EU) 2016/679 General Data Protection Regulation - European Parliament and Council of the European Union (EUR-Lex)
- Principles for effective risk data aggregation and risk reporting (BCBS 239) - Basel Committee on Banking Supervision, Bank for International Settlements
- Apache Iceberg Table Specification - Apache Software Foundation
- SLSA Provenance specification v1.1 - Open Source Security Foundation (OpenSSF), Linux Foundation
- Commission Regulation (EC) No 1205/2008 implementing Directive 2007/2/EC as regards metadata, Annex Part B section 6 (Quality and validity) - legislation.gov.uk (The National Archives), republishing Commission Regulation (EC) No 1205/2008
Open questions
- Retrieve ISO 19115-1 LI_Lineage, LI_ProcessStep and LI_Source through an accessible route after the ISO catalogue returned HTTP 403, so the structured process-step account is grounded directly rather than indirectly through a retained-EU-law republication.
- Consult ISO/IEC 11179 data element registration to give field-node identity a registry-theoretic basis instead of relying solely on table-format field IDs and column lineage facet naming.
- Consult MLCommons Croissant and the SPDX AI and dataset profiles to determine whether machine-learning feature, training-run and model-artifact lineage belongs in this model or in a sibling.
- Find cross-system normative grounding for row-level and record-level lineage; the only retrieved evidence is Iceberg-specific (_row_id, _last_updated_sequence_number) and does not generalise.
- Model streaming and continuous-processing lineage properly: windowing, watermark and reprocessing semantics currently reduce to an observation window plus a snapshot-versus-accumulative pattern code.
- Locate or commission a normative confidence grading scale for inferred lineage, which is currently declared as a required element with no source and left to the adopting Dimension.
- Determine whether a tombstoned lineage record can itself retain personal data through location-derived node names, file paths or field names; no question in retention-and-disposition covers residue in the tombstone, which weakens the erasure resolution.
- Enumerate non-EU obligation equivalents to populate the obligation binding register, since all legal anchors are currently EU instruments and BCBS 239 is used as a generalisation of supervisory expectation.
- Decide whether declaration, registration and attestation functions belong in the function set, given that five artifacts — perimeter declaration, column mapping specification, obligation binding register, namespace registry entry and accuracy attestation — have no producing function.
- Row-level and record-level lineage semantics are not modelled beyond a declarable granularity level; the only primary evidence retrieved is table-format specific and does not generalise.
- Streaming and continuous-processing lineage is handled only through an observation window and a snapshot-versus-accumulative pattern code; windowing, watermark and reprocessing semantics are not modelled and would need a dedicated treatment.
- Machine-learning feature and model lineage (feature stores, training-run to model-artifact derivation, evaluation dataset lineage) is not modelled; MLCommons Croissant and the SPDX AI and dataset profiles were not consulted in this research pass.
- ISO 19115-1 LI_Lineage, LI_ProcessStep and LI_Source could not be retrieved directly because the ISO catalogue returned HTTP 403, and the INSPIRE lineage requirement is cited via a retained-EU-law republication rather than the authentic EUR-Lex text; the structured process-step model those standards define is therefore represented only indirectly.
- ISO/IEC 11179 data element registration was not consulted, so field-node identity is grounded in table-format field IDs and column lineage facet naming rather than in a registry-theoretic account of data element identity.
- Registry identifiers for the logging and audit-record model and for the access control and classification model are unresolved; both composition targets are named descriptively and must be rebound once those registry entries exist.
- Confidence grading for inferred lineage has no normative source in the retrieved material. The grading scale is declared as a required element but its content is left to the adopting Dimension and is marked as a gap rather than presented as canonical.
- Cost, latency and volume propagation along lineage paths, and lineage-driven access propagation, are excluded as they belong to operational and access models respectively.
Machine files
Provenance
world-models research · reviewable-draft
Built from: models/wm-dat-006-data-lineage/spec.yaml, ver-cy/world-models/card-supplements/wm-dat-006-data-lineage.json