Dataset
Describe a dataset as a catalogued, identifiable data asset so that an agent can register it, understand its scope and structure, serve it through distributions and services, judge its quality and fitness, and govern its rights, versions and retirement independently of storage format or interface.
Bundle → Layer → Finding → Questions Filled
7 bundles · 15 layers · 27 findings · 117 questions
Identity and registration What this dataset is, how it is named and identified, how its catalogue entry is registered and how it is classified for discovery.
Designation and identity
Identifiers, titles, descriptions and languages that make the dataset referable and comprehensible.
Dataset identity and designation
The identifier set and human designation of the dataset: the authoritative internal identifier held by the master system, any governed persistent identifier such as a DOI or resolvable IRI, alternate identifiers, and multilingual title/description that name the asset for humans.
- Which system is the master for this dataset's identifier, and what is the identifier it assigns? identity
- Does a governed persistent identifier (DOI, Handle, ARK, resolvable IRI) exist, who registered it, and what does it resolve to? authority
- Does the identifier denote the abstract dataset or one specific version of it? definition
- Which alternate or legacy identifiers must be kept resolvable to avoid breaking existing references? interoperability
- What title, description and language variants are authoritative for this dataset? definition
Registration and classification
The catalogue record that lists the dataset and the controlled terms that make it findable.
Catalogue record and listing state
The registration facet: which catalogue lists the dataset, when the record was created and modified, the record's own state (draft, listed, hidden, deleted) and the fact that record metadata changes independently of the dataset.
- In which catalogue or register is this dataset listed, and under what record identifier? identity
- What is the current listing state of the record, and who may change it? state
- How do record creation and modification times differ from the dataset's own issued and modified times? temporal
- Is the same dataset listed in more than one catalogue, and which listing is authoritative? relationship
Thematic classification and keywords
Controlled-vocabulary themes, keywords, resource type and any mandated classification codes that place the dataset in a discovery taxonomy.
- Which controlled vocabularies are mandatory for themes in the governing profile, and which terms apply? classification
- What general resource type best describes this asset, and does the target profile constrain the allowed values? classification
- Which free-text keywords are maintained, in which languages, and who curates them? quality
- Is the dataset designated under a regulatory category such as an EU high-value dataset thematic category? requirement
Scope and structure What the dataset covers - population, time, space - and what shape it has, including its schema declaration and its membership in a series.
Coverage
The observational, temporal and spatial extent and granularity of the data.
Observation unit and granularity
The population or universe covered, the unit of observation or analysis, the variables measured, and the resolution at which records are expressed.
- What population or universe does the dataset claim to cover, and is coverage a census, a sample or opportunistic? definition
- What is the unit of observation of one record, and does it differ from the unit of analysis? composition
- Which variables are measured, and by what measurement technique? measurement
- What spatial and temporal resolution do the records carry, and what is the smallest reportable granularity? constraint
Temporal coverage and currency
The period the data describes, the reference date of the content, and the currency of the asset expressed through issued, modified and next-expected-update times.
- What start and end instants does the data content cover, and is the interval open-ended? temporal
- What is the reference date of the content, as distinct from the date the metadata was recorded? provenance
- When was the dataset last updated and when is the next update expected? state
- Which time zone or offset applies to the timestamps inside the data, and is it recorded explicitly? constraint
Spatial coverage and reference system
Geographic extent, named places, coordinate reference system and spatial resolution, applicable when the data has a location dimension.
- What geographic extent does the dataset cover, expressed as named places, codes or a bounding geometry? spatial
- Which coordinate reference system do the coordinates use, and is it stated by registry identifier? interoperability
- What is the spatial resolution or smallest geographic unit released, and does disclosure control coarsen it? constraint
- If the dataset has no spatial dimension, is that recorded explicitly rather than left empty? exception
Structure and membership
The declared schema the data claims to follow and the dataset's place in a series or aggregate.
Declared schema reference
The reference from the dataset to the schema, data dictionary or data contract it conforms to, with the version of that schema and whether conformance is asserted or verified.
- Which schema, profile or data contract does the dataset declare conformance to, and at which version? composition
- Is conformance asserted by the publisher or verified by a validator, and where is the evidence? evidence
- Where is the field-level data dictionary published, and does it cover code lists and missing-value conventions? interoperability
- How is divergence between the declared schema and the served data detected and reported? validation
Series and aggregate membership
Whether the dataset is an edition or slice of a continuing series, part of a larger aggregate, or a standalone asset, and the ordering relations among members.
- Is this dataset a member of a dataset series, and what identifies the series? relationship
- Which member precedes and follows this one, and on what ordering property? composition
- Is a new release a new series member or a new version of the same dataset, and which rule decides? decision
- Does the dataset aggregate identifiable parts or subsets that are themselves catalogued? composition
Representation and access How the dataset is materialised as distributions and served through endpoints, and under what access conditions.
Distributions and packaging
Concrete serialisations of the dataset with format, size, fixity and packaging conventions.
Distribution manifest and fixity
Each accessible form of the dataset - its media type and format, byte size, download and access URLs, compression or packaging profile, and the checksum that lets a consumer verify the bytes received.
- Which distributions exist for this dataset, and what media type and format does each use? composition
- What checksum algorithm and value is published for each distribution, and when was fixity last verified? validation
- Is the distribution packaged in a container profile such as a data package or research object crate, and where is its descriptor? interoperability
- Are the distributions semantically equivalent renderings of the same content, or do some carry reduced or transformed content? constraint
- What byte size and record volume should a consumer plan for, and how volatile is that size? measurement
Access channels and conditions
Services that serve the data and the conditions under which access is granted.
Access service and protocol
Endpoints that serve the dataset - APIs, query services, streams or bulk download points - with their protocol, declared conformance class, documentation and operational limits.
- Which services serve this dataset, and what is the endpoint URL and description of each? composition
- What protocol or API standard does each endpoint implement, and which conformance classes does it declare? interoperability
- What rate limits, pagination, quota or payload limits apply to programmatic access? constraint
- Is a bulk download path available alongside the API, as required for designated public-sector datasets? requirement
Access conditions and embargo
The access level of the asset, authentication and authorisation requirements, request and approval processes for restricted data, and time-bound embargoes.
- What is the declared access level of this dataset, using the governing profile's vocabulary? access
- What authentication and authorisation does each channel require, and who can approve a grant? security
- If access is restricted, what is the request process, expected decision time and appeal route? process
- Is an embargo in force, when does it lift, and what becomes available at that moment? temporal
- When the data itself is restricted, does descriptive metadata remain publicly visible? exception
Rights and accountability Who is accountable for the dataset, under what licence and restrictions it may be used, and what protections apply to personal or sensitive content.
Rights and licensing
Licence, rights holder and machine-readable use restrictions.
Licence and rights statement
The licence under which each distribution is offered, expressed by a governed licence identifier, together with the rights holder, attribution requirement and any separate rights statement for metadata versus data.
- Which licence applies to each distribution, identified by a governed licence identifier rather than free text? authority
- Who holds the rights being licensed, and on what basis do they hold them? ownership
- What attribution text or citation must a re-user reproduce? requirement
- Does the metadata carry a different licence from the data itself? constraint
- Does a legal obligation constrain the licence choice, such as an open licence mandate for designated datasets? requirement
Use restrictions and machine-readable policy
Permissions, prohibitions and duties that a licence alone does not express: third-party rights, contractual limits, purpose limitations, redistribution and text-and-data-mining conditions.
- Which actions are permitted, prohibited or subject to a duty, and for which assignee parties? constraint
- Do third-party rights encumber part of the content, and which part? ownership
- Are there purpose limitations, such as research-only use or a prohibition on training models? privacy
- Is the policy expressed in a machine-readable form an agent can evaluate before use? interoperability
Accountability and protection
Named roles accountable for the asset and the controls applied to personal or otherwise sensitive content.
Holder, publisher and contact roles
The accountable data holder, the publisher that lists the dataset, creators or contributors, and the contact point through which questions and access requests are handled.
- Which party is accountable for the content and service level of this dataset? ownership
- How do the holder, publisher, creator and catalogue operator differ for this asset, and are any the same party? relationship
- What contact point is published, and is it a durable role address rather than an individual? access
- Who is the operational steward responsible for updates, quality checks and metadata upkeep? authority
Personal and sensitive data controls
Whether the dataset contains personal, confidential or otherwise sensitive content, what legal basis and protective measures apply, and what disclosure control has been performed.
- Does the dataset contain personal or otherwise sensitive data, and which components carry it? privacy
- On what legal basis or consent is the sensitive content processed and published? authority
- What de-identification or statistical disclosure control has been applied, and with what residual risk? process
- Could this dataset be combined with others to re-identify individuals, and is that risk documented? security
- If special-category data is retained solely for bias detection, are the required safeguards documented? requirement
Lifecycle and provenance Where the dataset came from, how it changes over time, and how it is retired, preserved or deleted.
Provenance
Origin, derivation and production method at asset level.
Source and derivation
The upstream datasets, systems or collection activities this dataset was derived from, the agent responsible, and the generation event that produced this version.
- Which upstream datasets or sources was this dataset derived from, and at which of their versions? provenance
- Which activity generated this version, when did it start and end, and which agent was responsible? event
- To whom is the dataset attributed, as distinct from who published it? ownership
- How far back does asset-level lineage go here before the question is delegated to the lineage model? composition
Production method and processing
How the data was collected or computed: collection mode, instruments, sampling procedure, processing and cleaning steps, and the assumptions those steps embed.
- By what mode and instrument was the data collected or generated? process
- Which preparation operations - cleaning, joining, labelling, aggregation - were applied, and in what order? process
- What assumptions or design choices does the production method embed that a re-user must know? evidence
- Can the production be re-executed, and what would be needed to reproduce this exact output? validation
Versioning and release
Version identity, change records and the cadence and status of releases.
Version identity and change record
How versions are identified and chained, what changed between them, whether a change is breaking, and which version is current.
- What versioning scheme is used, and does a new version get a new persistent identifier? identity
- Which version precedes this one, which supersedes it, and which is currently authoritative? relationship
- What changed relative to the previous version, and is the change breaking for consumers? lifecycle
- Which event triggers a new version rather than an in-place correction? decision
Release cadence and status
The declared update frequency, planned release calendar, punctuality against that plan, and the maturity or availability status of the asset.
- What update frequency is declared, using which controlled vocabulary? temporal
- How does actual release timing compare with the announced calendar? measurement
- What is the current status of the asset - draft, in development, stable, deprecated - and who sets it? state
- Is any part served in real time or near real time, and what latency is guaranteed? requirement
Retirement and retention
Deprecation, withdrawal, tombstoning, retention scheduling and deletion.
Deprecation and withdrawal
How a dataset is marked superseded or removed from service, what replaces it, what consumers are told, and what remains resolvable after removal.
- Is the dataset deprecated or withdrawn, when did that take effect and on whose decision? lifecycle
- What supersedes the withdrawn dataset, and how are consumers redirected? relationship
- What does the tombstone record state, and does it keep the citation and identifier resolvable? evidence
- How is removal signalled to downstream catalogues that previously harvested this record? interoperability
Retention, preservation and deletion
The retention rule governing how long the dataset and its distributions are kept, the preservation commitment, and how deletion is executed and evidenced without destroying the descriptive record.
- Which retention rule applies, what is the retention period and what triggers its start? retention
- What preservation commitment has the holder made, and for which distributions? requirement
- When data is deleted, what is deleted, what metadata is retained, and where is the deletion evidenced? process
- How are conflicts resolved between an erasure obligation and an archival or legal-hold obligation? exception
Quality and fitness Measured quality attached to the asset and the judgement of whether it is fit for a stated purpose, including AI training use.
Quality measurement
Measured quality results attached to the dataset or its distributions.
Quality measurements
Individual measurements of a quality metric computed on this dataset or distribution, each with its value, metric reference, dimension, method and measurement time, plus certificates or user feedback where they exist.
- Which metric was computed, in which quality dimension, and what value resulted? measurement
- What exactly was measured - the whole dataset, one distribution, or a subset? composition
- When was the measurement computed, and against which dataset version? temporal
- What threshold makes the result acceptable, and who set it? constraint
- Do certificates or user feedback annotations exist alongside computed measurements? evidence
Fitness and limitations
Judgements about suitability for stated purposes and documented limitations, including AI-specific readiness.
Fitness for use and known limitations
Statements of intended and discouraged uses, known caveats, coverage gaps, comparability breaks, and who made the fitness judgement for which purpose.
- For which stated purpose has fitness been assessed, and what was the verdict? decision
- What known limitations, caveats or coverage gaps must a re-user account for? quality
- Which uses are explicitly discouraged or unsupported by the holder? constraint
- Are there breaks in series or definitional changes that prevent comparison across time or units? temporal
ML readiness and training use
The durable documentation an ML consumer needs: record structure and field typing, label provenance and annotation process, representativeness and bias examination, and any restriction on using the dataset to train models.
- Is the dataset described in a machine-readable ML format that resolves files to record sets and typed fields? interoperability
- How were labels or annotations produced, by whom, and with what agreement or review? provenance
- What is documented about representativeness relative to the target population and about examined biases? evidence
- Is training a model on this dataset permitted, restricted or prohibited, and by which instrument? authority
- Are canonical splits or benchmark subsets defined at asset level, or left to the consuming training run? composition
Interoperability and exchange How the dataset entry conforms to external profiles, is validated, is exchanged with other catalogues, and is cited and linked.
Profiles and validation
Declared application profiles and the validation evidence for them.
Application profile conformance and validation
Which application profiles the catalogue entry claims to satisfy, the obligations those profiles impose, the validation performed against them, and the mapping losses when projecting into each profile.
- Which application profiles does this entry declare conformance to, at which versions? interoperability
- Which properties are mandatory under each declared profile, and are all of them populated? requirement
- What validation was run, against which rule set, and what were the violations? validation
- What information is lost or coerced when projecting this entry into a given profile? constraint
- Where two declared profiles disagree on a property's obligation or vocabulary, which prevails? exception
Exchange and linkage
Harvesting between catalogues and typed links to other resources.
Harvesting and federation
How the entry is exchanged with other catalogues: source of truth, harvest datestamps and increments, deleted-record signalling, duplicate detection and the rules for locally overriding harvested values.
- Was this entry authored locally or harvested, and from which source catalogue? provenance
- Which datestamp drives incremental harvesting, and at what granularity is it exposed? temporal
- What deleted-record support does the source declare, and how are removals propagated? lifecycle
- May local edits override harvested values, and how is the divergence recorded and reconciled? decision
- How are duplicate entries for the same dataset across federated catalogues detected and merged? identity
Citation and related identifiers
The citation form of the dataset and its typed relations to publications, software, source datasets, successors and other resources.
- What is the recommended citation for this dataset, and which version does it pin? requirement
- Which typed relations to other resources are asserted, using which relation-type vocabulary? relationship
- Are related references qualified enough for a machine to know the direction and nature of each link? interoperability
- Who must be credited in the citation, and how are contributor roles distinguished from creators? ownership
Classifiers Filled
- Family
- World Models
- Category
- Information and virtual systems
- Entry kind
- entity
- Navigation path
- NAV.INF.DAT.SET
- Domain
- INF.DAT.SET
- Industry
- Cross-industry
- Tags
- datasetinf.dat.set
- Also called
- N3
What it is Filled
WM-DAT-001 covers the asset-level semantics of a dataset: identity and designation, catalogue registration, classification, scope and coverage, declared structure, distributions and access services, rights and accountability, provenance, versioning and retirement, quality and fitness evidence, and profile-level interoperability. It treats DCAT 3, DCAT-AP, DCAT-US, schema.org/Dataset, DataCite, Data Package, RO-Crate, Croissant, DQV and PROV-O as alignments, not as the model's own semantics. It stops at the boundary of the data content itself: the meaning of individual variables, the full derivation graph, the full quality assessment apparatus and the commercial offering wrapped around the data belong to sibling models.
In scope
- Dataset identity, persistent identifiers and designation (title, description, language)
- Catalogue record and registration state, distinct from the dataset it describes
- Thematic classification, keywords and controlled-vocabulary terms
- Scope and coverage: observation unit, granularity, temporal and spatial extent
- Declared schema reference, data dictionary pointer and dataset series membership
- Distributions: format, media type, byte size, checksum, packaging profile
- Access services, endpoints, protocols and access conditions including embargo
- Licence, rights statement, use restrictions and machine-readable policy
- Holder, publisher, creator, contact point and steward accountability
- Personal and sensitive data controls at asset level
- Source lineage, production method and processing summary
- Version identity, change record, release cadence and status
- Deprecation, withdrawal, tombstoning, retention and deletion
- Quality measurements, fitness-for-use statements and known limitations
- ML training readiness and dataset documentation for AI use
- Application profile conformance, validation reports, harvesting and citation
Out of scope
- Field-level semantics, datatypes, code lists and contract enforcement (WM-DAT-004 schema / data contract)
- The full lineage graph across processing steps and systems (WM-DAT-006 lineage)
- The quality rule engine, metric definitions and assessment runs (WM-DAT-007 data quality)
- The commercial or consumer-facing offering, SLA and pricing (WM-DAT-008 data product)
- Organization and agent master data for holders, publishers and operators (organization model)
- Identifier scheme governance and allocation rules (identifier / naming model)
- Statistical production semantics such as estimation, weighting and disclosure methodology (official statistics model)
- Model training runs and evaluation results that consume the dataset (WM-AI-006)
- Study design, protocol and ethics approval of the producing research activity (WM-ACT-036)
- Physical storage, indexing and query engines; file systems and databases as such
Why it exists Filled
Describe a dataset as a catalogued, identifiable data asset so that an agent can register it, understand its scope and structure, serve it through distributions and services, judge its quality and fitness, and govern its rights, versions and retirement independently of storage format or interface.
Distinguishing features Filled
- Describes the dataset as a catalogued asset; the catalogue record about it is a separate thing.
- Distributions are concrete files or services of one dataset; field-level schema belongs to the data contract model.
- Licence, rights and steward accountability are part of the asset, not of a storage location.
What robots and AI may and may not do Filled
Must not
- Use or redistribute data against its licence or embargo.
- Combine personal data across datasets in ways the stated purpose does not allow.
- Overwrite a published version instead of issuing a new one.
Only with a human decision
- Publishing a dataset containing personal or sensitive data.
- Retiring a dataset that others depend on.
May
- Register a dataset and its distributions in a catalogue.
- Read data within the stated licence and access conditions.
- Report quality and fitness evidence.
Moral aspects Filled
- Datasets about people need lawful basis, minimization and protection from re-identification.
- Bias in collection carries into every model trained on the data.
- Open licences should be honoured, including attribution.
Who is affected
- Data subjects
- Data creators and publishers
- Users and those affected by derived models
Owners Filled
Steward
A Dimension adopting WM-DAT-001 must name the accountable data holder role and the system of record that mints dataset identifiers before any entry is created.
Roles
- Data holder
- Accountable for content, access decisions, licence choice and service level; Approves withdrawal, deletion and access-level changes; Owns fitness-for-use statements and documented limitations
- Data steward
- Maintains descriptive metadata, classification and coverage statements; Runs profile validation and resolves violations; Monitors release cadence, punctuality and schema drift
- Catalogue operator
- Operates the register, its records and its harvesting endpoints; Enforces profile obligations and duplicate detection across federated sources; Publishes tombstones and propagates deletion signals
- Quality assessor
- Computes and records quality measurements against defined metrics; Issues quality certificates and records user feedback annotations; Escalates threshold failures to the holder and steward
- Privacy and compliance reviewer
- Classifies sensitivity and reviews de-identification and disclosure control; Verifies legal basis, retention rules and erasure or hold conflicts; Reviews training-use permissions and bias documentation for AI reuse
- Consuming agent
- Resolves identifiers, verifies fixity and pins the version it consumed; Evaluates access conditions and machine-readable policy before use; Reports defects and reacts to deprecation and breaking-change signals
Links to other meta-models Filled
references
- WM-DAT-004 (schema / data contract) - The dataset declares conformance to a schema or data contract; the schema's field semantics, datatypes and enforcement stay in WM-DAT-004 so there is one master for the contract.
- WM-DAT-008 (data product) - Inbound: a data product catalog entry references one or more datasets. The offering, its consumers and its service levels are not modelled here.
- WM-AI-006 (training run) - Inbound: a training run references dataset versions it consumed. Split selection and evaluation results belong to the run; only durable ML-readiness documentation is held here.
- WM-ACT-036 (research study) - Inbound: a study produces datasets. The registry records this as PRODUCES; expressed here as a reference because study design, protocol and ethics remain in the study model.
- Organization / agent model (registry identifier to be confirmed) - Holder, publisher, creator, contributor and service operator are agent references; only the role assignment and its effective period are local.
- Identifier and naming scheme model (registry identifier to be confirmed) - Persistent identifier schemes, allocation authority and resolution policy are governed externally; the dataset records only the identifier value and its scheme.
composes
- WM-DAT-006 (lineage) - Asset-level derivation here is a projection of the full lineage graph; WM-DAT-006 supplies activities, agents and qualified derivations beyond the immediate sources.
- WM-DAT-007 (data quality) - Quality measurements attach to the dataset while metric, dimension and threshold definitions and the assessment machinery are owned by WM-DAT-007, following the DQV split between measurement and metric.
- W3C DQV - Quality measurement facet applied to datasets and distributions, kept as a mix-in so quality can be absent without breaking the entry.
- W3C PROV-O - Provenance facet for derivation, attribution and generation time, including the distinction between generation and recording instants.
- W3C ODRL 2.2 - Machine-readable permissions, prohibitions and duties where a licence identifier alone cannot express the terms.
aligned
- W3C DCAT 3 - Primary alignment for Dataset, Distribution, DataService, DatasetSeries and CatalogRecord, and for checksum and version chaining. Alignment, not conformance: profile obligations are validated per deployment.
- DCAT-AP 3.0.0 and DCAT-US - Jurisdictional profiles that add mandatory properties and controlled vocabularies; recorded as declared profiles with validation evidence rather than as built-in obligations.
- DataCite Metadata Schema 4.6 - Citation, persistent identifier, typed related identifiers and tombstone practice for research and repository contexts.
- Frictionless Data Package 2 and RO-Crate 1.2 - Packaging profiles for distributions, providing descriptor conventions, resource-level integrity and payload-plus-metadata containers.
- OAI-PMH 2.0 - Exchange semantics for federated catalogues: selective harvesting by datestamp and declared deleted-record support.
extends
- MLCommons Croissant 1.0 (with RAI extension) - Extends the catalogue entry with ML-loadable record set structure and responsible-AI documentation for training use.
neighbor
- WM-DAT-004 Schema / data contract - This model records only that a dataset declares conformance to a schema or contract, mirroring dcterms:conformsTo and dcat:Distribution/dcterms:conformsTo. The schema's internal structure, field semantics and enforcement live in WM-DAT-004; duplicating them here would create two masters for one contract.
- WM-DAT-006 Lineage - The dataset carries an asset-level derivation summary (prov:wasDerivedFrom, dcterms:source, schema.org isBasedOn) sufficient for catalogue use. The full activity-entity-agent graph, including intermediate entities and qualified derivations, is composed in from WM-DAT-006.
- WM-DAT-007 Data quality - DQV separates the measurement (attached to the dataset) from the metric and dimension definitions. This model holds the attached measurement, certificate or annotation; metric and dimension definitions and the assessment machinery are owned by WM-DAT-007.
- WM-DAT-008 Data product - A data product is an offering with consumers, terms and service levels; a dataset is the catalogued asset. DCAT's DataService is the serving surface referenced here, but product packaging, pricing and consumer contracts stay in WM-DAT-008.
- Catalogue / register (dcat:Catalog and dcat:CatalogRecord) - DCAT 3 distinguishes the resource from the catalogue record about it. This model keeps the registration facet as a finding on the dataset entry rather than modelling the catalogue itself as an entity; a catalogue with its own governance is a separate registry model.
- WM-AI-006 Training run - Training-time selection, splits used and evaluation outcomes belong to the run. The dataset holds only durable ML-readiness documentation such as Croissant RecordSets and RAI statements that remain true independent of any run.
- Organization model - Holder, publisher, creator and contact point are references to agents governed elsewhere; only the role assignment and accountability statement is local to the dataset.
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- Authoritative master-system identifier: the identifier assigned by the system of record that owns the dataset entry always takes precedence.
- Governed global identifier or IRI: a DOI, Handle, ARK or other registered persistent identifier or resolvable IRI, used when no master-system identifier exists or as the public alternate key.
- UUID or ULID assigned by the adopting Dimension: used only when neither of the above exists, and recorded as locally minted.
- Checksums, titles, URLs and version strings are attributes, never primary identifiers; a date or period is never an identifier.
Direct properties not applicable Not applicable
Not applicable
Institutional or informational subject: no invented physical properties.
Recognition optional Filled
- A dataset has a persistent identifier, a title, a publisher, a licence and one or more distributions.
Capabilities and actions required Filled
- Register dataset: Create the dataset entry and its catalogue record, assigning the master-system identifier and, where applicable, requesting a persistent identifier.
- Publish distribution: Attach an accessible form of the dataset with format, size, access URLs and checksum, and expose it through the declared channels.
- Verify fixity: Recompute the checksum of a distribution and compare it with the published value.
- Issue version: Create a new dataset version, chain it to its predecessor and record what changed and whether the change is breaking.
- Record quality measurement: Attach a computed metric result to a stated target and version with its dimension, method and measurement time.
- Validate against profile: Run the rule set of a declared application profile against the catalogue entry and record the violations.
- Evaluate access request: Decide a request for a restricted dataset against its access conditions and machine-readable policy.
- Harvest entry: Fetch or refresh a catalogue entry from a source catalogue using selective harvesting and apply local override rules.
- Withdraw dataset: Remove a dataset from service, publish a tombstone and notify downstream catalogues while keeping identifiers resolvable.
- Execute retention action: Apply the retention rule at its trigger point: extend, archive or delete the payload, and record the evidence.
- Project entry to a target format: Render the neutral dataset entry into a target serialisation or profile, recording any lossy mapping.
- Bind data service: Attach a data service endpoint that serves the dataset and optionally backs a distribution.
- Attach dataset to series: Add a separately published dataset to an explicit DatasetSeries and update first, previous and last navigation.
- Record generation or derivation: Point the dataset at generating activities and source entities, composing WM-DAT-006 for the full graph.
- Cite dataset: Build a citation from mandatory DataCite-style properties and the chosen persistent identifier.
Hazards and failure modes required Filled
- Privacy breach through re-identification.
- Wrong conclusions from outdated or unfit data.
Standards and interfaces required Filled
- W3C DCAT 3 and DCAT-AP.
- schema.org/Dataset.
- DataCite.
- Croissant and Data Package.
Context of use required Filled
- EU-specific: Implementing Regulation (EU) 2023/138 obligations (free of charge, machine-readable, API and bulk download, open licence) apply only to designated high-value datasets held by public sector bodies under Directive (EU) 2019/1024; they are modelled as a designation question, not as a universal rule.
- EU-specific: DCAT-AP mandates EU Publications Office authority lists (data themes, file types, frequency, corporate bodies) that have no force outside the EU context.
- US-specific: accessLevel, bureauCode and programCode come from the US federal profile and are meaningless elsewhere.
- EU-specific: AI Act Article 10 data-governance duties bind high-risk AI system providers, not dataset holders generally; the model records the documentation, not the obligation.
- Indigenous and community data governance (for example CARE-style collective rights) is not represented; jurisdictions applying it need an additional rights facet.
- Statistical quality dimensions are taken from the European Statistical System framework; other national statistical offices use overlapping but differently named dimension sets.
- European exchanges using DCAT-AP 3.0 are assumed to require NAL data-theme for dcat:theme and the DCAT-AP controlled vocabularies for frequency, licences and access rights.
- United States federal open-data catalogues often follow DCAT-US three-tier Catalog/Dataset/Distribution practice; DatasetSeries adoption may lag DCAT 3.
- Research repositories that mint DataCite DOIs treat Creator, Title, Publisher, PublicationYear and ResourceType as mandatory regardless of DCAT optionality.
- Geospatial agencies using ISO 19115-1/19157 and INSPIRE remain more constraint- and lineage-heavy than generic open-data portals.
- Personal-data and statistical-disclosure rules are jurisdiction-specific; this model records flags and policies but does not encode a single privacy statute.
Sources Filled
- Data Catalog Vocabulary (DCAT) - Version 3 - World Wide Web Consortium (W3C)
- Data Quality Vocabulary (DQV) - World Wide Web Consortium (W3C)
- PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
- DCAT Application Profile for data portals in Europe (DCAT-AP) 3.0.0 - European Commission - SEMIC / Publications Office
- Croissant Format Specification 1.0 - MLCommons
- DataCite Metadata Schema 4.6 - DataCite e.V.
- Data Package standard (v2) - Frictionless Data / Open Knowledge Foundation
- DCAT-US Schema (Project Open Data metadata schema) - U.S. General Services Administration - resources.data.gov
- RO-Crate 1.2 Specification - Research Object / RO-Crate community (w3id.org/ro/crate)
- FAIR Principles - GO FAIR International Support and Coordination Office
- schema.org type: Dataset - Schema.org Community Group (W3C)
- RFC 3339: Date and Time on the Internet: Timestamps - Internet Engineering Task Force (IETF)
- Commission Implementing Regulation (EU) 2023/138 laying down a list of specific high-value datasets and the arrangements for their publication and re-use - European Union - EUR-Lex (Official Journal)
- European quality standards / Quality Assurance Framework of the European Statistical System - Eurostat - European Commission
- ODRL Information Model 2.2 - World Wide Web Consortium (W3C)
- The Open Archives Initiative Protocol for Metadata Harvesting, Version 2.0 - Open Archives Initiative
- Tombstone Pages (DataCite Support) - DataCite e.V.
- Article 10: Data and Data Governance, Regulation (EU) 2024/1689 (AI Act) - EU Artificial Intelligence Act explorer (Future of Life Institute)
- DDI-Codebook (DDI-C) - DDI Alliance
- Data Catalog Vocabulary (DCAT) - Version 3 - World Wide Web Consortium
- Data on the Web Best Practices: Data Quality Vocabulary - World Wide Web Consortium
- DataCite Metadata Schema Documentation for the Publication and Citation of Research Data and Other Research Outputs - DataCite e.V.
- PROV-O: The PROV Ontology - World Wide Web Consortium
- ISO 19115-1:2014 Geographic information — Metadata — Part 1: Fundamentals - International Organization for Standardization
- ISO 19115-1 Metadata Best Practice Guide - Intergovernmental Committee on Surveying and Mapping (Australia / New Zealand)
Open questions
- ISO 19115-1 constraint, extent and lineage packages and ISO 19157 quality measures: both providers were blocked by the paywall, so geospatial alignment currently rests on public summaries and tier-3 guidance and must be re-derived from the primary texts.
- A primary security standard (ISO/IEC 27001-family or equivalent) covering security classification of asset bytes versus catalogue metadata, endpoint hardening, key management and transport security, to close the base's declared security gap.
- OAIS / ISO 14721 information-package semantics (SIP, AIP, DIP and representation information); both providers list preservation packaging as unmodelled, and the base currently represents preservation only as a commitment statement.
- Dynamic and continuously updated dataset citation: snapshot versus time-stamped query versus a living work that is not citable as a fixed entity, including event-stream semantics such as watermarks and late-arriving data.
- Landing-page provenance as a discovery property — distinguishing the original provider's landing page from aggregator and catalogue pages — which is absent from the base and materially affects harvesting integrity.
- Catalogue and register as its own model: dcat:Catalog governance, membership, nested catalogues and the empty-catalogue signal, which the accepted boundary excludes from this entry.
- Identifier reuse after retirement (prohibition and successor-identifier resolution) and retention of prior checksums as evidence of what was removed at withdrawal.
- Funding and award obligations recorded as DataCite FundingReference — deposit, embargo and licence conditions — and their explicit separation from holdership.
- Domain-specific dataset profiles not enumerated by either provider: SDMX, STAC, DDI variable-level description, clinical and earth-observation profiles, and an Indigenous or CARE-style collective rights facet.
- Field-level and variable-level semantics: only a data dictionary pointer and record-set reference are held here; DDI variable descriptions and schema internals are delegated.
- Domain-specific dataset profiles (geospatial ISO 19115/INSPIRE, statistical SDMX, clinical, earth observation) are not enumerated; they would enter as declared profiles.
- ISO/IEC 25012, ISO 8000, ISO/IEC 11179 and ISO 14721 were not directly consultable (paywalled or blocked, HTTP 403), so quality dimensions lean on the ESS framework and DQV instead; a Dimension using ISO vocabularies must re-verify the mapping.
- Streaming and continuously updating datasets are handled only through cadence, real-time flags and resolution; event-stream semantics such as watermarks and late-arriving data are not modelled.
- Synthetic data provenance, differential-privacy budgets and data clean-room arrangements are not covered.
- Cost of access, quotas as commercial constructs and SLA penalties are excluded by boundary.
- Dataset discovery ranking, usage metrics and download telemetry are not modelled; DQV user feedback is referenced but usage analytics are out of scope.
- ISO 19115-1 full element dictionaries and ISO 19157 quality measures were not read from the ISO paywalled texts; geospatial constraint and lineage packages are aligned from public summaries and ICSM guidance.
- FAIR Guiding Principles (Wilkinson et al., Scientific Data 2016) and GDPR Article 17 erasure were not retrieved as live primary pages in this run; persistence versus erasure is modelled from DataCite withdrawal, DCAT status and PROV invalidation only.
- DDI Codebook, SDMX dataflow, Frictionless Data Package, STAC, Croissant ML metadata and Datasheets for Datasets are emerging or domain profiles not modelled as core.
- OAIS SIP/AIP/DIP packaging, CKAN extras and physical-media distributions are not modelled.
- Agent and organisation master data, identifier-scheme registries and official-statistics methodology are sibling concerns without model IDs in the supplied registry extract.
- Continuously updated databases versus snapshot datasets are only covered through the dynamic-dataset versioning question.
- Confidential-computing clean rooms and dataset pricing belong elsewhere.
Machine files
Provenance
world-models research · reviewable-draft
Built from: models/wm-dat-001-dataset/spec.yaml, ver-cy/world-models/card-supplements/wm-dat-001-dataset.json