World Models · public research draft

Dataset

Describe a dataset as a catalogued, identifiable data asset so that an agent can register it, understand its scope and structure, serve it through distributions and services, judge its quality and fitness, and govern its rights, versions and retirement independently of storage format or interface.

AI YAMLAGENTS.mdResearch evidence
Research draft. The Claude + Grok synthesis is public for review and use with caution. It passed structural validation but is not yet a canonical Vercy release because the source and coverage holds below remain open.
Catalogue IDWM-DAT-001
Version0.3.0-research.1
Previous version-
Typeentity
ValidationPassed
Synthesis digestsha256:fb689c9ea2b3a7fa…
25Sources
7Bundles
15Layers
27Findings
117Questions
27Artifacts
Format-independent logical structure

Bundles → Layers → Findings → Questions + Artifacts

identity-and-registrationIdentity and registration2 layers

What this dataset is, how it is named and identified, how its catalogue entry is registered and how it is classified for discovery.

designation-and-identityDesignation and identity1 findings

Identifiers, titles, descriptions and languages that make the dataset referable and comprehensible.

dataset-identity-and-designation

Dataset identity and designation

The identifier set and human designation of the dataset: the authoritative internal identifier held by the master system, any governed persistent identifier such as a DOI or resolvable IRI, alternate identifiers, and multilingual title/description that name the asset for humans.

Questions
  1. Which system is the master for this dataset's identifier, and what is the identifier it assigns?identity
    Expected answer
    • Master system reference
    • Master-system dataset identifier
    • Identifier scheme reference
  2. Does a governed persistent identifier (DOI, Handle, ARK, resolvable IRI) exist, who registered it, and what does it resolve to?authority
    Expected answer
    • Persistent identifier value
    • Registration agency
    • Resolution target URL
    • Registration timestamp
  3. Does the identifier denote the abstract dataset or one specific version of it?definition
    Expected answer
    • Identifier granularity code (abstract | version | distribution)
    • Concept-identifier reference
  4. Which alternate or legacy identifiers must be kept resolvable to avoid breaking existing references?interoperability
    Expected answer
    • Alternate identifier list
    • Scheme per alternate identifier
    • Deprecation status per identifier
  5. What title, description and language variants are authoritative for this dataset?definition
    Expected answer
    • Title per language tag
    • Description per language tag
    • Primary language code
Artifacts
  • Dataset metadata recordThe canonical descriptive record for the dataset, serialisable as DCAT/DCAT-AP RDF, DataCite XML/JSON, schema.org JSON-LD or Croissant JSON-LD without changing its semantics.
registration-and-classificationRegistration and classification2 findings

The catalogue record that lists the dataset and the controlled terms that make it findable.

catalogue-record-and-listing

Catalogue record and listing state

The registration facet: which catalogue lists the dataset, when the record was created and modified, the record's own state (draft, listed, hidden, deleted) and the fact that record metadata changes independently of the dataset.

Questions
  1. In which catalogue or register is this dataset listed, and under what record identifier?identity
    Expected answer
    • Catalogue reference
    • Catalogue record identifier
    • Listing URL
  2. What is the current listing state of the record, and who may change it?state
    Expected answer
    • Record state code
    • State changed timestamp
    • Authorised role
  3. How do record creation and modification times differ from the dataset's own issued and modified times?temporal
    Expected answer
    • Record created timestamp
    • Record modified timestamp
    • Dataset issued timestamp
    • Dataset modified timestamp
  4. Is the same dataset listed in more than one catalogue, and which listing is authoritative?relationship
    Expected answer
    • Listing reference list
    • Authoritative listing flag
    • Duplicate-of reference
Artifacts
  • Catalogue recordThe registration object about the dataset (dcat:CatalogRecord), separable from the dataset metadata record and independently timestamped for harvest.
thematic-classification

Thematic classification and keywords

Controlled-vocabulary themes, keywords, resource type and any mandated classification codes that place the dataset in a discovery taxonomy.

Questions
  1. Which controlled vocabularies are mandatory for themes in the governing profile, and which terms apply?classification
    Expected answer
    • Theme vocabulary URI
    • Theme term URIs
    • Profile reference
  2. What general resource type best describes this asset, and does the target profile constrain the allowed values?classification
    Expected answer
    • Resource type code
    • Type vocabulary reference
  3. Which free-text keywords are maintained, in which languages, and who curates them?quality
    Expected answer
    • Keyword list per language
    • Curation owner
    • Last curated timestamp
  4. Is the dataset designated under a regulatory category such as an EU high-value dataset thematic category?requirement
    Expected answer
    • Regulatory designation code
    • Legal instrument reference
    • Designation effective date
Artifacts
  • Classification term setThe resolved set of vocabulary terms bound to the dataset, each with vocabulary URI and version so drift in the source vocabulary is detectable.
scope-and-structureScope and structure2 layers

What the dataset covers - population, time, space - and what shape it has, including its schema declaration and its membership in a series.

coverageCoverage3 findings

The observational, temporal and spatial extent and granularity of the data.

observation-unit-and-granularity

Observation unit and granularity

The population or universe covered, the unit of observation or analysis, the variables measured, and the resolution at which records are expressed.

Questions
  1. What population or universe does the dataset claim to cover, and is coverage a census, a sample or opportunistic?definition
    Expected answer
    • Universe statement
    • Coverage mode code
    • Sampling procedure reference
  2. What is the unit of observation of one record, and does it differ from the unit of analysis?composition
    Expected answer
    • Unit of observation term
    • Unit of analysis term
    • Record count
  3. Which variables are measured, and by what measurement technique?measurement
    Expected answer
    • Variable name list
    • Measurement technique statement
    • Unit of measure per variable
  4. What spatial and temporal resolution do the records carry, and what is the smallest reportable granularity?constraint
    Expected answer
    • Spatial resolution value
    • Temporal resolution duration
    • Minimum reportable granularity
Artifacts
  • Coverage statementA structured statement of universe, unit, variables and resolution, suitable for publication as a codebook section or as Croissant RecordSet documentation.
temporal-coverage-and-currency

Temporal coverage and currency

The period the data describes, the reference date of the content, and the currency of the asset expressed through issued, modified and next-expected-update times.

Questions
  1. What start and end instants does the data content cover, and is the interval open-ended?temporal
    Expected answer
    • Coverage start timestamp
    • Coverage end timestamp
    • Open-ended flag
  2. What is the reference date of the content, as distinct from the date the metadata was recorded?provenance
    Expected answer
    • Content reference date
    • Metadata record timestamp
    • Ingestion timestamp
  3. When was the dataset last updated and when is the next update expected?state
    Expected answer
    • Last modified timestamp
    • Next expected update timestamp
    • Currency status code
  4. Which time zone or offset applies to the timestamps inside the data, and is it recorded explicitly?constraint
    Expected answer
    • Offset or time-zone identifier
    • Explicit-offset compliance flag
Artifacts
  • Temporal coverage recordMachine-readable interval and currency block covering content period, issue, modification and expected update.
spatial-coverage-and-reference-system

Spatial coverage and reference system

Geographic extent, named places, coordinate reference system and spatial resolution, applicable when the data has a location dimension.

Questions
  1. What geographic extent does the dataset cover, expressed as named places, codes or a bounding geometry?spatial
    Expected answer
    • Place term URIs
    • Bounding geometry
    • Geometry encoding
  2. Which coordinate reference system do the coordinates use, and is it stated by registry identifier?interoperability
    Expected answer
    • CRS registry identifier
    • CRS registry reference
    • Axis order statement
  3. What is the spatial resolution or smallest geographic unit released, and does disclosure control coarsen it?constraint
    Expected answer
    • Spatial resolution in metres
    • Smallest geographic unit
    • Coarsening rule reference
  4. If the dataset has no spatial dimension, is that recorded explicitly rather than left empty?exception
    Expected answer
    • Not-applicable flag
    • Justification note
Artifacts
  • Spatial extent recordGeometry or place-code block with an explicit CRS reference, publishable in a geospatial metadata profile.
structure-and-membershipStructure and membership2 findings

The declared schema the data claims to follow and the dataset's place in a series or aggregate.

declared-schema-reference

Declared schema reference

The reference from the dataset to the schema, data dictionary or data contract it conforms to, with the version of that schema and whether conformance is asserted or verified.

Questions
  1. Which schema, profile or data contract does the dataset declare conformance to, and at which version?composition
    Expected answer
    • Schema reference
    • Schema version
    • Declaration scope (dataset | distribution)
  2. Is conformance asserted by the publisher or verified by a validator, and where is the evidence?evidence
    Expected answer
    • Conformance status code
    • Validation report reference
    • Validated timestamp
  3. Where is the field-level data dictionary published, and does it cover code lists and missing-value conventions?interoperability
    Expected answer
    • Data dictionary URL
    • Code list references
    • Missing value convention
  4. How is divergence between the declared schema and the served data detected and reported?validation
    Expected answer
    • Drift check method
    • Last drift check timestamp
    • Open drift findings
Artifacts
  • Schema declarationThe binding between the dataset (or a distribution) and an external schema artefact, holding the reference, version and conformance status only.
series-and-membership

Series and aggregate membership

Whether the dataset is an edition or slice of a continuing series, part of a larger aggregate, or a standalone asset, and the ordering relations among members.

Questions
  1. Is this dataset a member of a dataset series, and what identifies the series?relationship
    Expected answer
    • Series reference
    • Series identifier
    • Membership role
  2. Which member precedes and follows this one, and on what ordering property?composition
    Expected answer
    • Previous member reference
    • Next member reference
    • Ordering property
  3. Is a new release a new series member or a new version of the same dataset, and which rule decides?decision
    Expected answer
    • Release policy rule
    • Decision outcome code
    • Rule owner
  4. Does the dataset aggregate identifiable parts or subsets that are themselves catalogued?composition
    Expected answer
    • Part references
    • Part granularity
    • Aggregation rule
Artifacts
  • Series membership recordEdge record binding a dataset to a series with ordering and role, kept separate so the series can be governed by its own holder.
representation-and-accessRepresentation and access2 layers

How the dataset is materialised as distributions and served through endpoints, and under what access conditions.

distributionsDistributions and packaging1 findings

Concrete serialisations of the dataset with format, size, fixity and packaging conventions.

distribution-manifest-and-fixity

Distribution manifest and fixity

Each accessible form of the dataset - its media type and format, byte size, download and access URLs, compression or packaging profile, and the checksum that lets a consumer verify the bytes received.

Questions
  1. Which distributions exist for this dataset, and what media type and format does each use?composition
    Expected answer
    • Distribution identifier list
    • Media type per distribution
    • Format term per distribution
  2. What checksum algorithm and value is published for each distribution, and when was fixity last verified?validation
    Expected answer
    • Checksum algorithm
    • Checksum value
    • Fixity verified timestamp
  3. Is the distribution packaged in a container profile such as a data package or research object crate, and where is its descriptor?interoperability
    Expected answer
    • Packaging profile reference
    • Descriptor filename
    • Descriptor URL
  4. Are the distributions semantically equivalent renderings of the same content, or do some carry reduced or transformed content?constraint
    Expected answer
    • Equivalence flag
    • Reduction description
    • Authoritative distribution reference
  5. What byte size and record volume should a consumer plan for, and how volatile is that size?measurement
    Expected answer
    • Byte size
    • Record count
    • Size measured timestamp
Artifacts
  • DistributionAn accessible form of the dataset: file, archive, package or dump, described by a manifest entry with format, size, fixity and access URLs.
access-channelsAccess channels and conditions2 findings

Services that serve the data and the conditions under which access is granted.

access-service-and-protocol

Access service and protocol

Endpoints that serve the dataset - APIs, query services, streams or bulk download points - with their protocol, declared conformance class, documentation and operational limits.

Questions
  1. Which services serve this dataset, and what is the endpoint URL and description of each?composition
    Expected answer
    • Service reference list
    • Endpoint URL
    • Endpoint description URL
  2. What protocol or API standard does each endpoint implement, and which conformance classes does it declare?interoperability
    Expected answer
    • Protocol identifier
    • Conformance class list
    • Specification reference
  3. What rate limits, pagination, quota or payload limits apply to programmatic access?constraint
    Expected answer
    • Rate limit value
    • Page size limit
    • Quota policy reference
  4. Is a bulk download path available alongside the API, as required for designated public-sector datasets?requirement
    Expected answer
    • Bulk download URL
    • Obligation reference
    • Availability status
Artifacts
  • Access service descriptorMachine-readable description of the serving endpoint, including protocol, conformance and documentation links; the API definition document itself is referenced, not owned here.
access-conditions-and-embargo

Access conditions and embargo

The access level of the asset, authentication and authorisation requirements, request and approval processes for restricted data, and time-bound embargoes.

Questions
  1. What is the declared access level of this dataset, using the governing profile's vocabulary?access
    Expected answer
    • Access rights code
    • Vocabulary reference
    • Justification note
  2. What authentication and authorisation does each channel require, and who can approve a grant?security
    Expected answer
    • Authentication method
    • Authorisation role
    • Approver reference
  3. If access is restricted, what is the request process, expected decision time and appeal route?process
    Expected answer
    • Request procedure reference
    • Decision service level
    • Appeal route
  4. Is an embargo in force, when does it lift, and what becomes available at that moment?temporal
    Expected answer
    • Embargo end timestamp
    • Post-embargo access level
    • Embargo reason code
  5. When the data itself is restricted, does descriptive metadata remain publicly visible?exception
    Expected answer
    • Metadata visibility flag
    • Suppressed field list
    • Policy reference
Artifacts
  • Access condition statementHuman- and machine-readable statement of access level, conditions and request route, publishable on a landing page even when the data is closed.
rights-and-accountabilityRights and accountability2 layers

Who is accountable for the dataset, under what licence and restrictions it may be used, and what protections apply to personal or sensitive content.

rights-and-licensingRights and licensing2 findings

Licence, rights holder and machine-readable use restrictions.

licence-and-rights-statement

Licence and rights statement

The licence under which each distribution is offered, expressed by a governed licence identifier, together with the rights holder, attribution requirement and any separate rights statement for metadata versus data.

Questions
  1. Which licence applies to each distribution, identified by a governed licence identifier rather than free text?authority
    Expected answer
    • Licence identifier
    • Licence list reference
    • Licence URL
  2. Who holds the rights being licensed, and on what basis do they hold them?ownership
    Expected answer
    • Rights holder reference
    • Basis of rights statement
    • Effective from timestamp
  3. What attribution text or citation must a re-user reproduce?requirement
    Expected answer
    • Attribution statement
    • Required citation string
  4. Does the metadata carry a different licence from the data itself?constraint
    Expected answer
    • Metadata licence identifier
    • Data licence identifier
    • Divergence note
  5. Does a legal obligation constrain the licence choice, such as an open licence mandate for designated datasets?requirement
    Expected answer
    • Legal instrument reference
    • Mandated licence family
    • Compliance status
Artifacts
  • Licence bindingRecord binding a distribution or the dataset to a licence identifier, rights holder and attribution text, with its own effective period so licence changes are auditable.
use-restrictions-and-policy

Use restrictions and machine-readable policy

Permissions, prohibitions and duties that a licence alone does not express: third-party rights, contractual limits, purpose limitations, redistribution and text-and-data-mining conditions.

Questions
  1. Which actions are permitted, prohibited or subject to a duty, and for which assignee parties?constraint
    Expected answer
    • Permitted action list
    • Prohibited action list
    • Duty list
    • Party reference
  2. Do third-party rights encumber part of the content, and which part?ownership
    Expected answer
    • Encumbered component reference
    • Right holder reference
    • Restriction description
  3. Are there purpose limitations, such as research-only use or a prohibition on training models?privacy
    Expected answer
    • Purpose limitation code
    • Scope of limitation
    • Source of limitation
  4. Is the policy expressed in a machine-readable form an agent can evaluate before use?interoperability
    Expected answer
    • Policy document reference
    • Policy language identifier
    • Evaluation result
Artifacts
  • Use policyMachine-readable policy set attaching permissions, prohibitions, duties and constraints to the dataset as an asset.
accountability-and-protectionAccountability and protection2 findings

Named roles accountable for the asset and the controls applied to personal or otherwise sensitive content.

holder-and-role-assignment

Holder, publisher and contact roles

The accountable data holder, the publisher that lists the dataset, creators or contributors, and the contact point through which questions and access requests are handled.

Questions
  1. Which party is accountable for the content and service level of this dataset?ownership
    Expected answer
    • Holder reference
    • Accountability statement
    • Effective from timestamp
  2. How do the holder, publisher, creator and catalogue operator differ for this asset, and are any the same party?relationship
    Expected answer
    • Role to party mapping
    • Same-party flags
    • Role vocabulary reference
  3. What contact point is published, and is it a durable role address rather than an individual?access
    Expected answer
    • Contact channel
    • Role-address flag
    • Response service level
  4. Who is the operational steward responsible for updates, quality checks and metadata upkeep?authority
    Expected answer
    • Steward reference
    • Delegated duties
    • Escalation route
Artifacts
  • Role assignment recordEdge record binding a party reference to a role on the dataset with an effective period, so accountability history survives party changes.
personal-and-sensitive-data-controls

Personal and sensitive data controls

Whether the dataset contains personal, confidential or otherwise sensitive content, what legal basis and protective measures apply, and what disclosure control has been performed.

Questions
  1. Does the dataset contain personal or otherwise sensitive data, and which components carry it?privacy
    Expected answer
    • Sensitivity classification
    • Affected component references
    • Assessment timestamp
  2. On what legal basis or consent is the sensitive content processed and published?authority
    Expected answer
    • Legal basis code
    • Consent record reference
    • Jurisdiction
  3. What de-identification or statistical disclosure control has been applied, and with what residual risk?process
    Expected answer
    • Control method
    • Applied timestamp
    • Residual risk assessment
  4. Could this dataset be combined with others to re-identify individuals, and is that risk documented?security
    Expected answer
    • Linkage risk statement
    • Known linkable datasets
    • Mitigation measure
  5. If special-category data is retained solely for bias detection, are the required safeguards documented?requirement
    Expected answer
    • Safeguard list
    • Necessity justification
    • Deletion commitment
Artifacts
  • Sensitivity assessmentDated assessment recording sensitivity classification, applied controls, residual risk and reviewer, retained even after the dataset is withdrawn.
lifecycle-and-provenanceLifecycle and provenance3 layers

Where the dataset came from, how it changes over time, and how it is retired, preserved or deleted.

provenanceProvenance2 findings

Origin, derivation and production method at asset level.

source-and-derivation

Source and derivation

The upstream datasets, systems or collection activities this dataset was derived from, the agent responsible, and the generation event that produced this version.

Questions
  1. Which upstream datasets or sources was this dataset derived from, and at which of their versions?provenance
    Expected answer
    • Source dataset references
    • Source version identifiers
    • Derivation type
  2. Which activity generated this version, when did it start and end, and which agent was responsible?event
    Expected answer
    • Generating activity reference
    • Activity start timestamp
    • Activity end timestamp
    • Responsible agent reference
  3. To whom is the dataset attributed, as distinct from who published it?ownership
    Expected answer
    • Attributed agent reference
    • Attribution role
    • Publisher reference
  4. How far back does asset-level lineage go here before the question is delegated to the lineage model?composition
    Expected answer
    • Lineage depth policy
    • Delegation boundary reference
    • Lineage graph reference
Artifacts
  • Provenance recordAsset-level provenance bundle linking this dataset version to sources, generating activity and responsible agents, exportable as a PROV graph.
production-method

Production method and processing

How the data was collected or computed: collection mode, instruments, sampling procedure, processing and cleaning steps, and the assumptions those steps embed.

Questions
  1. By what mode and instrument was the data collected or generated?process
    Expected answer
    • Collection mode code
    • Instrument reference
    • Collection period
  2. Which preparation operations - cleaning, joining, labelling, aggregation - were applied, and in what order?process
    Expected answer
    • Operation sequence
    • Operation parameters
    • Executing system reference
  3. What assumptions or design choices does the production method embed that a re-user must know?evidence
    Expected answer
    • Assumption statements
    • Design choice rationale
    • Documented-by reference
  4. Can the production be re-executed, and what would be needed to reproduce this exact output?validation
    Expected answer
    • Reproducibility status
    • Required inputs list
    • Code or workflow reference
Artifacts
  • Method documentationMethodology description covering collection, sampling and processing, published alongside the dataset as a codebook or datasheet section.
versioning-and-releaseVersioning and release2 findings

Version identity, change records and the cadence and status of releases.

version-identity-and-change

Version identity and change record

How versions are identified and chained, what changed between them, whether a change is breaking, and which version is current.

Questions
  1. What versioning scheme is used, and does a new version get a new persistent identifier?identity
    Expected answer
    • Version scheme code
    • Version string
    • New-identifier policy
  2. Which version precedes this one, which supersedes it, and which is currently authoritative?relationship
    Expected answer
    • Previous version reference
    • Next version reference
    • Current version reference
  3. What changed relative to the previous version, and is the change breaking for consumers?lifecycle
    Expected answer
    • Change summary
    • Breaking change flag
    • Affected component list
  4. Which event triggers a new version rather than an in-place correction?decision
    Expected answer
    • Versioning rule
    • Trigger event type
    • Rule owner
Artifacts
  • Change log entryPer-version entry recording what changed, when it was released and whether it breaks consumers; the change log is the append-only series of these entries.
release-cadence-and-status

Release cadence and status

The declared update frequency, planned release calendar, punctuality against that plan, and the maturity or availability status of the asset.

Questions
  1. What update frequency is declared, using which controlled vocabulary?temporal
    Expected answer
    • Frequency code
    • Vocabulary reference
    • Declared from timestamp
  2. How does actual release timing compare with the announced calendar?measurement
    Expected answer
    • Planned release timestamps
    • Actual release timestamps
    • Punctuality delta
  3. What is the current status of the asset - draft, in development, stable, deprecated - and who sets it?state
    Expected answer
    • Status code
    • Status set timestamp
    • Authorising role
  4. Is any part served in real time or near real time, and what latency is guaranteed?requirement
    Expected answer
    • Real-time flag
    • Latency target
    • Obligation reference
Artifacts
  • Release calendarPublished schedule of planned releases with announced and actual instants, enabling punctuality measurement.
retirement-and-retentionRetirement and retention2 findings

Deprecation, withdrawal, tombstoning, retention scheduling and deletion.

deprecation-and-withdrawal

Deprecation and withdrawal

How a dataset is marked superseded or removed from service, what replaces it, what consumers are told, and what remains resolvable after removal.

Questions
  1. Is the dataset deprecated or withdrawn, when did that take effect and on whose decision?lifecycle
    Expected answer
    • Retirement state code
    • Effective timestamp
    • Deciding authority
  2. What supersedes the withdrawn dataset, and how are consumers redirected?relationship
    Expected answer
    • Successor reference
    • Redirect target
    • Migration guidance
  3. What does the tombstone record state, and does it keep the citation and identifier resolvable?evidence
    Expected answer
    • Tombstone URL
    • Unavailability statement
    • Retained citation string
  4. How is removal signalled to downstream catalogues that previously harvested this record?interoperability
    Expected answer
    • Deleted-record support level
    • Deletion datestamp
    • Notified consumers
Artifacts
  • Tombstone recordPersisting record containing the full citation, the identifier in human- and machine-readable form and a statement of unavailability; it survives deletion of the payload.
retention-and-deletion

Retention, preservation and deletion

The retention rule governing how long the dataset and its distributions are kept, the preservation commitment, and how deletion is executed and evidenced without destroying the descriptive record.

Questions
  1. Which retention rule applies, what is the retention period and what triggers its start?retention
    Expected answer
    • Retention rule reference
    • Retention period
    • Trigger event type
  2. What preservation commitment has the holder made, and for which distributions?requirement
    Expected answer
    • Preservation commitment statement
    • Covered distributions
    • Committing party
  3. When data is deleted, what is deleted, what metadata is retained, and where is the deletion evidenced?process
    Expected answer
    • Deletion scope
    • Retained metadata set
    • Deletion evidence reference
    • Deletion timestamp
  4. How are conflicts resolved between an erasure obligation and an archival or legal-hold obligation?exception
    Expected answer
    • Conflicting obligations
    • Resolution decision
    • Approving authority
Artifacts
  • Retention schedule entryBinding of the dataset to a retention rule with computed retain-until instant, hold status and executed-deletion evidence.
quality-and-fitnessQuality and fitness2 layers

Measured quality attached to the asset and the judgement of whether it is fit for a stated purpose, including AI training use.

quality-measurementQuality measurement1 findings

Measured quality results attached to the dataset or its distributions.

quality-measurements

Quality measurements

Individual measurements of a quality metric computed on this dataset or distribution, each with its value, metric reference, dimension, method and measurement time, plus certificates or user feedback where they exist.

Questions
  1. Which metric was computed, in which quality dimension, and what value resulted?measurement
    Expected answer
    • Metric reference
    • Dimension reference
    • Measured value
    • Expected data type
  2. What exactly was measured - the whole dataset, one distribution, or a subset?composition
    Expected answer
    • Computed-on reference
    • Subset definition
    • Sample size
  3. When was the measurement computed, and against which dataset version?temporal
    Expected answer
    • Measured timestamp
    • Dataset version reference
    • Method reference
  4. What threshold makes the result acceptable, and who set it?constraint
    Expected answer
    • Threshold value
    • Pass or fail outcome
    • Threshold owner
  5. Do certificates or user feedback annotations exist alongside computed measurements?evidence
    Expected answer
    • Certificate references
    • Feedback annotation references
    • Issuing party
Artifacts
  • Quality measurement recordImmutable record of one metric computation against a stated target and version, retained as an evidence series.
fitness-and-limitationsFitness and limitations2 findings

Judgements about suitability for stated purposes and documented limitations, including AI-specific readiness.

fitness-for-use

Fitness for use and known limitations

Statements of intended and discouraged uses, known caveats, coverage gaps, comparability breaks, and who made the fitness judgement for which purpose.

Questions
  1. For which stated purpose has fitness been assessed, and what was the verdict?decision
    Expected answer
    • Purpose statement
    • Fitness verdict
    • Assessing party
    • Assessment timestamp
  2. What known limitations, caveats or coverage gaps must a re-user account for?quality
    Expected answer
    • Limitation statements
    • Affected components
    • Severity
  3. Which uses are explicitly discouraged or unsupported by the holder?constraint
    Expected answer
    • Discouraged use list
    • Reason
    • Source of statement
  4. Are there breaks in series or definitional changes that prevent comparison across time or units?temporal
    Expected answer
    • Break points
    • Definitional change description
    • Comparability note
Artifacts
  • Fitness assessmentDated assessment binding a purpose to a verdict with supporting limitations and evidence references; commonly published as a datasheet or quality report section.
ml-readiness

ML readiness and training use

The durable documentation an ML consumer needs: record structure and field typing, label provenance and annotation process, representativeness and bias examination, and any restriction on using the dataset to train models.

Questions
  1. Is the dataset described in a machine-readable ML format that resolves files to record sets and typed fields?interoperability
    Expected answer
    • ML description document reference
    • Format version
    • Record set inventory
  2. How were labels or annotations produced, by whom, and with what agreement or review?provenance
    Expected answer
    • Annotation process description
    • Annotator population
    • Agreement measure
  3. What is documented about representativeness relative to the target population and about examined biases?evidence
    Expected answer
    • Representativeness statement
    • Bias examination result
    • Mitigation measures
  4. Is training a model on this dataset permitted, restricted or prohibited, and by which instrument?authority
    Expected answer
    • Training permission code
    • Instrument reference
    • Scope of permission
  5. Are canonical splits or benchmark subsets defined at asset level, or left to the consuming training run?composition
    Expected answer
    • Split definitions
    • Split ownership
    • Delegation note
Artifacts
  • ML dataset descriptionMachine-readable description resolving distributions to record sets, fields and types, with responsible-AI statements attached; loadable by ML frameworks without bespoke code.
interoperability-and-exchangeInteroperability and exchange2 layers

How the dataset entry conforms to external profiles, is validated, is exchanged with other catalogues, and is cited and linked.

profiles-and-validationProfiles and validation1 findings

Declared application profiles and the validation evidence for them.

profile-conformance-and-validation

Application profile conformance and validation

Which application profiles the catalogue entry claims to satisfy, the obligations those profiles impose, the validation performed against them, and the mapping losses when projecting into each profile.

Questions
  1. Which application profiles does this entry declare conformance to, at which versions?interoperability
    Expected answer
    • Profile identifier list
    • Profile version
    • Declaration timestamp
  2. Which properties are mandatory under each declared profile, and are all of them populated?requirement
    Expected answer
    • Mandatory property list
    • Missing property list
    • Compliance percentage
  3. What validation was run, against which rule set, and what were the violations?validation
    Expected answer
    • Validator identifier
    • Rule set reference
    • Violation list
    • Validated timestamp
  4. What information is lost or coerced when projecting this entry into a given profile?constraint
    Expected answer
    • Lossy field list
    • Coercion rule
    • Reversibility flag
  5. Where two declared profiles disagree on a property's obligation or vocabulary, which prevails?exception
    Expected answer
    • Conflicting profiles
    • Conflicting property
    • Precedence decision
Artifacts
  • Validation reportMachine-readable report of a validation run against a named profile rule set, with per-rule outcomes and timestamps; retained as evidence for conformance claims.
exchange-and-linkageExchange and linkage2 findings

Harvesting between catalogues and typed links to other resources.

harvesting-and-federation

Harvesting and federation

How the entry is exchanged with other catalogues: source of truth, harvest datestamps and increments, deleted-record signalling, duplicate detection and the rules for locally overriding harvested values.

Questions
  1. Was this entry authored locally or harvested, and from which source catalogue?provenance
    Expected answer
    • Origin code
    • Source catalogue reference
    • Source record identifier
  2. Which datestamp drives incremental harvesting, and at what granularity is it exposed?temporal
    Expected answer
    • Harvest datestamp
    • Granularity
    • Last successful harvest timestamp
  3. What deleted-record support does the source declare, and how are removals propagated?lifecycle
    Expected answer
    • Deleted-record support level
    • Deletion signal
    • Propagation lag
  4. May local edits override harvested values, and how is the divergence recorded and reconciled?decision
    Expected answer
    • Override policy
    • Overridden field list
    • Reconciliation rule
  5. How are duplicate entries for the same dataset across federated catalogues detected and merged?identity
    Expected answer
    • Matching keys
    • Duplicate-of reference
    • Merge decision
Artifacts
  • Harvest log entryPer-run record of what was fetched, changed, deleted or rejected during a harvest, supporting replay and divergence analysis.

Publication holds

  • Source verification is incomplete: all 19 base source URLs must be re-fetched live and version-pinned before publication, and the DataCite divergence resolved — the base cites schema 4.6 while the second provider reports 4.7 released 3 March 2026, which affects citation properties and relation types.
  • The base reports that EUR-Lex full text (Implementing Regulation (EU) 2023/138) and ISO catalogue pages could not be fetched (timeout and HTTP 403); every claim resting on those sources is currently limited to indexed titles, dates and dimension names and must be re-verified against full text before publication.
  • The AI Act Article 10 material is sourced from a tier-3 secondary rendering (artificialintelligenceact.eu), not the Official Journal; no obligation-shaped statement derived from it may be published until the primary text is read.
  • Multi-profile validation has not been performed by either provider: the merged entry must be validated against at least DCAT-AP 3.0.0 SHACL, DCAT-US, DataCite and one geospatial profile (ISO 19115-1 / INSPIRE) before publication, since profile conformance is asserted as a declaration, not a verified state.
  • Domain-profile applicability is unverified: the second provider asserts uniform application across open data portals, research repositories and internal registers, but neither provider tested the model against a worked example in each of those three deployment contexts.
  • The security dimension is a declared gap in the base — no primary security standard was consulted — and must be published as a stated gap, not silently presented as covered by access-service attributes.

Deferred research

  • ISO 19115-1 constraint, extent and lineage packages and ISO 19157 quality measures: both providers were blocked by the paywall, so geospatial alignment currently rests on public summaries and tier-3 guidance and must be re-derived from the primary texts.
  • A primary security standard (ISO/IEC 27001-family or equivalent) covering security classification of asset bytes versus catalogue metadata, endpoint hardening, key management and transport security, to close the base's declared security gap.
  • OAIS / ISO 14721 information-package semantics (SIP, AIP, DIP and representation information); both providers list preservation packaging as unmodelled, and the base currently represents preservation only as a commitment statement.
  • Dynamic and continuously updated dataset citation: snapshot versus time-stamped query versus a living work that is not citable as a fixed entity, including event-stream semantics such as watermarks and late-arriving data.
  • Landing-page provenance as a discovery property — distinguishing the original provider's landing page from aggregator and catalogue pages — which is absent from the base and materially affects harvesting integrity.
  • Catalogue and register as its own model: dcat:Catalog governance, membership, nested catalogues and the empty-catalogue signal, which the accepted boundary excludes from this entry.
  • Identifier reuse after retirement (prohibition and successor-identifier resolution) and retention of prior checksums as evidence of what was removed at withdrawal.
  • Funding and award obligations recorded as DataCite FundingReference — deposit, embargo and licence conditions — and their explicit separation from holdership.
  • Domain-specific dataset profiles not enumerated by either provider: SDMX, STAC, DDI variable-level description, clinical and earth-observation profiles, and an Indigenous or CARE-style collective rights facet.