Synthetic Data Product
Describe a governed generated-data product through release-specific lineage, fitness evidence, privacy claims and distribution conditions.
Bundle → Layer → Finding → Questions Filled
3 bundles · 4 layers · 5 findings · 10 questions
Generation How the synthetic data was produced and from what.
Source and generator
The source data and the generation method used.
Source data basis
The real data the generator learned from, with its legal basis.
- Which source dataset and version was the generator trained or fitted on?
- Under which legal basis and agreement was the source data used?
Generation method
The generator, its configuration and any privacy mechanism.
- Which generation method and configuration produced this release?
- Was a formal privacy mechanism applied, and with which parameters?
Evidence What is known about fitness and protection of the release.
Fidelity and utility
How closely the synthetic data matches the source for intended uses.
Fidelity assessment
Comparison of distributions and relationships with the source.
- Which fidelity measures were computed, and against which source version?
- For which analyses or uses was utility tested and found adequate?
Privacy risk
Residual disclosure risk of the release.
Disclosure assessment
Tests for membership inference, attribute disclosure and record copying.
- Which disclosure attacks or tests were run on the release?
- Were any synthetic records found to replicate real records?
Release How the product is offered and labelled.
Publication terms
Labelling, licence and intended use of the release.
Synthetic labelling
Clear marking of the data as synthetic with its limits.
- Is the release clearly labelled as synthetic in its metadata and files?
- Which uses does the release state as unsuitable?
Classifiers Filled
- Family
- World Models
- Category
- Information and virtual systems
- Entry kind
- entity
- Navigation path
- NAV.INF.DAT.SYN
- Domain
- INF.DAT.SYN
- Industry
- Cross-industry
- Tags
- syntheticdataproductinf.dat.syn
What it is Filled
A synthetic data product is a dataset generated by a model or procedure to mimic the statistical properties of real source data, released together with evidence of its fidelity and privacy protection. It is a data product with a generation lineage; real data that has only been masked or de-identified is outside this class.
In scope
- Product identity and synthetic classification with explicit generated, retained-real or mixed component boundaries
- Release-pinned generation basis, intended use, schema constraints, utility and privacy evidence references
- Local claim status, release decision references, distribution bindings and withdrawal or reassessment notices
Out of scope
- Running generators, training models, executing attacks or computing quality and privacy assessments
- Owning source data, generic data catalog operations, schema definitions, lineage graphs or evaluator lifecycles
- Automatic anonymisation certification, legal conclusions, universal utility thresholds, real-world observations inferred from generated records
- Operational generation or use of dangerous-subject data; any such domain extension remains policy-level and separately reviewed
Why it exists Filled
Describe a governed generated-data product through release-specific lineage, fitness evidence, privacy claims and distribution conditions.
Distinguishing features Filled
- Its records are generated, not collected, so no row corresponds to a real unit by design.
- It carries fidelity and privacy evidence as part of the product, unlike an ordinary dataset release.
- Distinct from de-identified or masked data, which still consists of transformed real records.
- Distinct from test fixtures written by hand, which do not model a real source distribution.
What robots and AI may and may not do Filled
Must not
- Present synthetic data as real observations or mix it with real data without labelling.
- Release synthetic data without a disclosure-risk assessment when the source holds personal data.
- Claim a privacy guarantee that the generation method does not provide.
- Use the source data beyond the purpose and agreement under which it was supplied.
Only with a human decision
- Approving the release of synthetic data derived from personal or confidential data.
- Accepting the residual disclosure risk stated in the assessment.
May
- Register a synthetic release with its generator, source reference and evidence.
- Compute and report fidelity and disclosure-risk measures.
- Recommend whether a release suits a stated analytic use based on its evidence.
Moral aspects Filled
- Synthetic data can still leak information about real people if the generator memorises records.
- Low fidelity can lead to wrong conclusions, especially about small or minority groups.
- Unlabelled synthetic data can be mistaken for real evidence in research and decisions.
Who is affected
- Data subjects in the source data
- Analysts and model developers using the release
- Source data holders
- People affected by decisions based on the data
Owners Filled
Steward
The data product owner who commissions the generation answers for the release and its evidence.
Roles
- Product steward
- Maintain synthetic classification, use context and accountable references
- Generation custodian
- Attest externally produced run and configuration bindings without exposing restricted inputs
- Utility reviewer
- Interpret release-specific fitness evidence and rejected uses
- Privacy reviewer
- Qualify disclosure and formal privacy evidence with explicit uncertainty
- Release authority
- Approve audience and conditions through an external decision record
- Records custodian
- Apply scoped retention and withdrawal policy through authorized systems
Master systems
- Data catalogues
- Data product registries
Links to other meta-models Filled
extends
- WM-DAT-008 Data Product / Product Catalog Record - Candidate registry parent specialization: bind one product master and generic product policy; add only synthetic-specific claims. No executable inheritance is certified.
references
- WM-DAT-001 Dataset - Bind generated output snapshots and optional input datasets without owning their lifecycle.
- WM-DAT-004 Data Schema / Data Contract - Pin schemas and constraints used to interpret generated content; definitions stay external.
- WM-DAT-005 Data Pipeline - Reference generation runs if a pipeline is used; no execution or scheduling ownership.
- WM-DAT-006 Data Lineage - Bind external lineage evidence; local run selection is not a full lineage graph.
- WM-DAT-007 Data Quality Evaluation - Bind assessment results with product-specific interpretations; no evaluation execution.
aligned
- W3C DCAT 3, PROV-O and DQV - Conceptual catalogue, provenance and quality mappings; implementer must pin and test a profile before claiming conformance.
neighbor
- WM-DAT-008 Data Product / Product Catalog Record - Registry parent is candidate semantic specialization. Bind its product master identity and generic terms; local content supplies synthetic-specific claims and evidence. Executable inheritance is not ratified.
- WM-DAT-001 Dataset and WM-DAT-004 Data Schema / Data Contract - Content snapshots and schemas are version-pinned references. A product is not each byte file, source dataset or schema; distributions and format conversions do not automatically create a new product.
- WM-DAT-005 Data Pipeline and WM-DAT-006 Data Lineage - Reference generation runs and provenance records. This model records which run supports a release and never orchestrates training, execution or the full lineage lifecycle.
- WM-DAT-007 Data Quality Evaluation - Reference assessments and retain a product-specific interpretation. Evaluators own test execution and measurement records; a utility observation is not universal approval.
- Masked data, simulation outputs and mixed products - Masking or sampling real rows alone does not establish synthetic origin. Rule-based generation can qualify with a declared target and use. Partially generated products must label retained real components; classification alone implies no privacy guarantee.
parent
- WM-DAT-008
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- A release is identified by a persistent dataset identifier and a version, as for any catalogued dataset.
- Its lineage references the source dataset version and the generator configuration.
Direct properties not applicable Not applicable
Not applicable
A synthetic data product is an information artefact with no physical properties to measure.
Recognition optional Filled
- A synthetic release names its generator, its source reference and its fidelity and privacy evidence.
- Often confused with anonymised data, sample data or hand-made test fixtures.
Capabilities and actions required Filled
- New releases can be regenerated from a newer source version or generator configuration.
- Evidence can be recomputed when new assessment methods become available.
Hazards and failure modes required Filled
- Membership inference or attribute disclosure about people in the source data.
- Misleading analysis from synthetic data that does not preserve key relationships.
- Loss of the synthetic label as the data is copied into other datasets.
Standards and interfaces required Filled
- W3C DCAT for catalogue description.
- W3C PROV-O for generation lineage.
- schema.org/Dataset for public discovery.
Context of use required Filled
- Used for sharing, testing and model training where real data cannot be released.
- Where the source holds personal data, data protection law applies to the generation step.
Sources Filled
- Data Catalog Vocabulary (DCAT) - Version 3 - World Wide Web Consortium
- PROV-O: The PROV Ontology - World Wide Web Consortium
- Data on the Web Best Practices: Data Quality Vocabulary - World Wide Web Consortium
- Guidelines for Evaluating Differential Privacy Guarantees - National Institute of Standards and Technology
- De-Identifying Government Datasets: Techniques and Governance - National Institute of Standards and Technology
- Synthetic data policy - Office for National Statistics
- Introduction to anonymisation - Information Commissioner's Office
- NIST SP 800-188 De-Identifying Government Datasets (NIST)
- ISO/IEC 20889 Privacy enhancing data de-identification terminology and classification of techniques (ISO/IEC)
Open questions
- Pin source editions and errata, run the coordinator HTTP checker, and independently verify claim support and regional legal applicability.
- Define adopting instance schemas and negative fixtures for mixed origin, stale digests, incomplete claims, audience expansion and incomplete withdrawal.
- Review multimodal, longitudinal, graph and multi-table profiles with purpose-specific utility and qualified privacy evidence.
- Restore independent external review before any canonical or publishable-draft promotion.
- No executable nested instance schemas, external adapters, thresholds or benchmark fixtures are supplied.
- No generator is run, dataset evaluated, formal privacy accountant implemented or source licence adjudicated.
- Modality-specific text, image, longitudinal, graph, simulation and multi-table evaluation requirements remain deferred.
- No independent external review or direct HTTP status measurement in this sandbox.
Machine files
Provenance
world-models research · reviewable-draft
Built from: models/wm-dat-012-synthetic-data-product/spec.yaml, ver-cy/world-models/card-supplements/wm-dat-012-synthetic-data-product.json