Genome / Genomic Assembly
Versioned genome-level biological reference
Bundle → Layer → Finding → Questions Filled
3 bundles · 3 layers · 5 findings · 10 questions
Assembly identity Which assembly, of which organism and in which version.
Source and version
The organism, sample and assembly version.
Source organism
The taxon and sample from which the assembly was built.
- Which taxon and which specimen or sample was sequenced?
- Is the source sample deposited in a collection or biobank with an identifier?
Version lineage
The assembly version and what changed from the previous one.
- Which assembly version is this, and which version did it replace?
- Are patches or coordinate changes documented between versions?
Assembly quality How complete and correct the assembly is.
Contiguity and completeness
Measures of contiguity, completeness and error.
Quality metrics
Contiguity, completeness and base accuracy measures.
- What contig and scaffold N50, gap count and total length are reported?
- How complete is the assembly against a conserved gene set, and how was this measured?
Method
Sequencing technologies and assembly method.
- Which sequencing technologies and coverage depth were used?
- Which assembly method and version produced the assembly?
Use and access How the assembly is deposited, cited and used.
Deposition and terms
Archive deposition and the terms of use.
Public record
The archive accession and any use restrictions.
- Under which archive accession is the assembly deposited?
- Does the assembly come from a human individual, and what consent and access terms apply?
Classifiers Filled
- Family
- World Models
- Category
- Physical world and living systems
- Entry kind
- standalone-mm
- Navigation path
- NAV.PHY.LIV.MOL
- Domain
- PHY.LIV.MOL
- Industry
- Cross-industry
- Tags
- genomegenomicassemblyphy.liv.mol
What it is Filled
A genome assembly is a versioned reconstruction of the genome sequence of an organism or individual from sequencing reads, organized into contigs, scaffolds and chromosomes and used as a reference for coordinates and annotation. The class covers the assembly as a reference with its quality and version history; the physical DNA, single sequencing runs and gene annotations built on the assembly are separate subjects.
Why it exists Filled
Versioned genome-level biological reference
Distinguishing features Filled
- It is a versioned reference sequence, not the physical DNA or a single sequencing run.
- Coordinates are meaningful only with the assembly version, so annotations and variants must cite it.
- It differs from a genome annotation, which describes features placed on the assembly.
- Assemblies of human individuals carry personal data duties that species reference assemblies usually do not.
What robots and AI may and may not do Filled
Must not
- Mix coordinates from different assembly versions without conversion.
- Share an assembly of an identifiable human individual outside its consent and access terms.
- Attempt to re-identify individuals from genomic data.
- Overwrite a published assembly version instead of issuing a new one.
Only with a human decision
- Releasing a human individual assembly or granting controlled access to it.
- Adopting a new reference assembly version for clinical or production use.
- Resolving disputes about sample origin and benefit sharing for genetic resources.
May
- Retrieve public assemblies by accession and report their version and quality metrics.
- Convert coordinates between assembly versions using published mappings.
- Compare assemblies and report differences in contiguity and completeness.
- Flag annotations or variants that do not state their assembly version.
Moral aspects Filled
- Human genome data is identifying and reveals information about relatives as well as the donor.
- Genetic resources from a country or community can carry access and benefit-sharing obligations.
- Reference assemblies that represent few populations can bias research and clinical results.
Who is affected
- Donors of human samples and their relatives
- Countries and communities providing genetic resources
- Researchers and clinicians using the reference
Owners Filled
Steward
The sequencing consortium or research group that produced the assembly, with archive curators maintaining the deposited record.
Master systems
- International nucleotide sequence archives
- Genome assembly databases
- Controlled-access archives for human genomic data
Links to other meta-models Filled
parent
- WM-LIV-002
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- An assembly is identified by its archive accession with version, such as an INSDC assembly accession.
- Well-known references also carry assembly names; the source organism is identified by a taxonomy identifier.
Direct properties not applicable Not applicable
Not applicable
A genome assembly is a digital sequence reference; physical measurements belong to the DNA sample, and assembly quality metrics are recorded in the structure.
Recognition optional Filled
- An assembly is recognized by its sequence files, an accession with version and an assembly report.
- Often confused with raw sequencing reads, with a gene annotation or with the physical specimen.
Capabilities and actions required Filled
- Serves as a coordinate system for genes, variants and other annotations.
- Can be compared with other assemblies to find structural differences.
- Can be updated by patches or new versions with documented coordinate changes.
Hazards and failure modes required Filled
- Wrong variant calls or annotations when coordinates from different versions are mixed.
- Re-identification of donors from human genomic data.
- Misleading results from contaminated or misassembled regions.
Standards and interfaces required Filled
- FASTA for sequences and AGP for the placement of contigs into scaffolds.
- INSDC archive submission and retrieval services.
- GA4GH standards such as refget for sequence retrieval and data access.
Context of use required Filled
- Used in genomics research, breeding, biodiversity science and clinical genetics.
- Human genomic data is subject to data protection law, and genetic resources may fall under the Nagoya Protocol.
Sources Filled
- International Nucleotide Sequence Database Collaboration (INSDC)
- Genome Reference Consortium
- Global Alliance for Genomics and Health (GA4GH) standards
Open questions
- Planned model: boundary questions, research and every section remain to be written.
Machine files
Provenance
planned (registry candidate) · todo
Built from: models/runtime-index.json, ver-cy/world-models/card-supplements/wm-liv-011-genome-genomic-assembly.json
Planned entry, hidden from the catalogue until researched.