Protein / Biomolecule
Represent one scoped molecular kind with scientific master references, form definitions and evidence-qualified properties, structures and biological claims.
Bundle → Layer → Finding → Questions Filled
3 bundles · 3 layers · 4 findings · 9 questions
Identity and sequence What the molecule is.
Identifiers and sequence
Repository identifiers and the defining sequence or chemical structure.
Master record
The molecule is anchored to a master repository record with a stable accession.
- Which repository accession identifies this molecule, and in which release?
- Which organism and gene is the protein attributed to?
- Is the record reviewed or automatically annotated?
Isoform and modification
The exact sequence form, isoform and modifications are stated.
- Which isoform or sequence version is meant?
- Which post-translational modifications are recorded?
Structure How the molecule is shaped.
Three-dimensional structure
Experimental and predicted structures.
Structure evidence
A structure states whether it is experimental or predicted, with method and quality measures.
- Is the structure experimental or computationally predicted?
- Which method and resolution or confidence score apply?
Function What the molecule does.
Function and interactions
Molecular function, pathways and interaction partners.
Annotated function
Function claims carry evidence codes and references.
- Which functions are annotated, and with which evidence codes?
- Which pathways and interaction partners are recorded?
Classifiers Filled
- Family
- World Models
- Category
- Physical world and living systems
- Entry kind
- entity
- Navigation path
- NAV.PHY.LIV.MOL
- Domain
- PHY.LIV.MOL
- Industry
- Cross-industry
- Tags
- proteinbiomoleculephy.liv.mol
What it is Filled
A protein or biomolecule is a molecular entity produced by living organisms, such as a protein, nucleic acid, lipid or metabolite, described by its sequence or chemical structure, its three-dimensional structure and its function. The subject is the molecular kind as recorded in scientific master repositories, not a physical sample of it and not the gene that encodes it.
In scope
- Scoped molecular identity, accessions and ambiguity-aware crosswalks.
- Sequence or chemical definitions, form differences and qualified properties.
- Structure, recognition, function, context and interaction evidence references.
- Local assertion provenance, curation continuity and governed exchange.
Out of scope
- Physical specimen, batch, reagent inventory, custody and experimental execution.
- Gene, transcript, genome, organism and pathway lifecycle or process simulation.
- Molecular engineering, production, optimization, clinical decisions and regulatory approvals.
- A complete chemical ontology, all biomolecule subclasses, raw experimental archives or external master curation.
Why it exists Filled
Represent one scoped molecular kind with scientific master references, form definitions and evidence-qualified properties, structures and biological claims.
Distinguishing features Filled
- It is the molecular kind, not a physical specimen, batch or reagent lot.
- Distinct from the gene or genomic feature that encodes it.
- Identity rests on independent scientific master repositories rather than any one owner.
- Experimental and predicted structures are kept apart because their evidence differs.
What robots and AI may and may not do Filled
Must not
- Present a predicted structure or function as experimentally established.
- Provide design, production or enhancement guidance for toxins or other harmful biomolecules.
- Merge records for different isoforms or species without saying so.
- Cite an accession without the repository release it came from.
- Make medical claims about a molecule beyond its evidence.
Only with a human decision
- Curating or correcting a master repository record.
- Using molecule data in clinical or regulatory decisions.
- Handling requests that touch biosecurity concerns.
May
- Resolve and cross-reference molecule identifiers across repositories.
- Summarise annotated function and structure with their evidence.
- Flag conflicting or outdated annotations.
Moral aspects Filled
- Some biomolecules are toxins or allergens, so information about them carries dual-use concerns.
- Open repositories are a shared scientific resource whose accuracy many depend on.
- Molecules derived from genetic resources may fall under access and benefit-sharing rules.
Who is affected
- Researchers and clinicians
- Patients relying on research results
- Repository curators
- Countries and communities providing genetic resources
Owners Filled
Steward
A scientific repository consortium that curates the master records and their identifiers.
Roles
- Scientific data steward
- Own the local identity profile and source-version decisions.
- Scientific curator
- Assess claim evidence, form ambiguity and context restrictions.
- Evidence reviewer
- Challenge mappings, contradictions and method applicability.
- Access and retention custodian
- Apply the adopting policy to linked artifacts and local payload disposal.
- Model agent
- Perform only authorized local read, validation and proposal operations; escalate unresolved scientific or policy decisions.
Master systems
- Protein sequence knowledge base
- Macromolecular structure archive
- Chemical entities of biological interest database
Links to other meta-models Filled
references
- WM-LIV-002 - Candidate organism context only when an identified source individual is relevant. Reject registry parent_ids as biological inheritance or containment: a molecular kind is not an organism individual.
- WM-LIV-013 - Candidate gene context; gene identity, genomic coordinates and its lifecycle remain external. Binding version must be reviewed before implementation.
- WM-LIV-023 - Candidate specimen context; custody, consent and specimen measurements remain externally mastered, not molecular-kind properties.
- WM-LIV-016 - Candidate pathway context; this root holds participation evidence, not pathway dynamics or process execution.
- WM-MAT-001 - Candidate chemical-substance boundary reconciliation for nonprotein forms. Preserve one external chemical master; avoid duplicate substance, batch or inventory lifecycles.
aligned
- PDBx/mmCIF Version 5 and ModelCIF - Selected molecular entity and structure-evidence mapping only; entry, entity, chain and assembly identifiers are distinct. Exact dictionary pins and conformance deferred.
- ProForma 2.0 and 2.1 - Explicit version selection for form notation. The 2.0 research paper supports concepts; 2.1 syntax and parser support require a separate conformance check.
- Gene Ontology annotations and PSI-MITAB 2.8 - Preserve function evidence and interaction qualifiers using versioned vocabulary bindings; no automatic whole-schema equivalence.
- PROV-O 2013-04-30 - Attribution and derivation of local assertions; rights, truth, audit enforcement and runtime policy are not supplied by this ontology.
neighbor
- WM-LIV-002 - Registry parent is contextual only, not entity inheritance; a molecule kind is not an organism individual.
- WM-LIV-013 - Encoding reference does not make a protein the same subject as its gene.
- WM-LIV-023 - A specimen is a physical occurrence with custody; this root describes a molecular kind.
- WM-MAT-001 - Shared chemical identity may refer to the same external master; substance and material lifecycle is not copied here.
- Structure archive record - One entry can describe multiple entities and representations; entry identity cannot substitute for the molecular subject.
parent
- WM-LIV-002
What else AI and robots need to interact with it Filled
Identity and identifiers required Filled
- Proteins are identified by UniProt accessions, structures by Protein Data Bank identifiers and small biomolecules by ChEBI identifiers.
- Enzymes can also be classified by EC numbers, which name a reaction class rather than one molecule.
Direct properties required Filled
- Molecular mass in daltons.
- Sequence length in residues for proteins and nucleotides for nucleic acids.
- Isoelectric point as a pH value, measured or computed.
- Structure resolution in angstroms for experimental structures, or a confidence score for predicted ones.
Recognition required Filled
- A molecule is recognised by sequence, mass spectrometry, structure determination or binding assays.
- Often confused with the gene, with another isoform or with a homologous protein from another species.
Capabilities and actions required Filled
- Sequences can be compared and aligned across species.
- Structures can be determined experimentally or predicted computationally.
- Functions can be annotated with evidence codes from the literature.
Hazards and failure modes required Filled
- Biological toxins and allergens with health effects.
- Misannotation that propagates through automated pipelines.
- Dual-use misuse of information about harmful molecules.
Standards and interfaces required Filled
- UniProt for protein sequences and annotation.
- Worldwide Protein Data Bank archive and the PDBx/mmCIF format for structures.
- Gene Ontology annotations for function.
Context of use required Filled
- Used in life science research, drug discovery, diagnostics and biotechnology.
- Work with harmful biomolecules is subject to biosafety and biosecurity rules.
Sources Filled
- Accession - UniProt Consortium
- How do I get the nucleotide sequence that corresponds to the UniProtKB sequence? - UniProt Consortium
- About ChEBI - EMBL-EBI
- Proteomics Standards Initiative ProForma 2.0: Unifying the Encoding of Proteoforms and Peptidoforms - HUPO Proteomics Standards Initiative and Consortium for Top-Down Proteomics
- PDBx/mmCIF User Guide - Worldwide Protein Data Bank
- Stand-alone wwPDB Validation Service - Worldwide Protein Data Bank
- ModelCIF dictionary documentation - ModelCIF Working Group
- Introduction to GO annotations - Gene Ontology Consortium
- PSI-MI TAB 2.8 Format - HUPO Proteomics Standards Initiative
- PROV-O: The PROV Ontology - World Wide Web Consortium
- ProForma version and implementation status - HUPO Proteomics Standards Initiative
- UniProt Knowledgebase (UniProt Consortium)
- Worldwide Protein Data Bank (wwPDB)
- Chemical Entities of Biological Interest, ChEBI (EMBL-EBI)
Open questions
- Run the coordinator source checker, retrieve complete limited-access source text and pin source versions; separately assess claim support and applicable licenses.
- Develop specialist nonprotein identity profiles and reconcile chemical-substance overlap without duplicating scientific master identity.
- Implement and test nested schemas and bindings with accession splits, ambiguous forms, partial structures, complex expansion, negative annotations and permitted payload disposal.
- Restore independent scientific and external provider review before any canonical or publishable-draft promotion.
- Independent external provider review is absent; a separate local Codex self-audit cannot replace it.
- Direct HTTP checks are not run because the sandbox is blocked; current release pins and complete source text remain unverified for selected sources.
- Specialist nucleic-acid, glycan, lipid, heterogeneous polymer and complex identity profiles need further research.
- Nested instance schemas, precise neighbor bindings, parser support and adversarial conformance fixtures are not implemented.
- Clinical, regulatory, jurisdictional access and benefit-sharing requirements require qualified profile research; no universal rule is asserted.
Machine files
Provenance
world-models research · reviewable-draft
Built from: models/wm-liv-014-protein-biomolecule/spec.yaml, ver-cy/world-models/card-supplements/wm-liv-014-protein-biomolecule.json