{"schema":"https://ver.cy/schemas/card/1.0.0","id":"vr.wm-ai-009","code":"wm-ai-009-evaluation-dataset-benchmark","url":"https://ver.cy/models/wm-ai-009-evaluation-dataset-benchmark/","name":"Evaluation Dataset / Benchmark","alternateNames":[],"kind":"world-model","status":"published","version":"0.3.0-research.1","language":"en","classifiers":{"family":"World Models","category":"Information and virtual systems","entryKind":"aggregate","plane":"","domain":["INF.AI.DAT"],"industry":["Cross-industry"],"navPath":"NAV.INF.AI.DAT","tags":["evaluation","dataset","benchmark","inf.ai.dat"],"facets":{}},"whatItIs":"Owns benchmark identity, family, suite, task and release; intended use and population; versioned items, schemas, labels, reference answers, rubrics, annotations, splits, sampling and hidden-test controls; source lineage, rights, privacy, security and access; protocol, environment, metrics, scoring, statistics, validity, contamination, subgroups, robustness, limitations, lifecycle, result bindings, retention, audit and projections. Evaluated models, evaluation runs, results, leaderboards, claims, policy, credential, provenance, audit and records masters remain external.","purpose":"Represent one governed, versioned evaluation dataset or benchmark definition that makes population, content, labels, splits, protocol, metrics, validity and release semantics machine-readable without absorbing evaluation runs or result masters.","scope":{"in":["Benchmark identity, release, purpose, task, population, content, items, labels, references, rubrics, annotations, splits, sampling, lineage, rights and access","Protocol, interface, environment, metrics, scoring, uncertainty, quality, leakage, contamination, subgroup, robustness, safety, lifecycle, result bindings and projections"],"out":["Owning or executing evaluated models, evaluation runs, results, leaderboards, deployment decisions, policy, credentials, provenance, audit or records masters","Treating a score, leaderboard position, checksum, split label or publication as proof of safety, fairness, validity or fitness","Autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup"],"boundaries":[{"neighbor":"WM-DAT-001 Dataset","distinction":"The catalogue parent signal identifies the dataset boundary, but no frozen relation-ledger edge grants inheritance or mutation authority. The benchmark adds task, split-role, protocol, metric and validity semantics while source datasets remain external masters."},{"neighbor":"WM-AI-003 Evaluation Run","distinction":"The incoming candidate reference expresses evaluation reproducibility. A run owns execution inputs, environment, outputs and measurements; this model owns the benchmark definition and non-owning run or result bindings."},{"neighbor":"AI model, result, leaderboard and claim","distinction":"Models are evaluated subjects; results and leaderboards are version-qualified observations and presentations. None becomes part of the benchmark definition merely by reference."},{"neighbor":"DCAT, Croissant, SPDX, Hugging Face, MLPerf, PROV and SKOS","distinction":"These are catalog, dataset, supply-chain, card, benchmark, provenance and classification profiles with different scopes. Every mapping requires a release pin and information-loss declaration."}]},"distinguishingFeatures":["Defines a versioned evaluation dataset with protocol and metrics, not the evaluation runs or the results.","Makes split roles, hidden tests and contamination controls explicit.","Differs from a general dataset by declaring intended use, population and scoring.","Treats a leaderboard score as an observation, not as proof of validity."],"structure":{"bundles":[{"id":"benchmark-identity-scope-version-and-definition","name":"Benchmark identity, scope, version and definition","description":"Groups governed context for benchmark identity, scope, version and definition.","layers":[{"id":"benchmark-root-family-suite-task-and-release","name":"Benchmark root, family, suite, task and release","description":"Groups benchmark context for benchmark root, family, suite, task and release.","findings":[{"id":"benchmark-identity-namespace-owner-revision-and-current-head","name":"Benchmark identity, namespace, owner, revision and current head","description":"Records benchmark identity, namespace, owner, revision and current head as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish benchmark identity, namespace, owner, revision and current head?","id":"benchmark-identity-namespace-owner-revision-and-current-head-q01","kind":"identity"},{"text":"Who may create, assert, review, approve, correct or disclose benchmark identity, namespace, owner, revision and current head, under which purpose and authority?","id":"benchmark-identity-namespace-owner-revision-and-current-head-q02","kind":"composition"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify benchmark identity, namespace, owner, revision and current head?","id":"benchmark-identity-namespace-owner-revision-and-current-head-q03","kind":"privacy"}]},{"id":"family-suite-task-profile-version-alias-and-equivalence","name":"Family, suite, task, profile, version, alias and equivalence","description":"Records family, suite, task, profile, version, alias and equivalence as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish family, suite, task, profile, version, alias and equivalence?","id":"family-suite-task-profile-version-alias-and-equivalence-q01","kind":"classification"},{"text":"Who may create, assert, review, approve, correct or disclose family, suite, task, profile, version, alias and equivalence, under which purpose and authority?","id":"family-suite-task-profile-version-alias-and-equivalence-q02","kind":"evidence"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify family, suite, task, profile, version, alias and equivalence?","id":"family-suite-task-profile-version-alias-and-equivalence-q03","kind":"lifecycle"}]}]},{"id":"purpose-population-domain-and-coverage","name":"Purpose, population, domain and coverage","description":"Groups benchmark context for purpose, population, domain and coverage.","findings":[{"id":"intended-use-question-domain-modality-language-and-jurisdiction","name":"Intended use, question, domain, modality, language and jurisdiction","description":"Records intended use, question, domain, modality, language and jurisdiction as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish intended use, question, domain, modality, language and jurisdiction?","id":"intended-use-question-domain-modality-language-and-jurisdiction-q01","kind":"requirement"},{"text":"Who may create, assert, review, approve, correct or disclose intended use, question, domain, modality, language and jurisdiction, under which purpose and authority?","id":"intended-use-question-domain-modality-language-and-jurisdiction-q02","kind":"ownership"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify intended use, question, domain, modality, language and jurisdiction?","id":"intended-use-question-domain-modality-language-and-jurisdiction-q03","kind":"quality"}]},{"id":"target-population-sampling-frame-inclusion-exclusion-and-coverage","name":"Target population, sampling frame, inclusion, exclusion and coverage","description":"Records target population, sampling frame, inclusion, exclusion and coverage as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish target population, sampling frame, inclusion, exclusion and coverage?","id":"target-population-sampling-frame-inclusion-exclusion-and-coverage-q01","kind":"measurement"},{"text":"Who may create, assert, review, approve, correct or disclose target population, sampling frame, inclusion, exclusion and coverage, under which purpose and authority?","id":"target-population-sampling-frame-inclusion-exclusion-and-coverage-q02","kind":"measurement"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify target population, sampling frame, inclusion, exclusion and coverage?","id":"target-population-sampling-frame-inclusion-exclusion-and-coverage-q03","kind":"security"}]}]}]},{"id":"content-items-labels-splits-and-distributions","name":"Content, items, labels, splits and distributions","description":"Groups governed context for content, items, labels, splits and distributions.","layers":[{"id":"item-schema-input-reference-and-annotation","name":"Item schema, input, reference and annotation","description":"Groups benchmark context for item schema, input, reference and annotation.","findings":[{"id":"item-case-identity-input-prompt-context-response-and-schema","name":"Item or case identity, input, prompt, context, response and schema","description":"Records item or case identity, input, prompt, context, response and schema as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish item or case identity, input, prompt, context, response and schema?","id":"item-case-identity-input-prompt-context-response-and-schema-q01","kind":"identity"},{"text":"Who may create, assert, review, approve, correct or disclose item or case identity, input, prompt, context, response and schema, under which purpose and authority?","id":"item-case-identity-input-prompt-context-response-and-schema-q02","kind":"exception"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify item or case identity, input, prompt, context, response and schema?","id":"item-case-identity-input-prompt-context-response-and-schema-q03","kind":"retention"}]},{"id":"label-reference-answer-rubric-annotation-adjudication-and-uncertainty","name":"Label, reference answer, rubric, annotation, adjudication and uncertainty","description":"Records label, reference answer, rubric, annotation, adjudication and uncertainty as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish label, reference answer, rubric, annotation, adjudication and uncertainty?","id":"label-reference-answer-rubric-annotation-adjudication-and-uncertainty-q01","kind":"quality"},{"text":"Who may create, assert, review, approve, correct or disclose label, reference answer, rubric, annotation, adjudication and uncertainty, under which purpose and authority?","id":"label-reference-answer-rubric-annotation-adjudication-and-uncertainty-q02","kind":"provenance"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify label, reference answer, rubric, annotation, adjudication and uncertainty?","id":"label-reference-answer-rubric-annotation-adjudication-and-uncertainty-q03","kind":"interoperability"}]}]},{"id":"splits-sampling-hidden-tests-and-distributions","name":"Splits, sampling, hidden tests and distributions","description":"Groups benchmark context for splits, sampling, hidden tests and distributions.","findings":[{"id":"train-validation-calibration-evaluation-split-role-and-leakage-control","name":"Train, validation, calibration and evaluation split role and leakage control","description":"Records train, validation, calibration and evaluation split role and leakage control as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish train, validation, calibration and evaluation split role and leakage control?","id":"train-validation-calibration-evaluation-split-role-and-leakage-control-q01","kind":"validation"},{"text":"Who may create, assert, review, approve, correct or disclose train, validation, calibration and evaluation split role and leakage control, under which purpose and authority?","id":"train-validation-calibration-evaluation-split-role-and-leakage-control-q02","kind":"process"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify train, validation, calibration and evaluation split role and leakage control?","id":"train-validation-calibration-evaluation-split-role-and-leakage-control-q03","kind":"decision"}]},{"id":"sampling-stratification-weight-hidden-test-distribution-and-digest","name":"Sampling, stratification, weight, hidden test, distribution and digest","description":"Records sampling, stratification, weight, hidden test, distribution and digest as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish sampling, stratification, weight, hidden test, distribution and digest?","id":"sampling-stratification-weight-hidden-test-distribution-and-digest-q01","kind":"security"},{"text":"Who may create, assert, review, approve, correct or disclose sampling, stratification, weight, hidden test, distribution and digest, under which purpose and authority?","id":"sampling-stratification-weight-hidden-test-distribution-and-digest-q02","kind":"validation"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify sampling, stratification, weight, hidden test, distribution and digest?","id":"sampling-stratification-weight-hidden-test-distribution-and-digest-q03","kind":"state"}]}]}]},{"id":"provenance-rights-privacy-security-and-quality","name":"Provenance, rights, privacy, security and quality","description":"Groups governed context for provenance, rights, privacy, security and quality.","layers":[{"id":"source-collection-transformation-and-lineage","name":"Source, collection, transformation and lineage","description":"Groups benchmark context for source, collection, transformation and lineage.","findings":[{"id":"source-acquisition-collection-generation-transform-and-release-lineage","name":"Source, acquisition, collection, generation, transform and release lineage","description":"Records source, acquisition, collection, generation, transform and release lineage as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish source, acquisition, collection, generation, transform and release lineage?","id":"source-acquisition-collection-generation-transform-and-release-lineage-q01","kind":"provenance"},{"text":"Who may create, assert, review, approve, correct or disclose source, acquisition, collection, generation, transform and release lineage, under which purpose and authority?","id":"source-acquisition-collection-generation-transform-and-release-lineage-q02","kind":"privacy"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify source, acquisition, collection, generation, transform and release lineage?","id":"source-acquisition-collection-generation-transform-and-release-lineage-q03","kind":"identity"}]},{"id":"contributor-tool-pipeline-derivation-version-and-current-release","name":"Contributor, tool, pipeline, derivation, version and current release","description":"Records contributor, tool, pipeline, derivation, version and current release as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish contributor, tool, pipeline, derivation, version and current release?","id":"contributor-tool-pipeline-derivation-version-and-current-release-q01","kind":"ownership"},{"text":"Who may create, assert, review, approve, correct or disclose contributor, tool, pipeline, derivation, version and current release, under which purpose and authority?","id":"contributor-tool-pipeline-derivation-version-and-current-release-q02","kind":"lifecycle"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify contributor, tool, pipeline, derivation, version and current release?","id":"contributor-tool-pipeline-derivation-version-and-current-release-q03","kind":"classification"}]}]},{"id":"rights-consent-privacy-security-and-access","name":"Rights, consent, privacy, security and access","description":"Groups benchmark context for rights, consent, privacy, security and access.","findings":[{"id":"license-ip-data-right-consent-purpose-export-and-use-term","name":"License, IP, data right, consent, purpose, export and use term","description":"Records license, ip, data right, consent, purpose, export and use term as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish license, ip, data right, consent, purpose, export and use term?","id":"license-ip-data-right-consent-purpose-export-and-use-term-q01","kind":"authority"},{"text":"Who may create, assert, review, approve, correct or disclose license, ip, data right, consent, purpose, export and use term, under which purpose and authority?","id":"license-ip-data-right-consent-purpose-export-and-use-term-q02","kind":"quality"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify license, ip, data right, consent, purpose, export and use term?","id":"license-ip-data-right-consent-purpose-export-and-use-term-q03","kind":"relationship"}]},{"id":"personal-sensitive-data-deidentification-security-access-and-disclosure","name":"Personal or sensitive data, deidentification, security, access and disclosure","description":"Records personal or sensitive data, deidentification, security, access and disclosure as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish personal or sensitive data, deidentification, security, access and disclosure?","id":"personal-sensitive-data-deidentification-security-access-and-disclosure-q01","kind":"privacy"},{"text":"Who may create, assert, review, approve, correct or disclose personal or sensitive data, deidentification, security, access and disclosure, under which purpose and authority?","id":"personal-sensitive-data-deidentification-security-access-and-disclosure-q02","kind":"security"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify personal or sensitive data, deidentification, security, access and disclosure?","id":"personal-sensitive-data-deidentification-security-access-and-disclosure-q03","kind":"authority"}]}]}]},{"id":"protocol-metrics-scoring-and-statistics","name":"Protocol, metrics, scoring and statistics","description":"Groups governed context for protocol, metrics, scoring and statistics.","layers":[{"id":"execution-protocol-interface-and-environment","name":"Execution protocol, interface and environment","description":"Groups benchmark context for execution protocol, interface and environment.","findings":[{"id":"evaluated-interface-task-contract-prompt-template-and-tool-access","name":"Evaluated interface, task contract, prompt template and tool access","description":"Records evaluated interface, task contract, prompt template and tool access as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish evaluated interface, task contract, prompt template and tool access?","id":"evaluated-interface-task-contract-prompt-template-and-tool-access-q01","kind":"requirement"},{"text":"Who may create, assert, review, approve, correct or disclose evaluated interface, task contract, prompt template and tool access, under which purpose and authority?","id":"evaluated-interface-task-contract-prompt-template-and-tool-access-q02","kind":"retention"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify evaluated interface, task contract, prompt template and tool access?","id":"evaluated-interface-task-contract-prompt-template-and-tool-access-q03","kind":"requirement"}]},{"id":"inference-setting-environment-repetition-randomness-resource-and-timing","name":"Inference setting, environment, repetition, randomness, resource and timing","description":"Records inference setting, environment, repetition, randomness, resource and timing as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish inference setting, environment, repetition, randomness, resource and timing?","id":"inference-setting-environment-repetition-randomness-resource-and-timing-q01","kind":"temporal"},{"text":"Who may create, assert, review, approve, correct or disclose inference setting, environment, repetition, randomness, resource and timing, under which purpose and authority?","id":"inference-setting-environment-repetition-randomness-resource-and-timing-q02","kind":"interoperability"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify inference setting, environment, repetition, randomness, resource and timing?","id":"inference-setting-environment-repetition-randomness-resource-and-timing-q03","kind":"constraint"}]}]},{"id":"metrics-thresholds-scoring-aggregation-and-uncertainty","name":"Metrics, thresholds, scoring, aggregation and uncertainty","description":"Groups benchmark context for metrics, thresholds, scoring, aggregation and uncertainty.","findings":[{"id":"metric-definition-direction-unit-threshold-rubric-and-valid-range","name":"Metric definition, direction, unit, threshold, rubric and valid range","description":"Records metric definition, direction, unit, threshold, rubric and valid range as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish metric definition, direction, unit, threshold, rubric and valid range?","id":"metric-definition-direction-unit-threshold-rubric-and-valid-range-q01","kind":"measurement"},{"text":"Who may create, assert, review, approve, correct or disclose metric definition, direction, unit, threshold, rubric and valid range, under which purpose and authority?","id":"metric-definition-direction-unit-threshold-rubric-and-valid-range-q02","kind":"decision"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify metric definition, direction, unit, threshold, rubric and valid range?","id":"metric-definition-direction-unit-threshold-rubric-and-valid-range-q03","kind":"event"}]},{"id":"item-score-aggregation-weight-statistic-interval-significance-and-uncertainty","name":"Item score, aggregation, weight, statistic, interval, significance and uncertainty","description":"Records item score, aggregation, weight, statistic, interval, significance and uncertainty as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish item score, aggregation, weight, statistic, interval, significance and uncertainty?","id":"item-score-aggregation-weight-statistic-interval-significance-and-uncertainty-q01","kind":"quality"},{"text":"Who may create, assert, review, approve, correct or disclose item score, aggregation, weight, statistic, interval, significance and uncertainty, under which purpose and authority?","id":"item-score-aggregation-weight-statistic-interval-significance-and-uncertainty-q02","kind":"state"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify item score, aggregation, weight, statistic, interval, significance and uncertainty?","id":"item-score-aggregation-weight-statistic-interval-significance-and-uncertainty-q03","kind":"temporal"}]}]}]},{"id":"validity-contamination-robustness-fairness-and-limitations","name":"Validity, contamination, robustness, fairness and limitations","description":"Groups governed context for validity, contamination, robustness, fairness and limitations.","layers":[{"id":"quality-validity-leakage-and-contamination","name":"Quality, validity, leakage and contamination","description":"Groups benchmark context for quality, validity, leakage and contamination.","findings":[{"id":"completeness-label-accuracy-schema-conformance-duplicate-and-quality","name":"Completeness, label accuracy, schema conformance, duplicate and quality","description":"Records completeness, label accuracy, schema conformance, duplicate and quality as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish completeness, label accuracy, schema conformance, duplicate and quality?","id":"completeness-label-accuracy-schema-conformance-duplicate-and-quality-q01","kind":"validation"},{"text":"Who may create, assert, review, approve, correct or disclose completeness, label accuracy, schema conformance, duplicate and quality, under which purpose and authority?","id":"completeness-label-accuracy-schema-conformance-duplicate-and-quality-q02","kind":"identity"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify completeness, label accuracy, schema conformance, duplicate and quality?","id":"completeness-label-accuracy-schema-conformance-duplicate-and-quality-q03","kind":"composition"}]},{"id":"overlap-leakage-contamination-memorization-gaming-and-detection","name":"Overlap, leakage, contamination, memorization, gaming and detection","description":"Records overlap, leakage, contamination, memorization, gaming and detection as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish overlap, leakage, contamination, memorization, gaming and detection?","id":"overlap-leakage-contamination-memorization-gaming-and-detection-q01","kind":"security"},{"text":"Who may create, assert, review, approve, correct or disclose overlap, leakage, contamination, memorization, gaming and detection, under which purpose and authority?","id":"overlap-leakage-contamination-memorization-gaming-and-detection-q02","kind":"classification"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify overlap, leakage, contamination, memorization, gaming and detection?","id":"overlap-leakage-contamination-memorization-gaming-and-detection-q03","kind":"evidence"}]}]},{"id":"subgroups-robustness-safety-and-limitations","name":"Subgroups, robustness, safety and limitations","description":"Groups benchmark context for subgroups, robustness, safety and limitations.","findings":[{"id":"subgroup-coverage-performance-fairness-representativeness-and-disparity","name":"Subgroup coverage, performance, fairness, representativeness and disparity","description":"Records subgroup coverage, performance, fairness, representativeness and disparity as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish subgroup coverage, performance, fairness, representativeness and disparity?","id":"subgroup-coverage-performance-fairness-representativeness-and-disparity-q01","kind":"measurement"},{"text":"Who may create, assert, review, approve, correct or disclose subgroup coverage, performance, fairness, representativeness and disparity, under which purpose and authority?","id":"subgroup-coverage-performance-fairness-representativeness-and-disparity-q02","kind":"relationship"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify subgroup coverage, performance, fairness, representativeness and disparity?","id":"subgroup-coverage-performance-fairness-representativeness-and-disparity-q03","kind":"ownership"}]},{"id":"adversarial-robustness-privacy-security-safety-limit-and-misuse","name":"Adversarial robustness, privacy, security, safety, limit and misuse","description":"Records adversarial robustness, privacy, security, safety, limit and misuse as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish adversarial robustness, privacy, security, safety, limit and misuse?","id":"adversarial-robustness-privacy-security-safety-limit-and-misuse-q01","kind":"constraint"},{"text":"Who may create, assert, review, approve, correct or disclose adversarial robustness, privacy, security, safety, limit and misuse, under which purpose and authority?","id":"adversarial-robustness-privacy-security-safety-limit-and-misuse-q02","kind":"authority"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify adversarial robustness, privacy, security, safety, limit and misuse?","id":"adversarial-robustness-privacy-security-safety-limit-and-misuse-q03","kind":"measurement"}]}]}]},{"id":"lifecycle-results-governance-and-projections","name":"Lifecycle, results, governance and projections","description":"Groups governed context for lifecycle, results, governance and projections.","layers":[{"id":"release-validation-correction-and-result-bindings","name":"Release, validation, correction and result bindings","description":"Groups benchmark context for release, validation, correction and result bindings.","findings":[{"id":"draft-candidate-validated-released-active-deprecated-withdrawn-and-superseded","name":"Draft, candidate, validated, released, active, deprecated, withdrawn and superseded","description":"Records draft, candidate, validated, released, active, deprecated, withdrawn and superseded as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish draft, candidate, validated, released, active, deprecated, withdrawn and superseded?","id":"draft-candidate-validated-released-active-deprecated-withdrawn-and-superseded-q01","kind":"lifecycle"},{"text":"Who may create, assert, review, approve, correct or disclose draft, candidate, validated, released, active, deprecated, withdrawn and superseded, under which purpose and authority?","id":"draft-candidate-validated-released-active-deprecated-withdrawn-and-superseded-q02","kind":"requirement"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify draft, candidate, validated, released, active, deprecated, withdrawn and superseded?","id":"draft-candidate-validated-released-active-deprecated-withdrawn-and-superseded-q03","kind":"exception"}]},{"id":"evaluation-run-result-leaderboard-baseline-claim-correction-and-binding","name":"Evaluation run, result, leaderboard, baseline, claim, correction and binding","description":"Records evaluation run, result, leaderboard, baseline, claim, correction and binding as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish evaluation run, result, leaderboard, baseline, claim, correction and binding?","id":"evaluation-run-result-leaderboard-baseline-claim-correction-and-binding-q01","kind":"relationship"},{"text":"Who may create, assert, review, approve, correct or disclose evaluation run, result, leaderboard, baseline, claim, correction and binding, under which purpose and authority?","id":"evaluation-run-result-leaderboard-baseline-claim-correction-and-binding-q02","kind":"constraint"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify evaluation run, result, leaderboard, baseline, claim, correction and binding?","id":"evaluation-run-result-leaderboard-baseline-claim-correction-and-binding-q03","kind":"provenance"}]}]},{"id":"access-retention-audit-and-interoperability","name":"Access, retention, audit and interoperability","description":"Groups benchmark context for access, retention, audit and interoperability.","findings":[{"id":"role-purpose-access-hidden-test-audit-retention-hold-and-tombstone","name":"Role, purpose, access, hidden test, audit, retention, hold and tombstone","description":"Records role, purpose, access, hidden test, audit, retention, hold and tombstone as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish role, purpose, access, hidden test, audit, retention, hold and tombstone?","id":"role-purpose-access-hidden-test-audit-retention-hold-and-tombstone-q01","kind":"retention"},{"text":"Who may create, assert, review, approve, correct or disclose role, purpose, access, hidden test, audit, retention, hold and tombstone, under which purpose and authority?","id":"role-purpose-access-hidden-test-audit-retention-hold-and-tombstone-q02","kind":"event"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify role, purpose, access, hidden test, audit, retention, hold and tombstone?","id":"role-purpose-access-hidden-test-audit-retention-hold-and-tombstone-q03","kind":"process"}]},{"id":"dcat-croissant-spdx-huggingface-mlperf-prov-skos-projection-and-loss","name":"DCAT, Croissant, SPDX, Hugging Face, MLPerf, PROV and SKOS projection and loss","description":"Records dcat, croissant, spdx, hugging face, mlperf, prov and skos projection and loss as versioned benchmark-definition context while evaluation runs, models, results, leaderboards, policies and records masters remain external.","questions":[{"text":"What stable identity, benchmark release, declared value and explicit unknown establish dcat, croissant, spdx, hugging face, mlperf, prov and skos projection and loss?","id":"dcat-croissant-spdx-huggingface-mlperf-prov-skos-projection-and-loss-q01","kind":"interoperability"},{"text":"Who may create, assert, review, approve, correct or disclose dcat, croissant, spdx, hugging face, mlperf, prov and skos projection and loss, under which purpose and authority?","id":"dcat-croissant-spdx-huggingface-mlperf-prov-skos-projection-and-loss-q02","kind":"temporal"},{"text":"Which effective, observed, recorded, ingested and knowledge times, evidence and uncertainty qualify dcat, croissant, spdx, hugging face, mlperf, prov and skos projection and loss?","id":"dcat-croissant-spdx-huggingface-mlperf-prov-skos-projection-and-loss-q03","kind":"validation"}]}]}]}]},"agentConduct":{"may":["Register a benchmark release with items, splits and protocol.","Check models and training data for contamination with test items.","Compute scores under the declared protocol.","Report coverage gaps in population or tasks."],"mustNot":["Disclose hidden test items or reference answers.","Change items, labels or metrics within a released version.","Tune a model on test splits and report the score as held-out.","Release a benchmark with unclear rights or personal data.","Claim fairness or safety from a single benchmark score."],"requiresHuman":["Releasing or withdrawing a benchmark version.","Adjudicating disputed reference labels.","Granting access to hidden test sets."]},"ethics":{"considerations":["Benchmarks shape which systems are deployed, so unrepresentative populations hide harm to groups left out.","Annotators and people whose content forms items deserve fair pay, credit and privacy.","Contaminated benchmarks mislead users about real capability."],"affectedParties":["Groups represented or missing in the items","Annotators and content creators","Users who rely on reported scores"]},"owners":{"steward":"Dimension owner, benchmark mandate and accountable AI system owner","roles":[{"name":"Benchmark owner and accountable AI system owner","responsibilities":["Own purpose, scope, risk acceptance, release boundaries and accountable use."]},{"name":"Dataset and item steward","responsibilities":["Supply source-qualified items, schemas, splits, lineage, rights and quality evidence."]},{"name":"Annotator and adjudicator","responsibilities":["Create reference answers, labels and rubrics while preserving disagreement and uncertainty."]},{"name":"Evaluation methodologist and statistician","responsibilities":["Define reproducible protocol, metrics, aggregation, uncertainty and validity tests."]},{"name":"Independent safety, security and fairness reviewer","responsibilities":["Review contamination, gaming, subgroups, robustness, privacy, security, safety and limitations."]},{"name":"Release and disclosure authority","responsibilities":["Make attributable release, exception, correction, withdrawal and disclosure decisions."]},{"name":"Legal, privacy and records steward","responsibilities":["Own rights, consent, export, privacy, access, hold, retention and disposition profiles."]}],"masterSystems":[]},"relations":[{"target":"WM-DAT-001 Dataset","type":"references","note":"Represent the unfrozen parent boundary as a non-owning dataset reference until an extension edge is approved."},{"target":"WM-AI-003 Evaluation Run, AI Model, Result and Leaderboard","type":"references","note":"Resolve evaluated subjects, executions and outcomes without absorbing their masters."},{"target":"Policy, credential, provenance, audit and records models","type":"references","note":"Resolve authoritative control and evidence records without transferring ownership."},{"target":"DCAT 3, Croissant 1.0, SPDX 3.0.1 Dataset, Hugging Face Dataset Cards, MLPerf policies, PROV-O and SKOS","type":"aligned","note":"Project release-pinned catalog, dataset, supply-chain, card, benchmark, provenance and classification views with declared loss."},{"target":"WM-DAT-001 Dataset","type":"neighbor","note":"The catalogue parent signal identifies the dataset boundary, but no frozen relation-ledger edge grants inheritance or mutation authority. The benchmark adds task, split-role, protocol, metric and validity semantics while source datasets remain external masters."},{"target":"WM-AI-003 Evaluation Run","type":"neighbor","note":"The incoming candidate reference expresses evaluation reproducibility. A run owns execution inputs, environment, outputs and measurements; this model owns the benchmark definition and non-owning run or result bindings."},{"target":"AI model, result, leaderboard and claim","type":"neighbor","note":"Models are evaluated subjects; results and leaderboards are version-qualified observations and presentations. None becomes part of the benchmark definition merely by reference."},{"target":"DCAT, Croissant, SPDX, Hugging Face, MLPerf, PROV and SKOS","type":"neighbor","note":"These are catalog, dataset, supply-chain, card, benchmark, provenance and classification profiles with different scopes. Every mapping requires a release pin and information-loss declaration."},{"target":"WM-DAT-001","type":"parent"}],"interaction":{"identity":{"applicability":"required","items":["Authoritative master-system identifier for each benchmark, release, item, assertion, decision or projection, qualified by issuer, namespace and record kind.","Governed globally resolvable benchmark-release IRI.","Dimension UUID or ULID when neither preceding identifier exists."]},"properties":{"applicability":"not-applicable","items":[]},"recognition":{"applicability":"optional","items":["A benchmark names a family, a release version, tasks, splits, metrics and a protocol.","Often confused with a training dataset, an evaluation run, a leaderboard or a test suite in software."]},"capabilities":{"applicability":"required","items":["Register a benchmark definition: Governed operation to register a benchmark definition without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Compose or import versioned items: Governed operation to compose or import versioned items without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Define split roles and sampling: Governed operation to define split roles and sampling without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Annotate and adjudicate reference data: Governed operation to annotate and adjudicate reference data without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Define the evaluation protocol: Governed operation to define the evaluation protocol without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Define metrics, scoring and aggregation: Governed operation to define metrics, scoring and aggregation without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Validate quality, leakage and contamination: Governed operation to validate quality, leakage and contamination without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Release, deprecate, withdraw or supersede: Governed operation to release, deprecate, withdraw or supersede without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Bind evaluation results and leaderboard observations: Governed operation to bind evaluation results and leaderboard observations without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup.","Correct, project, disclose, retain and audit: Governed operation to correct, project, disclose, retain and audit without autonomous release, hidden-test disclosure, rights waiver, access expansion, benchmark manipulation or destructive cleanup."]},"hazards":{"applicability":"required","items":["Test leakage into training data.","Label errors that reward wrong answers.","Overclaiming capability from narrow tasks."]},"interfaces":{"applicability":"required","items":["MLCommons Croissant metadata format.","W3C DCAT 3.","SPDX 3.0 Dataset profile.","W3C PROV-O.","W3C SKOS for task and label taxonomies."]},"context":{"applicability":"required","items":["Rights, consent, privacy, security, export, testing, disclosure, retention and high-risk AI obligations depend on jurisdiction, sector and use case.","The EU AI Act and GDPR are European Union profiles; NIST publications are voluntary United States guidance unless adopted by policy or contract.","DCAT, Croissant, SPDX, Hugging Face, MLPerf, PROV and SKOS are versioned profiles and are not assumed lossless or universally applicable."]}},"sources":[{"title":"Artificial Intelligence Risk Management Framework (AI RMF 1.0)","url":"https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10","note":"National Institute of Standards and Technology"},{"title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","url":"https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf","note":"National Institute of Standards and Technology"},{"title":"Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations","url":"https://www.nist.gov/publications/adversarial-machine-learning-taxonomy-and-terminology-attacks-and-mitigations-0","note":"National Institute of Standards and Technology"},{"title":"Regulation (EU) 2024/1689 Artificial Intelligence Act","url":"https://eur-lex.europa.eu/eli/reg/2024/1689/oj","note":"European Union"},{"title":"Regulation (EU) 2016/679 General Data Protection Regulation","url":"https://eur-lex.europa.eu/eli/reg/2016/679/oj","note":"European Union"},{"title":"Data Catalog Vocabulary (DCAT) Version 3","url":"https://www.w3.org/TR/vocab-dcat-3/","note":"World Wide Web Consortium"},{"title":"PROV-O: The PROV Ontology","url":"https://www.w3.org/TR/prov-o/","note":"World Wide Web Consortium"},{"title":"Croissant ML-ready Dataset Metadata Format","url":"https://docs.mlcommons.org/croissant/","note":"MLCommons"},{"title":"MLCommons Benchmarks","url":"https://mlcommons.org/benchmarks/","note":"MLCommons"},{"title":"MLPerf Inference Rules","url":"https://github.com/mlcommons/inference_policies/blob/master/inference_rules.adoc","note":"MLCommons"},{"title":"Dataset Cards","url":"https://huggingface.co/docs/hub/datasets-cards","note":"Hugging Face"},{"title":"Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI","url":"https://research.google/pubs/data-cards-purposeful-and-transparent-dataset-documentation-for-responsible-ai/","note":"Google Research"},{"title":"SPDX Specification Dataset Profile","url":"https://spdx.github.io/spdx-spec/v3.0.1/model/Dataset/Dataset/","note":"SPDX"},{"title":"SKOS Simple Knowledge Organization System Reference","url":"https://www.w3.org/TR/skos-reference/","note":"World Wide Web Consortium"},{"title":"Date and Time on the Internet: Timestamps","url":"https://www.rfc-editor.org/info/rfc3339/","note":"Internet Engineering Task Force"}],"openQuestions":["Approve or reject the WM-DAT-001 parent boundary and register evaluation-run, model, result, leaderboard, policy, provenance and records relations.","Create task, modality, domain, population, language, culture, accessibility, psychometric and scientific validity profiles.","Validate jurisdiction and organization-specific source rights, license, consent, IP, export, privacy, security, hidden-test disclosure, safety and retention policies.","Test release-pinned DCAT, Croissant, SPDX, Hugging Face, MLPerf, PROV and SKOS mappings with conformance, round-trip and information-loss evidence.","Refresh NIST AI RMF and mutable MLPerf policy mappings after normative changes and obtain supplemental independent external review before canonical promotion.","Claude and Grok each timed out on one bounded attempt; no independent external result was admitted.","The unified parent_ids WM-DAT-001 has no approved frozen extension edge; this result records only a non-owning reference.","The incoming candidate WM-AI-003 reference grants no ownership of evaluation runs, results, leaderboards or claims.","Task-specific psychometric, scientific, cultural, accessibility and domain-validity profiles require separate research."],"resources":{"spec":"/models/wm-ai-009-evaluation-dataset-benchmark/spec.yaml","agents":"/models/wm-ai-009-evaluation-dataset-benchmark/AGENTS.md","source":"https://github.com/ver-cy/world-models/tree/feat/mega-model-registry/publications/wm-ai-009-evaluation-dataset-benchmark"},"provenance":{"origin":"world-models research","builtFrom":["models/wm-ai-009-evaluation-dataset-benchmark/spec.yaml","ver-cy/world-models/card-supplements/wm-ai-009-evaluation-dataset-benchmark.json"],"providers":["Codex"],"researchStatus":"reviewable-draft","generatedAt":"2026-09-06T12:06:45Z","builder":"tools/build_cards.py@1.0.0"},"completeness":{"sections":{"classifiers":"filled","whatItIs":"filled","purpose":"filled","distinguishingFeatures":"filled","structure":"filled","agentConduct":"filled","ethics":"filled","owners":"filled","relations":"filled","interaction.identity":"filled","interaction.properties":"not-applicable","interaction.recognition":"filled","interaction.capabilities":"filled","interaction.hazards":"filled","interaction.interfaces":"filled","interaction.context":"filled","sources":"filled"},"notes":{"interaction.properties":"Institutional or informational subject: no invented physical properties.","_supplement":"Sections authored in card supplement 1.0.0 by Claude (Opus 5.5) (2026-10-05, unreviewed). Written from the published specification and established practice in the field; no new sources were read. Unreviewed."},"score":1.0}}