← Back to catalogue
Published

AI Inference / Agent Run

vr.wm-ai-004 · wm-ai-004-ai-inference-agent-run

Provide the format-neutral context structure an AI agent needs to open, execute, observe, close, audit, cost and retain a single AI inference or agent run, so that what was asked, what was configured, what was invoked, what was produced, at whose authority, at what cost and with what evidence can be reconstructed after the fact.

World Models Information and virtual systems INF.AI.RUN

Bundle → Layer → Finding → Questions Filled

7 bundles · 15 layers · 30 findings · 115 questions

Run identity and classification What this run is, how it is named, how it is correlated to surrounding executions, and which operational and regulatory class it falls into.

Identity and correlation

The authoritative identifier for the run and the correlation keys that place it inside a trace, a conversation and a chain of related runs.

Authoritative run identifier

Which identifier authoritatively names this run, which system issued it, how it behaves under retries and asynchronous retrieval, and which vendor-native identifiers are kept alongside it. OpenTelemetry marks gen_ai.response.id as Required only for fetch_response operations, Anthropic returns a msg_ prefixed message id, and Amazon Bedrock records a requestId; none of these is a governed global identifier, so the adopting Dimension must state its priority order.

  1. Which system is the master of record for this run's identifier, and what is the exact literal form of that identifier? identity
  2. When the executing provider returns no run identifier, what identifier does the adopting Dimension mint, and how is uniqueness guaranteed? identity
  3. Does the identifier stay stable across retries, streaming resumption and later asynchronous result retrieval, or does each attempt receive a new one? constraint
  4. Which vendor-native identifiers are retained beside the authoritative identifier, and what is each one good for? interoperability

Trace, conversation and parent-run correlation

The keys that place the run inside a wider execution graph: W3C trace-id and span-id, gen_ai.conversation.id for session grouping, and the parent-run edge for delegated executions. These are correlation, not identity: trace-ids may be unsampled or absent, and a conversation groups many runs.

  1. Which trace identifier and span identifier bind this run into the surrounding distributed trace, and in which propagation format? relationship
  2. Which conversation, session or thread identifier groups this run with the runs that preceded and followed it? relationship
  3. How are correlation keys preserved when the run crosses a process, vendor or network boundary? interoperability
  4. What is recorded when the trace is not sampled, telemetry is dropped, or no conversation exists for this operation type? exception

Operation typing and risk classification

The operational shape of the run and the regulatory or risk class that determines which obligations attach to it.

Operation type and invocation mode

Which operation the run represents and how it was invoked. OpenTelemetry enumerates chat, embeddings, retrieval, fetch_response, create_agent, invoke_agent, invoke_workflow, plan and execute_tool; the run must also record whether execution was synchronous, streaming, batched or asynchronously retrieved, because that changes where the run's temporal boundaries lie.

  1. Which operation type does this run represent, drawn from a closed and versioned code list? classification
  2. Was the run synchronous, streaming, batched or asynchronously retrieved, and where exactly do its temporal boundaries fall? process
  3. Does this run represent a single model call or an agentic loop that may contain many model and tool calls? composition

Regulatory and oversight classification

Whether the AI system executing this run falls into a regulated class, which minimum log content that triggers, and what level of human oversight is mandated. EU AI Act Article 12 obligations attach only to high-risk systems, and the Article 12(3) minimum log set applies only to Annex III point 1(a) remote biometric identification, so over-application is a real modelling error.

  1. Under which regulatory regime and which risk category is the AI system performing this run classified, and by whose determination? classification
  2. Does the applicable regime impose a minimum log content set for runs of this class, and is every mandated element actually present? requirement
  3. What degree of human oversight must be exercised over runs of this class before their effects take hold? authority
  4. Which jurisdictions' obligations apply to this run, given where it executed and where its effects land? spatial
Invocation context and execution binding Everything that was fixed at the moment of invocation: what was sent in, where it came from and how far it was trusted, which model and provider served the run, which decoding parameters applied, and which capability surface was exposed.

Request content and input trust

The content actually presented to the model, how much of it is captured verbatim, and what is known about where each segment came from.

Input messages, system instructions and payload handling

What was sent to the model - system instructions, chat history, prompt variables and attachments - and under what capture rule. OpenTelemetry marks gen_ai.input.messages and gen_ai.system_instructions as opt-in with explicit sensitive-data warnings, and Bedrock externalises any body over 100 KB to object storage, so the run record must distinguish inline content from a reference plus digest.

  1. What exactly was presented to the model as system instructions, chat history, prompt variables and attachments? composition
  2. Is verbatim input content captured, and under which opt-in, truncation, filtering or sampling rule was that decided? privacy
  3. Where content is too large or is binary, how is it externalised and referenced back from the run record? constraint
  4. How can the exact input be reconstructed later if the referenced payload store has since changed? validation

Input segment provenance and trust level

For each segment of the model's context, where it came from and how far it may be trusted. MCP states that tool annotations and tool results must be treated as untrusted unless the server is trusted, and directs clients to validate tool results before passing them to the model; a run record that flattens user text, retrieved documents and tool output into one undifferentiated context loses the evidence needed to investigate an injection.

  1. For each input segment, what is its origin - end user, retrieved document, tool result, upstream agent or system operator? provenance
  2. Which input segments were treated as untrusted data that must never be interpreted as instructions? security
  3. How is a suspected prompt-injection or untrusted-instruction attempt recorded against the run? event

Model, parameter and capability binding

The exact execution substrate the run was bound to: model and provider, decoding parameters and configuration snapshot, and the tools and data sources made available.

Model, provider and endpoint binding

Which model was requested, which model actually answered, and through which provider, endpoint and region. OpenTelemetry keeps gen_ai.request.model and gen_ai.response.model separate precisely because a gateway or alias may resolve to a different model, and Bedrock records both a modelId and the serving region.

  1. Which model was requested and which model identifier actually produced the response? identity
  2. Which provider, endpoint address and serving region handled the run? spatial
  3. Which model version, snapshot or inference profile is recorded so a later reader knows which weights answered? provenance

Decoding parameters and reproducibility limits

Which sampling and decoding parameters were in force and what actually prevents an identical re-execution. Anthropic marks temperature, top_p and top_k as deprecated for newer models, so a model that assumes these are always present and always determinative will mis-record reproducibility. The run references an immutable configuration snapshot rather than restating its content.

  1. Which sampling and decoding parameters were actually transmitted and honoured for this run? constraint
  2. Can this run be re-executed to produce an identical result, and what specifically prevents bit-identical reproduction? validation
  3. How is a parameter recorded when the provider has deprecated, capped or silently ignored it for the selected model? exception
  4. Which configuration snapshot does the run reference, and is that snapshot immutable and content-addressed? provenance

Capability surface and granted scopes

Which tools, resources and data sources the model could reach at each point in the run, and under which authorization scopes. MCP requires tool sets to be deterministic per authorization but permits them to change over time via list-changed notifications, and requires RFC 8707 resource indicators binding tokens to a specific server, so the surface is a time-varying, scope-bounded fact rather than a static list.

  1. Which tools, resources and data sources were offered to the model at run start, and in which order were they presented? composition
  2. Which authorization scopes and resource indicators bounded what the run could reach? authority
  3. Did the available capability set change during the run, and how is each change point recorded? state

Executing agent and exact configuration references

Invoke and create agent spans record gen_ai.agent.id, name, description and version. For hosted agents the id should be the provider-assigned stable identifier (Bedrock ARN, GCP registry id, OpenAI assistant id). In-memory instance ids are not recommended. PROV associates the activity with an agent and optionally a plan. OpenInference records agent.name. The run also records the configuration identifiers needed for reproducibility (prompt version, tool definitions version, policy version, decoding profile) as references to WM-AI-005 rather than embedding the configuration catalogue.

  1. Which stable executing agent identifier, name and version performed this run? relationship
  2. Which PROV association, attribution or delegation links this activity to the responsible agent, organisation or human overseer? provenance
  3. Which exact configuration identifiers were in force for prompts, tools, decoding, routing and safety policy? relationship
  4. Which organisation, deployer or system owner is accountable for this execution? ownership
Execution trajectory What the run actually did, step by step: its internal decomposition and deliberation, its interactions with tools and retrieved data, and any work it delegated to sub-runs.

Step trajectory and deliberation

How the run decomposes into ordered, nested steps, and how internal reasoning and planning output is treated.

Step decomposition and deliberation capture

The ordered, nested set of steps making up the run, and the treatment of intermediate reasoning. OpenTelemetry defines plan spans for task decomposition and counts inference_calls and tool_calls per agent invocation; Anthropic returns thinking blocks as first-class content. Reasoning content is high-sensitivity and often subject to a shorter retention rule than the final output, so its capture policy must be recorded rather than assumed.

  1. How is the run decomposed into ordered steps, and what identifies each step uniquely within the run? composition
  2. What parent-child relationships hold between orchestration steps, model-call steps and tool-call steps? relationship
  3. Are intermediate reasoning or planning outputs captured, and under which disclosure and retention rule? retention
  4. How many model calls and tool calls did the run make, and were any iteration limits reached? measurement

Tool and environment interaction

Calls the run made out to tools, servers and knowledge sources, and the evidence retained about each.

Tool call record and effect character

For each tool invocation: what was called, with what arguments, what came back, whether it failed, and what side-effect character the tool declared. MCP separates protocol errors (JSON-RPC error) from tool execution errors (isError true in the result) and warns that annotations are untrusted unless the server is trusted, so the record must keep the declaration and the trust judgement apart. Retries, timeouts and multi-round-trip input-required flows must be linked to one logical call.

  1. For each tool call, which tool on which server was invoked, with which arguments, and what did it return? process
  2. Did the call fail at protocol level or return an execution error, and how is that distinction preserved in the record? exception
  3. What side-effect character did the tool declare - read-only, destructive, idempotent, open-world - and was that declaration treated as trusted? security
  4. How are retries, timeouts, cancellations and input-required round trips of the same logical call linked together? state

Retrieval and grounding evidence

Which knowledge sources were consulted and which retrieved items entered the model context. OpenTelemetry defines gen_ai.data_source.id for retrieval and grounding, but nothing normative pins the retrieved item to a resolvable version, so the run must record item references with digests and retrieval instants if the grounding is to be re-examined later.

  1. Which data sources or knowledge bases were queried to ground this run's output? relationship
  2. Which specific retrieved items entered the model context, and can each be re-resolved at the version that was used? evidence
  3. How fresh and how authoritative was the retrieved content at the instant it was used? quality

Delegation and sub-runs

Work this run handed to other agents or nested runs, and how authority, cost and outcome flow across the delegation boundary.

Sub-run and multi-agent delegation

Which sub-runs the run spawned, on whose authority each acted, and how their results roll up. OpenTelemetry provides invoke_workflow for coordinated multi-agent processes but defines no rollup semantics; PROV-O supplies actedOnBehalfOf for the responsibility chain. Unbounded recursion and authority widening across hops are the two failure modes the record has to make visible.

  1. Which sub-runs or delegated agent invocations did this run spawn, and in which order? composition
  2. On whose authority did each sub-agent act, and was the granted authority narrowed at every delegation hop? authority
  3. How are the cost, failure and output of a sub-run rolled up into the parent without double counting? measurement
  4. What terminates a delegation chain, and what prevents unbounded recursion or fan-out? constraint
Outcome, state and time How the run progressed through its states, why it ended, when everything happened, and what it produced.

Lifecycle state and temporal semantics

The run's state machine, its termination outcome, and the time model that governs every instant recorded on it.

Run state machine and termination outcome

The legal states of a run and how it ended. OpenTelemetry records gen_ai.response.finish_reasons as an array and a Stable error.type when an operation ends in error; Anthropic returns stop_reason values end_turn, max_tokens and stop_sequence. A truncated run is neither a clean success nor a failure, and conflating the two destroys audit value.

  1. What is the complete set of run states, and which transitions between them are legal? lifecycle
  2. In which state did the run terminate, and for what stated reason? state
  3. How is a partially completed or truncated run distinguished from an outright failure for audit purposes? classification
  4. Which low-cardinality error type is recorded when the run ends in error, and where does the full diagnostic live? exception

Temporal semantics and latency measures

Every instant on the run, in a single governed format, with event time kept apart from observation, ingestion and billing-bucket time. RFC 3339 requires an explicit numeric offset or Z and reserves -00:00 for an unknown local offset. Anthropic's usage reporting is bucketed with a stated freshness lag, so cost timestamps are demonstrably not the same clock as execution timestamps.

  1. What are the run's start and end instants, and in which format and offset are they expressed? temporal
  2. How is the instant at which the run occurred distinguished from the instants at which it was observed, ingested and billed? provenance
  3. Which latency measures are recorded, and from whose vantage point are they measured? measurement
  4. Which clock source produced these instants, and what skew is tolerated before ordering becomes unreliable? constraint

Retries, cancellation, compaction and fetch-without-inference

Rare but specified endings: automatic retries folded into one logical span; cancellation or timeout; gen_ai.conversation.compacted indicating the effective context is a compacted view of a prior conversation and must not be set to false; and fetch_response, which loads a previously generated response by identifier without performing inference and must not report token usage. These cases are easy to omit from naive chat-only schemas.

  1. Were transient automatic retries folded into this logical run, and how many attempts occurred? process
  2. Was the run cancelled or terminated by timeout before a model finish reason was produced? lifecycle
  3. Was the effective conversation context a compacted view of a prior conversation, and was the compacted attribute left unset rather than set to false when unknown? constraint
  4. Is this a fetch of a stored response without inference, and were token usage attributes and metrics omitted as required? classification

Outputs and generated assets

What the run returned, how that return is evidenced, and what provenance travels with any media it produced.

Output content, structure and evidence

The content blocks the run produced, whether they validated against a declared schema, and how they are digested and referenced. MCP requires servers to make structured results conform to a declared outputSchema and clients to validate them; Bedrock externalises outputs above 100 KB. Which parts a human actually saw is a distinct and audit-relevant fact.

  1. Which content blocks did the run produce, in which modalities and in which order? composition
  2. Was the output schema-constrained, and did it validate against the declared output schema? validation
  3. How is the output stored, digested and referenced so a later reader can prove what was actually returned? evidence
  4. Which parts of the output were presented to a human, and which were consumed only by downstream machinery? access

Generated asset provenance and AI disclosure

For media assets the run produced, whether a signed content-provenance manifest is attached and what it asserts. C2PA 2.2 binds assertions and a claim through a claim signature with a hard binding to the content, and the c2pa.actions assertion carries a digitalSourceType disclosing generative-AI involvement. No normative C2PA field carries a run identifier, so the binding back to this model is an adopting-Dimension convention.

  1. For each media asset produced, is a signed provenance manifest attached, and which assertions does it carry? provenance
  2. Is the asset marked as AI-generated using a recognised digital source type, and does that marking survive downstream processing? interoperability
  3. How is the manifest bound back to this run and to the ingredients that were consumed to create the asset? relationship
Accountability, authority and provenance Who authorised the run, who intervened in it, and to whom its activity, entities and effects are attributed.

Actors, authority and oversight

The principals behind the run, the scope they conferred, and the human decisions taken during execution.

Principals, credentials and delegation of rights

Which principal initiated the run and whose credentials were presented to each downstream resource. MCP requires tokens to be validated as issued specifically for the receiving server and forbids servers accepting or transiting foreign tokens; Bedrock records identity.arn for the calling principal. The record captures who and with what scope, never the credential material itself.

  1. Which principal initiated this run, and whose credentials were presented against each downstream resource? authority
  2. Which scopes were granted, and was each access token bound to the specific resource it was used against? security
  3. Where the agent acted on behalf of a person or organisation, how is that on-behalf-of chain recorded? provenance
  4. How is an end user identified in the run record without persisting direct personal identifiers? privacy

Human oversight, approval and intervention

Human decisions taken inside the run. MCP states there SHOULD always be a human able to deny tool invocations and that clients should show tool inputs before calling; EU AI Act Article 12(3)(d) requires identifying the natural persons involved in verifying results for Annex III point 1(a) systems. The record must capture what the decider was shown, not only what they decided.

  1. Which actions in this run required explicit human approval before they could execute? decision
  2. Who approved or denied each gated action, at what instant, and what information were they shown at the time? evidence
  3. Was the run paused, stopped or overridden by a human, and what effect did that have on the trajectory? event
  4. How is a missing but required approval detected, escalated and recorded as a non-conformity? exception

Provenance and accountable party

Formal provenance typing of the run and the identification of who owns the record and who answers for its effects.

Provenance attribution, ownership and accountability

How the run maps onto a formal provenance model and who is answerable. PROV-O types the run as an Activity with used and wasGeneratedBy edges to its inputs and outputs and wasAssociatedWith to its agent. Separately, EU AI Act Article 19 puts the log-keeping duty on the provider for logs under their control, which differs from the deployer who is answerable for effects, so record owner and accountable party are two distinct fields.

  1. Which provenance classes and properties does this run map to, and what are the resolvable IRIs for its activity and agents? provenance
  2. Who owns the run record, and who is the accountable deployer answerable for the run's effects? ownership
  3. Which outputs were derived from which inputs and retrieved items, and is that derivation edge explicit rather than inferred? relationship
  4. Under which tenant, workspace, project or cost centre is this run attributed? classification
Measurement, cost and quality evidence What the run consumed, what it cost, how good its output was judged to be, and which policy checks it passed or failed.

Consumption and cost

Resource consumption attributable to the run and the monetary cost derived from it.

Token and resource consumption

What the run consumed. OpenTelemetry defines gen_ai.client.token.usage with token types input and output only, while providers report a richer set including cache-creation, cache-read and reasoning tokens, plus server-side tool use. Provider-reported counts and locally computed counts routinely disagree, so the reconciliation rule must be explicit.

  1. How many input, output, cache-read, cache-creation and reasoning tokens did the run consume? measurement
  2. Which non-token resources were consumed, and how is that set enumerated for this operation type? composition
  3. How are provider-reported counts reconciled with locally computed counts when the two disagree? validation
  4. Into which measurement time bucket is usage assigned when the run spans a bucket boundary? temporal

Cost attribution and reconciliation

What the run cost and how confident that figure is. No GenAI telemetry standard defines a cost metric, and authoritative cost is published as daily aggregates in a single currency with a data-freshness lag and known exclusions such as priority-tier pricing. Per-run cost is therefore an estimate until reconciled, and the record must say which it is.

  1. What monetary cost is attributed to this run, in which currency and against which price basis? measurement
  2. Is the recorded cost a local estimate or an authoritative billed amount, and how is that status flagged and later upgraded? provenance
  3. To which API key, workspace, service tier and cost centre is this run's cost charged? ownership
  4. How are sub-run costs and cached-token savings allocated so that nothing is counted twice? constraint

Quality and safety evidence

Evaluations, feedback and policy decisions that qualify how good and how safe this run was.

Evaluation results and outcome quality

Assessments bound to the run. OpenTelemetry defines gen_ai.evaluation.result as its own event, capable of being emitted independently of the trace, which means evaluation is frequently asynchronous and arrives after the run has closed. Human feedback is a distinct signal from automated evaluation and must not be merged into the same score field.

  1. Which evaluations were performed against this run's output, by which evaluator and at which evaluator version? quality
  2. What score, label or verdict did each evaluation produce, and on what scale is it interpretable? measurement
  3. Was the evaluation performed inline during the run or asynchronously afterwards, and how late may it still arrive? temporal
  4. How is human feedback on the run captured and kept distinguishable from automated evaluation? evidence

Guardrail and policy decisions

Policy checks applied to inputs and outputs, their verdicts, and the policy version in force. A blocked or refused run creates a genuine tension: the record must prove the block occurred without persisting the prohibited content, so the model records decision metadata and digests rather than the content itself.

  1. Which policy or guardrail checks were applied to this run's inputs, intermediate steps and outputs? security
  2. What did each check decide - allow, transform or block - and what exactly was modified? decision
  3. Which policy version was in force at the instant of each decision? provenance
  4. How is a blocked or refused run evidenced without persisting the prohibited content itself? privacy
Governance of the run record How the record of the run is retained, protected, exposed and exported once the run itself is over.

Retention and record integrity

How long the record is kept, under what basis, how it is disposed of, and how alteration is made detectable.

Statutory log content, retention and disposal

The regulated content and lifespan of the record. Article 19 requires providers to keep automatically generated logs under their control for a period appropriate to the intended purpose and at least six months unless other Union or national law provides otherwise, with financial institutions keeping them under sectoral rules. That floor collides with data-minimisation and erasure duties, so the reconciliation rule must be recorded rather than left to the operator.

  1. Which regulated minimum log elements must this run's record contain, and is each one demonstrably present? requirement
  2. For how long must the record be kept, counted from which start point, and on which legal basis? retention
  3. Who holds the logs when provider and deployer differ, and how is the phrase under their control demonstrated? ownership
  4. How are erasure requests, legal holds and statutory minimum-retention duties reconciled when they conflict? exception
  5. What exactly is destroyed at end of life, and what tombstone remains to prove the record once existed? lifecycle

Record integrity and tamper evidence

How a closed run record is made alteration-evident. C2PA demonstrates the pattern that matters here: assertions gathered into a claim, hashed, and covered by a claim signature verifiable against a trust anchor. A run record without a seal cannot support an accountability claim, and corrections must be recorded as new assertions rather than overwrites.

  1. How is the run record sealed at close so that any later alteration becomes detectable? security
  2. Which fields become immutable when the run closes, and which may still legitimately be appended afterwards? constraint
  3. How is a correction recorded without overwriting the original assertion? lifecycle
  4. Who can verify the seal, against which trust anchor, and what happens when verification fails? validation

Access control and interoperability

Who may see which parts of the record, and how the record moves into and out of external schemas.

Access scoping and confidentiality of run content

Separating access to run metadata from access to run content. OpenTelemetry warns explicitly that input messages, output messages, system instructions, prompt variables and tool definitions are likely to contain sensitive and personal data, which makes field-level rather than record-level access control the default position. Residency constraints on the record may differ from those on the inference itself.

  1. Who may read run inputs and outputs, as opposed to run metadata and measures only? access
  2. Which fields are classified as sensitive, and what masking or truncation applies by default? privacy
  3. Which accesses to the run record must themselves be logged, and what does that access log capture? security
  4. Which residency or cross-border constraints bind the record, and how do they differ from the constraints on the inference itself? spatial

Export profiles and schema mapping

Projecting the run into external schemas and importing foreign run records without silently mis-stating their origin. Because gen_ai.* attributes are still in Development status, and providers differ in identifiers and usage fields, every mapping is provisional and lossy in known ways; the model records the loss rather than claiming conformance.

  1. Into which external schemas can a run record be projected, and which profile governs each projection? interoperability
  2. Which fields have no target in a given external schema, and how is that loss declared to the consumer? constraint
  3. How are vendor-specific attributes carried without polluting the canonical model? definition
  4. How is an imported third-party run record marked as externally asserted rather than locally observed? provenance

Classifiers Filled

Family
World Models
Category
Information and virtual systems
Entry kind
event
Navigation path
NAV.INF.AI.RUN
Domain
INF.AI.RUN
Industry
Cross-industry
Tags
aiinferenceagentruninf.ai.run

What it is Filled

One bounded execution occurrence of an AI system: from the moment an invocation is accepted until the run reaches a terminal state, plus the durable record of that occurrence. The run is an occurrent (a PROV Activity) that references, but does not define, the agent that performed it, the configuration it was bound to, or the AI system it belongs to. It covers single model calls, agentic loops, orchestrated multi-agent workflows and asynchronous or batched executions whose result is retrieved later. It is storage- and interface-neutral: OTLP spans, JSON log records, Git-tracked files, MongoDB documents and MCP resources are projections of the same semantics.

In scope

  • Run identity, correlation to traces, conversations and parent runs
  • Invocation inputs, system instructions and their trust provenance
  • Binding to a specific model, provider, endpoint and immutable configuration snapshot
  • Capability surface offered to the run (tools, resources, data sources) and its authorization scopes
  • Ordered step trajectory including model calls, planning steps, tool calls and delegated sub-runs
  • Run state machine, termination reason, error classification and partial completion
  • Outputs, structured results and provenance manifests attached to generated assets
  • Principals, on-behalf-of delegation chains and human oversight interventions
  • Token, resource and monetary consumption attributable to the run
  • Evaluation results, guardrail decisions and other quality evidence bound to the run
  • Statutory record-keeping, retention, legal hold, deletion, access scoping and record integrity
  • Export profiles and lossy-mapping declarations to external telemetry and provenance schemas

Out of scope

  • Definition, capabilities, persona and lifecycle of the agent itself (WM-AI-002)
  • Prompt templates, parameter defaults, weights and versioned configuration content (WM-AI-005)
  • AI system registration, model cards, training data and conformity assessment (WM-AI-001)
  • Conversation, session or thread as a container entity that groups many runs
  • Evaluation campaign design, benchmark definition and dataset curation
  • Incident management, post-market monitoring cases and regulatory reporting workflows
  • Pricing schedules, rate cards, invoices and contractual billing terms
  • Person and organisation master data for users, approvers and deployers
  • Compute infrastructure inventory, capacity planning and rate-limit configuration
  • Physical actuation, robot motion planning and safety envelopes

Why it exists Filled

Provide the format-neutral context structure an AI agent needs to open, execute, observe, close, audit, cost and retain a single AI inference or agent run, so that what was asked, what was configured, what was invoked, what was produced, at whose authority, at what cost and with what evidence can be reconstructed after the fact.

Distinguishing features Filled

  • Records one inference or agent run, not the agent, its configuration or a whole conversation.
  • Captures what was asked, configured, invoked and produced with authority, cost and evidence.
  • Unlike a distributed trace, it is a governed audit record with retention and redaction rules.
  • Treats run bodies as immutable; only annotations, redactions and seals are added.

What robots and AI may and may not do Filled

Must not

  • Persist credentials or raw tokens in run records.
  • Capture full inputs and outputs by default without a governed decision.
  • Edit or delete a sealed run body.
  • Report estimated cost as reconciled cost.
  • Treat alignment with tracing conventions as conformance.

Only with a human decision

  • Granting an oversight decision requested during a run.
  • Approving release of run content to third parties.

May

  • Open and close a run with references to agent, configuration and principal.
  • Record each step, tool call and oversight request in order.
  • Attach evaluation results and cost estimates with their confidence.
  • Redact personal content under a recorded legal basis.

Moral aspects Filled

  • Run content often contains personal and confidential data from users.
  • Run records may be needed to explain or contest a decision affecting a person.
  • Retention of prompts and outputs must be limited to what is needed.

Who is affected

  • Users whose inputs are recorded
  • People affected by the run's outputs
  • Operators and auditors

Owners Filled

Steward

A Dimension adopting WM-AI-004 must name an accountable owner for run records who is distinct from the operator that generates them, and must state whether that owner acts as provider, deployer or both for each AI system in scope.

Roles

Model steward
Maintain the canonical structure, code lists and export profiles of WM-AI-004 and version them under the compatibility rules; Adjudicate boundary disputes with WM-AI-001, WM-AI-002 and WM-AI-005 and record the outcome as a boundary note; Re-assess external alignments whenever an aligned convention changes stability status
Run record custodian
Operate the stores holding run records and externalised payloads, and guarantee their immutability and reachability; Compute and apply integrity seals at close and preserve superseded seals across corrections; Execute retention decisions, redactions and disposals, and keep the retention and hold register current
Accountable deployer
Answer for the effects of runs performed under their control and confirm which regulatory class applies; Ensure required human oversight is genuinely available and that gated actions cannot execute unapproved; Approve the content-capture policy and the retention schedule for each run class in scope
Auditor or conformity assessor
Reconstruct past runs and verify that mandated minimum log elements are present; Verify integrity seals against the declared trust anchor and report verification failures; Test that access to content-tier fields is scoped, logged and justified
Privacy and data protection officer
Set and review masking, pseudonymisation and truncation rules for sensitive run content; Adjudicate conflicts between erasure requests and statutory minimum retention duties and record the resolution; Approve cross-border transfers and residency placement of run records
Cost and usage controller
Reconcile estimated per-run costs against authoritative bucketed reporting and publish superseding statements; Maintain allocation rules for sub-runs, cached tokens and shared capacity so nothing is double counted; Attribute consumption to tenants, workspaces and cost centres consistently with the identity rules

Links to other meta-models Filled

child

  • WM-AI-001 - The run belongs to the AI system whose registration, intended purpose and risk classification determine which record-keeping obligations attach to each execution. The system model owns the classification; the run inherits it as a coded reference.

references

  • WM-AI-002 - Every run references the agent that performed it, mirroring the PROV-O separation of Activity from Agent and carrying the observed agent identifier, name and description without redefining the agent.
  • WM-AI-005 - Every run references the immutable configuration snapshot it was bound to, by reference and digest, so reproducibility questions resolve against a single governed snapshot rather than a copy embedded in each run.

aligned

  • OpenTelemetry GenAI semantic conventions (spans, metrics, events) - Field-level alignment for operation naming, provider discrimination, model binding, token usage, agent and tool attributes, and content-capture events. Alignment only: the gen_ai.* attribute set is in Development status, so no conformance is claimed.
  • W3C PROV-O (http://www.w3.org/ns/prov#) - Typing alignment: the run maps to prov:Activity with used, wasGeneratedBy, wasAssociatedWith and actedOnBehalfOf edges, giving a vendor-neutral lineage projection for outputs and delegation chains.
  • W3C Trace Context Level 1 - Correlation alignment: traceparent and tracestate supply the trace and span identifiers that place a run inside a distributed execution, without ever being treated as the run's identity.
  • Regulation (EU) 2024/1689 Articles 12 and 19 - Obligation alignment for the subset of runs performed by high-risk AI systems: automatic event recording over the system lifetime, the Annex III point 1(a) minimum log set, and provider retention of at least six months. Applies conditionally, not universally.
  • Model Context Protocol 2026-07-28 (tools and authorization) - Interaction alignment for tool call structure, the protocol-error versus execution-error distinction, untrusted annotation handling, human-in-the-loop denial, and audience-bound authorization scopes recorded on the run.
  • C2PA Specification 2.2 - Downstream provenance alignment for media assets a run produces: signed manifests, ingredient chains and digitalSourceType disclosure. The run-to-manifest binding is an adopting-Dimension convention because C2PA defines no run identifier field.

neighbor

  • WM-AI-002 AI Agent - The agent is a continuant that bears responsibility; the run is the occurrence it performs. PROV-O separates prov:Agent from prov:Activity and links them with prov:wasAssociatedWith. The run stores only a reference plus the agent identity attributes observed at execution time (gen_ai.agent.id, gen_ai.agent.name); it never redefines the agent.
  • WM-AI-005 AI Configuration - The run references an immutable, content-addressed configuration snapshot. Parameter semantics, defaults and version history belong to the configuration model; the run records only which snapshot was in force and any per-call overrides actually sent.
  • Distributed trace and span - A trace is a correlation mechanism, not the run. OpenTelemetry trace-ids are 16-byte values that may be unsampled, absent, shared across many runs, or restarted at a process boundary; the run keeps trace-id and span-id as correlation keys, never as identity.
  • Conversation, session or thread - gen_ai.conversation.id groups runs that share a history; it is Conditionally Required telemetry, not run identity. A conversation container entity, if modelled, is a sibling aggregate that references runs; embeddings and classification runs have no conversation at all.
  • Evaluation result - OpenTelemetry defines gen_ai.evaluation.result as a separate event that may be emitted asynchronously after the run. The run holds references to evaluation results and their verdicts; evaluator definitions, rubrics and campaign scope stay outside.
  • Content provenance manifest (C2PA) - A C2PA manifest asserts provenance of an asset, not of an execution. It is attached to outputs of the run as a downstream artifact; there is no normative C2PA field that carries a run identifier, so the binding is an adopting-Dimension convention and must be recorded as such.
  • Regulatory logging obligation - EU AI Act Articles 12 and 19 impose automatic-logging and retention duties on providers of high-risk AI systems. The run record is the technical substrate that can satisfy those duties, but the obligation, its addressee and its assessment belong to the AI system governance model, not to this model.
  • MCP server and tool definition - MCP defines the wire contract for tools/list and tools/call. The run records which tool was called with which arguments and what came back; tool schemas, annotations and server capabilities are external declarations that the run snapshots by reference and digest.

parent

  • WM-AI-001

What else AI and robots need to interact with it Filled

Identity and identifiers required Filled

  • Authoritative master-system identifier: the run identifier assigned by the system of record that actually executed the run - for example the provider response or invocation identifier - recorded together with the identity of the issuing system.
  • Governed global identifier or IRI: a resolvable IRI minted in the adopting Dimension's governed namespace, used when several master identifiers compete or when a cross-system lineage anchor is required.
  • UUID or ULID assigned by the adopting Dimension, used only when neither an authoritative master identifier nor a governed IRI exists, and always marked as locally assigned.
  • Never identity: a date or timestamp, a trace identifier, a conversation identifier, a sequence ordinal, a filename or a content digest. These are correlation, ordering or integrity values and are recorded as such.

Direct properties not applicable Not applicable

Not applicable

Institutional or informational subject: no invented physical properties.

Recognition optional Filled

  • A run has a run identifier from the executing system, a start and end, a principal, a configuration reference and steps.
  • Often confused with a conversation thread, a trace span, an agent definition and an evaluation result.

Capabilities and actions required Filled

  • Open run: Accept an invocation, assign the authoritative run identifier, bind the run to a configuration snapshot, capability surface and principal, and record the start instant.
  • Record step: Append one ordered step to the run trajectory: a planning step, a model call, a tool call or a delegated sub-run reference, with its payload references, timings and error disposition.
  • Request oversight decision: Suspend a gated action, present the exact proposed inputs to an authorised human, and record the approval, denial or timeout together with what was shown.
  • Close run: Move the run to a terminal state, record the finish reason or error type, fix the end instant, and finalise consumption counters and completeness classification.
  • Seal run record: Canonicalise the closed run record together with its payload digests, compute a digest and apply a verifiable signature so that later alteration is detectable.
  • Attach evaluation: Bind an evaluation result or human feedback signal to a closed run by reference, without mutating the sealed run record.
  • Reconstruct run: Rebuild the full trajectory, inputs, bindings and outputs of a past run for audit, incident investigation or reproduction, reporting explicitly on any evidence that can no longer be resolved.
  • Redact run content: Remove or mask sensitive content from a run's payloads while preserving the record skeleton, digests and audit structure required by retention duties.
  • Apply retention decision: Evaluate a run record against its retention class, active holds and erasure requests, then hold, redact, transfer or destroy it and register the disposition.
  • Export run: Project a run record into a named external profile, emitting the mapped payload together with its field mapping table and an explicit declaration of unmapped fields.
  • Apply content capture and redaction: Decide whether to store payloads inline, truncated, hashed, redacted or externally, and record the manifest.

Hazards and failure modes required Filled

  • Leakage of personal data or secrets through stored prompts and outputs.
  • Unreconstructable decisions when runs are not retained.
  • Unauthorised tool actions not traced to a principal.

Standards and interfaces required Filled

  • OpenTelemetry and its generative AI semantic conventions.
  • W3C Trace Context.
  • C2PA content provenance manifests.
  • Model Context Protocol for tool invocation.

Context of use required Filled

  • EU AI Act Articles 12 and 19 bind only providers and deployers of high-risk AI systems placed on the Union market; they are not a global baseline and must not be applied to every run.
  • The enumerated minimum log set - period of use, reference database, matching input data, identity of verifying natural persons - applies specifically to Annex III point 1(a) remote biometric identification systems, not to high-risk systems generally.
  • Providers that are financial institutions subject to Union financial services law keep these logs as part of their sectoral documentation, which can override the general regime.
  • NIST AI RMF 1.0 and the Generative AI Profile are voluntary United States guidance with no legal force, and they inform structure here rather than impose obligations.
  • Data-residency and inference-geography dimensions are provider-specific vocabularies with values such as global or a country code; they are not standardised across vendors and cannot be assumed comparable.
  • Currency and cost reporting in the cited first-party source is USD-only with costs expressed in minor units; multi-currency operation requires a conversion policy this model does not supply.
  • Retention floors, erasure rights and legal-hold mechanics vary by jurisdiction; the six-month figure cited is an EU provider minimum and is not a global default.
  • EU logging, oversight and six-month retention are mandatory only where Regulation (EU) 2024/1689 classifies the system as high-risk and the actor is in scope (including certain third-country providers whose output is used in the Union).
  • NIST AI RMF 1.0 and NIST AI 600-1 are voluntary United States frameworks used here as monitoring and documentation alignments, not as legal duties.
  • ISO/IEC 22989 is a paid international terminology standard; only publicly visible definitions (inference as process and result, agent, ML model) are relied upon.
  • OpenInference cost examples assume USD; other currencies require an explicit code and are not specified by OTel.
  • Server address does not by itself prove data-residency region.

Sources Filled

  1. Semantic conventions for generative AI spans - OpenTelemetry Authors (Cloud Native Computing Foundation)
  2. Semantic conventions for generative AI agent spans - OpenTelemetry Authors (Cloud Native Computing Foundation)
  3. Semantic conventions for generative AI metrics - OpenTelemetry Authors (Cloud Native Computing Foundation)
  4. Semantic conventions for generative AI events - OpenTelemetry Authors (Cloud Native Computing Foundation)
  5. Article 12: Record-keeping, Regulation (EU) 2024/1689 (Artificial Intelligence Act) - European Commission - AI Act Service Desk
  6. Article 19: Automatically Generated Logs, EU Artificial Intelligence Act - Future of Life Institute (EU AI Act Explorer)
  7. W3C Trace Context, Level 1 - World Wide Web Consortium (W3C)
  8. PROV-O: The PROV Ontology - World Wide Web Consortium (W3C)
  9. Model Context Protocol specification 2026-07-28: Tools - Model Context Protocol maintainers
  10. Model Context Protocol specification 2026-07-28: Authorization - Model Context Protocol maintainers
  11. RFC 3339: Date and Time on the Internet: Timestamps - Internet Engineering Task Force (IETF)
  12. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 - National Institute of Standards and Technology (NIST), U.S. Department of Commerce
  13. Claude Messages API reference - Anthropic PBC
  14. Usage and Cost API - Anthropic PBC
  15. Monitor model invocation using CloudWatch Logs and Amazon S3 - Amazon Web Services
  16. C2PA Specification, version 2.2 - Coalition for Content Provenance and Authenticity (C2PA)
  17. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act) - European Union
  18. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 - National Institute of Standards and Technology
  19. PROV-DM: The PROV Data Model - World Wide Web Consortium
  20. OpenInference Semantic Conventions - Arize AI
  21. ISO/IEC 22989:2022 Information technology — Artificial intelligence — Artificial intelligence concepts and terminology - ISO/IEC JTC 1/SC 42
  22. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 - National Institute of Standards and Technology
  23. Trace Context - World Wide Web Consortium

Open questions

  • Agent memory-step operations (gen_ai memory create, search, update, delete, upsert): add a scoped memory finding under the tool and environment layer. Deliberately not imported this pass because the carrying grok node also duplicated retrieval, planning and nested-agent structure already in the base.
  • OpenTelemetry MCP semantic conventions: listed by the GenAI repository but ingested by neither provider. Ingest and reconcile MCP-flavoured tool spans against the base tool-call finding, which currently relies on the MCP protocol specification rather than the telemetry convention.
  • Billed-versus-consumed token counting: fold the OpenTelemetry rule that billed counts must be reported when both billed and consumed counts exist into the existing token and resource usage finding, alongside the current provider-reported versus locally-computed reconciliation rule.
  • Encrypted and redacted reasoning-block preservation for stateless replay (message content signatures, encrypted_content, redacted_thinking, tool-call reasoning signatures): evaluate as a refinement to the step deliberation and output content findings, given the tension with the base rule that reasoning content carries a shorter retention rule.
  • Proxy versus upstream vendor provider discrimination: extend the model and provider binding finding so a recorded provider name that identifies a gateway or relay is distinguishable from the upstream model vendor.
  • Sampling versus statutory logging duty: determine, against primary regulatory text, whether a run sampled out of full telemetry can still satisfy an automatic-recording obligation, and where the resulting rule belongs.
  • Emerging alignments neither provider ingested: NIST AI 800-4 post-deployment monitoring (March 2026), ISO/IEC 42001 and ISO/IEC 23894, IETF SCITT AI agent action receipts, and the draft ISO/IEC 22989 generative-AI amendment.
  • Cross-vendor run identity: no governed global identifier exists and both providers fall through to an adopting-Dimension minted identifier. Track whether any standards body publishes a run identifier before the identity priority order is frozen.
  • Training-time provenance, dataset lineage and model card content, which belong to the parent AI system model.
  • Conversation, session or thread as a container entity: the run references a conversation identifier but does not define what a conversation is or how it is bounded.
  • Streaming chunk-level telemetry below the step granularity; time-per-output-chunk is referenced as a measure but individual chunk records are out of scope.
  • Rate limiting, quota and capacity management, which shape whether a run may start but are properties of the serving platform.
  • Incident management, post-market monitoring case files and regulatory notification workflows that consume run records downstream.
  • Pricing schedules, rate cards and invoices; only the derived cost attributable to a run is in scope.
  • Federated or fully on-device runs where no network telemetry exists and neither server address nor provider-side usage counters are obtainable.
  • Physical actuation and robotics safety envelopes, consistent with the registry's zero robotics factor for this model.
  • Person and organisation master data for initiators, approvers and deployers, which is referenced by governed identifier only.
  • OpenTelemetry MCP semantic conventions were listed by the GenAI repository but not ingested in this research budget, so MCP-flavoured tool spans are a gap pending that document.
  • ISO/IEC 42001 AI management-system controls and ISO/IEC 23894 risk management were not fetched and are not claimed.
  • IETF SCITT AI-agent action receipts are an emerging draft, not an adopted alignment.
  • Provider-specific cost APIs, invoices, reservations and spot GPU economics are omitted beyond run-level quantities.
  • Multi-agent A2A directories, long-running unsupervised agent loops without a closable span, and hardware/accelerator traces are omitted.
  • C2PA/watermark packaging of generated media is omitted.
  • NIST AI 800-4 post-deployment monitoring (March 2026) appeared during search but was not fully ingested; treat as a likely future alignment.
  • Amendment 1 to ISO/IEC 22989 for generative AI is still a draft and not used as terminology beyond the published 2022 edition.

Machine files

Provenance

world-models research · reviewable-draft

Built from: models/wm-ai-004-ai-inference-agent-run/spec.yaml, ver-cy/world-models/card-supplements/wm-ai-004-ai-inference-agent-run.json