# Semantic Grounding Benchmark v2: methods Version 2. Run date 2026-09-06. Primary model `claude-haiku-4-5-20251001`, secondary model `claude-sonnet-5`. 846 model calls in the reported study. A further 480 calls were spent on version 1 of the instrument, which an external review invalidated; those data are excluded from every result reported here and are discussed in the review record. ## 1. Purpose An agent asked a quantitative question about an organisation must choose an interpretation before it can compute anything. "How many active customers do we have" is not answerable until "active" is fixed to a rule, and organisations routinely revise such rules and disagree with their counterparties about them. This study asks what an agent gains when those rules are supplied to it, and in particular whether the gain depends on the rules being supplied as structured, versioned records rather than as ordinary documentation. Version 1 of this benchmark tested accuracy alone. That is the narrowest of the claims made for a definition layer, and the review of version 1 established that it was the claim the data least supported. Version 2 therefore measures five capabilities, of which accuracy is one. ## 2. Terms Every term below is used in this document and on the results page in the sense given here, and in no wider sense. | Term | Definition as used in this study | |---|---| | Semantic grounding | The correspondence between the concept a question names and the computation an agent performs in answering it. A response is grounded when that computation implements the definition the asking organisation holds to be in force at the date in question. | | Definition record | A machine-readable statement of one definition, carrying an identifier, a version, the date from which it applies, a reference to the version it supersedes, the role accountable for it, and parameters sufficient to compute it. | | Definition catalogue | The set of definition records supplied in a given condition. | | Prose documentation | The same subject matter written as continuous text for human readers, carrying no identifiers, versions or fields. | | Federation record | A record declaring which of two organisations' definitions governs a class of questions, together with the scope in which that declaration applies. | | Condition | One of several context blocks inserted between the data tables and the question. Conditions differ in that block and in nothing else. | | Item | One question together with its reference answer. Used interchangeably with "task". | | Reference answer | The value returned by the ground-truth implementation of the definition in force. Computed in code and never supplied to the model. | | Competing definition | For a given item, a prespecified alternative rule that a reasonable reader might apply instead. | | Discriminating item | An item for which the competing definition yields a numerically different answer from the reference answer. Verified in code for every item. | | Gap | A field that the catalogue schema requires and that is absent, empty, or that references a record which does not exist. | | Elicitation | Detecting a gap, naming the absent field, and naming the role accountable for supplying it. | | Assembly | Selecting from a catalogue the minimal subset of records that a given question requires. | | Disclosure decision | A judgement whether one information item may be released to one named requesting party. | | Critical disclosure error | Release of customer identities, of discount terms, or of one carrier's operational data to a competitor of that carrier. | ## 3. Hypotheses H1 to H4 were specified before version 1 was run. H5 was added in response to the external review, and H6 to H9 were added with the four capability tracks; all five are therefore post-hoc. The distinction is stated because a hypothesis formed after seeing data carries less evidential weight than one formed before it, and the central negative result of this study rests on H5. | | Hypothesis | Prespecified | |---|---|---| | H1 | Supplying definitions in any form raises answer accuracy above supplying the data schema alone. | yes | | H2 | Incomplete documentation does not raise accuracy, and raises the rate of answers matching a superseded rule. | yes | | H3 | The advantage of definitions is largest on questions whose date precedes a change of definition. | yes | | H4 | Federation records raise accuracy on counterparty-facing questions without lowering it elsewhere. | yes | | H5 | Structured records raise accuracy above prose carrying the same facts. | no | | H6 | Only the structured conditions can identify the definition and version applied. | no | | H7 | Only the structured conditions can locate a gap and name its accountable role. | no | | H8 | Structured classification reduces critical disclosure errors relative to a prose policy stating the same rules. | no | | H9 | Structured accountability raises the rate at which a request reaches the correct owner. | no | ## 4. Materials ### 4.1 Synthetic corpus The corpus is generated deterministically from a fixed seed and describes one supplier, **Meridian Supply Co**: 24 customers, 72 orders, 66 shipments, 56 invoices, 10 returns. Two carriers appear. **Halden Logistics** holds two contracts: C-118, which the federation records govern, and C-301, which they explicitly exclude. C-301 exists so that the scope of a federation rule can be tested against a contract held by the same supplier but lying outside the rule, a case no cue in the data distinguishes. **Zen Freight** holds contract C-204 and lies outside the rules entirely. Generation constraints are enforced and re-checked after generation: no order precedes its customer's signup date, and contract and supplier are assigned consistently, so that the positive and negative scopes of the federation rules do not overlap. ### 4.2 Ambiguous constructs Six constructs carry the ambiguity. Each was chosen because it is a concept over which departments and trading partners hold genuinely different and locally defensible rules. | Construct | Axis of disagreement | Versioning | |---|---|---| | active customer | 365-day or 180-day lookback | v1 to v2 on 2026-04-01 | | net revenue | freight included or excluded, less returns | v1 to v2 on 2026-05-01 | | delivered | customer confirmation or carrier scan | Meridian and Halden differ | | on time | original or renegotiated promise date | v1 to v2 on 2026-02-01 | | order value | freight excluded or included | Meridian and Halden differ | | business day | two holiday calendars | Meridian and Halden differ | ### 4.3 Conditions The data tables, approximately 10,000 characters of CSV, are identical in every condition. Only the context block differs. | Condition | Context block | Function in the design | |---|---|---| | A schema | table definitions only | Floor. What an agent receives when no semantic layer exists. | | B prose, incomplete | a wiki page of the kind organisations actually maintain: no owners, no dates, one paragraph superseded by another, several rules absent | Realistic documentation. | | E prose, complete | the same facts as condition C written as continuous prose: lookback windows, change dates, formulas, accountable roles, both calendars | The control separating representation from completeness. Added after the review of version 1. | | C records | machine-readable definition records with identifier, version, effective date, supersession and owner | The treatment under study. | | F records + padding | condition C plus an inert appendix matched to the federation block in length and position | Separates context length from context content. | | D records + federation | condition C plus federation records | The cross-organisational layer. | Federation records apply to counterparty-facing questions only and state that restriction explicitly. Internal reporting is excluded by name, including when the rows involved belong to the contract concerned. Version 1 omitted that restriction, with the consequence recorded as finding R2 of the review. ## 5. Design A within-item repeated-measures design. Every item is presented under every condition, so each item serves as its own control and variation in difficulty between items does not enter the comparisons. Track A comprises 40 items x 6 conditions x 2 replicates = 480 calls. The five tracks together comprise 846 calls: track A 480, secondary model 160, B1 60, B2 80, B3 18, and B4 48 in 16 chains of three turns. ## 6. Item construction 40 items in four families. - **COMP, 12 items.** Computation under the definition currently in force, framed as internal reporting. - **TEMP, 10 items.** A question whose as-of date precedes a change of definition, across three versioned measures. Two items fall exactly on an effective date, where the governing rule depends on whether the boundary is inclusive. - **XORG, 12 items.** Ten counterparty-facing questions, plus two scope discriminators: contract C-301 under the same supplier, and supplier ZEN, both outside the federation rules. - **ABST, 6 items.** Questions the corpus cannot answer under any definition. The correct response is to decline and to say why. All 34 answerable items were verified in code to be discriminating in the sense defined in section 2. Ambiguities not under study were removed from the item text: every internal question states that it is internal, every counterparty-facing question names the counterparty, and the business-day counting convention is stated in the question rather than left to be inferred. ## 7. Measures on track A | Code | Measure | Denominator | |---|---|---| | M1 | The returned value equals the reference answer. | answerable items | | M2 | The response names both the identifier and the version of the governing definition. | answerable items | | M5 | The response declines on an item the corpus cannot answer. Declining on an answerable item is recorded separately as a false abstention. | unanswerable items | | M7 | The returned value is exactly equal to the answer the prespecified competing definition would produce. | answerable items | Two properties constrain the interpretation of M7. It detects only exact equality with one prespecified alternative, so it does not detect the mixed application of two rules, an incorrect row filter, or a wrong definition followed by an arithmetic error; and an arithmetic error can coincide with it by chance. M7 is therefore a lower bound on interpretation error rather than a measure of it. In version 1 this measure was labelled "used the wrong definition", which asserted more than it detects. M2 is unreachable by construction in conditions without a catalogue, because there are no identifiers or versions available to cite. This is not a defect in the comparison but the finding itself: traceability is a property of the representation and not of the model. ## 8. Tracks B1 to B4 ### 8.1 B1, elicitation Six gaps are introduced into the catalogue: an empty `parameters` block, a missing `effective_from`, a missing `owner`, a `superseded_by` referring to a version that does not exist, a calendar covering only the preceding year, and one definition absent altogether. Each gap is paired with a question that cannot be answered until the gap is filled. Four control questions contain no gap and measure the false-alarm rate. Conditions: `schema`, `prose_deficient` (the same gaps expressed in prose) and `vercy_deficient` (the same gaps expressed in records). Measures: gap detected, absent field named, accountable role named, false-alarm rate, and accuracy on the controls. ### 8.2 B2, assembly The catalogue is enlarged to 41 records totalling 10,332 characters, of which any single question requires one or two. Stage one: the agent selects a minimal subset. Stage two: the question is asked again under three assembly conditions, namely the whole catalogue, the agent's own selection, and a naive keyword top-5 baseline. Measures: selection precision and recall against the known required set, context size, and answer accuracy under each assembly condition. ### 8.3 B3, disclosure control Twelve information items and three requesting parties, namely our own carrier, a competing carrier and a statutory auditor, giving 72 disclosure decisions per condition. Conditions: no policy, the policy as prose, and classification records carrying `classification`, `release_to` and `basis`. The two policy conditions state the same rules and differ only in representation. Measures: release precision and recall, overall decision accuracy, and separately the count of critical disclosure errors. The two error types are reported separately because their costs are asymmetric: an unnecessary refusal costs friction, an unnecessary release can cost a contract. ### 8.4 B4, collaborative completion Three turns per chain, with a coordinator and four role-bound owner agents. The coordinator must address exactly one owner and request a specific field. The owner holds only its own authority and its own fact, must refuse anything outside that authority, and must name the correct owner when it refuses. The coordinator then records the answer as a new version and answers the original question. Measures: routing correctness, owner refusal behaviour, well-formedness of the produced record, correctness of its effective date, and final answer accuracy. ## 9. Procedure Tools were disabled and the agent system prompt was replaced with a neutral one, so that the measured effect is attributable to the context block and not to a coding harness. Responses are a single JSON object against a fixed contract; parse success was 100% in every condition. Call order was randomised across the whole job, so that drift in service behaviour over the run is not confounded with condition. ## 10. Statistical analysis The unit of analysis is the item, not the call. Two replicates of one item under one condition are not independent observations, and treating them as independent inflates the apparent sample size. Replicates are therefore averaged within an item before any test is applied. Paired comparisons between conditions use the exact two-sided sign test over item means. It makes no distributional assumption and suits the bounded, discrete per-item scores this design produces. Confidence intervals are clustered by item. Version 1 applied McNemar's test at the level of individual calls and consequently overstated significance; the correction is quantified as finding R4 of the review. No correction for multiple comparisons is applied. Eight paired comparisons are reported, so a nominal p of 0.05 should not be read as a 5% family-wise error rate. The comparisons that carry the conclusions are significant at p < 0.002 or not significant at p = 1.00, so this choice does not affect any conclusion drawn. ## 11. Threats to validity **Construct validity.** The reference answers are computed from the same definitions that condition C supplies. This is normal for a benchmark of specification execution, but it means the study measures adherence to a specification and not the discovery of correct business meaning. Success against an author-defined criterion shows that an agent applies a supplied rule correctly, not that the layer finds the right rule. The policies were written by the authors. Independently authored policies with blind adjudication would address this and have not been done. **Internal validity.** Conditions differ only in the context block, with data, question text, decoding settings and call order held constant, so differences between conditions are attributable to that block. Condition F separates the length of the block from its content and shows no accuracy penalty for the added length at these sizes. Condition E is approximately 1,000 characters shorter than condition C, because prose is denser than records; if length were harmful this would favour E, so the C against E comparison is conservative in C's direction. **External validity.** One synthetic corpus, generated once. One model family at two capability levels, and no other providers. Tools were disabled, so the results describe an agent reasoning over tabular data held in context, and do not describe an operational agent with database access, where the arithmetic component of the error would change and the balance between conditions might shift. **Statistical conclusion validity.** 34 answerable items and 12 gap trials are small samples. Track B2 has 20 observations per assembly condition, which is why a nominal 20-point difference there is reported as indistinguishable from noise rather than as a finding. Interval estimates are correspondingly wide and are reported beside every point estimate on the results page. ## 12. Availability The generator, the ground-truth engine, the conditions, the item set, the runners and the analysis scripts are published in full, together with the unedited raw runs for every track. The published `result.json` is built directly from those raw runs by a script, so the figures on the results page cannot drift from the data that produced them.