# External adversarial review of version 1 ## 1. Review protocol Version 1 of this benchmark was submitted to an adversarial external review before publication. The reviewer was **Codex (GPT-5.2)**, run locally in read-only mode over the instrument directory, with a single instruction: act as a hostile and competent methodologist, rank findings by how severely each threatens the stated conclusion, and cite file and line for each. The review returned two findings that changed the result and five that changed what could be claimed or how it was worded. The version 1 figures were discarded rather than corrected in place, because two of the findings affected the instrument itself and not only its interpretation. This document records each finding, its evidential basis, the remedy applied, and the effect on the conclusions. A benchmark whose failures are not published is a marketing artefact rather than a measurement. Findings are classified by the type of validity they threaten, following the standard four-way division used in section 11 of the methods. ## 2. Summary | ID | Validity threatened | Finding | Disposition | |---|---|---|---| | R1 | Construct | The manipulation confounded representation with completeness | Remedied by adding a condition; original hypothesis refuted | | R2 | Internal | The cause of the federation decrement was misattributed | Remedied; instrument was at fault, model was correct | | R3 | Construct | A measure was named beyond what it detects | Renamed and bounded | | R4 | Statistical conclusion | Significance computed over non-independent observations | Reanalysed; one claimed effect did not survive | | R5 | Construct | Criterion circularity | Accepted as a limitation, not remedied | | R6 | Internal | Item-level defects in four places | All remedied and re-verified | | R7 | External | Claim stated more broadly than the design supports | Claim narrowed | ## 3. Findings ### R1. The manipulation confounded representation with completeness **Finding.** The prose condition did not contain facts that the structured condition contained: not the lookback window then in force, not the date on which it changed, not the holiday calendars, and not the counterparty's valuation formula. **Evidence.** Direct comparison of the two context blocks. No model infers an arbitrary 180-day policy from transaction rows, so the missing facts were not recoverable from the data held constant across conditions. **Assessment.** The comparison established that a complete specification outperforms an incomplete one, which was not the claim being made. The claim that versioning helps in itself was not identified by the design at all, because no otherwise-equivalent unversioned condition existed. **Remedy.** Condition E was added: complete prose carrying exactly the facts of condition C. This makes representation the only variable in the C against E contrast. **Effect on conclusions.** Decisive and adverse. Condition E matched condition C to the decimal at p = 1.00. Hypothesis H5 was refuted and the version 1 headline claim did not survive. ### R2. The cause of the federation decrement was misattributed **Finding.** Version 1 reported that the federation condition lost accuracy through arithmetic degradation caused by a longer context. That attribution was wrong. **Evidence.** The reviewer read the raw responses. On two items the model returned exactly the value obtained by applying the counterparty's rule to that counterparty's rows and our own rule to the remainder, and stated as much in its own explanatory note. The behaviour was rule application, not arithmetic failure. **Assessment.** The defect was in the instrument. The federation rule read "applies to any question about contract C-118 or supplier HAL", so internal reporting questions fell legitimately within its scope, while the reference answers applied our own definition throughout. Contract and supplier were also generated independently, so the positive and negative scopes of the rule overlapped. The model followed the rule it was given; the reference answer did not. **Remedy.** Federation records now apply to counterparty-facing questions only and say so explicitly. The corpus was regenerated so that scopes cannot overlap. Condition F was added to separate context length from context content, which the original explanation had conflated. **Effect on conclusions.** With the defect corrected, the federation condition is neutral overall and helps on the questions it is intended for. The length-based explanation was withdrawn entirely; condition F shows no accuracy penalty for added length at these sizes. ### R3. A measure was named beyond what it detects **Finding.** The measure labelled "used the wrong definition" detects only exact equality with one prespecified alternative answer. **Evidence.** Inspection of the scoring function. **Assessment.** The measure cannot detect the mixed application of two rules, an incorrect row filter, or a wrong definition followed by an arithmetic error, and an arithmetic error can coincide with it by chance. The accompanying claim, that arithmetic errors produce arbitrary values while interpretation errors produce the exact alternative, does not follow. **Remedy.** Renamed to "the answer is exactly equal to the prespecified competing definition's answer", with its limits stated wherever it appears. It is now described as a lower bound on interpretation error. **Effect on conclusions.** No numeric change. The measure's direction and magnitude stand; only the strength of the inference drawn from it was reduced. ### R4. Significance was computed over non-independent observations **Finding.** Three replicates of one item under one prompt were treated as three independent observations, and the test pooled all calls. **Evidence.** Reanalysis at item level, performed independently by the reviewer and reproduced by us. | Analysis | structured against federation | |---|---| | per call, as published in version 1 | p = 0.027 | | exact sign test over 34 item means | p = 0.082 | | McNemar on the majority of three per item | p = 0.057 | **Assessment.** The 14.7-point difference between the two conditions is unchanged by the correction; its significance is not. Under the corrected analysis it does not reach the conventional threshold. The comparisons that carry the study's main claim, structured against schema and against incomplete documentation, survive the correction at p = 0.0005 and p = 0.0015 respectively. **Remedy.** Replicates are averaged within an item, paired tests run across items using the exact sign test, and confidence intervals are clustered by item. **Effect on conclusions.** One claimed effect was withdrawn. The principal effect was unaffected in magnitude and remains significant. ### R5. Criterion circularity **Finding.** The reference answers are derived from the same definitions that the structured condition receives. **Evidence.** Inspection of the ground-truth engine and the prompt template. **Assessment.** There is no answer leakage: the prompt interpolates data, context, question and date, and never the reference answer. But the structured condition receives what is close to an executable specification, so selecting the interpretation reduces largely to retrieval, and the residual work is the join and the aggregation. Success against an author-defined criterion therefore does not demonstrate that the layer discovers correct business meaning. **Remedy.** None. The limitation is stated in the methods and on the results page. **Effect on conclusions.** The scope of every accuracy claim is narrowed to adherence to a supplied specification. Independently authored policies with blind adjudication remain the appropriate next step. ### R6. Item-level defects **Finding.** Four defects in the item set and corpus. 1. One scope discriminator did not discriminate: the correct answer was also reachable through a partially incorrect application of the rule. 2. The business-day counting convention was never stated, so an interpretation of the interval endpoints could be mistaken for a choice of calendar. 3. The temporal family ran on two measures only, and four items shared a single answer value. 4. Orders were generated with dates preceding their customers' signup dates. **Remedy.** The discriminator was moved to a contract under the same supplier but outside the rule; the counting convention is now stated in the question; a third versioned measure was added; and order dates are now constrained at generation. Re-checked after regeneration, with zero violations. **Effect on conclusions.** Item-level noise reduced. No published figure from version 2 depends on the defective items, which no longer exist. ### R7. The claim was stated more broadly than the design supports **Finding.** Tools were disabled throughout, so the study measures reasoning over tabular data held in context. **Assessment.** The results do not describe an operational agent with database access. With SQL available, the arithmetic component of the error changes and the balance between conditions may shift. **Remedy.** The defensible claim is now stated narrowly wherever results appear: on this item set, supplying a precise definition catalogue improves this model's exact answers relative to absent or incomplete documentation. ## 4. What the review did not overturn The advantage of supplying definitions at all, relative to the schema alone and relative to incomplete documentation, survived the corrected analysis by several orders of magnitude in p, and the reviewer confirmed this with an independent recomputation. The disagreement was never about whether the improvement existed. It was about what to attribute it to, and condition E answered that question against the original hypothesis. ## 5. Open after version 2 - One synthetic corpus. Several independently generated corpora are needed for external validity. - Policies authored by the study's authors. Independently written policies with blind adjudication are needed to address R5. - One model family at two capability levels. Other providers are untested. - No run with tools enabled.