Vercy · reproducible study

Structure did not improve accuracy over complete prose.

Objective
Determine whether supplying an agent with versioned, structured definitions improves its answers, and whether any improvement is attributable to the structure or to the completeness of what is supplied.
Design
One synthetic supplier corpus with six ambiguous constructs whose rules change on known dates and differ between two organisations. 40 items asked under six context conditions in a within-item design, plus four tracks measuring gap detection, context assembly, disclosure control and completion across roles. 846 model calls. The unit of analysis is the item, not the call.
Principal result
Completeness accounted for the accuracy gain and representation did not. Continuous prose carrying the same facts as the structured records scored identically, 76.5% against 76.5%, exact two-sided sign test over 34 items p = 1.00.
Secondary results
The structured conditions were the only ones that cited the rule and version applied (100% against 0%), located a missing definition and named its owner (83% against 0%), routed a request to the correct owner (100% against 50%), and made all 72 disclosure decisions without a critical error (0 against 2).
Conclusion
On this item set, accuracy is a function of what the agent is told rather than of how it is represented. What distinguishes the structured representation is traceability, gap detection, disclosure control and routing, none of which the prose conditions achieved at any level.
76.5%complete prose
76.5%versioned records
p = 1.00no difference

Terms used on this page

TermMeaning in this study
ConditionA context block inserted between the data and the question. Conditions differ in that block and in nothing else.
Definition recordA machine-readable statement of one rule, carrying an identifier, a version, an effective date, the version it supersedes and the accountable role.
Complete proseThe same facts as the records, written as continuous documentation without identifiers, versions or fields.
Discriminating itemAn item for which a plausible competing rule yields a different correct answer. Verified in code for all 34 answerable items.
Competing answerA response exactly equal to the value the prespecified competing rule would produce. A lower bound on interpretation error, not a measure of it.
Critical disclosure errorRelease of customer identities, discount terms, or one carrier's operational data to a competitor of that carrier.

Design

Six constructs carry the ambiguity: active customer, net revenue, delivered, on time, order value and business day. Each was given rules that change on known dates and that the supplier and its carrier define differently. 40 items were then asked under six conditions with the underlying data tables identical throughout, so that any difference between conditions is attributable to the context block alone. Every answerable item was verified in code to be discriminating.

Track A. Accuracy on an agreed meaning

ConditionAccuracy (M1)95% CICited rule and version (M2)Competing answer (M7)
schema only26.5%11-420%16.2%
incomplete documentation29.4%14-440%35.3%
complete prose76.5%65-880%2.9%
versioned records76.5%65-88100%1.5%
records + inert padding85.3%78-93100%0.0%
records + federation rules77.9%67-8998.5%1.5%

Supplying definitions in any form raised accuracy from 26.5% to 76.5%, which survives a paired exact sign test over 34 items at p = 0.0005. The condition carrying the same facts as continuous documentation scored identically to the structured one. Under this design accuracy is determined by the completeness of what is supplied and not by its representation.

Two further results qualify the first. Incomplete documentation performed no better than none and doubled the rate of answers matching a superseded rule, from 16.2% to 35.3%: a document describing a rule no longer in force is worse than no document. And the complete-prose condition, which holds every fact the structured condition holds, cited the governing rule in 0% of responses against 100%. It produces the correct value and cannot state which rule produced it.

Capabilities other than accuracy

Accuracy is one of the five capabilities measured. On the remaining four the prose conditions did not perform worse than the structured ones; they did not perform at all. The comparisons below are therefore differences between presence and absence rather than differences of degree.

B1. Detecting a missing definition

Six gaps were introduced into the catalogue: an empty parameter block, a missing effective date, a missing owner, a reference to a version that does not exist, a calendar ending a year early, and one definition absent altogether. Each was paired with a question that cannot be answered until the gap is filled. Four control questions contained no gap and measure the false-alarm rate.

ConditionGap detectedField namedOwner namedFalse alarms
schema only0%0%0%0%
prose, same gaps0%0%0%0%
records, same gaps83.3%83.3%83.3%0%

Prose containing the same six gaps detected none of them in twelve trials. The structured catalogue detected ten of twelve, naming the absent field and the accountable role in each case, with no false alarm on the four controls. Both failures were the same gap, the definition absent from the catalogue altogether: a record with a missing field can be identified, an absent record cannot. This bounds the capability and indicates the next step, a register of expected concepts rather than only of present ones.

B2. Assembling only the records a question requires

The catalogue was enlarged to 41 records and the agent was asked to select a minimal subset before answering. The same question was then asked under three assembly conditions.

Assembly conditionAccuracyContext supplied, chars
whole catalogue80.0%10,332
agent's own selection60.0%473 (5%)
naive keyword top-560.0%1,989 (19%)

Selection was near perfect: 100% recall of the required records, 95% precision, and a mean of 1.2 records chosen from 41, reducing the context to a twentieth of the catalogue. Accuracy under the reduced context was nominally lower, but with 20 observations per condition the difference is not distinguishable from noise (p = 0.34). The supported claim is a twentyfold reduction in context at no measurable cost to accuracy. A gain in accuracy is not supported.

B3. Deciding what may be disclosed

Twelve information items and three requesting parties, our own carrier, a competing carrier and a statutory auditor, giving 72 disclosure decisions per condition. The prose policy and the classification records state the same rules and differ only in representation.

Policy representationRelease precisionRelease recallCritical errors, of 72
no policy39.1%45.0%5
policy as prose66.7%100%2
classification records100%100%0

Both errors under the prose policy were the same failure, releasing one carrier's operational data to a competitor of that carrier. The prose policy prohibits this in a sentence; the structured policy encodes it in a field, and made no error across all 72 decisions. This is the result with the most direct operational consequence, because the two error types are not symmetric: an unnecessary refusal costs friction, an unnecessary release can cost a contract.

B4. Completing a definition across roles

Three turns. A coordinator must address exactly one owner and request a specific field; that owner holds only its own authority and its own fact and must refuse anything outside it; the coordinator then records the answer as a new version and answers the question.

ConditionRouted to correct ownerWell-formed version recordFinal answer correct
prose, same gaps50.0%100%50.0%
records, same gaps100%100%50.0%

Routing, the capability specific to working across roles, doubled. The prose condition misdirected the requests about the active-customer window and about the holiday calendar, where accountability is expressed as a phrase rather than as a field. Final accuracy was identical in both conditions and was limited by arithmetic rather than by representation.

Conclusion

The claim that a versioned definition layer makes an agent more accurate is not supported by these data. Accuracy is a function of completeness, and well-written documentation achieves it equally. What documentation did not achieve at any level of quality is the identification of a missing definition and its owner, the citation of the rule and version behind a value, and the disclosure decision made without error. Those are the claims these data support, and they are stated in place of the broader one.

Limitations

One synthetic corpus generated once, one model family at two capability levels, no other providers, and tools disabled throughout. The reference answers are derived from the same definitions the structured condition receives, so the study measures adherence to a supplied specification and not the discovery of correct business meaning. Version 1 of this instrument contained a scope defect in its federation rules and computed significance over non-independent observations; both were identified by external adversarial review, both are documented, and the affected figures were discarded rather than corrected in place.

Open result.jsonRead the methodsRead the external reviewThe runnerRaw track A run

Materials

The whole instrument is published. world.py generates the synthetic organisations deterministically and holds the ground-truth engine, one implementation per definition. layers.py holds the six context conditions, tasks.py the 40 items, tracks.py the four capability tracks. run.py and run_tracks.py drive the calls; analyze.py and analyze_tracks.py perform the statistics at item level. Every raw run is published unedited: track A, Sonnet, B1, B2, B3, B4.