Vercy · reproducible study
Structure did not improve accuracy over complete prose.
- Objective
- Determine whether supplying an agent with versioned, structured definitions improves its answers, and whether any improvement is attributable to the structure or to the completeness of what is supplied.
- Design
- One synthetic supplier corpus with six ambiguous constructs whose rules change on known dates and differ between two organisations. 40 items asked under six context conditions in a within-item design, plus four tracks measuring gap detection, context assembly, disclosure control and completion across roles. 846 model calls. The unit of analysis is the item, not the call.
- Principal result
- Completeness accounted for the accuracy gain and representation did not. Continuous prose carrying the same facts as the structured records scored identically, 76.5% against 76.5%, exact two-sided sign test over 34 items p = 1.00.
- Secondary results
- The structured conditions were the only ones that cited the rule and version applied (100% against 0%), located a missing definition and named its owner (83% against 0%), routed a request to the correct owner (100% against 50%), and made all 72 disclosure decisions without a critical error (0 against 2).
- Conclusion
- On this item set, accuracy is a function of what the agent is told rather than of how it is represented. What distinguishes the structured representation is traceability, gap detection, disclosure control and routing, none of which the prose conditions achieved at any level.
Terms used on this page
| Term | Meaning in this study |
|---|---|
| Condition | A context block inserted between the data and the question. Conditions differ in that block and in nothing else. |
| Definition record | A machine-readable statement of one rule, carrying an identifier, a version, an effective date, the version it supersedes and the accountable role. |
| Complete prose | The same facts as the records, written as continuous documentation without identifiers, versions or fields. |
| Discriminating item | An item for which a plausible competing rule yields a different correct answer. Verified in code for all 34 answerable items. |
| Competing answer | A response exactly equal to the value the prespecified competing rule would produce. A lower bound on interpretation error, not a measure of it. |
| Critical disclosure error | Release of customer identities, discount terms, or one carrier's operational data to a competitor of that carrier. |
Design
Six constructs carry the ambiguity: active customer, net revenue, delivered, on time, order value and business day. Each was given rules that change on known dates and that the supplier and its carrier define differently. 40 items were then asked under six conditions with the underlying data tables identical throughout, so that any difference between conditions is attributable to the context block alone. Every answerable item was verified in code to be discriminating.
Track A. Accuracy on an agreed meaning
| Condition | Accuracy (M1) | 95% CI | Cited rule and version (M2) | Competing answer (M7) |
|---|---|---|---|---|
| schema only | 26.5% | 11-42 | 0% | 16.2% |
| incomplete documentation | 29.4% | 14-44 | 0% | 35.3% |
| complete prose | 76.5% | 65-88 | 0% | 2.9% |
| versioned records | 76.5% | 65-88 | 100% | 1.5% |
| records + inert padding | 85.3% | 78-93 | 100% | 0.0% |
| records + federation rules | 77.9% | 67-89 | 98.5% | 1.5% |
Supplying definitions in any form raised accuracy from 26.5% to 76.5%, which survives a paired exact sign test over 34 items at p = 0.0005. The condition carrying the same facts as continuous documentation scored identically to the structured one. Under this design accuracy is determined by the completeness of what is supplied and not by its representation.
Two further results qualify the first. Incomplete documentation performed no better than none and doubled the rate of answers matching a superseded rule, from 16.2% to 35.3%: a document describing a rule no longer in force is worse than no document. And the complete-prose condition, which holds every fact the structured condition holds, cited the governing rule in 0% of responses against 100%. It produces the correct value and cannot state which rule produced it.
Capabilities other than accuracy
Accuracy is one of the five capabilities measured. On the remaining four the prose conditions did not perform worse than the structured ones; they did not perform at all. The comparisons below are therefore differences between presence and absence rather than differences of degree.
B1. Detecting a missing definition
Six gaps were introduced into the catalogue: an empty parameter block, a missing effective date, a missing owner, a reference to a version that does not exist, a calendar ending a year early, and one definition absent altogether. Each was paired with a question that cannot be answered until the gap is filled. Four control questions contained no gap and measure the false-alarm rate.
| Condition | Gap detected | Field named | Owner named | False alarms |
|---|---|---|---|---|
| schema only | 0% | 0% | 0% | 0% |
| prose, same gaps | 0% | 0% | 0% | 0% |
| records, same gaps | 83.3% | 83.3% | 83.3% | 0% |
Prose containing the same six gaps detected none of them in twelve trials. The structured catalogue detected ten of twelve, naming the absent field and the accountable role in each case, with no false alarm on the four controls. Both failures were the same gap, the definition absent from the catalogue altogether: a record with a missing field can be identified, an absent record cannot. This bounds the capability and indicates the next step, a register of expected concepts rather than only of present ones.
B2. Assembling only the records a question requires
The catalogue was enlarged to 41 records and the agent was asked to select a minimal subset before answering. The same question was then asked under three assembly conditions.
| Assembly condition | Accuracy | Context supplied, chars |
|---|---|---|
| whole catalogue | 80.0% | 10,332 |
| agent's own selection | 60.0% | 473 (5%) |
| naive keyword top-5 | 60.0% | 1,989 (19%) |
Selection was near perfect: 100% recall of the required records, 95% precision, and a mean of 1.2 records chosen from 41, reducing the context to a twentieth of the catalogue. Accuracy under the reduced context was nominally lower, but with 20 observations per condition the difference is not distinguishable from noise (p = 0.34). The supported claim is a twentyfold reduction in context at no measurable cost to accuracy. A gain in accuracy is not supported.
B3. Deciding what may be disclosed
Twelve information items and three requesting parties, our own carrier, a competing carrier and a statutory auditor, giving 72 disclosure decisions per condition. The prose policy and the classification records state the same rules and differ only in representation.
| Policy representation | Release precision | Release recall | Critical errors, of 72 |
|---|---|---|---|
| no policy | 39.1% | 45.0% | 5 |
| policy as prose | 66.7% | 100% | 2 |
| classification records | 100% | 100% | 0 |
Both errors under the prose policy were the same failure, releasing one carrier's operational data to a competitor of that carrier. The prose policy prohibits this in a sentence; the structured policy encodes it in a field, and made no error across all 72 decisions. This is the result with the most direct operational consequence, because the two error types are not symmetric: an unnecessary refusal costs friction, an unnecessary release can cost a contract.
B4. Completing a definition across roles
Three turns. A coordinator must address exactly one owner and request a specific field; that owner holds only its own authority and its own fact and must refuse anything outside it; the coordinator then records the answer as a new version and answers the question.
| Condition | Routed to correct owner | Well-formed version record | Final answer correct |
|---|---|---|---|
| prose, same gaps | 50.0% | 100% | 50.0% |
| records, same gaps | 100% | 100% | 50.0% |
Routing, the capability specific to working across roles, doubled. The prose condition misdirected the requests about the active-customer window and about the holiday calendar, where accountability is expressed as a phrase rather than as a field. Final accuracy was identical in both conditions and was limited by arithmetic rather than by representation.
Conclusion
The claim that a versioned definition layer makes an agent more accurate is not supported by these data. Accuracy is a function of completeness, and well-written documentation achieves it equally. What documentation did not achieve at any level of quality is the identification of a missing definition and its owner, the citation of the rule and version behind a value, and the disclosure decision made without error. Those are the claims these data support, and they are stated in place of the broader one.
Limitations
One synthetic corpus generated once, one model family at two capability levels, no other providers, and tools disabled throughout. The reference answers are derived from the same definitions the structured condition receives, so the study measures adherence to a supplied specification and not the discovery of correct business meaning. Version 1 of this instrument contained a scope defect in its federation rules and computed significance over non-independent observations; both were identified by external adversarial review, both are documented, and the affected figures were discarded rather than corrected in place.
Materials
The whole instrument is published. world.py generates the synthetic organisations deterministically and holds the ground-truth engine, one implementation per definition. layers.py holds the six context conditions, tasks.py the 40 items, tracks.py the four capability tracks. run.py and run_tracks.py drive the calls; analyze.py and analyze_tracks.py perform the statistics at item level. Every raw run is published unedited: track A, Sonnet, B1, B2, B3, B4.