Vercy · reproducible study

Structure did not beat complete notes. Removing any part of it did.

Objective
Determine what the representation of one person's own records does for an assistant reading them, and which part of a structured representation carries which capability.
Design
One synthetic person: 34 facts with validity intervals and sources, 7 contacts, 8 journeys, 5 live conflicts between sources of unequal reliability, 3 expired documents and 2 still valid. Six representations, all rendered from the same fact table. 43 items in seven families, presented under every representation. Two further tracks on disclosure and on recording an update. 570 model calls. The unit of analysis is the item.
Principal result
The claim that structure helps was not established. Chronological notes carrying every fact, every date and every source scored 93.4% against 100% for the structured dimension, three items apart, exact sign test p = 0.25.
What did separate
Every ablation of the structure was decisive. Removing validity intervals took historical accuracy from 100% to 0% and expiry judgements from 100% to 20%, while leaving current-state questions untouched at 100%. Removing provenance took conflict resolution from 100% to 33%, also leaving current-state questions untouched. Only the arms carrying a conflict policy could name the rule they applied, 100% against 0%.
Correction to the previous study
Traceability is not a property of structure. When the notes carry source identifiers, they cite them on 96.4% of answers. The earlier benchmark reported 0% for prose because its prose had no identifiers to cite, not because it was prose.
Conclusion
For a personal record set of this size, well-kept notes are close to sufficient, and the value of structure is not in the average answer. It is that the fields are individually load-bearing: each one is the sole support of a capability that collapses without it, and a representation that drops any of them fails in a specific, predictable way rather than degrading gently.
93%notes, all facts present
100%structured dimension
0.25paired sign test, p

Terms used on this page

TermMeaning in this study
DimensionThe records describing one person: facts, the people and journeys they refer to, the sources they came from, and the rule for resolving disagreement.
Validity intervalThe dates between which a fact was true. An open interval means it is still true.
Source tierHow far a kind of source is trusted: document scan 1, device sync 2, email 3, self-reported 4.
Conflict policyCP-01: where two records covering the same date disagree, the lower tier wins; within a tier, the later capture wins.
Naive recencyThe rule a summary without provenance can actually implement: the most recently captured value wins, whatever its source.
OverclaimAnswering a question about a past date with the value that holds today.
AblationAn arm with one field class deliberately removed, used to find which field class carries which capability.

Design

The subject is one synthetic person whose records look like a personal assistant's: addresses, employers, phone numbers, salary, identity documents, devices and warranties, subscriptions, health data, bank details and diet. Each fact carries the dates between which it was true and the source it came from, ranked by how far that kind of source is trusted. Five attributes carry two records that both cover today and disagree; three documents have expired and two have not.

The previous benchmark was criticised, correctly, for a prose condition that silently lacked facts the structured condition had. Here every representation is rendered from one fact table, and a check that runs before every execution fails the study if any complete representation is missing any value, date or source that another one carries. That check found seven omissions in the notes while the instrument was being built, which were fixed before the first call. Arms A, B and D are complete; arms C, E and F are ablations that remove one field class each, and are labelled as such.

The main grid

RepresentationCompleteAnswer accuracy95% CICited the recordOverclaimFabricationContext
A notesyes93.4%86-10096.4%2.6%5.3%4,664
B notes + profile summaryyes93.4%86-10092.9%4.0%5.3%5,242
C flat recordsablation47.4%31-63-6.6%1.3%2,097
D dimensionyes100.0%100-100100.0%0.0%0.0%10,038
E dimension minus provenanceablation86.8%76-98-5.3%5.3%4,903
F dimension minus validityablation55.3%39-7193.3%0.0%1.3%6,744

The headline comparison is arm A against arm D, the two complete representations. Notes answered 93.4% of items correctly, the dimension 100%. Across 38 answerable items the two disagree on three, all in the dimension's favour, which an exact sign test puts at p = 0.25. On this item set the structured representation is not shown to be more accurate than complete notes, and the honest statement of the headline result is that the claim was not established rather than that it was refuted.

Three differences do survive the small sample. The dimension is the only arm with no overclaim, meaning it never answered a question about a past date with today's value, and the only arm that never produced a value absent from the records. Its context is also the largest, at 10,038 characters against 4,664 for the notes, so whatever it buys, it buys with roughly twice the context.

By question family

Representationcurrenthistoricalconflictexpirysame-namemulti-recordunanswerable
A notes100.0%100.0%75.0%100.0%100.0%83.3%100.0%
B notes + profile summary100.0%100.0%75.0%100.0%100.0%83.3%100.0%
C flat records100.0%0.0%33.3%20.0%100.0%33.3%100.0%
D dimension100.0%100.0%100.0%100.0%100.0%100.0%100.0%
E dimension minus provenance100.0%100.0%33.3%100.0%100.0%83.3%100.0%
F dimension minus validity100.0%0.0%83.3%20.0%100.0%33.3%100.0%

The averages hide the shape of the result. Every representation answers current-state questions perfectly, including the one that holds nothing but current state. The differences are entirely in the families that need a field the ablated arms do not have: history, conflict and expiry. Flat records, which keep only what is true today, answer none of the historical questions and one in five of the expiry questions, and decline 44.7% of the items they were asked.

Which field carries which capability

Each ablation is paired with a control family where the removed field is predicted to make no difference. This is what separates a claim about a field from a claim about size or structure in general.

Field class removedFamilyAblatedFull dimensionp
validity intervalshistorical questions0.0%100.0%0.00781
validity intervalscurrent questions100.0%100.0%1.0
provenance and policyconflict questions33.3%100.0%0.125
provenance and policycurrent questions100.0%100.0%1.0

Both predictions hold and both controls hold. Validity intervals are what make historical questions answerable at all; without them the accuracy on that family is zero, and current-state accuracy is unchanged. Provenance is what makes conflict resolution work; without it accuracy on conflicts falls by two thirds, and again current-state accuracy is unchanged. The conflict contrast has only six items and does not reach the conventional threshold on its own, but its direction and its control are both as predicted.

Expiry, false alarms and declining

Three documents have expired and two have not. The two valid ones are controls: a representation that makes expiry visible must not start flagging records that are still good.

RepresentationExpiry judged correctlyFalse expiry alarmsDeclined a question it could answer
A notes70.0%3.0%0.0%
B notes + profile summary60.0%0.0%0.0%
C flat records50.0%0.0%44.7%
D dimension90.0%0.0%0.0%
E dimension minus provenance90.0%0.0%5.3%
F dimension minus validity60.0%9.1%39.5%

The dimension judges expiry correctly on 90% of trials with no false alarms, against 70% for notes and 60% for the profile summary. The ablated arms look respectable on this metric only because they decline so often: flat records declined 44.7% of the answerable items and the arm without validity intervals 39.5%, against 0% for both complete representations. Declining is the correct response to information you do not have, and it is why the accuracy of those arms should be read as a measure of what their representation cannot express rather than of how the model reasons.

Deciding what may be released

Twelve pieces of information and three requesters: the subject's sister as next of kin, the employer's HR department, and an unverified web form. The prose policy and the classification records state the same rules and differ only in representation.

Policy representationRelease precisionRelease recallDecision accuracyCritical errors
no policy95.2%58.8%79.2%1 of 72
policy as prose100.0%79.4%90.3%0 of 72
classification records100.0%88.2%94.4%0 of 72

The prediction that structured classification would reduce critical errors was not supported. Both policy forms produced none: a short, clearly written prose policy was enough. What the structured form did was refuse less unnecessarily, releasing 88.2% of what the requester was entitled to against 79.4% for prose, with precision at 100% in both. The single critical error in the whole track came from the condition with no policy at all, which released a health detail to the employer. This is a different outcome from the previous study, where a prose policy leaked twice; the policy here is shorter and covers fewer cases, and that is the most likely reason.

Recording a change

Six pieces of new information arrive: a move, a renewed passport, a cancelled subscription, a new number, a laboratory correction and a renewed insurance policy. The agent must record each without destroying what was true before.

RepresentationNew value rightStart date rightClosed the right recordEnd date rightHistory keptAll four right
A notes100.0%83.3%0.0%100.0%100.0%0.0%
C flat records100.0%100.0%0.0%75.0%100.0%0.0%
D dimension100.0%100.0%75.0%75.0%100.0%66.7%

Every arm produced the correct new value and every arm said it had kept the history. Only the dimension could name the record it was superseding, on 75% of scenarios, and it was the only arm to get all four parts of the update right, on 66.7%. The notes and flat arms score zero on that column because their records have no identifiers to name, which is a property of the representation rather than of the model's reasoning: an update cannot be expressed as a supersession when there is nothing to supersede by name. Even the dimension gets the whole update right only two thirds of the time, so this is a capability the representation makes possible and does not guarantee.

Conclusion

Two studies now agree on the same uncomfortable point: on average accuracy, a complete and well-kept unstructured record is hard to beat. What this study adds is where the difference actually lives. Each field class in the structure is the sole support of one capability, and the failure when it is missing is not a gentle decline but a collapse to zero on exactly the questions that field serves. A personal assistant whose memory keeps only current state will answer every question about today and cannot answer a single one about last year. That is the claim these data support.

Two defects found and fixed during the run

The provenance metric originally accepted only the source identifier. In the arms that also carry fact identifiers the model cited the fact, which answers the question asked exactly as unambiguously, so the metric was penalising the arms with more identifiers; it reported 71.4% for the dimension where the corrected metric reports 100%. Separately, an abstention returned from the scorer before the expiry judgement was recorded, so the ablated arms were being scored on the two or three items where they happened not to decline. Both were found by reading raw responses rather than summaries, both were fixed, and both were recomputed from the stored responses without repeating a call. The raw files are published unmodified and the corrected scorer is the one in the instrument.

What this does not show

One authored subject, one model, tools disabled, 34 facts. A dimension of this size is the scale of a personal assistant's memory and not of a document corpus, and nothing here transfers to large collections. The reference answers come from the same fact table the structured arms display, so the study measures retrieval and correct application of supplied records, not the discovery of what is true about a person. The dimension is also the longest context and the flat records the shortest; there is no padding control separating length from content, which the previous study had and this one does not. The rule attributed to the summary and flat arms, that the most recently captured value wins, is a modelling assumption about how such representations are built.

Open result.jsonRead the methodsThe dimensionThe six renderersRaw grid run

Materials

The whole instrument is published. dimension.py holds the subject, every fact with its interval and source, and the ground-truth engine. views.py renders all six representations from that one table and carries the completeness check. items.py holds the 43 items and the two tracks, run.py the runner and the scorer, analyze.py the statistics at item level. Every raw run is published unedited: the grid, disclosure and write-back.