Vercy · reproducible study

Where governance lives did not matter. Whether it exists mattered a lot.

Objective
When several teams write into one shared memory, does it matter whether ownership, the conflict rule and the release list live inside the record or in a governance layer beside it?
Design
One fictional organisation, six teams writing into one memory: 30 records, 8 contested business concepts, 6 rival definitions written by teams that do not own them, 3 superseded versions and 5 records restricted to named teams. Six representations, all rendered from the same record table. 35 items in seven families, plus a cross-team release track and a restatement track. 510 model calls. The unit of analysis is the item.
Principal result
The Vercy-specific claim was not supported. Putting the governance fields inside the record scored 95.2% against 93.5% for the same graph with a separate governance catalogue beside it, two items each way, p = 1.00. A plain team wiki carrying the same facts scored 96.8%, also p = 1.00.
What was supported
Carrying that information at all is worth a great deal. Adding a governance catalogue to a bi-temporal graph took it from 77.4% to 93.5%, ten items better and two worse, p = 0.039. Removing the governance fields from the dimension cost 21 points, p = 0.016, and took conflict resolution from 100% to 33%.
The leak result
Every representation carrying release rules made all 60 cross-team release decisions correctly, with no over-shares. The bi-temporal graph without them released only 59.1% of what each team was entitled to and made two critical over-shares.
Conclusion
For shared organisational memory, the field class that decides the outcome is ownership, the conflict rule and the release list. The packaging is interchangeable: a wiki, a catalogue beside a graph and fields inside a record all work. What does not work is the shape several memory products actually ship, which carries time and provenance but none of that.
94%graph with a catalogue beside it
95%governance inside the record
77%graph with neither

Terms used on this page

TermMeaning in this study
Shared memoryOne record set that several teams write into and every agent reads.
Concept ownerThe single team accountable for what a business term means. Six teams write; for each concept exactly one of them decides.
CP-ORGThe conflict rule used here: where two records cover the same date and disagree, the owning team's record prevails; within one team the later effective date prevails. A non-owner's record is kept as that team's view, never silently overwritten.
Governance fieldsConcept owner, accountable role, conflict policy and release list, carried inside each record rather than in a separate catalogue.
Rival team's valueAn answer that reports the figure written by a team that does not own the concept.
Critical over-shareReleasing salary bands, litigation or acquisition talks to a team not entitled to read them.

What was compared, and what was not

Three of the six representations are reimplementations of documented product shapes. No product was executed. Nothing ran against a vendor's service, and retrieval, ranking, ingestion and storage are all out of scope. What is compared is the shape of the record an agent is given to read, which is one component of a memory system and not the system. A vendor could fairly object that their retrieval would change what the agent sees; that objection is correct and is the main limitation here.

Arm B follows a vector-first store that keeps write timestamps but supports no as-of query. Arm C follows a bi-temporal knowledge graph whose fact edges carry valid-at and invalid-at beside the episode they came from. Arm D adds the separately governed glossary that catalogue vendors position next to such a graph, which is the architecture the industry currently recommends and the strongest competitor in this design. Arms A, D and E carry everything and are verified complete in code before every run; B, C and F drop a field class on purpose.

The main grid

RepresentationCompleteAnswer accuracy95% CINamed the ownerCited the ruleCited the recordTook the rival team's valueContext
A shared wikiyes96.8%92-100100.0%100.0%100.0%0.0%6,618
B vector memory with timestampsablation56.5%40-7377.1%0.0%0.0%3.2%2,694
C bi-temporal knowledge graphablation77.4%66-8979.2%0.0%100.0%4.8%8,932
D graph plus a governance catalogueyes93.5%86-100100.0%100.0%100.0%0.0%10,621
E dimensionyes95.2%88-100100.0%100.0%100.0%0.0%11,191
F dimension minus governance fieldsablation74.2%59-8968.8%0.0%100.0%6.5%6,474

Three representations cluster at the top and are statistically indistinguishable from one another: a team wiki at 96.8%, the dimension at 95.2% and a graph with a separate catalogue at 93.5%. Every pairwise comparison among them returns two items each way or fewer, p between 0.62 and 1.00. What they share is that each states who owns a concept, how competing entries are settled, and who may read what. Where that information sits appears not to matter.

ComparisonFromToItems betterItems worsep
A shared wiki against E dimension96.8%95.2%221.0
D graph plus a governance catalogue against E dimension93.5%95.2%221.0
F dimension minus governance fields against E dimension74.2%95.2%700.01562
C bi-temporal knowledge graph against E dimension77.4%95.2%1120.02246
B vector memory with timestamps against E dimension56.5%95.2%1400.00012
C bi-temporal knowledge graph against D graph plus a governance catalogue77.4%93.5%1020.03857
A shared wiki against D graph plus a governance catalogue96.8%93.5%130.625
B vector memory with timestamps against C bi-temporal knowledge graph56.5%77.4%1130.05737

The three representations without it sit twenty to forty points lower. The gap is not spread evenly: they answer questions about the present and about unrecorded facts as well as anything else, and they fail on the questions that only arise because several teams write into the same place.

By question family

Representationauthorityconflicthistoricalsupersededroutingprovenanceunrecorded
A shared wiki100.0%91.7%100.0%100.0%100.0%87.5%100.0%
B vector memory with timestamps66.7%33.3%70.0%70.0%90.0%0.0%100.0%
C bi-temporal knowledge graph75.0%41.7%90.0%80.0%90.0%100.0%100.0%
D graph plus a governance catalogue100.0%91.7%100.0%90.0%100.0%75.0%100.0%
E dimension100.0%100.0%100.0%70.0%100.0%100.0%100.0%
F dimension minus governance fields83.3%33.3%100.0%60.0%80.0%100.0%100.0%

Conflict is where the difference concentrates. Asked which of two live records stands when two teams disagree, the representations carrying an ownership rule answer 92% to 100% correctly; those without it answer 33% to 42%, and they do not fail by declining. They pick the other team's number and state it as the company's.

Removing the governance fields, with a control

Arm F is the dimension with concept owner, accountable role, conflict policy and release list removed. The historical family is the control: those fields have nothing to do with what was in force on a date, so removing them should make no difference there.

FamilyRoleGovernance fields removedFull dimensionp
authoritytarget83.3%100.0%1.0
conflicttarget33.3%100.0%0.125
routingtarget80.0%100.0%1.0
historicalcontrol100.0%100.0%1.0

The control holds exactly: historical accuracy is 100% with and without the governance fields. The three target families all fall, most sharply on conflict, from 100% to 33%. With six items in that family the contrast does not reach the conventional threshold on its own, but the whole-grid comparison does, at p = 0.016.

A check that the competitor was not made a straw man

Arms C and F remove the same field classes by two different routes: C by reimplementing a competitor's shape, F by deleting fields from the dimension. If the reimplementation were unfair, C would score well below F. It does not: 77.4% against 74.2%, five items better and six worse, p = 1.00. The two agree, which is the evidence available here that the competitor shape was rendered faithfully.

What may cross a team boundary

Five restricted records, six requesting teams, 30 release decisions per representation per replicate. Salary bands, litigation and acquisition talks count as critical if released to a team not entitled to read them.

RepresentationRelease precisionRelease recallDecision accuracyCritical over-shares
A shared wiki100.0%100.0%100.0%0 of 60
B vector memory with timestamps95.2%90.9%95.0%1 of 60
C bi-temporal knowledge graph86.7%59.1%81.7%2 of 60
D graph plus a governance catalogue100.0%100.0%100.0%0 of 60
E dimension100.0%100.0%100.0%0 of 60

This is the sharpest result in the study. Every representation carrying release rules was perfect: 100% precision, 100% recall, no over-shares across 60 decisions, whether the rules sat in a wiki page, in a catalogue or in a field. The bi-temporal graph, which carries time and provenance but no release list, made two critical over-shares and refused 41% of what each team was entitled to see. It is wrong in both directions at once, which is the characteristic failure of guessing.

One team restating another team's record

Five pieces of news arrive, each a team revising something a different team had recorded. The earlier record must survive and the other team's entry must not be discarded.

RepresentationNew value rightEffective date rightRecorded to the right teamClosed the right recordHistory keptOther team's entry kept
A shared wiki100.0%80.0%100.0%0.0%100.0%100.0%
C bi-temporal knowledge graph100.0%80.0%100.0%100.0%100.0%100.0%
E dimension100.0%80.0%100.0%100.0%100.0%100.0%

Every representation produced the right new value, recorded it to the right team and said it had kept the history. The wiki could not name the record it was superseding on any scenario, because its entries have no identifiers; the graph and the dimension named it every time. This is a property of having identifiers rather than of reasoning, and it is the one place in this study where prose has no way to express what is being asked.

Conclusion

The claim this study set out to test was that governance belongs inside the memory record rather than in a layer beside it. These data do not support it: the two arrangements are indistinguishable, and so is a well-kept wiki. The claim the data do support is about the category rather than about us. A memory that carries validity intervals and provenance but no ownership, no conflict rule and no release list is missing the field class that decides most of the questions a shared memory is actually asked, and adding it is worth sixteen points and two avoided leaks. That is a finding for whoever standardises this, not an argument for one implementation.

Two defects, one caught before the run and one during it

The provenance measure originally accepted only the document identifier, which would have penalised the arms that also carry record identifiers for citing the record instead. That was the same defect found in the previous study, and here it was fixed before the first call. During the run, three of the five historical items were found to ask for a version number, a field only some representations carry: the family was measuring which arms have a version counter rather than which can say what was in force on a date. The items were reworded to ask for the content in force, a match rule was added that rejects an answer quoting the later wording, and the family was re-run under all six arms. Both runs are published and the analysis uses the corrected one.

What this does not show

One authored organisation of 30 records, one model, tools disabled, and retrieval excluded, which is the component several of the modelled products consider their main contribution. The competitor shapes are reimplementations from public documentation and may not match a current release. Reference answers come from the same record table the arms display, so this measures correct application of supplied records and not the discovery of what is true in an organisation. The conflict rule is ours. Thirty-one answerable items in families of four to six is a small sample, and no correction for multiple comparisons is applied.

Open result.jsonRead the methodsThe organisationThe six renderersRaw grid run

Materials

The whole instrument is published. org.py holds the organisation, every record with its validity interval, writing team and source, and the ground-truth engine. views.py renders all six representations from that one table and carries the completeness check. items.py holds the 35 items and both tracks, run.py the runner and the scorer, analyze.py the statistics at item level. Every raw run is published unedited, including the superseded first run of the historical family: the grid, the corrected historical family, the release track and the restatement track.