Why a Claim Ledger
Why Y2 records research as append-only, evidenced claims with three clocks, and why a shared name never merges two subjects
Most knowledge graphs store the current answer: one node per thing, properties overwritten as new data arrives, and a merge whenever two records look alike. That shape is fast to query and easy to get wrong in ways you cannot see later. Y2's entity ledger stores something different: what a public source said, when it was retrieved, how strongly it supports a claim, and when that claim is taken to hold. The picture you read is computed from those claims.
The subject of a claim is not its string
A person, organization, or system is a subject: a row opened with a class and nothing else. Everything that can change arrives as a claim about it. A name, a handle, a domain, or a CIK is a designator: one public string, stored once.
Keeping the string and the subject apart is what prevents the most common identity error in
collected intelligence. Two people who share a published name, or two companies whose aliases
overlap, are two subjects that happen to be designated by one string. Graph platforms that derive an
entity's identity from its name and aliases collapse them. In the ledger, a shared designator is only
a shared string. Subjects merge only by a human decision, made by reading a same-as hypothesis that
the ledger still holds as a hypothesis; when a distinct-from is also live, the pair is contested
and both rows remain.
Every change is a new line
Claims are appended, never edited. A correction supersedes the current head of its claim key; a retraction clears the grade and interval but still cites the observation that says why. Each line is sealed to the one before it with SHA-256, so an edited history fails verification.
This costs some storage and buys three things:
- Accountability. You can show what you believed, on what evidence, before a correction.
- Safe collaboration. A line must name the head it replaces, so two analysts working from stale views cannot silently overwrite each other; the second write is rejected.
- Portability. The canonical log is a plain NDJSON file that anyone can check with the reference checker, independent of Y2.
Verification belongs to the claim
A person is not "verified." A specific handle claim is located because one page ties it, or
cross-linked because two different pages identify it. The ledger enforces that the state matches
the evidence: cross-linked needs two retrieved pages; analytic means the only support is an
analyst's note, which is how hypotheses and display labels are graded. This follows the Berkeley
Protocol's separation of provenance, source evaluation, and verification, and keeps the grade on the
thing that was actually checked.
Three clocks
| Clock | Field | Question |
|---|---|---|
| Record time | recorded_at | What had the ledger written by then? |
| Valid time | valid_from, valid_to | When is the claim taken to hold in the world? |
| Retrieval time | retrieved_at | When was the page fetched? |
Record time selects a prefix of the log: what the research believed at that moment. Valid time filters the heads of that prefix. Asking "what did we believe on March 20 about April 10" sets both. A missing bound means unknown, never a guessed date: a page that says "since 2023" is quoted in the excerpt until a source gives a day. When the only evidence is that a page showed something between two retrievals, the interval is a sighting interval and the note says so.
Passive by construction
The method vocabulary is closed and public: web pages, profiles, DNS, certificate transparency, archives, registries, code hosts, and analyst notes that cite them. Interaction with the subject, authentication, and non-public access have no name in the vocabulary, so they cannot be recorded as if they were valid sources. The ledger keeps a short excerpt and its hash, not page bodies, and Y2 never writes a personal name into a sealed line.
Where the design comes from
The ledger borrows its publishing pattern from MITRE ATT&CK (stable IDs, every statement sourced), its interchange from STIX 2.1 (identities, observables, relationships with confidence and time bounds), its provenance fields from W3C PROV-O, and its method duties from the Berkeley Protocol. It deliberately differs from graph platforms that upsert on a deterministic ID and deduplicate nearby edges: the ledger keeps every correction. See STIX 2.1 export for the projection, and Catalog and ledger alignment for how the ledger relates to Y2's shared catalog.