Y2 Elite workspaces are rolling out for teams
Y2Y2Docs
Ontology & Fusion

Why a Claim Ledger

Why Y2 records research as append-only, evidenced claims with three clocks, and why a shared name never merges two subjects

Most knowledge graphs store the current answer: one node per thing, properties overwritten as new data arrives, and a merge whenever two records look alike. That shape is fast to query and easy to get wrong in ways you cannot see later. Y2's entity ledger stores something different: what a public source said, when it was retrieved, how strongly it supports a claim, and when that claim is taken to hold. The picture you read is computed from those claims.

The subject of a claim is not its string

A person, organization, or system is a subject: a row opened with a class and nothing else. Everything that can change arrives as a claim about it. A name, a handle, a domain, or a CIK is a designator: one public string, stored once.

Keeping the string and the subject apart is what prevents the most common identity error in collected intelligence. Two people who share a published name, or two companies whose aliases overlap, are two subjects that happen to be designated by one string. Graph platforms that derive an entity's identity from its name and aliases collapse them. In the ledger, a shared designator is only a shared string. Subjects merge only by a human decision, made by reading a same-as hypothesis that the ledger still holds as a hypothesis; when a distinct-from is also live, the pair is contested and both rows remain.

Every change is a new line

Claims are appended, never edited. A correction supersedes the current head of its claim key; a retraction clears the grade and interval but still cites the observation that says why. Each line is sealed to the one before it with SHA-256, so an edited history fails verification.

This costs some storage and buys three things:

  • Accountability. You can show what you believed, on what evidence, before a correction.
  • Safe collaboration. A line must name the head it replaces, so two analysts working from stale views cannot silently overwrite each other; the second write is rejected.
  • Portability. The canonical log is a plain NDJSON file that anyone can check with the reference checker, independent of Y2.

Verification belongs to the claim

A person is not "verified." A specific handle claim is located because one page ties it, or cross-linked because two different pages identify it. The ledger enforces that the state matches the evidence: cross-linked needs two retrieved pages; analytic means the only support is an analyst's note, which is how hypotheses and display labels are graded. This follows the Berkeley Protocol's separation of provenance, source evaluation, and verification, and keeps the grade on the thing that was actually checked.

Three clocks

ClockFieldQuestion
Record timerecorded_atWhat had the ledger written by then?
Valid timevalid_from, valid_toWhen is the claim taken to hold in the world?
Retrieval timeretrieved_atWhen was the page fetched?

Record time selects a prefix of the log: what the research believed at that moment. Valid time filters the heads of that prefix. Asking "what did we believe on March 20 about April 10" sets both. A missing bound means unknown, never a guessed date: a page that says "since 2023" is quoted in the excerpt until a source gives a day. When the only evidence is that a page showed something between two retrievals, the interval is a sighting interval and the note says so.

Passive by construction

The method vocabulary is closed and public: web pages, profiles, DNS, certificate transparency, archives, registries, code hosts, and analyst notes that cite them. Interaction with the subject, authentication, and non-public access have no name in the vocabulary, so they cannot be recorded as if they were valid sources. The ledger keeps a short excerpt and its hash, not page bodies, and Y2 never writes a personal name into a sealed line.

Where the design comes from

The ledger borrows its publishing pattern from MITRE ATT&CK (stable IDs, every statement sourced), its interchange from STIX 2.1 (identities, observables, relationships with confidence and time bounds), its provenance fields from W3C PROV-O, and its method duties from the Berkeley Protocol. It deliberately differs from graph platforms that upsert on a deterministic ID and deduplicate nearby edges: the ledger keeps every correction. See STIX 2.1 export for the projection, and Catalog and ledger alignment for how the ledger relates to Y2's shared catalog.