What structured memory costs, and what closes the gap.
This is a follow-on to the memory testing story and the architecture drawn post from four weeks ago. Those covered how the memory subsystem got tested and how it was shaped. This one covers what the tests then said, and what four RFCs did about it.
The paired benchmark that measures this now runs 199 questions over the same corpus, one variable changed: whether the answerer reads distilled facts, or whether it reads the raw session turns those facts were extracted from. Same retriever, same reader, same evaluator, same questions. Here is what it says.
On temporal questions specifically ("which medication was I on in April"), the split is worse: facts 0.036, turns 0.873. The paired McNemar shows 111 questions turns-only-right against 3 facts-only-right. This is not a rounding artifact. Structured memory on its own, on real long-form questions, loses to reading the transcript.
That is the honest surprise. If a memory subsystem sells the pitch that a compressed structured store answers the same questions as the raw log, at 7.2–9.0× compression, and it turns out the number is 0.157 to 0.738, then the pitch is wrong and the number is what matters. The last four weeks of loomcycle work is what closes that gap. It does not close it by walking the pitch back; it closes it by naming the four things structured memory was missing.
The short version. Four RFCs, one arc. RFC CV collapses the two-halves-of-a-fact duplication (a k/v row plus a separate entity chunk) into subject-document identity: a subject IS its Document, and its facts are its children, sitting in the Path tree at /entities/<type>/<name>. Placement is one decision, and both halves land together or neither does. RFC DF ships recall_attach_traces, a runtime-supplied grant that reattaches the source turns to a facts recall (measured +24.5pp on LongMemEval and +25.9pp on LoCoMo, twice). The two other shapes for delivering the same grant, a tool parameter and an imperative prompt, both failed measurably: the tool-parameter version was passed on 51 of 128 calls; the imperative-prompt version reached 0.0034 accuracy at 1.0 abstention because a mandated protocol competes with the task. A cascade with a verifier routes questions past the expensive gate cheaply: 9 percent of calls buy 40 percent of the gain; 39 percent buy 60 percent. RFC DM (landing in v1.100) decomposes a chunk into derived search units so a heading path, a code block, and a transclusion are separately findable. Postgres-only. Corpus-dependent thresholds.
The paired benchmark, and why the number hurts
Loomcycle's paired retrieval benchmark is described in docs/MEMORY-ARCHITECTURE.md under "What has actually been measured." It loads a real long-form multi-session corpus, runs each conversation through the extractor to produce the structured store, then poses 199 labelled questions to a fresh session in two arms.
- Facts arm. The answerer reaches into memory only. Distilled facts, notes, and documents come back through the normal recall path.
- Turns arm. The same answerer reaches into the raw session turns, retrieved by the same semantic index over the same content, at the same top-k.
Both arms retrieve, then answer. The variable is what the retrieved content is: a compressed structured summary of what happened, or the actual thing that happened. The facts arm compresses at 7.2:1 (LongMemEval slice) and 9.0:1 (LoCoMo slice), so the numeric win a facts-only system would need to justify itself is a comparable answer at a fraction of the context.
Instead the accuracy split is 0.157 to 0.738, the abstention split is 0.764 to 0.146 (facts abstain five times as often as turns), and on single-hop questions where a distilled fact should shine, facts read 0.199 against turns' 0.775. The one slice where the arms tie is open-domain (0.312 both), where retrieval quality is not what is being measured. Everywhere the corpus tests memory of the conversation, turns win.
This is not because the facts are wrong. Consolidation quality had already been measured (see the memory testing story) and the extractor turns real conversations into real facts. The reason for the split is that a distilled fact throws away the linguistic detail the reader needed to answer the actual question. "Was the payment for the annual subscription or the monthly one?" is a fact question in the abstract; on a real transcript, the reader needs the sentence Alice wrote, not the label the extractor put on it.
What RFC CV changes: a subject is its Document
The pre-CV picture (the one drawn in the architecture post) had a fact stored twice. A k/v row under memory/<class>/<slug> that semantic recall searched, and a separate entity-graph chunk that a graph walk reached. The two halves were supposed to travel together, and mostly did, but they were coordinated through two independent grants (memory_scopes and sql_scopes) and two independent code paths.
RFC CV collapses that. There is one representation now: the subject is its Document. Facts about the subject are the Document's children. Placement is a single decision that lands the whole subject-Document in one scope. The Document sits in the Path tree at /entities/<type>/<name> alongside every other Document, using the dirent/inode split described in docs/PATH.md. The k/v row is retired. The graph walk seeds from a semantic match on the subject-Document rather than a title hit.
flowchart LR
subgraph BEFORE["Before RFC CV"]
F1["Fact about Ada"] --> KV["k/v row
memory/person/ada"]
F1 --> CHUNK["entity chunk
separate table"]
KV -.->|semantic recall
reaches this| RECALL1["recall result"]
CHUNK -.->|graph_recall
reaches this| GRAPH1["graph_recall result"]
KV -. two grants
can disagree .-> CHUNK
end
subgraph AFTER["After RFC CV"]
F2["Fact about Ada"] --> SUBDOC["subject-Document
/entities/person/ada"]
SUBDOC -->|semantic recall
reaches THIS| RECALL2["recall result"]
SUBDOC -->|graph_recall
seeds from a
semantic match on this| GRAPH2["graph_recall result"]
end
style BEFORE fill:#fff5f5
style AFTER fill:#f0fff4
style SUBDOC fill:#e8ffef,stroke:#4a9e60
The measurements the collapse produces are two: on a two-user corpus, duplicate copies fell from 30 to 23 (a 23 percent reduction, not elimination; the honest figure is that cross-user duplication was eliminated and replaced by tenant-owner duplication). And 94 facts became readable by both users where none had been before. graph_recall's rebuilt seeding path measures 30 ms against the old 928 ms on a 2,942-row scope. About 31 times faster.
Two operator-visible pieces that shipped with the same arc. First, the consolidator gained a collapse_facts step that refuses to lose a fact when merging: if two extractions produce facts that overlap, the merge is required to include everything from both, and a failure aborts the write. Second, an unknown subject is now proposed for operator adoption (Memory op=propose_subject) rather than minted on the spot, and the tenant entity registry is curator-gated. A rogue extraction cannot silently share a subject.
What RFC DF changes: the trace grant
CV closes the placement gap but not the answer-quality gap. If the facts arm reads a distilled sentence that dropped the wording the reader needs, no amount of placement discipline recovers that wording.
RFC DF (recall_attach_traces) is the shape that does. It is an opt-in per-agent grant: when set, a facts recall reattaches the source turns each fact was drawn from to the returned payload, as a separate source_turns block on the response, alongside the fact bodies themselves. Not a re-rank, not a summary, not a synthesis. The reader gets both.
The measurement is +24.5 percentage points on LongMemEval and +25.9 percentage points on LoCoMo. Measured twice.
The interesting part of RFC DF is not that the grant helps. It is that the grant only helps when it is runtime-supplied, and the other two shapes for delivering the same behaviour both failed measurably.
- Tool parameter. The most obvious shape was to make attach-traces a boolean argument the model passes on the Memory recall call. Measured: the model passed it on 51 of 128 calls. It forgot on the other 77. A parameter the model has to remember to set is a parameter that competes with the task.
- Imperative prompt. The next obvious shape was to make the attach-traces protocol part of the extraction instructions, so the model would always attach traces because the instructions said so. Measured: the accuracy dropped to 0.0034 at 1.0 abstention. Adding an operational protocol to a reasoning task steals from the reasoning. The model became so preoccupied with the trace protocol that it stopped answering the question.
- Runtime-supplied. The version that shipped is an operator grant. Set once in the agent def, applied by the runtime on every recall in scope, invisible to the model. This is the version measured at +24.5pp / +25.9pp.
One consequence of the runtime shape: recall_attach_traces is notOverridable per-run. A per-run override would let a compromised prompt reach transcripts through the memory path around history_scope. The operator decides once, per agent, and no in-run instruction can widen it.
What the cascade changes: the expensive decider is not needed on every call
Trace attachment is not free. A recall that returns fact bodies plus source turns is a bigger payload than one that returns fact bodies alone. If every recall attaches traces on every call, the memory subsystem effectively becomes a slower and more expensive path than reading the transcript, which is what the arc was supposed to avoid.
The cascade is what makes it cheap enough to keep. A recall returns its top-1 similarity score. On a threshold-clear case ("the top hit is well above the rest, the answer is right here"), the cheap path answers the question and stops. On a threshold-unclear case, a verifier reads the retrieved material and decides whether the material actually answers the question the reader was asked. The verifier is a same-class local model on the same store: cheap enough to run on the ambiguous cases, expensive enough that it should not run on every one.
flowchart TB Q["Question"] --> R["Memory recall
top-k with scores"] R --> T{"top-1 score
above threshold?"} T -->|yes: clear case| ANS1["answerer reads
the retrieved material
and answers"] T -->|no: ambiguous case| V["Verifier
same-class local model
reads retrieved material"] V --> D{"can this material
actually answer
the question?"} D -->|yes| ANS2["answerer answers"] D -->|no| ABSTAIN["answerer abstains
with the honest reason"] style ANS1 fill:#e8ffef,stroke:#4a9e60 style ANS2 fill:#e8ffef,stroke:#4a9e60 style ABSTAIN fill:#fff8e0,stroke:#c99a3c style V fill:#e8f4ff,stroke:#4a90e2
Two efficiency numbers are worth naming. 9 percent of calls buy 40 percent of the gain. 39 percent of calls buy 60 percent. That is the shape of the win: a small fraction of questions are worth the verifier's time; the rest can be answered from the threshold alone. Running the verifier on every question would cost roughly 10× as much for almost no additional accuracy.
One ruled-out shape worth naming, because a lot of teams would reach for it first: set-homogeneity gates (measure how much the top-k results agree with each other; skip verification when they agree strongly, run it when they disagree). The measurement on the same corpus: homogeneity correlates with top-1 at r ≈ −0.45, so it is roughly redundant with the threshold already in use, and adding it costs 0.005 accuracy out of fold. It looked promising and did not survive contact with the data.
What RFC DM changes: derived search units
The three RFCs above sit on top of a retrieval index that was still, structurally, one embedding per chunk body. A chunk with a heading, a code block, and inlined transclusions embedded once, so a query for the code block competed with the heading for the same vector. On the reference deployment, 186 of 3000 heading-only chunks were invisible before this arc, because they carried no body content and produced no meaningful embedding.
RFC DM (landing in v1.100, unreleased) separates "the chunk" from "the search unit derived from the chunk." One chunk generates multiple derived search units: a heading path, a code block, a transclusion each get their own embedding and their own row in the resolution index. Resolution knows which chunk each unit belongs to, so a hit on a code block returns the whole chunk with the code block highlighted rather than the code block alone.
flowchart LR CHUNK["a Document chunk
heading + prose + code
+ transclusion"] --> DERIVE["derive search units"] DERIVE --> U1["unit: heading path"] DERIVE --> U2["unit: prose body"] DERIVE --> U3["unit: code block"] DERIVE --> U4["unit: transclusion"] U1 --> EMB1["embed"] U2 --> EMB2["embed"] U3 --> EMB3["embed"] U4 --> EMB4["embed"] EMB1 --> IDX["resolution index
each unit points
back to the chunk"] EMB2 --> IDX EMB3 --> IDX EMB4 --> IDX IDX --> HIT["query hit on ANY unit
returns the whole chunk
with the matching unit highlighted"] style CHUNK fill:#f8ecff,stroke:#8b6cbf style HIT fill:#e8ffef,stroke:#4a9e60
Storage and resolution shipped in RFC DM C1. The document-side generator that decides which units to derive per chunk shipped in C2 alongside a fresh set of benchmark runs under bench/results/locomo/2026-09-24-rovemark-rating. The comparison to the pre-DM state is what the C2 report will publish; expect it in the release notes for v1.100.
What we measure now
A summary table of what the retrieval subsystem now reports, on real corpora, at v1.93.0.
| Measurement | Number | Source |
|---|---|---|
| Facts arm accuracy (paired, 199 questions) | 0.157 | MEMORY-ARCHITECTURE.md |
| Turns arm accuracy (paired, 199 questions) | 0.738 | same |
| Temporal slice (facts / turns) | 0.036 / 0.873 | same |
| Paired McNemar p | 2.4e-29 | same |
| Compression (LongMemEval / LoCoMo) | 7.2:1 / 9.0:1 | same |
recall_attach_traces gain (LongMemEval) | +24.5pp | RFC DF report |
recall_attach_traces gain (LoCoMo) | +25.9pp | same |
| Verifier recall on LoCoMo unanswerables (57/59) | 0.966 | rovemark-rating rig |
Verifier recall on LongMemEval _abs (29/29) | 1.000 | same |
| Cascade efficiency (calls that pay for gain) | 9% → 40%, 39% → 60% | same |
| Cross-session synthesis, traversal (RFC DB) | 58.6% | bench/synthesis |
| Cross-session synthesis, single-hop (control) | 5.8% | same |
| Cross-session synthesis, shuffled-relation (control) | 0.8% | same |
| Cross-user duplication (2-user corpus) | 30 → 23 | RFC CV report |
| Facts shared across users after placement | 0 → 94 | same |
graph_recall predicate speedup (2,942 rows) | 928 ms → 30 ms | RFC BU §6 |
Two of these deserve a note. The cross-session synthesis row (RFC DB) uses a benchmark that isolates whether traversal-quality is coming from the relations or from the volume of facts in prompt. Same code, same depth, same 28 facts, only edge targets permuted: 58.6% falls to 0.8%. That is a good rule-out. It says the traversal is doing what it looks like it is doing, not just amortising over a large context.
The graph_recall predicate speedup row is a bug fix wearing a benchmark hat. When document chunk bodies started embedding into the same per-scope vector plane (RFC BU §6), the guarantee that Memory recall does not reach documents quietly broke: a memory recall would return document hits and thrash on unrelated data. The predicate now excludes documents from that call. Restoring the guarantee also restores the speed.
What this costs, honestly
Three limits are load-bearing enough to name.
The retrieval tier is Postgres-only. SQLite has no working vector tier in practice, so a deployment that runs on SQLite reads the store as if it were pre-arc: recall on strings only, no derived search units, no cascade verifier gate. The rest of loomcycle runs fine on SQLite (single-node development, laptop-local agents), but the retrieval quality numbers in the table above do not apply. Postgres or one of the other supported vector-capable backends is required.
The verifier threshold is corpus-dependent. The 0.966 and 1.000 recall figures come from LoCoMo and LongMemEval respectively, each with a calibrated top-1 threshold. A different corpus with different similarity-score distributions needs its own calibration. Do not lift the LoCoMo threshold to a new deployment and expect the same recall; run the calibration first. The cascade wins because the numbers are honest, not because there is a magic cutoff.
Memory add with infer=true does not store a retrievable row. This is called out as "the single most surprising property" in docs/MEMORY-ARCHITECTURE.md, and it will bite a deployment before anything else in this post does. The default write path enqueues content for the consolidator, which extracts facts asynchronously; the row is not queryable until the consolidator has run. A deployment that expects a read-after-write memory (write a fact, immediately recall it) needs either infer=false to bypass the consolidator, or a wait on the consolidator's cursor. And a deployment that never turns on context.recall reads only the facts arm, so it lives at the 0.157 number until it turns on the trace grant.
What this bought
The one-sentence version: a structured memory subsystem that is honest about what it loses when it compresses, and that closes the gap by reattaching what it lost only when the answer would have been worse without it. The facts-alone number stays 0.157; that is what facts alone can do. What ships is a runtime that does not have to run facts alone.
The Sept 3 architecture post drew the shape of the memory subsystem. This post fills in what actually happens when the shape meets a real corpus and gets measured. The memory testing story is where the arc began.
Follow-ons queued. A source selector on Recall (distinguishing spans harvested from prior runs from operator-authored facts) is the shape RFC BW gave the Memory search API and is being retrofitted here. The team-orchestrator Σ handoff (from the same distillation post) still moves strings end to end and lands next. Loomboard canvas primitives for agent-team construction and debugging (RFC CY / RFC DJ) are in progress and are the subject of the next arc.
Reference material: docs/MEMORY-ARCHITECTURE.md in the loomcycle repo for the paired-benchmark methodology and per-slice numbers; bench/results/locomo/2026-09-* for the RFC DM C2 runs; bench/synthesis/ for the RFC DB isolation study; docs/PATH.md for the tree the subject-Documents sit in.