Skip to main content
loomcycle
§ story · documents indexing

When one embedding per chunk stops working.

Any system that does semantic search over documents has to decide what to embed. The obvious choice is one embedding per chunk. Split the document into meaningful pieces, embed each piece as a vector, run cosine similarity at query time. That is how most retrieval systems work. It is how ours worked for months.

Then we noticed some chunks were invisible.

On our reference test store, out of 3,000 chunks, 186 of them never came back in any query. Not ranked poorly. Never returned. When we looked closer, the pattern was clear: every invisible chunk was a heading-only chunk. A chunk whose body was empty because the meaningful content was in the sub-chunks below it.

186 / 3,000
chunks invisible before
928 → 30 ms
memory recall side-effect fix
v1.100
where this lands

That was on us. When you embed the empty body of a chunk, you get either an empty vector or a garbage one, and either way, no query matches it. The heading, which was the meaningful content of the chunk, was never seen by the embedder because the embedder only saw the body.

Fixing that part was one line: a heading-only chunk should embed its title. Done. The 186 invisible chunks came back.

The bigger issue was the chunks that were not invisible.

What a real chunk actually looks like

Here is a chunk from a real project handbook:

### How to run the deploy script

Before you run this, make sure your `DEPLOY_TOKEN` is set in `.env.local`.

```bash
./deploy.sh --env production
```

See also: [[rollback-procedure]]

This chunk has four things in it:

When you embed this chunk as one vector, all four things get averaged into a single point in the embedding space. Now imagine two queries:

The problem is not that the retrieval algorithm is bad. The problem is that the chunk is not one thing, and treating it as one thing loses the parts of it that do not dominate the average.

What we did

We stopped generating one embedding per chunk. We now generate one embedding per derived search unit. A search unit is a piece of a chunk that has its own semantic identity. For the chunk above, that means four search units:

Each unit gets its own embedding. The resolution index knows which chunk each unit belongs to. If a query matches the code unit, the result is the whole chunk with the code block highlighted. Not the code block alone. The reader gets context; the search gets precision.

flowchart TB
  subgraph BEFORE["Before"]
    C1["one chunk
heading + prose + code + transclusion"] --> E1["one embedding
average of everything"] E1 --> Q1["query matches whichever piece
dominates the average"] end subgraph AFTER["After"] C2["one chunk
heading + prose + code + transclusion"] --> DERIVE["derive 4 search units"] DERIVE --> U1["heading unit → embed"] DERIVE --> U2["prose unit → embed"] DERIVE --> U3["code unit → embed"] DERIVE --> U4["transclusion unit → embed"] U1 --> IDX["resolution index
each unit → its chunk"] U2 --> IDX U3 --> IDX U4 --> IDX IDX --> Q2["hit on ANY unit
returns whole chunk
with matching unit highlighted"] end style BEFORE fill:#fff5f5 style AFTER fill:#f0fff4 style Q1 fill:#ffe8e8,stroke:#c86a5c style Q2 fill:#e8ffef,stroke:#4a9e60
Same chunk, two indexing strategies. The averaged version loses parts of the chunk that do not dominate. The derived version has one chance per piece to match.

The numbers so far

The 186 heading-only chunks that were invisible before are now findable. That is a bug fix, not a surprise.

The more interesting numbers are the ones we cannot fully quote yet, because the retrieval-quality report for derived search units lands with the next release. It runs the same paired benchmark we describe in the memory retrieval post: same corpus, same reader, same evaluator, same questions. What earned this the priority slot in the schedule is that the internal runs on the reference deployment showed a real improvement on queries that hit mixed-content chunks, without hurting queries that hit prose-only chunks. The public report will name the numbers with the release.

One number we can quote from the fix itself: on a chunk that carries a heading plus a code block plus a paragraph, a query for the code returns the whole chunk. Not just the code line. The chunk is the unit of context; the search unit is what matched. Keeping those two ideas separate is what the design costs and what it buys.

The side effect nobody asked for

Along the way, we caught a bug that was costing us 900 milliseconds per memory recall on a 2,942-row scope.

Here is what happened. A while back, we started embedding document chunk bodies into the same per-scope vector plane as the memory tier. That was good for document search. It also meant that Memory recall, which is supposed to reach only memory facts, was accidentally reaching into the document embeddings too. So a query that should have returned three facts about a person came back with those three facts plus a dozen document chunks that mentioned similar terms. The agent got noisier context, and the query took a lot longer.

We added a predicate that excludes documents from Memory recall. On the reference store, that took the recall from 928 milliseconds to 30 milliseconds. About 31 times faster.

flowchart LR
  subgraph BEFORE2["Before"]
    Q3["Memory recall query"] --> VP1["one per-scope
vector plane"] VP1 --> F1["fact rows"] VP1 --> D1["document chunk rows
accidentally reached"] F1 --> R1["noisy result set
928 ms"] D1 --> R1 end subgraph AFTER2["After"] Q4["Memory recall query"] --> VP2["one per-scope
vector plane"] VP2 -->|"predicate:
exclude documents"| F2["fact rows only"] F2 --> R2["clean result set
30 ms"] end style BEFORE2 fill:#fff5f5 style AFTER2 fill:#f0fff4 style R1 fill:#ffe8e8,stroke:#c86a5c style R2 fill:#e8ffef,stroke:#4a9e60
The memory-recall predicate fix. Same vector plane, filtered at query time so memory recall reads only what its name says.

The speedup is not the interesting part. The interesting part is that Memory recall now does what its name says. A caller asking for facts about a subject does not have to explain that they do not also want document chunks about vaguely similar things.

What this does not fix

Two things worth naming, because they will bite a deployment before anything else in this post does.

First, not every chunk has clear semantic pieces. A chunk that is one long paragraph of prose does not benefit from derived search units, because there is nothing to derive. It still embeds as one vector. That is the correct behaviour, but a system that promises "multiple embeddings per chunk" might imply it works on everything. It works on structured chunks. Homogeneous prose stays where it was.

Second, this all runs on Postgres. Our SQLite deployments do not have a working vector tier in practice. If you are running loomcycle on SQLite, the derived search units are computed but not queryable through the vector path. The full-text path still works, and full-text on a code block often does fine anyway, but the retrieval-quality numbers we are about to publish do not apply.

What comes next

The derived search units land in the next release. The retrieval-quality report will land with it, running the same paired benchmark we used for memory. The near-term follow-on is teaching the derive step to be smarter about which pieces of a chunk are worth their own unit and which pieces are noise. A chunk with three sentences and a footnote does not need four search units. That is a calibration knob we have not finished tuning.

The whole memory-and-documents arc we started in September is coming to a natural close with this. The next arc is about agent teams: state machines for multi-agent workflows, review gates for team outputs, and a visual canvas for constructing and debugging them. Different story, coming next.

Companion reading: what structured memory costs, and what closes the gap; the memory architecture drawn; the memory testing story; compact context, still recall.