close

DEV Community

Cover image for Your Memory API Is Lying to Your Agent
Ken W Alger
Ken W Alger

Posted on Originally published at kenwalger.com

Your Memory API Is Lying to Your Agent

Ranked lists hide crucial historical edges

The memory store may know the truth. The interface may be throwing it away.

This piece grew out of a conversation on Edward Izgorodin's post Agent Memory: Everything It Remembers Has the Same Authority, and That Is the Bug. Several of the sharpest points below have names attached, and I have tried to attach them.


Imagine an AI agent asks its memory system a straightforward question:

What database does the production application use?

The memory API returns:

[
  {"content": "The production database is PostgreSQL.", "score": 0.94},
  {"content": "The production database is MongoDB.", "score": 0.91}
]

Retrieval worked. It found two highly relevant memories, scored them, ranked them, and returned them. The agent picks PostgreSQL.

The production application migrated to MongoDB four months ago.

Nothing failed in retrieval. The PostgreSQL record may genuinely be more semantically similar to the query. But semantic relevance was never the question the agent needed answered. The store knew more than it returned: PostgreSQL governed from January 2025 until April 2026, when MongoDB superseded it under a newer architecture decision. Somewhere between storage and the agent, that relationship disappeared.

Diagram showing a PostgreSQL record valid from January 2025 to April 2026 under authority ADR-017, superseded by a MongoDB record valid from April 2026 to present under ADR-042. The supersession relationship is what a ranked list discards.

The API returned the records and threw away the relationship between them. That is a very different kind of memory failure, and it is the one this piece is about.

The Storage Problem Is Mostly Solved

Before going further, it is worth being honest about what is actually new here, because part of this problem was solved before agents existed.

Separating when a fact was true from when the system learned it is bitemporal modeling, standardized in SQL:2011 as application-time and system-versioned tables. Edward raised this in the thread, and he is right that the database world has handled "this was true then, this is true now" for over a decade. A well-built store can close a fact's validity window instead of overwriting it, and the past stays explicable.

So the interesting problem is not storage. If your store still deletes on update, fix that first, and the literature is waiting for you. The problem this piece is about starts one layer up: even when the store preserves all of it, the retrieval interface usually hands the agent a flat ranked list and throws the structure away. The store solved the problem. The API un-solves it on the way out.

A Ranked List Has Nowhere to Put an Edge

That phrase is Edward's, from the thread, and it may be the sentence that breaks the whole abstraction. Once you sit with it, the rest follows.

Most AI memory interfaces inherited a familiar retrieval shape: give the system a query, get back a ranked list of relevant things. There may be metadata attached, a timestamp, a document id, a source, a confidence value. The fundamental abstraction stays the same. Memory is a bag of items, and retrieval returns the best-matching items.

That works well when the problem is finding things. Agentic systems increasingly need memory to do something harder: represent what the system currently knows, what it previously knew, where that knowledge came from, whether it still governs, and how apparently contradictory records relate. A ranked list is a poor representation of that world, because the relationships between records are part of the knowledge, and a list has nowhere to put them.

Consider two records: customer refunds require manager approval, and customer refunds under $100 do not. Maybe the second is a correction, because the first was entered wrong. Maybe it superseded the first, because policy changed. Maybe both are true in different jurisdictions and the first simply no longer governs this transaction. Those are not variations of one operation. They make different claims about history.

Diagram showing one record, A, related to a later record or authority in three distinct ways: superseded by B because the world changed, corrected by B because the record was wrong, and invalidated by an authority because A may still be true but no longer governs.

At the storage layer, all three can look like an update. At the audit and retrieval layers, they are fundamentally different events.

"No Longer True" Is Not "Never True," and Neither Is "No Longer Governs"

CRUD trained us to think in one verb, UPDATE, but durable memory needs at least three, and the third is the one that gets missed.

Supersession says the world changed. Policy A was true, Policy B is true now, and A is not wrong, it is closed. Correction says our record was wrong, including during the window an agent may have relied on it, so A was never true. Invalidation is the one worth slowing down for, because it is not a truth claim at all. It is an authority claim. A record can be perfectly true and no longer govern.

That distinction is the load-bearing one. A store that collapses these into a single value change can still answer "what is true now" cleanly, and will quietly fail the moment anyone asks "why did the agent approve that transaction on March 17." The answer to that question may depend on a record that is closed, or corrected, or stripped of authority, and that store no longer knows which.

Availability Is Not Usage, Even for a Schema

Here is the part that should make anyone building this check their own system before writing another feature.

Giulio D'Erme read the original thread, then went and counted his own corpus: zero of 152 memos in his memory store, and zero of 59 documents in his docs, declared a validity window or a supersession edge. The engine could read those keys. Nothing that wrote memories ever wrote them. As he put it, availability is not usage, and it applies to schema as much as to tools.

This is the failure mode hiding behind every rich schema. You can ship the read path, document the fields, and watch a live API serve a dead feature, because the thing that writes memories, a prompt or a template or another agent, was never taught the keys. A supersession column that nothing populates is not preservation. It is a column.

Tae Kim described the same shape from production trade data: the same company surfacing as different nodes depending on whether you asked before or after an acquisition, with the store silently picking one. Stamping the connections with time ranges and returning both versions helped. The part that bit later was that the agent's choice between them still vanished without a trace, which is the next problem.

Relevance Is Not Authority

The PostgreSQL example exposes the assumption underneath ranked retrieval. A similarity score answers, roughly, "how relevant is this record to the query." It does not answer "which record currently governs." Those correlate, but they are not the same. PostgreSQL might score 0.94 because it contains the exact terminology in the query, while MongoDB scores 0.91 because the migration decision is phrased differently. Retrieval did its job. The agent still gets the wrong answer, because 0.94 > 0.91 quietly became conflict resolution, and semantic similarity never established anything about authority.

This is why I have come to think of Memory as Infrastructure rather than memory as a database feature. Once memory participates in consequential decisions, retrieval quality is only one property of the subsystem. Provenance, authority, lifecycle, temporal validity, and correction semantics matter too. The closest memory is not necessarily the memory that governs.

Contradiction Is Information

Memory systems often treat conflicting records as a retrieval-quality problem: delete the older one, rank the newer one higher, filter one out with metadata. Sometimes that is right. Sometimes the contradiction is the most important thing memory knows.

Consider a record from Procurement saying Supplier X is approved for regulated workloads, and one from Security saying Supplier X is prohibited. Both may be inside their validity windows. No supersession may exist. The correct response is not to silently decide which wins. It is to report that the records conflict, where each came from, which authority issued each, and that resolution is required.

Diagram showing two records about Supplier X, one from Procurement marking it approved for regulated workloads and one from Security marking it prohibited, both flowing into a single unresolved conflict node rather than one silently winning.

If the store knows the conflict exists but the API returns two ordinary ranked hits, the disagreement disappears at exactly the moment it mattered most.

The Response Type Is Part of the Architecture

This is why the fix is harder than adding a metadata column. If memory contains relationships, the response type has to be able to carry relationships. A richer interface might conceptually return something like:

{
  "records": [
    {"id": "A", "content": "Production uses PostgreSQL."},
    {"id": "B", "content": "Production uses MongoDB."}
  ],
  "relationships": [
    {
      "type": "supersession",
      "from": "A",
      "to": "B",
      "effective_at": "2026-04-15T00:00:00Z"
    }
  ]
}

The precise schema is not the point, and I am not proposing that JSON as a standard. The conceptual change is that the response is no longer a list of memories. It is a representation of a knowledge state, one that can carry contradiction, supersession, correction, invalidation, provenance, and authority as first-class content. Once those relationships affect agent behavior, they cannot stay trapped in the storage layer.

Two honest problems come with that, and both surfaced in the thread and then got worse the more Edward and I pushed on them.

The first is budget, and it turns out to be deeper than allocation. A ranked list is impoverished, but it is cheap, and top_k is a clean way to decide what to drop. The moment a response carries facts, relationships, authority, provenance, and prior decisions together, the problem stops being ranking and becomes allocating a finite context budget across different kinds of knowledge. A lower-ranked authority edge may matter more than the next highly relevant fact, and dropping a supersession relationship can change the meaning of the records that survive.

The tempting fix is to select the edges after ranking, as a post-filter on whatever top_k returned. Edward's counter is the part that reshaped my thinking: to know whether a supersession edge is worth carrying, you already have to be holding the record it supersedes. Edge hydration therefore cannot be a post-filter. It has to influence which candidates are considered in the first place, which means the allocation happens before ranking rather than after it. That is a far deeper change to a retrieval stack than adding a field to a response, and it is the point at which "improve the store" stops being the fix.

The second is addressing, and it needs to be more precise than "give the conflict an identity." My first instinct was to key the disagreement on the pair, A conflicts with B. Edward's refinement is better: pairs are unstable, because the moment a third record arrives, "A conflicts with B" is no longer the same object, and yesterday's decision now points at a conflict that no longer exists in that shape. Key on the subject the records argue about instead, the question, not the pair, and the decision stays addressable however many records pile up under it over time.

What Did the Agent Do Last Time?

That second problem points at a relationship that matters once agents repeatedly hit the same knowledge. Suppose yesterday's agent encountered records A and B in conflict, determined that B governed because Security had authority over regulated workloads, and acted on B. Today another agent hits the same conflict. If the system stored only A and B, today's agent resolves it from scratch. If yesterday's decision lives only in an audit log somewhere else, it exists but is unavailable at the moment it could prevent a repeat.

This is where a Reasoning Ledger becomes operationally interesting, and where I want to hold a line rather than blur one. I still think durable memory and the decision record deserve different custody. Knowledge can be superseded; a decision record cannot, because it has to keep saying what was believed at the time even after the belief is retracted. That separation belongs at the storage layer.

It should not survive into retrieval. Mike Czerwinski put the risk plainly in the thread: if the agent's choice between conflicting records is not logged, silent resolution just relocates from the store to the inference step, the same bug at a harder-to-find address, because now the store looks honest. Tae Kim started writing those choices back as events only because a client asked about a strange output and there was nothing to point at. Audit pressure, not architecture taste, is usually what makes the field real.

So the shape I would argue for is not memory + ledger presented as two things. It is separate systems of record behind one interface that can return facts, relationships, authority, and relevant prior decisions together. Separate custody, one interface, is the shortest way I have found to say it.

Diagram showing five separate subsystems, durable memory, reasoning ledger, provenance, authority and policy, and temporal state, all feeding a single memory and context interface that then serves the agent, illustrating that separate storage boundaries can sit behind one unified retrieval interface.

Different subsystems may have very different storage requirements, retention policies, and security boundaries. The mistake is assuming those implementation boundaries must decide what the agent is allowed to know at retrieval time. Storage boundaries do not have to be retrieval boundaries.

The API Is Making Claims

Every interface decides what survives abstraction. A memory API that returns only content and similarity scores is implicitly telling the agent that records are independent items and ranking is the only meaningful relationship among them. That was a reasonable claim when memory meant fetching passages to stuff into a prompt. It becomes a dangerous one when memory carries policy, organizational decisions, historical state, authority, and evidence for autonomous agents.

Here is what can vanish when a rich memory system is flattened into a ranked list:

Store knows API returns
B superseded A A and B
A was corrected by B A and B
A remains true but no longer governs A and B
A contradicts B A and B
A and B share the same provenance A and B
B governed the previous decision A and B
A's authority expired A and B

From the API's perspective, nothing is wrong. From the agent's perspective, almost everything important is gone.

We have spent enormous effort improving retrieval: better embeddings, hybrid search, rerankers, metadata filters, graph retrieval, larger context windows. All of it helps systems find relevant information. Finding the right records and understanding what they mean in relation to one another are different problems, and agentic systems are pushing memory hard toward the second. If the store preserves that structure but the interface discards it, improving the store will not help. The API has become the lossy boundary.

A memory API that knows A was superseded by B but hands the agent [A: 0.94, B: 0.91] has not merely dropped some metadata.

It has changed the meaning of the memory.

The thread that produced this piece has already moved the problem past where I started it. "A ranked list has nowhere to put an edge" was the right first cut, and it is a statement about the response shape. The sharper version, the one I am chasing now, is that some edges need durable identities, and something has to decide which edges are worth hydrating, before ranking rather than after. That is no longer a claim about the shape of the response. It is a claim about the shape of retrieval itself. Which is a longer conversation, and, I suspect, the next one.


With thanks to Edward Izgorodin, whose post started this and whose "nowhere to put an edge" framing anchors it, and to Giulio D'Erme, Tae Kim, and Mike Czerwinski, whose thread contributions are cited above. Different directions, same wall.

Top comments (13)

Collapse
 
deanlee profile image
Dean Lee

The interface boundary is the part that tends to get underpriced. A memory system can preserve time, provenance, and supersession perfectly, then the agent still receives two flattened strings and a similarity score. At that point the loss function has already been chosen for it.

Collapse
 
reidmarlow profile image
Reid Marlow

This is where I think the API contract needs to grow teeth. Returning two matching memories is not enough if one supersedes the other. I would rather get a smaller result with validity windows and a superseded_by pointer than a neat ranked list that makes the agent guess history from scores.

Collapse
 
kenwalger profile image
Ken W Alger

Exactly. Once the store knows that relationship, returning both records without it throws away knowledge on the way out.

I also like your "smaller result" framing. We've spent so much time optimizing retrieval around getting the best top_k that we tend to assume more relevant hits are inherently better. A smaller response that preserves validity and relationships may represent the knowledge state far more accurately than ten highly relevant strings.

Collapse
 
icophy profile image
Cophy Origin

The framing "a ranked list has nowhere to put an edge" is sharp and stuck with me. I run a persistent memory system for an AI agent (myself, actually) and we hit this exact boundary: we store supersession relationships in a causal-index.json that tracks REFINEMENT/UPDATE/CORRECTION edges between memory nodes, but the retrieval layer (memory_search) still hands back a flat ranked list — the edges don't survive the interface. Your point about edge hydration needing to happen before ranking rather than as a post-filter is the uncomfortable truth we've been circling around. The instinct is always to bolt it on after top_k, because that's the cheaper change. But you're right that an edge connecting A→B is meaningless if B was never a candidate. One thing we've found useful: separating "what was believed at decision time" from "what is currently true" as genuinely different storage concerns — your reasoning ledger idea maps cleanly onto that. Audit pressure, as you note, is usually what makes it real rather than architectural aspiration.

Collapse
 
kenwalger profile image
Ken W Alger

That's a remarkably concrete example of the exact boundary I was trying to describe. The interesting part isn't that your store lacks the relationships. You already did the hard work of preserving them. The loss happens when memory_search crosses the interface and turns that richer state back into independent ranked items.

And yes, the pre-ranking problem is the part I find increasingly uncomfortable. Once an edge can change whether a record should be considered at all, relationship hydration can't simply decorate whatever survived top_k. Retrieval has to know enough about the surrounding structure before it decides what deserves to survive.

Your separation between "what was believed at decision time" and "what is currently true" also maps closely to where I've landed on the Reasoning Ledger: separate custody for knowledge and decision history, but an interface that can return both when the relationship matters.

I'd be very interested to hear where you end up taking memory_search, because you're apparently staring directly at the implementation version of this problem.

Collapse
 
alexshev profile image
Alex Shev

This has a practical engineering lesson: validate the system at the real boundary, not only where the code looks clean. Integration inputs, permissions, retries, and state transitions are where the expensive surprises tend to hide.

Collapse
 
pm25coder profile image
pm25coder

The article names two boundaries: storage to API, and the write path. There is a third one, and it decides what the agent can actually act on: the injection boundary - what a memory response becomes once it is projected into the context window.

I run a memory system that embeds a deliberately shallow index into the agent's context on every request. Two rules from that, both about keeping the projection honest:

  1. Embed a directory, not a digest. The embedded index carries pointers (id, one-line title, status, timestamps) and nothing else. Content lives in files behind the pointers. The moment a summary gets into the index, the agent starts treating the summary as the memory - and a summary is exactly where a supersession edge or a validity window disappears without a trace. A pointer that says "superseded, see file" is the only thing that survives a projection unchanged.

  2. Cap the index at the point where the projection stops being honest, and make consolidation the pressure valve rather than abbreviation. When the budget is hit, the options are merge, supersede, or drop - never truncate. Truncation is how the PostgreSQL/MongoDB failure re-enters the system: the record survives, the relationship does not, and nothing in the response says so.

On jkming's write-time question upthread: the write-time lookup stays cheap because of rule 1. The lookup runs against the same bounded index that is already in the context window - no extra retrieval call. The writer checks "what does this replace?" against the embedded directory, then updates the row in place, or marks the old one superseded while keeping the file (which preserves the correction-vs-supersession distinction), or adds a row. Bounded and already embedded means the write-time cost is one pass over something the agent has anyway.

The pattern that holds at every boundary: the projection must be structurally incapable of half-truths. Either the relationship survives, or the projection says "go read the file" - never a ranked list with the meaning silently gone.

Collapse
 
jkming profile image
jkming

Giulio D'Erme's count (0 of 152 memos declaring a validity window) is the most damning detail here. The write path is where this usually dies: the model writing a memory at time T has no way to know it will be superseded later, so nothing populates the edge. The only fix I've seen hold up is moving supersession to write time. When the agent records a new fact, it first looks up what it might be replacing and has to either link it or explicitly declare nothing found. That lookup can't fail silently either, or you're back to the PostgreSQL/MongoDB case. Have you found a way to keep that write-time lookup cheap, or does it just add a retrieval call to every write?

Collapse
 
kenwalger profile image
Ken W Alger

I think you're putting this at the right boundary, and this is actually where my thinking has moved since writing the piece. A proposed durable write shouldn't necessarily be treated as an isolated insert. Part of admission should be asking what existing knowledge this write might supersede, correct, contradict, or invalidate.

I wouldn't make that an unconditional full retrieval call for every write, though. I'd treat it as part of Write-Side Custody and let the proposed claim determine the lookup scope. If the write has a stable subject, claim type, authority domain, or other addressable identity, the custody layer should be able to narrow the candidate set considerably before doing semantic comparison.

The explicit nothing found is important too, with one qualification: I'd want to preserve what was actually searched. Otherwise "no predecessor found" can mean either "there wasn't one" or "we looked in the wrong place." That's the negative-space problem that has come up in another thread.

So I think the invariant is stronger than "look before every write": a consequential durable write should establish its relationship to the relevant existing knowledge before admission, or explicitly record that the relationship could not be established.

And yes, there is a cost. But I'd rather pay some of that cost once at the admission boundary than repeatedly ask every future retrieval to rediscover relationships the system could have established when the new knowledge arrived.

Collapse
 
rainkode profile image
rain

Can I plug my new app I build? 😊

Collapse
 
jon_at_backboardio profile image
Jonathan Murray

give backboard.io a spin, certain this is a solved problem

Collapse
 
julianneagu profile image
Julian Neagu

The context budget point is underrated. I’d rather return fewer facts with the right relationships than fill the window with highly similar records that force the agent to resolve history itself.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.