Deletion Propagation for Production RAG in October 2026: Remove Deleted Source Documents From Vector Indexes, Caches, and Logs
Most RAG pipelines are built to add knowledge. A connector pulls new documents, a chunker splits them, an embedding model turns the chunks into vectors, and the vector store grows. Very few teams design the reverse path with the same care. Then someone deletes a confidential file from the shared drive, a customer asks you to erase their data, or legal pulls an outdated policy, and the assistant keeps quoting it for weeks.
This guide covers how to make deletions flow through a production RAG system: from the source of truth, through the chunk and vector layers, into caches, and out of the places people forget, like evaluation sets and logs. The goal is a system where "deleted upstream" reliably means "gone from answers" within a time window you can state out loud.
Why deletes are harder than inserts
An insert has one obvious output: new vectors in the index. A delete has to find every derived copy of a document, and a typical pipeline makes a lot of them:
- Chunks. One source file usually becomes dozens or hundreds of chunks, each with its own ID.
- Vectors. Each chunk has an embedding, sometimes several if you run hybrid search or keep an older embedding model around during a migration.
- Keyword indexes. If you use BM25 or a full-text engine alongside vectors, the text lives there too.
- Caches. Semantic caches, exact-prompt caches, and reranker caches can hold answers that quote the deleted text.
- Summaries and memories. Agent memory stores, document summaries, and pre-computed FAQ answers often paraphrase the source.
- Logs, traces, and eval sets. Retrieved passages get written into request logs, tracing tools, and golden datasets.
If any one of these keeps the content, the deletion is incomplete. The fix is not a clever delete query. It is bookkeeping: knowing, for every source document, exactly what was derived from it.
Step 1: Give every source document a stable ID and keep a lineage table
Start by assigning each upstream document a stable identifier that does not change when the file is renamed or edited. Most sources already have one: a Google Drive file ID, a Confluence page ID, a database primary key, or an S3 object key plus bucket. Use that, prefixed with the source system, for example gdrive:1AbC... or confluence:482913.
Then keep a small lineage table, in an ordinary relational database, that maps each source ID to everything derived from it:
source_id | chunk_id | index_name | embedding_model | ingested_at
gdrive:1AbC | gdrive:1AbC#0001 | docs_v3 | embed-2026-08 | 2026-10-01T14:02Z
gdrive:1AbC | gdrive:1AbC#0002 | docs_v3 | embed-2026-08 | 2026-10-01T14:02Z
gdrive:1AbC | gdrive:1AbC#0001 | docs_bm25 | n/a | 2026-10-01T14:02Z
Two habits make this table pay off. First, build chunk IDs deterministically from the source ID plus a position or content hash, so you can reconstruct them even if the lineage table is damaged. Second, store the source ID as metadata on every vector and every keyword-index document. That gives you a second way to find derived records with a metadata filter when the lineage table is missing something.
Step 2: Detect deletions at the source
You cannot propagate a delete you never hear about. There are three common ways to detect them, and production systems often combine two.
Change events
Many SaaS sources offer change feeds or webhooks that include delete or trash events. These are the fastest path, often seconds to minutes. Treat them as hints rather than guarantees: webhooks get dropped, and some systems report "moved to trash" differently from "permanently deleted." Decide which of those should remove content from search. For most teams, trashed content should disappear from answers right away.
Periodic reconciliation
On a schedule, list every source ID the connector can currently see and compare it with the source IDs in your lineage table. Anything in your index that the source no longer returns is a candidate for deletion. This catches missed webhooks, permission changes that hide a document from your service account, and connectors that never supported delete events at all.
Run reconciliation with a safety brake. If a sync suddenly reports that 60 percent of documents vanished, the likely cause is an expired credential or an API outage, not a mass deletion. Set a threshold, for example "more than 5 percent of a source in one run," above which the job pauses and alerts a human instead of deleting.
Explicit erasure requests
Privacy requests, such as a right-to-erasure request under GDPR, may target a person rather than a file. Those need a lookup step: find every source document that mentions or belongs to that person, then feed those source IDs into the same deletion path. Keep that mapping logic separate and well tested, because it is the part most likely to miss something.
Step 3: Delete in a fixed order, with a tombstone first
When a deletion is confirmed, write a tombstone record before touching any index:
{
"source_id": "gdrive:1AbC",
"reason": "upstream_deleted",
"requested_at": "2026-10-06T22:10:04Z",
"status": "pending"
}
The tombstone does two jobs. It lets the retrieval layer filter the document out immediately, before the slower cleanup finishes. And it stops a late-arriving ingestion job from re-adding the document, which is a surprisingly common bug when a re-index and a delete run at the same time.
Then work through the derived stores in a predictable order:
- Query-time filter. Add the source ID to a deny list that the retriever checks on every request. This is your instant kill switch.
- Vector index. Delete vectors by chunk ID from the lineage table, then run a metadata-filter delete on
source_idto catch strays. - Keyword index. Same pattern: delete by ID, then by source metadata.
- Caches. Invalidate any cache entry whose answer cited the source. This is much easier if cache entries store the list of source IDs they were built from.
- Derived artifacts. Regenerate or delete summaries, memories, and FAQ answers that used the document.
- Raw copies. Remove the stored original and extracted text from object storage.
Mark the tombstone complete only after every step reports success. If a step fails, retry it; every step should be idempotent, so running it twice does no harm.
Step 4: Understand what "deleted" means inside your vector store
Many vector databases and search engines do not physically remove data the moment you call delete. Graph-based indexes such as HNSW, and segment-based engines in general, often mark records as deleted and skip them at query time, then reclaim space later during compaction or segment merges. For search results that is fine: a soft-deleted vector should not come back. For a privacy or legal requirement, it may not be enough on its own.
Read your vector store's documentation on deletes, compaction, backups, and snapshots, and write down the answers to three questions:
- After a delete call, can the record still appear in any query result?
- When is the data physically removed from disk, and can you trigger that?
- How long do backups and snapshots that contain the record live?
Those answers become the honest deletion window you can promise, for example "removed from answers within 15 minutes, purged from primary storage within 7 days, and from backups when they expire after 30 days."
Step 5: Don't forget logs, traces, and evaluation data
Request logs that capture full retrieved passages are useful for debugging, and they are also copies of your documents. A few practical options:
- Log references, not text. Record chunk IDs and source IDs in traces, and fetch the text on demand from the live index. A deleted chunk then simply cannot be fetched.
- Short retention for full payloads. If you must log full prompts, keep them for days, not months, and document the period.
- Tag eval items with source IDs. Golden questions built from a document should carry that document's ID, so the deletion job can retire or rewrite them.
Step 6: Test the delete path like any other feature
Deletion bugs are silent. Nobody files a ticket saying "the bot didn't mention the file I deleted." So build tests that prove it works:
- Ingest a canary document that contains a unique, made-up phrase, such as a random string.
- Confirm a query for that phrase retrieves it.
- Delete the document at the source.
- Poll retrieval until the phrase stops appearing, and record how long it took.
- Search every derived store, including caches and logs, for the phrase.
Run this end to end in staging on every pipeline change and on a schedule in production. Alert when the time-to-disappear exceeds your stated window or when the phrase shows up anywhere after cleanup.
A practical checklist
- Every chunk, vector, and keyword record carries a stable
source_id. - A lineage table maps sources to every derived record and cache entry.
- Deletions are detected by events and by scheduled reconciliation with a mass-delete safety brake.
- A tombstone and query-time deny list hide content immediately and block re-ingestion.
- Cleanup steps are ordered, idempotent, and retried until complete.
- You know how your vector store handles soft deletes, compaction, and backups.
- Logs and eval sets reference IDs instead of storing full text where possible.
- A canary test measures time-to-disappear and alerts on regressions.
None of this is glamorous, but it is what lets you tell a customer, a security team, or a regulator exactly how long deleted content can live in your AI system. Build the delete path at the same time as the ingest path, and it costs very little. Retrofit it after an incident, and it costs a lot more.
Comments
Post a Comment