RAG Incremental Updates and Deletion Consistency: Activating Complete Revisions and Excluding Obsolete Evidence

A successful vector write does not complete a RAG knowledge-base update. Chunks, lexical and vector indexes, and answer caches may represent different revisions, while deleting a source file may leave retrieval copies intact. This article separates acknowledgment, visibility and activation, then develops a revision manifest, a prepare–validate–activate protocol and a deletion high-water mark, with failure scenarios and measurable acceptance criteria.

Retrieval-augmented generation, or RAG, retrieves material before a model uses it to answer. Common enterprise failures include answers citing an obsolete policy, an updated section being combined with older chunks, and cached answers surviving source deletion. A verifiable revision-release process addresses these failures. The design below distinguishes platform guarantees from application responsibilities. All examples are constructed; no experiment or measured deployment result is claimed.

Separate acknowledgment, search visibility and business activation

A document update has at least three milestones: the pipeline receives the change, derived indexes make it searchable, and the business permits that revision to serve as evidence. They may occur close together, but a successful API response does not make them identical. Qdrant documents that the default or wait=false acknowledged response confirms receipt; wait=true waits for completion and vector search availability.[1] It still cannot establish that every chunk of a document and every other index is ready.

Elasticsearch's refresh parameter controls search visibility: refresh=wait_for waits for a refresh, whereas refresh=true actively refreshes and adds indexing work.[2] The platforms have different semantics. Check deployed versions, replication and read-consistency settings separately rather than treating a waiting parameter as a cross-system transaction guarantee.

Here, activation means the application's decision to allow a document revision to be used by its retriever. It requires complete content, visible retrieval paths and valid access policy. An older revision may remain available during preparation where business rules permit that overlap. Revoked or explicitly invalid material must instead be denied first; availability does not justify continued use of invalid evidence.

Describe a complete revision through a manifest

Give each document a stable documentId and each source revision an immutable revisionId. Identify chunks by document, revision and position, and retain the source fingerprint, parser version, chunking-rule version, embedding-model identifier and access-policy version. A revision identifier can be an unordered unique value, but a separate reliable source order is needed to establish precedence. Arrival time is not a source revision.

A revision manifest records the expected chunk-identifier set, count, fingerprints, source location, validity and processing status for each index. Twenty stored vectors do not establish completeness: some chunks may have been duplicated while others are missing. Compare identifier sets and fingerprints, and check that parsing has not omitted tables, attachments or material passages. An intentionally empty document needs an explicit state so it cannot be confused with parsing failure.

Build the manifest from the current parsing result rather than historical counts. If a policy changes from twenty chunks to eighteen, revision filtering must exclude the two leftovers from the old revision. Recording a revision on every chunk enables answers to trace back to a complete source version and avoids cleanup based on approximate text matching.

Make prepare–validate–activate a recoverable protocol

First, prepare the revision. Write it to a separate chunk namespace or revision-tagged records instead of overwriting serving chunks in place. Deterministic chunk identifiers let repeated submission of identical content converge; different content under the same revision identifier should raise a conflict. Existence in storage does not yet make the new data eligible for retrieval.

Second, validate it. Reconcile the manifest with actual chunk sets, confirm visibility through all required lexical and vector paths, check access metadata, and run queries covering changed clauses. Probe queries establish only the paths they exercise. Manifest reconciliation and platform consistency semantics are still needed to assess completeness. Failed validation leaves a pending state; whether the old revision remains available depends on its business validity.

Third, activate it. An authoritative metadata store maintains activeRevision for each documentId. Use a transaction or conditional update to recheck source precedence, deletion state and access policy at cutover, preventing an older job from overwriting a newer activation. If the condition fails, reread state rather than blindly overwrite it on retry. A policy package whose documents must become effective together can use a release generation selecting all revisions. Individual document pointers do not provide a package-wide consistent snapshot.

Clean up old revisions after cutover. Retaining them supports diagnosis and rollback, but costs storage and requires continued access control. Rollback is itself a business activation: verify that the older content remains valid. Technical recoverability is not permission to restore it. High-throughput systems can prepare and validate in batches and reduce pointer-switch frequency to contain refresh and coordination costs.

Enforce revision eligibility and handle cutover during a query

Before passing evidence to the model, check every candidate chunk against the allowed revision, deletion state and access policy. Filter early where the retrieval platform supports it. If it cannot express a dynamic revision map, retrieve candidates, check them against the authoritative manifest and replenish eligible results. Post-filtering reduces effective recall, so bound replenishment and report insufficient evidence instead of filling gaps with obsolete chunks.

A query may straddle activation. One option binds the request to an explicit release generation and uses the same revision set throughout. Another rechecks revisions before context assembly and retrieves again if needed. The first requires retained version snapshots; the second adds latency and retries. Either should prevent mixing two revisions of the same document in one answer. A request snapshot cannot override a later access revocation, which requires its own mandatory check.

Revision and permission are related eligibility conditions. The article on permission-aware enterprise RAG explains how identity, evidence and access scope connect; the additional concern here is maintaining revision completeness during updates. If the authoritative revision or revocation service is unavailable, high-risk knowledge should fail closed or be routed to a person. A policy must explicitly govern whether ordinary material may use a degraded response with a stated cutoff time.

Deny deleted content first, then clean up derived copies

A missing source file does not establish that its indexed material is gone. Azure AI Search's storage-indexer documentation distinguishes change detection from deletion detection, requires a deletion policy from the first run, and requires processing within the retention window when native soft deletion is used.[3] These are conditions for that indexer, not capabilities automatically supplied by every RAG platform.

We recommend recording deletion or revocation in authoritative state first, with an ordered deletion high-water mark: the latest deletion revision processed for that document. Retrieval uses it to deny older revisions while cleanup removes vectors, lexical records, stored chunks, derived summaries and related caches. Each cleanup target needs a receipt and retry state. Query exclusion and completion of physical-copy cleanup are separate acceptance outcomes.

The high-water mark prevents a late older update from reviving a document only when source revisions are comparable. A controlled sequence or another explicit order can work within one source. Timestamps or log positions from unrelated systems cannot simply be compared. Retention must cover the permitted replay horizon. Very old history, incomparable ordering or a new source generation calls for quarantining the job and rebuilding a baseline rather than guessing precedence.

Restoration requires an explicit new source action and renewed permission and revision validation. An old upload still waiting in a retry queue is not a restoration request. Text already delivered to a reader cannot be technically withdrawn, so specify whether a revocation promise covers subsequent evidence use, cleanup of particular stored copies, or both.

Answer caches and conversation context are derived data too

Cache keys need to cover query meaning, tenant, access scope and relevant revision information. Entries should retain evidence dependencies, which are checked for validity on a hit. A model version alone cannot detect a changed policy. Sharing entries solely by query text can also confuse different users' access scopes.

When dependencies are numerous, invalidating an entire release generation is simpler but reduces hit rate and increases regeneration cost. Precise document-based invalidation requires reverse-dependency tracking. Cache hits must not bypass revocation checks. Recheck dependencies before the next generation in a long conversation; remove affected context or restart the conversation where necessary. Deleting a vector does not erase text already supplied to the model.

Include query-result caches, reranking results and pregenerated summaries. A summary combining several documents needs traceable dependencies or conservative batch invalidation. Record separate retention, cleanup and restore rules for logs and backups. Online retrieval denial is not evidence that every historical copy has disappeared.

A constructed failure shows how missing chunks create mixed answers

In this constructed example, document D7 revision V4 contains twelve chunks. V5 has eleven and changes maintenance approval conditions. Only ten new chunks are ready in the vector path, while the lexical path still serves V4. If the application declares the update complete upon a write acknowledgment, vector retrieval may return the new explanation and lexical retrieval the old approval clause. The answer can combine incompatible evidence.

Under this protocol, the V5 manifest fails validation because a chunk is missing and another retrieval path is not ready. V4 stays active if still valid; otherwise the system temporarily declines a definitive answer about that policy. After completing V5, reconcile its chunk set, visibility and policy, then conditionally switch activeRevision. Twelve, eleven and ten are illustrative counts, not performance measurements.

Now suppose V6 deletes the document and a delayed V5 upload arrives afterward. The deletion high-water mark must prevent V5 activation, even before cleanup finishes. If a query obtained V5 before deletion, a revocation check before generation should block further use. Record the remaining race between that check and output, and use request cancellation or tighter coordination according to the promised behavior. One check does not provide absolute instantaneous revocation.

Prove updates with four kinds of observation and failure tests

Freshness lag runs from source-change commit to complete activation of the new revision. Record receipt, parsing, index visibility and validation stages separately, and report percentiles and the longest failed backlog. With no source changes, an old last-document timestamp is not necessarily a pipeline fault; use heartbeats and source progress as well. Business timeliness determines targets, not a universal guarantee of a few seconds.

Revision completeness rate uses revisions due for activation as its denominator and revisions passing manifest, index and policy validation as its numerator. Mixed-revision violations count answers whose evidence contains disallowed revision combinations for the same document. Deletion propagation lag runs from authoritative deletion to consistent rejection of old evidence at serving entry points; record physical cleanup completion separately. A zero count means no violation was observed within test coverage, not proof that none can ever occur.

Acceptance cases should cover missing chunks, repeated writes, reordered updates, replay after deletion, cache hits on old answers, simultaneous access and content changes, and process failure around activation. Fix the source revision, expected eligible evidence and failure behavior for each case. Check retrieval and the revisions cited in the final answer. Retain manifests, activation receipts, a deletion-cleanup ledger and replayable test inputs so that 'knowledge base updated' becomes an auditable state.

References

  1. [1] Qdrant — Points, Awaiting Result
  2. [2] Elasticsearch — The refresh parameter
  3. [3] Azure AI Search — Change and delete detection using indexers for Azure Storage
Back to insights
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.