Skip to main content
AI strategyProduct engineeringAgentic operationsResponsible deploymentGlobal deliveryAI strategyProduct engineeringAgentic operationsResponsible deploymentGlobal deliveryAI strategyProduct engineeringAgentic operationsResponsible deploymentGlobal delivery
AIoverflow.tech
All posts
Knowledge assistantsRAGAI operations

Keep Knowledge Assistants Current With Versioned Releases

AIoverflow5 min read
Keep Knowledge Assistants Current With Versioned Releases

Keep an internal knowledge assistant accurate by treating changed source material as a controlled release—not just another indexing job. Track authoritative versions, publish updates only after validation, retire superseded content, and invalidate answers that depend on it. When current evidence is unavailable, the assistant should say so rather than silently reuse an old rule. The practical design question is: when does a source change become safe to answer from?

Define what counts as current before building the pipeline

Retrieval-augmented generation (RAG) supplies a language model with relevant material retrieved at question time. Updating that material avoids relying solely on the model’s learned knowledge, but retrieval alone does not establish which policy applies.

A recently uploaded document might be a draft, a future policy or a copy of an obsolete handbook. Define authority separately from upload time. For each governed document, record:

  • A stable document identifier and a distinct version identifier.
  • The accountable owner and approval status.
  • Effective-from and, where applicable, effective-until dates.
  • The version it replaces, plus its withdrawal status.
  • Source location, access rules and last successful verification time.

Assign freshness targets by business consequence. A facilities guide might tolerate an overnight refresh; a withdrawn financial approval rule may need immediate exclusion. These are operating decisions for source owners, not assumptions the model should make.

Also distinguish current-policy questions from historical ones. An archived version can remain useful for an explicitly historical request without being eligible evidence for today’s instructions.

Publish complete knowledge versions, not partial updates

Microsoft’s RAG design guidance separates document processing, retrieval and answer generation, and recommends evaluating individual stages as well as the final response. That separation is useful for diagnosing stale answers: the failure might precede the model call entirely.

A recommended release sequence is:

  1. Detect and reconcile. Use source change notifications where available, plus periodic comparisons against the source inventory to catch missed edits and deletions.
  2. Stage the candidate. Split the document into chunks—smaller searchable passages—and attach version metadata to each. Generate embeddings, numerical representations used for similarity search, where required.
  3. Validate completeness. Check that required sections and tables survived extraction, all passages belong to the expected version, and approval metadata is present.
  4. Activate together. Make the complete version eligible for retrieval in one controlled switch. Do not expose a mixture of old limits and new exceptions.
  5. Retire dependencies. Exclude superseded passages and invalidate cached answers that used them.

Maintain a release manifest: a record linking source versions, indexed passages and activation status. For cached answers, store the source-version dependencies. If that is impractical, invalidate the affected collection’s cache when publishing.

Conversation history also needs attention. A prior answer is not fresh evidence. Follow-up questions about current policy should retrieve again rather than inherit an earlier answer’s facts.

Worked example: a travel-policy change

Illustrative example: an organisation lowers its domestic hotel limit from £180 to £150 per night, effective Monday. The new policy is approved Friday, but includes an exception for bookings approved before Monday.

On Friday, the pipeline stages version 8 while version 7 remains current. Validation checks that both the new limit and the booking exception appear in searchable passages. At Monday’s activation, retrieval switches to version 8 and cached answers supported by version 7 become ineligible for current-policy questions.

An employee asks: “Can I claim £170 for Tuesday’s hotel?” The assistant should not simply reject the amount. It should explain the £150 limit, cite version 8, and ask whether the booking was approved before Monday. That missing fact determines whether the exception applies.

Now suppose extraction drops the exception table. The completeness check should block activation. Once Monday arrives, continuing to serve version 7 as current would still be wrong. Mark the affected topic unavailable for definitive answers, explain that the current policy cannot be verified, and route the question to the travel-policy owner.

Rollback can restore working software or an index structure; it must not silently restore a withdrawn business rule.

Handle conflicting and suspicious updates explicitly

If two approved sources disagree, apply a documented authority rule or escalate to their owners. Do not ask the model to choose whichever passage sounds more convincing. A failed connector should likewise produce an operational alert, not an indefinitely reassuring “last synced” badge.

Updates can also introduce prompt injection: instructions embedded in source content that attempt to redirect the model. OWASP’s prevention guidance identifies retrieved documents as an attack channel and warns that formatting data separately from instructions is not an enforcement boundary.

Screen changed material, quarantine suspicious candidates for review, keep permissions enforced outside the model, and validate outputs before display. If a candidate is quarantined, preserve the withdrawal rules for its predecessor. Security review must not accidentally extend an expired policy’s validity.

Measure release correctness, not just answer quality

Build a change-focused test set with before-and-after questions, future-effective policies, deletions, conflicting sources, failed extraction and malicious document instructions. Use these concrete acceptance criteria:

  • Freshness delay: measure source approval or effective time to verified availability; set and monitor a target per source class.
  • Obsolete retrieval: require zero withdrawn passages in the current-policy test results.
  • Evidence support: reviewers verify that each material rule in an answer is supported by the cited, applicable version.
  • Exception completeness: require every critical policy-exception test to include the exception or request the missing fact.
  • Safe failure: all designated unavailable-evidence cases must abstain and name an escalation route; also track unnecessary refusals on answerable questions.

Repeat these checks after parser, retrieval or model changes. Connect failures to named owners and alerts, following the approach in Production LLM Observability: From Traces to Action.

If you need help defining knowledge-release boundaries and acceptance tests, contact AIoverflow to scope the workflow and its operating controls.

Sources & further reading

Prepared with AI assistance using the sources above and AIoverflow’s service context. Examples are illustrative; validate implementation decisions against your own requirements. Suggest a correction.

Got a workflow that might fit AI?

We start with an honest discovery call — and tell you straight whether it's worth building.