Article summary
- A useful internal AI assistant needs more than fluent answers: it must find the right source, respect permissions, show uncertainty and fail safely.
- Create questions from real work rather than polished demonstrations. Include straightforward requests, vague wording, outdated terminology, multilingual queries, conflicting documents and questions whose answer is not available. Record the expected source, expected answer and required escalation for each case.
- Test with users who have different roles. A person must not receive information through the assistant that they cannot access in the source system. Include deliberately restricted documents, departed-user scenarios and role changes. Retrieval rights should be checked before content is sent to the model.
- Launch only when the agreed critical tests pass and remaining limitations are documented. Keep a regression set for future document, prompt and integration changes. Monitor failures, permission incidents, source freshness and user feedback after release. A small reliable scope is more valuable than a broad assistant nobody can trust.
Key takeaways
- Start by defining what the assistant must do, what it must never do and which errors are unacceptable. For a policy assistant, a correct answer may require the current approved document, a clear citation and a refusal when evidence is missing. Separate useful behaviour from pleasant wording.
- Evaluate retrieval before judging the final response. First ask whether the system found the correct, current and permitted source. Then assess whether the answer accurately reflects that evidence. This distinction helps teams fix indexing, document quality, retrieval rules or prompting instead of treating every failure as a model problem.
- The assistant should cite important answers, distinguish facts from suggestions and state when evidence is incomplete. Test questions with no answer, contradictory sources and sensitive implications. Define when the assistant must stop, ask for clarification or hand the task to a named person.
- Launch only when the agreed critical tests pass and remaining limitations are documented. Keep a regression set for future document, prompt and integration changes. Monitor failures, permission incidents, source freshness and user feedback after release. A small reliable scope is more valuable than a broad assistant nobody can trust.
Define acceptance criteria before testing
Start by defining what the assistant must do, what it must never do and which errors are unacceptable. For a policy assistant, a correct answer may require the current approved document, a clear citation and a refusal when evidence is missing. Separate useful behaviour from pleasant wording.
Build a representative question set
Create questions from real work rather than polished demonstrations. Include straightforward requests, vague wording, outdated terminology, multilingual queries, conflicting documents and questions whose answer is not available. Record the expected source, expected answer and required escalation for each case.
Measure source and answer quality separately
Evaluate retrieval before judging the final response. First ask whether the system found the correct, current and permitted source. Then assess whether the answer accurately reflects that evidence. This distinction helps teams fix indexing, document quality, retrieval rules or prompting instead of treating every failure as a model problem.
Test permissions and restricted information
Test with users who have different roles. A person must not receive information through the assistant that they cannot access in the source system. Include deliberately restricted documents, departed-user scenarios and role changes. Retrieval rights should be checked before content is sent to the model.
Check uncertainty, citations and escalation
The assistant should cite important answers, distinguish facts from suggestions and state when evidence is incomplete. Test questions with no answer, contradictory sources and sensitive implications. Define when the assistant must stop, ask for clarification or hand the task to a named person.
Run workflow tests with real users
Let intended users complete real tasks in a controlled pilot. Observe whether the assistant reduces searching, creates extra verification work or changes how decisions are made. Capture failed queries, misleading answers and confusing interface behaviour. Do not expand the scope until recurring failures have owners and fixes.
Create a launch and monitoring decision
Launch only when the agreed critical tests pass and remaining limitations are documented. Keep a regression set for future document, prompt and integration changes. Monitor failures, permission incidents, source freshness and user feedback after release. A small reliable scope is more valuable than a broad assistant nobody can trust.
Internal AI assistants · AI integrations · Business automation · Services and pricing · Contact
Want a website that is fast, measurable and yours?
Book a practical website consultation and we will look at ownership, SEO, analytics and conversion tracking together.