Define what a good answer means
A document assistant should help a specific audience answer a specific set of questions. Write down which documents it can use, which version is authoritative, and what to do when the collection has no answer. Without that boundary, evaluation becomes a matter of whether the output sounds plausible.
For an illustrative internal-policy assistant, a good answer could state the applicable rule, cite the supporting passage, and identify a relevant exception. The expected result should come from the document collection, not from the model's general knowledge.
Build a small, deliberate question set
Collect straightforward lookups, questions requiring multiple passages, ambiguous requests, and questions whose answers are absent. Include realistic wording rather than only repeating phrases from the documents. Record a reference answer or the correct abstention behavior for each question.
Keep some questions outside the development loop. Repeatedly tuning against the same examples can make a system appear stronger than it is on new requests. Record the source-document version alongside each expected answer.
Inspect retrieval separately
Before judging the final prose, check whether the needed evidence was retrieved. If the passage is missing, changing the answer prompt may not fix the underlying problem. Investigate document parsing, chunk boundaries, query interpretation, and ranking.
For each example, store the retrieved passages and their positions. This makes it possible to distinguish a missing source from a generation error and provides a concrete basis for comparing changes.
Check claims against citations
A citation is useful only when its passage supports the nearby claim. Review the answer sentence by sentence. Look for omitted conditions, overgeneralization, conflicting sources, and references that mention the topic without supporting the conclusion.
Include an explicit check for unsupported questions. A clear statement that the available documents do not contain an answer may be the correct result. Track unnecessary refusals separately from unsupported answers.
Measure the whole interaction
Record response time, approximate model cost, and the amount of human correction needed. Use consistent settings when comparing versions. Note the model, retrieval configuration, document snapshot, and evaluation date so another person can repeat the comparison.
Choose thresholds for the intended workflow. A casual knowledge search and a high-consequence operational decision need different review procedures. Avoid presenting a single small test set as proof that the system is reliable in every context.
Turn errors into the next iteration
Group failures by cause and fix the largest relevant group first. Parsing errors call for better ingestion; missing evidence calls for retrieval work; unsupported synthesis calls for answer-generation or review changes. Rerun the held-out examples after each meaningful change.
This is a proposed evaluation method, not a published benchmark. Explore RAG and evaluation services or describe your document workflow.