Build an AI Evaluation Set from Real Business Work

Build a representative AI evaluation set by sampling actual work across its main categories, preserving the information available at the decision point, and having domain experts define acceptable outcomes. Keep ordinary workload measurement separate from rare-risk testing. The result should tell you both how often the workflow succeeds and where it must stop—not merely whether a model can handle a few convincing examples.
Define the unit of work before collecting records
An evaluation set is a collection of test cases with explicit expectations used to measure a system’s behavior. Start with one bounded decision: extracting an invoice, proposing a case response, or answering an internal policy question.
Write down what the system receives, what it may access, what it must produce, and what it must never do. For a multi-step workflow, include expected routing and actions, not just the final text.
Each case should contain:
- A stable identifier and the original record’s provenance: where it came from and when.
- Input documents, conversation history and relevant system state available at that moment.
- Applicable policy version, user permissions and permitted actions.
- Expected fields or acceptable response criteria, supporting evidence and escalation requirements.
- Category labels, reviewer status and sensitivity restrictions.
Avoid giving the system later information that the original operator could not see. A resolved case may contain an answer discovered only after a follow-up call; including that answer in the input makes the test unrealistically easy.
Sample for frequency and consequence separately
Choose a collection window that covers meaningful business variation, such as month-end processing or seasonal requests. Inventory the workload by dimensions likely to affect performance: task type, language, document format, source channel, customer category and exception status.
Use stratified sampling: divide work into relevant groups, then randomly select records within each group. Do not collect only clean examples volunteered by enthusiastic users. Include abandoned, corrected and escalated work where accessible.
Maintain two collections:
- Workload sample: reflects the observed mix of eligible work. Use this to estimate routine performance.
- Challenge sample: deliberately concentrates rare exceptions, security threats and high-consequence decisions. Report it separately rather than pretending it reflects everyday frequency.
For retrieval-augmented generation, or RAG—answering with information retrieved from external sources—Microsoft’s design and evaluation guidance calls for representative source material and queries, including questions the material cannot answer. In practice, sampling questions without their underlying documents and access conditions leaves important failure modes untested.
Use authorized records, minimize personal data and restrict access. When masking identifiers, preserve relationships and formatting needed for the task. Record exclusions so stakeholders can see what the evaluation does not cover.
Worked example: an invoice intake evaluation set
Illustrative example: an operations team wants AI to extract invoice fields and route discrepancies for review, without authorizing payments.
Suppose its workload review finds 70% standard digital invoices, 20% scanned invoices and 10% credit notes. A first workload sample of 200 records would contain 140, 40 and 20 records respectively, selected randomly within those categories. These counts illustrate a starting design, not a universal minimum or proof of readiness.
The team adds 40 separately reported challenge cases covering duplicate submissions, missing purchase orders, conflicting totals and suspicious embedded instructions. Constructed variants are labelled synthetic and linked to their originating cases.
One case contains an invoice requesting payment of 1,200 while the approved purchase order permits 1,000. The expected outcome is accurate extraction of both amounts, identification of the discrepancy and routing to an exception queue. Silently replacing the invoice amount with 1,000 fails, even though it matches the purchase order.
If the purchase-order service is unavailable, the expected outcome changes: preserve the extracted invoice, mark validation as incomplete and defer the decision. Do not treat a failed lookup as proof that no purchase order exists. For implementation detail, see controlled document workflows.
Establish trustworthy answers and protect the test
Historical human decisions are evidence, not unquestionable truth. Ask domain reviewers to label expected outcomes using a written scoring rubric—a shared set of rules for deciding whether an output passes. Have a second reviewer assess ambiguous and high-risk cases, and resolve disagreements before scoring models.
When policy itself is unclear, mark the case unresolved and assign a policy owner. Do not force a convenient answer merely to complete the dataset.
Separate development cases from a held-out set: cases reserved for final assessment rather than prompt tuning. Keep duplicates, related conversation threads and synthetic variants in the same partition to prevent leakage, where prior exposure makes results look better than genuine generalization.
Include prompt injection tests: attempts to redirect the AI through malicious instructions embedded in inputs. OWASP’s guidance identifies documents, emails and tool outputs as potential carriers. Test that these inputs cannot trigger unauthorized actions or disclosure; a refusal sentence alone is insufficient evidence if a tool still executes.
Set release criteria and maintain coverage
Agree on criteria before comparing systems. For the illustrative invoice workflow, candidate gates could include:
- At least 98% exact-match accuracy on required fields, reported separately by document category.
- Every designated must-escalate challenge case reaches review without a payment action.
- No unauthorized writes or sensitive-data disclosures in the security suite.
- Routine unnecessary escalations remain below an agreed review-capacity limit.
- Processing time and cost per completed case remain within operating budgets.
These are illustrative gates, not industry standards. Report counts alongside percentages: 20 credit notes provide limited evidence. Zero observed security failures does not establish zero risk.
Version the dataset, rubric, model configuration and results. Add newly observed failures to development tests, replenish the held-out set with fresh records, and reassess the workload mix after policy or channel changes.
If you need help turning operational records into a testable first workflow, contact AIoverflow to discuss scope, evaluation criteria and implementation boundaries.
Sources & further reading
Prepared with AI assistance using the sources above and AIoverflow’s service context. Examples are illustrative; validate implementation decisions against your own requirements. Suggest a correction.