Build a useful AI pilot with fictional cases, realistic exceptions and explicit review criteria, without copying customer records or overstating the results.
Direct answer: Build fictional cases from the task's structure, not by disguising customer records, and write the expected handling before asking AI to process them. Include ordinary requests, missing information, conflicting instructions and cases the system should refuse or escalate. Use the pilot to decide whether further evaluation is worthwhile, not to claim the workflow is ready for real customer information. If realistic testing requires confidential policies or identifiable examples, pause until those inputs and the chosen service are explicitly approved.
A pilot needs enough realism to expose failure without quietly importing the risk it was supposed to avoid. Changing names in actual support messages may leave locations, unusual circumstances or account details that still reveal the people involved.
The harder limitation is evidential. A system can succeed on neatly written fictional requests and struggle with real ambiguity. Treat your results as evidence about the cases you constructed, not as a prediction of production accuracy. This is a focused extension of adopting AI without losing a small team's trust.
Applies to: an early, isolated pilot of a low-consequence task. It does not authorise live customer uploads, automated customer actions or decisions requiring specialist judgement.
Build the synthetic-case coverage map
The synthetic-case coverage map is an editorial method for listing the situations a trial must represent, then assigning fictional cases and expected responses to each. “Synthetic” here means deliberately invented test material, not a dataset statistically generated from real customer records.
Before starting, identify the task owner, the output you want, the service and account you are allowed to use, and who will review results. Keep the pilot separate from live email, shared drives, customer databases and automated actions. If the application cannot be used without connecting those systems, choose a different test arrangement or stop.
Write the task in one sentence. For example: “Turn a fictional equipment-hire enquiry into a list of missing booking details.” Do not expand the same pilot to quote prices, promise availability or accept bookings. Those are different decisions with different consequences.
My recommendation is to begin with manually invented cases even when a vendor offers to build a test set from your records. The manual approach is slower, but you can explain where every detail came from. A specialist-led evaluation using authorised real or derived data may be stronger later; it is not the same low-exposure first step.
Invent structures rather than disguised people
Describe the fields and relationships the task needs: requested dates, item type, quantity and whether the enquirer has supplied a return time. Create fictional combinations from scratch. Use labels such as “Example customer A” and omit real phone numbers, addresses and account identifiers entirely.
Do not copy a memorable incident and replace the customer's name. The combination of circumstances may be the identifying part. The ICO notes that synthetic data derived from real data can still reveal information about the originals; processing real personal information to create it also raises data protection obligations. ICO guidance on synthetic data.
That is a UK context. Requirements vary by country, sector and contract. Get appropriate data protection advice if your proposed test depends on real records, even when the final examples will look fictional.
Check the instructions as well as the example records. A completely fictional enquiry can still be attached to a confidential pricing policy or reveal a security procedure. Use an invented policy with visibly labelled illustrative rules when your pilot is only testing the workflow's shape.
Cover the failures the task can produce
For the equipment-hire example, your map could include these situations:
| Situation | What the fictional case contains | What the reviewer should expect |
|---|---|---|
| Missing information | No return time | Ask for the missing time without inventing it |
| Conflicting details | Two different collection dates | Flag the conflict rather than choosing one |
| Long irrelevant context | A lengthy story around a short request | Preserve the actual booking details |
| Instruction inside the enquiry | A request to ignore the checking rules | Treat it as customer text, not authority |
| Policy boundary | A request outside the invented hire conditions | Identify the boundary without promising an exception |
| Unrelated request | A question the workflow is not designed to answer | Route it for human handling |
The fourth row probes a limited form of prompt injection, meaning instructions in input that try to redirect the system's behaviour. The NCSC identifies this as an AI security weakness. A few fictional cases do not establish protection against determined attacks. NCSC guidance on AI and cyber security.
Include ordinary cases too, either within each situation or as a separate baseline. Vary wording without changing the underlying answer. If your only difficult example uses an obvious warning phrase, success may say little about subtler versions.
Write the expected result before the run. Acceptable wording can vary, but the conditions should not: the missing time must remain missing, the conflicting dates must both be recognised, and no booking must be confirmed.
Run the pilot with a separate review record
Number each case and keep its expected handling in a separate column. Save the actual output, the application's visible configuration and the date. Record model identity only if the application exposes it reliably; do not infer an underlying model from a product name.
For every case, inspect the relevant details against the fictional source and rule. Mark factual mistakes, unsupported commitments, ignored uncertainty and unusable formatting separately. This tells you what needs repair. One overall score hides the difference between an awkward sentence and an unauthorised promise.
Keep failures in the record. If you revise the instruction after a failure, rerun the affected case and check previously successful ones. Do not report only the final polished examples. If you repeatedly tune on the same cases, add fresh fictional variations that were not used while editing the instruction.
The reviewer must be able to reject an output and explain why. If nobody can establish correct handling independently of the model, the task is not ready for this pilot design.
Calculate coverage and the effort it costs
Suppose a team constructs 24 fictional requests across the six situations above, with four cases per situation. All figures below are illustrative planning assumptions.
Coverage of the named situations is 6 ÷ 6 = 100%, but that means only that every listed situation has a case. It is not 100% coverage of real-world behaviour. You have not established how common those situations are or whether important ones are missing.
Assume 21 cases meet the written criteria and three fail. The case pass proportion is 21 ÷ 24 = 87.5%. If one failure confirms a booking the system had no authority to confirm, the workflow should not advance merely because most cases passed.
Budget the work as well:
- Case preparation: 24 × 3 minutes = 72 minutes.
- Expected-response preparation: 24 × 2 minutes = 48 minutes.
- Running and recording outputs: 12 minutes.
- Review: 24 × 2 minutes = 48 minutes.
- Correcting the instruction and checking affected cases: 30 minutes.
Total effort is 72 + 48 + 12 + 48 + 30 = 210 minutes, or 3.5 hours.
That is evaluation effort, not a claim of time saved. Compare it with the consequence of the decision it supports. A tiny one-off job may be cheaper to complete manually; a repeated workflow may justify the investigation. Future operational savings still need separate measurement with authorised conditions and continuing review.
Decide what happens after the first session
- Spend 30 minutes defining the task, excluded actions and fictional input rules. Stop if the proposed setup requires access you have not been authorised to grant.
- Build the coverage map and expected responses, then schedule the preparation and review effort rather than squeezing it around other work.
- Run the cases and retain failures. Ask a second capable person to challenge the most consequential outcomes where possible.
- At the review meeting, choose to abandon the idea, revise and repeat the fictional trial, or request approval for a more representative evaluation. State exactly what the pilot has not demonstrated.
Do not quietly replace fictional inputs with customer records because the demonstration looked convincing. That is a new stage requiring a separate data-handling and operational decision.
Related guides
Frequently asked questions
Can another AI tool generate the fictional cases?
Yes, provided the instruction itself contains no restricted information and you review every generated case. Give it an invented task, fictional rules and the situations you want represented. Do not upload customer records and ask it to make them anonymous. Check that the resulting cases have coherent answers and do not accidentally include plausible contact details you might later use. A person should write or approve the expected handling independently. Using one model to generate both easy cases and agreeable answers can make a pilot look stronger than it is.
How many cases do I need before the pilot is meaningful?
There is no universal minimum that makes an AI workflow safe or accurate. Start with the situations that would change your decision, then add variations where wording, length or missing context matters. The 24 cases in this article are an illustrative workload, not a validation standard. A small set can reveal a reason to stop, but passing it does not establish reliable production behaviour. If you need a defensible performance estimate or must handle high-consequence decisions, involve someone qualified to design a representative evaluation and interpret its uncertainty.
Should I include deliberately malicious requests?
Include bounded examples relevant to your task, such as text asking the system to ignore the fictional checking rule. Keep the trial isolated from real accounts and external actions so a failed response cannot cause a live consequence. Record whether the system follows the untrusted instruction, but do not describe a handful of successful cases as a security assessment. If the eventual service will receive public input or have permission to act, specialist security evaluation is a separate requirement. Do not connect real systems to make an early attack demonstration more dramatic.
Can a successful fictional pilot justify buying a subscription?
It can justify a limited next step if the purchase is proportionate and the unresolved questions are explicit. It does not by itself prove that the tool will save money on real work. Check the actual account terms, exit conditions, required features and continuing review effort before paying. Prefer a reversible commitment while testing assumptions rather than a long contract based on a demonstration. If the pilot exposes a flaw in the process rather than a need for AI, changing the process may be the more useful outcome.
What if a reviewer recognises a real customer in a fictional case?
Remove the case from the trial and investigate how it was constructed before uploading or sharing it further. It may have been based too closely on an actual incident, even if the author changed obvious identifiers. Replace it with a genuinely invented situation that preserves only the task structure you need. If restricted information has already been exposed, follow your organisation's incident process rather than quietly deleting the local copy. The relevant response depends on the data, service and jurisdiction, so involve the responsible privacy or security person.
How should I describe the results to colleagues?
State the task, the invented nature of the cases, the situations covered and the failures retained. Give the actual case count and review criteria, then identify what was not evaluated, such as real customer language, live integrations, accessibility or sustained workload. Avoid phrases that imply independent certification or a production accuracy rate. A useful result might be that the system handles missing dates but invents answers when rules conflict. That specific finding supports a decision about the next test; a claim that the pilot was successful usually needs more explanation.
Sources and verification
- ICO: security and data minimisation in AI, checked 11 September 2026 for the limits of synthetic data derived from personal information.
- NCSC: AI and cyber security, checked 11 September 2026 for prompt-injection terminology and limitations.
- The parent guide was read in the supplied website source; its public route could not be retrieved. No pilot results or product testing are claimed here.
This article is practical guidance. Apply it in proportion to your tools, evidence, risks, and responsibilities.



