AI as an assistant, not an authority
AI tools can be genuinely useful and genuinely wrong, often in the same paragraph. Our approach is to treat them as assistants: helpful for first drafts, summaries and searching, always checked by a person on anything that matters.
Where we test AI
We test AI in places where a wrong answer is cheap to catch and a right one saves time:
- Summarising ServiceNow ticket and case history so a person picking up a ticket understands it quickly.
- Suggesting ServiceNow knowledge articles for a support analyst to consider.
- Drafting first versions of ServiceNow documentation, test cases and release notes for a person to edit.
- Classifying ServiceNow requests by likely category, with a person confirming.
- Searching and organising internal ServiceNow documentation.
AI inside ServiceNow workflows
Because most of our work is on ServiceNow, much of our AI testing is about how AI assistance fits inside an existing workflow. We look at where in a process a suggestion appears, who sees it, what they can do with it and how their decision is recorded. A suggestion that is easy to accept without reading is riskier than one that requires a deliberate choice, so we design the interface with that in mind.
How we evaluate
Real examples. We collect a set of real cases, with the answers a person would give, and compare the tool’s output against them.
Record the misses. We note where the tool was wrong, how often and how badly, because the pattern of errors matters more than the average.
Test unusual input. We try incomplete, ambiguous and unusual cases, because that is where tools fail.
Watch over time. We recheck periodically, because behaviour can change when data or tools change.
Data readiness
Most AI disappointments trace back to data. Before testing a tool, we check what data it needs, who owns it, who may see it and whether it is accurate enough. Sometimes the finding is that the data needs work before AI is worth trying, and we say so.
Human review, by design
For every AI-assisted step, we decide in advance what a person reviews, how they will know when to be suspicious and how their decision is recorded. We do not let an automated suggestion quietly become an automated decision.
What we do not do
We do not claim accuracy figures we have not measured, and we do not put confidential customer data into tools that have not been approved for it.
A small Salesforce note
We occasionally test AI assistance on CRM tasks, such as summarising an account’s history, using the same method.