You change a prompt. The assistant's next answer sounds better. You try another question, get a reasonable reply, and feel ready to ship.
Then comes the awkward part: explaining what "better" means.
More polite? More accurate? More likely to finish the user's request? Those are different promises. A few good conversations can't tell you which one you've kept.
An eval is a repeatable check of whether your AI product did what you expected. The useful part is deciding what you're willing to call a failure.
I'd begin with a task your assistant already handles. Pick a mistake that would leave a customer stuck, then work out what evidence would let you catch it. You can build a useful first check around that one mistake.
Imagine a delivery assistant. A customer writes:
Please change my delivery address to 14 Cedar Street. I'll be there tomorrow.
The assistant calls the address-update tool. The tool returns an error because the parcel has already left the warehouse. The assistant replies:
All set! Your package will arrive at 14 Cedar Street tomorrow.
That answer reads well. It also gives the customer a reason to wait at the wrong address.
For this situation, I'd start with one question: did the assistant claim the address changed without confirmation from the delivery system?
Now we have something concrete to check. If the tool rejected the update and the assistant said it succeeded, that's a failure. If the assistant explained that the update failed, it passes this particular check. Whether it offered a useful next step is a separate question.
Notice how much the tool result matters. Someone reading only the final reply could miss the problem entirely. The reviewer needs the customer's request, the attempted action, and the result behind the answer.
I'd also write down a few awkward cases before automating this check:
- The tool accepted the request but hasn't confirmed the change yet.
- The customer asked about changing the address but never authorized an update.
- The tool timed out, so the assistant doesn't know whether anything changed.
- The update succeeded, but for a different order.
These cases force a product decision. What should the assistant say when it doesn't know? What evidence is enough to promise that a task is done?
For the timeout case, a reasonable response might be: "I couldn't confirm the address change. Let me check the order before trying again." That gives the customer an honest account of the situation and avoids blindly repeating an action.
Once you have an automated reviewer, give it examples you've already checked yourself. Look closely at the failures it misses and the harmless answers it flags. A reviewer that rejects every answer would catch every bad promise, but it wouldn't help you decide what to ship.
Suppose a revised prompt now explains the rejected address change correctly. Good. Next, try a successful update. An overly cautious assistant might keep telling customers it can't confirm changes even when the delivery system says they worked. The fix has created a different problem.
This is why I like starting with a narrow question. When a check fails, you can explain the failure in ordinary language and investigate a specific behavior. "The assistant promised a delivery change after the tool rejected it" gives an engineer much more to work with than "quality dropped."
Real conversations help you find the next question worth asking. Perhaps customers keep supplying their order number twice. Perhaps the assistant handles one parcel correctly but gets confused when there are two. Each pattern gives you another behavior to investigate and, once understood, another check to add.
That's where Currai fits. I built it to help teams inspect real agent conversations, understand user intent, and investigate failures with the tool evidence attached. It groups recurring needs and checks agent behavior against rules you define, so you have concrete examples to bring back into your evaluation process.
If you're building an AI assistant, try Currai and start with a conversation where something went wrong. Work out what the assistant should have done. That's a useful place to begin your next eval.