Why evals alone won't fix AI assistants, and why I built Currai

An assistant can pass your tests and still leave users stuck. Why AI quality needs production evidence, and how that shaped Currai.

Imagine asking an AI assistant to move a meeting to Friday. It finds the meeting, calls the calendar tool, and replies, "Done. I've moved it to Friday."

Except the calendar tool rejected the update. Your meeting is still on Thursday.

The response is clear. The tone is friendly. An evaluator looking only at the final message might approve it. You now have a scheduling problem and a reason to stop trusting the assistant.

This is the gap I care about when I think about AI quality: the distance between an answer that looks right and a task that actually got done.

It's also why I don't believe evals, on their own, will solve the problems with AI assistants.

I want evals in the development process. They help catch regressions and compare changes before users experience them. A test that checks the calendar's final state could catch exactly the failure above. The trouble starts when passing the tests becomes enough evidence to declare the product reliable.

A test suite has boundaries. Users have very little interest in staying inside them.

Someone asks two questions at once. Someone changes their mind halfway through a booking. A returning user assumes the assistant remembers a detail from earlier. A tool returns an unfamiliar error, and the assistant continues as if everything worked.

You can write tests for these situations. You should. But you need a way to discover the situations you haven't thought to test yet.

That discovery work continues after release. Your knowledge base changes. An integration behaves differently. People start using the product for something you didn't design it to do. Even an unchanged prompt can be operating in a different environment from the one you evaluated.

There's another problem: deciding what counts as success.

Suppose a support assistant gives an accurate explanation of how to change an account setting. The user asked the assistant to change it for them. The explanation can be factually correct and still leave the request unfinished.

Or imagine an assistant that eventually completes a booking but asks for the same information four times. A completion check passes. The user has spent the conversation wondering whether anyone is listening.

These are different failures, and they need different evidence. The first requires understanding the user's intent and the assistant's ability to act. The second requires looking across the conversation. A score for the last response tells you very little about either experience.

Good evaluation practices already account for this. Anthropic's guide to agent evals describes evals as part of a broader process that includes production monitoring, transcript review, and user feedback. That is the approach I agree with. The mistake is expecting one part of that process to do all the work.

Even online evals depend on the evidence you give them. A judge reading a transcript cannot verify a calendar update if the tool result and calendar state are missing. And an evaluator's verdict needs checking too, especially when the definition of success is ambiguous.

I built Currai around the need to understand these production conversations. It brings user intent, agent responses, and tool evidence together so teams can investigate where the experience broke down. It also groups recurring needs and flags behavioral violations, with evidence attached to the findings.

The part that matters to me is being able to follow a claim back to what happened.

If an assistant says it completed an action, I want to inspect the result behind that statement. If a user keeps repeating a request, I want to understand what the assistant missed. If similar conversations keep failing, I want the team to see the pattern without relying on someone to report every incident.

That evidence changes the work you do next.

Go back to the calendar example. A failed update could mean the assistant lacked permission. It could mean the tool received an invalid date. It could mean the integration was unavailable. Adding "be more accurate" to the prompt won't repair a permission problem.

Once you understand the failure, you can fix the relevant part of the system and preserve the case as a regression test. Then you need to watch what happens after the fix reaches users. Does the same failure still occur? Does the assistant explain the problem clearly when an update is rejected? Can the user recover without starting over?

That's the workflow I want Currai to support: find the failure, inspect the evidence, make a specific correction, and check whether the experience improves.

It still takes judgment. An abandoned conversation doesn't prove that the assistant failed. A frustrated message doesn't tell you the cause. Automated findings are leads to investigate, and incomplete instrumentation leaves gaps. I don't think another score or dashboard makes those problems disappear.

But a team can make better decisions when it can see the conversation and the actions behind it.

I'll keep using evals. I also want to know what happens in the conversations those evals don't represent yet. That is the reason I built Currai, and the question I want it to help teams answer every day: did the assistant actually help this person finish what they came to do?