We built Currai because debugging an AI application often means working backward from an answer that looks wrong.
The request succeeded. The model returned a response. Nothing crashed. But the assistant misunderstood the user, ignored a tool result, or answered with confidence it had not earned.
To fix that, you need evidence: the prompt, the retrieved context, the tool calls, and the conversation that led to the answer. You also need a way to test whether your next change actually improves the system.
That is the problem Currai helps teams work through.
We open-sourced it because we believe the tools used to understand AI should be open to inspection too.
Trust needs something you can inspect
An evaluation score can influence whether a team ships a prompt change, switches models, or investigates a production failure. That makes the process behind the score important.
What was captured? What context reached the evaluator? What assumptions went into the result?
Documentation can explain the intended behavior. Access to the code gives developers another way to check it, investigate unexpected behavior, and ask more precise questions.
Open source does not automatically make software correct. It makes the implementation available for scrutiny. For a tool that helps teams judge the behavior of other systems, that matters.
Every AI application has its own definition of failure
A support assistant, a research agent, and a shopping assistant can produce similar-looking conversations while having very different responsibilities.
One needs to follow an escalation policy. Another needs to support its claims with sources. Another needs to use current product information and recognize when it cannot complete a request.
The useful questions depend on the application.
We have opinions about how traces, evaluations, and production evidence should connect. But those opinions should have room to evolve as developers bring different workflows and constraints.
Opening the code gives teams a path to investigate those differences, propose changes, and adapt the parts that need to work differently.
Real failures make better tools
We can test Currai against the systems we build and the examples we know. That still leaves plenty we have not encountered.
Other teams will bring different tool chains, conversation patterns, and failure cases. Some will expose bugs. Others will challenge an assumption that seemed reasonable in our own applications.
That feedback is valuable because AI quality depends heavily on context. A workflow that makes one failure easy to understand might leave another difficult to diagnose.
A useful contribution can be a fix or an integration. It can also be a reproducible example showing where the tool falls short. Those examples help us make better engineering decisions.
Publishing the code is a starting point
Open-sourcing Currai does not finish the work.
The code needs clear documentation. Issues need useful responses. Contributions need review. Developers need to understand how the project works and where it is going.
We want that work to happen alongside the people using Currai to investigate real AI behavior.
The goal remains the same: help teams move from a failed conversation to an explanation, from an explanation to a specific change, and from that change to evidence that the experience improved.
Opening Currai gives more people a way to examine and improve that process.
If you are building an AI application, start with a conversation where the assistant failed. Look at what it received, what it did, and what the user needed. Then ask whether your tools give you enough evidence to make the next version better.
That is the work we built Currai for. It is also the work we want to build on together.