The one-line version
Kodara builds custom AI assistants that talk, sell, and support on behalf of an expert. Before a client would let that AI loose on real customers, they froze, not because the AI wasn't good, but because they had no way to know it was good. I designed a scored readiness report that replaces that guesswork with evidence and turns the moment right before launch into a clear, confidence-backed decision.
Kodara claims to offer clients a “brain in a box.” Essentially, Kodara trains an AI to act, sound, and sell like the client’s hired expert. This is considered a high-trusted product. The client does not buy a certain feature. They give that expert’s reputation to something they have not even written.
A consistent and evidence-based pattern emerged during onboarding calls. Clients would reach the finish line of a fully functioning and tested AI Brain and, for various reasons, would not launch it. The reason would not be due to a defect in the AI Brain. Clients would open the chat and, during onboarding calls, would freeze from reading a response the client would have likely phrased differently.
Through 15 onboarding calls, 7 of 15 clients explicitly cited fidelity anxiety (“will it actually sound like me?”) and, during onboarding calls, would stall during the handoff. This stall is costly in the manner that it compounds. Everything downstream of launch, from the enrollment of end-users to the retention and upselling of users, will not begin until the client clicks the launch button. Clients have a finished product that sits unused for three weeks. This friction is not minor. It is a total funnel block.
The team's instinct going in was to build more dashboards and more visibility into "AI health." My first job was to separate two different problems being treated as one. Helping a client trust their AI enough to launch it is not the same as helping them understand how their AI is performing after launch. We named the first Launch Confidence and scoped it as its own initiative, the AI Readiness Report, rather than folding it into a general purpose analytics dashboard, which we put on hold once it was clear it was solving a different job.
The most confident moment in a client’s lifecycle is when a Brain has been built and tested. Testers use the Brain, everything we measure says it is ready, and the task is complete. However, that moment never had any artifact. We handed our client a chat link and said "looks good, go ahead."
A scored readiness report turns that silent handoff into a decision point. A readiness report is that much more valuable to a client who is ready. Instead of hearing that the system is ready, a client gets an evidenced green light to proceed. A readiness report provides a client who is not ready a short, specific statement of what needs to be addressed instead of telling that client that they need to review the entire AI to determine if it is ready.
There was a second, quieter benefit I made sure stayed visible in the PRD. A standardized readiness framework formalizes Kodara’s delivery team’s quality control process. It moves away from subjective assessments regarding whether or not a Brain is complete. This becomes more important as the client base expands beyond what the founder’s intuition can support.
The toughest problem to resolve early on was epistemic, not visual. How can we make a score trustworthy, rather than just decorative?
I refused to accept the framing that the report just needed to “give the sense of deep understanding.” The idea that we could survey thin measurement accompanied by a score that looked confident was not something I was willing to stake on. The sense of confidence is the same failure mode we were addressing, just at a different level, and it’s particularly dangerous to have a product that is trusted as the foundation of its value proposition; if a client publishes a report that is based on a false sense of confidence, and as a result, the first negative experience with the client’s AI that is not prepared, then Kodara would lose trust before we would from a report that would take longer and be honest.
The score needed to be built from true, defensible signal. I came up with three elements: Brain coverage, Tester Engagement, and Kodara Manual Review.
Brain coverage reflects how much the AI has been trained in the categories that are relevant to a launch decision (voice, method, boundaries, top questions, objections, pricing).
Tester Engagement reflects how much the testers actually interacted with the Brain, measured against something concrete. I had to remove a card, “your AI vs the others,” that was perceived as a comparison we could not defend, and designed it to allow a comparison.
Kodara Manual Review means that at least one person has reviewed the output to ensure the score was not self-graded.
Every one of those three had to trace back to something real and inspectable, not a number chosen because it looked reassuring on a screen.
The interaction model went through four real iterations, and each one taught me something about what this artifact actually needed to be.
Following our existing design system, I started to pull some of our components and assemble them in a dashboard style. Before that, I used AI (Claude/ChatGPT) to help me create some rough wireframes for initial guidance. I used Granola to record the meetings and the transcript + PRD to build those.


After designing these options, we decided to speak with the team and stakeholders to see which one looks more intuitive and meets the goals of the project.
After designing the foundation, I decided to go straight to vibe coding in Cursor to save time. As I had the design system built in the codebase, felt this was the fastest way to put this in the front of users.
I had Claude Design make the animated orbs. We used these, in this particular brand, to represent the knowledge of the coaches. Everything that connects to the brain builds the brain of the user. This helps us to design the user’s persona, and, perhaps the best part of this is that they can now use it to engage their users and clients in a way they haven’t before… through a chatbot!
When we tested the layout with the internal team, we found a discoverability problem. Reaching the deeper data required scrolling, but several team members did not realize that on first glance, so the information sat there unseen. We addressed this by exposing it through a button group instead, so the deeper data is visible immediately rather than depending on the user to scroll.
When we tested the layout with the internal team, we found a discoverability problem. Reaching the deeper data required scrolling, but several team members did not realize that on first glance, so the information sat there unseen. We addressed this by exposing it through a button group instead, so the deeper data is visible immediately rather than depending on the user to scroll.
If the user faces a score is below 90%, we will not allow the user to publish their agent, they will need to improve by answering some questions related to what content is missing. We call this step Interview Mode.
This allows the user to improve their AI without needing our help.
I made a deliberate point of finalizing the instrumentation plan before the visual design was fully locked, specifically so the definition of success could not quietly drift into something easier to hit, like counting button clicks, once the pressure to show progress arrived. The primary success metric for this feature is not the click rate on the call to action button. It is the time between a client opening the report and the client actually sharing their Brain link. A client can click a call to action and still stall afterward, so a click alone tells you very little about whether the underlying anxiety was actually resolved. Everything else we planned to track, whether the report was opened at all, how long a client spent before clicking anything, and which branch of the flow a given client ended up on, is diagnostic information that helps explain the outcome, not a substitute for the outcome itself.
After finishing the project, I published the branch and sent a PR to the dev team just to make sure everything was following the codebase. The team is currently testing this feature with new coaches. So far, so good we have positive feedback, but it is too early to have actual metrics






