The talk about evals is heating up. Rightly so. AI products have an interesting problem: they can work “perfectly” and still produce crappy results.
Evals are essentially ways of systematically testing whether AI output is any good. Pretty important when LLM output can vary from run to run.
I’ve been dealing with this problem inside Candor. The answer? Evals! (Sort of, and that’s the point.)
But not automated ones. There’s no shortage of eval frameworks, benchmarks, LLM-as-a-judge setups, automated regression tests, scorecards and increasingly sophisticated tooling. I’ll use more of that in the future. Today: I’m doing it mostly manual. I’m convinced this is critical, because before you automate things you need to answer a key question: What does good actually look like?
AI products can be wrong without being broken
Traditional software is relatively straightforward to test. You click a button and something should happen. You submit a form and the data should save. You call an API and you should get the expected response. Software testing can get incredibly sophisticated, but fundamentally you’re checking whether it did what it was supposed to do.
AI adds another layer. The workflow can succeed and the output can still be wrong.
Let’s dig into this using Candor as an example.
Candor is a synthetic user research platform. Someone defines an audience they want to learn from, and Candor creates synthetic users, interviews them and produces a research report. It’s heavily reliant on LLMs.
When Candor works, it creates 15-24 synthetic users, interviews them and produces a polished research report. No errors, no crashes, everything is smooth sailing. But were the users realistic, and did they actually match the audience the researcher wanted? Did the interviewer listen to what someone said and ask smart follow-up questions, or just plow through the script? Did the report overstate something or include a quote that nobody said?
You could call these bugs, but they’re different from traditional software bugs. The workflow succeeded. The output just wasn’t good enough. AI is good at hiding quality problems. A synthetic user can sound completely convincing and still be wrong. A report can read beautifully and contain something unsupported. An AI interviewer can ask perfectly reasonable questions while conducting a mediocre interview.
If you don’t look closely, everything seems fine. Until a customer notices. Or, more likely, they don’t tell you and just stop using the product.
The first version of an eval is you
When people talk about evals, the conversation moves quickly to automation. What model should judge the output? What’s the scoring rubric, the dataset, the pass rate? How do we run evals automatically on every change?
All useful questions. But I wasn’t ready to answer them. I didn’t have a rubric for product output. I didn’t know which things mattered most, and I didn’t know what kinds of failures I was going to find. So I’m starting manually.
I'm not the first to say this. Hamel Husain and Shreya Shankar call error analysis (reading the output yourself) "the most important activity in evals," and Teresa Torres makes the same case for product teams.
Here's what it looks like in practice for Candor:
Run the study, or inspect a real customer study.
Look at the output.
Find the problems. Claude Code helps me find issues, but I read the important parts myself.
Look for patterns.
Figure out what “better” would mean.
Make a change.
Test again.
Not very sexy, but it’s already finding real issues. This is the grind of building quality software. We found synthetic user profiles that included facts that didn’t make sense. We found the interviewer occasionally covering ground that had already been discussed, and not handling lukewarm or ambiguous answers as well as I’d like. We found quotes in reports that needed stronger verification against the underlying interviews.
So we fixed those things. Quotes are a good example. Candor now verifies every quote against the underlying transcript before the report is saved. If nobody actually said it, it doesn’t get presented as a quote. That’s a pretty obvious rule once you’ve found the problem. But first you have to find the problem.
That’s really what I’m doing by systematically evaluating everything. I’m discovering the eval. Eventually the patterns become criteria, the criteria become a rubric, and parts of the rubric become automated. Human judgment, then rubric, then automated eval. I don’t think I’d trust that last step nearly as much if I skipped the first one.
Not everything needs an eval
Sometimes the answer isn’t another LLM-based eval. It’s code.
Candor starts by researching an audience and creating synthetic participants. Let’s say a researcher asks for people who work at companies with more than 500 employees. That’s a requirement. I don’t want an LLM to get it right most of the time. Same with age ranges, seniority or other explicit criteria. So we’ve moved those things out of the AI’s judgment. The model can create the person. Code makes sure the person meets the requirements. I have a feeling that people are becoming overly dependent on LLMs to make decisions that are better served through other instruments.
I’ve started using this distinction: If something is a rule, don’t ask the AI nicely to follow it. Enforce it.
This applies way beyond Candor. If something:
has to be a particular format
must fall within a range
needs to match an enum
must contain exactly five items
has an explicit eligibility requirement
can be verified deterministically
…you may not need an eval. You need a check.
Save the AI for the stuff that actually requires judgment. Is this participant realistic? Is this answer vague? Did the interviewer ask a good follow-up? Does this evidence actually support the conclusion? Those are harder questions, and those need evals.
This distinction sounds obvious, but AI makes it incredibly tempting to throw another prompt at every problem. Sometimes the correct amount of AI is less AI.
Sometimes a failed eval is telling you to fix the UX
Candor’s output is not solely the responsibility of LLMs. A user’s input matters too. Candor has a very clear “garbage-in, garbage-out” problem. As people use Candor, I can evaluate the quality of their inputs and the impact it has on the outputs.
It’s tempting to say, “Well, that’s the user’s fault.” It isn’t. Or at least that isn’t a very useful way to think about it. If users regularly put the wrong thing into a field, the field is probably wrong. Or the example isn’t clear enough. Or we need to validate the input before running the research.
I wrote recently about people using AI to generate giant blobs of text and pasting them into Candor. People don’t read instructions. Maybe they never did. Now, I think it’s worse: they have an AI that can instantly create 700 words for a box that probably needs 50.
This is leading to a bunch of product changes. We’re adding better examples directly below fields, with clearer context. We’re warning people when an input looks suspicious, like a concept description that reads more like a research brief. Unfinished setups get saved so people don’t feel like they have to rush through everything.
That’s not an LLM fix. It’s a UX fix. And that’s changed how I think about evals. I’m not evaluating a model. I’m evaluating a product that has AI in it. The quality of the final output depends on the user, the interface, the inputs, the web research, the models, the prompts, the code around them and all the small decisions happening in between. If the final result sucks, asking “Which prompt do I need to fix?” is way too narrow.
Define “better” before you make the change
I’ve written before about defining “done” before I let Claude Code build something. Claude Code is very eager to declare victory. Ask it to build something, then ask, “Did you do it?” and you’ll get some version of, “Absolutely! Everything is working perfectly!” 😄 Instead, when coding, it’s best to define pass/fail criteria upfront. That way you can define what a feature should do, how it should work, and your expectations on how to validate whether it’s done properly or not.
I’m now using the exact same habit for AI behavior. Before we make a meaningful change, I want to know the current baseline, what we’re trying to improve, how we’ll measure it, how much better the new approach has to be, and whether there’s a cost or latency tradeoff. Then we test it.
This matters because AI makes experimentation incredibly cheap.
“Let’s do more web searches!” Sure.
“Let’s use that new, cool model!” Why not?
“Let’s have another AI check the first AI!” Sounds very AI. 🤣
Before you know it, your product is slower, more expensive and significantly more complicated. But is it better?
I built a blind taste test for search results
Candor builds its synthetic audiences, in part, using web research. If the research is weak, the participants are weak, and everything downstream inherits the problem.
As more people used Candor, I noticed two problems:
The system was generating too few synthetic users
The synthetic users were too similar
After some investigation, I had a hypothesis that Perplexity could be a better web search tool than Tavily. So I tested it.
I ran the same web search queries through Perplexity and Tavily and compared the results. The initial analysis was inconclusive.
Then I had Claude Code build me a blind review page. Each page shows one study brief and two lists of search results. Nothing tells me which approach produced which list, and the sides are shuffled.
I answer one question: Which list would better help me build realistic synthetic users for this audience?
I ran about 30 of these head-to-head tests.
Perplexity was stronger on some regional searches, especially when we were looking for something specific to one country. But it performed worse at finding broader trends that were still important for creating realistic synthetic users.
That led to another round of questions. Were the search queries themselves the problem? Should we run more searches? Fewer? Should each query focus on one question instead of several?
Down the rabbit hole we went. 😄
The key is that it’s OK, and probably necessary, to evaluate things by hand.
It takes time, and you’re only one judge. But early on, someone has to establish the quality bar. You need to decide what “better” actually means before you ask a system to evaluate it for you.
The whole experiment cost $7.54. A few dollars and a bit of my time saved us from making a change blindly.
A hot new model is exactly when you need evals
This came up again with Jev, a new model from TypeSafe that’s getting a lot of attention. I’m working on a future post about Jev now.
The short version: Jev isn’t designed to generate text like ChatGPT or Claude. TypeSafe calls it a “System One” model. You give it a decision to make (yes/no, pick an option, score something) and it gives you a structured, probabilistic answer. That maps nicely to a bunch of small decisions Candor makes behind the scenes. Is this answer vague? Has the participant already covered this topic? Do we have enough evidence on this audience, or should we keep researching?
New model, interesting architecture, super fast. Let’s put it everywhere!
Nah. Let’s test it first. We took 261 real answers from our test environment and had a stronger model grade each one. That became our answer key. Then we compared Jev against what Candor was already doing.
Some of the results were pretty dramatic. For deciding whether a participant had already told us something, our old keyword rule was right 1 time out of the 26 times it fired. For detecting vague answers, the keyword approach caught 6 of 60. Jev was significantly better at both because it judges meaning instead of matching words. We also tested the decision Candor makes about whether it has enough audience research to continue. Our existing small model wrongly stopped the research 48% of the time across 490 replayed cases. Jev did that 11.5% of the time.
Then we tested Jev for fact-checking participant profiles, and it wasn’t good enough. At least not yet. Depending on the test, it would have missed somewhere between a sixth and a quarter of the real errors. So we’re not using it there. Same model, same product, different job, different result.
That’s why “Is Jev better?” is the wrong question. Better at what? On which data? Against what baseline, at what cost, and how good does it have to be before you’re willing to put it into production?
“New” isn’t a reason to use something. If anything, the opposite is probably true. And “everyone on X is talking about it” isn’t a reason either. Hype is hype. Test things on your actual problem.
Evals tell you what not to ship
In the last couple of weeks we’ve tested more than a dozen ideas that didn’t make it into Candor: different search providers, different ways of creating queries, different models for small decisions and extra verification steps.
Some worked.
A lot didn’t.
This is the part of evals I’m appreciating the most. Most discussion about evals focuses on making AI more reliable, and that’s important. But evals have another job: they tell us what not to ship.
That’s increasingly valuable because the cost of building things is collapsing.
For most of my career, engineering capacity was constrained. You had ten things you wanted and enough capacity to build two. That forced prioritization. We definitely didn’t always prioritize correctly, but there was friction. Building something cost enough that you generally had to make a case for it.
AI is stripping that friction away.
In the last couple of weeks alone I’ve merged more than 200 changes into Candor. That’s a lot compared to what I could have done a couple years ago (which, on my own, might have been zero!)
It’s also dangerous.
Shipping 200 changes doesn’t mean the product is 200 changes better.
AI coding tools are perfectly happy to keep going. There’s always another bug, optimization, prompt tweak, model, agent or verification step. They don’t get tired. They don’t have somewhere else to be. And they don’t particularly care if the product gets better.
I do.
Engineering capacity used to be a forcing function. If building is getting dramatically cheaper, we need a new one.
I think evals are part of that.
The question I’m asking is, “Can I build this?” Yup, I can.
It’s not even, “How quickly can I build this?” Usually, pretty quickly.
The harder question is: How will I know if this actually made the product better?
That’s what I’m trying to build evals for.
And for now, apparently, that means reading a lot of synthetic interviews by hand. 😬






