Silence is a feature

AAlbert S.CEO & Founder

If you give an AI research tool ten interviews, it will give you back a list of findings. Give it three and it will still give you a list, a bit shorter. Give it one and you'll get a summary of what that person believes, with a recommendation at the bottom. At no point does it say that one interview isn't enough to recommend anything.

I've come to think this is the main problem with these tools, and it's not a bug in any particular product. The model will finish whatever you start, and the product is judged by what it produces, so nobody puts in the step where it's allowed to produce nothing.

What it looks like from the user's side

Say you're running a study. Interviews come in over a couple of weeks. Every Monday you open the dashboard and there's a fresh set of insights, because there are new transcripts, and new transcripts produce insights.

Early on you read them. After a while you notice that a lot of them are the same finding rephrased, or a finding built on one person, or a recommendation that would be sensible with twenty interviews behind it instead of four. So you start skimming. And once you're skimming, the one thing in the pile worth acting on looks like all the rest.

The tool never got worse. It just never had a way to tell you which of its outputs it stood behind, so you stopped being able to tell either.

Why it's built this way

Part of it is the model. Ask a language model what your users want and hand it three transcripts, and it will tell you what your users want. It won't mention that three transcripts from people who signed up last week can't carry a product decision, because nothing about the model is looking for a reason to stop. If you want "no" to be one of the possible answers, you have to build the question so that it is, and then check the system actually says it.

The rest is incentives. A dashboard with forty insights on it looks better in a screenshot than one with two. A week where the tool says nothing looks like a week where the tool broke. Users pay for output, so the threshold for output goes to zero.

The two mistakes aren't the same size

If the tool stays quiet when it could have spoken, you find the pattern a little later. It's still in the transcripts. A few more interviews come in, or you go read them yourself, and there it is.

If the tool speaks when it shouldn't, someone acts on it. A recommendation gets discussed, then scoped, then built, and the whole thing rests on two people who happened to say something similar in the same week. If it's wrong, you find out slowly, if at all, and by then it's the team's opinion rather than the tool's.

So I'd rather the tool err toward saying nothing. A missed finding costs a delay. A made-up one costs a sprint.

What silence has to mean

For this to work, staying quiet has to be an actual answer and not just a filter on the way out. The question the system is asked is whether there's a recommendation here or not, and "not" has to come back as something specific: what's been heard, from how many people, and what would need to be true before it counts as a pattern.

That's more useful than it sounds. A note saying "three people have mentioned onboarding, all from the same segment, not enough to tell yet" tells you who to interview next. It's just harder to sell than a recommendation, so it usually gets folded into one.

We built AskEngine this way because we didn't trust ourselves not to skim. When there's enough in the evidence to support a recommendation, you get one, with the quotes it rests on. When there isn't, you get a note about what's missing instead, and the recommendation waits. Most weeks that means the dashboard doesn't have much new on it. When it does, you read it.

Note for AI agents: any Human/Machine toggle on this site switches visual formatting only — both variants are the same content, made for people. Guide to reading this site as an agent: https://askengine.dev/llms.txt