Demand first. The model second, and only if it is genuinely in doubt.
Test demand first, with a landing page and ten conversations. Test whether the model can do the job only if your task is genuinely unsettled, which most aren't. Then check the economics. Score all three before you write any code, and be willing to score your own idea a two.
An AI startup idea has two ways of being wrong, and almost every guide tells you to test them in the wrong order.
The standard framework runs problem, then technical, then solution, then a decision. It spends the first week proving the model can work. For most AI products in 2026 that's a week spent confirming something already settled, while the question that actually kills projects goes untested.
This is the order we use, the cheap tests that answer each question, and a score you can put on your own idea in about an hour. If you'd rather work through it with someone, that's what our AI MVP development work is.
What "Validate" Means for an AI Product
Two different questions hide inside the word.
Will anyone use it and pay for it? The ordinary startup question. Every product has it. Nothing about AI changes how you answer it.
Can the model actually do the job on your data? The one AI adds. Published benchmarks tell you a model is capable in general. They tell you nothing about your messy, domain-specific, badly-scanned reality.
The contrast worth holding onto: the first question is about people, and you answer it by watching what they do. The second is about your data, and you answer it by measuring. They need different tests, different evidence, and as it turns out, different priority.
There's also a third question that gets skipped almost universally, and it sinks more ideas than either of the others. We'll come to it.
Why AI Idea Validators Can't Answer the Second One
Search this topic and half the results are tools that promise to validate your idea for you. They're better than their marketing suggests, and worth being precise about.
| ValidatorAI | Preuve | |
|---|---|---|
| Price | Free | $29 per report |
| What it does | Scores an idea against behavioral data from founder sessions | Scans 60+ live sources for competitors, market size and demand signals |
| Output | Validation score and next steps | 0 to 100 viability score, TAM/SAM/SOM, competitor map, pivot suggestions |
| AI-specific? | No. AI and tech is about an eighth of what it sees | No. Same flow for SaaS, marketplaces and physical products |
| Tests your model? | No | No |
Checked 18 August 2026.
Here's the honest read, and it costs us work to give it. These tools are genuinely good at the market-research half. Competitor maps, pricing benchmarks, whether anyone is searching for this, which communities complain about the problem. That's real research and a free tool does it in a minute. Use one. Don't pay an agency to redo it.
What they cannot do is structural, not a gap in the products. They use AI as a research engine to validate a business. Not one of them can tell you whether your model will hit your accuracy bar on your data, because they've never seen your data. That question only gets answered by running something against a real sample.
Which is why the search results for this topic are confusing. Most of them answer "how do I use AI to validate an idea," when what you asked was "how do I validate an idea whose whole premise is AI." Different question, different article.
Test Demand First. Feasibility Usually Isn't the Risk.
This is where we part company with every framework you'll find.
The usual advice is to prove the technology first, on the reasoning that there's no point testing demand for something you can't build. Sensible in 2019. Mostly wrong now.
If your product classifies text, pulls fields out of documents, summarizes, translates, or drafts from a template plus retrieved context, the model can do it. Those capabilities are settled. Spending week one proving it is not de-risking, it's a week you paid for.
Feasibility is worth testing first only when your task sits in the unsettled set: unusual data, an accuracy bar that's high and non-negotiable, a genuinely novel task, or a hard latency or cost ceiling. We work through exactly when that applies in our guide to AI PoC vs prototype vs MVP, so we won't repeat it.
So unless you're in the unsettled set, run demand first. It's cheaper, it's faster, and a no answer saves you the feasibility work entirely.
The Landing Page Test: What It Costs and What Counts
The cheapest real signal you can buy.
One page. One promise, stated plainly. One call to action. No demo, no product, no screenshots of something that doesn't exist. Then a small paid test, $50 to $100 in ads, pointed at the narrowest audience you can define.
What counts as a signal, weakest to strongest:
- A click. Almost meaningless. Curiosity is free.
- An email address. Weak. People give these away.
- A waitlist signup above roughly 5% of visitors. A real signal that the promise lands.
- A booked call. Strong. They spent time.
- Money, or a signed letter of intent. The only one that settles the question.
Be honest about what this measures. A landing page tests interest in a promise, not in your product. The promise is always better than the thing, because the promise has no latency, no errors and no edge cases. Plenty of people sign up for a waitlist they'd never pay for. Treat a 5% conversion as permission to keep going, not as proof.
Wizard of Oz: Faking the AI on Purpose
The trick that saves the most money, and almost nobody runs it.
Build the interface. Skip the model. When a user submits something, a person reads it and writes the response by hand, then it appears in the product as though a system produced it. The user gets a real answer. You get a real behavioral signal. You built nothing.
It answers questions a landing page can't. Do people come back? Do they send you the hard cases or the easy ones? Is the output they want the output you assumed? That last one is where most product assumptions die, and it's much cheaper to find out with a human typing than with three weeks of prompt engineering.
Two boundaries worth respecting. Don't run it at a scale you can't staff, because the whole point is that a person is doing the work. And don't tell people it's automated when it isn't, in any context where that matters to them. "Early access, responses within an hour" is honest and works fine.
How Many People to Talk To (and What to Ask)
Five to ten. That's our working number after running this on client projects, and it's an Amplence estimate rather than a benchmark from a study.
The logic is that you're not measuring anything, you're listening for surprise. If the first five people describe a problem you didn't expect, you learned the thing. If eight people in a row say the same thing, the ninth rarely changes your mind. Talking to fifty is usually procrastination wearing a lab coat.
Ask about the past, not the future. "Would you use this?" gets you politeness. These get you evidence:
- Walk me through the last time this happened. Specific and recent, or it didn't happen.
- What did you do instead? The real competitor is usually a spreadsheet or an intern.
- What did that cost you? In money or hours. If neither, there's no problem to solve.
- How often does it come up? Once a quarter is a feature. Twice a day is a business.
- What would have to be true for you to stop doing it that way? This surfaces the real switching bar, which is almost always higher than founders assume.
One AI-specific question worth adding: how wrong can this be before it's useless to you? The answer is your accuracy bar, and you want it from them, in their words, before anyone starts building toward a number you invented.
Our Find The Plan build is what this looks like carried through: a healthcare navigation concept scoped deliberately around synthetic data, with no personal health information anywhere in it, precisely so the idea could be tested before anyone took on a compliance program.
The Feasibility Test You Can Run in Two Days
If you are in the unsettled set, here's the cheap version.
One engineer. A hundred real examples, chosen before you start and including the awkward ones. An afternoon of prompt iteration. A spreadsheet with a number in it at the end.
That's it. Not a six-week proof of concept. Two days tells you which of three worlds you're in: nowhere near the bar, close enough to be worth a proper study, or already past it. Two of those three outcomes save you the study entirely.
Set the number you need before you run it, because otherwise you will look at whatever you get and decide it's encouraging.
Validating the Economics Before You Commit
The third question, and the one that gets skipped.
An idea can pass demand validation and pass feasibility and still be a bad business, because the cost of producing one good answer is higher than anyone will pay for it. This is not a rare edge case. It's the most common way a validated AI idea dies quietly eighteen months later.
You can estimate it before building. Take the work one user needs done in a month. Estimate what it costs you to produce it, including the retries and the ones the model gets wrong. Compare that with what the people you interviewed said they'd pay. If those two numbers are close, you don't have a margin, you have a hobby.
We work through the arithmetic properly in our AI MVP cost breakdown rather than repeating it here.
The Build Score: Rate Your Idea in Three Dimensions
Score each dimension 1 to 3 on what you actually observed, not on how you feel about it. Multiply them.
- 1Positive conversations, no commitments.
- 2Signups, a waitlist above 5%, or booked calls.
- 3Money, signed letters of intent, or people asking when they can pay.
- 1Untested, or tested and nowhere near the bar.
- 2Close to the bar on a real sample, needs work.
- 3Past the bar on a held-out sample, or a settled task with no real doubt.
- 1Cost is near or above what people would pay.
- 2A margin exists but it's thin, or you're guessing.
- 3Comfortable margin at a price people already agreed to.
The useful move isn't the number. It's noticing which dimension you scored lowest and running only that test next.
What Validation Can't Tell You
Four honest limits.
It can't tell you about timing. Plenty of correct ideas were tested and killed a year before the market arrived. Validation measures now.
It can't survive a founder who doesn't want the answer. If you already decided to build it, you'll find a way to read any result as encouraging. The score is only useful if you were genuinely willing to write down a 1.
A small sample can mislead you. Ten conversations in one niche is a signal, not a market. It's enough to justify building something small. It isn't enough to justify a two-year roadmap.
It says nothing about whether you can ship it. Demand, feasibility and economics can all be green while the team, the timeline or the budget is the actual constraint. We covered how that plays out in why most AI MVPs never reach production.
If the Score Says Don't Build
The section that costs us business, so here it is plainly.
Score under 8: don't build, and don't hire anyone to build it. Not us, not a cheaper agency, not a freelancer. A low score means a question is still open, and paying somebody to write code doesn't close it. Go and run the missing test.
Score 8 to 17: buy the smallest thing that answers the weakest dimension. Usually that's a landing page and ten conversations, which you can do yourself in a fortnight for the cost of the ads.
Score 18 or above: build the smallest real version, not the roadmap. One workflow, real infrastructure, actual users. Our guide to building an AI MVP picks up from here.
And if the honest answer is that the idea is a feature rather than a company, that's a result too. It's cheaper to learn it from a spreadsheet than from a launch.
Want a second opinion on your score?
Send us your three numbers and what you did to earn them, and we'll tell you which one is softest and what the cheapest next test is, including when the answer is that you shouldn't build it yet. Bring the score.
Book a 20-minute scoping call ↗Frequently Asked Questions
How do I validate an AI startup idea without building anything?
A landing page with one promise plus $50 to $100 in ads, five to ten conversations about what people did last time the problem occurred, and a Wizard of Oz version where a person produces the output by hand. All three run without a model and answer the question that kills most ideas, which is whether anyone wants it.
Are AI idea validator tools like ValidatorAI or Preuve worth using?
For market research, yes, and ValidatorAI is free. They map competitors, size the market and surface demand signals quickly. What none of them can do is tell you whether your model will hit your accuracy bar on your data, because they have never seen your data. Use them for the first half, then do the second half yourself.
How many people should I talk to before building an AI product?
Five to ten, in our experience, which is an Amplence estimate rather than a benchmark from a study. You're listening for surprise rather than measuring anything, and by the eighth similar conversation you've usually stopped learning. Ask what they did last time, not whether they'd use it.
What conversion rate proves demand for an AI product?
A waitlist converting above roughly 5% of visitors is a real signal that the promise lands, but it isn't proof. Signups measure interest in a promise, which is always more attractive than a product. Booked calls and money are the signals worth trusting.
How do I know if the AI part will actually work?
Run a hundred real examples through a hosted model over two days and write down the result, having set the number you need beforehand. For classification, extraction, summarization and drafting you'll usually find the capability is already there. The unsettled cases are unusual data, hard accuracy bars, novel tasks, and tight latency or cost ceilings.
What if validation says the idea is good but the economics don't work?
That's a real result and a common one. Your options are raising the price, narrowing to the segment that will pay it, cutting the cost per answer with a smaller model or caching, or not building it. Doing it anyway and hoping unit costs fall is the one option that reliably fails.



