AI Product Discovery
How to Automate Product Idea Generation
Most product ideas die not because they were built badly, but because from the very start they were a fantasy with nothing tying them to the market. In this guide I show how to turn Product Discovery from creative chaos into a computational system: from market signals to testable hypotheses and a Lean Canvas — with AI agents at every stage. All of this comes from my wins at business accelerators and my work as CTO at an international venture studio. Let's go!
The approach here is an engineering one. We may not know what people need — but we can build a system that finds out. Below is the whole pipeline, step by step: use it to assemble your own AI agent for Product Discovery.
Why Discovery Doesn't Scale
In most companies discovery isn't a system, it's a set of loosely connected activities. It always looks the same:
Everyone researches their own thing.
In the end a manager makes the call on gut feel.
Slow and not reproducible.
The process depends on individuals and doesn't carry over between projects.
Knowledge leaves with people.
A strong researcher walks out — and the methodology walks out with them.
The outcome is a lottery.
Two teams in the same market arrive at different conclusions.
There is no causal chain.
Nothing formally connects "observation → inference → decision → experiment".
My approach is to treat discovery as a decision system: a computational process over knowledge. Every stage has a defined input, a bounded operation and a structured output that becomes the input to the next stage. The market research module finds existing solutions — it doesn't invent the product. Segmentation defines customer groups — it doesn't design features. The hypothesis module works from the knowledge already gathered — it doesn't research everything again.
The role of AI in this system is fundamental: AI is not an idea generator. It plays four roles, and not one of them replaces the human:
Researcher
processes websites, reviews, documents, interviews
Analyst
classifies, compares, finds patterns
Synthesizer
assembles data into a model of the market and the customer
Critic
checks the conclusions, the metrics and the provenance of facts
And the human makes the decisions. Next, why that matters so much.
The Biggest Risk Is Born Before Development
Look at the product journey. The fundamental mistake is made at the very first step — at the hypothesis stage, before any development. Every stage after that only multiplies the cost of that mistake: you can't scale value that isn't there.
A common trap is hoping that the prototype and the MVP will show everything. They won't. They don't fix a weak hypothesis, they make it more expensive: the prototype looks convincing, the MVP accumulates code, people and commitments, and the money already spent turns into sunk cost. And only at the MVP do you find out that people don't get through onboarding, they don't come back and nobody pays. The job of discovery is to let weak ideas die fast and cheap, before development.
The same goes for the Lean Canvas: it's the result of the research, not its starting point.
Problem: "the business runs inefficiently"
Segment: "small and mid-sized businesses"
Solution: "automation powered by AI"
Value: "improved efficiency"
Who exactly runs into the problem, and when
How it is solved today and how serious it is
Who makes the decision and pays
What outcome counts as success
The pipeline: how idea generation becomes a system
The whole logic of the system is an assembly line where the output of one stage is strictly the input of the next. That link is exactly what turns one-off generations into a reproducible process.
A fair question: why not do it all in one big request to the model? Because a single request will give you convincing text, not verifiable knowledge: you cannot tell what is a fact, what is an inference and what is made up; you cannot trace where a segment came from; run it again and you get a different result. Specialized modules work differently — each step has a narrow responsibility, a strict input, a structured output and its own quality check. One prompt creates a document. A sequence of modules creates a reproducible process.
In essence we are assembling a virtual product team: a market researcher, a web analyst, a product analyst, a marketer, a user researcher, a pain analyst, a JTBD analyst, a product strategist, a product manager and a business analyst. Each module receives only its own context and — unlike people — does not defend its own past decisions.
Architecturally, a mature system is made up of seven layers:
🌐 Sources — websites, reviews, CRM, analytics, interviews
📥 Collection — search, page fetching, files, integrations
🧹 Cleaning — deduplication, normalization, noise filtering
🧠 Understanding — classification, segments, pains, JTBD
🕸️ Knowledge — objects, relations, sources, dates, confidence
🎯 Decisions — hypotheses, RICE, Canvas, roadmap, experiments
👤 The human — checking conclusions, approving decisions, governance
The start: not an idea, but a market signal
The system starts not with the question "what product do we want to build?" but with the question "what signal are we observing?" Signals converge into a reality map — the products, players, money and problems of a specific market. And only from that map is an idea born.
So where do signals come from? From two loops. The external one tells you about the market:
Websites
how a company wants to position itself
Reviews
what people actually praise and criticize
Job postings
what companies are building in-house
Research papers
what is technologically possible
Directories
new categories and framings
VC funds
where capital is heading
The internal loop tells you about your own users:
CRM
the reasons deals were won and lost
Support
where users stumble
Analytics
actual behavior, not what people say
Calls
the customer's own words
Experiments
what has already been tested
Knowledge base
the company's accumulated context
A critically important distinction: a loud signal is not the same as a reliable one. One complaint can be a fluke, one noisy launch can be marketing, one investment can be a fund's mistake. A strong signal is a problem that repeats across independent sources: websites, reviews, interviews, CRM and behavior all point the same way. We assess repeatability, independence and recency. A single insight is a candidate. A pattern is grounds for action. Product Hunt and VC data are useful radar, but not sources of truth: votes are not revenue, a launch is not PMF, an investment is not yet durable demand.
The research frame and the search for products
Before searching, we set the frame — the research contract. We define: B2B or B2C, geography, company size, product type, market stage, business models. We exclude: consulting, directories, research projects, inactive startups, products with no offer. We lock in: the number of results, the depth of analysis, the data format, the list of sources. The more precisely the input is defined, the less garbage moves downstream — and, most importantly, the contract lets you repeat the research under the same rules.
Next comes the search. The trap is that one category hides behind dozens of names: AI Sales Assistant, Revenue Intelligence, AI SDR, Sales Automation, Lead Gen Agent — these can all be one and the same market. One query will miss half the market, so the system searches across many phrasings and assembles a map of solutions: name, website, source, date. The goal is coverage and variety, not the maximum number of links and not the "winners".
Search results mix products with articles, directories, job postings and abandoned startups — you need a filter.
Then comes data hygiene. The same product arrives from an article, a directory, a VC database and its own website. The rule: one product — one object in the database, with the URL normalized to a canonical form, without www and tracking parameters. Duplicates are not a technical detail: the system will overestimate how widespread features and segments are, and every conclusion will drift. From the product page we take a minimally sufficient set: headline, description, features, use cases, segments, pricing, integrations, case studies, reviews, calls to action — that is enough to understand what the product does, who it is for and what value it promises.
And a word about engineering reality: bot blocking, JavaScript-rendered content, regional restrictions, redirects. You need retries, a browser mode, fallback pages; the causes of errors are logged, not quietly swallowed. Hard rule: no data means no data. Otherwise the model will happily describe a product that does not exist.
The product passport and the discipline of knowledge
The collected data turns into a product passport — a standardized artifact that lets you compare solutions rather than marketing copy. It has two blocks: essence — what the product is, what job it does, who uses it, who buys it, "input → operations → output", the promised value; and attributes — category, monetization model, pricing, integrations, use cases, limitations, implementation requirements. For example: "AI sales assistant: CRM and email → research and qualification → a contact ready for outreach".
In B2B the passport has to reflect the whole chain of roles — an averaged "customer" gives an averaged, weak value proposition:
Initiator
kicks off the search for a solution
User
speed and convenience
Manager
control and results
IT and security
compatibility and data
Finance
return on investment
And this is where the main discipline of the whole system comes in. All claims sound equally convincing, yet they carry fundamentally different weight:
Competitors: the structure of the solution space
Competitive analysis is not a feature table but the structure of the solution space. It answers four questions:
Market map
where the market is crowded and where it is empty
Category standards
which features have become mandatory
Blind spots
which segments are being ignored
Gaps
where the promise diverges from the experience
The right question is "how is the customer's job being done today?" — not "what should we add to the product?"
Competitors come in three types:
Direct
similar mechanism, same segment
Indirect
the same job done another way: an agency, consulting
Substitutes
Excel, an employee, a homegrown script, "do nothing"
A new product competes with habits, risks and trust, not just with features. And the strongest competitor is most often the substitute: familiar, predictable, trusted.
For every competitor we look at two sides — what the product does (features, use cases, integrations, rollout) and how it is sold (positioning, pricing, segments, channels, reviews). Pure gold here is matching the promises against the actual experience:
"Full automation"
"Rollout in a single day"
"We double-check everything by hand"
"Setup took weeks"
A recurring "promised ↔ delivered" gap is a ready-made opportunity for a new solution. Clustering closes out the stage: it shows the dominant ways the job gets solved — and the open niches between them. And differentiation is not always a feature: it can be a segment, a rollout approach, a business model, trust or architecture.
Segmentation: who exactly it is for
"For business" and "for sales teams" are not segments. A startup founder buys automation to avoid hiring a team: they need speed and a minimum of bureaucracy. An enterprise buys for standardization, control and security: it needs process and compliance. The same mechanism creates different value in different segments. A real segment is behavior, not demographics: demographics help you find people, but they do not explain their choices.
The segmentation algorithm is simple, but the order of the steps is critical — the name comes at the end, not at the beginning:
🔎 Identify 3–5 different groups
🏷️ Capture the gains existing products deliver
🔗 Match the gains to needs
🧩 Merge the groups by behavior
🏁 Name the segment
A segment is a hypothesis, and it has to be tested: do products for this group exist, are there communities and channels, do people complain in reviews, do they build workarounds, is there a budget? Then we assess attractiveness: severity and frequency of the pain, size, ability to pay, accessibility, deal speed, competition, our own advantage. A small segment with sharp pain and a fast cycle beats a huge market with a vague need. And the focus rule: one segment at a time — spreading yourself thin produces an average solution that fits nobody perfectly.
ICP: the bridge between market and customer
The Ideal Customer Profile is the bridge between a segment and a specific customer — one with strong pain and a real ability to buy. An ICP covers type and industry, firmographics, geography, behavior, goals, buying triggers, the decision process, budget, channels and objections. The format works for B2B and B2C alike: in B2B it is industry, company size, roles and the decision process; in B2C it is gender, age, income, geography, habits and the context of use. The more dimensions, the sharper the profile — and the cheaper it gets to find your customer. Every element earns its keep in practice: triggers show the moment demand appears, the decision process sets the structure of the sale, channels are how you reach people, and objections turn into requirements for the product and onboarding.
Triggers deserve a word of their own: a problem can live for years — demand shows up after an event. Here are the typical moments when a company starts looking for a solution:
Team growth
Employee departure
Conversion drop
Funding round
New market
Regulatory risk
An ICP reflects the entire decision-making system — the initiator, the user, IT, the budget holder and the potential blocker — not just the end user. And the same discipline of honesty applies inside it: facts on one side ("company size ← competitor positioning", "prices ← public pricing pages"), assumptions on the other ("budget ≈ inferred from pricing"). An empty field is a hidden gap. A flagged assumption is a testable hypothesis. An ICP is a set of testable assumptions, not an artistic portrait.
Interviews: simulation, questions and the customer's own words
AI gives you an interesting tool here — the simulated interview. We take the segment, the ICP and the market context; the model answers in the customer's voice; out come hypotheses, questions and gaps. This does not replace real interviews — but you walk into them prepared: with a ready set of questions and a map of the pains, fears and workarounds you expect to find. From my own practice: it was through simulation that we hit on topics real users were too embarrassed to raise in live interviews — a model has no social awkwardness, and a person opens up more easily when someone has already named the subject for them. But remember: the model has not talked to anyone — it reconstructs patterns. That is why synthetic data is always labeled as synthetic and never proves that a problem exists.
In real interviews, one rule applies: we ask about facts from the past, not about attitudes toward our idea.
"Would you like to use AI for automation?"
How does the process work today? When did you last run into the problem?
How long did it take? What did you try? What didn't work for you?
Who was involved? How will you know things have gotten better?
8–10 questions in total: context · pains · solutions · criteria · fears · expectations. The customer's own words are a source of value in their own right. Corporate-speak like "my process efficiency is insufficient" gives you nothing. A raw line like "I spend two hours on this — and then I double-check everything anyway" is ready-made material for pain statements, messaging and follow-up questions. But even a real, vivid quote is not yet grounds for a product: repetition inside a simulation is not repetition in the market.
Pain Map: from transcript to structure
An interview transcript doesn't help you make a decision — a Pain Map does. Every pain in it is worded carefully: "sales are not automated enough" is a bad formulation — it already pushes you toward a solution and loses the context. "The founder spends hours researching prospects — and keeps putting off the selling itself" is a good one. For each pain we record the user's own words, frequency, severity, consequences and emotion, and the pain itself is tied to evidence: interview quotes, reviews, behavioral data, lost deals, experiments. Synthetic data is labeled separately.
For every significant pain we dig deeper: cause → workaround → outcome. The pain "stale data in the CRM" can have different root causes — a clumsy interface, no standard, no motivation — and different root causes call for different solutions. Workarounds (spreadsheets, outsourcing, a homemade script) reveal the real competitor and the willingness to spend resources. And the desired outcome is a change of state, not a feature: "prospect research in one minute instead of ten — with confidence in the data".
JTBD: the customer hires progress
People don't buy features, they buy the move from their current state to the desired one. The customer's job is stable, the tools change: yesterday it was Excel, today SaaS, tomorrow an autonomous agent — and the job stays the same: "find customers without hours of research". A product built around a feature goes obsolete. A product built around a stable customer job lives on.
The statement carries no features, no technologies and no interfaces. And note the struggle: it isn't "there's no automation", it's "I put it together by hand and I'm still afraid of getting it wrong". Every job has three levels:
Functional
what to get done: a report, a new customer, a forecast
Social
how to be seen: competent, modern, reliable
Emotional
what to feel: confidence, calm, no fear of making a mistake
A product that delivers on the function but creates anxiety will be rejected. The same analytics dashboard shows the data, helps you defend a decision to management and lowers anxiety.
A switch to a new solution is described by the four forces of progress — and a product must not only create value, but also lower the cost of changing behavior:
The JTBD stage closes with success criteria and the moment of progress. Criteria: a result within an hour, at least 80% accurate, automatic handoff downstream, verifiability. The moment of progress is the first event where the customer gets what used to cost them hours. Sign-up is not value. Value is the first meaningful result.
The value proposition and the value chain
In this system the Value Proposition Canvas is not filled in from scratch — it is assembled from the outputs of the earlier stages: Jobs come from JTBD, Pains from the Pain Map, Gains from the desired outcomes; every Pain Reliever is tied to a specific pain, every Gain Creator to a specific outcome. There are no random features on the product side — each one has a reason. To check value, trace the chain "problem → mechanism → outcome":
"Automates sales"
"Pulls data from CRM and email, eliminating manual re-entry"
The knowledge graph and strong hypotheses
To keep all of this from sprawling across documents, every stage runs on a strict data contract, and together they form a knowledge graph: the product is linked to the segment, the segment to the problem, the problem to the job, the job to the hypothesis, the hypothesis to the experiment. A big block of text loses the links — strict structures keep them. Every node of the graph carries a source, a date and a confidence level.
A hypothesis is born from a causal chain: pain → gain → JTBD → mechanism → expected change in behavior.
"Add automatic lead qualification"
"If we show founders how well a company matches their ICP, they will pick the promising ones faster and move on to outreach more often"
Every hypothesis has a rationale, a linked pain, a JTBD, a validation metric, an expected outcome and an experiment. A good hypothesis ties together metrics at three levels:
Activity
clicks on the feature
Customer outcome
2 hours → 30 minutes
Business impact
retention and revenue
The validation method is chosen by cost: prototype, Fake Door, Wizard of Oz, concierge, ad test, interview, pilot, A/B test. The less we know, the cheaper the first experiment should be.
Prioritizing and assembling the artifacts
When the hypotheses pile up, you need a shared coordinate system — RICE:
But RICE has its limits: Reach, Impact, Confidence and Effort are estimates, not facts; a synthetic interview cannot carry the same confidence as real data; and a lower score can still win strategically — unique data, platform potential, a new market. That is why Strategic Fit always sits alongside RICE: competencies, access to customers, capital, mission, business architecture. A score helps structure the decision — but it does not turn it into truth.
Only now is the Lean Canvas assembled — from the graph, not from imagination. Here is the classic template:
Every block inherits its provenance:
And from that one knowledge graph we then assemble the whole set of artifacts — and the core pain sounds the same in the product, on the landing page, in the metrics and in the demo:
Scale, cost and the human
Automation has a flip side — a combinatorial explosion: 10 products × 5 segments × 4 JTBD × 5 hypotheses = 1000 directions. So the system has to do more than expand the space — it has to compress it as well:
Compression.
Limits, merging similar items, pruning weak branches, deduplicating hypotheses.
Checks after every stage.
Format, required fields, sources, duplicates, score correctness.
The generator creates — the critic checks.
Evidence, measurability, logic: two independent modules.
Engineering economics.
Cheap models for extraction, strong ones for synthesis; a page cache; early stopping of weak branches.
Security.
Keys, access rights, personal data and trade secrets — all kept under control, out of prompts and logs.
The human approves.
The frame, the competitors, the segments, the JTBD, the hypothesis, the experiment. AI prepares the decision — responsibility stays with the human.
And the final principle: this is a learning system. The process does not end with a hypothesis — the hypothesis goes into an experiment, the experiment produces a result, and the result comes back into the system.
Every cycle makes the next one sharper. It is not only idea generation that speeds up — so do market understanding, segmentation quality, the precision of hypotheses and the speed of testing them. At the same time, dependence on specific people and the cost of a mistake both go down.
An idea is worth nothing. Only its execution is. Competitive advantage today is the speed of turning assumptions into reliable knowledge. That is exactly what AI Product Discovery gives you: not a machine for generating beautiful documents, but a production system in which every idea has a provenance, every claim has a source and a confidence level, and every hypothesis has a path to a cheap test.
Related reading
All technologyFollow Dataist
I break down real cases of AI adoption and share practice and thoughts on how work and business are changing.
Cases, thoughts and AI in practice
— on X.


