| Dimension | Score |
|---|---|
| Output reliability | 3/5 |
| Workflow fit | 4/5 |
| Context handling | 3/5 |
| Learning curve vs. payoff | 4/5 |
| Failure transparency | 3/5 |
| Overall | 3.4/5 |
What AI does well
Drafting PRDs and specs. Claude or ChatGPT turn a messy voice-note braindump into a structured PRD in minutes; purpose-built ChatPRD does the same with PM-specific templates. Observed gain: first drafts in 15 minutes instead of a half day. Caveat: the draft inherits your framing errors at machine speed, and a wrong problem statement gets beautifully formatted.
Synthesizing user feedback. Feeding a quarter of support tickets, NPS verbatims, and interview notes into a long-context model (Claude handles the longest inputs best) produces theme clusters that used to take a research sprint. Caveat: it over-weights articulate complainers and under-counts silent churn; check counts against the raw data.
Competitive and market research. ChatGPT with browsing or Perplexity compresses a day of tab-hopping into an hour of directed questions. Caveat: market-size numbers and pricing details are hallucinated often enough that every figure needs a source click.
Meeting overhead. Notetakers plus a workspace AI (Notion AI or Coda AI) handle recaps, action items, and status updates. This is the least glamorous and most reliable win: pure secretarial load, gone.
Where it fails
Prioritization. Given a backlog and asked what matters, models produce confident, generic rankings that optimize whatever scoring framework you named, not your actual strategy. The failure is invisible because the output looks like reasoning. A concrete example: asked to RICE-score a backlog, a model invents reach numbers with three significant digits for features that have no usage data at all.
Stakeholder judgment. AI cannot know that engineering is burned out, that the CEO’s pet feature is political, or that a customer’s ask contradicts their behavior. PMs who paste AI-written tradeoff analyses into exec reviews get caught by the first follow-up question.
Anything that requires saying no. Models are agreeable by training. A roadmap assistant that never pushes back is a yes-machine with formatting.
My take
My take (August 2026, Bernat Sampera)
The SERP for this query is courses and vendor pitches, because the honest answer does not sell either: AI is a drafting engine for product managers, not a product manager. Used that way, the gain is real and boring. The writing tasks above compound into roughly a day a week, mostly from PRD drafts, feedback synthesis, and meeting overhead. Used as a decision engine, it is worse than nothing, because the failures look like reasoning and survive review. The verdict is falsifiable: track PRD cycle time and decision reversal rate for a quarter. Cycle time should drop while reversal rate stays flat. If reversal rate climbs, AI is making your decisions, and you should stop. Tool choice matters less than discipline here: any assistant above drafts well enough. The skill is never shipping an AI-written tradeoff you cannot defend from memory in an exec review.
What I would build
The gap in every tool above is memory. Each session starts from zero: the model does not know your product, your users, or the decisions you already made. What I would build is a persistent context layer for PM work: the product strategy, past decisions with their reasons, and customer segments with their real usage, stored where every AI conversation can read them. With that in place, drafting gets sharper, because the PRD inherits your strategy instead of a generic template, and the failure modes above shrink, because the model can be checked against your recorded decisions. Until something like that exists in your stack, the workaround is manual: keep one living context document and paste it into every session that matters.
Verdict (August 2026, Bernat Sampera): AI now does the writing half of product management well (PRDs, summaries, updates) and still cannot do the deciding half (prioritization, tradeoffs, saying no); a PM who treats it as a drafting engine gains a day a week, and one who treats it as a decision engine ships worse roadmaps faster. Overall: 3.4/5.