Product Field Work All writing
WritingCustomer insight

The customer in the loop

If we can build almost anything, the question that matters is which two things customers would actually pay for, and what the best version of each looks like. AI can’t replace talking to customers. But it can take the real signal we have, amplify it, and put a trustworthy panel of our customers inside every decision, from ranking ideas to testing screens while they’re being built.

The problem

Enterprise customers are busy. Getting an hour with the right buyer or user can take weeks, and a usability study takes longer to set up than some features now take to build. So decisions get made on a handful of conversations, the loudest request, or a guess, and usability problems show up after the code is written, when fixing them means unpacking work.

The tempting shortcut is to ask a model, “Would a CFO buy this?” That produces fluent, confident, generic answers. The better path is harder and far more useful: build proxies of our customers from what real customers have actually told us, check them against real outcomes, and keep real people in the loop where they matter most.

85%as accurate as people are at repeating their own survey answers two weeks later: AI agents built from two-hour interviews with 1,052 real people
Too shallowNielsen Norman Group’s verdict on generic synthetic users compared against three real studies. They praised every concept
32%of statistically different findings in one comparison of AI-generated and real survey data pointed in the opposite direction
How I’d know a customer panel is trustworthy
  1. Agreement with real studies, checked every time real research runs
  2. False positives: ideas the panel loved that real customers didn’t
  3. Evidence behind each answer, how many real signals it rests on
  4. Rework after real usability tests, which should fall
Why I wrote it

I build for enterprise customers in a regulated industry, where every hour with a user is precious. The most valuable thing AI could do for product work isn’t writing code; it’s telling us, with honest confidence, what’s worth writing.

The ideaAmplify, don’t invent. Build customer panels from verified human signal, our own conversations first and the far larger volume of public customer voices after, calibrate them against real outcomes, make them argue rather than agree, and plug them into the engine so every idea and every screen gets the same scrutiny, with real people confirming the calls that matter.

1From idea to confidence, then and now

One feature idea, two paths

Typical elapsed time on one shared scale. Orange marks rework caused by learning too late; gold marks the real-customer checks that remain.

Research, then build, then test

The customer in the loop

Illustrative. The new path assumes a panel already built from the company’s customer signal.

The real change isn’t just speed. Usability stops being a phase that happens two weeks after building. It happens while the screens are generated, so problems are fixed before they become code anyone has to unpack.

2Real signal is everywhere

Grounding a panel in real customers doesn’t mean limiting it to the conversations we personally had. Working hard, a product team might talk to ten target customers in a good week. Meanwhile those same kinds of people are describing their work, their frustrations and their tools in public every day: reviews of our competitors on sites like G2 and Capterra, community forums, conference talks, podcast interviews, blog posts, even job postings that spell out what a role is responsible for.

That’s real signal from real customers. We didn’t have to extract it ourselves; agents can find it, transcribe it, check who’s speaking, and fold it in. Our own conversations stay the strongest evidence. Public signal gives them a hundred times more company.

Where real customer signal comes from

Illustrative monthly volume for a mid-size B2B category, and how much weight each source should carry. Switch to see how much evidence the panel can stand on.

Before public signal counts, agents check
  • Who’s speaking: is this our target role, company size and industry, or someone else entirely?
  • When: recent enough to reflect how the work is done now.
  • Why: incentivized reviews, vendor-written posts and angry one-offs are weighted down or set aside.
  • How often: one loud voice counts once; a pattern across many sources counts more.
  • Whether we may: public doesn’t mean unrestricted. Respect each source’s terms, and keep personal details out.

3How a synthetic panel earns trust

“Synthetic users” covers everything from a one-line prompt to a carefully grounded model of a real person. The research draws a sharp line between them. Generic personas are fluent and agreeable; panels built from real interviews and records get close to how those people actually answer. Trust has to be climbed, rung by rung.

The grounding ladder

How much to trust a panel, depending on what it’s built from. Bars are my judgment, informed by the studies cited; only rung 2 has a published benchmark.

A synthetic customer is only as good as the real customers behind it. Amplify signal; never invent it.

4Ranking a thousand possible features

The dream for any product leader or founder: put ten ideas in, get back a ranked list with honest confidence, and know why a customer would pick ours over a competitor’s. A grounded panel can get surprisingly close, as long as it shows its evidence and admits when it doesn’t have enough.

A panel scores eight ideas

An illustrative panel of 48 grounded proxies across four personas, built from our interviews, calls, tickets and usage, plus verified public signal. Run it, then pick an idea to see the evidence.

Notice the idea ranked lowest-confidence: the panel doesn’t know enough, so it says so and asks for five real interviews instead of guessing. That behavior matters more than the ranking itself.

5Usability while it’s being built

The same panel can sit beside the coding agent. As screens are generated, each persona walks through the real tasks it performs and flags where it would hesitate, misread or give up, tied to how that kind of customer actually works. The agent fixes the easy things immediately; a person decides the rest.

A virtual usability pass on one screen

Bulk actions on a list screen. Switch between the first generated version and the revision after the panel’s feedback.

After the revision, three real users confirm the flow in a 20-minute session. The panel narrows what to test; people confirm it.

6Making the panel argue

The biggest danger isn’t that synthetic customers are wrong. It’s that they’re agreeable, and agreeable feedback looks like validation. A useful panel is designed to find the reasons an idea will fail.

Ground itEvidence or silenceEvery answer cites the real signals behind it, our own conversations first, then verified public ones. No evidence, no opinion.
Force trade-offsChoose, don’t rateAsk which of three options a buyer would fund this quarter, not whether each is “useful.”
Ask for moneyPay, switch or ignoreWould they pay more, switch from a competitor, or ignore it? Enthusiasm isn’t willingness to pay.
Invite dissentA skeptic in every panelOne proxy’s job is to find the objection the buying committee would raise.
CalibrateKeep scoreEvery real study is also a test of the panel. Track where it was right and wrong, and retrain it.
Know the edgesSay “unknown”Where signal is thin, a new market or a group we rarely hear from, the panel should decline to answer.
Where real people stay essential
  • Finding needs nobody has written down yet. Panels can only amplify what’s already been said.
  • Confirming the top one or two bets before real money goes into them.
  • Anything with emotional or high stakes: trust, fear of change, what happens when something goes wrong.
  • Buying committees, politics and budgets, which rarely show up in the data.
  • Calibrating the panel itself. Every real conversation makes it better.

7Plugged into the engine

The panel is most valuable when it isn’t a separate step. Exposed to coding agents the same way the knowledge server is, it becomes a tool they call while they work: before building, “which version would a coordinator choose?”; while building, “would this screen confuse an admin?”; after shipping, “does real usage match what the panel predicted?”

Every feature idea runs through the same process, with the same scrutiny, every time. And every real conversation, ticket and outcome flows back in, so the panel gets more accurate the longer it runs.

CalibrationAgreement between panel predictions and real studies
False positivesIdeas the panel favored that real customers rejected
Evidence depthReal signals behind each recommendation
Idea to decisionTime from an idea to a ranked, evidence-backed call
Late reworkChanges forced by usability issues found after building
Real contactHours with real customers, which should not fall

Choices

Grounded, not genericPanels are built from real customers’ words and behavior, ours first and verified public voices after, not from a model’s idea of “a CFO.”
Confidence you can seeEvery ranking shows its evidence and its uncertainty, and low confidence triggers real research.
More real contact, not lessThe panel narrows what to ask real people. It never becomes the reason to stop asking.
Where this can go wrong

A panel that’s trusted too much becomes a mirror: it reflects what customers have already said and misses what they haven’t. Built from a biased sample, it amplifies the bias, and public signal has its own skews: review sites attract extremes, some reviews are incentivized, and loud communities aren’t always representative. Customer data used to build proxies needs consent and careful handling. And the numbers in my panel example are illustrative; any real panel has to earn its trust against real outcomes before anyone ranks a roadmap with it.

Sources

  1. Park et al., “Generative Agent Simulations of 1,000 People,” 2024: agents built from two-hour interviews replicate survey answers 85% as accurately as people replicate themselves, and reduce bias compared with agents given demographic descriptions.
  2. Summary of Nielsen Norman Group’s comparisons: synthetic users “too shallow to be useful,” praising every concept; and a survey comparison where a third of significant differences flipped direction.
  3. Userbrain, synthetic users experiment: synthetic results only become trustworthy when compared with real ones.
  4. Ballpark research glossary: use synthetic users to prepare for real participants, never instead of them.
  5. Anthropic, Model Context Protocol: how agents can call a panel as a tool while they build.