Customer Interviews With AI: Use AI Without Losing the Evidence

Updated on

Use AI for interview preparation, live assistance, moderation, transcription and synthesis while preserving source evidence, participant context and human judgment.

Share:

AI can help prepare, transcribe, translate, organize, challenge and even conduct parts of customer-interview research. It does not make the research unbiased. Every AI step adds a transformation between the participant and the Product decision, and that transformation can introduce new errors as well as remove work.

Use this operating model:

Research decision → AI role → Data gate → Bounded AI task → Source-linked output → Human verification → Evidence status → Product decision

The central rule is simple:

AI output is not customer evidence merely because it was generated from customer evidence.

A transcript produced from audio is a transformation. A translated transcript is another transformation. A cluster, summary or stakeholder brief is further downstream again. The more transformations you add, the more important source traceability becomes.

If you need the complete interview workflow — method fit, participants, running sessions, synthesis and deciding what happens next — start with Customer Interviews for Product Discovery. This page owns the narrower job: how AI may enter that workflow without corrupting evidence, context, participant handling or research judgment.

Do not treat “using AI” as one research method

The risk changes with the job you give the system. A model that critiques a draft guide is not doing the same thing as a model that asks a participant questions, and neither is doing the same thing as a model that generates a synthetic participant response.

AI roleUseful jobsHuman still owns
Colleague / criticChallenge a plan, flag question risks, suggest alternative interpretationsResearch decision, method, participant definition, final guide
Silent live assistantTrack objectives, suggest possible probes, flag uncovered topicsWhether to ask, wording, timing, rapport, interpretation
Moderator / interviewerAsk and adapt questions with a real participantResearch design, controls, quality evaluation, escalation, evidence interpretation
AnalystTranscribe, translate, extract, code, cluster, compare, search contradictionsFinding validity, context, source verification, interpretation
BrieferDraft a research summary or stakeholder digestWhat is true enough to communicate and what uncertainty remains
Synthetic respondentRehearse, generate edge cases, pressure-test a guide or analysis pipelineThe explicit boundary between simulation and customer evidence

AAPOR's 2026 responsible-AI work describes AI as entering multiple parts of the research lifecycle — including interviewing, processing, analysis and reporting — while emphasizing rigor, transparency and trust. The practical implication is that AI risk depends on where the system enters the workflow and what authority it receives, not on the generic label “AI research.” See the AAPOR Task Force report announcement.

Human–AI Research Control Map

Use this map whenever AI materially touches a real customer-research round. Each stage has a human owner, a bounded AI job, a predictable failure mode, a verification step and an artifact that makes the handoff auditable.

Stage 1 — FRAME

Human owns

  • the Product decision;
  • the unknown blocking that decision;
  • whether an interview is the right method;
  • the participant definition;
  • what evidence would actually change the decision.

AI may help

  • critique the scope;
  • expose hidden assumptions;
  • generate alternative interpretations of the same known evidence;
  • ask what observation would make the team less confident.

Failure mode

A fluent model invents customer context that was never observed or turns a vague Product goal into a confident research question.

Verify

Can every factual premise in the research brief be traced to an actual Product or research source? Is the unknown genuinely unknown, rather than a generated conclusion presented as fact?

Audit artifact

A short research brief with Decision / Known evidence / Unknown / Method fit / Participant / Evidence needed.

Stage 2 — PREPARE

Human owns

  • the interview objective;
  • participant relevance;
  • the final discussion guide;
  • which questions are necessary to produce decision-relevant evidence.

AI may help

  • flag leading wording;
  • inspect double-barreled questions;
  • identify assumption-heavy phrasing;
  • suggest conditional probes;
  • simulate possible conversation branches;
  • challenge whether a question is likely to produce evidence for the stated unknown.

Failure mode

AI makes the guide sound generically neutral while removing the specificity needed to understand a real event, or it inserts the team's preferred explanation into a “helpful” probe.

Verify

For every important question, ask: What evidence is this trying to produce? What assumption does it smuggle in? What answer would challenge our preferred interpretation?

Audit artifact

A versioned guide with the research objective and material AI-assisted changes visible.

For the full question-selection and repair system, use Customer Interview Questions That Reveal Real Behavior. AI critique belongs here; the full question methodology belongs there.

Stage 3 — CONDUCT

There are three legitimate modes. Do not collapse them into one “AI interview” category.

Human-led, no live AI

The researcher conducts the session without AI influencing the conversation in real time. AI may still support preparation and post-session work.

This can be the better choice when rapport, emotional sensitivity, nonverbal/contextual cues, safeguarding, cultural nuance or unexpected disclosures matter materially.

Human-led + silent AI assistance

AI can track objectives, flag an uncovered topic or suggest a possible probe. The researcher decides whether the suggestion is relevant, neutral and worth interrupting the natural conversation for.

A live suggestion such as:

“Was security your biggest concern?”

is not automatically a good interview question. It presupposes both the mechanism and its importance. A human may reject it and instead ask:

“What were you considering at that moment?”

AI-moderated real participant

The AI directly asks questions to a real person. This is a different research design, not “the same interview but scalable.” It requires stronger piloting, disclosure decisions, interaction logging, quality review, failure handling and escalation.

Failure mode

Live AI narrows the conversation around its own hunch, over-pursues a topic, repeats itself, misses participant context or changes interviewer attention and rapport.

Verify

Review consequential probes and departures from the guide. If AI moderated the session, sample and inspect interaction quality rather than assuming the system behaved consistently because it used the same prompt.

Audit artifact

Mode used, guide/system-prompt version where material, interaction record, exceptions and escalation notes.

A 2026 peer-reviewed study of 74 chatbot-led interviews found that a generative-AI interviewer could support adaptive qualitative interviewing, while the authors still called for careful oversight, critical reflexivity and ongoing methodological development. That is evidence for evaluation, not a claim that AI moderation is equivalent to an expert human moderator. See Jack, Cooper & Flower (2026).

Stage 4 — CAPTURE

Human owns

  • what source material is retained;
  • which passages require source verification;
  • how transcript/translation uncertainty is handled;
  • which data may be processed by which system.

AI may help

  • transcribe;
  • diarize;
  • translate;
  • structure notes;
  • timestamp excerpts;
  • identify passages that need review.

Failure mode

Speaker attribution changes, a negation is missed, technical language is misheard, a number changes, an idiom is flattened, or translated wording sounds more definite than the source.

Verify

Use risk-based verification. For decision-relevant quotes, specialist terminology, consequential numbers, ambiguous passages and translation-sensitive claims, check the original audio/video or source-language material where feasible and material.

Audit artifact

Source recording/reference, transcript version, translation provenance and marked corrections.

Stage 5 — SYNTHESIZE

Human owns

  • whether a candidate finding is valid;
  • how participant context changes meaning;
  • how contradictions and minority cases affect interpretation;
  • the boundary between observation and inference.

AI may help

  • extract candidate observations;
  • cluster similar material;
  • compare participants;
  • search for contradictions;
  • propose themes;
  • draft alternative interpretations;
  • prepare a first-pass summary.

Failure mode

The model merges distinct segments because the wording is similar, suppresses a low-frequency severe case, invents connective detail, or produces a polished theme structure that anchors the human analyst.

Verify

Trace every decision-relevant candidate finding back to real sources. Inspect evidence for and against it. Review participants that do not fit. Check whether the model summarized, interpreted or generated a hypothesis rather than merely extracting material.

Audit artifact

The AI-Assisted Research Synthesis Ledger below.

Stage 6 — DECIDE

Human owns

  • what qualifies as Evidence, Inference, Assumption, Unknown or Decision;
  • the decision threshold;
  • the Product choice;
  • the uncertainty the team is willing to carry forward.

AI may help

  • challenge the reasoning;
  • propose alternative interpretations;
  • expose assumptions;
  • ask what evidence would reverse the decision;
  • draft a decision note after the human has classified the evidence.

Failure mode

The team mistakes a confident summary for research confidence, treats AI's preferred interpretation as authority, or lets the model turn qualitative evidence into roadmap priority.

Verify

Can the team defend the Product decision from source evidence and business context without referring to “what the AI thinks”? Are unsupported causal, prevalence or market-size claims still visible as unknowns?

Audit artifact

A decision record that explicitly separates Evidence → Inference → Assumption → Unknown → Decision.

AI Research Data Gate: decide before you upload or automate

This is a practical research gate, not legal advice. Use it before giving an AI system audio, video, transcripts, participant details or live access to a research session.

# AI Research Data Gate

Purpose
- What exact research task does the AI need to perform?
- Would a smaller input or a non-AI workflow achieve the same job?

Data
- What will the system receive: audio, video, transcript, name, employer,
  account details, free-text disclosure, sensitive information, internal Product context?

Participant expectation
- What has the participant been told about recording, transcription, translation,
  AI assistance, AI moderation and data processing?

Tool approval
- Is this service approved by the organization and, where relevant,
  security, privacy, legal, research-governance or ethics owners?

Provider behavior
- What do the CURRENT provider terms/policies say about retention, model training/reuse,
  subprocessors, human access, deletion, region/transfers and enterprise controls?

Data minimization
- Does the model need the whole transcript?
- Could selected excerpts, pseudonymized records or fewer fields do the job?

Re-identification
- Could the remaining context still identify the participant after names are removed?

Sensitivity
- Could this session contain sensitive personal information, confidential customer data,
  trade secrets, vulnerable-participant information or unexpected disclosures?

Escalation
- Does this study require specialist legal, privacy, security, ethics or safeguarding review?

Stop condition
- If the proposed tool cannot satisfy the required handling, DO NOT process the data with it.

Do not guess provider behavior from brand reputation. Vendor retention, training/reuse, access and transfer terms change. Check the service and account tier you are actually using.

Removing names is not the same thing as proving data is anonymous or safe. The EDPB has emphasized that AI-model anonymity needs case-by-case assessment, while ICO guidance emphasizes purpose, data minimization, retention and understanding how personal data moves through AI systems. These are useful control principles, not a substitute for the requirements that apply to your organization and study. See the EDPB opinion summary and ICO data-minimization guidance.

Participant disclosure: describe the session truthfully

Replace simplistic rules such as “always request recording consent” with a more accurate operating rule:

Tell participants truthfully how the session is being conducted and follow the recording, consent, privacy, research-ethics and data-processing requirements that apply to your jurisdiction, organization and study.

For each study, explicitly decide whether and how the participant must be told that AI will:

  • listen to or record the session;
  • transcribe or translate it;
  • suggest live questions;
  • moderate the interview;
  • analyze the material;
  • retain or otherwise process the data.

Do not disguise AI moderation as human moderation.

For EU-facing work, current transparency obligations under Article 50 of the EU AI Act apply from 2 August 2026 and can require people to be informed when they are directly interacting with certain AI systems, subject to the provision's scope and exceptions. Verify whether those rules apply to the actual system and study rather than copying generic compliance text. See the European Commission's Article 50 transparency guidelines.

Use AI to critique the guide — not define the research

“Bias-proof the script” is the wrong promise. No tool bias-proofs an interview.

A useful AI question-risk critique can look for:

  • leading wording;
  • double-barreled questions;
  • hypothetical future claims that could be replaced with relevant real experience;
  • hidden solution assumptions;
  • implied urgency or willingness to pay;
  • questions that collapse two research jobs;
  • missing concrete-incident probes;
  • language that became vague or generic during rewriting.

Then ask four questions yourself:

  1. What evidence is this question trying to produce?
  2. What assumption does the wording smuggle in?
  3. What answer would falsify or weaken the researcher's preferred interpretation?
  4. Did AI make the wording more neutral but less specific?

AI can homogenize language as easily as improve it. The goal is not “neutral-sounding text.” The goal is a guide capable of producing relevant evidence without revealing the answer the team wants.

When human moderation should remain the default

Do not treat AI moderation as the “mature” destination of every research program. Human moderation may be the more appropriate design when:

  • the topic is emotionally sensitive;
  • participant distress or safeguarding risk is material;
  • nonverbal or contextual signals materially affect meaning;
  • high-trust rapport is central to the research;
  • unexpected disclosures require judgment;
  • language or cultural nuance is central;
  • the population is vulnerable;
  • an AI failure could materially harm the participant or the quality of the research.

The decision is about research design and participant risk, not enthusiasm for or resistance to AI.

AI-Moderated Interview Quality Gate

If AI talks directly to real participants, define how you will inspect quality before you deploy it. Do not compress these dimensions into one fake “AI Moderator Quality Score.”

Coverage

Did the interaction cover the research objectives without turning the guide into a rigid checklist?

Probe relevance

Did follow-ups respond to what the participant actually said, or did the system chase its own hypothesis?

Neutrality

Did probes insert the expected answer, a preferred mechanism or emotionally loaded language?

Context sensitivity

Did the moderator notice when role, workflow, organization, culture or prior experience changed the meaning of an answer?

Repetition

Did it loop, ask already-answered questions or keep requesting detail after the participant had exhausted the topic?

Participant burden

Was the interaction frustrating, unnatural, over-persistent or unnecessarily long?

Safety and escalation

What happens when a participant shares sensitive information, becomes distressed, asks a question the model should not answer, or enters a situation the automated moderator cannot safely handle?

Evidence quality

Did the session produce concrete behavioral/contextual detail relevant to the research question, or only fluent but shallow responses?

Pilot these dimensions against the actual population and study. “The vendor says it scales” is not an evaluation plan.

Transcription is already an AI transformation

Do not treat a transcript as perfect ground truth simply because the words look clean.

Audio/video → transcript can fail through:

  • speaker attribution errors;
  • names and domain terminology;
  • accents and speech variation;
  • interruptions and incomplete utterances;
  • numbers;
  • negation;
  • loss of emphasis, hesitation or prosody.

You do not need to re-listen to every word by default. Use risk-based verification: check passages that materially affect a Product claim, quoted language, specialist term, number or ambiguous interpretation.

Translation is not lossless

AI translation can remove substantial manual work. It can also flatten idiom, emotional register, ambiguity, local concepts and culture-specific language.

When translated meaning materially affects a decision:

  • preserve the source-language text;
  • retain which tool/workflow produced the translation;
  • verify consequential passages with a capable human or bilingual reviewer when appropriate;
  • distinguish participant wording, translated wording and analyst interpretation.

A 2026 methodological study of Roman Urdu qualitative interviews found useful surface-level LLM coding alongside loss of cultural and emotional nuance, reinforcing the need for human interpretation in context-sensitive analysis. See the study in the Asian Journal of Psychiatry.

Preserve provenance after the interview

The central synthesis rule is:

Every decision-relevant AI finding must be traceable back to real source evidence.

That does not mean “attach one quote to every insight.” A credible finding may need several excerpts, participant context, an observation, contrary evidence and a clear interpretation boundary.

Weak:

“AI found that enterprise admins primarily fear security risks.”

Stronger:

“AI proposed security/access uncertainty as a candidate theme. Human review confirmed relevant source excerpts from these participant contexts and found two cases that did not fit the simple pattern.”

The model can propose the structure. The model must not become the source.

AI-Assisted Research Synthesis Ledger

Use this for any candidate finding that could materially affect a Product decision.

# AI-Assisted Research Synthesis Ledger

Candidate claim / theme
[What did the AI propose?]

AI operation
[Extraction / clustering / summary / comparison / contradiction search / other]

Source participant + context
[Which real research records does this come from? Include the context needed to interpret them.]

Supporting source evidence
[Exact source excerpts or references — not generated reconstructions.]

Contrary / disconfirming evidence
[What real evidence pushes against the simple interpretation?]

Minority / severe cases
[What low-frequency or segment-specific signal could aggregation suppress?]

Missing context
[What can neither the source nor the AI safely establish?]

AI transformation boundary
[Extracted from source / summarized / interpreted / generated hypothesis]

Human evidence status
[Evidence / Inference / Assumption / Unknown / Decision]

Decision impact
[What Product choice could this change?]

Next evidence
[What remains unresolved and what source/method could reduce it?]

The ledger exists to stop a polished model output from silently becoming a “research finding.” It forces the team to expose the transformation between source material and decision.

A generated quote is not a customer quote

This is a hard trust boundary:

A generated or reconstructed quote is not a customer quote.

AI may find an excerpt, extract it and return the source location. It must not rewrite, merge or paraphrase participant language and then put the result inside quotation marks as if the participant said it.

If your team cleans a quote for readability, label the editing according to your research standards and preserve the original source. A paraphrase can be useful; it simply is not verbatim participant language.

AI themes are candidate interpretations, not ground truth

Avoid:

“AI detected the top five customer themes.”

Prefer:

“AI proposed five candidate clusters from this input corpus.”

Then inspect:

  • examples inside each cluster;
  • cases near the cluster boundary;
  • whether distinct segments were mixed;
  • outliers;
  • contradictions;
  • whether similar wording caused two different mechanisms to be merged;
  • whether one mechanism was split because people used different words for it.

A cluster is an analytical convenience. It becomes decision-relevant only after a human inspects what it actually contains.

Theme counts do not establish market prevalence

AI can rapidly count mentions, tagged excerpts or cluster membership. That does not automatically tell you the percentage of customers who experience the problem.

Before interpreting a count, inspect:

  • what corpus was included;
  • how participants were recruited;
  • the unit being counted — mention, excerpt, session, participant or account;
  • whether duplicate participants or repeated mentions exist;
  • the tagging/inclusion rule;
  • whether every session gave the issue a comparable chance to surface;
  • whether the research design supports a prevalence claim at all.

If the remaining question is “how common is this?”, route it to the appropriate quantitative evidence rather than asking AI for a more confident summary. For interview-count and stopping logic, use How Many Customer Interviews Are Enough?. Do not calculate research confidence from the number of interviews.

Search explicitly for contradictions and minority signals

Summarization is compression. Compression can erase the exact signal that matters.

Ask AI to surface:

  • strongest evidence supporting the candidate finding;
  • strongest evidence against it;
  • participants who do not fit;
  • severe minority cases;
  • segment-specific differences;
  • unresolved cases.

Then inspect those source records manually. The model's self-audit is not enough.

A low-frequency issue can still matter because of severity, strategic relevance, contractual risk, safety or the value of the affected segment. Frequency is not priority.

Protect yourself from automation bias

Once a model produces a clean taxonomy and polished narrative, humans can anchor on it. NIST's Generative AI Profile explicitly treats human over-reliance and automation bias as AI risks and notes their interaction with confabulation, bias and homogenization. See NIST AI 600-1.

Practical countermeasures:

  • inspect raw evidence before accepting the final synthesis;
  • ask for a materially different clustering or alternative interpretation;
  • explicitly search for disconfirming evidence;
  • keep AI-draft and human-approved findings separate;
  • record meaningful analyst corrections;
  • for high-stakes decisions, consider comparing a human-first interpretation with the AI-assisted one rather than always anchoring the human on the model's first structure.

Record enough AI lineage to audit consequential changes

Do not create bureaucratic logging for trivial assistance. But when AI materially influences research, preserve enough provenance to understand the transformation.

Useful fields can include:

  • tool/model;
  • date;
  • material model/configuration version when available;
  • system/research prompt;
  • analysis prompt;
  • corpus included;
  • exclusions;
  • processing sequence;
  • human corrections.

A changed model, prompt, corpus, context window or input order can change an analysis. Do not promise perfect reproducibility. Preserve enough lineage to investigate consequential differences.

A synthetic user is not a customer interview

Synthetic-user output can be useful for:

  • rehearsing an interview;
  • pressure-testing a guide;
  • practicing probing;
  • brainstorming edge cases worth investigating;
  • generating hypotheses;
  • checking whether a concept explanation is superficially understandable;
  • piloting an extraction or analysis pipeline.

It is generated data, not direct evidence of:

  • lived customer experience;
  • actual workflow;
  • real unmet need;
  • true willingness to pay;
  • actual switching behavior;
  • market prevalence.

Do not:

  • add synthetic responses to the interview sample N;
  • quote them as customers;
  • label them “customer evidence”;
  • use them to claim interview saturation.

AI moderator vs synthetic respondent

AI moderatorSynthetic respondent
Who answers?A real participantThe model
Who/what is the evidence source?The participant, subject to normal research limitationsA model simulation
Primary riskThe AI changes how real evidence is elicitedGenerated output is mistaken for real-world evidence
Correct labelAI-moderated interview with a real participantSimulation / synthetic response

These are different epistemic situations. Do not mix them in the same evidence ledger without an explicit boundary.

Worked example: B2B onboarding interviews

This example is hypothetical. It shows the evidence trace, not a claim about a real company.

Product decision

Should the team change the setup experience for new workspace admins?

Unknown

Analytics shows that some admins delay connecting production data. The cause is unknown.

1. Human frames the research

The team records the known behavioral signal and defines the unknown without inventing a cause.

Evidence: setup delay exists in the relevant analytics view.

Unknown: why the delay happens and which mechanisms matter in which admin contexts.

2. AI critiques the draft guide

AI flags:

“Why is connecting data confusing?”

The question presupposes confusion. The researcher rewrites it:

“Walk me through what happened when you first reached the data connection step.”

AI improved the draft only because the human checked the assumption inside the suggestion.

3. Human conducts the interview with silent AI assist

The live assistant suggests:

“Was security a concern?”

The researcher rejects it as leading and asks:

“What were you considering at that moment?”

Lesson: AI suggestion ≠ interview question automatically.

4. AI transcribes the session

A transcript mishears the acronym SSO. Because the passage affects the interpretation of an access-control issue, the researcher checks the audio and corrects the source transcript.

5. AI proposes clusters

Candidate clusters:

  • security/access uncertainty;
  • unclear value;
  • missing permissions.

Human source review finds one severe enterprise case that the broad clustering made easy to overlook. The team preserves it as a minority/severity signal instead of deleting it for being “only one interview.”

6. AI drafts the synthesis

The draft says:

“Security is the primary reason admins delay setup.”

The team rejects that strength. The current qualitative corpus does not establish population prevalence or a single primary cause.

A bounded evidence statement becomes:

“Security/access uncertainty appeared repeatedly among interviewed enterprise-admin contexts and warrants targeted follow-up. The current qualitative evidence does not establish prevalence across all admins.”

7. Human makes the Product decision

The team does not “ship what AI recommended.” It:

  • updates a trust/security explanation prototype;
  • runs targeted follow-up with the relevant context;
  • checks the corresponding behavioral segment data;
  • keeps other mechanisms visible as unresolved.

The chain remains inspectable:

Source evidence → AI transformation → Human verification → Evidence status → Product decision.

When AI research goes wrong

Fluent fabrication

Failure: the model creates plausible supporting detail absent from the transcript.

Action: verify the source; remove the invented claim; inspect whether other outputs inherited it.

Quote mutation

Failure: AI rewrites participant words inside quotation marks.

Action: restore the exact source or label the text as a paraphrase.

Theme collapse

Failure: distinct segments with similar language are merged into one theme.

Action: restore participant/context metadata and inspect the mechanism inside each segment.

Minority erasure

Failure: a low-frequency but severe case disappears from the summary.

Action: inspect outliers and severity separately from frequency.

Leading live probe

Failure: AI suggests a question that contains the desired answer.

Action: reject or rewrite the probe before it reaches the participant.

Synthetic evidence contamination

Failure: generated participant responses enter the same sheet or repository as real interviews.

Action: separate them immediately and label them as simulation, not customer evidence.

Privacy leakage

Failure: a raw sensitive transcript is pasted into an unapproved AI service.

Action: stop processing and follow the organization's applicable privacy/security/incident procedure. Do not keep experimenting with the same data.

Cultural flattening

Failure: translation or summarization removes material nuance.

Action: inspect the source-language passage and involve appropriate language/cultural expertise when the interpretation matters.

Automation anchoring

Failure: the team accepts the first AI theme structure because it is coherent and presentation-ready.

Action: inspect raw sources, alternative interpretations and disconfirming evidence before approving the finding.

AI can assist the decision; it should not become the Product authority

End the workflow by classifying what you actually have.

Evidence

A participant described delaying setup until a security approver granted access.

AI extraction

“Security approval delayed setup.”

This is still checked against the source.

Inference

Security/access uncertainty may contribute to setup delay in some enterprise-admin contexts.

Assumption

Changing onboarding copy will remove the delay.

Unknown

How common the mechanism is and whether explanation alone changes behavior.

Decision

Test a clearer trust/security explanation with the relevant population and combine the qualitative evidence with the behavioral data needed for the next decision.

That is the control boundary: AI can generate candidate interpretations and Product hypotheses. It should not become the authority that chooses the roadmap.

If you need to decide whether you have enough interview coverage, use the sample-size and saturation decision system. If the next question requires combining interviews with analytics, surveys, support or another source, use Research Triangulation.

Practical next step

Before adding AI to a live research session, make the human-owned research job explicit. Build the discussion guide from the Product decision, unknown, participant and evidence need with the User Interview Script Generator. Then use the AI Product Management workflow guide to decide where AI assistance belongs across the broader Product workflow.

Turn the research plan into a human-owned guide

Build the interview guide before adding AI assistance

Use the research decision, evidence need, and participant context to build the actual discussion guide. AI can critique or assist the workflow, but the guide should still be anchored in the research job.

Portrait of Andrea Mezzadra, author of the blog post

Andrea Mezzadra@____Mezza____

Published on August 28, 2025 • Updated on September 17, 2026

Ex Product Director turned Independent Product Creator.