AI agent demos are easy. Reliable agent behavior in production is hard. The core Product mistake is treating "agent" as a product category that is inherently better than simpler mechanisms, and then measuring fluency instead of end-to-end task reliability.
This guide is for Product teams designing or shipping an agentic workflow. It answers one job:
Does this task actually need agentic autonomy? What exactly is delegated, which actions need approval, how do we evaluate end-to-end reliability and recovery, and what evidence justifies expanding, or reducing, autonomy?
The durable thesis: agentic Product design is delegation design. Reliability comes from a bounded job, explicit context and permissions, observable actions, risk-matched approval points, representative evals, failure recovery, and Product-outcome evidence, not from how autonomous or fluent the agent looks.
For the broader discipline, deciding whether AI belongs at all, building evals for any AI feature, connecting system quality to Product outcomes, economics, and iteration, use the AI Product Management hub. That hub owns the full AI Product Operating System; this page owns the agentic deep dive: delegation, autonomy, permissions, action traces, tool-use reliability, partial failure and recovery, agent-specific evals, and autonomy expansion or reduction. For harm scenarios, governance, and release risk, see the responsible-AI release guide.
The short answer
Before designing anything agentic, run this chain:
User job → simplest mechanism → delegated task → context and state → tools and permissions → plan and actions → confirmation and handback → end state → trace and eval → Product metric → autonomy decision
Most candidate "agent" workflows fail at the second step: a deterministic workflow or a one-shot assistant solves the job with less cost, latency, risk, and review burden. An agent must earn its complexity by doing something those mechanisms structurally cannot: adapt a multi-step plan, choose among tools, observe intermediate state, and recover or re-plan. If your workflow does not need that, do not build an agent.
There is no universal autonomy ladder to climb, no universal rollout cadence, and no universal success-rate threshold. Every one of those decisions depends on stakes, reversability, error detectability, permission scope, and economics of your specific task.
1. Agent Suitability Gate: does this job need an agent?
Compare four mechanisms before committing. The output is a decision with rationale and open uncertainties, never a readiness percentage.
| Dimension | Deterministic workflow | One-shot assistant / copilot | Agentic workflow | Human process |
|---|---|---|---|---|
| Best when | Rules are stable, state transitions known, exact validation matters, paths enumerable | One generation or transformation solves the job; the user stays the actor | The system must adapt a plan, choose tools, observe state, recover or re-plan across variable paths | Hidden context dominates, stakes are high, errors hard to detect |
| Tool use | Fixed calls in fixed order | None or one retrieval step | Autonomous multi-step tool selection | Human judgment |
| Failure shape | Predictable, testable | One bad output the user sees and discards | Wrong plan, wrong tool, wrong parameters, partial completion, loops, side effects | Slow, expensive, inconsistent |
| Cost profile | Cheapest to run and verify | Cheap; user absorbs review | Multiplied: model calls, retrieval, tool calls, retries, review | Human time |
Work through these suitability dimensions explicitly:
- Task variability: how different is each instance from the last?
- Adaptive planning need: does the path depend on what intermediate steps reveal?
- Tool and action need: must the system act on external state, or only produce text?
- Observability: can you see what happened step by step?
- Error detectability: will the user or reviewer realistically catch a wrong action?
- Reversibility: can a wrong action be undone?
- Stakes: what breaks if the agent is wrong?
- Hidden context: how much of "correct" depends on information the system cannot see?
- Permission scope: what is the blast radius of the tools involved?
- Latency tolerance: can the user wait for multi-step execution?
- Cost per successful task: including retries, retrieval, and human review.
- Human review burden: does reviewing the agent cost more than doing the work?
Decision rule: use the simplest mechanism that meets the quality bar. An agent is justified only when variability plus action plus recovery need exceeds what deterministic logic or a copilot can cover, and the economics survive review burden.
2. Delegation Contract: make the delegation inspectable
Every agentic workflow gets a written contract before launch. Copy and fill this in; it is the single artifact that makes delegation reviewable by PM, engineering, design, and risk stakeholders.
# Agent Delegation Contract
User job:
[...]
Why an agent instead of deterministic workflow / copilot:
[...]
Trigger (what starts the workflow):
[...]
Delegated task (exact scope):
[...]
Start state:
[...]
Expected end state:
[...]
Allowed context / data sources:
- [...]
Trusted vs untrusted inputs:
- [...]
Allowed tools:
- [...]
Allowed actions:
- [...]
Forbidden actions:
- [...]
Permission scope:
[...]
Actions requiring user approval:
- [...]
Abstain / stop conditions:
- [...]
Escalation / human handback:
[...]
Rollback / undo:
[...]
Success definition:
[...]
Failure states:
- [...]
Latency / cost constraints:
[...]
Trace / evidence retained:
[...]
Owner:
[...]
Version / change trigger:
[...]
Keep it to one page. If filling it in feels like bureaucracy, the scope is too broad, narrow the delegated task until the contract fits.
3. Autonomy is a design choice, not a maturity ladder
Do not present autonomy as levels every product should climb. A highly reliable agent can correctly require confirmation for an irreversible action forever; a low-stakes reversible action can warrant full autonomy with light evaluation. "More autonomy" is not "more mature."
Common modes to choose from, not stages to graduate through:
- Assist / retrieve: surface context, user acts.
- Draft: produce a draft the user edits and sends.
- Recommend: propose an action with rationale, user decides.
- Prepare action: stage everything short of execution for one-click confirmation.
- Execute after confirmation: run only the approved action.
- Execute bounded reversible actions autonomously: act freely inside a permission fence, escalate the rest.
- Escalate high-impact actions: autonomous for the routine, human for the consequential.
Choose per action class using these dimensions: stakes if wrong, reversibility, user ability to verify, permission scope, error detectability, frequency, time sensitivity, human review cost, and side effects.
4. Autonomy / Approval Matrix: map controls to stakes
For each action class in your workflow, fill one row. Teach the team to argue from stakes and reversibility, never from a universal permission table.
| Action class | Read | Draft | Recommend | Execute | Confirmation required? | Undo available? | On uncertainty | Escalation owner |
|---|---|---|---|---|---|---|---|---|
| Search / read data | Yes | , | , | Scoped reads | No | N/A | Stop and report | Agent owner |
| Modify draft | Yes | Yes | Yes | No | User sends | Version history | Narrow scope, ask | User |
| Send external communication | Yes | Yes | Yes | Only approved content | Yes, preview required | Recall window if any | Do not send | User / on-call |
| Publish content | Yes | Yes | Yes | Only to staging | Yes for public | Revert plan | Hold in staging | Content owner |
| Modify account state | Scoped | , | Yes | Bounded fields only | Yes | Compensating action | Roll back, escalate | Support / ops |
| Spend money / trigger billing | Scoped | , | Yes | Never autonomous | Yes, explicit amount | Refund path | Block and escalate | Finance owner |
| Delete data | Scoped | , | Yes | Never autonomous | Yes, explicit scope | Backup / restore | Block and escalate | Data owner |
| Change permissions | Scoped | , | Yes | Never autonomous | Yes | Revert change | Block and escalate | Security owner |
| Trigger external system action | Scoped | , | Yes | Allow-listed only | Depends on reversibility | Compensating action | Verify state first | Agent owner |
The pattern: irreversible, broad-scope, or hard-to-detect actions need confirmation, validation, and an undo or compensation path regardless of how "smart" the agent is. Reversible, narrow, easily verified actions can carry more autonomy with proportionally lighter evaluation.
5. Context, state, and source contract
Agents usually fail on context, not on reasoning. Before planning, the Product team must answer:
- What context is required before the agent may plan at all?
- Which source is authoritative when sources conflict?
- Which sources may contain untrusted instructions or content?
- How fresh must the context be, and what happens on stale data?
- What user and account state must be preserved across steps?
- What happens when context is missing or contradictory: stop, narrow, ask, or escalate?
- May the agent act on inferred state, or only on verified state?
- How does it communicate uncertainty to the user?
Separate trusted inputs (your database, your policy docs, verified user input) from untrusted inputs (external pages, pasted content, third-party API responses, retrieved documents that could contain instructions). Untrusted content must never silently widen permissions. (Conceptual grounding for context design is covered by the context-engineering references; this page uses it only to explain agent reliability.)
6. Tools are permissions, not just capabilities
Giving an agent a tool changes what can go wrong. For each tool and action, inspect:
- Read vs write, and exact scope.
- Argument validation: what counts as a legal call?
- Authentication and authorization: whose credentials, what limits?
- Side effects and whether they are idempotent.
- Reversibility and the compensation path.
- Rate limits, timeouts, and error shapes.
- What data comes back, and whether it may be untrusted.
- What a retry means: safe to repeat, or double-charge and double-send risk?
Product principle: every tool added is a new failure class accepted. Prefer least-necessary permissions, and route deep implementation security to the responsible-AI release guide rather than solving it in the Product spec.
7. Planning quality is not task success
Never grade an agent on reasoning text, "good plan," or fluent final output. An agent can plan well and call the wrong tool, pick the right tool with the wrong customer ID, complete the action while creating a harmful side effect, partially complete without reporting it, or write a perfect summary of a backend state it got wrong.
The evaluation target is always the end-to-end task and the resulting state transition, what changed in the world, not what the agent said about it.
8. Agent Evaluation Stack
No universal score, no universal pass threshold. Thresholds are tied to task stakes and to the decision the eval must unlock (ship, narrow, add approval, stop).
Layer 1, Deterministic checks. Where ground truth exists, assert it: valid schema, required end state reached, correct tool called, forbidden action absent, expected record created or updated, no duplicate side effect.
Layer 2, Task-quality rubric. Human or domain judgment where deterministic ground truth is insufficient. Write the rubric before running cases, not after seeing outputs.
Layer 3, Model-assisted review. A model judge may pre-screen traces, but only for criteria it has been validated on, with known limitations stated. A model judge is never ground truth.
Layer 4, Trace-level evaluation. Inspect the action sequence: tool selection, parameters, retries, context actually used, approvals obtained, and failure handling. A correct final answer can hide an incorrect action trace.
Layer 5, Product and online outcome. Successful end state, user correction burden, escalation rate, abandonment, time to outcome, repeat use and trust where relevant, and the actual workflow outcome.
Layer 6, Operational and economic. End-to-end latency, tool and model retries, cost per successful task, human review time, and failure-support burden.
Offline evals give repeatability, regression detection, and configuration comparison. Online Product measurement gives usefulness, trust, workflow impact, and economics. You need both, and they answer different questions.
9. Representative eval sets
Happy paths alone prove nothing. Cover the real task distribution plus important failures:
- Normal cases and common variants.
- Edge cases and ambiguous instructions.
- Missing, contradictory, or stale context.
- Tool or API errors, permission denials, timeouts.
- Partial completion and duplicate requests.
- High-stakes must-not-act cases.
- Malicious or untrusted input where the agent consumes external content.
- Recovery cases: after the failure, does the agent detect, contain, recover, and report?
There is no magic number of cases. Coverage should reflect how often each situation occurs plus how badly it hurts when it does.
10. Tool-use evals
Evaluate explicitly whether a tool should have been used at all, whether the correct tool was selected, whether arguments were correct and necessary, whether the result was interpreted correctly, whether the follow-up action matched the result, whether side effects were acceptable, and whether retries were safe. A failed first call followed by correct recovery can still be acceptable depending on cost and stakes, but only if the eval looks at the trace, not just the final message.
11. Failure & Recovery Matrix
Reliability is a Product system: for each failure class, define detection, prevention, containment, recovery, user notification, retry-vs-stop-vs-escalate policy, and whether it becomes a regression case.
| Failure class | How detected | Prevent | Contain | Recover | Notify user? | Retry, stop, abstain, or escalate? | Regression case? |
|---|---|---|---|---|---|---|---|
| Wrong task interpretation | Trace review, user correction | Narrower trigger, better contract | Limit action scope | Re-run with narrowed scope | Yes if action taken | Stop, clarify | Yes |
| Wrong plan | Trace review, eval cases | Representative plan coverage | Approval checkpoints | Re-plan from verified state | If visible | Retry once, then escalate | Yes |
| Wrong tool selected | Deterministic tool assertion | Tighter tool descriptions, fewer tools | Permission fence | Undo, redo with correct tool | If side effect occurred | Stop, escalate | Yes |
| Wrong tool parameters (e.g. wrong customer ID) | Entity validation, state check | Parameter validation, confirmation preview | Dry-run / preview | Compensating action | Yes | Stop, never blind-retry writes | Yes |
| Missing context | Pre-plan context gate | Required-context checklist | Refuse to plan | Ask, narrow, or hand back | Yes | Abstain | If recurring |
| Stale context | Freshness check, version stamp | TTL and re-fetch policy | Re-read before acting | Re-run on fresh state | If decision used stale data | Retry with fresh context | If recurring |
| Unsupported assumption | Claim-to-source check | Evidence-link requirement | Draft-only mode | Correct and re-verify | Yes | Stop, ask | Yes |
| Hallucinated action or result | State verification after action | Verify-then-report discipline | Read-back of system state | Compensate, correct record | Yes | Stop, escalate | Yes |
| Permission failure | Auth error signal | Least-privilege review | Graceful denial path | Hand back with clear next step | Yes | Hand back, do not retry-privilege-escalate | If recurring |
| Downstream API error | Error signal, timeout | Timeout and fallback design | Bounded retry with backoff | Resume from verified state | If user-visible | Retry by failure type, then escalate | If recurring |
| Timeout after action may have happened | State read-back | Idempotency keys where possible | Verify-before-repeat | Reconcile, never assume | Yes | Verify state first, then decide | Yes |
| Partial completion presented as success | End-state assertion | Explicit end-state definition | Per-step status | Complete, roll back, or hand back remainder | Yes, with what remains | Escalate remainder | Yes |
| Duplicate action | Idempotency check, audit log | Idempotent design | Dedupe guard | Compensate duplicate | Yes | Stop | Yes |
| Loop / stuck state | Repeated-equivalent-action detection | Bounded attempts, stop conditions | Kill switch, attempt cap | Escalate with trace | Yes | Stop and escalate | Yes |
| Unsafe or unintended action | Permission fence, preview | Narrower scope, approval | Revoke, revert | Compensate, incident path | Yes | Stop and escalate | Yes |
| Silent failure | End-state and heartbeat checks | Mandatory outcome reporting | Alert on missing outcome | Investigate from trace | Yes | Escalate | Yes |
| Bad human handoff | Handoff quality rubric | Structured handoff format | Keep resumable state | Re-handoff with full context | Yes | Hand back better | Yes |
| Failed rollback | Rollback verification | Tested compensation paths | Incident path | Manual remediation | Yes | Escalate | Yes |
12. Partial success must be explicit
Multi-step tasks do not collapse into success or failure. Track outcome states such as: not started, blocked before action, partially completed, completed with recoverable warning, completed correctly, completed with unacceptable side effect, rolled back, and escalated. Define what each state means for your workflow, who sees it, and what happens next.
13. Human review is not automatically safety
"Human in the loop" with no operational definition is theater. For every approval point, specify: what exactly the human reviews, at what point in the flow, with what evidence shown, whether the human can realistically detect the error class in question, what happens on approve, edit, reject, or no response, and whether review time destroys the economic value.
Watch for rubber-stamping: if approvers accept nearly everything without reading, the checkpoint provides legal cover, not safety. Product mechanisms that help: preview diffs, source-evidence links, structured action summaries, explicit irreversible-action warnings, undo windows, and approval history. But none of them guarantee safety, measure whether review actually catches seeded errors, and track review burden as a cost.
14. Abstention, escalation, and handback
A strong agent knows when not to act. Define when it must ask for missing input, narrow scope, abstain, escalate, hand back to the user or operator, stop after tool failure, or request approval. Then measure handback quality: does it explain what was completed, what remains, why it stopped, what the user must do next, and what state is safe to resume from? A clean handback on a task the agent should not attempt is a success, not a failure. This handoff UX is core Product work.
15. Loops, retries, and stop conditions
Agent loops multiply latency, cost, side effects, user wait, and rate-limit risk. Require bounded attempts, explicit stop conditions, idempotent actions where possible, retry policy by failure type (a read timeout and a failed payment deserve different policies), detection of repeated equivalent actions, and escalation when progress stalls. Do not publish a universal max-step number, set bounds from the task's cost, stakes, and reversibility.
16. Agent Trace Review
Use this artifact in eval sessions and incident reviews. It focuses on observable actions, state, and results, never on hidden chain-of-thought, which evaluators should not require.
# Agent Trace Review
Task / user job:
[...]
Expected end state:
[...]
Context available:
- [...]
Tools / permissions available:
- [...]
User approvals required:
- [...]
Observed action trace:
1. [...]
2. [...]
Tool failures / retries:
- [...]
Unexpected side effects:
- [...]
Final system state:
[...]
User-visible result:
[...]
Task outcome:
[Complete / Partial / Escalated / Rolled back / Failed]
Correctness / quality issues:
- [...]
Permission / control issues:
- [...]
Recovery quality:
[...]
Latency / cost / review burden:
[...]
Root failure class:
[...]
Add to regression set?:
[Yes / No + why]
Product decision:
[...]
17. Observability: answer Product questions, not just logging
A PM should be able to answer: what did the agent attempt, which tools and actions ran, what state changed, what failed and was retried, what the user approved, what was escalated, how long it took, what it cost, and whether the result could be undone. Distinguish three views: the engineering trace, the operator and debug view, and the user-visible activity history. Respect privacy and security boundaries, internal reasoning internals are not user content.
18. Cost per successful task
Never optimize token cost alone. Count model calls, context and retrieval, tool calls, retries, fallback paths, long-running state, human review time, support burden, and failure remediation. A cheaper run with lower task success is often more expensive per successful outcome. No universal cost threshold exists; compare against the baseline the agent replaces.
19. Latency and waiting UX
Agentic workflows can take far longer than chat generation. Design for it: progress visibility, background execution, resumability, notification on completion, partial results, cancellation, and timeout and failure UX. A spinner is not a delegation experience. Keep this at Product-pattern depth, async-job engineering belongs elsewhere.
20. Change management and regression
Agent behavior can shift materially when any of these change: model, provider, or version; system instructions; prompts or templates; context or retrieval corpus; tool definitions; external APIs; routing; memory or state logic; permissions; safety policy; structured-output schema; or fallback logic. Treat material behavior changes as Product changes: keep regression cases, compare against baseline, maintain version awareness, and monitor after rollout. Match rigor to risk and behavior impact, not every copy tweak needs certification.
21. Rollout: evidence-based, not ritual-based
There is no universal four-stage rollout and no "promote after two stable weeks" rule. Rollout scope follows risk and evidence: stakes, permission scope, reversibility, error detectability, eval coverage, observed production failures, user ability to recover, support readiness, economics, and monitoring. Possible scopes include internal only, sandbox or read-only, opt-in beta, low-risk actions only, approval-required production, bounded autonomy, and expanded population or action scope. Expansion is a Product decision gated on evidence, not a maturity ritual over time.
22. Autonomy Expansion Contract
Before granting more autonomy, answer in writing:
- Which new action or permission is being added?
- Which failure becomes newly possible?
- Can that failure be detected, and is it reversible?
- Do eval cases cover it, including failure and recovery?
- Is production evidence representative of the expanded scope?
- Does human review remain effective at the new scope?
- Does Product value improve enough to justify added risk and cost?
Moving from "draft" to "send" is a major Product change even when the model is unchanged. The same contract in reverse governs reducing or removing autonomy: state which failure, cost, or correction burden triggered the narrowing.
23. Security and prompt-injection boundary
Agents that consume external or untrusted content and can take actions face instruction-injection and tool-abuse risks. At Product depth: separate trusted from untrusted sources, grant least-necessary permissions, require approval for sensitive actions, validate actions before execution, bound secret and data access, include malicious-input test cases, and define a monitoring and incident path. This page stays at Product implications; deeper security engineering and responsible-release work belongs with the responsible-AI release guide.
24. Worked cases
Five hypothetical cases show how the system above produces different verdicts. All scenarios are synthetic illustrations, not benchmark results.
Case A, Support triage: category assignment only
Job: assign each incoming ticket a category and priority. Tempting agent candidate, wrong architecture. A deterministic rules layer plus a classifier, or one-shot AI categorization, covers it with lower latency, cost, and failure surface. Verdict: do not use an agent. Agentic planning adds nothing when the output is a single label.
Case B, Support workflow: account lookup, policy search, ticket action
Job: route the ticket and prepare the next action, which requires inspecting customer and account state, checking known incidents, searching policy docs, deciding whether data is sufficient, drafting a reply, and creating or escalating an issue, including choosing whether to send or request approval. Variability, multi-step tool use, and conditional planning justify an agent. The Delegation Contract defines read vs write tools, approval before any customer-visible send or state change, abstention on missing account context, and handback with completed-vs-remaining state. Eval cases cover stale account data, contradictory policy sources, permission denials, and duplicate tickets. Verdict: bounded agent with approval-gated actions.
Case C, Release notes from merged PRs: the overkill detector
Job: draft release notes from merged PRs. Retrieval plus deterministic PR selection plus one-shot generation is sufficient; autonomous multi-step planning adds little value and new failure modes. Verdict: do not use an agent. This is the key information-gain example: agentic architecture must earn its complexity, and here it does not.
Case D, Customer summary with traceability
Job: produce an account brief from CRM data, recent tickets, usage summaries, and meeting notes. Risks dominate: stale state, cross-account leakage, unsupported claims, incorrect aggregation, source conflict. The source contract names the authoritative system per field, requires evidence links for every non-trivial claim, grants read-only permissions, and hands back on missing context. Eval dimensions include leakage probes (never with real customer data), conflict cases, and correction burden, because a brief the user must fully re-verify saves nothing. Privacy and risk depth routes to the responsible-AI release guide. Verdict: read-only bounded agent, draft mode, human sends.
Case E, External action and B2B remediation
Job: investigate a failed customer data sync and take bounded remediation action across systems, potentially sending email, modifying account state, or triggering a workflow. Output quality is insufficient as a metric; what matters is action preview, permission scope, confirmation logic, validation, audit history, undo or compensation, and escalation. Long-running execution needs progress visibility, resumability, and stall detection; cost and time accounting decides whether the autonomy pays. Verdict: bounded autonomy for reversible steps, confirmation for everything else, full trace review before any scope expansion.
25. Product metrics: outcome, not smartness
Fluency is not success. Track four layers, with no universal targets and no single blended "agent score":
- Task and reliability: eligible attempts, successful end states, partial completions, escalations, rollbacks, unacceptable side effects, recovery success.
- User and workflow: correction burden, time to outcome, abandonment, successful handoffs, trust and continued use where relevant.
- Operational and economic: latency, retries, cost per successful task, human review time, support incidents.
- Product and business: the outcome the workflow exists to move, resolution time, sync recovery rate, brief acceptance, defined per workflow.
Denominators matter
A completion rate is meaningless without its denominator. Distinguish eligible tasks, tasks delegated, tasks attempted, tasks completed, tasks completed without correction, and tasks requiring escalation. Define exclusions before analysis, never filter out difficult cases after the fact. (Measurement-implementation depth belongs to the analytics instrumentation guide; this page owns what agent reliability must measure.)
Severity beats averages
A 95% success rate is unacceptable when the 5% deletes data; a lower rate can be fine for low-stakes drafts with easy detection and correction. Preserve failure categories, inspect performance by important slice, weight decisions by consequence rather than aggregate percentage, and never invent weighted-risk formulas that fake precision. Deeper harm analysis routes to the responsible-AI release guide.
26. Common failure modes
- Agent by default: multi-step autonomy chosen where a deterministic workflow or copilot would win.
- Fluency equals quality: good conversation hides wrong system state.
- Permission sprawl: broad write access granted for convenience.
- Human-in-the-loop theater: approval exists but no human can effectively verify.
- Happy-path evals: no tool failure, ambiguity, or edge cases.
- Final-answer-only eval: actions and side effects ignored.
- Silent partial completion: the agent stops halfway but presents success.
- Retry theater: retries multiply side effects and cost.
- No handoff contract: the user cannot understand or resume a failed task.
- Aggregate success theater: rare severe failures hidden by a high average.
- Token-cost optimization: retries, review, and support ignored.
- Autonomy as status: more autonomy treated as inherently better.
- No regression contract: prompt, model, or tool changes alter behavior silently.
27. FAQ
What is an AI agent in Product Management? A system delegated a bounded multi-step task: it plans across steps, uses tools that read or change state, observes results, and recovers or escalates. The Product job is defining that delegation boundary, not the underlying model.
What is the difference between an agent and an automated workflow? A workflow executes fixed steps in a fixed order; an agent adapts its plan, chooses among tools, and responds to intermediate state. If your paths are enumerable and validation is exact, a workflow wins on cost, latency, and testability.
When should a Product use an AI agent instead of a chatbot or copilot? When the job structurally needs multi-step planning, tool use across variable paths, and recovery, and a one-shot generation with the user as actor cannot cover it. Cases A and C above show where the answer is "do not use an agent."
How do you evaluate an AI agent? End to end: deterministic checks on state and actions, a task-quality rubric, validated model-assisted review at most as a screen, trace-level inspection of tool choice and recovery, online Product outcomes, and operational economics. Judge the state transition, not the final text.
What metrics should an AI agent use? Task reliability (success, partial, escalation, rollback, side effects), user burden (correction, time to outcome, abandonment), economics (latency, retries, cost per successful task, review time), and the workflow's own Product outcome. No universal thresholds.
How much autonomy should an AI agent have? As much as stakes, reversibility, detectability, and economics justify per action class, set in the Autonomy / Approval Matrix. There is no ladder to climb; "draft forever" can be the correct end state.
What should require human approval? Irreversible, broad-scope, or hard-to-verify actions: external sends, publishing, state changes, spending, deletion, permission changes. Approval must include inspectable evidence and a real ability to detect the error.
How do you handle agent tool failures? By failure type: verify state before repeating any write, use idempotency where possible, bound retries with backoff, and escalate with the trace when progress stalls. Timeouts after a possible write are reconcile-first, never blind-retry.
How do you stop an agent from looping? Bounded attempts, explicit stop conditions, detection of repeated equivalent actions, and escalation on stall, sized from the task's cost and stakes, not from a universal step limit.
What is agent observability? The ability to answer what was attempted, which tools ran, what state changed, what failed and was retried, what was approved or escalated, what it cost and took, and whether it can be undone, across engineering, operator, and user-visible views.
How do Product Managers monitor agent cost and latency? Per successful task, including retries, retrieval, tool calls, and human review, compared against the baseline the agent replaces, with waiting UX (progress, backgrounding, resumability, cancellation) designed as part of the Product.
When should an agent hand back to a human? On missing or contradictory context, permission denial, tool failure it cannot recover from, high-stakes uncertainty, or any stop condition, with completed work, remaining work, reason, next action, and resumable state stated.
28. What to do next
- If you have not defined the AI task and baseline yet, work through the AI Product Management hub first.
- If risk, safety, or release governance dominates, use the responsible-AI release guide.
- If the question is which AI tools help you do PM work faster, see the AI PM tools guide, this page is about building agentic products, not tool selection.
- If trace and Product metrics need implementation, follow the analytics instrumentation guide.
- If you are evaluating an assistant-style agent for PM workflows, the Molbot field test shows what draft-first usage realistically delivers.
- When an agentic intervention is ready for causal testing, run it through the A/B test plan generator.
Last materially reviewed: October 3, 2026. Agent tooling moves quickly; principles above are durable, while provider-specific behaviors should be verified against current official documentation (including Anthropic's agent-building and agent-eval guidance, OpenAI's agent evaluation guidance, the NIST AI Risk Management Framework and Generative AI Profile, and OWASP guidance on agentic security) at implementation time.
