AI Agents in Product Management: Delegation, Evals & Production Reliability

Updated on

For PMs shipping agentic workflows: suitability gate, delegation contract, autonomy and approval design, end-to-end evals, failure recovery, and evidence-based autonomy decisions.

Share:

AI agent demos are easy. Reliable agent behavior in production is hard. The core Product mistake is treating "agent" as a product category that is inherently better than simpler mechanisms, and then measuring fluency instead of end-to-end task reliability.

This guide is for Product teams designing or shipping an agentic workflow. It answers one job:

Does this task actually need agentic autonomy? What exactly is delegated, which actions need approval, how do we evaluate end-to-end reliability and recovery, and what evidence justifies expanding, or reducing, autonomy?

The durable thesis: agentic Product design is delegation design. Reliability comes from a bounded job, explicit context and permissions, observable actions, risk-matched approval points, representative evals, failure recovery, and Product-outcome evidence, not from how autonomous or fluent the agent looks.

For the broader discipline, deciding whether AI belongs at all, building evals for any AI feature, connecting system quality to Product outcomes, economics, and iteration, use the AI Product Management hub. That hub owns the full AI Product Operating System; this page owns the agentic deep dive: delegation, autonomy, permissions, action traces, tool-use reliability, partial failure and recovery, agent-specific evals, and autonomy expansion or reduction. For harm scenarios, governance, and release risk, see the responsible-AI release guide.

The short answer

Before designing anything agentic, run this chain:

User job → simplest mechanism → delegated task → context and state → tools and permissions → plan and actions → confirmation and handback → end state → trace and eval → Product metric → autonomy decision

Most candidate "agent" workflows fail at the second step: a deterministic workflow or a one-shot assistant solves the job with less cost, latency, risk, and review burden. An agent must earn its complexity by doing something those mechanisms structurally cannot: adapt a multi-step plan, choose among tools, observe intermediate state, and recover or re-plan. If your workflow does not need that, do not build an agent.

There is no universal autonomy ladder to climb, no universal rollout cadence, and no universal success-rate threshold. Every one of those decisions depends on stakes, reversability, error detectability, permission scope, and economics of your specific task.

1. Agent Suitability Gate: does this job need an agent?

Compare four mechanisms before committing. The output is a decision with rationale and open uncertainties, never a readiness percentage.

DimensionDeterministic workflowOne-shot assistant / copilotAgentic workflowHuman process
Best whenRules are stable, state transitions known, exact validation matters, paths enumerableOne generation or transformation solves the job; the user stays the actorThe system must adapt a plan, choose tools, observe state, recover or re-plan across variable pathsHidden context dominates, stakes are high, errors hard to detect
Tool useFixed calls in fixed orderNone or one retrieval stepAutonomous multi-step tool selectionHuman judgment
Failure shapePredictable, testableOne bad output the user sees and discardsWrong plan, wrong tool, wrong parameters, partial completion, loops, side effectsSlow, expensive, inconsistent
Cost profileCheapest to run and verifyCheap; user absorbs reviewMultiplied: model calls, retrieval, tool calls, retries, reviewHuman time

Work through these suitability dimensions explicitly:

  • Task variability: how different is each instance from the last?
  • Adaptive planning need: does the path depend on what intermediate steps reveal?
  • Tool and action need: must the system act on external state, or only produce text?
  • Observability: can you see what happened step by step?
  • Error detectability: will the user or reviewer realistically catch a wrong action?
  • Reversibility: can a wrong action be undone?
  • Stakes: what breaks if the agent is wrong?
  • Hidden context: how much of "correct" depends on information the system cannot see?
  • Permission scope: what is the blast radius of the tools involved?
  • Latency tolerance: can the user wait for multi-step execution?
  • Cost per successful task: including retries, retrieval, and human review.
  • Human review burden: does reviewing the agent cost more than doing the work?

Decision rule: use the simplest mechanism that meets the quality bar. An agent is justified only when variability plus action plus recovery need exceeds what deterministic logic or a copilot can cover, and the economics survive review burden.

2. Delegation Contract: make the delegation inspectable

Every agentic workflow gets a written contract before launch. Copy and fill this in; it is the single artifact that makes delegation reviewable by PM, engineering, design, and risk stakeholders.

# Agent Delegation Contract

User job:
[...]

Why an agent instead of deterministic workflow / copilot:
[...]

Trigger (what starts the workflow):
[...]

Delegated task (exact scope):
[...]

Start state:
[...]

Expected end state:
[...]

Allowed context / data sources:
- [...]

Trusted vs untrusted inputs:
- [...]

Allowed tools:
- [...]

Allowed actions:
- [...]

Forbidden actions:
- [...]

Permission scope:
[...]

Actions requiring user approval:
- [...]

Abstain / stop conditions:
- [...]

Escalation / human handback:
[...]

Rollback / undo:
[...]

Success definition:
[...]

Failure states:
- [...]

Latency / cost constraints:
[...]

Trace / evidence retained:
[...]

Owner:
[...]

Version / change trigger:
[...]

Keep it to one page. If filling it in feels like bureaucracy, the scope is too broad, narrow the delegated task until the contract fits.

3. Autonomy is a design choice, not a maturity ladder

Do not present autonomy as levels every product should climb. A highly reliable agent can correctly require confirmation for an irreversible action forever; a low-stakes reversible action can warrant full autonomy with light evaluation. "More autonomy" is not "more mature."

Common modes to choose from, not stages to graduate through:

  • Assist / retrieve: surface context, user acts.
  • Draft: produce a draft the user edits and sends.
  • Recommend: propose an action with rationale, user decides.
  • Prepare action: stage everything short of execution for one-click confirmation.
  • Execute after confirmation: run only the approved action.
  • Execute bounded reversible actions autonomously: act freely inside a permission fence, escalate the rest.
  • Escalate high-impact actions: autonomous for the routine, human for the consequential.

Choose per action class using these dimensions: stakes if wrong, reversibility, user ability to verify, permission scope, error detectability, frequency, time sensitivity, human review cost, and side effects.

4. Autonomy / Approval Matrix: map controls to stakes

For each action class in your workflow, fill one row. Teach the team to argue from stakes and reversibility, never from a universal permission table.

Action classReadDraftRecommendExecuteConfirmation required?Undo available?On uncertaintyEscalation owner
Search / read dataYes,,Scoped readsNoN/AStop and reportAgent owner
Modify draftYesYesYesNoUser sendsVersion historyNarrow scope, askUser
Send external communicationYesYesYesOnly approved contentYes, preview requiredRecall window if anyDo not sendUser / on-call
Publish contentYesYesYesOnly to stagingYes for publicRevert planHold in stagingContent owner
Modify account stateScoped,YesBounded fields onlyYesCompensating actionRoll back, escalateSupport / ops
Spend money / trigger billingScoped,YesNever autonomousYes, explicit amountRefund pathBlock and escalateFinance owner
Delete dataScoped,YesNever autonomousYes, explicit scopeBackup / restoreBlock and escalateData owner
Change permissionsScoped,YesNever autonomousYesRevert changeBlock and escalateSecurity owner
Trigger external system actionScoped,YesAllow-listed onlyDepends on reversibilityCompensating actionVerify state firstAgent owner

The pattern: irreversible, broad-scope, or hard-to-detect actions need confirmation, validation, and an undo or compensation path regardless of how "smart" the agent is. Reversible, narrow, easily verified actions can carry more autonomy with proportionally lighter evaluation.

5. Context, state, and source contract

Agents usually fail on context, not on reasoning. Before planning, the Product team must answer:

  • What context is required before the agent may plan at all?
  • Which source is authoritative when sources conflict?
  • Which sources may contain untrusted instructions or content?
  • How fresh must the context be, and what happens on stale data?
  • What user and account state must be preserved across steps?
  • What happens when context is missing or contradictory: stop, narrow, ask, or escalate?
  • May the agent act on inferred state, or only on verified state?
  • How does it communicate uncertainty to the user?

Separate trusted inputs (your database, your policy docs, verified user input) from untrusted inputs (external pages, pasted content, third-party API responses, retrieved documents that could contain instructions). Untrusted content must never silently widen permissions. (Conceptual grounding for context design is covered by the context-engineering references; this page uses it only to explain agent reliability.)

6. Tools are permissions, not just capabilities

Giving an agent a tool changes what can go wrong. For each tool and action, inspect:

  • Read vs write, and exact scope.
  • Argument validation: what counts as a legal call?
  • Authentication and authorization: whose credentials, what limits?
  • Side effects and whether they are idempotent.
  • Reversibility and the compensation path.
  • Rate limits, timeouts, and error shapes.
  • What data comes back, and whether it may be untrusted.
  • What a retry means: safe to repeat, or double-charge and double-send risk?

Product principle: every tool added is a new failure class accepted. Prefer least-necessary permissions, and route deep implementation security to the responsible-AI release guide rather than solving it in the Product spec.

7. Planning quality is not task success

Never grade an agent on reasoning text, "good plan," or fluent final output. An agent can plan well and call the wrong tool, pick the right tool with the wrong customer ID, complete the action while creating a harmful side effect, partially complete without reporting it, or write a perfect summary of a backend state it got wrong.

The evaluation target is always the end-to-end task and the resulting state transition, what changed in the world, not what the agent said about it.

8. Agent Evaluation Stack

No universal score, no universal pass threshold. Thresholds are tied to task stakes and to the decision the eval must unlock (ship, narrow, add approval, stop).

Layer 1, Deterministic checks. Where ground truth exists, assert it: valid schema, required end state reached, correct tool called, forbidden action absent, expected record created or updated, no duplicate side effect.

Layer 2, Task-quality rubric. Human or domain judgment where deterministic ground truth is insufficient. Write the rubric before running cases, not after seeing outputs.

Layer 3, Model-assisted review. A model judge may pre-screen traces, but only for criteria it has been validated on, with known limitations stated. A model judge is never ground truth.

Layer 4, Trace-level evaluation. Inspect the action sequence: tool selection, parameters, retries, context actually used, approvals obtained, and failure handling. A correct final answer can hide an incorrect action trace.

Layer 5, Product and online outcome. Successful end state, user correction burden, escalation rate, abandonment, time to outcome, repeat use and trust where relevant, and the actual workflow outcome.

Layer 6, Operational and economic. End-to-end latency, tool and model retries, cost per successful task, human review time, and failure-support burden.

Offline evals give repeatability, regression detection, and configuration comparison. Online Product measurement gives usefulness, trust, workflow impact, and economics. You need both, and they answer different questions.

9. Representative eval sets

Happy paths alone prove nothing. Cover the real task distribution plus important failures:

  • Normal cases and common variants.
  • Edge cases and ambiguous instructions.
  • Missing, contradictory, or stale context.
  • Tool or API errors, permission denials, timeouts.
  • Partial completion and duplicate requests.
  • High-stakes must-not-act cases.
  • Malicious or untrusted input where the agent consumes external content.
  • Recovery cases: after the failure, does the agent detect, contain, recover, and report?

There is no magic number of cases. Coverage should reflect how often each situation occurs plus how badly it hurts when it does.

10. Tool-use evals

Evaluate explicitly whether a tool should have been used at all, whether the correct tool was selected, whether arguments were correct and necessary, whether the result was interpreted correctly, whether the follow-up action matched the result, whether side effects were acceptable, and whether retries were safe. A failed first call followed by correct recovery can still be acceptable depending on cost and stakes, but only if the eval looks at the trace, not just the final message.

11. Failure & Recovery Matrix

Reliability is a Product system: for each failure class, define detection, prevention, containment, recovery, user notification, retry-vs-stop-vs-escalate policy, and whether it becomes a regression case.

Failure classHow detectedPreventContainRecoverNotify user?Retry, stop, abstain, or escalate?Regression case?
Wrong task interpretationTrace review, user correctionNarrower trigger, better contractLimit action scopeRe-run with narrowed scopeYes if action takenStop, clarifyYes
Wrong planTrace review, eval casesRepresentative plan coverageApproval checkpointsRe-plan from verified stateIf visibleRetry once, then escalateYes
Wrong tool selectedDeterministic tool assertionTighter tool descriptions, fewer toolsPermission fenceUndo, redo with correct toolIf side effect occurredStop, escalateYes
Wrong tool parameters (e.g. wrong customer ID)Entity validation, state checkParameter validation, confirmation previewDry-run / previewCompensating actionYesStop, never blind-retry writesYes
Missing contextPre-plan context gateRequired-context checklistRefuse to planAsk, narrow, or hand backYesAbstainIf recurring
Stale contextFreshness check, version stampTTL and re-fetch policyRe-read before actingRe-run on fresh stateIf decision used stale dataRetry with fresh contextIf recurring
Unsupported assumptionClaim-to-source checkEvidence-link requirementDraft-only modeCorrect and re-verifyYesStop, askYes
Hallucinated action or resultState verification after actionVerify-then-report disciplineRead-back of system stateCompensate, correct recordYesStop, escalateYes
Permission failureAuth error signalLeast-privilege reviewGraceful denial pathHand back with clear next stepYesHand back, do not retry-privilege-escalateIf recurring
Downstream API errorError signal, timeoutTimeout and fallback designBounded retry with backoffResume from verified stateIf user-visibleRetry by failure type, then escalateIf recurring
Timeout after action may have happenedState read-backIdempotency keys where possibleVerify-before-repeatReconcile, never assumeYesVerify state first, then decideYes
Partial completion presented as successEnd-state assertionExplicit end-state definitionPer-step statusComplete, roll back, or hand back remainderYes, with what remainsEscalate remainderYes
Duplicate actionIdempotency check, audit logIdempotent designDedupe guardCompensate duplicateYesStopYes
Loop / stuck stateRepeated-equivalent-action detectionBounded attempts, stop conditionsKill switch, attempt capEscalate with traceYesStop and escalateYes
Unsafe or unintended actionPermission fence, previewNarrower scope, approvalRevoke, revertCompensate, incident pathYesStop and escalateYes
Silent failureEnd-state and heartbeat checksMandatory outcome reportingAlert on missing outcomeInvestigate from traceYesEscalateYes
Bad human handoffHandoff quality rubricStructured handoff formatKeep resumable stateRe-handoff with full contextYesHand back betterYes
Failed rollbackRollback verificationTested compensation pathsIncident pathManual remediationYesEscalateYes

12. Partial success must be explicit

Multi-step tasks do not collapse into success or failure. Track outcome states such as: not started, blocked before action, partially completed, completed with recoverable warning, completed correctly, completed with unacceptable side effect, rolled back, and escalated. Define what each state means for your workflow, who sees it, and what happens next.

13. Human review is not automatically safety

"Human in the loop" with no operational definition is theater. For every approval point, specify: what exactly the human reviews, at what point in the flow, with what evidence shown, whether the human can realistically detect the error class in question, what happens on approve, edit, reject, or no response, and whether review time destroys the economic value.

Watch for rubber-stamping: if approvers accept nearly everything without reading, the checkpoint provides legal cover, not safety. Product mechanisms that help: preview diffs, source-evidence links, structured action summaries, explicit irreversible-action warnings, undo windows, and approval history. But none of them guarantee safety, measure whether review actually catches seeded errors, and track review burden as a cost.

14. Abstention, escalation, and handback

A strong agent knows when not to act. Define when it must ask for missing input, narrow scope, abstain, escalate, hand back to the user or operator, stop after tool failure, or request approval. Then measure handback quality: does it explain what was completed, what remains, why it stopped, what the user must do next, and what state is safe to resume from? A clean handback on a task the agent should not attempt is a success, not a failure. This handoff UX is core Product work.

15. Loops, retries, and stop conditions

Agent loops multiply latency, cost, side effects, user wait, and rate-limit risk. Require bounded attempts, explicit stop conditions, idempotent actions where possible, retry policy by failure type (a read timeout and a failed payment deserve different policies), detection of repeated equivalent actions, and escalation when progress stalls. Do not publish a universal max-step number, set bounds from the task's cost, stakes, and reversibility.

16. Agent Trace Review

Use this artifact in eval sessions and incident reviews. It focuses on observable actions, state, and results, never on hidden chain-of-thought, which evaluators should not require.

# Agent Trace Review

Task / user job:
[...]

Expected end state:
[...]

Context available:
- [...]

Tools / permissions available:
- [...]

User approvals required:
- [...]

Observed action trace:
1. [...]
2. [...]

Tool failures / retries:
- [...]

Unexpected side effects:
- [...]

Final system state:
[...]

User-visible result:
[...]

Task outcome:
[Complete / Partial / Escalated / Rolled back / Failed]

Correctness / quality issues:
- [...]

Permission / control issues:
- [...]

Recovery quality:
[...]

Latency / cost / review burden:
[...]

Root failure class:
[...]

Add to regression set?:
[Yes / No + why]

Product decision:
[...]

17. Observability: answer Product questions, not just logging

A PM should be able to answer: what did the agent attempt, which tools and actions ran, what state changed, what failed and was retried, what the user approved, what was escalated, how long it took, what it cost, and whether the result could be undone. Distinguish three views: the engineering trace, the operator and debug view, and the user-visible activity history. Respect privacy and security boundaries, internal reasoning internals are not user content.

18. Cost per successful task

Never optimize token cost alone. Count model calls, context and retrieval, tool calls, retries, fallback paths, long-running state, human review time, support burden, and failure remediation. A cheaper run with lower task success is often more expensive per successful outcome. No universal cost threshold exists; compare against the baseline the agent replaces.

19. Latency and waiting UX

Agentic workflows can take far longer than chat generation. Design for it: progress visibility, background execution, resumability, notification on completion, partial results, cancellation, and timeout and failure UX. A spinner is not a delegation experience. Keep this at Product-pattern depth, async-job engineering belongs elsewhere.

20. Change management and regression

Agent behavior can shift materially when any of these change: model, provider, or version; system instructions; prompts or templates; context or retrieval corpus; tool definitions; external APIs; routing; memory or state logic; permissions; safety policy; structured-output schema; or fallback logic. Treat material behavior changes as Product changes: keep regression cases, compare against baseline, maintain version awareness, and monitor after rollout. Match rigor to risk and behavior impact, not every copy tweak needs certification.

21. Rollout: evidence-based, not ritual-based

There is no universal four-stage rollout and no "promote after two stable weeks" rule. Rollout scope follows risk and evidence: stakes, permission scope, reversibility, error detectability, eval coverage, observed production failures, user ability to recover, support readiness, economics, and monitoring. Possible scopes include internal only, sandbox or read-only, opt-in beta, low-risk actions only, approval-required production, bounded autonomy, and expanded population or action scope. Expansion is a Product decision gated on evidence, not a maturity ritual over time.

22. Autonomy Expansion Contract

Before granting more autonomy, answer in writing:

  • Which new action or permission is being added?
  • Which failure becomes newly possible?
  • Can that failure be detected, and is it reversible?
  • Do eval cases cover it, including failure and recovery?
  • Is production evidence representative of the expanded scope?
  • Does human review remain effective at the new scope?
  • Does Product value improve enough to justify added risk and cost?

Moving from "draft" to "send" is a major Product change even when the model is unchanged. The same contract in reverse governs reducing or removing autonomy: state which failure, cost, or correction burden triggered the narrowing.

23. Security and prompt-injection boundary

Agents that consume external or untrusted content and can take actions face instruction-injection and tool-abuse risks. At Product depth: separate trusted from untrusted sources, grant least-necessary permissions, require approval for sensitive actions, validate actions before execution, bound secret and data access, include malicious-input test cases, and define a monitoring and incident path. This page stays at Product implications; deeper security engineering and responsible-release work belongs with the responsible-AI release guide.

24. Worked cases

Five hypothetical cases show how the system above produces different verdicts. All scenarios are synthetic illustrations, not benchmark results.

Case A, Support triage: category assignment only

Job: assign each incoming ticket a category and priority. Tempting agent candidate, wrong architecture. A deterministic rules layer plus a classifier, or one-shot AI categorization, covers it with lower latency, cost, and failure surface. Verdict: do not use an agent. Agentic planning adds nothing when the output is a single label.

Case B, Support workflow: account lookup, policy search, ticket action

Job: route the ticket and prepare the next action, which requires inspecting customer and account state, checking known incidents, searching policy docs, deciding whether data is sufficient, drafting a reply, and creating or escalating an issue, including choosing whether to send or request approval. Variability, multi-step tool use, and conditional planning justify an agent. The Delegation Contract defines read vs write tools, approval before any customer-visible send or state change, abstention on missing account context, and handback with completed-vs-remaining state. Eval cases cover stale account data, contradictory policy sources, permission denials, and duplicate tickets. Verdict: bounded agent with approval-gated actions.

Case C, Release notes from merged PRs: the overkill detector

Job: draft release notes from merged PRs. Retrieval plus deterministic PR selection plus one-shot generation is sufficient; autonomous multi-step planning adds little value and new failure modes. Verdict: do not use an agent. This is the key information-gain example: agentic architecture must earn its complexity, and here it does not.

Case D, Customer summary with traceability

Job: produce an account brief from CRM data, recent tickets, usage summaries, and meeting notes. Risks dominate: stale state, cross-account leakage, unsupported claims, incorrect aggregation, source conflict. The source contract names the authoritative system per field, requires evidence links for every non-trivial claim, grants read-only permissions, and hands back on missing context. Eval dimensions include leakage probes (never with real customer data), conflict cases, and correction burden, because a brief the user must fully re-verify saves nothing. Privacy and risk depth routes to the responsible-AI release guide. Verdict: read-only bounded agent, draft mode, human sends.

Case E, External action and B2B remediation

Job: investigate a failed customer data sync and take bounded remediation action across systems, potentially sending email, modifying account state, or triggering a workflow. Output quality is insufficient as a metric; what matters is action preview, permission scope, confirmation logic, validation, audit history, undo or compensation, and escalation. Long-running execution needs progress visibility, resumability, and stall detection; cost and time accounting decides whether the autonomy pays. Verdict: bounded autonomy for reversible steps, confirmation for everything else, full trace review before any scope expansion.

25. Product metrics: outcome, not smartness

Fluency is not success. Track four layers, with no universal targets and no single blended "agent score":

  • Task and reliability: eligible attempts, successful end states, partial completions, escalations, rollbacks, unacceptable side effects, recovery success.
  • User and workflow: correction burden, time to outcome, abandonment, successful handoffs, trust and continued use where relevant.
  • Operational and economic: latency, retries, cost per successful task, human review time, support incidents.
  • Product and business: the outcome the workflow exists to move, resolution time, sync recovery rate, brief acceptance, defined per workflow.

Denominators matter

A completion rate is meaningless without its denominator. Distinguish eligible tasks, tasks delegated, tasks attempted, tasks completed, tasks completed without correction, and tasks requiring escalation. Define exclusions before analysis, never filter out difficult cases after the fact. (Measurement-implementation depth belongs to the analytics instrumentation guide; this page owns what agent reliability must measure.)

Severity beats averages

A 95% success rate is unacceptable when the 5% deletes data; a lower rate can be fine for low-stakes drafts with easy detection and correction. Preserve failure categories, inspect performance by important slice, weight decisions by consequence rather than aggregate percentage, and never invent weighted-risk formulas that fake precision. Deeper harm analysis routes to the responsible-AI release guide.

26. Common failure modes

  • Agent by default: multi-step autonomy chosen where a deterministic workflow or copilot would win.
  • Fluency equals quality: good conversation hides wrong system state.
  • Permission sprawl: broad write access granted for convenience.
  • Human-in-the-loop theater: approval exists but no human can effectively verify.
  • Happy-path evals: no tool failure, ambiguity, or edge cases.
  • Final-answer-only eval: actions and side effects ignored.
  • Silent partial completion: the agent stops halfway but presents success.
  • Retry theater: retries multiply side effects and cost.
  • No handoff contract: the user cannot understand or resume a failed task.
  • Aggregate success theater: rare severe failures hidden by a high average.
  • Token-cost optimization: retries, review, and support ignored.
  • Autonomy as status: more autonomy treated as inherently better.
  • No regression contract: prompt, model, or tool changes alter behavior silently.

27. FAQ

What is an AI agent in Product Management? A system delegated a bounded multi-step task: it plans across steps, uses tools that read or change state, observes results, and recovers or escalates. The Product job is defining that delegation boundary, not the underlying model.

What is the difference between an agent and an automated workflow? A workflow executes fixed steps in a fixed order; an agent adapts its plan, chooses among tools, and responds to intermediate state. If your paths are enumerable and validation is exact, a workflow wins on cost, latency, and testability.

When should a Product use an AI agent instead of a chatbot or copilot? When the job structurally needs multi-step planning, tool use across variable paths, and recovery, and a one-shot generation with the user as actor cannot cover it. Cases A and C above show where the answer is "do not use an agent."

How do you evaluate an AI agent? End to end: deterministic checks on state and actions, a task-quality rubric, validated model-assisted review at most as a screen, trace-level inspection of tool choice and recovery, online Product outcomes, and operational economics. Judge the state transition, not the final text.

What metrics should an AI agent use? Task reliability (success, partial, escalation, rollback, side effects), user burden (correction, time to outcome, abandonment), economics (latency, retries, cost per successful task, review time), and the workflow's own Product outcome. No universal thresholds.

How much autonomy should an AI agent have? As much as stakes, reversibility, detectability, and economics justify per action class, set in the Autonomy / Approval Matrix. There is no ladder to climb; "draft forever" can be the correct end state.

What should require human approval? Irreversible, broad-scope, or hard-to-verify actions: external sends, publishing, state changes, spending, deletion, permission changes. Approval must include inspectable evidence and a real ability to detect the error.

How do you handle agent tool failures? By failure type: verify state before repeating any write, use idempotency where possible, bound retries with backoff, and escalate with the trace when progress stalls. Timeouts after a possible write are reconcile-first, never blind-retry.

How do you stop an agent from looping? Bounded attempts, explicit stop conditions, detection of repeated equivalent actions, and escalation on stall, sized from the task's cost and stakes, not from a universal step limit.

What is agent observability? The ability to answer what was attempted, which tools ran, what state changed, what failed and was retried, what was approved or escalated, what it cost and took, and whether it can be undone, across engineering, operator, and user-visible views.

How do Product Managers monitor agent cost and latency? Per successful task, including retries, retrieval, tool calls, and human review, compared against the baseline the agent replaces, with waiting UX (progress, backgrounding, resumability, cancellation) designed as part of the Product.

When should an agent hand back to a human? On missing or contradictory context, permission denial, tool failure it cannot recover from, high-stakes uncertainty, or any stop condition, with completed work, remaining work, reason, next action, and resumable state stated.

28. What to do next

Last materially reviewed: October 3, 2026. Agent tooling moves quickly; principles above are durable, while provider-specific behaviors should be verified against current official documentation (including Anthropic's agent-building and agent-eval guidance, OpenAI's agent evaluation guidance, the NIST AI Risk Management Framework and Generative AI Profile, and OWASP guidance on agentic security) at implementation time.

Recommended courses

From the blog

Portrait of Andrea Mezzadra, author of the blog post

Andrea Mezzadra@____Mezza____

Published on September 23, 2025 • Updated on October 3, 2026

Ex Product Director turned Independent Product Creator.