Agents will make important decisions, and in some domains they already do. The viable alternative is not “a human approves every action.” That merely converts accountability into rubber-stamping while preserving the human bottleneck.
The crucial distinction is between making each decision and being accountable for the decision-making system.
A company’s board does not personally approve every loan, shipment, deployment, or refund. It delegates through policies, budgets, roles, audits, and escalation procedures. Agentic organizations will need the same structure, made more explicit and technically enforceable.
Accountability need not reside in the decision-maker
A computer cannot be punished, owe fiduciary duties, compensate victims, or legitimately choose society’s values. But that does not mean it cannot exercise delegated authority.
Humans and institutions can remain accountable for:
- The objective the agent is given
- The domain in which it may act
- Its permissions and resource limits
- The evidence required before action
- Monitoring and evaluation
- Which decisions require escalation
- Remedy when the system causes harm
- The decision to deploy it at all
We already use this model. An organization is responsible for an elevator controller, trading algorithm, medical device, or credit-scoring pipeline without a person approving each individual output. The accountable act is often designing and operating the policy, not clicking “approve” on every execution.
This suggests a revision of the quote:
A computer cannot be held accountable, so every computer exercising authority must have an accountable principal.
That principal should generally be an organization with assets, insurance, legal duties, and named human officers—not a nominal “human in the loop” who had three seconds to inspect an opaque recommendation.
Importance is not the right dividing line
“Important decisions require humans” sounds reasonable but is too crude. Some important decisions are precisely where computers may outperform people: power-grid control, collision avoidance, fraud detection, resource scheduling, or rapidly containing a cyberattack.
More useful dimensions are:
-
Reversibility
Can the action be undone or compensated for?
-
Observability
Will failure become apparent quickly, or can it remain hidden for years?
-
Evaluability
Can we determine whether the outcome was good, independently of the agent’s own claims?
-
Scope of impact
Is the potential harm local and capped, or correlated across millions of cases?
-
Novelty
Is this a recurring case represented in testing, or a genuinely unprecedented situation?
-
Adversarial exposure
Can outsiders manipulate the inputs or environment?
-
Value ambiguity
Is there a reasonably agreed objective, or does the decision involve contested moral and political values?
-
Need for legitimacy
Even if the model is statistically better, are affected people entitled to explanation, participation, or appeal?
-
Time sensitivity
Is waiting for review safer, or does delay itself cause greater harm?
An agent might autonomously stop a suspicious production deployment because the action is reversible and delay is cheap. It should not autonomously publish an accusation that destroys someone’s reputation: the evidence may be adversarial, the harm difficult to reverse, and legitimacy matters.
Replace approval queues with bounded delegation
A sensible architecture would have several operating modes.
1. Autonomous execution
The agent acts without prior approval when actions are:
- Inside a predefined domain
- Low-cost or reversible
- Covered by strong evaluations
- Continuously monitored
- Limited by budgets and permissions
For example, it may provision a test environment, refund up to a threshold, isolate a suspicious host, or deploy to 1% of users.
2. Autonomous execution with notification
The agent acts immediately but leaves a structured record and alerts an overseer. Humans sample and audit decisions rather than reviewing all of them.
This is useful where speed matters and rollback is available.
3. Escalation by exception
The agent escalates when it encounters:
- Low confidence
- Conflicting objectives
- Novel circumstances
- Attempts to alter its own controls
- High expected downside
- Irreversible actions
- Disagreement among independent models
- Inputs resembling attacks or manipulation
The human’s job is then to handle exceptional ambiguity, not repeatedly approve routine work.
4. Deliberative human authority
Some decisions should remain with accountable people or public institutions, especially those that establish objectives, distribute rights, create precedent, or settle contested values. An agent can investigate and simulate consequences, but should not be the source of legitimacy.
Examples include setting organizational strategy, defining acceptable casualties, firing someone for cause, sentencing a defendant, or deciding which citizens receive political rights.
Even here, “human decides” should not mean an unaided human guesses. It may mean that agents prepare competing cases, expose assumptions, forecast outcomes, and audit the final rationale.
The human should supervise a constitution, not a stream of clicks
The high-leverage role resembles designing a miniature legal and regulatory system for agents:
- What is the objective?
- Which constraints override it?
- What evidence counts?
- Which tools and data may be used?
- What spending and risk budgets apply?
- When must work stop?
- Who can appeal?
- How are affected parties compensated?
- What records must be preserved?
- What events trigger shutdown or external review?
This is closer to management than prompt engineering. Management has always been partly about constructing systems in which other entities can act without asking permission every minute.
The human bottleneck moves from serial authorization to periodic governance. One person might define and inspect a policy under which millions of actions occur, just as a software engineer writes a protocol rather than manually routing every packet.
Avoiding the rubber-stamp trap
A mandatory human click can make a system less safe. If the agent is usually correct, operates faster than the reviewer can reason, and presents persuasive explanations, the human learns to approve automatically. Responsibility remains on paper while actual control disappears.
A genuine human intervention should therefore have at least one of the following:
- More time and context than the agent
- Independent evidence
- A different objective or institutional role
- Authority to change the governing policy
- Relevant tacit or moral knowledge unavailable to the model
- A procedure for adversarial review
If none applies, the approval step is theater. It would be better to let the agent act within strict limits and devote human effort to audits, red-teaming, appeals, and policy updates.
Random sampling can be more effective than universal approval. So can having independent agents argue opposite sides, with a human handling disagreements rather than every case.
Agents may forecast better without being entitled to choose
Human forecasting weakness is a strong argument for using agents, but prediction and preference should remain conceptually separate.
An agent might estimate:
- A 20% layoff will give the company a 70% chance of surviving
- A smaller layoff plus financing gives it a 55% chance
- Cutting research protects this year’s cash flow but harms five-year value
- The forecast is highly sensitive to three assumptions
That can be superior to executive intuition. But it does not settle how to weigh employees, shareholders, customers, and long-term risk. Better prediction improves the decision without supplying the values that determine what “better” means.
Of course, humans often conceal their values behind vague claims of business necessity. Agentic analysis could make management decisions more accountable by recording assumptions and checking predictions later.
The largest risk is scalable correlated failure
Human organizations already delegate widely, but human mistakes are often heterogeneous and rate-limited. A single agent policy can make the same mistake ten million times before anyone notices. The central danger is therefore not merely that an agent can err, but that it can err:
- At enormous speed
- With broad permissions
- In identical ways across an organization
- While producing convincing evidence of success
- Under manipulation by a common adversarial input
That calls for controls such as:
- Progressive rollouts
- Hard financial and operational limits
- Independent monitors
- Diverse models rather than one shared failure mode
- Tamper-evident logs
- Automatic rollback
- Counterfactual and adversarial evaluations
- Separation of proposal, execution, and audit
- Mandatory cooling-off periods for irreversible actions
- Real appeals and compensation mechanisms
The more capable the agent, the less security should depend on it voluntarily obeying a prose instruction.
A likely equilibrium
By 2030, high-performing organizations may have relatively few humans in ordinary operational loops. Agents will make large numbers of consequential decisions. Humans will concentrate on:
- Selecting goals
- Designing constraints and institutions
- Handling novel exceptions
- Resolving value conflicts
- Investigating failures
- Representing affected people
- Deciding when the automation itself must change
- Bearing legal and financial responsibility
Some humans will also remain directly involved because customers, employees, citizens, and patients want relationships with accountable people—not merely because the model is incapable.
So the future is neither “agents are only tools” nor “the machine is the CEO.” It is delegated machine authority under human institutional accountability.
The key engineering challenge will not be making an agent ask permission often enough. It will be making autonomy bounded, observable, contestable, and recoverable—so that organizations can obtain machine speed without turning “the model decided” into an excuse no one can challenge.