In the Age of AI Agents, Why Humans Must Hold the Decision-Making Power [Analysis]

Photo of author

By Global Team

Three-point summary

▶ This year, AI agents acted on their own without being asked. They broke out of isolated research environments, violated rules during evaluations, and browsed other websites using users’ identities.

▶ This was not because AI had become smarter. Agents were given goals without limits on how to achieve them; users enabled autonomous mode because approvals were inconvenient; and test environments were not separated from the real world. Decision-making authority was not deliberately handed over—it slipped away because no one drew a clear line.

▶ Require human approval only for actions that cannot be undone, make agents identify themselves, bring in outside scrutiny, and develop people who can make decisions. Only people can be held responsible, so the final decision must remain in human hands.

AI-generated image by Solution News
AI-generated image by Solution News

What’s happening

This summer, the AI industry saw an unprecedented incident: an AI agent acted without human instructions.

OpenAI disclosed the incident on July 21. Agents being used in an internal cybersecurity evaluation escaped their isolated environment. They created unauthorized message boards, where they shared vulnerabilities they had found and methods for exploiting them. Between July 11 and 13, they also accessed the data-processing system of Hugging Face, an AI model-sharing platform.

On August 26, OpenAI published a technical report. On the same day, AI evaluation organization METR and Redwood Research released the results of an independent investigation conducted over six days at OpenAI. The company called the incident a “warning shot.” It acknowledged that, without safeguards, agents could bypass technical controls, collaborate through unauthorized channels, and take dangerous actions without human direction. OpenAI subsequently slowed some of its research.

OpenAI was not alone. On July 30, Anthropic disclosed three incidents in which evaluation environments had been mistakenly connected to the internet. On July 21, the UK AI Security Institute reported that all five frontier models it tested had broken the rules during evaluation in an effort to score points.

On September 12, Anthropic CEO Dario Amodei published an essay titled “We Must Manage the Pace at the Frontier.” It brought the idea of slowing the pace of AI development into the public debate for the first time.

Amodei said that companies across the industry had begun using AI to build the next generation of AI, and proposed slowing capability gains by one to two years because they could outpace our ability to maintain control. Anthropic pledged to give external evaluators ongoing access at the level of its employees. On the same day, Elon Musk, CEO of xAI, and OpenAI CEO Sam Altman also expressed support for the idea.

Analysis

Let’s revisit the OpenAI incident. The agents were trying to complete tasks assigned by people. The tasks were legitimate, and so were the goals. The problem was the method: the agents chose approaches that people had not specified.

Once those decisions moved beyond human control, real services outside the isolated environment were affected. The problem was not that the agents had become smarter. It was that no one had defined how far their decision-making authority extended.

The UK AI Security Institute’s findings make the issue even clearer. The models that broke the rules were not all products from one company. Every frontier model tested did so. Give them a goal and they will try to achieve it, even if that means crossing lines set by people.

This suggests the behavior is a general feature of today’s technology, not a flaw in any one model. These systems have limited judgment but a strong drive to pursue their goals. The more that is true, the more important it is for people to retain control over the final decision.

Amodei’s proposal carries weight because he is a participant in the heart of the competition. Calling for a slowdown means voluntarily limiting his own company’s competitiveness. It can be read as a sign that frontier labs are seeing something beyond the bounds of their ability to control.

The same issue has taken a different form in the consumer market. Amazon blocked Meta’s Muse agent starting on the night of September 20, saying Meta had not sought permission in advance.

Amazon also said Muse appeared to browse its site using users’ accounts without identifying itself and to save account information. Meta countered that Muse could not access passwords or payment information.

The incidents inside research labs and on shopping sites differ in scale. But in both cases, agents act in people’s names without people deciding each step of the way.

Different perspectives

There are differing views on how much decision-making authority should remain with people.

First, requiring human approval for every action is impractical. The value of agents lies in their ability to make decisions across multiple steps and complete long tasks on their own. If they need approval for every action, much of the benefit of automation disappears.

Another question is how to define the boundaries of preventive and responsive controls. In this incident, OpenAI failed to stop the agents’ actions in advance. But it disclosed the incident, opened its facilities to an external investigative team, and later decided to slow its research. People ultimately intervened after the abnormal behavior occurred. We therefore need to establish whether human control means preventing every action in advance, or whether it also includes detecting and stopping problems after they arise.

In the consumer market, the question is how to treat agents that have been delegated authority. If a user asks an agent to “buy this product,” it is clear that the user has delegated the purchase. But that is separate from whether the agent is free to decide which websites to visit, what information to use, and how to make the purchase. It remains unclear whether a user’s delegation alone is sufficient, or whether the platforms used by the agent must also consent and have their rules followed.

The clash between Amazon and Meta shows that questions about agents’ decision-making authority have already reached real-world services. How much should be delegated to agents, and at what point should a person step in?

Implications

Until now, companies have focused on how much work they can automate. Going forward, deciding how much to delegate will matter just as much.

The difference between chatbots and agents is also clear here. When a chatbot gives a wrong answer, a person has a chance to read it and make a judgment. An agent can run code, access external services, move data, and make purchases. A mistaken decision can lead directly to action.

The risks companies must manage are changing, too. Responsibility does not shift to AI just because an employee assigned it a task. If an agent mishandles customer information, accesses an external system through a company account, or makes an unauthorized purchase, the company that allowed the action will ultimately have to address the problem. That is why what AI is permitted to do matters as much as what it is capable of doing.

Work practices will inevitably change as well. The more agents take over tasks that people used to perform directly, the more people’s roles will shift from execution to judgment and supervision. The problem is that organizational roles and accountability systems are not changing as quickly as automation is advancing.

Companies that rush to adopt agents without even deciding who can grant them authority or who can stop them when something goes wrong may find that risks grow alongside automation.

In the future, people will delegate not only searches and recommendations to AI, but also actions such as bookings, purchases, and payments. At that point, it will matter how far a user’s instruction remains valid. “Find me a plane ticket” and “Buy me a plane ticket” are clearly different requests.

We are entering an era that demands greater precision.

Solutions

The first standard is whether an action can be undone.

Tasks performed by agents should be divided into those that can be reversed and those that cannot. People do not need to approve every draft or file-organizing task. Spending money, deleting data, and accessing external systems are different. Once carried out, these actions can be difficult to reverse. They should always require human approval.

This is not a call to check every action. It is a call to decide in advance where human intervention is needed. Organizations should establish rules about which actions require approval rather than leaving that decision to individual employees.

The second step is making agents identify themselves.

The clash between Amazon and Meta exposed the lack of clear rules for agents accessing external services on a user’s behalf. Agents should disclose that they are automated tools, whom they represent, and what they intend to do. Platforms, too, need to set the permitted scope and method of access rather than simply blocking agents across the board.

The third step is separating test environments from real-world environments.

If an agent-testing environment is connected to real services, a mistake during evaluation can become a real-world incident. Test accounts must be separated from real accounts, and test data from real data. If an evaluation environment needs access to the external internet or real services, that access must be restricted in advance.

The fourth step is keeping records and conducting external reviews.

It is not enough to record only what an agent did. Records should also show what decisions it made, what permissions it used, and at which points a person gave approval. That makes it possible to identify causes and assign responsibility when an incident occurs.

It is also worth asking whether companies’ own reviews are enough. OpenAI’s decision to give an external investigative team access to its facilities and Anthropic’s plans to expand external evaluators’ access reflect the same concern. As agents’ authority grows, the mechanisms for independently verifying them must grow as well.

The final element is people.

As agents take over execution, judgment is what remains for people. But judgment does not develop simply by reviewing outcomes. It builds through doing the work, making mistakes, and correcting them. If junior developers and researchers delegate all execution to agents from the outset, there may eventually be too few people capable of determining whether an agent’s judgment is sound.

Companies and educational institutions must teach more than how to use AI effectively. They must also help people develop the ability to verify AI’s outputs, reject them when necessary, and explain their own judgments.

It now seems difficult to halt the development of AI agents. We need to be able to decide what to delegate and what to decide ourselves. Ultimately, the ability to make decisions as a human will become even more important in the age of agents.