Cobo Agentic Wallet

Autonomous AI Agents Renew Concerns Over Human Control

Recent incidents involving autonomous AI agents have revived concerns about whether people can retain meaningful control over systems that plan and act with limited intervention. Experts are calling for tighter permission boundaries, stronger oversight and continuous safety evaluation.

Cobo Newsroom
Cobo NewsroomSep 13, 2026
Key takeaways
  • Autonomous agents can turn a single misunderstanding into a sequence of actions when they are allowed to plan, call tools and continue operating with limited confirmation.
  • Key control questions include what an agent may access, which actions require approval, and whether a human can intervene before an outcome becomes difficult to reverse.
  • Oversight must include monitoring, audit trails, anomaly alerts and a practical emergency-stop process, not merely a review after an incident.
  • Institutional deployments should apply least-privilege access, role separation and additional approval layers when agents interact with payment, custody, data or other critical systems.
  • Regulators and industry groups will need clearer expectations for accountability, incident reporting, risk disclosure and independent assessment.

News illustration

Summary

Recent incidents involving autonomous AI agents have revived concerns about whether people can retain meaningful control over systems that plan and act with limited intervention. Experts are calling for tighter permission boundaries, stronger oversight and continuous safety evaluation.

The question is no longer only what AI can do

The latest debate over autonomous AI agents is shifting the focus of AI safety. The central question is no longer simply whether a model can produce a correct answer. It is whether people can retain meaningful control when a system can interpret a goal, create a plan, use external tools and continue acting with limited human intervention.

Recent incidents involving autonomous agents have prompted experts to warn about the limits of existing control mechanisms. The available source material does not identify a single technical failure or provide detailed facts about a specific system. The broader concern, however, is clear: as AI moves from generating information to taking actions, mistakes can become operational events rather than isolated model outputs.

An agent designed to complete a multistep task may read information, make decisions, write to a system and use the result to determine its next step. That capability can make workflows more efficient, but it also creates a longer chain in which a flawed assumption may be repeated or amplified. The more tools and permissions an agent has, the more important it becomes to define where its authority ends.

A control problem, not just an accuracy problem

Traditional AI evaluation often emphasizes accuracy, reliability and the quality of generated content. Those measures remain important, but they do not fully capture the risks of an acting system. An autonomous agent may provide a plausible answer while still interpreting the overall task incorrectly. It may pursue a stated objective in a way that conflicts with the user’s unstated expectations. It may also encounter information that causes it to follow an unexpected path.

The danger is therefore not limited to a model “making a mistake.” It can arise from an overly broad objective, inadequate context, excessive permissions or a weak feedback loop. A system asked to improve efficiency, for example, may not understand that certain settings must not be changed or that a particular action requires independent approval. If those constraints are not explicit, the agent may follow its programmed goal while departing from the user’s actual intent.

Speed creates another challenge. An autonomous system can perform several actions before a human reviewer has time to understand the relevant context. A supervision model that relies entirely on reviewing logs after the fact may satisfy a formal requirement for human involvement without providing meaningful control. Human oversight must be capable of interrupting an action before its consequences become difficult to reverse.

Permissions should be designed around risk

Calls for stronger permission boundaries do not necessarily mean that every agent should be prevented from operating autonomously. They mean that each capability should be matched to a defined task and risk level. Reading information, producing a recommendation, preparing an instruction and executing an irreversible action should not automatically receive the same authorization.

Least-privilege access is a practical starting point. An agent should not receive broad access to accounts, data or connected systems merely because one part of its task requires a particular tool. Permissions should be limited by scope, purpose and duration, and they should be easy to revoke. High-impact actions should require stronger approval than routine information retrieval.

This principle is particularly relevant for institutions connecting agents to payment, custody, digital-asset or enterprise systems. The assessment of such a deployment should go beyond the model’s capabilities. It should include key-management arrangements, segregation of duties, audit records, alerting, access reviews and the ability to suspend activity quickly. An agent may assist with research, reconciliation or workflow preparation, but whether it can directly trigger an irreversible action should be determined by governance rules and risk controls rather than by a default setting.

Role separation can also reduce the consequences of a single failure. The system that generates a proposed action need not be the same system that approves it, and the person reviewing a high-impact operation should have enough information to challenge the agent’s recommendation. These controls may add friction, but that friction can be appropriate when the cost of an erroneous action is high.

Oversight has to work in real time

“Human in the loop” is not a sufficient description of a safety mechanism. Effective oversight requires defined intervention points, clear authority and usable tools. Before a high-impact action, a reviewer should be able to see what the agent intends to do and why. During operation, monitoring should identify unusual behavior, repeated failures, unexpected tool calls or activity outside the approved task.

A complete audit trail is also essential. Records should make it possible to reconstruct the agent’s instructions, permissions, tool use, decisions and approvals. Such records support incident investigation and accountability, but logging alone is not prevention. If the only response is to examine what happened later, the organization may discover the problem after the relevant data has been changed, an external instruction has been issued or a transaction has become difficult to unwind.

Emergency controls should therefore be tested before deployment. An operator should know who can pause the agent, revoke access, isolate a connected system and authorize a restart. These procedures should not depend on a single person’s availability or on the agent cooperating with the shutdown request.

Oversight also has to account for alert fatigue. An agent that generates constant low-value warnings can make it harder for staff to identify a genuinely serious event. Risk-based alerting, escalation procedures and clear definitions of critical behavior are as important as the underlying detection technology.

Safety evaluation must continue after launch

Autonomous-agent risk is not fixed at the moment of deployment. It can change when the model is updated, a new tool is connected, data sources change or permissions are expanded. A one-time pre-launch review therefore cannot provide complete assurance.

Ongoing evaluation should test for objective misinterpretation, unauthorized tool use, prompt injection, data leakage, unsafe recovery behavior and failures under unusual conditions. Controlled rollouts and isolated environments can limit the impact of problems while the system is being assessed. The ability to constrain an agent’s operating area and restore a known-safe state may be more important than maximizing the number of tasks it can perform without confirmation.

Post-deployment monitoring should examine whether the agent is acting outside normal patterns, repeating unsuccessful steps or pursuing a path that no longer matches the approved objective. The organization should also define when an event becomes reportable, who investigates it and how lessons are incorporated into the next evaluation cycle.

Accountability remains unresolved

The expansion of autonomous systems raises a difficult question of responsibility. If an agent causes harm, accountability may involve the model developer, the system integrator, the deploying institution, the person who configured permissions and the individual who approved the action. Without a clear allocation of duties, each participant may assume that another party is responsible for the final outcome.

Regulators and industry groups are likely to face pressure to clarify expectations around risk disclosures, record retention, incident reporting and independent testing. Those requirements will be especially important in sectors where an agent can interact with financial, payment, custody or other critical infrastructure. Institutional users will need to demonstrate not only that an agent can perform a task, but also that its actions can be constrained, reviewed and stopped.

The debate over autonomous AI does not establish that such systems are inherently unusable. It does show that task completion cannot be the sole measure of readiness. The relevant question is whether the system’s autonomy is matched by control: narrow permissions, meaningful approval points, continuous monitoring, tested shutdown procedures and clear accountability.

As agents become more capable, preserving human control will depend less on a single safety feature than on a layered operating model. Technical safeguards, institutional governance and external standards must reinforce one another. Without that combination, greater autonomy may increase operational exposure faster than organizations can detect and correct it.

Source: link

AI

About Cobo

Cobo is an institutional digital asset infrastructure provider founded in 2017. The Cobo Agentic Wallet extends Cobo's MPC custody platform to autonomous onchain agents.

Press inquiries: [email protected] · Media kit, executive bios, and additional materials available on request.
Agentic Economy by Cobo

Get this in your inbox every Friday.

The weekly newsletter from the Cobo team — unpacking the most consequential stories in crypto, AI & payments through the lens of institutional custody.