In spring 2026, during the war between the United States and Iran, an alarming intelligence report circulated within the US military. It claimed that a Chinese ship passing through the Middle East was carrying components linked to a nuclear weapons programme. The report prompted preparations to seize the vessel: armed forces were placed on alert, and US aircraft had already been launched. Only a further review, conducted at an advanced stage of the preparations, found that the claim was false.
According to a CNN investigation based on four sources familiar with the incident, an analyst at the US military’s special operations command in the Pacific used an AI chatbot to analyse the ship’s cargo. The analyst included both open-source information and classified signals intelligence in the query. The system connected the pieces of information and reached the incorrect conclusion that the cargo was linked to a military nuclear programme. That conclusion was incorporated into an intelligence report circulated to other officials, without adequate scrutiny of its sources or how it had been reached.
This was not an autonomous weapon system deciding to open fire. The AI did not launch a weapon, nor did it independently decide to attack the ship. On the face of it, humans remained in the decision chain at every stage. Even so, the system’s erroneous output nearly led to a military operation against a Chinese vessel, with the potential to cause a direct crisis between two major powers.
The incident involving the Chinese ship shows that the problem does not begin only when an autonomous system fires a weapon. It begins earlier, when AI interprets intelligence, connects disparate pieces of information, attributes intent and recommends who or what should be considered a threat. Without independent verification, access to sources, measures of confidence and a person able to halt the process, a single computational error can become a military decision with strategic consequences within minutes.
This is a different, and sometimes more complex, risk than the familiar debate over “killer robots.” Systems such as Project Maven were originally designed to scan drone and satellite imagery, identify objects and reduce analysts’ workloads. Since then, the Maven Smart System has developed into a decision-support system that combines information from multiple sensors and sources, helps identify and prioritise targets, and passes them along for further action in the targeting process. Officially, humans remain responsible for selecting and approving targets. But as the system processes more information and presents its conclusions quickly and confidently, the risk grows that the human operator will shift from scrutinising its findings to simply approving them.
US Department of Defense Directive 3000.09 governs weapon systems that, once activated, can select and engage targets without further input from an operator. It requires commanders and operators to be able to exercise “appropriate levels of human judgment” over the use of force, as well as reliability testing, intervention mechanisms and legal review. But the directive focuses primarily on the weapon system itself. It does not fully address AI systems operating earlier in the process, when a target is still being constructed from images, signals, documents and intelligence assessments.
The phrase “human in the loop” can also be misleading. If that person receives a conclusion without access to the underlying material, does not know which parts were produced by AI and is given no clear assessment of uncertainty, their presence does not guarantee meaningful oversight. Time pressure, information overload and the tendency to regard computer output as objective can reduce human review to a rubber stamp.
Oversight therefore cannot begin only before someone pulls the trigger. It must run through the entire decision chain: clearly labelling AI-generated content, preserving sources and the reasoning behind conclusions, having another analyst cross-check the findings, presenting alternative explanations, and separating those who produce an assessment from those who approve an operation. Otherwise, responsibility may formally remain with a human even though the decision has, in practice, already been made at an earlier stage by a system whose reasoning no one fully understands.
Explore stories by subject ↗
