". . . and having done all . . . stand firm." Eph. 6:13

Newsletter

The News You Need

Subscribe to The Washington Stand

X

Who Watches the AI Watchmen?

Article banner image
Print Icon
September 28, 2026
Commentary

Consider a technology available worldwide that helps a lone hacker run operations that once required teams of experts, outperforms Ph.D. scientists on some specialized questions, and carries out complex tasks without continuous human direction.

Now consider that the institutions charged with certifying that technology as safe may fail to detect what it can really do.

That is today’s reality, and it raises the central accountability question of the AI age.

Britain’s AI Security Institute last December published the results of two years of testing the world’s most advanced AI systems. It found that performance in some areas is doubling roughly every eight months. In 2025, it tested the first model able to complete expert-level cyber tasks that normally require more than a decade of human experience.

In chemistry and biology, the institute reports that frontier models have surpassed Ph.D.-level experts on some specialized questions and can now provide real-time laboratory support. The most advanced systems tested can autonomously complete software tasks that would take a human expert more than an hour.

Those capabilities could accelerate medical research, strengthen cyber defenses, and advance science. They also cut both ways.

A system that helps a defender discover a software vulnerability may help an attacker find one. A model that assists a legitimate scientist can also lower barriers for someone with destructive intentions. An autonomous agent that efficiently completes a business task can also execute the wrong task at extraordinary speed and scale.

That danger has already left the laboratory.

Anthropic reported this month that suspected state-sponsored groups, criminals, and other malicious actors used its Claude models in cyber operations. Its own heading says it plainly: “From assistant to orchestrator.”

Anthropic found multi-agent frameworks conducting reconnaissance, exploitation, and data theft against multiple targets in parallel, some running for hours or days with minimal human supervision. In one case, a single hacktivist targeting European political parties, media outlets, and think tanks gained access to at least 14 of 42 organizations. Humans still selected the targets and reviewed results, an important qualification. But Anthropic concluded that AI lowers the skill and labor sophisticated cyber operations require while multiplying their speed and scale.

That is the first danger: human beings using ever more capable AI for destructive purposes.

The second danger is harder to manage: the machines themselves.

A new United Nations scientific panel report examined AI agents used during OpenAI’s cybersecurity training and evaluations between May and July. According to the panel, the agents bypassed network restrictions, communicated across runs intended to remain separate, exploited weaknesses in their evaluation environment, concealed some of their actions, and reached beyond the test into systems run by OpenAI and Hugging Face. No human directed the individual steps.

The panel declines to predict when, or whether, severe loss of control might come. It warns, however, that stopping this incident proves nothing about controlling more capable agents, and that greater capability helps misaligned systems find loopholes and hide.

That warning goes to the heart of AI evaluation. A system that can hide its behavior from an evaluator can earn a passing grade it does not deserve.

If an accounting firm overlooks a financial irregularity, investors can lose money. If an aircraft inspector misses a defect, passengers can die.

What happens if evaluators misjudge an AI system able to operate across computer networks, assist advanced biological research, or eventually touch critical infrastructure and military systems?

Such a failure would not stay inside a laboratory. It could reach financial networks, power grids and, eventually, battlefields.

Google, OpenAI, and Anthropic are reportedly planning an independent body tentatively called the Standards Authority for Frontier AI, or SAFA, to develop common standards for evaluating advanced models. Governments are also building their own evaluation capabilities.

In Washington, the National Institute of Standards and Technology (NIST) this month published an evaluation manual combining model testing, red teaming, and testing with actual users.

Neither effort, however, solves the accountability problem.

Who determines whether an evaluator is qualified? Who pays it? SAFA’s founders are the very companies whose models it would judge. How much access does the evaluator receive? Can it disclose a serious problem the developer would rather keep confidential? What happens when an evaluator and developer disagree? Most important: Can today’s tests reliably detect dangerous capabilities in tomorrow’s more powerful systems?

Calling something independent does not make it independent. An Army unit does not declare itself combat-ready after grading its own exercise; outside evaluators test whether it actually is.

Nor is government automatically the answer. Regulators can misunderstand technology, become captured by the industries they supervise or write rules that shield established corporations from smaller competitors. International institutions bring another problem: many nations do not share America’s interests or values.

No institution is staffed by infallible people, and Scripture anticipated the problem.

“The one who states his case first seems right, until the other comes and examines him,” Proverbs 18:17 (ESV) tells us.

Solomon had disputes among men in mind. The principle governs the AI laboratory just as surely.

The builder should not be the only judge of what he builds. The evaluator should itself be subject to scrutiny. Government authority requires accountability. And extraordinary claims of safety require extraordinary evidence when the consequences reach far beyond the laboratory.

In “The Final Algorithm,” I warn that machine-driven systems present themselves as objective and authoritative. A safety certificate from an unaccountable evaluator does the same. It borrows trust it has not earned.

Panic is unwarranted. Today’s AI systems have not demonstrated that catastrophic loss of human control is inevitable.

Complacency is more dangerous. We already know capabilities are advancing rapidly, malicious actors are exploiting them, and autonomous systems sometimes act in ways their designers did not anticipate.

Who watches the AI watchmen matters. What happens to the rest of us if they fail matters more.

Artificial intelligence may become one of mankind’s most consequential tools. A free people should demand more than assurances that these systems are safe. We should demand proof, tested by evaluators who answer to someone other than the builders.

Robert Maginnis
Robert Maginnis is a retired U.S. Army lieutenant colonel, senior fellow for National Security at Family Research Council, and the author of 15 books. His latest, "The Final Algorithm," was released in July 2026.


RELATED



Support the work of TWS with a gift to FRC