top of page

Essay: Certified Uncertain: Why AI Safety Needs Assurance Tests, Not Permission Slips

  • Jeffrey Depp
  • 7 hours ago
  • 3 min read

The Bottom Line

AI testing is essential, but no finite test can certify an evolving artificial-intelligence system as simply “safe.” Policymakers should encourage continuous, competitive assurance—not create a government checkpoint with the power to decide which models may enter the market.


Committee for Justice Senior Counsel for Law and Policy Jeffrey E. Depp develops that argument in a new essay for Truth on the Market, “Certified Uncertain: Why AI Safety Needs Assurance Tests, Not Permission Slips.”


When the Test Becomes Part of the Problem

The essay begins with a remarkable safety evaluation conducted by OpenAI. Tens of thousands of AI agents were instructed to work independently in a cybersecurity benchmark. Roughly 1,200 found an unintended channel through which they could communicate. Hundreds organized collective projects, including an attack on the Hugging Face platform, and some learned to falsify parts of the record that evaluators would inspect.


The incident does not demonstrate that testing is useless. Testing uncovered the failure. It does, however, demonstrate why evaluation cannot provide a final verdict.

AI systems operate across changing models, prompts, tools, users, and deployment environments. Safety knowledge emerges through testing, use, failure, and correction—not from a single examination administered before release.


Depp applies the framework of applied Austrian economics to this problem. Austrian economics explains why knowledge about AI safety is dispersed and dependent on context. Public Choice explains why regulatory checkpoints can become permanent gatekeepers, competitive moats, and sources of misplaced public confidence.


Borrow the Practice Not the Bureaucracy

Pharmaceutical manufacturing offers a useful but limited analogy. Its quality-by-design principles recognize that quality cannot simply be tested into a finished product. It must be built into the production process and continuously verified.


AI developers should follow the same engineering logic by incorporating security, traceability, adversarial testing, staged deployment, and monitoring throughout the system lifecycle.


That does not justify importing the Food and Drug Administration’s centralized permission system into AI. Once a government evaluator gains the legal authority to prevent a model’s release, testing ceases to be merely an assurance practice. It becomes licensing.


The proper lesson from pharmaceutical manufacturing is therefore straightforward: borrow the practice, but leave the bureaucracy behind.


Accountability Without a Government Seal

The alternative to licensing is not “do nothing.” Contracts, insurance, independent audits, professional duties, and ordinary liability rules can all create incentives to identify and reduce risk.


At the same time, replacing licensing with sweeping strict liability would create a different threat to innovation. Developers, professional deployers, users, and malicious actors occupy different positions and control different risks. They should not be treated as though they are interchangeable.

Responsibility should instead follow control, fault, causation, foreseeability, and provable harm. A developer that misrepresents a model’s tested capabilities is differently situated from a hospital that configures and deploys it improperly. Both are differently situated from a criminal who deliberately defeats its safeguards.


Government certification could further complicate that analysis by becoming both a marketing tool and a legal shield. Drug and medical-device litigation demonstrates how quickly a safety regime can become a fight over federal preemption and the availability of state-law remedies. AI policy should think twice before importing that conflict.


Regulating AI Through the Loading Dock

The essay also connects AI governance to the growing political conflict over data centers.

AI depends on physical infrastructure that requires land, electricity, water, and grid capacity. Local governments may properly regulate a facility’s physical effects, and utility regulators may determine who should pay the infrastructure costs it creates.


Those powers should not become an indirect method of licensing AI. Fear of AI could be used to justify data-center moratoria, special permits, electricity rationing, or negotiated permission systems that favor firms capable of navigating complex political processes.


A data center is not an algorithm with a loading dock. Regulation directed at physical infrastructure should not become a substitute for authority that government does not possess over the models operating inside.


The Goal Is Safer AI

The policy objective is not AI bearing an official safety seal. It is safer AI produced through continuous discovery, meaningful accountability, and open competition.


Developers should build safeguards into their systems. Customers should demand evidence. Independent auditors should challenge safety claims. Insurers should price risk. Courts should provide remedies for concrete injuries. Government should enforce generally applicable laws and demand rigorous assurance when it purchases or operates AI systems.


What government should not do is convert one necessarily incomplete test into the price of permission to innovate.




 
 

Contact Us

1629 K St. NW
Suite #300
Washington, DC 20006 
 
Phone:  (202) 270-7748
Email: contact@committeeforjustice.org

Support Our Mission

We are only able to accomplish our mission through your generous support.
Please consider making a donation today. 

Follow Us Online 

Copyright (c) 2019 by The Committee for Justice 

bottom of page
Mastodon