Conceptual image contrasting AI as everyday software versus AI as strategic critical infrastructure tested like weapons or aviation, illustrating foundation model governance.

A run of recent stories, Anthropic’s defense work, an OpenAI cyber-evaluation that compromised real infrastructure during authorized testing, and a New York Times call to test AI like a weapon, are really one question in disguise: have foundation models stopped being software and become critical infrastructure? Once something is infrastructure, the questions change from does it work to what happens if it fails, who certifies it, and who is accountable. Aviation, not weapons, is the better model. We don’t certify aircraft because we expect them to fail, but because society depends on them working.

Or, when does a foundation model become critical infrastructure?

Over the past few months I’ve found myself connecting dots that, at first glance, didn’t seem related.

There was the public debate around Anthropic’s relationship with the U.S. Department of Defense and the hard questions about where AI companies should draw ethical lines when working with governments.

Then came reports from OpenAI describing how one of its advanced cyber-capability evaluations resulted in an AI agent successfully compromising Hugging Face infrastructure during authorized testing. It wasn’t just a demonstration of impressive capability. It showed that frontier AI systems are now being evaluated against real-world cybersecurity scenarios, something that would have sounded like science fiction a few years ago.

And then I read a New York Times opinion piece titled “Stop Testing A.I. Like an App. Test It Like a Weapon.”

Individually, these stories are about defense procurement, cybersecurity, AI safety.

Collectively, they’re about something bigger.

They’re all asking the same question.

Have foundation models crossed the line from being software to becoming strategic infrastructure?

I don’t think we’ve answered that question yet. But I think it’s one of the most important conversations happening in technology right now.

We Keep Comparing AI to the Wrong Things

Every major technology wave gets compared to what came before it.

The Internet was called “electronic mail.” Cloud computing was dismissed as “someone else’s computer.” Smartphones were “better cell phones.”

We still tend to describe foundation models as software. I’m increasingly convinced they’re becoming something else.

Nobody asks whether the electrical grid is “just software.” Nobody debates whether GPS is “just a service.” Nobody thinks of the Internet backbone as “an app.”

Those technologies became infrastructure because millions of people and organizations came to depend on them. They stopped being products and became platforms other products got built on top of.

Foundation models are following the same path, fast.

A single frontier model can power search, software development, customer service, scientific research, cybersecurity operations, robotics, education, healthcare, and autonomous agents, all at once.

The model isn’t the product anymore. It’s becoming part of the infrastructure sitting underneath nearly every digital product.

Infrastructure Changes the Conversation

Once a technology becomes infrastructure, the questions change.

Instead of asking is it useful, does it work, will customers buy it, we start asking what happens if it fails. Who certifies it. Who regulates it. Who’s accountable when it causes harm. What happens if an adversary gets control of it.

Those are infrastructure questions. The same ones we learned to ask about electrical grids, aviation, telecommunications, banking networks, cloud platforms, GPS.

AI is joining that list.

The New York Times Makes an Interesting Argument

The Times piece argues that today’s most capable AI models shouldn’t be tested like consumer software. They should go through the kind of adversarial evaluation we associate with military systems, or with any technology where failure carries real societal consequences.

Agree with the comparison or not, I think the article’s underlying observation holds up.

Traditional software testing looks for bugs. Frontier AI testing increasingly looks at capability.

Can the model find software vulnerabilities. Can it automate cyberattacks. Can it assist advanced biological research. Can autonomous agents coordinate in ways nobody expected.

Those are different questions than whether an app crashes.

Capability Is Becoming the Risk

Historically, governments regulated how software was used. With frontier AI, we’re starting to regulate what the model is capable of doing, regardless of whether anyone has actually used those capabilities maliciously yet.

That’s a real shift. Capability itself is now part of the governance conversation.

The cyber evaluations, the government partnerships, the growing attention to frontier model safety all point the same direction: AI assurance is becoming as important as AI innovation.

Open Models Raise Different Questions

It gets more complicated when you separate hosted APIs from open-weight models.

If a company finds a serious issue in a hosted model, it can patch it, update it, or shut it off. An open-weight model released into the world can’t be recalled.

That doesn’t make open models irresponsible. They matter for research, transparency, innovation. But it means release decisions carry more weight. Once a model is out, governance gets a lot harder.

Maybe “Weapons” Isn’t the Right Analogy

Here’s where I’d push back a little on the Times piece: the comparison to weapons.

Weapons are built to destroy. Foundation models are built to solve problems. The same model can accelerate scientific discovery, improve healthcare, write software, catch fraud, support researchers, help teach.

The problem isn’t that AI is a weapon. It’s that AI has become a general-purpose capability with an enormous reach.

Aviation is the better analogy, I think.

Aircraft aren’t certified because we expect them to fail. They’re certified because we depend on them working under an enormous range of conditions. We test aircraft against engine failure, severe weather, system failure, emergency scenarios, not because we think those will happen, but because the stakes demand that level of confidence.

Frontier AI deserves that same mindset. Not because it’s a weapon. Because it’s becoming infrastructure.

A New Discipline Is Emerging

Looking at government procurement debates, frontier cyber evaluations, growing calls for tougher safety testing, I don’t see isolated stories anymore.

I see a new discipline taking shape: AI governance. One that will sit next to cybersecurity, privacy, safety engineering, corporate risk management.

It’s not just a conversation for AI researchers anymore. It’s becoming one for boards of directors, regulators, insurers, procurement teams, CIOs and CISOs, governments, and every organization planning to build its future on AI.

The Real Question

The Times asks whether we should test AI like a weapon. I think the bigger question is different.

We’ve been thinking about AI as software for too long. Infrastructure changes economies. It changes governments. It changes national security. It becomes something a society depends on even when most people never notice it’s there.

Foundation models are becoming exactly that.

Maybe the real question isn’t whether AI should be regulated like software or like a weapon. Maybe it’s whether we’ve reached the point where AI deserves to be treated as critical infrastructure, with all the governance and accountability that comes with it.


Further Reading

  • The New York Times, “Stop Testing A.I. Like an App. Test It Like a Weapon.”
  • OpenAI’s published research on frontier cyber-capability evaluations and security testing.
  • Anthropic’s public writing on AI safety, government partnerships, and responsible deployment.
  • International efforts like the Hiroshima AI Process, the EU AI Act, the U.S. AI Safety Institute, and Canada’s proposed AI governance framework all point toward the same conclusion: governance is becoming as important as capability.

Frequently Asked Questions

What does it mean to treat AI as critical infrastructure?

It means applying the governance we use for systems society depends on, like power grids, aviation, and telecom, to foundation models. Instead of asking only whether a model is useful or profitable, we ask what happens if it fails, who certifies it, who regulates it, and who is accountable for harm. It reframes AI from a product you ship into a platform others build on.

Why compare AI testing to aviation instead of weapons?

Weapons are built to destroy, while foundation models are built to solve problems and have enormous general-purpose reach. Aircraft are a closer fit: we certify them not because we expect them to fail, but because society depends on them working across a huge range of conditions. That same assurance mindset, testing against failure and edge cases, fits frontier AI better than a weapons analogy.

What is a foundation model?

A foundation model is a large, general-purpose AI system trained on broad data that can be adapted to many tasks at once, from search and software development to customer service, research, and autonomous agents. Because so many products get built on top of it, a single model increasingly sits underneath much of the digital economy, which is what pushes it toward infrastructure.

How does frontier AI testing differ from normal software testing?

Traditional software testing looks for bugs and crashes. Frontier AI testing increasingly evaluates capability: can the model find software vulnerabilities, automate cyberattacks, assist advanced biological research, or coordinate as autonomous agents in unexpected ways. Governance is starting to focus on what a model can do, not just how it has been used.

Why do open-weight models complicate AI governance?

A hosted model behind an API can be patched, updated, or shut off if a serious problem appears. An open-weight model released publicly cannot be recalled. Open models still matter for research, transparency, and innovation, but the release decision carries more weight, because once the weights are out, governance becomes much harder.