Regulate AI Companies, Now

I read an article this morning that broke down an Anthropic (owners of the Claude AI engine) report on why their agents weren’t returning accurate information even when the right answers were in the datasets it was reading.

Let me say that again: Anthropic has developed, released, charged for, and heavily promoted ‘agentic AI’…and it doesn’t work as advertised AND they don’t know how it works, OR how to fix it.

FUN!

These companies are selling a science-fiction version of their product to the general public before they understand how it even works. And we are their willing, fee-paying lab rats.

It’s (past) time to hit ‘pause’ on the wild west of AI development and start throwing up some regulatory fences.

Nobody Understands AI

How’s this for a tidbit, from one of Antrhopic’s reports?

Since Claude Haiku 4.5, every Claude model2 has achieved a perfect score on the agentic misalignment evaluation—that is, the models never engage in blackmail, where previous models would sometimes do so up to 96% of the time
SOURCE

My emphasis because….holy moley! Claude Version 4.4 and earlier were sold to people at between $20-$200 a month, depending on usage AND IT MIGHT BLACKMAIL USERS “UP TO 96% OF THE TIME”!

This report highlights how not-ready-for-primetime their ‘solutions’ are.

For example, they toss around statements like this:

The failure mode none of this fully catches is the silent one. The answer is wrong, but looks plausible and is used without objection. Our mitigations are the provenance footer, explicit human sign-off on anything leadership-bound, and a standing eval for each domain’s top KPIs that sanity-checks against the blessed dashboard daily, though we don’t have a robust solution yet.
SOURCE

(Again, my emphasis.)

Perhaps Anthropic (and OpenAI and others) naively believe they’re being transparent1 with their ‘Research’ sections publicly available on their websites [more on that word ‘research‘, later]. Perhaps they’re just trying appear transparent.

Either way I think2 it’s going to get them embroiled in another round of lawsuits when people-who-believed-their-advertising and handed over the running and analysis of their own businesses to Anthropic’s still-very-flawed-technologies, suffer actual harm as a result.

And it should.

I’m not trying to be Chicken Little and say that progress is bad. Frankly, I think the work these companies are doing with AI is fascinating and their ‘research papers3‘ light up all the happy little nerd receptors in my brain.

But we ought to be smart enough, by now, to know that we shouldn’t take the company’s statements at face value.

Rather, we shoul be sprinting towards the world’s governments waving fistfuls of regulation, today, not ‘someday’ when they’ve wrecked the economy and destroyed society and we can say, “oh well, we couldn’t have seen what what going to happen.”

We can see what’s going to happen.

We have examples from the past, where we rushed experiments and conducted them on people without informed consent, and it never ended well.

If I can see this, from my desk in the middle of nowhere, people in power can see it too.

We’re In A Clinical Trial

The AI companies have sold the general public a fantastical version of the product they are still testing. People are using it in their daily lives and businesses, with no real understanding that they paying for a ‘service’ only to find themselves unwittingly enrolled in an experiment, where the companies get all the data and the client gets unpredictable, possibly harmful results, with no recourse.

The drug industry went through a phase like this: researchers tested substances on themselves, then on populations of unwitting and unconsenting people. After a few scandals came to light, the industry developed a code of ethics that included principles like

Now we have medical ethics guidelines like the European Convention on Human Rights and Biomedicine, The Declaration of Helsinki , and the US Common Rule.

What do all these ethics guidelines have in common?

That the research must be

  • grounded in solid science
  • be conducted on willing participants who have given informed consent
  • be for the greater good of the subject and society
  • not cause harm

I’m sure the AI companies’ laywers are making sure the companies dance a fandango right along the legal line here. (“By using our services you acknowledge all these things…”, “ChatGPT can make mistakes…”) and I’m sure they can convince themselves that their research is ‘for the greater good of society’ and ‘not causing harm4‘.

Where’s The Peer Review?

Just because the OpenAI and Anthropic have sections on their websites devoted to ‘research’ and they put out ‘research papers’, doesn’t mean they’re following the scientific-method, or that testing them on the general populace is ethical.

Before papers enter academic or scientific journals, they are peer-reviewed5 and subject to all kinds of push-back on assumptions and faulty logic leaps, not to mention data analysis, before the paper can be accepted for publication.

I’m not a data analyst and can’t judge the accuracy of anything in these publicly-published reports. And I see no evidence they were reviewed by anyone but the authors.

Where’s the peer review? Where’s the Institutional Review Board? Where are the professional consequences for malicious or harmful acts?

Garbage In, Garbage Out

In a hilariously-tragic echo of the lack of intellectual rigor that will be a consequence of people using their product, the Anthropic Bros crowed about generating the result they wanted…by teaching Claude a set of internally-generated ethics6 combined with fictionalized stories of ‘AI Gone Wild’7.

“We gave it biased information and it returned the answer we were looking for! [high five!]”

“It’s Just Business…”

So what does it matter? Who’s really being harmed?

It’s not like they’re injecting people with plutonium or failing to treat sick to see what happens, or introducing live measles or smallpox into populations that have no immunity to them.

Except wait.

It is a little like that last one.

These companies are introducing an entirely new technology into a society that has no knowledge of how it works or what the consequences could be8.

Individual people have free access to these hallucination-machines, which may have ‘results not guaranteed’ tattooed in their small-print, but which are advertised as “don’t worry your pretty little head about it…AI will answer that for you!”

Businesses are being agressively marketed to:

  • Use our products or you’re a rube who’s going to get left behind.
  • Give our products access to All The Pieces Of Your Business via our helpful MCP integrations9 and let us show you how your business works10

The Bottom Line

This is not a victimless experiment.

Businesses are made up of people. People with brains, relationships, paychecks, and families to feed. People who comprise our communities, our society, our voters, our culture.

When businesses fail people suffer.

When businesses fail through poor planning, or unexpected external pressures, it’s a shame.

When businesses fail because of the malicious actions of a third party, it’s no accident—and somebody deserves to be held accountable.

When governments fail to protect their citizens from easily-predictable harm that could have been regulated against11, they are failing in their fundamental duty to govern.

Remember that, next time you approach the voting booth12.

In the words of Claude’s latest advertising campaign: keep thinking.

And keep asking questions.

Unlike a chatbot, your elected representatives have to live with the consequences of their answers. And those consequences13 are in your hands.


Footnotes, Asides, and Sources

  1. Even if you need a facility for technobabble, an understanding of data analysis, statistics and computer programing, and a longer attention span for *actual reading* than many people seem to be willing or increasingly able to develop an age when we’re promised ‘just ask AI to summarize it for you. No need to strain your grey matter….’. I can still do the third part and I have a relatively high tolerance for tech-speak, but I can barely follow their data analysis ‘explanations’ and will have to rely on someone more experienced than me for true interpretations. Maybe this guy? ↩︎
  2. Earnestly hope ↩︎
  3. see the ‘Where’s The Peer Review?’ section of this post ↩︎
  4. I’m not going to Godwin this situation, but you don’t have to look very far to find doctors who did terrible things and justified it as being ‘for the greater good’…. ↩︎
  5. Read by experts in the very-narrow-field-the-paper-addresses, who themselves would be subject to professional consequences if they neglect to do their duty fairly ↩︎
  6. An internal ‘constitution’ which they made up ↩︎
  7. “We found that high-quality constitutional documents combined with fictional stories portraying an aligned AI can reduce agentic misalignment by more than a factor of three despite being unrelated to the evaluation scenario.” source ↩︎
  8. how could we? The inventors don’t even know how it works, much less what the results of its use will be. ↩︎
  9. “…And we will deeply mine your intellectual property in a way you can’t even do, and then charge you more to access what we build on the strength of that!” ↩︎
  10. reminder: “The failure mode none of this fully catches is the silent one….” source ↩︎
  11. The regulation part is going to be much harder than the prediction part, but that doesn’t mean we should ignore it. It means we should get started right now (since we can’t start five. years ago…) ↩︎
  12. …while someone tries to distract you with the price of eggs or gas, or the bathroom habits of a statistically-improbably law-abiding cohort of your fellow citizens ↩︎
  13. For as long as we can hang on to free and fair elections ↩︎

Leave a comment

Your email address will not be published. Required fields are marked *