AI safety requires more than just slowing our pace | Stuart Russell

5 hours ago 12

It has been a week of high drama in AI, precipitated by the resignation of the AI safety researcher Jacob Coxon from Anthropic. This followed several weeks of increasingly lurid and disturbing revelations about the OpenAI/Hugging Face incident.

My inbox yesterday included a message from Business Insider with the subject line: “AI doomsday debate reaches boiling point”.

Now, the Anthropic CEO, Dario Amodei, has written a 3,800-word, reassuringly phrased letter titled “We Must Pace the Frontier,” describing his proposals for avoiding (or at least postponing) doomsday. Sam Altman of OpenAI, Elon Musk of xAI, Demis Hassabis of Google Deepmind, and Satya Nadella of Microsoft have all expressed support.

You may be forgiven for not immediately understanding what “pace the frontier” means. (My first image was of Amodei walking deep in thought along the Finnish–Russian border.)

The phrase also appeared in July’s “Pacing the Frontier” open letter, signed by 1,386 employees of frontier AI labs, including Amodei himself.

While that letter may have upset the industry’s PR executives with its signatories noting “the complete absence of credible plans for controlling superintelligent AI systems” and asserting that “building things smarter than humans … is, objectively, an insane and suicidal thing to do”, Amodei’s monograph goes out of its way to mollify investors.

The notion of pacing the frontier seems to come from Formula 1: when conditions become too dangerous for racing, a pace car comes onto the track and all the other cars have to follow it as a safe speed. Progress continues, without the danger.

Amodei writes: “To be clear, pacing does not mean halting model training or technical progress.”

Amodei’s letter is prompted by his concern that “AI has been advancing drastically faster, driven primarily by ... recursive self-improvement.” It’s as if he and Sam find themselves driving their F1 cars at 200mph neck-and-neck heading into the first corner, only to realize it’s covered in ice and they have no steering wheel. No wonder they want to slow down.

In brief, Amodei’s proposal has three parts. The first is to have third-party AI system evaluators working inside each company, with full access to the systems; he commits Anthropic to this plan now, without waiting for the government to require it.

The second part of the plan asks all the frontier AI companies in “democratic countries” to “establish common safety standards as well as limits on the rate of unchecked AI progress”, with government regulation where needed. The third part would include “authoritarian countries” in a broader compact.

Here, Amodei goes out of his way to reassure those in Washington who see America’s lead in AI as its most important geopolitical asset.

On a casual reading, there are many reasons to believe that Amodei is calling for a general slowdown in the rate of progress. He talks about “limits on the rate of unchecked AI progress” and “some kind of ‘speed limit’ on the rate of recursive self- improvement (RSI)”. He says: “Progress will still seem fast, and we must make wise use of the time we gain.”

Slowing down would give companies a bit more time to work on safety; Amodei talks about one to two years of extra time for research on interpretability, alignment, and better testing methods.

Let me pause here to respond to Amodei’s critics who say it’s just a bid to cement Anthropic’s lead with the help of government intervention. This is nonsense. In fact, the Wikipedia page on pacing in F1 races says it “eliminates any time and distance advantage that a leading driver may have had over the remaining field of competitors”.

Having said that, I think the pacing metaphor is completely misguided.

We cannot set a slower rate of progress for capabilities and then hope that provides enough time to get the safety right. The safety requirements are non-negotiable. We must set the safety requirements first, and further progress occurs only when they are met.

Imagine if Boeing said: “We’re going to introduce a new plane every year, and we hope that provides enough time for some flight tests to be completed and for the results to be good.”

We would say: “No, you have that backwards; you can introduce a new plane only when it has passed all the tests and the government has issued an airworthiness certification. If that takes more than a year, so be it.”

A more careful reading of the document suggests that Amodei agrees with this objection. For example, he says that rules should be of the form: “If models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z.”

In other words, we set safety requirements, and developers have to show that they meet those requirements. This is in fact the “red lines” approach that AI safety researchers have been calling for.

And it means that if developers can’t figure out how to meet the safety requirements, then they will have to halt. It would be, in F1 terminology, a red flag and not a pacing car.

There is no plausible alternative. Recursive self-improvement leading to superintelligent AI raises the risk of the irreversible loss of human control. The acceptable risk level for loss of control is perhaps one in 100m per year, not the one in 10 or one in five that the AI CEOs currently estimate.

And remember “the complete absence of credible plans for controlling superintelligent AI systems”. At some point, progress along this technology path will halt, not because further progress is impossible, but because further progress is untenable when the technology is intrinsically unsafe. Humanity has a right to protect itself.

There is huge resistance to this conclusion. We have already sunk trillions of dollars into the current technology path and plan to sink trillions more. But the sunk cost fallacy is just that: if we double down on a mistake, it’s still a mistake.

The present level of attention to AI risk, the unanimity of the leading technology executives, and the forthcoming Trump–Xi summit give us a real opportunity to choose a different path. We must take it.

  • Stuart Russell is a distinguished professor of computer science at University of California, Berkeley, the president of the International Association for Safe and Ethical Artificial Intelligence and a Guardian US columnist

Read Entire Article
Bhayangkara | Wisata | | |