The Finish Line Will Be Announced Shortly

Anthropic says AI safety requires winning the race. But who defines victory, draws the finish line, and decides when responsibility means surrendering the lead?

Every race needs a starting gun, a finish line, and someone empowered to disqualify the contestants. The AI race appears to have dispensed with the last two, while distributing starting guns to everyone with a data centre.

At POLITICO’s Decoded summit on September 16, Anthropic’s head of public policy, Sarah Heck, explained that “you can’t do safety from second place.” She was discussing America’s lead over China and arguing for chip export controls. She also called for government regulation and rejected industry self-policing. That context matters. Her argument deserves better than a cartoon of a technology executive demanding permission to accelerate through a nursery. POLITICO

Still, it is a remarkable sentence from a company whose identity is so closely associated with safety. Apparently, responsibility has qualifying heats.

The strongest version of the argument is reasonable. A technological lead could give democratic countries influence over standards and room to slow development without surrendering strategic power. Anthropic CEO Dario Amodei has proposed precisely this combination: slower capability advances, embedded external evaluators, coordination among democracies, and attempts at international agreements. Anthropic has committed to admitting embedded evaluators; the broader coordination remains a proposal. This is more substantial than putting a seatbelt icon on the accelerator. Amodei’s proposal

But the proposition still needs interrogating. How much lead buys how much safety? Six months? Two model generations? A warehouse of chips visible from space? If no measurable answer exists, the breathing room could become an indefinitely renewable excuse to keep running.

First, Ms Heck: what exactly are we racing towards?

Better medical research, reliable software, military advantage, profitable automation, and machines capable of improving themselves are different objectives. A system might excel at one and be disastrous at another. Compressing them into a single leaderboard is wonderfully convenient for speeches. It is less useful for deciding whether humanity is doing well.

There is real competition for talent, computing power, customers, and strategic influence. Calling that competition a race, however, smuggles in assumptions: one track, one direction, one winner. It makes asking whether a particular destination is desirable sound like complaining about your running shoes.

And who sets the finish line? If the answer is artificial general intelligence, who certifies its arrival? If the answer is superiority over China, when does superiority become sufficient? An enduring requirement to stay ahead describes a permanent security contest. Nobody gets a medal and goes home. The podium is a treadmill.

The language can also help create the predicament it claims merely to describe. If every laboratory believes restraint hands victory to a less responsible rival, each acquires a moral justification for moving faster. Collectively, they can produce the danger each says it wants to prevent. Everyone is the designated adult; somehow, the car is still going down the stairs.

Then there is the curious relationship between national leadership and corporate success. Heck’s claim concerned the United States; it did not establish that Anthropic itself must win. Those are different propositions. A country could benefit from several capable laboratories, independent researchers, and institutions strong enough to disappoint all of them.

Would Anthropic welcome a safer competitor becoming the standard? Would it support an independent finding that its own newest model should wait while someone else’s proceeds? These are useful questions because they make the principle expensive. Almost anyone can support safety when it arrives with increased market share and complimentary conference catering.

The harder question concerns stopping. What finding would require a delay even if China continued advancing? Who would make that finding, and who could enforce it?

Amodei’s proposal gives external reviewers access and publication rights, subject to specified limits. Those are meaningful commitments. But access, disclosure, and authority to halt development are separate powers. An evaluator can have a desk, a badge, and a deeply troubling report while the training run continues upstairs. The visitor’s pass needs an institutional answer behind it. Proposed evaluator powers

None of this requires assuming bad faith. People can sincerely fear a technology, sincerely believe their organisation is unusually responsible, and sincerely conclude that the best remedy is to give their organisation more influence. Sincerity does not resolve the conflict of interest. It merely ensures that everyone sleeps badly for respectable reasons.

Nor does geopolitical advantage settle technical safety. Leadership might help governments enforce standards or deter misuse. It does not, by itself, make a model controllable. A system pursuing an unintended objective will not become reassuring because its servers are located in a democracy. The flag outside the building is doing very little work inside the mathematics.

There is also the awkward question of everybody else. Who represents countries without a frontier laboratory, workers whose livelihoods may change, or communities asked to host the infrastructure? Are they participants in choosing the destination, or spectators expected to applaud whichever team secures the naming rights to the future?

A credible safety policy would make these questions answerable: defined risks, independent scrutiny, enforceable limits, and evidence that restraint reduces danger. It would distinguish the benefits of leadership from the hazards of speed. It would also explain what happens when those interests collide, since that is precisely when a policy earns its keep.

Anthropic may have serious answers. Its willingness to discuss regulation and outside evaluation gives us something concrete to examine. But the most revealing test of being the responsible contestant is what happens when responsibility threatens the lead.

Before asking us to cheer, perhaps show us the rulebook. Especially the page explaining who can stop the race.

No comments yet