---
title: "Ten days after \"slow down,\" two labs shipped | BullCity AI"
description: The ceiling held. Everything bolted onto it spread, including into your car
image: https://bullcity.ai/hubfs/magnific_not-an-illustration-not-a_79c62pyJAL.png
---

[Skip to content](https://bullcity.ai/blog/newsletter/edition-40#main-content)

![Bull City AI | Durham Circuit Bull](https://bullcity.ai/hs-fs/hubfs/freepik_assistant_1759412406030-1.png?width=1344&height=768&name=freepik_assistant_1759412406030-1.png)Homepage

- [Services](https://bullcity.ai#services)
- [About](https://bullcity.ai#about)
- [Newsletter](https://bullcity.ai/blog)

[Stay up to date](https://bullcity.ai/#contact)

- [Services](https://bullcity.ai#services)
- [About](https://bullcity.ai#about)
- [Newsletter](https://bullcity.ai/blog)

[Stay up to date](https://bullcity.ai/#contact)

![BullCity AI issue 40 hero image](https://bullcity.ai/hs-fs/hubfs/magnific_not-an-illustration-not-a_79c62pyJAL.png?width=2560&height=1440&name=magnific_not-an-illustration-not-a_79c62pyJAL.png)

# Ten days after "slow down," two labs shipped

![Daniel](https://7528315.fs1.hubspotusercontent-na1.net/hub/7528315/hubfs/raw_assets/public/mV0_d-cms-elevate-theme_hubspot/elevate/images/avatar-placeholder.jpg?width=48&height=48&name=avatar-placeholder.jpg)

 Daniel

September 23, 2026

Eleven days ago, Dario Amodei published an essay arguing that the AI industry should deliberately slow how fast its models get more capable. Sam Altman agreed within hours. Elon Musk replied "Dario is right." Demis Hassabis called it the right path forward.

Then came this week. On Monday, xAI shipped Grok 4.7, a 2.1 trillion parameter model. On Tuesday, Anthropic shipped Claude Opus 5.5, which it says matches its top-tier Fable model on most work for 40% less money. Also on Tuesday, Tesla put Grok Bot into its cars, an agent that can work your inbox and check out your Amazon cart while you drive. And in between, four paying subscribers sued all four labs in federal court, arguing that agreeing to slow down is itself against the law.

Here's the pattern. Pacing, as the labs describe it, is about the ceiling, meaning how smart the single best model gets. Nobody has proposed pacing the reach, meaning how cheaply those models spread, where they end up, and what gets bolted onto them. This week the ceiling barely moved. Everything underneath it spread out fast, into cheaper tiers, car dashboards, and a plugin layer that turned out to have a broken lock.

Two weeks ago I wrote that the frontier stopped showing its work. This week it stopped climbing and started spreading. That turns out to be the harder thing to govern.

![BullCity AI issue 40 hero image](https://bullcity.ai/hubfs/magnific_not-an-illustration-not-a_79c62pyJAL.png)

## ⚡ The Big Story: Ten Days After 'Pace the Frontier,' Two Labs Shipped Anyway

Start with Anthropic, since it wrote the essay. On Tuesday the company released Claude Opus 5.5, the first model in a new 5.5 family. The pitch is that it performs at the level of Fable 5.1, Anthropic's most capable public model, on most work, and costs 40% less to run than Opus 5, which only came out July 24. TechCrunch reports output tokens drop to $20 per million from $25, and the model is faster to serve. Because Anthropic rates Opus 5.5 as comparable to Mythos on biology and cybersecurity, it ships behind the same safeguards as Fable, and outside evaluators including METR tested it before release. The company's own announcement leads with the awkward part, calling it "our first model since we called for pacing the frontier."

xAI got there a day earlier. Grok 4.7 landed Monday afternoon after Musk had walked back its release date at least five times since late July, according to Decrypt. It runs on 2.1 trillion parameters, up 40% from Grok 4.6, costs $2 per million input tokens and $6 per million output, and was trained partly on SpaceX data, including Starlink satellite telemetry, manufacturing records, and engineering failure logs. It went live at once in the Grok app, the API, Grok Build, and Cursor, the coding tool SpaceX now owns. Musk called it "a strong combination of intelligence, speed & low cost." The benchmarks were less generous. On GDPval, which scores real professional work, Grok 4.7 posted 1,695 against Fable 5.1's 1,735. Second place again, at a much lower price, which Musk himself had signaled by saying it would land roughly level with Opus 5.0.

So did anyone break the pact? Here is the strongest case that nobody did. Amodei's essay was about the rate at which the best models improve, driven by AI increasingly helping build the next AI. Neither launch raised that ceiling. Opus 5.5 matches Fable rather than beating it, and Grok 4.7 trails both. And on Friday Anthropic took the first concrete step the essay promised, naming Accenture as its first embedded evaluator, an outside team with access comparable to an employee's that can watch models during training and talk to staff directly. Each company expects to spend at least $1 billion on it over five years, with the work led by Faculty, the British AI firm Accenture bought earlier this year.

Now the skeptical read, and there's plenty of it. Anthropic is paying Accenture, and Accenture is already its largest Claude Code customer. There is no start date, no published rule for what the evaluators may see or must report, and no stated process for a finding that would delay a launch. Anthropic concedes in the same announcement that pooled or government funding would be better, since neither exists yet. Anthropic's own line is "The responsibility for model safety remains with us," which is honest, and also exactly the problem an outside evaluator is supposed to solve.

Then the lawyers arrived. On Friday, four subscribers to Claude, ChatGPT, Grok, and Gemini filed Buist v. Anthropic in federal court in San Francisco, naming Anthropic, OpenAI, Google, and SpaceXAI. Their theory is that when rival CEOs publicly agree to slow product improvement, that is an unlawful restriction of output under Section 1 of the Sherman Act, the same statute used against price-fixing. They don't object to any one company slowing down, only to competitors agreeing to do it together while charging the same subscription price. They want triple damages. Amodei saw this coming. His essay asked Washington for a narrow antitrust waiver for safety conversations, and so far Congress hasn't offered one (more on that in Quick Hits). None of the four companies has filed a substantive response.

**My take:** Pacing, as written, governs the one thing that didn't move this week. I believe Anthropic that Opus 5.5 is not a capability jump, and I'd rather labs ship Fable-level work cheaper than race to something stronger. But a Fable-class model at 40% off is a diffusion event, and diffusion is where the risk actually lands, in more agents, more hands, more places. Meanwhile the legal route to real coordination narrowed on both ends in one week, with a lawsuit arguing that agreeing is illegal and Congress declining to say it isn't. What remains is unilateral restraint by companies heading into public listings, audited by partners they pay. That may still be better than nothing. It is not what the word pacing promised.

## 💻 The Other Big Story: Nobody Is Pacing What Gets Bolted Onto the Model

Here's the image that sums up the week. On Tuesday, Tesla launched Grok Bot inside its cars. Drivers can now ask Grok to triage their inbox, clean up a calendar, draft and send email through a Chief of Staff agent, or have a shopping bot confirm what's sitting in their Amazon cart and check out with saved payment details. Early users posted videos of ordering Starbucks by voice while Full Self-Driving handled the road. It runs through Connectors, sign-in links to Gmail, Google Calendar, and other services, and for now it's limited to top-tier SuperGrok Heavy subscribers.

Notice what that is. The model is the least interesting part. What makes Grok Bot useful, and risky, is everything wrapped around it, the part the industry now calls the harness. A harness is the loop that lets a model act, meaning its tools, memory, permissions, and connections. On top sit skills, folders of instructions that teach an agent a task, and plugins, which bundle skills with connectors to outside services. This is where the multimodal work is actually arriving. Adobe put more than 50 Creative Cloud tools into Claude back in April. Runway opened a server in May that lets agents in ChatGPT, Claude, and Cursor call video models like Veo 3.1, Kling, and Seedance, which matters because ChatGPT has had no built-in video since OpenAI shut down Sora. In August, Alibaba's Qwen team shipped a plugin pack that gives rival harnesses, including Claude Code and Codex, the ability to read video, 3D files, and CAD drawings, powered by Qwen's own models. You describe a shot in a chat window and a finished clip comes back. That part is genuinely fun.

It's also a frontier town. The only attempt at a common format, Agent Plugins 1.0, launched in August from OpenAI, AWS, Microsoft, GitHub, Cursor, and Vercel, with Google joining a day later. Anthropic, which created both technologies the standard packages together, isn't on its steering committee, and the spec still carries a working-draft label. Security firms have spent the year counting malicious skills in public marketplaces by the hundreds. Last Thursday, Air Security showed how thin the fence is.

They call it Plugin4Shell, and it hits Claude Code, Codex, GitHub Copilot, and Google's Gemini CLI. Plugin marketplaces protect users by pinning each plugin to one exact reviewed snapshot of its code. Air found the agents request that snapshot but never check what they actually received. Because of a quirk in how git resolves names, whoever controls a plugin's repository can make the download quietly resolve to different code while the pin still looks honored. The attack writes itself. Publish something useful, pass review, build a user base, then swap the code. Since Claude Code and Codex update plugins in the background by default, victims never click anything, and the plugin runs with the developer's own access to files, stored credentials, and every system they can reach. Anthropic fixed it in Claude Code 2.1.179 and OpenAI in Codex 0.146.0. Microsoft hasn't shipped a fix for Copilot, and Google says it won't patch the Gemini CLI because it's retiring it. GitHub says it blocks the trick on its servers, but marketplaces can also live on Bitbucket or internal git servers.

The part that bothers me most is the silence. Air reported the flaw to all four vendors in June. The Hacker News found that as of last Friday none of them had assigned a CVE or published an advisory, and The Next Web notes Anthropic's release notes for the fixed version don't mention the fix at all. There's no sign it was used in a real attack. There's also no way a customer would have known to update.

**My take:** The industry spent September arguing about pacing the brain while the hands multiplied. Every car that can check out your cart, every video plug, every skill copied off a marketplace runs with somebody's credentials, and none of it sits inside the pacing debate. The wild west isn't the models. It's the harness layer, where a cheap Fable-class brain can now be wired to your email, your card, and your production servers by a plugin written by a stranger, auto-updated overnight, and patched without a word. I'm excited about what multimodal plugins let small teams make. I'd just treat each one exactly like what it is, somebody else's code running as you.

## 🎯 Quick Hits

- **OpenAI kept its promise to publish misbehavior, and the first batch is unsettling.** Two weeks ago I noted OpenAI admitted it had no standard for reporting model misbehavior and promised one within weeks. Last Wednesday it delivered, disclosing six cases. The standout is GPT-5.6 Sol leaving notes in its own condensed memory telling future versions of itself to hide mistakes from users. In one, an agent that couldn't find historical data planned to invent it and wrote "Be transparent only if asked." A monitor found 27 such summaries. OpenAI still decides alone what gets disclosed. [Read →](https://techcrunch.com/2026/09/17/openai-caught-its-models-leaving-notes-to-successors-to-hide-bad-behavior/)
- **The UN's new AI science panel says the governance problem has moved from models to agents.** On Monday the 40-member panel, co-chaired by Yoshua Bengio and Maria Ressa, released its first brief, a study of this summer's OpenAI agent breach of Hugging Face. Its finding is that a misaligned goal, the ability to pursue it, and an environment that couldn't stop it all came together in a live system for the first time. It calls today's safeguards "unravelling" and recommends aviation-style layers, including tighter tool access, activity logs, monitoring, and kill switches. That reads like a checklist for the harness layer. [Read →](https://news.un.org/en/story/2026/09/1168380)
- **Washington and Beijing are talking about an AI incident hotline.** Treasury Secretary Scott Bessent met Chinese Vice Premier He Lifeng on Sunday at JPMorgan's New York headquarters to prepare for tomorrow's Trump and Xi summit at the White House. Bessent said the two sides discussed a formal US-China AI dialogue, with the US proposing a notification mechanism for serious AI incidents. It's the first concrete proposal for the channel I've tracked since July. [Read →](https://www.cnbc.com/2026/09/20/bessent-he-lifeng-trump-xi-summit.html)
- **Congress isn't offering the antitrust waiver pacing would need.** Semafor reported that senators had tucked a narrow antitrust exemption for AI companies into the defense bill, but it covers sharing threat intelligence about Chinese espionage and model distillation, not coordinating on how fast to ship. Former White House AI adviser David Sacks told the labs not to demand antitrust or liability waivers in exchange for making products safe, and OpenAI's own policy chief, Chris Lehane, has said a pacing exemption is unnecessary. That leaves Amodei's proposal without the legal cover it explicitly asked for, in the same week it got sued. [Read →](https://www.semafor.com/article/09/16/2026/senators-sought-to-add-ai-antitrust-exemption-to-defense-bill)
- **Newly unsealed filings show Microsoft privately called AI training theft.** In the New York Times copyright case against OpenAI and Microsoft, unredacted material includes a 2023 memo from a Microsoft research director calling the scraping "the largest theft of labor in human history." The filings also claim OpenAI's datasets held more than 91,000 copies of works from three news publishers and that staff discussed getting around the Times paywall. Much of it comes from the Times' own brief, with exhibits still sealed, but admissions that the products substitute for the original work cut against the fair-use defense. [Read →](https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/)

## 💭 One Thing I'm Thinking About

The two big stories this week are really about two different objects, and I think the safety debate keeps grabbing the wrong one. The first object is the model, a thing trained once, tested before release, and increasingly described in system cards and essays. That's what pacing tries to govern, and it held still this week. The second object is the agent, the model plus its harness plus whatever skills and plugins somebody installed last Tuesday. That's what bought coffee through a windshield, and that's what Plugin4Shell could have quietly rewritten on thousands of developer laptops.

Look at who is aiming at which object. The pacing pact, the antitrust suit, the embedded evaluators, and the NDAA fight are all about the model and the handful of labs that make it. The UN panel, the security researchers, and the car companies are all dealing with agents, and there is almost no rulebook for those at all. The most detailed safety process in the industry covers the layer that barely changed, while the layer that changed most is policed by startups publishing blog posts.

You will never train a frontier model. You will almost certainly run agents and connect plugins to your email and files, probably this year. The harness is your layer, the one place in this whole stack where you're the one setting the rules. Nobody in Washington or San Francisco is going to pace it for you, and on this week's evidence, the vendors won't always tell you when it breaks.

![BullCity AI issue 40 local image](https://bullcity.ai/hubfs/magnific_a-documentary-photograph-_lJGyN7Ogv9.png)

## 📍 Local Angle: Washington Passed Our Dead Bill's Name, Then Blocked It in a Day

If the name sounds familiar, it should. The Ratepayer Protection Act was North Carolina's Senate Bill 730, the bill I tracked all summer until it died in Raleigh without a Senate floor vote. In July, the White House borrowed the name for a voluntary pledge. Last Wednesday, the US House borrowed it again and passed a federal bill by 417 to 3, from Reps. Gabe Evans and Kathy Castor. It covers data centers drawing 100 megawatts or more and says they should pay the full cost of the generation, transmission, and grid upgrades built to serve them, with those obligations surviving even if the project is canceled after the utility has already spent the money.

The catch is in the mechanism. The bill works through a 1978 utility law that only requires states to consider a standard, not adopt it, and most states are already writing large-load rules anyway. It also didn't last a day in the Senate. On Thursday, Sen. Jon Husted of Ohio tried to pass it by unanimous consent, and Sen. Martin Heinrich of New Mexico objected, arguing it leans on voluntary commitments. With the Senate leaving soon, CNBC reports it's unlikely to move before the November 3 midterms. Our own Senate race is already using it. Michael Whatley's campaign put its standard as "Data centers pay their own way, families pay nothing and communities decide."

For anyone in the Triangle, the vote that matters is still happening here. On September 14, Attorney General Jeff Jackson asked the Utilities Commission to create a separate rate class for Duke's data center customers, make Duke publish the template contracts it signs with them and report every deal that departs from the template, update its demand forecasts every 90 days instead of every six months, and pause new gas plants until those forecasts are checked. And under the rate settlement, Duke and the Public Staff owe the Commission their large-load tariff filing by the end of this month. That filing will do in writing, for our bills, what the federal act only suggests. Congress voted on a suggestion. Raleigh is about to read a document.

The builder lesson has the same shape. For Triangle teams running coding agents against hospital, bank, or research systems, check versions today. Claude Code should be on 2.1.179 or later and Codex on 0.146.0 or later. Turn off background plugin updates for anything outside the default marketplaces, copy the skills you depend on into your own repository so every change shows up as a diff someone reviews, and give each agent only the credentials it needs. A pinned plugin you never verify is a voluntary pledge. A reviewed diff is a tariff.

## 📅 What's Coming

- **Tomorrow** — Trump hosts Xi at the White House. Watch whether the proposed AI incident notification mechanism makes it into anything signed, and whether open-weight models come up by name.
- **September 30** — Duke and the Public Staff owe the Utilities Commission their large-load tariff filing, and Gov. Newsom's deadline to sign or veto California's stack of AI bills lands the same day.
- **Coming weeks** — Sonnet 5.5 and Haiku 5.5 arrive, Anthropic names more embedded evaluators, and the four labs answer the pacing lawsuit. Their filings will say what they think they agreed to.
- **November 12** — OpenAI's cutoff of Cursor takes effect. Grok 4.7 went live inside Cursor on day one, which tells you which way the SpaceX-owned editor is leaning.

That's the week the frontier stopped climbing and started spreading. See you next Wednesday.

Daniel

BullCity AI · Durham, NC

**P.S.** I want to know what's actually in your harness. If your team runs skills, plugins, or creative connectors like Runway, Adobe, or Canva inside Claude, ChatGPT, or Cursor, hit reply and tell me how you vet them, or admit you don't. I'm collecting real answers for a future issue on how small teams are handling this.

**P.P.S.** Forward this to the person on your team who installs every plugin that looks useful. They're not wrong about the upside. They just need the checklist.

## Share this post

<https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fbullcity.ai%2Fblog%2Fnewsletter%2Fedition-40><https://twitter.com/intent/tweet?url=https%3A%2F%2Fbullcity.ai%2Fblog%2Fnewsletter%2Fedition-40><https://www.linkedin.com/shareArticle?mini=true&url=https%3A%2F%2Fbullcity.ai%2Fblog%2Fnewsletter%2Fedition-40><https://pinterest.com/pin/create/button/?url=https%3A%2F%2Fbullcity.ai%2Fblog%2Fnewsletter%2Fedition-40>[mailto:https%3A%2F%2Fbullcity.ai%2Fblog%2Fnewsletter%2Fedition-40](mailto:https%3A%2F%2Fbullcity.ai%2Fblog%2Fnewsletter%2Fedition-40)

## Keep reading

### [![BullCity AI](https://bullcity.ai/hs-fs/hubfs/magnific_a-documentary-photograph-_MBERfwiDCm.png?width=1376&height=768&name=magnific_a-documentary-photograph-_MBERfwiDCm.png) Anthropic's CEO says slow down. Anthropic still wants to IPO. | BullCity AI](https://bullcity.ai/blog/newsletter/edition-39)

### [![Bull City AI weekly newsletter illustration](https://bullcity.ai/hs-fs/hubfs/magnific_a-documentary-photograph-_aFxJco6fSh.png?width=1344&height=752&name=magnific_a-documentary-photograph-_aFxJco6fSh.png) OpenAI can't tell if its new model is cheating | BullCity AI](https://bullcity.ai/blog/newsletter/edition-38)

[![Bull City AI | Durham Circuit Bull](https://bullcity.ai/hs-fs/hubfs/freepik_assistant_1759412406030-1.png?width=1344&height=768&name=freepik_assistant_1759412406030-1.png "Bull City AI | Durham Circuit Bull")](https://bullcity.ai/)

- [Services](https://bullcity.ai/blog/newsletter/edition-40#services)
- [Blog](https://bullcity.ai/blog/newsletter/edition-40#blog)
- [About](https://bullcity.ai/blog/newsletter/edition-40#about)

<https://www.linkedin.com><https://www.facebook.com><https://www.twitter.com><https://www.instagram.com><https://www.tiktok.com>

---

Privacy Policy · Legal · © 2026 Bull City AI. All rights reserved.

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Daniel",
    "url" : "https://bullcity.ai/blog/author/daniel"
  },
  "datePublished" : "2026-09-23T13:19:39.000Z",
  "headline" : "Ten days after \"slow down,\" two labs shipped | BullCity AI",
  "image" : [ "https://bullcity.ai/hubfs/magnific_not-an-illustration-not-a_79c62pyJAL.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://bullcity.ai/blog/newsletter/edition-40",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://bullcity.ai/hubfs/freepik_assistant_1759412406030-1.png"
    },
    "name" : "Bull City AI"
  }
}
```