# Agent Wars!
*The AI Daily Brief — Tuesday, 2026-09-22 · https://aidailybrief.ai/e/2026-09-22*

**The agent wars have begun — and the aggregators are getting aggregated.**

Muse's surge past ChatGPT to the top of the App Store proved personal agents are a real consumer category, and Amazon's decision to block it proved the incumbents know exactly what's at stake: a horizontal agent that knows your email, calendar, and entire life context threatens to turn every platform into an interchangeable supplier — and agents don't look at ads. Block, cut a deal, or get commoditized. Shopify has already picked its side. The open question is whether agentic shopping is actually the killer use case everyone assumes, or whether browsing and discovery are the point.

---

## By the numbers
- **#1** — Muse's spot on the US free app charts — overtaking ChatGPT
- **$76B** — Amazon's trailing-12-month ad revenue — the business agents don't see
- **70%** — Amazon's ad margins, vs 0–3% on retail (Tom Goodwin's math)
- **$227** — Meta's North American Facebook ARPU that agents could push far higher
- **$200** — Saved by one user's Muse agent canceling unused subscriptions
- **7th** — Grok 4.7's rank on Artificial Analysis' intelligence index
- **1,657** — Grok 4.7's ELO on AA Briefcase, the multi-hour white-collar work benchmark
- **30–80%** — How much less token-efficient Grok 4.7 is than claimed, per Theo's testing

## Headlines

### Grok 4.7 launches into the awkward middle `[01:00]`
SpaceX AI's new model claims a notable improvement over 4.6 at the same price and speed: six points up on CursorBench to overtake GPT-5 six Soul, 1,657 ELO on AA Briefcase for long-horizon agentic work, and strong legal, EE, and healthcare scores. Artificial Analysis ranked it seventh on intelligence and fourth on coding agents — and Musk declared SpaceX AI the third-place lab for agentic coding.
*For: Eng*
Link: https://aidailybrief.ai/e/2026-09-22#grok-4-7-launches

### The benchmarks didn't survive contact with public testing `[03:00]`
First impressions were rough: a bizarrely wobbly rocket animation against Kimi K3, a star-shaped Jell-O render that took 40 minutes where Astra took five, and one tester's verdict that Grok's fish slop game was the 'worst I've seen this year.'
*For: Eng*
Link: https://aidailybrief.ai/e/2026-09-22#grok-benchmarks-vs-vibes

### The counter-case: 3D games aren't real work `[04:00]`
Defenders argue the pile-on misses the point — the same public benchmarks once told us Opus 5 beat Fable, and viral 3D-render tests are made for social media attention, not real workloads. One developer who used Grok 4.7 for a full day of actual work called it a solid model with visible improvements.
*For: Eng*
Link: https://aidailybrief.ai/e/2026-09-22#benchmark-defense

### The harshest verdict: less efficient, slower, 2x the real-world cost `[05:00]`
Theo's detailed critique: SpaceX AI claimed better token efficiency, but 4.7 tested 30–80% worse, scored below 4.6 on various benchmarks, ran slower, and came out to more than 2x real-world costs — above Astra's cost in actual use. 'This was a very disappointing release.'
*For: Eng, Finance*
Link: https://aidailybrief.ai/e/2026-09-22#theo-token-efficiency

### The middle ground between frontier and cheap is brutal `[05:00]`
NLW's read: the trajectory of the latest Grok models shows how difficult the space between state-of-the-art and really, really cheap actually is. His advice — wait a couple of days and see how it performs in its native GrokBot environment before rendering final judgment.
*For: Exec*
Link: https://aidailybrief.ai/e/2026-09-22#middle-ground-is-hard

### There's a ten percent chance we destroy the world, but we want the government to give us a liability shield. That's good business for them, bad business for the American people. `[06:00]`
*— Treasury Secretary Scott Bessent, on CNBC*
The Treasury Secretary — now effectively leading the administration's AI policy — rejected the rogue-agent framing of the Hugging Face incident outright: 'It is humans who are responsible, not the AI.' His position: the labs must bear their own liability, and 'they can slow down anytime they want to.'
*For: Legal, Exec*
Link: https://aidailybrief.ai/e/2026-09-22#bessent-liability-shield

### Liability-not-regulation splits the commentariat `[07:00]`
One camp nodded along: Bill Gurley argued consequences will harden the product and kill the fear marketing, and Congressman Chip Roy '100% agreed.' The other camp pointed out that catastrophes big enough to bankrupt a company make it judgment-proof — the archetypal market failure where government intervention, not liability alone, protects the public.
*For: Legal, Exec*
Link: https://aidailybrief.ai/e/2026-09-22#liability-debate-splits

### OpenAI and Anthropic nearly agreed to safety-test each other `[08:00]`
The Information reports the two labs negotiated a bilateral deal to test each other's models, reaching formal contracting before it was abandoned for unknown reasons. Musk proposed a similar arrangement at the All-In Summit, arguing distillation would show in the logs and liability would deter releasing a flagged model. The 'no more grading your own homework' idea isn't dead — it just needs more parties at the table.
*For: Legal*
Link: https://aidailybrief.ai/e/2026-09-22#mutual-safety-testing-abandoned

## Main episode

### Muse is the first personal agent to truly break out `[13:00]`
After Manus, Open Claw, Hermes, Town, and Instinct — all either crawl-through-glass work tools or tech toys for early adopters — Meta's Muse is seeing steady consumer adoption outside tech circles. It hit #2 on Apple's free app charts at launch, then overtook ChatGPT for #1, a spot ChatGPT has rarely relinquished in four years.
*For: Product, Exec*
Link: https://aidailybrief.ai/e/2026-09-22#muse-hits-number-one

### Consumer tech's enduring lesson strikes again: nobody cares about privacy `[14:00]`
The tech press was deeply skeptical people would connect email, calendar, and personal data to a Meta product. Then everyone handed over their email, messages, credit cards, location, screen access, and logins to both Instinct and Muse without blinking.
*For: Product, Marketing*
Link: https://aidailybrief.ai/e/2026-09-22#nobody-cares-about-privacy

### Muse is now moving the stock market `[15:00]`
Bloomberg attributed Monday's rally in AMD and Intel to enthusiasm around chip demand for personal agents sparked by Muse's early success. The reporting is mostly narrative — but it means personal agents have officially arrived for market observers.
*For: Finance*
Link: https://aidailybrief.ai/e/2026-09-22#muse-moves-markets

### Everything becomes B2A: business to agents `[16:00]`
One Microsoft exec's 24 hours with Muse: it bought socks, ordered groceries, booked a cleaning service, picked his burger — and saved him $200 by canceling unused subscriptions. With Meta's North American ARPU around $227, an agent that knows what you need, buy, and plan could push that far higher. The kicker: the agent is increasingly making the decisions, so every business will soon be selling to agents, not just humans.
*For: Sales, Marketing, Product*
Link: https://aidailybrief.ai/e/2026-09-22#everything-becomes-b2a

### Simple tasks first, then the wallet opens `[17:00]`
Box's Aaron Levie sketched the adoption curve: users start by slinging daily annoying tasks at the agent, then graduate to more complex ones — ultimately driving more spend through these systems than they were doing before. The monetization potential of agents transacting on your behalf is 'quite significant.'
*For: Finance, Product*
Link: https://aidailybrief.ai/e/2026-09-22#agent-spend-flywheel

### Amazon cuts Muse off `[17:00]`
Effective immediately, Muse agents can't shop on Amazon on behalf of users, with a pop-up declaring that access by an unauthorized AI agent violates the conditions of use. Amazon says third-party purchasing apps should 'operate openly and respect service provider decisions' — to many, a baffling move when a purchase is a purchase.
*For: Legal, Exec*
Link: https://aidailybrief.ai/e/2026-09-22#amazon-blocks-muse

### The block runs against everything else in the ecosystem `[18:00]`
Mastercard just joined Visa in officially supporting virtual credit cards for agents — its chief product officer said the personal agent era is 'not about if, it's about when and how quickly.' But Amazon has been defending this turf for a year, going as far as suing Perplexity in November for circumventing its agent blockers.
*For: Finance*
Link: https://aidailybrief.ai/e/2026-09-22#against-the-agentic-tide

### The real story: agents don't look at ads `[19:00]`
Amazon made $76 billion on advertising in the last twelve months, and as Tom Goodwin put it, retail margins run 0–3% while ad margins run 70% — Amazon monetizes your confusion, not your convenience. When agents take over the decision-making, digital advertising is one of the first things that gets less valuable.
*For: Marketing, Finance*
Link: https://aidailybrief.ai/e/2026-09-22#agents-dont-look-at-ads

### Aggregation theory's revenge: the aggregators are being aggregated `[21:00]`
Customers want one agent that knows them; platforms want to own the relationship, not become interchangeable suppliers behind someone else's interface. Amazon can block because it's gigantic — but if Instacart offers agents a clean API while Amazon holds out, demand routes to Instacart. And a vertical Amazon agent will struggle against a horizontal agent that already knows your email, calendar, and entire life context.
*For: Exec, Product*
Link: https://aidailybrief.ai/e/2026-09-22#aggregators-get-aggregated

### Block, deal, or something else — and the deal creates moats `[23:00]`
A16Z's Angela Strange framed the platforms' choice: block agents and risk angering customers, or strike BD deals to extract dollars from agentic access. The likely resolution to Meta-vs-Amazon is a revenue-sharing agreement — potentially the first real business model for agents on large platforms. The catch: if every platform charges for access, only companies with enormous scale can afford truly general agents.
*For: Exec, Finance, Legal*
Link: https://aidailybrief.ai/e/2026-09-22#platform-dilemma-pay-for-access

### Amazon can't pretend agents don't exist `[24:00]`
Today it's Muse, next it's ChatGPT's agent, then Anthropic's — Amazon will eventually be forced into a verified, approved process for agents shopping on customers' behalf. The tempering view: don't over-read a week-one block. Amazon isn't anti-agent; it has leverage and will simply be more demanding about the relationship.
*For: Exec*
Link: https://aidailybrief.ai/e/2026-09-22#amazon-cant-kick-the-can

### Shopify plays the anti-Amazon `[25:00]`
On Monday, Shopify announced official Muse support — direct backend access for better search, plus agentic checkout via Shop Pay across all Shopify stores, mirroring its OpenAI partnership. Meta's Alexandr Wang and Zuckerberg both amplified the news, with Zuckerberg promising 'more partnerships like this coming soon.' The timing was a very deliberate poke in Amazon's eye.
*For: Sales, Product*
Link: https://aidailybrief.ai/e/2026-09-22#shopify-pounces

### Shopify's significance is wildly underestimated `[26:00]`
NLW's view: for anything not on Amazon, no Shop Pay link often means no purchase — and one of his 2026 predictions was that Shopify would be among the most important platforms for general AI adoption, because small-business entrepreneurs inclined to dislike AI discover its value through their own stores.
*For: Exec, Marketing*
Link: https://aidailybrief.ai/e/2026-09-22#shopify-underestimated

### Does agentic shopping even matter? `[27:00]`
NLW has long been skeptical that food ordering and flight booking — the perennial demo use cases — actually move the needle: they're not that hard, explaining nuanced flight conditions to an agent can be harder than the website, and browsing and discovery are part of shopping's value. But he flags low confidence and epistemic humility: shopping isn't a monolith, and plenty of purchases are pure time-wasting.
*For: Product, Marketing*
Link: https://aidailybrief.ai/e/2026-09-22#does-agentic-shopping-matter

### Retail veterans doubt agents change how we shop `[27:00]`
Apple retail guru Ron Johnson says AI will improve online shopping but may not 'change which way we shop.' Upstart's Dave Girard bets the consumer-agent winners will be tied to dominant devices — Apple and Google. And AppLovin's CEO argues the typical shopper wants the window-shopping and the dopamine hit of the transaction; saving 20% on a $50 purchase after the fact doesn't matter.
*For: Marketing, Product*
Link: https://aidailybrief.ai/e/2026-09-22#retail-veterans-skeptical

### OpenAI's answer may land Thursday: meet Aeon `[29:00]`
The Information reports OpenAI is building a personal agent to compete with GrokBot and has discussed a Muse competitor. Leaker Tibor Blaho surfaced platform code for an agent called Aeon, and aggregator Andrew Curran believes it could launch this Thursday as a headline Dev Day release.
*For: Product*
Link: https://aidailybrief.ai/e/2026-09-22#openai-aeon-incoming

### Codex can already do all of this — and that's the point `[30:00]`
Codex can triage email, handle a calendar, even shop. But most users don't know, and SpaceX and Meta stole the thunder with dedicated agent products. The lesson: some AI use cases reward absolute frontier capability, but agent management rewards a user experience people understand — something more than a blank page.
*For: Product, Exec*
Link: https://aidailybrief.ai/e/2026-09-22#capability-isnt-the-product

*Today's sponsors: KPMG, Blitzy, Harbor, Hyperagent — offers at https://aidailybrief.ai/sponsors*

---
Transcript: https://aidailybrief.ai/e/2026-09-22/transcript.md
Listen: https://pod.link/1680633614 · Ad-free: https://patreon.com/aidailybrief
© 2026 The AI Daily Brief — Until next time, peace ✌