# Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models — Transcript (2026-10-01)

https://aidailybrief.ai/e/2026-10-01 · Listen: https://pod.link/1680633614

---

[00:00:00] Coming into 2026, Google was looking pretty good in the AI race. 2025 had been a good year. A lot of the questions inside DeepMind had been answered. They were putting out competitive Gemini models. They were pushing forward with interesting new products

And all the natural advantages that they had always had in terms of consumer distribution and data and all those things had a lot of people very bullish on them as a contender

But then vibe coding happened

and with the increase in coding capabilities paired with the power of the new harnesses like Claude Code and Codex those new capabilities unlocked agents in a way that hadn't been possible before And all of a sudden, Google found itself very behind

Charitably, you would say that in 2026 the company has been playing catch up But even that's not really accurate to what's been happening

for most of this year, Google has been firmly outside of the conversation as a top model lab

And yet I think that those who didn't have a partisan bias towards one of the other labs would never be fully comfortable writing Google off This week, the company announced Gemini 4, their first new frontier model in more than six months

By the benchmarks, it looks like Google is so back [00:01:00] But is that the whole story?

So let's dig in to Gemini 4 Argon and where it lands in the current AI race

The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Robots and Pencils, Harbor, and Granola. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts.

To learn more about sponsoring the show, send us a note at sponsors@aidailybrief.ai.

Also note that we have our next cohorts of super intelligence agent training coming up. These are paid programs which in very short order will get you far ahead when it comes to your agentic understanding and your ability to use agents in your daily work We have both the executive catch-up program and the agent intensive, which we call the executive agent leadership program the next cohorts for those start next week.

and you can find links to all of that at the very top of aidailybrief.ai 

President

President Trump [00:02:00] really looked at OpenAI DevDay and said, "Nope, absolutely not. I don't want that to be the biggest thing happening in AI this week."



and invited basically every big AI CEO to the White House for what David Sax would later call the Bretton Woods of AI

Now, there was a lot of chatter that came out of this meeting, but one of the first things that people noticed

was that Trump pulled Anthropic CEO Dario Amodei to be the one to speak to the press following the meeting

Now, Dario insisted that he's just saying what he's always said, that AI has some incredible benefits and some very real risks

However, in an act which some saw as public deference

Others saw as 

Dario finally playing the game. He commented, " As the president has said, whoever wins AI wins. I think that's very important. The mechanism, how we address the risks, is still under discussion. But we all need to work together to make sure that we can win, and we can win safely. If we do this right and work with the president and everyone here, we can all win safely."

Now, Now, X was absolutely filled with people psychoanalyzing the moment, breaking down the expressions from other tech leaders, Dario's [00:03:00] anxious tics, and even dissecting his fashion sense



while some thought that Trump and the other tech leaders might be bullying Dario as he stepped forward to speak The president followed up with some kind words later in the press conference saying, " Dario has been fantastic. We had dinner the other night, and he agrees with everyone."

Whether or not this was Dario truly bending the knee, it was a significant moment for a CEO who previously told staff that Anthropic had been targeted for refusing to give, quote, "dictator-style praise to Trump."

At At least in the immediate wake, the press conference completely overshadowed the subject of the meeting, which was the head of the Frontier Lab signing an accord on super intelligence. The accord is a one-page commitment to develop internal controls for frontier model testing and deployment, as well as partnering with external auditors for verification of those controls Trump told the press that the document is, quote, "morally binding," suggesting a continuation of the voluntary approach to AI safety testing.

Trump added, " "The biggest people in the world signed that, and I signed it as president, and it really is a form of protection." He also floated the idea of a 10-member industry oversight committee during the press conference, [00:04:00] but for now seems satisfied, commenting, " " I'm seeing tremendous self-policing."

The general tone, at least among the people at the meeting, was that this was an important first step

Mark Zuckerberg told the press, The idea isn't that this is the only thing we will ever do, it's that this is a start and an accord the whole industry could come to."



media, mostly the media response to this was a complete Rorschach test about how one feels about Trump and how one feels about the tech CEOs, i.e. whether you're inclined to use the word oligarchs to describe them

But I do think that the skeptical take on this one might be best summed up by media commentator Chuck Todd, who wrote, " "If If you trust the tech companies to police themselves and you want them to decide how AI will impact society without any input from the public, then you'll be fine with this accord."

I think even Trying to find the AI moderate or AI realist position in this one

I think a person could agree that it's better

that these folks are frequently in rooms together with key leaders in the government rather than exclusively viewing their role through the lens of competition



while while also wanting to see more external involvement, not least of which an articulation of what the independent external auditor or [00:05:00] evaluator to carry out independent assessments is actually going to look like in practice

perhaps now that this accord has been signed



that particular detail is one that can get built out Following the lunchtime meeting, President Trump signed a pair of executive orders marking the beginning of the super intelligence age. the first made Trump's name change official, stating, " As these capabilities continue to improve, they increasingly represent not merely artificial intelligence, but a new era of super intelligence.

The terminology used by the federal government should reflect the transformative capabilities of these technologies and the limitless opportunities they create for the American people."

Much more functional was the second order, which established America.gov as an AI-powered single portal for government services Trump said, " "With With America.gov, the federal government no longer stands in your way, it only stands at your service. We're simplifying it.

We're glamorizing it. We're making it what it should be Alongside that second executive order, we got an Apple-style product unveiling keynote, complete with Secretary of State Marco Rubio playing the role of Steve Jobs. Rubio showed demos of planned future [00:06:00] functionality, like being able to use the portal to apply for a passport, change your name after getting married, or enroll for Medicare

The features kind of seem like an MCP for government. A user can provide their information once, then the America.gov agent will track down and complete all the necessary forms across numerous government departments

How well this works remains to be seen, but certainly it is a real problem The insane difficulty, for example, of changing your name after getting married, depending on which state you live in, is hard to overstate Rubio said that the functionality will be available from next year, with the executive order giving government departments ninety days to integrate their services into the portal.



For now, america.gov only functions as a knowledge base for government services, searching across twenty-nine thousand government websites to answer questions based on official sources

Now, a lot of the chatter on this centered around security and privacy concerns, particularly after the president said, " It cannot in theory be hacked into, and when they figure out a way to do that, we'll end it, because they always figure out a way, right?"

The The National Design Studio, which designed the website, came under fire in June for including activity tracking tools in other websites they've [00:07:00] designed

And while some raised questions about providing the personal information required to apply for government services, AI jailbreaker Pliny the Liberator poked and prodded at the chatbot and found its guardrails were pretty airtight. He wrote, " Interesting. The america.gov chatbot will flag anything that appears to be personal information, like any text vaguely resembling an address or password, and refuse to accept the query until the flagged text is removed.

Never seen that before."

Meanwhile



If you think that them showing up together at this meeting means these companies are going to escape all government scrutiny

It might be worth noting that the FTC has opened an investigation into OpenAI and Anthropic's rogue agents. According to agency sources speaking with the New York Post, the FTC opened an investigation into the two leading AI labs a few weeks ago. The scope will cover the Hugging Face incident and the dozens of so-called rogue agent events that have been disclosed since.

The investigation has advanced to the stage that the FTC is drafting civil investigation demands, which are similar to subpoenas. An FTC source said the the agency has plans to compel the executives at these firms to testify about their products and about the dangers they allege their products may have [00:08:00] to consumers, to Americans.

Reports state that third-party safety research lab Meter can expect a demand alongside the frontier labs



Now, Now, from very early on in the AI narrative, one big thread of discourse has been focused on the idea that new AI-specific regulation wasn't necessary, that government could simply enforce product safety legislation that's already on the books. Former FTC chair Lina Khan was of this view, writing last month: " Law enforcers already have authority to charge companies and their CEOs for creating and releasing dangerous, unvetted, or defective products.

We shouldn't let discussions about new legal regimes distract from the fact that there's no AI exemption from laws already on the books." This post was co-signed by former AI czar David Sacks in yet another example of the AI safety debate creating very strange bedfellows



still still based on comments from the FTC source, it's not clear that the agency is necessarily pushing for some sort of harsh penalty. They told the New York Post, " "We need We need to win this SI race absolutely, and we are winning, and that's fantastic."

Getting a political jab in, the source continued, " And obviously the other side, the Democrat Party, wants to destroy this technology, wants to surrender to our [00:09:00] enemies and competitors, and that's not something that the chairman's interested in doing, and that's not really something the president's interested in doing.

With that being said, the laws have to be followed. People talk about how we need new laws for these companies, and the chairman's been very clear, we have plenty of laws on the books."

So an investigation, that's good, right? That'll appease people who think that the labs are getting a free pass. eh, that's not exactly what the discussion suggests

Anti-monopoly activist Matt Stoller wrote, " It's a protection racket. The idea is Trump investigates and clears them so other investigators have hurdles in their probes."

Former Director of Public Affairs at the FTC, Douglas Farrar wrote, " It would be good if the FTC investigated these companies properly, better starting two years ago, but the public credibility of this new investigation is 0.0%, especially after yesterday's AI CEO and Trump love fest."

The optimistic take comes from Joel Thayer who writes, " "I I kind of love this trust but verify strategy. Even though it's mostly self-regulatory, if any signatories fall short of these commitments, it could trigger FTC enforcement under Section 5 of its UDAP authority

Makes FTC [00:10:00] Chairman Andrew Ferguson's presence at the meeting very apt

As with anything policy right now, it's worth having a whole bowl of salt when you look at this

But things continue to move in some sort of direction. For now, that's gonna do it for the headlines. Next up, the main episode If you're leading AI inside an enterprise, you already know that the gap right now isn't capability, but execution. That's why KPMG's You Can with AI is back with a new season featuring conversations with leaders like Serojia Chatterjee of Emma, May Habib of Writer, Ellery Fisher, and others focused on practical execution.

What's working, what's not, and what it actually takes to move from pilots to real scaled impact across strategy, data readiness, governance, workforce, and value. And of course, it's co-hosted by me, Nathaniel Whittemore. Go listen and subscribe at www.kpmg.us/aipodcasts. That's www.kpmg.us/aipodcasts. 



The best teams don't have a single star carrying everyone else. They [00:11:00] know their own strengths and each other's weaknesses and play to both. That's the team Robots and Pencils has built on purpose. Nobody there is grinding through busy work to pad a headcount number.

People come for the hard problems, and they stay because everyone around them is leveling up at the same time. In a market full of companies that are just trying to hire fast, that's worth a look. check out robotsandpencils.com/careers 

If you listen to this show, you likely have a thesis. Maybe it's enterprise adoption, maybe it's compute, maybe it's a specific lab. Harbor Capital's AI Lab Ecosystem ETFs let you express it via five actively managed ETFs, each seeking exposure to the ecosystem around one major lab: Anthropic, OpenAI, DeepMind, Meta, or SpaceX AI.

your view of the AI race in ETF form. Harbor Capital Advisors AI Lab Ecosystem ETF suite gives investors a way to invest in the AI ecosystem they believe is best positioned for success. Search Harbor AI Lab Ecosystems ETFs wherever you invest or follow @HarborCapital on X to learn more

Visit harborcapital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information.

Read and [00:12:00] consider it carefully before investing. Risks include principal loss and artificial intelligence related risks. Harbor ETFs are distributed by Foresight Fund Services LLC. Harbor is not affiliated with AI Daily Brief, and the funds are not affiliated with, sponsored by, or endorsed by any AI lab This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principal. 

When I'm in a meeting, I'm fully in it. I'm thinking about the iteration and creative back and forth it takes to actually push a goal forward. What I'm not thinking about is capturing takeaways, tracking to-dos, or any of that, and that's where Granola comes in

Granola is an AI-powered notepad that captures what happens in your meetings and turns it into clean, structured notes with the decisions and action items pulled out and easy to find. There's no setup and no configuration. It just fits into how you already work

For me, it means I get to stay in idea mode and Granola makes sure those ideas actually become action

Once you try Granola on a first meeting, it is hard to go without. You can try it totally free at granola.ai/brief. That's granola.ai/brief to [00:13:00] get your time back Welcome back to the AI, welcome back to the AI Daily Brief

And friends, it appears that pigs are flying. Hell has frozen over. choose your metaphor for incredulity Because after months of waiting, Google has announced Gemini 4

Although maybe those pigs aren't completely off the ground Because although they announced the model, we don't actually have access to it yet On Wednesday, Google unveiled Gemini 4 Argon. It's their first frontier model in over six months

Now, the 2026 struggles of Google have been well-documented we never got a Gemini 3.5 Pro, with the company basically deciding that it just wasn't good enough to release

and for a few weeks now, speculation has been coming that Gemini 4 was on the horizon, and Google DeepMind CEO, Demis Hassabis says that the model delivers quote, " frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense."



now now despite a long period where Google's future as a frontier lab [00:14:00] has seemed somewhat tentative, to judge only by the reported benchmarks Google looks to be right back in the game Gemini 4 Argon is the new state-of-the-art model across many benchmarks

In agentic knowledge work, Gemini 4 scored sixty-eight point nine percent on the VALS index, beating Fable 5.1, GPT-6 Astra, and exceeding the score from the current leader, Opus 5.5, by a couple of percentage points. The gap was even larger on Automation Bench and VALS Finance Agent. And on Harvey's legal agent benchmark, Gemini 4 more than tripled the score of current leader Fable 5.1, scoring nineteen point six percent

Agentic coding scores were a little more patchy, but Gemini 4 is definitely in the ballpark of frontier rivals. It's the new state-of-the-art on DeepSUI, scoring seventy-seven point nine percent compared to Opus 55 seventy-four point two percent. On FrontierSUI, Gemini 4 scored fifty-five percent, which placed it ten points behind current leader Astra seven points behind Opus 55, and one point three points behind Fable 5.1



On Frontier SWI, however, Gemini 4 scored 55%, which [00:15:00] placed it 10 points behind current leader Astra 6 Seven points behind Opus 5.5 and 1.3 points behind Fable 5.1. It was also bottom of the pack on TerminalBench 4.0 with a score of 57.4%, nine points behind current leader Opus 5.5

5.5 Now you can look at this in two ways. One, agentic coding is one of the most important, if not the most important use case, so failing to achieve a state-of-the-art score is kind of problematic. or you can look at this as a huge improvement for Google, which it undeniably is.

coding had been Gemini's biggest weak point for the past year. So to have Gemini 4 be in the same ballpark as the frontier models from OpenAI and absolutely represents a big catch-up. For computer use, Gemini 4 is nipping on the heels of Astra, which set a new standard for the category.

On OS World, it scored sixty-nine point two percent against Astra's seventy-two point six percent Artificial analysis confirmed that Google is back in the mix as a leading model developer. Gemini 4 Argon scored fifty-three on the intelligence index, putting it tied for third place with GPT-6 Astra and Fable 5.1, one GPT-61 Sol, and three and five [00:16:00] points behind Sonnet 5.5 and Opus 5.5 respectively In terms of cost and efficiency

It's pretty close to the Pareto frontier, costing $1.99 per task on the AA benchmark run, compared to $3.26 for Astra and $7.63 for Fable 5.1. However, that is with Google's 50% launch discount which will be available for an unstated period, meaning that the full price would make it more expensive than Astra

It's also in that uncomfortable middle ground where it's more than twice as expensive as GPT-61 sole, even with the discount, without a super clear improvement on the benchmarks

For the moment, however, none of this really matters because unfortunately, Google is not making Gemini 4 Argon generally available to the public Their stated reasoning has to do with cybersecurity. The model scored 68% on the CWE bench cybersecurity benchmark, tied with Grok 4.7 GPT-6 Astra, and slightly ahead of Opus 5.5, Fable 5.1, and This was reason enough, according to Google, to limit access to what they call a set of trusted cyber defenders to ensure the model is not [00:17:00] misaligned. They said that they're engaging with the US government's voluntary testing and pre-release access program and intend to gradually expand access over time, beginning with API and Google Ultra subscribers

In their blog post, Google said that the model is already powering internal workflows and boasting strong performance on long horizon coding tasks like large scale code base migrations



now now as to the choice to announce the model without actually releasing it

I'm not really sure

Honestly, for us old heads who have been watching this for a while, it kind of has echoes of the first time Google fell behind and felt the need back in December of 2023 to announce Gemini, even though the pro version of the model wouldn't be available for a number of months

Yet maybe it was a good call because there were plenty of people who were excited to see it

There was an entire genre of posts on X

That basically came down to Google. So back. Dr. Derya Anutmaz wrote: " Wow, wow, Google DeepMind is back with Gemini 4 Argon. Insane benchmarks. In one fell swoop, they've jumped to the very top of the AI frontier. As I've said before, never bet against Google in the age of AI." Although he did add, and this seems to me to be the key detail, "[00:18:00] Can't wait to try Gemini 4 ASAP Ethan Mollick writes, " And it's a three-way race again."



Nathan Lambert, who does open model research and is about as far away as you could be from a hypes- from a hipster, writes, "Love to see Google surprising people with Gemini 4. More labs at the frontier iswonderful for consumers because of competition and wonderful for the world because of reduction in concentration of power.

Excited to see how it goes in real world scenarios."

AI leaker I Rule The World writes, " I've been critical of Google models, but Gemini 4 looks strong on benchmarks



we've seen this before with Google, but I'm optimistic they can release a first-class model with a first-class harness like Codex. Great work to all involved with Gemini 4. Rooting for you

But then on the other side of the coin is Angel who writes, " I'm sorry, but I can't with all these Google is so back posts. We literally can't even use the model, and it's not like benchmark scores for Google's models have ever been reliable. do we really need to remember what happened with Gemini 3 Pro?

Seriously?"

Ishu Ishu Aggarwal put some numbers around that

Sharing an older series of benchmarks and adding By the way, these were the [00:19:00] official benchmarks for Gemini 3.1 Pro when it released, practically destroying Opus 4.6 across the board. I'm not saying Gemini 4 Argon will be bad, but it's crazy that no one has the slightest bit of skepticism given Google's track record

To their credit, Logan Kilpatrick, who honestly should get Google's MVP for hanging in there throughout all of these PR challenges this year, actually engaged with this writing, " We've gotten much better at testing our models at scale across Google now.



so assume most new Gemini revs go through thousands of software engineers for weeks before getting released. Hopefully has helped close the benchmark to reality gap by a real margin."



now now so far all we have to go on is the reported benchmarks along with some reports from grumbling inside the company

A Bloomberg piece called Google Grapples with Employee Skepticism AboutNew Gemini 4 says, " While Gemini 4 has performed well on benchmarks widely used to gauge model efficiency, it does less well when employees actually put it to work. The model struggles to handle certain coding tasks according to people with direct access to the model."

Google, for their part, denied the reporting, commenting that it would be inaccurate to claim that Gemini is underperforming [00:20:00] in some areas such as coding.

And one of Bloomberg's other sources said that these gripers are a minority, claiming that a, quote, "large consensus internally at the company believe that Gemini 4 is the frontier."

What makes the position of Google so challenging is that in the time that they've been away from the top, the nature of the race has fundamentally changed with much optimism, Peter Yang wrote: " Google cooked on Gemini 4. Now they just need to be more competitive on coding harness, Antigravity, and personal agent Spark."

In other words, the competition has become more than just models

And it's not like there already isn't a ton of model competition.

One model released in advance of OpenAI DevDay that we didn't have a chance to talk about yet is Claude Sonnet 5.5



and certainly if you feel yourself giving a bit of an eye roll or at least a raised eyebrow wondering if you should actually care



the recent history of the Sonnet & Opus series

certainly created some reason for skepticism. Opus 5, you remember, was strongly disliked, and Sonnet 5 seemed to deliver even worse results for a higher price because it was so token hungry

[00:21:00] However, if you had to pick just one theme from the last couple of weeks on this show, It's the beloved return to form for Anthropic with Opus 5.5 So does Sonnet 55 continue that trajectory? The short answer seems to be, in most people's estimation, yes

Anthropic said that Sonnet 5.5 is 30% faster and 30% cheaper than Sonnet 5. and the model demonstrates some significant jumps on the benchmarks. For agent decoding, it scored 70.6% on Terminal Bench 4.0, which is up from just 10.3% for Sonnet 5 and even beats Opus 5.5 at 66.4%.

Scores on Frontier Code and Cursorbench were very slightly behind Opus 5.5. The same was true for knowledge work benchmarks AA, and AA Briefcase, where Sonnet 5.5 closed the previously massive gap between Sonnet and open class models Artificial analysis affirmed the benchmarks, giving Sonnet 5.5 an intelligence index score of 56.

That puts it in second place behind Opus 5.5 and somehow ahead of Fable 5.1 and GPT-6 Astra. However, [00:22:00] they didn't find that Anthropic's claims about cost savings stood during their testing Sonnet 5.5 cost $7.60 per task and used significantly more tokens than Sonnet 5 at the same per token price.

This meant that Sonnet 5.5's benchmark run was almost as expensive as Fable 27% more expensive than Opus 5.5, and more than twice as expensive as GPT-6 Astra

Now to give Anthropic the benefit of the doubt, they claimed a 30% cost reduction on similar tasks compared to Sonnet. the main AA benchmark is run at max inference settings, but turning down the settings to extra high cut the cost by two-thirds

Still, the model impressed once people got their hands on it. it is a massive improvement over Sonnet 5 in the rendering tests that grab attention on social media Matthew Berman made a series of visual tests and game clones writing, " Sonnet 5.5 basically Opus 5.5 but 50% cheaper and much faster.

I've been early testing it, and it's incredible. If this is pacing the frontier, sign me up



but but for real world use, Builder Kun Chen found it performed best when it's the right [00:23:00] tool for the job. He wrote, " "Sonnet Sonnet is definitely not the same as Opus just fifty percent cheaper. That's true only for problems that don't need much wisdom.

There's a clear difference when they are asked to propose plans for ambiguous product problems. Opus is able to approach problems with more strategic thinking, i.e. what's the real goal here? while Sonnet is more just looking at tactically, how do I get this done?" Now, Kun's current approach is to use Anthropic models as a complete system, with Opus for planning, Sonnet for implementation, and calling on Fable when things go sideways Huren tested Sonnet on his bug fix benchmark and found it outperformed all other models, including Fable and Astra

He wrote, " "The The secret? Sonnet 55 Max might be cheap and fast for most tasks, but it's the least lazy model I tested. It leads in bug hunting by spending turns."



Fascinatingly, YouTuber and entrepreneur Theo's review was that Sonnet 5.5 is an incredible model that you probably shouldn't use. In his view, the model is a big improvement over Sonnet five, but on max reasoning, it [00:24:00] overthinks and risks getting things wrong because of it.

And on lower settings, there's no point where Sonnet is more cost-effective than Opus because of how token hungry it is. This changes a little when using Sonnet as a sub-agent, where it can be more efficient on certain tasks, but it's usually not optimal as a standalone model

So will these results stand after a couple weeks of testing? With Sonnet 5,5 being a technical marvel given that it can produce the results of Fable 5 just four months after the release, but being too token hungry to make sense for most



That's certainly what I'll be watching, and I'll report back as people get more reps in

Still, going back to the new world that Gemini 4 has to compete in, the point that Peter Yang again was making is that just being good on the model isn't enough. In fact, one really interesting question brought up by the success of Muse



is whether an AI product with a great user experience but a less than state-of-the-art model beats a product with a less good user experience but a state-of-the-art model Muse has hit 3 million weekly active users

Daily users who send at least one prompt per day have reached one million. [00:25:00] Now Now remember, the product has only been available for a little over three weeks, so this is a strong positive early indication. At the end of the first week, Muse had five hundred thousand weekly active users and two hundred and fifty thousand daily users

Now, on the one hand, this is a tiny fraction of Meta's distribution potential, which reaches half the people on Earth and is nowhere near as strong as ChatGPT's growth. But a better comparison might be the adoption curve for Codex



Codex took about three months to reach three million weekly active users. so Muse currently looks like a faster-growing AI app. Now, one might expect that given that it is a personal agent versus Codex, very developer-focused audience

But whatever the case, the early success of Muse continues on



and yet now that OpenAI's DOTs is here, we're going to start to have some direct answers to this question of good UX plus worse model or worse UX plus better model

Eureka co-founder Sina writes, " "After After using Muse for a while and loving it, I'm starting to notice the underlying model isn't smart enough. I'm I'm not sure it's context overload, but it gets stuff wrong a lot [00:26:00] recently. Maybe they've had such huge success that they've had to switch to a different, cheaper, worse model.



anywho, Muse is great, form factor is great, but as a consumer, my loyalty is to none. If OpenAI will drop Muse with better models, I will use whatever is best. Intelligence is still the most."



and and yet on the flip side A couple days in, I'm definitely seeing some gripes with the way that Dots works Habibi Slop on X writes Overall, the vibes are not good. This feels directionally wrong. I just wanna talk to the model. This tool wants you to be a plumber in a house where you're not allowed to touch the pipes.

Dots has so much friction built in that it feels like it's not meant to be used



and lest you think that this is just a general griper account, they then go on to include about a dozen examples



the big one that I've seen repeated comes in their final thoughts when they write, " "They've They've launched a new product that wants to help you manage your work in ChatGPT and Codex without having sufficient access to any of the tools and data it needs from those services."



I I am certainly at least not even close to being willing to declare

Which is the winning strategy between focusing on the model versus focusing on the product user experience?



One One last note as we look [00:27:00] across the state of the competition that Google finds itself in as it gets ready to release Gemini 4. In one surprise launch this week, DoorDash has announced that their AI agent is getting the Muse treatment and can now take orders via text.

DoorDash already had an agent on their site, but they're using some of the features of personal agents like Muse to take it a step further Customers can now text the DoorDash agent with a natural language prompt, and the agent will figure out the rest. DoorDash gave the example of a customer texting, Order my usual protein bowl to the office."

The agent then matches the phone number to their user profile, looks up the usual order and delivery address, then handles the payment



Now, Now, in light of Amazon's decision to block Muse agents, there's a huge open question of whether the best agent strategy is to partner with the category leaders or build a platform-specific agent

The logic for an agent from DoorDash seems pretty clear at this point. The company has said that agentic orders have a fifty percent higher basket value when buying groceries. And by building their own, DoorDash can attempt to hold the user relationship close

There's a risk, of course, that external agents like Muse could order directly from restaurants and cut DoorDash out as a [00:28:00] middleman. At the same time, user preference might reject platform-specific agents because they have much less control

It's unclear what the optimal strategy will be, but for now, DoorDash seems to be doing a bit of both, keeping their platform open to third-party agents while also building an internal agent. And for that reason, it's worth paying attention to, even if you don't much care who wins how you order lunch to the office Anyways, bringing it back to Gemini 4 Argon

I'm with Nathan Lambert when he argues that Google doing well is both better for consumers and better for competition. so put me firmly in the camp of people who are hoping that this release is great

Let's just hope we get it soon so we can actually decide for ourselves



rather than dining from the table scraps of self-reported benchmarks and gripey employees talking to Bloomberg That's gonna do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace ​ 

[00:29:00]
