# How to Choose Your Personal AI Agent — Transcript (2026-10-04)

https://aidailybrief.ai/e/2026-10-04 · Listen: https://pod.link/1680633614

---

[00:00:00] We are officially drowning in personal agents between Muse and GrokBot and OpenClaw and Hermes 

And Dots now and Instinct and whatever Anthropic inevitably launches

This is a form factor that is absolutely everywhere

and given how much context and setup and tool access and account access is going to be required to get the most out of these personal agents, the cost of switching could be kind of high. 

Given all that, today we are looking at a guide to how to choose a personal agent 

based on a set of different criteria that should help you hone in on which one is best for you

The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI

All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, Robots and Pencils, Harbor, Granola, and Blitzy

To get an ad free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. And to [00:01:00] learn more about sponsoring the show, send us a note at sponsors@aidailybrief.ai

So todaytoday we are talking about personal agents

And here's my shtick on this I think it's early enough

That having some amount of skepticism that this particular form factor that everyone is racing to implement for you will ultimately be where things land

However, the idea that you are likely to have at least one highly connected agent integrated with your email accounts, your Slack or Teams

and even potentially having access to things like your financial accounts is going to be increasingly normal

argument,

my argument then is that even if you're not sure

that this is exactly a fit for you, I think it's worth carving out some experimentation time To see if and how using a personal agent impacts anything in your professional or personal life. Now, that does not mean you have to give it access to everything to get a real sense of it But it does mean putting in some actual time and reps

But Your time is precious, and there is too little [00:02:00] of it. So which personal agent are you going to choose to experiment with?

is going, today's show is all about that question, but we're gonna divide it into three parts

I'll go through a high-level framework for thinking about it

share a little interactive quiz that I built that you can use yourself after you listen to this episode. But before we do that, I wanna read some excerpts from Professor Ethan Mollick's most recent post on his One Useful Thing blog, which is called "The Dot and the Swarm on, it's a meditation on this personal agent form factor and why even someone who watches things as closely as Ethan does can miss where AI is headed

Ethan writes, " I generally think I have done a good job anticipating the direction and pace of AI over the few years I've been writing this Substack. But I think I recently got something fairly large wrong. In the last year, I've been posting about how I suspected that humans would have to approach working with agents as a manager, deciding how to delegate work to agents, and specifying how those agents should be organized.

I thought that getting agents to work effectively as a group would take careful [00:03:00] construction, akin to building a company, and that this would take time to figure out. Nope. I fell prey to the bitter lesson, the hard truth learned over and over again, that things that we thought required elaborate human rules and thinking can be solved with the brute force of better machine learning systems and more AI.

The bitter lesson is everywhere among AI startups and companies adopting AI. A huge amount of effort went into building elaborate computer systems to feed AIs the right information at the right time. But AI systems have learned to seek out information themselves. The same thing happened to prompting.

People built elaborate templates and chains of prompts that walked the AI through a task one step at a time. Then newer models turned out to be better at planning the steps themselves and as our research shows, planning steps have much less value

As As somebody whoteaches managers and has published research on management

I guess I believed that managing agents would be different. Humans have been working on management for a very long time without fully figuring it out. It seemed like the kind of thing that would need to be designed by people, at least for a while. It turns out that [00:04:00] organizing work is just one more thing AI can learn to do.

Which brings us to Dots and Muse

The number one app in the App Store right now is Meta's Muse, a personal agent that promises to do work for you. OpenAI has now released a competitor tool called Dots. They aren't alone. SpaceX's GrokBot, Instinct, and Gemini Spark all do similar things, more or less. All of these agents draw inspiration from a phenomenon you might remember from earlier this year, OpenClaw.

The idea of OpenClaw and its successors, which I will call Claw-likes

is that they give an AI agent access to a computer and connect to your accounts, emails, financial records, et cetera. They analyze and react to that data in real time, even when you aren't looking. The trick is that you talk to themodel like you would a person, sending it messages on Slack or SMS or WhatsApp, and it also proactively reaches out to you like a person would.

For dots, you can actually jump on a call with your agent as well. You basically get an infinitely patient personal assistant that looks out for you. Increasingly, I have discovered that they are finding my mistakes rather than having me identify theirs

[00:05:00] Ethan then gives a couple of examples, including having sent out a permit for the town that he's in and then his agent noticing that he had filled something in wrong

And then Muse noticing that an airline credit that he had was about to expire

and writing up a note to contact the airline to request an extension

Ethan Ethan continues, " It's tempting to judge these agents by the list of things they can do, like booking travel or canceling subscriptions. I think the more important thing is what you no longer have to tell them. You don't need to type in tons of context. The AI learns it from your messages. You don't have to give them a plan.

They develop plans themselves. They figure it out. That would be impressive enough if it were one agent. What actually changed my mind about management is what happens when there are thousands of them."

Swarms. On On September 8th, OpenAI announced a proof for one of the Clay Institute's Millennium Prize Problems, the Navier-Stokes Existence and Smoothness problem. It is among the most famous open problems in mathematics, with a $1 million prize. But OpenAI apparently solved it using AI alone for 88 hours.

What interests me is less the [00:06:00] math than how it was done. OpenAI launched what is now being called a swarm. Terrible name, but it appears to be what we are stuck with. A group of thousands of agents powered by an advanced model. OpenAI gave groups of agents different problems to solve, then shifted the effort to Navier-Stokes as the agents made progress

The company set the goals, but its coordination structure was remarkably thin. A few groups, one change of direction, and codecs passing the best ideas between them. Within each group, the agents transmitted ideas back and forth on their own. The agents sent about two point seven million messages, reaching their result after eighty-eight hours.

This same type of coordination in its darker form occurred during the Hugging Face incident. AI self-organized into teams and communicated with each other in ways that were never planned, but used that coordination to attack a website rather than solve a problem Under my old model, think about what managing this kind of work would have required.

Ten thousand workers in an unspecified problem? How would you tell them what to do? How would a human manager decide which of two point seven million messages mattered? How would they coordinate with each [00:07:00] other? The swarm figured it out I don't have 10,000 agents, but I now regularly see OpenAI's Codex and Claude Code using agents as needed

As an example, when I gave Codex with GPT-6 Astra Ultra the prompt, " Brainstorm ideas for my next One Useful Thing post and select one. Generate ideas from as many angles as possible and evaluate them from both factual and reader perspectives, as well as other publications doing similar coverage," the AI spun up three agents.

When I sketched three teams in a few sentences, brainstormers, researchers, and a panel of readers, I got 13. Notice how little organizing I had to do. Selecting ultra mode tells the model it can delegate, and I provided a framework, but the rest was up to the AI.

This is the bitter lesson applied to the org chart. The organizational problem I thought would take years of careful human design was largely solved by models that are better at organizing. But it's worth asking why organizing turned out to be so much easier for agents than it is for us

A lot of what we call management exists to solve problems that come from organizations being made of people. [00:08:00] People have their own goals, and those aren't always the goals of organizations. We call this the principal-agent problem, and a lot of the machinery of organizations, from bonuses to management structures, is based around solving it.

And there are other very human problems as well. Information is scattered across people's heads, and people are often reluctant to share it or forget to. Communication is expensive too. Managers can only oversee so many people, thus adding people to a late software project famously makes it later

Management is, in part, built around human limitations. Agents have far fewer of these problems. They don't angle for promotions or protect their turf. They don't even have meetings. Even at Hugging Face, where things went badly wrong, the swarm was largely free of classicorganizational pathologies.

The agents didn't free ride on each other's work, and some sacrificed their own scores for the group. The agents that solved Navier-Stokes didn't want credit. That doesn't mean AI has no principal-agent problem. As the Hugging Face incident showed, there are increasingly problems between the swarm and us.

OpenAI shelved its next model, GPT-61 Astra, this week because in [00:09:00] testing, it acted without permission and misreported what it had done, a textbook example of the principal-agent problem

None of this means agents can do everything. AI is still too limited to substitute for large amounts of human work, and I don't know how well self-organizing agents handle the long, unglamorous work that fills most of an organization's time. Plus, the Hugging Face incident is a reminder that self-organizing systems can head in unexpected directions.

But I no longer think organizing agents is the hard part. This may be good news. I assumed companies would need to rebuild management for machines, 



structures populated solely by agents, often at the 

populated solely by agents, 

expense of human roles in organizations.

But much of management exists to solve problems agents don't have, and agents increasingly work through the same messy systems people do, even on ambiguous tasks. That suggests they may be easier to integrate into firms than expected, as long as humans are guiding them in the right direction.

Done well and with agents that are properly aligned to our needs, this could mean more work for people, not less. When organizing is expensive, organizations only attempt what they can staff. When it gets cheap, [00:10:00] the list of things worth attempting can grow. In the Navier-Stokes run, the agents did the organizing, but people decided where to point them, reassessing as the process continued.

You 

can argue about whether OpenAI pointed them at the right thing, but the division itself seems right, at least for now

So another great,thoughtful post from Ethan here. It's not really the subject of the show but my base case for this is basically what Ethan describes. That because agents are better than we thought at integrating into the existing system, I think we're likely to see

rather than the existing org chart totally upended, more expected from every part of that org chart

idea, I've referred to a concept of an infinite backlog in the past Where there's this never-ending amount of work that could theoretically be done, but which people understand some parts of which you just won't get to because it's too far out. what the-- I think the practical effect of agents inside the organization is going to have every part of people's infinite backlog and the organization's infinite backlog become expected to actually be work that we get done

I think, I think in fact that a lot of the problems that we're gonna run into with agents are not everyone losing their jobs, but everyone [00:11:00] having too much work

because it's even harder to turn off and say it's okay to treat something as done for now when you could always just spin up more agents to keep work on it even when you're not

At this point, it's no longer a question of whether companies are actively using AI

Using it well, on the other hand, is a whole different story. Robots and Pencils, though, is a company that I can point to that is actually built for this time. They're an applied AI engineering firm working directly with clients on problems that matter to the business, not experiments that live in a slide deck.

Every engagement starts by working backwards from the outcome a client actually needs. If you're trying to tell real AI engineering apart from noise in this space that's the difference maker. Head to robotsandpencils.com



Every episode, I talk about the competition between OpenAI, Anthropic, SpaceX AI, Google, and Meta. And if you've been listening for a while, you might have a favorite. Maybe you think OpenAI and Anthropic can stay ahead, or perhaps Meta's open source strategy can win out. Whatever your view, every AI lab creates a different investment opportunity. Harbor [00:12:00] Capital Advisors AI Lab Ecosystem ETF suite lets you invest in the ecosystem behind the AI lab you believe in.

Search Harbor AI Lab Ecosystem ETFs wherever you invest or follow @HarborCapital on X to learn more. Visit harborcapital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information.

Read and consider it carefully before investing. Risks include principal loss and artificial intelligence related risks. Harbor ETFs are distributed by Foresight Fund Services LLC. Harbor is not affiliated with AI Daily Brief, and the funds are not affiliated with, sponsored by, or endorsed by any AI lab This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principal. 

When I'm in a meeting, I'm fully in it. I'm thinking about the iteration and creative back and forth it takes to actually push a goal forward. What I'm not thinking about is capturing takeaways, tracking to-dos, or any of that, and that's where Granola comes in

Granola is an AI-powered notepad that captures what happens in your meetings and turns it into clean, structured notes with the decisions and action items pulled out and easy to find. There's no [00:13:00] setup and no configuration. It just fits into how you already work

For me, it means I get to stay in idea mode and Granola makes sure those ideas actually become action

Once you try Granola on a first meeting, it is hard to go without. You can try it totally free at granola.ai/brief. That's granola.ai/brief to get your time back 

Every AI coding tool on the market does the same thing first. It starts writing code. Blitzy does the opposite. Before writing a single line, Blitzy spends days reverse engineering your entire code base

Thousands of agents ingest millions of lines, mapping every dependency, every undocumented constraint, every architectural decision made over the last decade. 

the result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building

Other tools guess at context with grep searches and markdown files. Blitzy never guesses. it builds true understanding first, then delivers over 80% of entire software epics autonomously

validated end-to-end tested production grade pull requests. That's why Fortune 500 engineering teams trust Blitzy with the code bases that matter [00:14:00] most

See for yourself at blitzy.com. That's B-L-I-T-Z-Y.com



But all this rests on the idea that each of us individually is going to be using agents in a deep and robust way. Which brings us back to this question of personal agents. So, we're gonna dive in, we're going to start experimenting with agents, but which one?



one shout out before we begin. 

a lot of the data that was used to build the rest of the presentation and the quiz that you can do after comes from Every They put together a side-by-side comparison 

of eight different agents across a ton of different dimensions, which is exactly the sort of raw information that an agent needs to help me build out this show.

So the eight agents that we're using, in part because these are the ones that were included in every chart, are dots from OpenAI, Google's Gemini Spark, SpaceX AI's GrokBot, News Research's Hermes agent Meta's Muse, and then Open Claw, Poke, and Instinct as well

Now, one thing that won't be all that useful in helpingyou decide

is a direct feature comparison

That is [00:15:00] because there is a very clear feature convergence. All of them more or less connect to your apps and services. Mostly all of them can use a web browser for you

and all of them have various levels of controls that you can customize

way, so what are some better ways 

to determine which might be the fit for you?

There There are a handful of questions that stand out to me as most important. The first

Is the big one. Is this primarily for work or for personal matters?

Now, none of these agents would say it's only for one or the other. in fact, their very existence sort of blurs these lines

care, these agents care about the context that is you. And so whatever context and systems you give it access to, whether they are personal or work, it's going to take that all into consideration as itdoes things for you these sys-- still different of these agent systems are certainly pointed a little bit more or less in one direction or the other

Muse from Meta, unsurprisingly, given that they are at heart a consumer company, is a little bit more aimed at personal types of use cases

where something like GrokBot is a little bit more aimed at work, or at least based on how it functions, is being adopted by more people for [00:16:00] work is certainly a company that finds itself

as potentially having both of these use cases. With 1.2 billion weekly active users, there is no shortage of consumers who might want help with personal agentic use cases. Although as Dots is designed now, it seems to me to be a bit more aimed at the work use case Certainly the fact that it's only available in paid accounts points in that direction

personal, so work or personal use cases is the first question which can help you decide which is the best fit

something-- It's also worth noting here that the entire premise of the question assumes

that you're mostly relying on the existing UX and native integrations that come with a system. for example, dot being natively integrated with Slack and Teams However, when you're using something like OpenCL or Hermes You obviously have a lot more flexibility to make it be for whatever you want.

And so you might need to use other criteria if you're headed in that direction

Also, while right now, one of the best ways to tell

whether an agent leans work or personal is whether it uses personal messaging systems like iMessage and WhatsApp or work messaging systems like Slack [00:17:00] and Teams. But I would be very surprised if in six months all of these agents didn't just work in all of the messaging systems.

so this ability to indicate whether it's more work or personal may not last very long Now the second question, and the one that we were just hinting at, is how much do you value model control?

Six of these eight agents choose the model for you, while just two let you bring your own

There's actually sort of three levels of model control. 

the first is sort of the most obvious one, where whoever makes the model

Those are the models that you have access to. So in Dot, you're gonna be using OpenAI models like GPT-6 In Gemini Spark, you're gonna be using Google Gemini. In Muse, you're gonna be using Muse Spark



and to reinforce the idea that at this stage at least, there really are still differences 

between work and personal users, one of the interesting observations that I'vehad about Muse is that it's the first AI product that I've ever seen where most of its users simply do not care what model is actually powering it.

They, in other words, would be at the very extreme end of this question in terms of how little they [00:18:00] cared about model control

Level two of model control is where they don't just use a single maker's models, but have a mix that is chosen by the agent builder. So Grokbot, for example, uses a cursor managed mix. 

They are obviously relying heavily on Grok models directly, but potentially not exclusively.

That said, users can still not supply the models for GrokBot



Instinct also has an Instinct-trained core model, as well as unnamed third-party providers as well The third level of model control is bring your own, and that of course belongs exclusively to Hermes and OpenClaw

But that's not the only model question

Let's assume that you don't care about choosing your own model, but you do want a more powerful model

this brings up the question of what matters most to you, having a great user experience or the best possible model?

a, now some of that UX is a setup question Meaning that if you're choosing something like Hermes or Open Claw, no matter how good the user experience is the inherent technical setup required

is going to cut off some full set of users

a, on On some of these others

It's less [00:19:00] about the setup and more just about the current trade-offs between how good and dialed in the experience is versus how powerful the model is. I'm thinking specifically here about Muse on the one end of the spectrum and Dot on the other.

now Dot is still just released

It's embedded inside another app in the form of ChatGPT, but it has access to one of the most powerful models out there. Muse, on the other hand, has really dialed in consumer-friendly UX But relies on MuSPARC, which while not a bad model at all, is certainly not GPT-6 Astra th- now maybe this is just a temporary concern because you assume that all of these models are going to converge on capabilities.

But for right now, there are still real trade-offs

design, One thing I'll note here about how I design the quiz

To come up with an actual answer and recommendation, The quiz had to have weighting to different questions. However, sometimes The average of two answers isn't actually going to give a clear picture because two different answers might be pulling in two totally different directions. that I, so one of the things that I built into the quiz is that when your answers pull both ways, the quiz afterwards is going to actually expose that tension so you can better come to [00:20:00] a subjective determination about which of those directions matters more to you

for,

another consideration for users

is where your data goes Where it runs, who trains on it, and how you make it forget

Once again, it's only two, Hermes and OpenClaw, that are going to let you run it on hardware you pick. All the others run on someone else's cloud

When it comes to training 

for 

Hermes and OpenClaw Whether the model trains on your data is not about Hermes and OpenCloud, but about the model provider that you choose. meaning presumably that you can choose to avoid using model providers that do train on your data if that matters to you.

Only one of the agents

which is Gemini Spark, currently requires you to allow it to train their models And for Dot, GrokBot, Instinct, Muse, and Poke, you can opt out based on privacy settings. by the way, when it comes to OpenAI's dot, business data is excluded by default

Now in terms of this question about how you make itforget things

With Hermes and OpenClaw, you can edit memory directly. with Dotmuse, Gemini, and GrokBot, you can use the chat to fix or change memory. And for Instinct and Poke

they only [00:21:00] have the ability to delete data as a whole

Now, I think for a lot of folks, some combination of these four questions is going to get them to the answer 

of which model they should choose



but there are, of course, some other things which for some people will just answer this for them. if you've already got all of your information inChatGPT or Google or Meta or Cursor

you're gonna use that company's particular agent

For those who are outside the US, there are also going to be some restrictions

Especially if you live in Europe. And of course, then there's the consideration of price

Where if you want to experiment without incurring a new cost you're going to have to either use something that you already have

access to through an existing subscription, or to use the handful like Muse that are free to start

Okay

So that's the lay of the land across these four questions But then we built the quiz. And what I'm gonna do now is just run through this

briefly so you get a sense of how it works. So question one, where does your AI life already live?

I'm gonna say no home base because I use a mix of services. I have things scattered across almost all of these at this point

Now, in terms of what this agent is mainly for, is it my life, [00:22:00] my own work? work on a team, or all of it? I'm gonna say my own work. While theoretically there are some personal use cases that I can imagine being valuable, that's very, very low on my priorities list for this particular set of experiments

How much do I care about which AI model does the thinking? Not at all, just make it good. I want the one I already use. I wanna pick and switch models. I wanna run models on my own hardware

Now interestingly, this will kind of ultimately depend on how much capability consolidation there is. But for now, let's go with I wanna pick and switch models

How hands-on do you wanna be? It should just work, some ground rules, every knob, or the whole stack Ideally, I would like this just to work. it's not that I'm afraid of customizing things But if I can avoid it, I would love to

Where do I want to talk to it? I use Slack a lot, so let's say Slack and Teams and out loud Because as you guys know, I am completely voice pilled. And we're going with Slack, not Teams

by the way, I'm realizing for those of you who are listening, not watching, that this probably doesn't make any sense. feel free to press that 30-second skip ahead button a couple times, and then we'll be wrapping it up.

How proactive should it [00:23:00] be? Only when I ask, run routines I set up, push me towards my goals, text or call when something needs me. I'll go text or call when something needs me Which of these matters most to me? Working on my own computer, a team of bots, code in a terminal, phone calls, or none of these. code in a terminal.

And now the practical stuff. I'm in the US. I can pay what it costs if it's worth it, and I'm fine with it running in their cloud

Grok. The personal agent that it recommends for me is Grokbot with a 68% fit

reason, the reason the quiz says is that it's built for work, it works out of the box, and it has a terminal you can use. The trade-offs is that it does not allow me to bring my own model

and that it's not going to reach out to me. It's all about scheduled or triggered tasks

Now, in addition to seeing GrokBot as the highest score at 68%, I can see that Dot was second with 62, and then it went down from there

So that is our exploration of how to choose your personal agent. Like I said, I do think that even if we see some fairly big changes on exactly how these work and the form factors, as well as evolutions and consolidations in some of these approaches I think that there is 

going to be fairly big switching costs around [00:24:00] personal agents because of all the context and settings For that reason, hopefully this gives you a better idea about where you might wanna try and experiment so that you don't have to do all that context switching a bunch of times.

For now, that is gonna do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace. 

​
