// Friday · August 14, 2026

How to Decide What Work AI Should Do for You: The AI Deputization Audit

Two new features — GrokBot's teach-a-task and ChatGPT's Computer History — mark the moment the AI bottleneck officially moved from capability to context. But once your AI can learn how you work, you still have to decide what work it should actually do. NLW proposes a scoring system: the AI deputization audit.

Ad-free on Patreon
Today's sponsors — KPMG · Blitzy · Harbor · Hyperagent · all offers →
The One Idea

AI's bottleneck has moved from capability to context — now the question is what you deputize.

GrokBot lets you teach a task by recording yourself doing it; ChatGPT's Computer History watches how you work and learns over time. Deliberate demonstration and ambient observation both solve the same problem: models are capable enough, they just don't know your work. What's left is a decision framework — score each recurring process on frequency, teachability, checkability, stakes, and how much it really needs to be you. Deputize the 8-10s, defend the 0-3s, and duet on everything in the middle, which is where most knowledge work still lives.

// 01

By the Numbers

340 tok/s
Gemini 3.7 Flash's speed — more than twice as fast as GPT-56
750 tok/s
GPT-56 Sol in ultra fast mode — 14× speed, on Cerebras hardware
40¢
Per task for 3.7 Flash on the Artificial Analysis run — ~8× the ultra-cheap tier
+20%
GPT 5.6 Sol's quality edge over Kimi K3 — while also 13% cheaper
1/5
What Gemma 4 and Inkling cost vs GLM 5.2 at similar quality
9 mo
Denise Dresser's tenure as OpenAI CRO before Thursday's exit
24 hrs
For a plumbing company to go zero-to-automated dispatch with GrokBot — no engineers
8/10
The deputization audit score where a task becomes a hand-off candidate
// 02

The Brief

ModelsEng01:00

Google ships Gemini 3.7 Flash — and it's built for speed

Not the delayed 3.5 Pro or the anticipated Gemini 4, but a play in a different category: efficiency. At 340 tokens per second it's more than twice as fast as GPT-56 and even edges NVIDIA's new Nemotron 3.5 Lightning, with solid gains on the DeepSwee coding benchmark and prices cut in half.

AI Daily Brief
ModelsEngFinance03:00

Flash sits in an uncomfortable middle ground on cost-per-intelligence

Even after the price cut, 3.7 Flash costs 40 cents per task on the Artificial Analysis run — the same as MuSpark, more than Nemotron 3 Ultra or GLM 5.2, and around eight times more than ultra-cheap models like GPT 56-Luna. Most users are either paying a little more for a frontier model or looking for something much cheaper.

AI Daily Brief
ModelsEng04:00

The counter-case: for interactive coding, Flash's speed changes the math

Vercel's Brandon Galang says to ignore the FUD — 3.7 Flash sits on the Pareto frontier once you account for 5.6 Luna being a much smaller model — and analyst Max Weinbach is enjoying it in Anti-Gravity with 'insanely high' usage limits. For partially synchronous coding, where you sit and interact with the agent's output, speed can justify a premium even below frontier performance.

AI Daily Brief
◆ The TakeExec05:00

The 'model race' is branching into multiple races

There's still a race for the frontier — but also races for distribution, harnesses, and revenue. Within models alone, Google is betting speed is a dimension people will pay attention to, and the massive efficiency boost also makes 3.7 Flash a big upgrade for Gemini Spark, Google's personal agent.

The AI Daily Brief
EnterpriseFinanceOps06:00

Switching to a cheaper Chinese model won't necessarily save you money

An Alpha Sense study ran US and Chinese models through hundreds of real-world financial-analysis tasks. GPT 5.6 Sol completed the work around 13% cheaper than Kimi K3 with roughly 20% higher quality, while US open-weight models Gemma 4 and Inkling matched GLM 5.2's quality at less than one-fifth the cost. Sonnet 5 was the outlier: lower quality than Opus 4.8 at more than five times the price.

AI Daily Brief
EnterpriseFinance07:00

Some of the more expensive models actually ended up being less costly because they were more efficient in using tokens.

— Jack Coco, Alpha Sense CEO. The message for enterprise buyers and planners: price per token is a deceptive metric. The faster conventional wisdom shifts toward measuring efficiency on your actual tasks, the better off your budget will be.

The AI Daily Brief
ModelsEngProduct07:00

OpenAI answers on speed: ultra fast mode for GPT-56 Sol

The company claims frontier intelligence at fourteen times the speed — 750 tokens per second, more than double Gemini 3.7 Flash — leveraging its Cerebras optimizations. API-only for select customers, pitched at latency-sensitive work like real-time voice, commerce, coding, financial research, and security response. No word on the premium, but presumably it's not ultra cheap.

AI Daily Brief
BusinessExecHR08:00

OpenAI's CRO exits after nine months — the third senior departure in months

Denise Dresser, the Salesforce veteran and former Slack CEO brought in to ease investor concerns, is leaving; former Wiz president and COO Dolly Rajak replaces her. Coming days after Brad Lightcap's exit and following Fiji Simo's July departure, it looks like a pattern — and Axios sources suggest President Greg Brockman has been building up his own team of leaders.

AI Daily Brief
◆ The TakeExec09:00

Nine times out of ten, the Occam's razor explanation for personnel moves is personal

There's an outsized appetite right now to read every executive shift as a tea leaf for major problems inside OpenAI, and analyzing from outside is fraught. Still, the market is noticing — turnover is being flagged as a pre-IPO problem, though with the listing now delayed to next year, there's plenty of time to build another narrative.

The AI Daily Brief
EnterpriseExecProduct13:00

The bottleneck in AI has moved from capability to context

Two features launched this week make the shift concrete: just because a model can do something doesn't mean it has the information it needs to do it well relative to you personally. The new tools attack that gap directly — by watching you work.

AI Daily Brief
EnterpriseOpsProduct14:00

GrokBot lets you teach it a task by recording yourself doing it

In Cursor and SpaceX AI's simplified take on Open Claw, you hit a plus button in chat, record yourself doing something in the browser, and the bot watches — then theoretically can do it again. The Cursor team says many things GrokBot handles on its own, but if it's struggling, teaching it a task solves the problem in one fell swoop.

AI Daily Brief
EnterpriseOpsProduct15:00

ChatGPT's Computer History learns from everything you do on a computer

Per OpenAI's Ari Weinstein, it lets ChatGPT understand how you work, finish tasks you're in the middle of, and suggest skills and automations based on how you use your machine. Same goal as GrokBot's teach-a-task: show the AI how you work so it can do more of it.

AI Daily Brief
EnterpriseLegalOps15:00

This is Windows Recall — except this time people want it

Microsoft's 2024 screenshot-everything feature was branded a privacy nightmare and had to be recalled and reworked. Now users are openly handing AI their history, medical records, and bank statements, and creators who quietly wanted Recall are saying so out loud. Part of it is shifting privacy attitudes; a technical difference helps too — as Simon Smith notes, Computer History records interaction events, not screens or audio.

AI Daily Brief
◆ The TakeProductExec17:00

What changed isn't just privacy norms — it's the value proposition

Microsoft pitched Recall as solving 'finding something you've seen before on your PC' — not a big enough problem to risk the privacy trade. An AI agent that actually does your work for you is a much different value proposition, and now there's real choice in how agents get the context they need.

The AI Daily Brief
EnterpriseOpsProduct18:00

Two paradigms for AI learning your work: ambient observation vs. deliberate demonstration

Computer History is ambient — it watches across apps and builds context without effort on your part, which for many will be the killer feature. GrokBot's teach-a-task is deliberate — you intentionally decide 'I am going to teach you this skill,' show the workflow, and it saves the steps as a repeatable routine. Both patterns have real constituencies.

AI Daily Brief
◆ The TakeExec19:00

Deputize, don't automate

The word choice is intentional: deputizing AI to go do something in your stead is a different relationship than automating a task away and never thinking about it again. That framing shapes the whole audit.

The AI Daily Brief
EnterpriseOpsExec19:00

Step one: inventory your recurring processes

No matter how novel we want our work to be, a lot of it is stuff done day in, day out: email triage, weekly status reports, meeting prep, research briefs, CRM hygiene, content repurposing, scheduling, vendor portal chores, inbound lead qualification, metrics pulls. Recurring work is what deputization is built for.

AI Daily Brief
EnterpriseOpsExec20:00

Score each process 0-2 on five dimensions

One: is it worth it — how often does it happen and how long does it take? Two: teachability — could you show it in a ten-minute screen share? Three: checkability — how long does verifying the output take versus doing it? Four: stakes — how bad is it if the AI gets it wrong and nobody catches it? Five: how much you personally are the key factor in the output's quality.

AI Daily Brief
EnterpriseOpsExec23:00

Three tiers: deputize (8-10), duet (4-7), defend (0-3)

Eight-plus means low stakes, high frequency, highly teachable, and it doesn't need to be you — hand it over, spot check, and start there with Computer History or GrokBot. Zero to three means keep it: mistakes are expensive or the person on the other end expects you. The catch is that most knowledge work today falls in the duet middle — and the real question is whether AI learns well enough over time to move duets into deputized.

AI Daily Brief
EnterpriseOpsEng24:00

For everything you can't hand off, name the specific blocker

Because new tools change the equation blocker by blocker. Work trapped in legacy software with no API? Computer-use agents that click the same screens you do dissolve that. Processes you know how to do but can't write down? Now you can show rather than tell. Missing context might be solved by Computer History's ongoing observation — but probably not by a ten-minute GrokBot demo.

AI Daily Brief
EnterpriseOpsLegal25:00

Some blockers survive — and the new tools even create one

Work requiring taste or judgment, costly or irreversible mistakes, and tasks that depend on human relationships aren't solved by teach-a-task capabilities. And if privacy or security was the blocker, these tools can add one: recording your screen creates a new thing to secure.

AI Daily Brief
EnterpriseOpsEng26:00

GrokBot's day-two reviews: great for beginners, limiting for power users

Superintelligent trainer Nufar Gaspar found GrokBot is good for people who haven't built complete agent systems yet, but advanced users hit walls — no control over which folder holds the relevant context, no model choice. If you've been through Claw Camp, it may not be a perfect fit.

AI Daily Brief
EnterpriseOpsSalesCS27:00

Early GrokBot patterns: topic-per-bot with a chief of staff on top

John O'Neill, who owns a plumbing company, went from zero to automated dispatch and office chores in 24 hours with no engineers on payroll. Many users spin up individual bots for individual tasks but interact only through a main chief-of-staff bot that coordinates the rest — an early emerging best practice.

AI Daily Brief
EnterpriseOps27:00

The universal starter job: inbox/Slack tracker and morning brief

Matt Van Horne notes it's the same first job as for every agent generation before — except now it takes about a minute to set up. A good low-risk first experiment for anyone with GrokBot or Computer History access.

AI Daily Brief
◆ The TakeExecHR28:00

The biggest barrier is carving out time to work differently

Even highly AI-forward people struggle here — especially high-productivity workers whose workflows are so dialed in that changing them feels like a short-term waste of time, even when it would save a lot of time in the long run. The experiments are worth it; watch for GrokBot's cost to come down as access expands.

The AI Daily Brief
Machine-readable ▸Download .mdTranscript .md— feed it to your own agent

Got this from a colleague? Get the brief every day.