# Everything You Need to Know About AI Tokens — Transcript (2026-08-02)

https://aidailybrief.ai/e/2026-08-02 · Listen: https://pod.link/1680633614

---

[00:00:00] Today on the AI Daily Brief, an operator's cut episode with Nufar, everything you need to know about AI tokens. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. ​

All right, All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, Rackspace, Blitzy, Section, and Airtable

To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts send us a note at sponsors@aidailybrief.ai

All right, friends. Well, Nufar Gaspar is back today, and Nufar and I have been cooking up a lot recently. a whole slew of you have done our most recent program, have explored our most recent program, the choose your own adventure style AI summer adventure



plus we've been cooking up an expanded set of educational resources that we'll be telling you about soon.



But one of the realities that both Nufar and I have been [00:01:00] living in is every company we interact with dealing with the same questions

of AI tokens and token economics

We are now firmly in the agentic era of AI, where companies have to think not only about how to get adoption and how to maximize AI's value, but how to do so in a way that doesn't just totally break the bank and where the right intelligence is being used for the right problems

Anyone who's ever built an open clock can tell you

that getting the right models to do what you want them to without going off into endless cycles of spin, take some real consideration. today's episode is designed to be the ultimate primer on AI tokens.



what we're talking about when we say that term, what the new challenges are

And some of the key pitfalls to avoid, as well as strategies to maximize the way that you and your company use AI tokens 

All right, Nufar back with another Operator's Cut talking about the topic du jour, the topic on everyone's minds. We are talking tokens. How are you doing?

I'm good. very psyched to talk about tokens.

Yeah, I think it's, uh, I love this period in a [00:02:00] discourse where we've gone from sort ofpulling hair out, freaking out about the new change to actually settling into new tactics, new strategy, and I think this is a perfect fit with that. So tell us a little bit about what we're gonna be talking about and, uh, and let's dive in. 

Good. So the reason why I wanted to, do this episode is because every room that I walk into these days, like literally every room, have some version of the same token conversation. Some practitioners feel like they are being watched when they use an expensive model. The regular users wonder whether one ambitious prompt will eat their weekly allowance, and the leadership teams, they see a bill growing faster than expected, and then they start asking a lot of questions on whether all of these, tokens produced anything useful.

So I don't know where you guys are sitting, but there is a new anxiety around, using too much intelligence, and I actually wanna flip the conversation, first of all, to make sure that everybody understands what tokens are and what the bill actually means, and then how to spend them wisely rather than sparingly.

So that's, why I'm here and what I'm planning to do [00:03:00] today.



one of the places that I've found myself with this conversation is there's been such a visceral reaction now as the cost has gone up. my-- I have a bigger concern around people retreating back to known ROI biases and, not boring, but u-ultimately low stakes use cases, let's say,as compared to what AI can actually do, that I've found myself in the position of having to defend things like token maxing and token leaderboards just relative to where the, the tone has shifted.

I think obviously we'll get into today the smarter version of that conversation, so I'm excited for it.

Exactly. All right. because, there is like, a growing conversation around, that, I think that there is a very, like a better language around, the feeling of, tokens shouldn't be just a financial thing. And just recently, the OpenAI CFO, proposed a scorecard, and, it was called Useful Intelligence Per Dollar, and that's built around the one question: What does each successful task actually cost?

And that's the conversation that I think people should have, and [00:04:00] I want to help you read the whole story. what are tokens, what your work costs, and where usage creates value, and where it quietly leaks, value rather than adding. So to kick us off,and how, uh, we got in here, I want to walk you through, the four eras of token consumptions, and probably you will recognize where you are.

and we all started by being basically token oblivious. That was the all-inclusive era, where model companies, subsidized the usage, and the flat sub-subscriptions hid the meter, and many individual users, and many of them are still there, just see the ceiling, not a per token price. So that's where we started.

And then, like you said, we, got into the era of token maximizing. the-- that was the leaderboard era where usa-usage became the badge of AI maturity. And we all remember some of the conversations around Meta, who tracked employee AI usage on an internal leaderboard. they used to call it, Claudnomics, and they used roughly, between sixty to seventy-four trillion tokens in a single month.

And, according to the [00:05:00] data that was published, the top individual user used, two hundred and eighty billion, tokens. So to give a sense of how many this is, that is roughly, two point th-three million books, worth of text. So if you want to try and imagine that, that's about, fifty books every minute continuously for a month.

So that was the Meta story. And then Uber launched also an adoption leaderboard and burned through the entire twenty twenty-six AI coding budget in about four months. And there was another company, unnamed, but, uh, according to TechCrunch, they ran up, five hundred million, dollars of Claude bill with no usage limits in place.

So then, in that era, usage became the metric, and dashboard measured activity while claiming to, measure value. Obviously, this was unsustainable, and I know you have some opinions around that, but, I'll be curious to say-- to hear your points, but I, also wanna say that it actually got us to be, as always, the pendulum took us way too far to the era that I call, token anxious, which is where we are.

But if you have anything to say in defense of, leaderboards, I'm, [00:06:00] here to listen.

Yeah. Now the, look, the defense is less leaderboards are a great concept. They come with, I think, a set of very predictable challenges. In fact, so predictable that I would say that the hand-wringing around the idea of people gaming them has always struck me as a little absurd. of course, people are gonna game systems if you put real stakes around them, but that's pretty predictable and also fairly, from first principles, you could figure out a lot of ways to deal with that.

So I think, one, it's overwrought,the sort of people who are freaking out about that. Secondly, my bigger point was that a company that wildly overspends right now via a token leaderboard or anything else, I will bet any amount of money that they will be farther ahead than a company that underspends because they're overly concerned with proving out ROI or whatever it is, on a sort of year timescale.

Now, the Goldilocks scenario, which I think is what you're gonna get into, is being able to experiment, being able to learn, being able to build, and being [00:07:00] able to actually understand consumption while also not being afraid of it. but I agree. I th- I think that the pendulum swung so aggressively, far too aggressively back from token maxing and excitement to token anxious, and that's the paradigm that we've been living in,in the recent, few weeks, couple months, whatever it is.

Yeah. Good. So you, I think you love my,model of, how to use tokens wisely. But with regards to being token anxious, what I'm seeing in many companies now is that many employees are self-censoring themselves, basically trying to avoid, costs. And even if we look at the same companies that were token maxing, so Meta went from the leaderboard to sending a memo that constrained the AI usage, and now, the press is calling it token, minimizing instead of token maxing, and Uber caps employees at 1,500.

So even the same companies who are token maxing are now, significantly,shortening that. And when employees, self-censor, it gets them to feel like every prompt is an ROI conversation, and that's not something that we wanna have. And I think that [00:08:00] this is a very bad era to stay in, because I believe that the most expensive token is the one that your best person is afraid to spend.

So this is where I want to direct all of us to be in, and I call it the token smart era. Meaning that you need to spend wisely, not sparingly, and understand what creates value and where usage quietly leaks. That's, the entire kind of theme of this episode. And what I wanted to,go by in this era, in this, podcast, is to, talk about the,four elements of what a token actually is, why tokens, were not born equal, how to audit your own usage, and how to govern or do it better within your company.

And I'm trying to make it relevant to anybody, whether you're the practitioner that needs to, apply some cost engineering playbook and, be smart about that, or the executives and admins that need to be proactive and avoid, having a difficult conversation with the CFO without a proper response to, " Let's just cut the bill without, talking about the business implications," as you [00:09:00] just said.

So that's the plan for us today. So I wanna start, with,introducing you to the token, 'cause I feel that even though it's, the most used term in AI, not too many people truly understand what it is. because that's, in the root of every bill, quota, and rate limit with AI, in tokens.

so in a simple, words, token is a chunk of a text, that the model reads and writes. It's typically bigger than one character, and it's usually smaller than a word. And, if you, have never, ever seen a tokenizer or how tokens look in action, OpenAI has a very good, page that is open to everyone that you can just take a look at how tokens actually look.

So it looks something like that. you can paste the text, and then you will see how, words are being chunked. So you can see that some words are stayed, staying, like, as one token, while, others might be separated into multiple tokens. And interestingly, numbers, often are being chopped in the middle, um,and so on.

So,we'll put it in the show notes, but, a very interesting experim- have never seen how your [00:10:00] text looks. And by the way, if you paste a non, English or a non-Latin language, you will see that typically the amount of tokens is much, larger than an English language. So that's the, OpenAI tokenizer.

And, S- a few things just to land it, home. In general, uh, the ratio in English is around,three-quarters of a word to token, meaning that, if you have, a page of text, it's roughly one thousand tokens. and some languages, that, that are like Hindi, Thai, Greek, other languages like that might, get two to five X more tokens for the same content.

And because billing is per token, then some questions, if you ask them in other languages, cost you much more, and that's sometimes referred to as, language tax, with AI. with code also it's different, and it has its own way. indentation and brackets and white spaces, they all become tokens.

There are some newer, ways to tokenize text that, are more code friendly, in order to do that. But, still,numbers is a huge problem. So you've seen the one, two, three, four, five being [00:11:00] chopped in the middle. And by the way, that's also why, whenever everybody's doing like the strawberry, uh, test for AI, and it,very badly, fails in trying to count h- how many Rs are in the word strawberry.

In many cases, that's just a tokenization, feature rather than a failure, and, the model just have-- has never, ever seen, the individual letters. It just saw the straw and the berry as separate words, and that's why it's counting it off. so model models have, various workarounds, but many of the like, AI is so dumb memes are literally just, tokenizer issues.

So that's tokens. In terms of what everyday work costs, I think that's a good, kind of mental model to have. So for example, drafting an email is around five hundred to seven hundred, tokens. A page of text, as noted, is about one thousand tokens. You can see a longer text, it can be, more than that.



if you send a model or a tool to do like a, a AI web search, often, it will add a few thousand, more tokens [00:12:00] for the, result, sometimes much more. images, interestingly, are in many cases are, not that large in terms of how many tokens. They are roughly around, slightly more than one thousand tokens.

Interestingly, deep research, can very easily be seventy thousand or, hundreds of thousands of tokens. But just the other day, one of our learners in one of courses, had a yes/no question, and accidentally, instead of asking for a web search for the problem, he was asking for a do a-- like the agentic tool to do a deep research.

The tool spawned about one hundred sub-agents to do the research, and then his yes/no question cost, over four million tokens, just to answer this question. So it can very easily amount to much more than that. specifically, a few additional places where you can find very token-heavy workloads will be, data analysis.



that can easily get to, one million or more tokens per, task. And, heavy, coding can also, be very aggressive, similarly with, many agentic, working flows. So just to give you a [00:13:00] sense of where it is and to be a little bit more concrete here, the everyday stuff, as you've seen, like the emails and so on, is almost free.

So it's around half a cent, and nobody should ration, emails. it's not where their money goes. Search and research can multiply very quietly, so that can, be a place to look for efficiency. And the top of the ladder, that's a completely different sport. So if you compare like email to agentic coding, task, it can be a factor of a, a thousand or even more.



and another thing that you need to pay attention is that every conversation compounds. So the model doesn't remember your previous messages and as such, it sends all of the previous conversations within the same session back to the model. So by,let's say turn number ten, it may be processing so much of the earlier exchange alongside your new message that the total grows much faster than the number of turns suggests.



that even happened before the system prompt, and we'll talk a-about strategies in, later on, but this is one of the things that can very easily, just having very long sessions can very easily amount to a ton of [00:14:00] tokens, being consumed.

I think this is one of the reasons why this is such an important conversation is, another way to put this is that the more advanced and ultimately higher value use cases consume more tokens, which is intuitive that more intelligence is required for bigger challenges. But the direction of use cases is proceeding this way.

And so the reason that the token anxiety is going to create problems if not addressed, is that it will incentivize people to stay swimming 

In daily mails.

less sophisticated use cases. So th- this is the, the trajectory is clear. in terms of less token consumption.

The, you want as a leader group, your people to be doing more advanced, more useful things with AI. It's just how they do it well.



So you want them to do deep research where a deep research is required, but you don't want them to accidentally do a deep research on a yes or no question that they can Google in a second. Good. so speaking of the agentic or the more, advanced, capabilities,[00:15:00] those, can significantly grow, the amount of tokens, because agents work, autonomously in loops, and as such, they consume, by very widely cited industry estimates five to thirty times the tokens of a simple chat, and poorly designed, agentic loops, or agentic, harnesses, can be even worse than that, because a typical task involves between ten to twenty model calls carrying instructions and history and tool definition and previous results.

And I think according to McKinsey, they estimate that roughly around sixty percent of an agentic task's cost is tied to, the checking and refining and like regeneration of the answers after the first response. So the expensive part is often getting from the answer to the, accepted results.



so that's an interesting one. and now it gets even more complex because tokens were not born equal. So by the way, the point here is not to not use the agentic tool, just to know that in-- as you said, intelligence, cost. But it gets even more complex because, tokens were not, born equal, [00:16:00] and every model lab has its own tokenizer.

You've just seen the OpenAI, but, different, model labs have different tokenizers. So for example, the OpenAI current tokenizer has a vocabulary of about two hundred, thousand tokens. Gemini has around two hundred and fifty-six thousand. Llama by Meta has about half of that, and Claude is unpublished.

And the reason why we all should care is that the price per million tokens is denominated in each lab's own tokens, and often we don't know them. And the same document can be ten to twenty percent more tokens on one provider than another, and even more so for code and non-English text. So the model behavior,widens the gap, and one model may answer in a single pass, while another reasons longer and writes more and takes more agentic steps or, needs retries.

And the tool around the model add its own system and context and the loop design, so the two stacks doing the same task can have different token counts and different completion rates, and as a result, a completely different bill. So the per token price is kind of the [00:17:00] sticker, but the cost per accepted task is the operating metric 'cause otherwise there is no way for you, to compare between, different providers and different tools. All right. an im-important, uh, story that also illustrates that, what happened when Opus, four point f- seven came on board. the tokenizer basically under the hood changed, and, it was this April, and the price sheet was identical to the previous model, the same dollar per million token. But the model was shipped with a new tokenizer that produced by, Anthropics on documentation.

They didn't hide it. Roughly thirty percent more tokens for the same text. So there were quite a few independent analysis of over a, a million requests that found native tokens,they grew and the count grew by about thirty-two, all the way to forty-five percent. And the real world, bills grew by, twelve to twenty-seven percent because some of the, difference was, absorbed by caching.

even Simon Wilson, he measured one of his own prompts at a- at around almost one, and a half X, more tokens. So even though it was documented, [00:18:00] facto, we paid more for the same intelligence, and this is like a, a shrinkflation, right? The same sticker price, but a smaller candy bar. So nobody prints now thirty percent fewer words per dollar, which is the case that happened, there.



so that's something that is constantly changing. Every lab tunes the tokenizer and often for good reasons. But the operator lessons here is that we have to talk about dollars,per task and not dollar per token, because the budget is like a moving denominator, and it's not the way for you to try and understand, how much it's gonna cost.



let's talk about what tokens are used for by the AI tools. you have to understand that every AI request has three token layers, and they are priced very differently. We have the input tokens. those will be the prompts and the conversation history and the files and the tools definition and everything that,is part of the input.

This is what the model reads, and this is the cheapest per token, but can accumulate fast because if the history is being recent or if a lot of context is being read, that can cost quite a lot. Then we have the reasoning tokens. That's [00:19:00] the second layer. These are the tokens being used for the model internal thinking before answering.

For the most part, it's gonna be invisible to you, but, it's billed at the output rates, meaning at the high rate of per token cost, and those can add between four to twenty x, cost per, request. And finally, we have the output. That's the answer that you actually see, and this is typically three to five x,more expensive than the input price per token.



and I think the reasoning layer is the one layer that catches everybody by surprise because you might have a four hundred token answer, but under the hood, it carried like thir-- I don't know, four thousand, thinking tokens underneath because the model was having an internal monologue and doing a lot of thinking in order to give you the answer.

And if you want the analogy, it's like thinking about the part of the restaurant bill that is labeled the kitchen time. So you don't get to see it. It's not part of the dish, but you still have to pay a lot of it, for that. and the models with the high reasoning effort are often the one with the twenty x, amount of tokens being consumed versus the [00:20:00] lower reasoning effort.

it can be the same question with a significantly different price tag. And sometimes spending a higher reasoning does not get you better results. So some metrics even show that for a simple question, it's better to use lower reasoning because the overall cost per task will be significantly, lower, and the quality will be improved without, having the model overthink everything.

So it's not always that smarter or, spending more time thinking gets you better results.



All

One of the more interesting shifts in enterprise AI right now is how quickly the conversation is moving towards infrastructure and operations. As AI moves into core workflows, regulated data environments, and agentic systems, enterprises need governed infrastructure and inference that can operate reliably day to day with clear operational accountability built in from the start.

As those systems scale, the operating model increasingly becomes part of the AI strategy itself

Rack-Rackspace Technology is the operator of the full enterprise AI stack, from agents to infrastructure across private [00:21:00] cloud, hybrid cloud, and edge environments. Rackspace builds and operates governed AI infrastructure, inference, and production AI systems for organizations where sovereignty, compliance, and uptime are non-negotiable.

Their forward-deployed engineers stay embedded beyond deployment to help operationalize and run AI in live environments. to learn more about where enterprise AI runs and outcomes scale, go to rackspace.com Every 

AI coding tool on the market does the same thing first. It starts writing code. Blitzy does the opposite. Before writing a single line, Blitzy spends days reverse engineering your entire code base

Thousands of agents ingest millions of lines, mapping every dependency, every undocumented constraint, every architectural decision made over the last decade. 

the result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building

Other tools guess at context with grep searches and markdown files. Blitzy never guesses. it builds true understanding first, then delivers over 80% of entire software epics autonomously

validated end-to-end tested production grade pull requests. That's why Fortune 500 engineering teams [00:22:00] trust Blitzy with the code bases that matter most

See for yourself at blitzy.com. That's B-L-I-T-Z-Y.com



Here's a harsh truth. Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized. Half of companies have AI tools, but only 12% use them for business value. Most employees arestill using ai. To summarize meeting notes, if you're the one responsible for AI adoption at your company, you need section.

Section is a platform that helps you manage AI transformation across your entire organization.

It coaches, employees on real use cases

tracks who's using AI for business impact and shows you exactly where AI is and isn't creating value.

The result, You go from rolling out tools to driving measurable AI value. Your employees move from meeting summaries to solving actual business problems, and you can prove the ROI. Stop guessing if your AI investment is working. Check out section@sectionai.com.

That's S-E-C-T-I-O-N ai com. 

This episode of the AI Daily Brief is brought to you by Hyperagent where [00:23:00] you run fleets of agents your team can manage together New users get 1000 in inference Forget local agents and chat workflows waiting on your laptop to be prompted Hyperagent deploys alwayson agents in the cloud doing real work across the tools your team already uses Marketing's agent turns competitor moves into landing pages sales agent enriches leads drafts emails and updates the CRM ops agent chases the paperwork and tracks the budget Every agent has access to shared context and follows your rules about scope and approvals It's time you add agents that feel like teammates Hire yours at Hyperagent built by the team at Airtable Claim your 1000 in inference at hyperagent.com/aidailybrief. 

in inference at hyperagent.com/aidailybrief. 

All right. and I think that the gap between input and output pricing keeps widening at the frontier. if you look at the Fable 5, it's about, ten-- not about, it's ten dollar per million input tokens and fifty dollars per million with the, GPT 5.6 solids, 6X ratio. So we're seeing the gap even, [00:24:00] widening, and those effort levels, that's also something that highly, adds the complexity because these, frontier models increasingly letting you, dial the reasoning effort.

with higher efforts, the reasoning tokens are significantly higher, and that's probably the dial that you should even be more mindful of even beyond the models, 'cause those can easily, cost you ten to 12X, token increase between, high or extra high effort to the low or medium effort effort All All right.



one last thing, here, a price tag that, that came from Databricks because, a smarter model might not always be more expensive than a less expensive model. what Da-Databricks did is they tested coding agents on real engineering tasks from its own code base, and they were using Sonnet five.

It was one point seven times cheaper per token than Opus, four point eight. However, Sonnet cost, around two dollar per task or two, two point oh nine per task versus one nine four for Opus. so because Sonnet needed [00:25:00] more iterations and more reasoning and, had to, spend way more tokens to get to the same results, overall, Opus, which is sig-- significantly on paper more expensive model, it was cheaper to operate, which means that we shouldn't just reach to the cheapest model possible.

We need to reach to the right model, for the task and that's not easy to get, but something to be mindful. the other thing that matters to the bill is the tool itself. So Databricks in their same experiment ran the same model at the same thinking effort through different agent harnesses, and they saw that more than a two X difference in cost per task with the same quality, using different harnesses.

just because primarily one tool was, feeding the model roughly three times less, context than the others, and thereby the overall cost was lower. So very difficult bill to read and very difficult bill to navigate, and I'll try to help you as best I can. so bottom line, we're dealing with cost, per task and not, tokens because otherwise, we will not be able to actually compare apples to [00:26:00] apples.



and the, the cost will include the retries, the review, the, every, like a correction that needs and every additional iteration. and then you need to divide by the number of accepted results. That's your cost per accepted task. That's the metric that you should aim for and optimize for this also brings the conversation much more into return on investment and business value rather than just having a conversation around tokens that is very hard as hopefully by now you understand to meter.

Good. the practical task that you can do, you need to take between five to ten representative tasks of what you do. Run them through tool model tool options in order to have, a good, understanding of, in, within your option space what you should do. and then, hold the input and the quality bar very constant and compare first pass success attempts, human correction, and elapsed time and the total cost.

The winner is the stack that gets your actual work done reliably. I know it sounds like a lot, but if you have a good taxonomy that was optimized for yourself and you know for the task that [00:27:00] you do which models overall get you better results potentially with fewer tokens or if you can do that for your team or your company, and you will need to do that re-recurrently because things change quickly, then at least you can teach folks that if you're doing that type of research, the recommended to, model to get you to the overall, best quality, and the value per, task is the following, and so on.



that's the current reality that we live in.

Okay. Now I want to,give you a language on how to look at your own tokens and, hopefully using that you will be able to, distinguish between the tokens that add value to the ones that not so much. And, um, every token that you or your organization spends, in my opinion, is one of three kinds.



there are tokens that I call tokens that teach. and this is, running in both directions, meaning that you're, teaching yourself. as Nathaniel said before, we don't want to stop the experimentation. So the tokens that, include the experimentation, the failed workflows, and let me try these three different ways so I will [00:28:00] learn.

And, so these are tokens that are,worth spending because you get, learning out of them, and you can look at them as tuition. And by the way, those also include, what you are teaching AI about yourself. So those will be the identity files and the curated context and the knowledge packs. the memory can be, counted,as, those because, teaching your AI who you are, what is your context, and learning f-- what works for you in AI is, very critical for you to continue moving forward.



they look a little bit like, a waste on a dashboard or a lot like a waste on a dashboard, because no deliverable ship. But I claim that these are tokens that you need to defend fearlessly because if you won't defend those, you will very go back to just, help... getting AI's help to draft emails and translate between languages rather than moving towards the workflows that, matter and especially if your company and yourself has a lot of catch up to do on where AI is currently at, at.

So I want the tokens that teach to, be defended, because those [00:29:00] are the, the things that will move the needle, beyond the next category, which I call them the tokens that produce. So obviously, those are the most defensible ones because those are the tokens that you use to, create work that ships.



Uh, it can be the, like the final proposal or the research or the code. obviously, that's the thing that is much easier to show the ROI. but lastly, we also have the tokens that should be eliminated, and those are tokens that I call tokens that spin. those can be machines talking to themselves or automations that nobody is looking at their output or automations that are running too infrequently, idle agents, bloated context, misused, tools and context, using Fable to, write an email, so using the wrong, model, unoptimized, workflows and so on.



those will be activity without sufficient output. So the token smart move, if I need to summarize, is to kill, the tokens that spin, to tune the production to make sure that it is cost effective and protect the teaching. and that's the order. Like first go and do the [00:30:00] audit on your spin, tokens, and then do the rest.

I have a very embarrassing, tokens that spin, story, which I will share in a minute. but I do wanna also note with regards to tokens that teach that a failed experiment is as important as a successful experiment. So we-- you should definitely encourage your employees to fail, to, try, because otherwise, the tokens they produce will not, yield as much value as possible



So it's embarrassing, as an AI expert to, to talk about it, but, my OpenClau, was a chief of staff. Was because it's currently disabled. A chief of staff that I called Chloe, and it was using the Anthropic API. And because it was using an API, it was like auto-renewed, renewing all the time, and the bills were sent to a secondary inbox.

So I wasn't really monitoring them. And I was seeing that the charges seem quite high, but because I was getting a ton of value and because I was not paying attention to how frequently I'm getting a new bill, I wasn't noticing. And then early June, I was traveling, so I was not using my OpenClau at all, and still, I, see that the bill [00:31:00] kept coming.

So I was saying like, "Why am I still getting some bills?" So I opened the dashboard only to realize that I spent in two weeks, fifteen hundred dollars on an agent that I was not using. So I opened, the dashboard, like double-clicked, and I realized that I have almost four hundred million tokens in and almost zero tokens out.

So it was a ratio of almost three thousand to one, from input to output, and that's literally the definition of a machine talking to itself, and billing me for like an internal monologue that it was running with itself. looking further, there was a bunch of cron jobs that the OpenClau created for itself.



and like it was like a compaction job that ran every thirty minutes on empty sessions. And even worse, like the trend was going up. So I first of all closed my, OpenClau and only to optimize it differently. But, if it happens to me, in this setup, it can happen to literally everybody, and especially when the, the credit card is owned by your company and not by yourself.

Often, you will not pay attention 'cause you are not sitting on the billing.

Yep. and also the [00:32:00] fact that you have a bunch of other things that are working well that you might assume it's those things that are amounting for the cost. So one thing that I wanted to mention with spin is that I thinka lot of the framing of spin, if people were to pick this up, they might assume that it's only mistakes or errors that produce that spin.

But that's not always gonna be the case. Like sure, this is sort of ain-between example where it wasn't exactly an error because it was doing something that it was meant to, but you weren't really paying attention, so it was doing more of it than it needed to. But I think a lot of times spin will also be just ill-defining the parameters for, a job that you actually do want.

an example of this that I had is I had an OpenClaw going for a while that was perpetually researching new data sources in AI, that could help us figure out where the state of adoption metrics was, right? Every day, there's new studies that come out that measure this or measure that, and that tell you about data readiness or systems [00:33:00] integration or use cases or whatever.

And it's too much to monitor for humans, but, agents are really good at it. And so this OpenClaw agent was a researcher that its only job was to, on a set schedule based on, on its heartbeat, go out and check for new things. And it was never meant to stop. It was always, it was on a specific schedule, but it basically was this continuous research process that was crawling to the ends of the internet every day and it ended up just not being valuable enough for, the cost, but it was doing what it was supposed to.

And so I think part of the, auditing spin is also just figuring out what things have accidentally become spin, even if they started in the right area. and I think, that's, why this idea of auditing I think is a good framework 'cause, sometimes it's going to be about just updating or changing a process that was valuable, as well as catching mistakes.



I agree. And even more I see many automations that people, created because they think they will be useful. Like,"Oh, I, I can't read. I have too many Slack messages. Let me just create like a [00:34:00] Slack miner that runs every hour and reads my entire,set of channels." And that can easily become $1,000 in tokens that literally do a, a job that moves the needle for nobody.



or the morning brief that, you created wholeheartedly with the intention to read it every morning, but for some reason you don't find value and you don't read it. So these are the things that you should definitely, audit and kill. And my rule of thumb is if you created an automation and for one or two weeks you have never, ever used the output, you should definitely kill 'cause it's a definition of a spin.

Or if you are, using this automation, but there is, an, a very bad proportion between the value of, summarizing all of your Slack channels to the, bill at the end of the month, that's also something that I consider to be a spin and I think that now that everybody gets co-work or,GPT work, that even more and more within companies because it's so easy to build these automations.

And without sufficient literacy about, how to effectively use the tokens, people create a ton of these automations that look good on paper, but don't look so g-great on the paper of the bill at the end of the month.

All right, so let me [00:35:00] give you a list for the suspects for silent token spenders. First of all, it's gonna be your idle agents and the over frequent jobs. These are gonna be the things that run without any meaningful output or way, way too frequently. We also, in many cases see automations that nobody uses, so it can be like the weekly report or the dashboard that, nobody ever goes to read.



additional thing can be what we refer to often as the pre-prompt tax. So anything the model runs and reads before the very first prompt. So those i-include the always-on rules or instructions, the skill definition, the, the tool definition, and so on. And those can very easily, if not, properly organized, amount to many thousands of tokens, each run without you typing a single word.

So those amount significantly. many folks also hold the immortal conversation, that they will continue an endless session that keeps carrying old history and old context, often if even creating a poorer quality we also, in many cases, see users, [00:36:00] never filter their data retrieval.

So instead of just getting twenty rows from a database, they will pull, five hundred rows, or they will process the entire inbox to look for a specific mail that they know what was the subject line, and so on. many other folks will have their context all over the place, so the agent will have to read through a ton of documentation just to understand, what, are they talking about and what's the, source truth here, as well as, rework loops.

So any time that your agent or your skill or your just day-to-day usage gets you to do more iterations just to get the same result, this, n- is, is, like, just, more tokens being spent on, nothing. So these are the, like, immediate suspects, and I wanna show you, how you can try to potentially id-identify whether your system is, in a spin situation or that the spin-to, production, ratio of tokens is not well formulated.

And the thing here is that not everybody can detect in the same way. some folks have concrete meter. Those will be people who are [00:37:00] using the API version of the, models. They have the API console that they can use. Or, uh, if they are Claude Code or Cursor users, they have usage view. And of course, people with admin privileges, they have an admin So if you are one of those, you can do the following, things. One thing that you should definitely do is the weekend test, meaning that if you didn't do anything with AI, but you look at your bill and you see that your bill keeps, compounding, you know that there are things that, are adding to your value without-- to your, bill without any value.



that was what happening to me. also look for very extreme input-to-output ratio. so agentic work legitimately runs, with high ratio, but if, you get to a point where it's many thousands to one, between input and output, in many cases, that's, empty loops. In the, in my OpenClaude case, it was, twenty-six hundred to one, which is, ridiculous.



and if you see that your spend keep rising while the work or the value that you do stays, flat, that's also potentially an [00:38:00] indication that you are in a scenario of spin, and you need to go and further understand what's the case. However, there are many folks that don't have Direct Meter because they're not using one of these tools or they don't have the admin privileges, which is probably most of the,regular users.



for them, you should probably use proxies. So just go directly to list all of your automation and the scheduled jobs that you own and ask, which one of them added business value last week. if you don't know, that's a suspect. Then also, watch your, quota. And if you're, burning through your weekly quota extremely fast, especially if you compare to other people in your, setup or in your, in similar roles, that might be that you're doing something wrong there.



And if you are in an enterprise plan, your admin do have the view at least of how much you are consuming and also the typically the input output that you can just, ask them. and there are many places that you can look. There are specific like, slash context and slash usu-usage in Cloud Code.



there is also, in, application visualization now both in Cloud Code and in Cursor that you can just click [00:39:00] on, the usage, meter and try to understand there. so regardless of what and how visible it is for you, you should definitely put some, caps on how much you spend rather than letting the bill just, extend all the time, and, put some alerts if there is some, kind of a, significant, jump in how much you consume.

That can be an indication that something is up in your system. So that's, for identifying spin. And now the habits that we should all, adopt to mind our tokens. These are, several things that anybody can do, immediately that, typically, improves the token consumption without, reducing the business value.



new task is a new session. this one is an interesting one because, we will talk in a minute about, also,model routers. But, least for now, for the most part, be intentional about which model you use for what, task. Sometimes it's actually going up to, like an Opus or even Fable class models because they will get the job done in one iteration and overall reduce the spend.

In some other cases, it's not, doing a web [00:40:00] search with Fable, but rather going to the Haiku or the lower, cost of models. rightsize your context. Tell your AI what it needs to know. This is a classical Goldilock, not too much, not too little, but sufficient such that it will not go into endless, internal reasoning, token loops just to try and understand what you're talking about.



build reusable capabilities. Often when we're, just,vibing with our model and trying to use it, ad hoc rather than sitting down and creating the skills, creating the proper automation, creating the proper agents, we're just wasting a ton of tokens to re-ask the tools to do something again and again.

So systematizing and building proper systems often is, one of the best levers that you have to, use their tokens, wisely. And, filter, everything that you can. Tell it in which rows of the table the data exists, in which parts of the, project board the data accounts for, which, Slack channels, and so on.

The more you point the model to the right place, the better the results that you will get. And lastly- in many cases, we start doing the work, we realize that the [00:41:00] model is completely off. Maybe it's the wrong model, maybe it's missing something. Don't let it spin. Just kill the job early and start again w-while understanding what you do.

And this is one of the cases where looking at the model reasoning will go a long way to understanding that it's completely off in the wrong direction. So I would recommend whenever you, send the model to start doing something, especially if it's a significant portion of work, open the thinking to understand what the model is, un-understanding from the task that you gave it.

And if it seems to be off, stop, improve the instructions rather than letting it, So that's the habits for everyone. two additional levers that you should consider, and, some of them are very, new. So if you are a Cloud Code user, you can use the /doctor command. This will basically, check not only how much,like a past installations take on your machine, but also how are your token divided, whether you have stale skills, stale, tool configuration, whether your overall instructions are overly long or overlapping.

So it's a very good,command that, Entropy created for us that you can go and [00:42:00] execute if you are a cloud user. If you're not a cloud user, you can just have your AI, uh, tool, investigate what the, the /doctor command does, and basically recreate it for your own tool, 'cause it's not, like a very complex, thing to do.

It just audits all of your system for you and gives you a structured report with concrete recommendations of things that you can kill because you haven't run them for a while, or things that are duplicated or stale or contradictory that you can potentially, reduce significantly. And, with regards to routing.



a lot of the industry conversations sit right now around the model routing, and you were just talking, I think today or the other day, around, some interesting, M&A around model routing. Um,picking the right model is still one of the highest return, things that you can do, even if you are able to use like the cursor automated router or some of the other solutions that are coming our way, because it's not always gonna be-- even if you have likea router in the background, it's not always gonna be [00:43:00] as precise as you knowing, which model to use.

And in many cases, you still don't have in your existing tool, a good enough or even an existing, router.

Yeah, I think that we are very early in figuring out the right patterns around routing. Obviously, there are a million solutions coming to market. They're all taking slightly different approaches. you have independent experiments from enterprises who are building their own systems that route between, you know, custom models that they've trained as well as, you know, the premier model.

Like, it is... There's no one clear approach yet, and even when there do start to be clear, use cases and patterns, they may not fit everyone in every use case. I think it would be entirely unsurprising to me, or I expect that routing norms around certain types of software engineering get solved first, because it's more deterministic and clear, and, you can kind of actually have, you know, more sort of verified success or not.

I think when it comes to knowledge work tasks more broadly, [00:44:00] it's going to be immensely more complicated, especially considering how much of our personal model routing that we do right now is about, not what the benchmarks would say on a test, but how we like the particular nature of one type of response versus another for a particular context.

So I continue to believe that understanding different model capabilities and having model preferences is still a very high leverage activity and is going to be for quite some time.

I agree. And I think the ultimate test was when GPT 5 was auto-automatically routing us, and all super users or just, like more than a, than occasional users, we were all very frustrated by what we got, from the auto mode. I think that's the original test that we want control, and we will probably even with a great router for many things will continue to be opinionated, and rightfully so.

just, uh, for people who are also,building their own, obviously they have additional levers, like you can, create more caching and so on. but still, for them, it's much the same, physics. Like the more control you have,[00:45:00] the more you are able to, be smart about, the way you use the models.



that's the additional levers. we talked about, tokens that, teach, and I think that, up until now we were very much focused on, things that we can do to reduce the bill. But here I want to, fight, the good fight and say that we want to protect those tokens, 'cause those are, in many cases, the tokens that you spend in order to get much better return.

And it's not just about optimizing the bill to go, downwards, but rather to improve also the return that we're getting. And often to improve the return, we need to improve the tokens that teach. And we're talking about two ways, whether it's you teaching yourself, meaning that you run, the same task using three different models in order to get to this taste of which,models you like for each task, or you try the same task in three different ways un-until you learn which one works best, or you experiment with a new tool, or you try a new skill or a new automation, and it doesn't work, and you try something else.

So all of these, typically gets you overall to much better [00:46:00] results from AI, so those should be protected, firstly. And also the other side of you teaching AI who you are, building the systems, adding more context such that you will get much more personalized results or much more, organizational aware results.

Those are almost always, with direct correlation to how much value you get from AI, and so does, data. I've seen a study of twenty k developers that found that the heaviest AI users, were roughly twice as productive in terms the, of the amount of production code that was shift. so i-in many cases, it's actually becoming, much more sm-- like a smart user will get you to better results.

So weSo to summarize, what you need to do, in two sides. So for the individual users, these are the things that you should definitely do. go and see whether you have tokens that spin. I'm sure that all of us have those, those idle automations, or maybe some of us have, like an even worse scenarios of the amount of tokens being s- spinned [00:47:00] without any business value.



practice those six habits. You can even them on a Post-it and just get yourself to work more effectively with the tokens that you have. do spend the time to invest in reusable capabilities and improved context, that the model can be much more, selective and, discover the relevant context where it matters.



I also want you to,audit, the things on a schedule, meaning regularly go back to the system and see, um, what is now stale or maybe something that was working well has become a stale automation. Maybe you need to improve the context, the instructions. Maybe you can remove some of the instructions per the new, advice coming from Anthropic that the modern models need fewer instructions, not more.

And make sure that you protect the learning budget, and as needed, go and negotiate that with the people responsible for the budget to make sure that, you're not now being reduced to the amount of tokens that, leaves you with very little room for exploration. for the organizational side, make the usage visible, and then teach the [00:48:00] people, because when managers and employees see their own they, are much smarter about how they use.

But make sure that they are not being encouraged to spend as little as possible, but to spend smartly. And also, make sure that the budget is by workload and by individual. if someone is building skills and context and usable capabilities for their entire team, they need to get significantly higher budget than the person that just uses the tool as a extended Google.

And all the time, we need to make sure that it's, by,that you tier it up, and some organizations, and some individuals get, significantly higher, others potentially less, and not just one size fits all for the entire organization. And make sure that everybody listens to something like that or that you do an internal, um,training that, teaches people on how to be smart about tokens, but not how to spend as little as possible, but al-also how to be mindful about the ROI and aiming to use tokens for the things that move the needle for the company.

So that's the concrete, actions for you and the team. And if you want to be [00:49:00] even more token smart, so beyond the audit of your own usage, we created for you, a token gym that you can go and, learn and flex your, token, smart, muscles. And if you want to go even further, and to learn how to build and work with AI and agents properly, we do have our,existing, trainings, and the next, cohort start on, early September.

So we'd love to have you there, in the executive catch-up or the executive agent leadership that will, bring, you all the way to be very smart about AI or very smart about agents, depending where you are That's it

Awesome. I, yeah, look, I think that this We're always at the beginning when we're talking about things, this show, but this one is, I think, particularly inflection pointy, let's say, to use a word that doesn't exist. We are so clearly just at the beginning of figuring out how to organize the relationship between people and the compute and intelligence that they're going to consume, and it is going to be [00:50:00] iterative and messy, which is why I think so many of these ideas that you've presented are shared as frameworks, you know, patterns to explore, right?

it's a set of steps that you can take to try to get a handle on these problems, but every organization at the beginning is going to solve them, or not in different ways. so thank you for sharing, some starting points and, you know, we'll continue to evolve this conversation as the, the tools around us change too.



​
