// Wednesday · October 7, 2026

The Best Way to Test New AI Models

A webinar edition for the fall model glut: with a new frontier model landing roughly every 11 days, published benchmarks tell you less than ever about whether a release matters for you. NLW and Nufar Gaspar walk through a five-step system for building a personal AI benchmark — your real tasks, blind-scored, with a switch/split/stay decision at the end — and Nufar's own results surprised even her.

Ad-free on Patreon
Today's sponsors — KPMG · Robots and Pencils · Harbor · Blitzy · all offers →
The One Idea

The only benchmark that matters is the one you build on your own work.

Published benchmarks are in the training data, and 'state of the art' can still be worse at your version of a task. There's no shortcut to knowing whether Opus 5.5, Sol 6.1, or an open-weight model fits your stack other than actually using them — so build a repeatable personal benchmark: pick four to six real tasks (including one wish-list item), run each model blind in fresh chats, score side by side, then decide to switch, split, or stay. Do it once properly and you can rerun it every time the frontier moves.

// 01

By the Numbers

11 days
Average gap between frontier model releases, per NLW
7
Models Nufar blind-tested — 'way too much' to score; she recommends up to four
4–6
Recommended number of benchmark tasks covering your real work
8
Questions in the discovery prompt that interviews you to pick your use cases
0/5
LinkedIn drafts from seven models that Nufar would post as-is
10×
Fable's cost premium over GPT Sol at equivalent personal scores
$500/mo
The plan tier for OpenAI's new ultra-fast mode
20 min
Nufar's wait for one website generation — why speed just became a criterion
// 02

The Brief

ModelsProductOps▶ 00:00

The fall glut makes personal benchmarking urgent

Opus 5.5, Astra 6, Sol 6.1, faster tiers like Sonnet 5.5, and a wave of open-weight models have all landed in a few weeks — each with benchmark tables that say nothing about how a model fits your personal AI stack. The fix is a system for testing new models against your own work.

AI Daily Brief
◆ The Take▶ 02:00

Be skeptical of published benchmarks — they're in the training data

NLW's standard caution: it's not that labs are lying, it's that most benchmarks are now part of the canon that goes into training sets. The longer a benchmark has been around, the less it tells you — even with a fertile wave of new benchmarks trying to fix this.

The AI Daily Brief
◆ The TakeProduct▶ 03:00

The state-of-the-art model can be worse at your version of the task

Even with perfect benchmarks, model choice is now about how a model feels and how it fits your stack, not strictly better-or-worse rankings. NLW guarantees there will be times the chart-topper underperforms a lower-ranked model on your work — especially for subjective tasks like writing.

The AI Daily Brief
EnterpriseOpsProduct▶ 04:00

The five-step personal benchmark

Pick the tasks (the most consequential step), pick the candidates — which can be models, harnesses, or tool-plus-model combos — run the blind comparison, score anonymized results, then decide. A benchmark run that favors a new model doesn't automatically mean switching: changes are hard and other considerations apply.

AI Daily Brief
ModelsMarketing▶ 07:00

Every launch follows the same arc — real signal arrives a week later

Announcement, benchmark table, then a week of hot takes and influencer toy demos (some suspect models are optimized to look good in exactly those tests). Around seven days in, reports from real work surface — like 'it quietly dropped a condition in my contract review' — which matters far more than a replicated '90s game.

AI Daily Brief
ModelsProduct▶ 09:00

The answer to 'everything's good enough': wish-list unlocks

The frontier is very good, and there's more value to extract without switching. But a new release can suddenly unlock a use case that was only on the wish list — which alone justifies periodically re-benchmarking, on your own schedule rather than weekly.

AI Daily Brief
EnterpriseOpsExec▶ 10:00

Even company-locked users should benchmark

If your employer limits you to one tool, there's still a choice — fast vs. thinking modes and other configurations. And running a benchmark on mock data can give you the evidence to persuade the people who decide that another license is worth buying.

AI Daily Brief
EnterpriseOpsProduct▶ 11:00

Pick tasks you'd actually stake decisions on

Choose four to six tasks that mirror your real work-personal balance: diverse, regular, high-stakes enough that you'd act on the results, and ideally including one where you know what a costly mistake looks like. Always include at least one wish-list task AI has previously failed you on — that's where frontier jumps show up first.

AI Daily Brief
ModelsOps▶ 14:00

Don't test seven models — Nufar tried

Rating seven model options across six tasks was 'extremely tedious.' The recommendation: your existing model versus one new candidate, or at most four options — one of which should always be your current baseline so you know whether a switch is justified.

AI Daily Brief
Models▶ 15:00

Hide the model names — you're more biased than you think

Run every request in a fresh chat, anonymize which model produced which output, score one-to-five with a one-line observation, then do the reveal. Guessing which model wrote what is part of the fun — and Nufar guarantees some results will surprise you.

AI Daily Brief
EnterpriseEngOpsLegal▶ 16:00

A shared script automates the blind test — but mind your data

Nufar's script uses an OpenRouter API key to run all candidates, anonymizes outputs, and serves a locally hosted scoring UI. One warning: if your company is sensitive, don't push company data through OpenRouter — use mock data and mock use cases instead.

AI Daily Brief
ModelsProduct▶ 17:00

Include at least one visual task

Website design, app design, or image generation is far easier to score quickly than walls of text — if you actually use AI for visual work. Store the outputs in a dedicated folder so you can open each one and compare properly.

AI Daily Brief
Models▶ 18:00

Comparative scoring beats absolute scoring

Giving a standalone output a one-to-five score is hard; seeing several alternatives side by side and ranking most-to-least liked is much easier. And your taste is the primary judge — there's an X factor to preference you often can't verbalize.

AI Daily Brief
ModelsEng▶ 20:00

AI nepotism: judge with a model from a different family

Models often favor outputs from their own family, so the AI judge should come from outside the set being evaluated — Nufar used Gemini Pro since no Gemini model was a candidate. And evaluate the judge itself: if it doesn't correlate with your taste, throw it away.

AI Daily Brief
ModelsEng▶ 21:00

Same model, same prompt, different answer — so run it several times

If the benchmark informs a decision with real monetary stakes, run each request multiple times per model. Nufar saw it that day: a website she loved on the first generation came out much worse on the second, identical run.

AI Daily Brief
EnterpriseExecOpsFinance▶ 24:00

The decision is switch, split, or stay — and habits are expensive

A winning benchmark result doesn't mandate a switch: company policy, terms, plan limits, cost, latency, and beloved features all weigh in. Habit formation is costly, so without a significant boost, staying with your current setup has real merit.

AI Daily Brief
EnterpriseOpsProduct▶ 28:00

A prompt interviews you to build the benchmark itself

Run the shared discovery prompt in the tool with the most memory of you; after about eight questions it proposes your use cases, writes the benchmark prompt for each, and drafts a scoring rubric for the AI judge. Critical step: refine anything generic — if you don't trust the tasks represent your work, it's not a good benchmark.

AI Daily Brief
ModelsMarketingProduct▶ 35:00

The reveal upends the brand assumptions

Kimi, Fable, and Opus topped LinkedIn writing — Nufar was sure GPT would be there. GPT models won the pricing negotiation, GPT and Groq won exercise creation, Groq won week planning, and Opus 5.5 won the website. No model earned a five on the LinkedIn task — none was postable as-is.

AI Daily Brief
BusinessFinanceOps▶ 36:00

The judge's winner cost an order of magnitude more

Fable topped the AI judge's rankings but at a total cost far above everything else — and on Nufar's own scores it merely tied GPT Sol, which was roughly ten times cheaper. Kimi won on both cost and duration; Groq was cheap but extremely slow. The leaderboard isn't the decision.

AI Daily Brief
ModelsEng▶ 37:00

Human and AI judge agreed on exactly one task

Nufar and Gemini Pro aligned only on the pricing pushback — everywhere else they diverged. One structural reason: an API-based judge sees the website's code, not the rendered page, so it's a poor evaluator of anything visual. Her conclusion: she'll trust her own scores.

AI Daily Brief
ModelsExec▶ 38:00

The benchmark actually changed her mind

Nufar admits she's biased toward Claude and Cursor with Groq day-to-day, but based on her results she's contemplating moving GPT into her top two — and if forced to pick one tool today, she might consider going primarily GPT.

AI Daily Brief
EnterpriseEngOps▶ 39:00

At team scale, graduate to real eval platforms

For individuals, the manual or script approach is plenty. Teams doing standardized, regular evaluation should look at dedicated platforms like LangSmith, Braintrust, and Langfuse for systematic experimentation, monitoring, and evals.

AI Daily Brief
ModelsProduct▶ 42:00

Ultra-fast mode is resetting expectations about speed

NLW notes Sam Altman said he didn't appreciate how much speed mattered until he had OpenAI's ultra-fast mode — currently a $500-a-month feature. Nufar felt it immediately: after seeing the demo, waiting 20 minutes for a website generation suddenly felt unbearable. Your benchmark criteria will evolve as the frontier moves.

AI Daily Brief
ModelsEngFinance▶ 45:00

Open source is now good enough for most knowledge work

For table-stakes tasks like email drafting and basic research, any decent model — 'Kimi and above, or Mistral and GLM and above' — shows no noticeable difference from commercial frontier models, especially in a good harness. The gap persists mainly on the most sophisticated tasks.

AI Daily Brief
BusinessFinance▶ 47:00

Your $200/month plan is hiding the real token bill

Generous Pro and Teams tiers subsidize a much higher actual token cost, so per-task cost comparisons don't tell the full story — unless you hit rate limits or pay per consumption. Nufar's prediction: the generosity may not last, and then everyone will care about specific costs.

AI Daily Brief
Machine-readable ▸Download .mdTranscript .md— feed it to your own agent

Got this from a colleague? Get the brief every day.