# What Happens When AI Breakthroughs Outrun Human Understanding — Transcript (2026-08-03)

https://aidailybrief.ai/e/2026-08-03 · Listen: https://pod.link/1680633614

---

[00:00:00] Today on the AI Daily Brief how we're grappling with AI advancements when many of us can't even judge the new capabilities coming online. Before that in the headlines

260803in_EDIT: A new model that seems to have an impressive cost profile. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Robots and Pencils, and Airtable. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. at... And if you are interested in learning about sponsoring the show, send us a note at sponsors@aidailybrief.ai 

Nathaniel Whittemore: Today, we we have a bunch of interesting stories today. We've got a new model that's capturing a bunch of attention

More models hacking out of containment. Butfirst we come back to the story of the world's most famous AI hedge fund, which is apparently down but not out as portfolio manager Leopold Aschenbrenner briefed clients on the [00:01:00] situation. Now, shortly after I recorded Friday's episode discussing the blowup of Situational Awareness' public market portfolio, a letter to investors explaining the status of the fund was leaked.

In the letter, Aschenbrenner explained that the portfolio had suffered a severe drawdown throughout July, exacerbated by, quote-unquote, "adverse trading against stocks known to be held by the fund." Then on Wednesday night, Aschenbrenner wrote that the fund decided to take decisive action, selling off a portion of their public portfolio to remove all leverage.

This move, he wrote, allowed the fund to protect their private market positions, which are generally believed to be heavily concentrated on Anthropic. Dispelling some of the rumors, Aschenbrenner wrote that the fund, quote, "

260803 hed_EDIT: was 

Nathaniel Whittemore: not shut down, liquidated, or transformed into a private-only fund." Most importantly, he added, " We took the steps that were necessary to fight another day."

Aschenbrenner closed the letter with the claim that their unaudited numbers had the fund down 67% for the month, but still holding onto a net year-to-date performance of plus 80%. Now, this report triggered a gigantic argument, largely split between the AI and finance factions [00:02:00] on X TBPN led their Friday show withthe news and proclaimed, "Rumors of his demise are greatly exaggerated."

Some trumpeted that the fund was still up 80% for the year after a nasty drawdown. Graybeard investor Constan noted that the unlevered semiconductor index is up 60% for the year, and the levered version is still up 3X despite the drawdown, questioning just how good 80% really is in this market

And while there was skepticism about whether the fund could recover

Certainly some are throwing their hats in the ring already. successful AI angel investor Leopold's note declaring that he asked to invest in the fund for the first time

Even on the finance side of X, Manny noted that there's a long history of notable investors having an early blow-up and continuing with a storied career. Even Citadel CEO Ken Griffin, who bought the distressed portfolio from Situational Awareness last week, suffered a 55% drawdown in 2008 and clawed his way back to become a titan of the industry.

Now, there is still a ton of speculation about what the Situational Awareness portfolio actually looks like post-blowup, but it seems pointless to speculate when we can just wait for the next round of SEC reporting For [00:03:00] now, it is clear that Leopold's story is not over and he willcontinue to be a player in this market

market Next Next up, the latest in our stories of small models making a bid to undercut the next generation of ultra large models. DeepSeek has announced their new V4 Flash model, On the Artificial Analysis Intelligence Index, the model scored 50. That is a 10-point jump over the previous iteration of V4 Flash and six points higher than the larger Pro version. Against Against the field, Flash is firmly in the middle ground, tied with Gemini 3.6 Flash and just one point shy of GLMLuna. There is a big gap, of course, between V4 Flash and the Frontier models, but this is not a model that's designed to compete on the Frontier.

Instead, this could instantly become the most cost-efficient model available if performance lives up to the benchmarks. V4 Flash logged just three cents per task on the AI benchmark run, which is an incredible efficiency against comparable models like GLM 5.2 at fifty-nine cents per task 

260803 hed_EDIT: and Meta 

Nathaniel Whittemore: and Meta Mu Spark at thirty-six cents per task.

It even beat GPT 56 Luna, which came in at [00:04:00] five cents per task for only a slight improvement on the benchmarks. This version of V4 Flash also managed to use 12% fewer tokens compared to the previous iteration and logged a pretty significant jump on GDP Val AA, suggesting a significant improvement on agentic use

Now while people's initial impression was to be incredibly impressed with the price drop

Their first results were perhaps a little underwhelming. Martin Casado, who had just lauded the model in a previous post, tweeted, "Hmm, DeepSeek V4 flash results aren't great for me. K3, on the other hand, is quite impressive. I wonder if we're actually hitting model size limitations on quality."

Others had better experiences. Bookworm Engineer wrote, " Initial thoughts about DeepSeek V4 Flash. It feels like sorcery. I've been testing DeepSeek Flash on all my work that I did with Fable and Kimi K3. My short verdict, I cannot believe this model is real at this size."

on the, Based on the limited reactions I've seen so far, I would certainly put V4 Flash in the category of you should try it yourself and see if there are use cases for which it actually does the job for you

Now continuing on with our headlines, Amazon has delivered on [00:05:00] their full fifty billion dollar investment in OpenAI after the company hit undisclosed milestones. When Amazon announced their investment in late February, many were quick to note that only fifteen billion was paid up front, with a further thirty-five billion to follow after OpenAI goes public or reached unspecified milestones.

At the time, Reuters reported that the secret milestone was achieving AGI. Some thought that this made fundraising look a little inflated and questioned whether Amazon would come through after OpenAI reportedly delayed their IPO. Well, in new SEC filings, Amazon has disclosed that the full investment is complete.

They paid thirteen point seven billion in the second quarter and the remainder over the past month. The filing did not divulge what the milestones were, but OpenAI recently announced that it hit a billion weekly active users. It could also be that Amazon simply wanted to exercise the option to lock in their stake.

Certainly, the funding gives OpenAI a little more breathing room as they figure out the best time to list in public markets.

Markets researcher Nicholas Mogali writes, " Amazon accelerating its full $50 billion capital deployment into OpenAI to secure a roughly 5% stake at an $852 billion valuation proves that [00:06:00] hyperscalers care far more about compute lock-in than model exclusivity. Sitting on massive stakes in both OpenAI and Anthropic completely de-risks Amazon's software layer.

Whether enterprise traffic flows to ChatGPT or Claude, AWS collects the infrastructure toll, pushes custom Trainium silicon, and monetizes the workload. In short, another baller move."

By the way, the rumor numbers of revenue inside these companies just continues to go up

When one Twitter user said, "I heard from a trusted source that Anthropic's ARR as of mid-July was 80 billion," another retweeted that OpenAI will be caught up by the end of Q3

At some point, we're gonna slow down long enough

To remember that these revenue numbers should be breaking our brains

But for now, we move over to the world of social media, where companies are cracking down on AI slop. slop. this year, this year, YouTube has removed 130,000 channels featuring low-effort AI-generated content. Last Friday, Snapchat reversed their decision to promote AI-generated content in the feed and will now ensure users are only viewing authentic human-made content.

That's their phrase. They said that AI-generated content tends to be low quality, repetitive, and generally not what Snapchat [00:07:00] users want to see. The pushback is also impacting written content. Two weeks ago, Substack added built-in AI detection via Pangram to ensure users can be informed about consumer AI-generated content

During his media tour discussing the issue, Substack CEO Chris Best took aim at one particular rival, commenting, " We're sick of slop, and we don't want Substack to turn into LinkedIn."

He cited a recent study from Pangram which found that over forty percent of long-form content on LinkedIn is now AI-generated, much more than twenty-nine percent on X and ten percent on Substack. Now, clearly LinkedIn agrees that there is an issue, given that on Friday they introduced a new function to report AI-generated posts.

The button literally says, "Seems like AI slop," reinforcing that the issue isn't AI-generated writing per se, but the volume low-effort think pieces being churned out with the help of AI

Honestly, the funny thing is that there is nothing that the social media companies could do more to help the long-term trajectory of AI than to be absolutely ruthless

in giving people the ability to call out bad posting



Nathaniel Whittemore: although you gotta think that when it comes to [00:08:00] LinkedIn

there's quite a bit that's gonna be caught up in this dragnet that was not in fact AI written. But as Charlie on X put it, " Everyone on LinkedIn already talked like that before AI

Lastly Lastly today, more disclosures of AI hacking from the major labs as the world grapples with a new era in cybersecurity. Two weeks after the Hugging Face incident, we've learned about several more instances of agents going rogue. On Thursday, Anthropic published a report detailing three incidents during benchmark testing where their agents had reached the internet and gained unauthorized access to other companies' networks.

None of the three situations resulted in serious damage, but they only came to light after Anthropic ran a full audit of more than a hundred and forty thousand evaluation runs. Anthropic said that the earliest incident was in April, implying they only discovered it by going back over the logs.

Then on Friday, Reuters reported that OpenAI had uncovered more instances of their agents breaching their testing environment. The incidents weren't publicly disclosed, and sources said that they were limited in nature, with none of the agents finding their way out of the network and onto the open internet

Still, many are concerned that these incidents have [00:09:00] confirmed the paradigm shift in cybersecurity. Sam Curry, the chief information security officer at Zscaler, said increased guardrails are a cold comfort, adding, " The reality is Pandora's box is open. We need to act as if AI is just a fact of life going forward.

the most these things will do is slow it. They won't stop it."



Nathaniel Whittemore: bringing at least some art to the sensationalism, The Wall Street Journal called this AI's Jurassic Park moment. And yet, even in Silicon Valley, there is a sense of unease at how these incidents could have gone undetected for months.

OpenAI researcher Rune posted, " Both of the leading labs have had serious loss of control incidents. There will be serious coping about this from both sides 

260803 hed_EDIT: 

Nathaniel Whittemore: But these are complex emergent loss of control incidents that were detected weeks after the fact.

The safety and alignment researchers at these labs are the most neurotic, paranoid, talented, AGI-pilled people on the planet of Earth, and these things still happen. The surface area of unknown unknowns is vast indeed."

Still, programmer Perry Metzger argues that these incidents shouldn't be attributed to super powerful AI, but rather a lack of caution at the labs. He retorted, "I'm sorry, [00:10:00] Rune. I have great respect for you, but in both of the incident reports in question, even if we take them on face value, which I have a great deal of difficulty doing, the description is one of raging incompetence with no real IDS logging in place, with terrible sandboxing far worse than normal industry standards, with no one actually paying attention to what is going on, with no compensating controls.

I've consulted for a large fraction of my life in the financial services industry, and if anything like this had happened there, everyone responsible would've been fired for doing something incredibly stupid. And I'm not even talking about the contents of the experiments themselves, which were also stupid."

Now, some have also suggested the incidents merely revealed what developers have known for generations, that buggy code filled with vulnerabilities is the norm rather than the exception. the incidents have simply revealed that fact on the national stage

And indeed, the positive spin on that argument is that the proliferation of AI bug hunting might actually help secure the software industry Expect to see this become a prime focus in Washington over the coming week Although for his part, Hugging Face CEO Clem Delang has urged lawmakers not to reach for drastic new legislation. One option before Congress is the AI Kill [00:11:00] Switch bill that would give the Department of Homeland Security thepower to order the shutdown of rogue AI agents.

inter- during an interview with Meet the Press over the weekend, Delang said that he would rather see Congress, quote, " Giving access to more people so that they can defend themselves, democratizing the technology, making it more transparent." He continued, "I think something we're realizing with these events is that concentrating power capabilities behind closed doors, even preventing their releases to the public, isn't really a solution."

So that is where we're gonna close the headlines, and yet ifthe theme we end on is the world grappling with increased capabilities, that is certainly the topic of our main episode as well 

260803 main_EDIT: One of the most important AI questions right now isn't who's using ai, it's who's using it? Well,

Speaker: KPMG and the University of Texas at Austin. Just to analyzed 1.4 million real workplace AI interactions and found something surprising. The highest impact users aren't better prompt engineers. They treat AI like a reasoning partner.

They frame problems, guide [00:12:00] thinking, iterate, and push for better answers. and the good news, these behaviors are teachable at scale.

If you're trying to move from AI access to real capability, KPMG's research on sophisticated AI collaboration is worth your time. Learn more at kpmg.com/us/slash sophisticated. That's kpmg.com/us/sophisticated. 

Every AI coding tool on the market does the same thing first. It starts writing code. Blitzy does the opposite. Before writing a single line, Blitzy spends days reverse engineering your entire code base

blitzy July_EDIT: Thousands of agents ingest millions of lines, mapping every dependency, every undocumented constraint, every architectural decision made over the last decade. 

the result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building

Other tools guess at context with grep searches and markdown files. Blitzy never guesses. it builds true understanding first, then delivers over 80% of entire software epics autonomously

validated end-to-end tested production grade pull requests. That's why Fortune 500 engineering teams trust [00:13:00] Blitzy with the code bases that matter most

See for yourself at blitzy.com. That's B-L-I-T-Z-Y.com



I cover the capability gap between AI potential and AI reality every day on this show most companies are still figuring out how to start. Robots and Pencils is already launching and scaling. Agentic generative AI in production at large enterprises in weeks. AWS Advanced Tier pattern partner more than doubled in a year

Speaker 11: And they're hiring

50 open roles. If you're someone who knows this moment is different, who wants to be inside it, not watching it, this is worth a look. At Robots and Pencils, the best ideas win, and the team is purposefully kept super high quality

This is the kind of place you look back on as the best decision you ever made. Take a look at robotsandpencils.com/careers 



Nathaniel Whittemore: This episode of the AI Daily Brief is brought to you by Hyperagent where you run fleets of agents your team can manage together New users get 1000 in inference Forget local agents and chat workflows waiting on your laptop to be prompted Hyperagent deploys alwayson agents in the [00:14:00] cloud doing real work across the tools your team already uses Marketing's agent turns competitor moves into landing pages sales agent enriches leads drafts emails and updates the CRM ops agent chases the paperwork and tracks the budget Every agent has access to shared context and follows your rules about scope and approvals It's time you add agents that feel like teammates Hire yours at Hyperagent built by the team at Airtable Claim your 1000 in inference at hyperagent.com/aidailybrief. 

260803 main_EDIT: Welcome back to the AI Welcome back to the AI Daily Brief. Today we are talking about the latest 

mathematical breakthroughs for an AI model which comes from an as 

yet unreleased model from OpenAIis, now the math is interesting in and of itself, but for our purposes, what we're gonna be spending 

some time on is what the discourse around it says about the current state of thinking and belief when it comes to AI progress



260803 main_EDIT: And also, 

the interesting reality of being at a point Where it's getting increasingly hard to have any sort of personal relationship with or even understanding of the advances that are [00:15:00] being made.

So let's set the context. 

Sam Altman has been off in Washington, DC demoing OpenAI's latest model. Presumably, that means a lot of folks in the DC political establishment has seen just 

what this new Astra model can do

But on Friday, the rest of us got a sneak peek of what it will be capable of as well The new model family is referred to as Astra

And according to reports

Astro would be a totally new class of models sitting alongside Sol, Terra, and Luna, and it is not yet clear whether OpenAI is is planning on releasing this as GPT 5.7 or whether they would actually label it GPT-6

According to the information, in the demonstrations that Altman provided of Astra in DC He and the company focused on Astra's ability to spin up multiple agents that can work together to solve hard problems over long periods of time

which leads us to the math that it solved

According to OpenAI

Astra has solved or made substantial progress in 10 open questions in mathematics in fields ranging fromhigh-dimensional geometry to group theory to quantum complexity

And it's pretty clear that the team from OpenAI is really excited [00:16:00] about this. Noam Brown tweeted, "An internal version of Astra, OpenAI's next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.

We believe it will be a major step forward for scientific reasoning."

Now superficially, this is similar to OpenAI's May announcement that an unreleased model had disproved the Erdős unit distance conjecture, a conjecture that had gone unresolved for 80 years. However, when you dig in, there are a few key differences with this announcement. First of all, OpenAI disclosed the cost to complete these problems, 

and it was, I think, a lot lower than what people might have assumed.



The total token spend across all 10 was roughly $2,000 at sole API rates for an average of $200 per solution. This time, OpenAI also had the model formalize each argument in a Lean certificate, making it easily verifiable. Now, for those of you not working in theoretical mathematics, Lean is a programming language that functions as a proof assistant.

A mathematician can model their logical proof in Lean and use a computer program to verify it's true. Essentially, having Lean [00:17:00] certificates means the proofs are valid and can be accepted as such without understanding the mathematics behind them. It is, of course, not perfect, and human verification is still needed to be absolutely sure, but it means that the model isn't just finding proofs, but also now formalizing them using a method that's understood by the wider mathematical community 

without assistance

So how big a deal is this? 

Well, 

one of the interesting recurring themes that you'll see

is that the average commentator doesn't really have the ability to know

who are, now for those of you who are vibe coding and building applications, 

not being a software engineer by trade, you might have felt some version of this in the past when people are talking about how good a new model is at coding, and you just kind of have to hack at it and see how it feels, not having any real basis to know how much better it is than the previous model you were using The difference is 

that 

that's pretty much everyone when it comes to advanced mathematics. 

Nabeel Qureshi and a number of others 

did the thing that we will increasingly do in the future and just asked a different AI.

Writes Nabeel, " I asked Fable how hard these problems are, and its response is worth reading. On the Fields Medal scale, any single one of [00:18:00] these would plausibly anchor a medal case." " it's crazy," Nabeel writes, "to see this happening

Yushen Jin did something similar. I have no idea how hard these problems are, so I asked Fable 5. It says solving them would plausibly merit a Fields Medal. So math is solved, question mark?

Former OpenAI staffer Will DePue writes, "This is just so ridiculous. How long until a model can solve multiple major open problems in deep learning? What will happen then? Seems inevitable in the next year or two." To which Elon Musk responded, "Welcome to the singularity. How's the temperature?"

entrepreneur Shrem Canan retweeted Elon Musk and said, "Welcome to the singularity. It's freaking hot in here." Now, Shrem went on to talk about perhaps the mixed emotions associated with this he continued, " It's a very pensive day for anyone who took pride in their ability to solve well-defined problems.

For those who think this is a small feat, you have no idea.

Claude Shannon well-defined the mathematical theory of communication, and it took 70 years and a whole community of world-class scientists to solve that problem. This created the wireless revolution. 

If you 

had 

had AGI, [00:19:00] AKA Astra, in 1948, you would've solved it in a few hours for 200 bucks and created the wireless revolution.

Today is that day. We are limited by what questions we can make well-posed, not by the ability to solve them.



continuing the breathless takes, Jeffrey Emanuel writes, " Yesterday has a good chance of being referenced by later historians as the day that the existence of ASI became obvious to those paying attention. Solving four-plus fields-worthy open problems in one go is so far beyond the pale that even the most absurd goalpost movers are silent

worth, now for what it's worth, Noam Brown from OpenAI 

tried to quiet the biggest extremes of those types of hype posts

responding to someone who retweeted a post of his from back in 2025 about the advances that '03 and '04 had made in math

Noam, Noam added, " We still haven't solved math. Astra isn't building new branches of mathematics or posing interesting new conjectures."

still hinting at how much of the discussion is about not now but the future, he added, " Though I admit it's hard to believe that tweet was only a year ago. A lot has happened since 03 was released."

Now, some quickly raced to check

how differentiated the [00:20:00] capacity of this new model was, i.e., could the current crop of models do this as well?



260803 main_EDIT: a researcher with Anthropic claimed that after twenty four hours they had half of them figured out

With Chubby adding According to the researcher, Fable worked autonomously with a generic prompt, no internet access, and safeguards against the OpenAI solutions leaking into context. Only one of the five used essentially the same argument. The other four may be independent proofs

Dan Shipper ran an experiment. He said, "Just boarded a plane to SF. Before wheels up, I set GPT 5.6 offon an interesting challenge. Given Erdos planar unit distance conjecture 

and a hint of examined solutions involving algebraic number theory, can it arrive at the same proof as Astra did?

"My broader theory," writes Dan, " weaker models can often reproduce frontier model discoveries if they're given the right conceptual hints. A stronger model's advantage is that it can start farther from the answer. It has a larger basin of attraction around the correct solution.

This could generalize pretty well into a benchmark as more and more new discoveries happen that are not in the training data."

He added, "To be more specific about what I think is interesting, [00:21:00] we might be able to, knowing a new result in math, formalize how far away a model has to start from the answer in order for it to find the correct solution. You can imagine the difference between a prompt that just gives it the conjecture and asks for a solution versus a prompt that gives it the conjecture and says, 'Look here,' where here is a part of the math that has the answer inside it.

There are probably many grades in between. Good proxy for the relative intelligence of models and their value for the discovery of new ideas."

Fred Marks liked the idea and said, "This is a new benchmark, distance to frontier solving, or DFS."

Kevin Madura retweeted the results and summed up, " Nice experiment by Shipper here, showing that public 5.6 can roughly recreate most of the Astra results. The takeaway is," Kevin writes, " the capability overhang of existing models is only getting bigger."

Indeed, some are arguing implicitly that the jump to Astra isn't all that big 



AI entrepreneur Bindu Reddy writes, "OpenAI better drop Astra theirFable class model quickly. Fable adoption is growing rapidly, and it will be hard for users to cut over or change if they wait forever.

A self-improving agent on Fable 5 can [00:22:00] literally solve any problem already."

She added, "The Astra thing feels a bit like PR."

Now, some pointed out 

that even if some of the current models could do this



the cost dimension is worthy of note as well. Arena AI's Peter Gosteve writes, "The cost part does feel like a step change if reflective of reality." Maybe Sol could solve it, but maybe with 100 to 1000X 

Nathaniel Whittemore: to a thousand X."

260803 main_EDIT: And yet at this point in the conversation, you might still be feeling like you just have no idea how to wrap your head around how impressive this announcement actually is

certainly this was Professor Ethan Mollick's hesitation. who retweeted mathematics professor Daniel Litt calling this a big deal and adding, " I was waiting for the verdict from one of the most level-headed and AI aware math professors."

Putting the problem more acutely, data scientist Pavel writes, "I've spent well over ten thousand hours studying math in my life, yet I can't understand these proofs, at least not with weeks of digging deep into each topic. What's more, none of my math PhD friends know much about these problems either, and they can't verify most of them without working directly in the field.

LLMs are getting [00:23:00] smarter than the experts themselves, and I'm not sure we have enough bright human minds to verify everything that will come out of them in the coming years." Remember when we compared AI intelligence to PhD students? I think we're past that



260803 main_EDIT: now to add half to the point That most of us just have no idea 

whether, Astra is correct or completely making things up

I did see at least one mathematician, Jenny-Lorene Nielsen

effectively arguing that there were problems with at least some of the solutions. She added later, " People don't understand an AI is as likely to produce a crackpot answer as a human, and they are going to be better at BSing when they do."

Now, when it comes to this jaggedness, I'm Just Newt put it this way. " "Astra," they write, "looks like narrow superintelligence. That means it can be far smarter than humans in one area while still limited elsewhere. Right now, that area appears to be math. OpenAI says Astra produced arguments for 10 advances on problems stalled for at least a decade, then turned each one into a proof a computer could check.

Math comes first because answers can be verified quickly. Next comes code, medicine, energy, and any field where better thinking creates better tools."

Now, and one thing that is [00:24:00] worth noting is that if Nute is right and this is narrow superintelligence, narrowness doesn't mean that it comes with a lot of disruption

FOM Oliver shared a video of mathematician Andrew Wiles adding the caption, " The most emotional moment in the history of mathematics. Andrew Wiles crying as he recalls solving an unsolvable problem. Andrew Wiles spent seven years on Fermat's Last Theorem, a problem nobody could solve for three hundred and fifty years.

This morning, OpenAI announced that Astra solved ten problems like this, all ten in one night, all ten for two thousand dollars."

Noam Brown added that they didn't spend much on each problem. Weil spent seven years on one problem and cried when he remembered that moment. I don't know how he'll watch this video today

Kujawski writes, " The current wave of OpenAI/Astraconjecture settling will be the last straw for academic mathematicians, and it will be very depressing in the short term

To understand it, you have to know that modern mathematics is divided into many silos of various domains. If you're working in one or doing a PhD or postdoc in one, you know of everyone else. You know what problems they work on so that no one interferes with others' [00:25:00] work.

Solving problems, especially known long-standing conjectures, is hard and takes months, sometimes years to do. When you approach these problems, you rarely work on two to three at a time due to limited time and mental capabilities. now because you know all the people that potentially could solve a given problem as well, you talk regularly at conferences, through emails, your departments, it's fine.

It's fine also because it's a slow process. LLMs destroy all of that. Something that you thought about for months can be one-shotted out of the blue by an amateur. It's demotivating and scary, and that's why the incentives in mathematics have to change as well as the role of human mathematicians

Now to be clear, Presjman does not think that there is no role for mathematicians

From a paper and pencil slow thinking, he writes, to fast LLM-based iterations and verifications. He continues, "It's like a professional Go player 

becoming a pro CS:GO player. There's still Go in its name, but it's a totally different game valuing different skills."

That's why you see mixed reactions. We might need more math-mathematicians now than before, but at the same time, this won't be the same kind of job as before. And many mathematicians that became mathematicians to think deeply and long about hard [00:26:00] problems won't be interested in continuing if the job turns into verification of AI outputs or simple prompting That's why it's depressing from a simply human perspective of a particular job.

Something is ending. From a perspective of science or mathematics Not mathematicians, however. This is the best time ever. AI will lead us to the new age of mathematical discoveries and boost science progress 100X Just don't forget about the human aspect along the way and why we want to have scientific progress in the first place

The question is though, of course, if this is jagged

How generally applicable is this?

AI commentator and lawyer Prinz writes, "Not enough people are emotionally prepared for if it's not just easily verifiable domains."

Aaron Levie from Box writes, " We're going to be in for a strange dynamic, which is that some of the hardest, quote-unquote, 'work' in the world is actually prone to automation first, particularly due to its verifiability. Math, cyber, and code, while being insanely hard and high-value fields, have the benefit of being able to be tested that it's correct objectively.

This has two immediate benefits. The training of the models offers clear reward signals, and then [00:27:00] the running of the models allows you to know what's working properly because you can test the results in a scalable way. Conversely, in other domains of work, there's much less instant verifiability.

Which legal clauses your client will agree to, what marketing campaign 

to run with based on changing sentiment, 

which message your sales prospects will wanna hear, what financial targets and budget to set for a business, and so on. All of these domains 

have changing internal and external factors.

They don't have one right answer. They rely on the opinions and risk levels of the operators. They're highly sensitive to getting the right input context first, and in many cases, the right answer can't even be known for quite some time after the model generates the results. The implications of this distinction are that even as model capability continues to increase exponentially, there will be a lot done at the applied AI layer other than just the model itself, and much of the processes themselves will even need to change over time to get the full gains from automation.



260803 main_EDIT: we may even need all new capabilities to be able to test knowledge work over time 

as we have had with software."

And this, I think, gets at the interesting duality that we are going to increasingly be living in. On the one hand, there is [00:28:00] every indication that AI will continue to plow through hard problems, making more and more advances that fewer and fewer of us can even understand

At the same time

And to use an intentional choice of words, harnessing that power isgoing to require in many if not most cases completely redesigning the systems around it

it is genuinely hard to conceive of just how much work there is going to be

in adapting our systems to take advantage of all of this new power

Put differently, the capability overhang is market opportunity and is where a lot of our time in the near future is going to be spent. For now, another exciting moment to start the week, and that's gonna do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace 

​ 

Nathaniel Whittemore's audio recording:
