
1034: 2026年9月の見逃し配信(ICYMI)――判断力、ワード・グラビティ、AIエージェントのマネジメントとトークン消費
1034: In Case You Missed It in September 2026
1034: In Case You Missed It in September 2026
Super Data Science: ML & AI Podcast with Jon Krohn
要約
Jon Krohnが9月の5つの対話から注目部分を再編集した回。Aishwarya Srinivasanは、コード生成が安価になり、ボトルネックが実行から判断へ移ったと述べる。Luis Serranoはアテンションを言葉が空間を曲げる「ワード・グラビティ」として説明する。Katie Maloneは人のマネジメント経験がエージェント運用に生きると語り、Ish ShahはエージェントのトークンのコストとMETRのチャートに触れ、Dilani Kahawalaは常時稼働エージェント「Anna」の難しさを語る。
- ●Srinivasanは、コード生成が安価になりボトルネックが実行から判断に移ったと述べ、意図や設計が不明確だとモート(競争優位)を築けないと指摘した。
- ●エージェントの評価はモデル出力だけでなく、ツール呼び出しやAPI障害などを含むエンドツーエンドで行うべきで、画一的な方法はないと語った。
- ●Serranoは、アテンションを単語が埋め込み空間で互いに引き合う動きと捉え、物理学者との共同論文で時空の曲率という緩い類推に発展させた。
- ●Maloneは、タスク定義、完了基準、trust but verifyといったマネジメントの考え方がエージェント活用に転用できると述べた。
- ●Shahは、サブエージェントの分業で作業量が膨らみ、トークン単価が下がっても使用量の増加が追い越すというジェボンズのパラドックス的な見方を示した。
- ●Kahawalaは、UIがないこと、ユーザーがエージェントに慣れていないこと、信頼性の3点が難所だと述べ、並行処理される会話の割り込み制御も課題に挙げた。
章立て
判断力がボトルネックになる時代
Aishwarya Srinivasanが、コード生成コストの低下、AI評価の重要性、従来のセキュリティの必要性を語る。
ワード・グラビティとアテンション
Luis Serranoが、アテンションを言葉が互いを引き寄せる動きとして捉え、時空の曲率という類推に発展させた経緯を説明する。
ポッドキャスト再開とエージェント管理
Katie Maloneが、マネジメント経験とエージェント活用の共通点や、大企業での導入の難しさを語る。
トークン消費の急増とMETR
Ish Shahが、サブエージェントによる消費増と、METRの能力チャートの限界を説明する。
常時稼働エージェントAnnaの課題
Dilani Kahawalaが、UIなしの製品設計、ユーザー教育、信頼性、会話の割り込み処理を語る。
解説記事
2026年9月に配信された5つのインタビューのハイライトを、Jon Krohnが一本にまとめた回である。AIシステムの内側の仕組みから、日々AIと働く際の実務まで、「AIと共に暮らす」ことを軸に構成されている。
実行から判断へ:Aishwarya Srinivasanの指摘
Srinivasanは、コードを書くコストが急落したため、ボトルネックは実行ではなく判断に移ったと述べる。ただし、ソフトウェアエンジニアが不要になるのではなく、役割が変わるという見方である。数分で作れるものは他者も数分で作れるため、何を、なぜ、誰のために作るのかという意図がなければモートはないという。
エージェントの失敗については、AI評価が重要だと強調する。モデルの非決定性に、ツール利用やループ、ハーネス(エージェントを動かす周辺の仕組み)が重なり、不確実性が積み上がる。そのため、入力からツール呼び出し、APIの失敗、最終応答までをエンドツーエンドで評価する必要があり、万能な型はないという。DDoS攻撃、不正ログイン、プロンプトインジェクションといった従来型のセキュリティ観点も欠かせないと語った。
Transformerを「空間が曲がる」と見る:Luis Serrano
Serranoは、クエリとキーの検索表やソフトマックスの式といった標準的な説明では、アテンションがしっくり来なかったと話す。ある動画で、単語が他の単語との線形結合(係数の合計は1)に置き換わると知り、単語が埋め込み空間内で互いに向かって動くイメージを得た。これを「ワード・グラビティ」と呼ぶようになったという。例として、financial な意味の bank が river に引かれて自然のほうへ寄る場面を挙げた。
トロントでの講演後、物理学者のRicardo Di Sipioから、力としての重力ではなく時空の曲率として捉える方が近いと指摘された。そこからジオデシック(測地線)の数学と実験を進め、論文にまとめた。Serranoは、これは物理の厳密な意味での相対性理論ではなく、緩い類推にすぎないと断っている。理解を助ける視覚的な見方という位置づけである。
マネジメント経験はエージェント運用に生きる:Katie Malone
Linear Digressionsを5年半休止して再開したMaloneは、データサイエンスがマネジメント寄りになったことが休止の一因だったと振り返る。いまは、その経験こそAIエージェント活用に通じると述べる。複数の作業ストリームを切り替え、タスクと完了基準を定義し、誤解や不完全な実行を確認する。これらは人のマネジメント概念がそのまま当てはまるという。
一方で、優れた個人貢献者だったエンジニアが、レビュー中心の仕事でフローに入れず苦労している状況にも触れた。大企業は航空母艦のように方向転換が難しく、変革管理が重要だとも述べる。Tom Davenportの「process slop」(AI生成物が往復するだけの業務プロセス)にも言及した。
トークン消費の膨張:Ish Shah
Shahは、結婚記念日の贈り物として作ったポケモン風のゲームを例に、週末だけで数十億トークンを消費した経緯を語る。エージェントが作業を分割してサブエージェントを生み出すため、スループットは上がるが、作業全体の量も膨らみやすいという説明である。
さらに、AIが人間の介在なしにこなせる作業量を示すMETRのチャートに触れた。数十時間分の作業をこなすモデルが登場し、人間のベンチマークが取れないためチャートから外れつつあると述べている。トークン単価が下がっても、使用量が連鎖的に増えるため、ジェボンズのパラドックスのように総コストは下がらないという見立てを示した。
UIのない常時稼働エージェント:Dilani Kahawala
家族向けの常時稼働AIアシスタントAnnaを開発するKahawalaは、プロダクト責任者としての約10年の経験の大半を捨てる必要があったと話す。難所は3つ。UIという支えがないこと、一般ユーザーがまだ「頼むだけ」というモデルに慣れていないこと、そして家庭の文脈を学習していないモデルに対して、消費者は誤りへの忍耐が低いため信頼性が求められることである。
Annaは待機せず、裏でメールや予定を処理し続ける。そのため、ユーザーがレストラン予約を頼んでいる最中に、会議変更が送迎と衝突すると割り込んで伝える必要がある。音声では、複数の依頼が並行して処理されるため、会話の順序を管理するキューが欠かせない。数か月前の音声モデルでは難しく、最近のGemini Liveの進化で可能になったとも語った。
まとめ
5つの話に共通するのは、AIが作業を速くする一方で、判断、評価、管理、運用設計といった人間側の仕事が重みを増すという点だ。日本の読者にとっては、エージェント導入で評価設計とセキュリティの基本を省かないこと、マネジメント経験を技術スキルとして再評価すること、トークンコストを使用量の側から見積もることが注目点になるだろう。
文字起こし(英語・自動生成)
This is episode number 1034, our In Case You Missed It in September episode. Welcome back to the Super Data Science Podcast. I'm your host, John Krohn. This is an In Case You Missed It episode that highlights the best parts of conversations we had on the show over the past month. My first clip is from episode number 1023, where I speak with the wildly popular data scientist and entrepreneur, Ishwarya Srinivasan. Ash has over a million followers who clamor to read her posts about shipping agentic AI in production. With the cost of writing code collapsing for all of us, the bottleneck in shipping has moved from execution to judgment. So I ask Ash, which failures a person ought to live through themselves before anyone should trust them to deploy an agent? Let's get into the nitty gritty of some of the stuff that you teach on agentic AI at the Gen Academy. I think that that's probably one of the most useful ways for our audience to spend their
time learning about the most cutting edge things. So you argue that AI shifts the bottleneck from execution to judgment. So if the kind of the first generation of AI education taught people how to make models execute agentic AI may require teaching them how to judge when systems are not working. So what does a curriculum centered on failure literacy look like? And which kinds of failures should our listeners be most looking out for? Should they maybe even encounter deliberately before they're trusted to deploy an agent? Very much. And this is something that I've also shared in the past in one of my previous sessions at Open Source Conference. And I was talking about this, that the cost of building code has become so cheap. And that is one of the reasons people misinterpret the fact that software engineering is going to become obsolete or like people are not going to need software engineers anymore.
I think the entire role is shifting. The amount of time that you used to spend on writing import statements for your code or fixing the intendation of your code is not the same. So the time that you invest in doing different parts of your job for a software engineer is not going to look the same. it's obvious it's dead obvious that nobody is writing code by hand anymore right everybody is doing aka white coding uh i don't know how i feel about that terminology it's good and bad at the same time because i feel it's great because it like lowers down the floor of what you can do and going from like an ideation to building something but at the same time i feel if it's not interpreted correctly, it gives away a feeling that building code and running code in production is as easy as thinking about an application, which it's not. So I feel having an understanding of where you cannot let go of not just engineering fundamentals,
but also knowing what exactly are you even trying to build, the optimization part of it, the architecture part of it, the decision-making part of it, if you don't know what an IAM is, you'll not know what an IAM is. So that's why, like, having that literacy of that software engineering fundamentals is very crucial if you really call yourself an AI builder. If you want to really not be using, like, Codex to, like, build these cute little applications which run within your Codex browser, and then as soon as you deploy it on Vercel app and you give it out to like the first thousand customers, it's going to start breaking, right? So I think that is a huge difference. So as it reduces the floor, as it reduces the barrier to entry for more people to produce code, it reduces the barrier to entry for the cost of producing code. People are also generating spaghetti code. It is a shit ton of spaghetti going all the way. And that's the AI flop,
even for code that's happening everywhere. and people don't know how to really make use of it. So what used to take you a few weeks to build something which was more mindful is now taking you 10 minutes to build but is mindless. So it's only accurate to say that while the execution has become so much cheaper, if you're not intentful about what you want to build, then you definitely don't have a moat. Because what you can build in 10 minutes, somebody else can build in 10 minutes too. So if you're not intentful about what you're building, why you're building, who is it for, how is it going to be used, how is it going to improve over time, how is it going to compete in the market, that's still classic business. That's still classic product. That's still classic engineering. So that's something which definitely hasn't changed. Yeah, and I think anyone who's interested in getting involved in a business or in a product, they're going to want to know that somebody has been mindful, that a lot of thought has gone into this,
that there's a moat, that it's going to be secure, that it's going to be compliant, that it's going to scale. Like you said, if you just all of a sudden have a thousand users and the system's not set up for that, then you're going to run into a lot of trouble. And yeah, I think it is so easy today to create not just software, but just about anything. I think I actually recently told a story online recently, but I think it aligns with what you say. A friend of mine sent a pitch deck for a new business that he's creating, and it was obvious that the whole deck was AI generated. And I just didn't want to read it. I wrote back, I was like, I feel like I'm wasting my time when I'm sent a document where I don't know if you spend more than two minutes on this. So I would prefer to have a terrible-looking Google slide that's just like 10 slides with a white background, but like some diagram, like a few diagrams and a few bullets that explain what your business idea is. And I know that you thought through it. That would be better to me than this like 20 page,
amazing formatted, all this detail. Like, I don't know. I mean, it's a trend, you know, like I feel the unfiltered raw things are more authentic now and authenticity is being credited for. and authenticity is what like people are resonating with because they've had enough of this log there are enough like beautiful looking presentations which is meaningless in in text so i i think people are recognizing that yeah yeah yeah anyway uh you gave a great answer but there was there was uh you know the question that i asked you now a few minutes ago but there was one part of it that i i'm not sure if we got to the kind of i asked a very long question but the end of it was which failures should people encounter deliberately before they're trusted to deploy an agent specifically. So not just like kind of general apps, but when you're thinking about training people up to create and deploy agents, what are the kinds of experiences that they need to have that they need to see go wrong in order to like be trusted in real production?
So I would say that's one of the areas where AI evaluations are such an important topic. And that's also like a huge, huge area. Because of the non-determinism nature of the models, that compounds with the tool use that you give it access to. That compounds with more loops that you're creating with it, more drafts that you're creating with it, the agent harness that you're creating along with it. That has a huge compounding effect of uncertainty in that entire system. That's the reason understanding how these models work in different scenarios is what is AI evaluations, which is you're not just evaluating the models output. It's not that you're giving a model input through a LLM API and you get a response and you're seeing how the text response looks like. You're actually deploying and evaluating it end to end, which is right from when a user puts in an input all the way to all the tools being called, all the failure modes being addressed.
if there is an API fail, if there is a tool login issue, all of that being considered till the end of it, when you're actually getting the final response, all of those breaking points need to be evaluated. And that is a crucial aspect. And it is not something which is run size fits all. And that's what's very important for people to understand that a lot of times people are looking for prescriptive approaches. while these prescriptive approaches or the frameworks can help you give a certain direction, which is just a generic direction, or even these metrics can give you a general sense of how your agent is performing. Every single attachment that you do to your model, every single of those joints can introduce incredible amounts of challenges, incredible amounts of like failure mode. And that is only something that you will be able to assess if you are mindful of making those joins and making those connections, knowing how many tool access does it need to have access to?
How are you going to manage the role-based access to different users? What kind of database does it need access to? When should it be able to go right back into a database or not? So having a deep understanding on what are you really giving the control for a specific agent is very important. And that's like one part of AI evaluations, right? surrounding that is the traditional software engineering thing that I'm coming back to, which is not something which is just new to AI agents. It is something that has existed even with traditional softwares. At the end of the day, if you have evaluated your AI agents inside that box, where you have like all possible different combinations of things that can go wrong, at the end of the day, that becomes a software. That is the new software. Any of the new softwares which are being built in 2026 are not non-AI powered. Everything has that non-determinism introduced to it. Now, that's the new software. That's the definition of a software now. So when you think about that as a software and you're deploying it, you come up with the traditional AI engineering constraints of what happens when you have a DDoS attack?
What happens if an unauthorized user is trying to log in? What happens during a prompt injection? What happens if a certain user who's unauthorized is able to get access to a database? And so on and so forth. So that comes back to like traditional security and safety, which is not entirely new to AI agents But knowing that is also equally important as much as knowing AI or agent stuff Ash keeps coming back to fundamentals knowing what is actually happening underneath the thing you are building My next guest took that instinct about as far as it will go. In episode number 1025, Dr. Luis Serrano, founder of Serrano Academy and author of the bestselling book, Grokking Machine Learning, tells me that none of the standard explanations of attention ever clicked for him. Not the query and key search table, not the formula. So he went off and built his own picture of what's happening inside a transformer. Then, a physicist, sitting in the front row of one of his talks, came up afterwards and told him
it was something far bigger. Something else that you published is actually a new paper. So you published last November with, I'm going to try not to butcher their names, Ricardo Di Sipio and Jairo Diaz-Rodriguez. Yes. Very good. I really put a lot of effort and thought into that. You guys wrote a paper together about the curved space-time of transformer architectures. Yes. And that is pretty mind-blowing. I think we're going to spend a bunch of time on that right now, because we'll learn about transformer architectures in a way, but I think we're also going to learn about space-time and relativity and these kinds of concepts. Yes. This was definitely very exciting. Definitely very exciting to work on. And yeah, definitely for the physicists listening, we use the word relativity, but in a very loose way. It's basically a space-time curvature analogy, a weak analogy of what's happening inside a transformer.
But I think it opens the door to what's happening underneath. And the fact that physics-related things start appearing, I found it mind-blowing. The story of that is that in order to understand attention, it never clicked to me. People say it's like a search table with a query and a key never made sense to me. Then they said, oh, the words pay attention to other words never made sense to me. Then they gave me the formula made even less sense. It's a soft max of KQ divided by square root of DK times V. there's nothing for me so i started looking at videos and looking at other things and looking at uh just writing and and um like like watching videos and somewhere somewhere in the process somebody which i bless his soul i don't remember which channel was this i think it was a kind of an underrated channel that didn't have very subscribe but this this person is just just It kind of like made a, like, this is a beautiful description where they were at some point,
they did a linear combination of the words, and they said this word becomes more like that one. And the linear combination added to one, the coefficients added to one because there's a softmax. So, you know, you turn your word apple into, you know, 70% of apple and 30% of orange, if you said the two words consecutively, say orange, apple, you know what I mean? so the word becomes a percentage of itself and and the rest of percentage of another word and to me that's moving in a line right like if you have if I have two points and then I take a percentage of one of the position of one point and the other I'm moving in the line between them and so I thought maybe words are moving in a line and I started rewriting all the equations as in like words are in a position in space because embeddings words are in a position in space and then attention just moves them in a line toward each other. And I immediately thought, oh my God, that's gravity or magnetism, right? Words pull each other. And I thought that makes a lot of sense because if I'm
saying the quintessential example is the river bank. Bank is a bank in the financial sector of the embedding around stocks and bonds. And then you say river bank and the bank just becomes a nature thing. So the word river just pulled it towards itself into the nature region of the embedding because in the nature region of the embedding lives a river and tree and stream and sea and all that stuff. And so it just pulled it. It infused itself with it, right? It infused some properties of it. For example, the nature property. It didn't infuse all the properties. There are some that don't, but the nature property, it moved in that. So the moving actually works in different directions. It's not towards it because of the value matrix. But anyway, the fact is, words, I started calling it word gravity. And as I said, any physics words that I say is a very loose analogy. But I started calling it word gravity, word gravity, word gravity. And then one day I gave a talk in the Toronto Machine Learning Summit. And Ricardo Discipio,
physicist. I also, by the way, I hope I'm pronouncing it well because I know Italian. And Ricardo is a physicist who was sitting in the first row, and then he came to me and said, hey, what you have is actually, you know, when Newton was talking about gravitation, and then Einstein came, which is a force between objects, he died and he had no idea why this force happened. and then Einstein came hundreds of years later and he said, it was not a force. The masses are bending the space. Like, you know, the reason you fall towards the earth or the earth falls towards the sun is not because the sun exerts a force, it's because the sun bends space and all of a sudden the line in which the earth should be flying in a straight line, it's curved because of the sun and it happens to be curved around the sun and that's why we're there, you know? So he said, I think it's the same concept. Like you're talking about where gravity is like a force between words. I think it's more like it's probably a space-time curvature thing. Like it probably works our bending space in a way that they just pull towards each other, right?
And so we started working on that. And he worked out the math a lot. He knows the physics a lot more than me. So he actually worked out the geodesics. And very much like the matrices that appear in the geodesics appear in are the key query and value matrices. and then we started working with another friend hired as a professor at York University in data science and statistics to run a lot of experiments so these two guys are wonderful they actually know a lot more about that than me and working on both the physics and a bunch of experiments that really study this analogy and so we're very excited actually of this I mean it provides an analogy I think it's more of a visual work we meant it as a visual work. We've gotten notices of like labs that are working with it for something else. So I'd love to see applications of it. But as of right now, we thought of it as like a fun analogy, like physics analogy of what's happening inside Chad Gubitier,
inside the brain of these models. That's a view from inside the model. My next clip is about what it feels like to work alongside these systems all day. In episode number 1029, Dr. Katie Malone, host of the very popular Linear Digressions podcast, a show she stopped making for five years and is only now brought back with the help of AI automations. In this clip, Katie tells me which part of her career turned out to be the best possible preparation for working with AI agents. Her answer is reassuring for some of us and rather less so for the engineers who are the strongest individual contributors. An interesting part of your podcast journey is that you started in 2015, so 11 years ago, and you ran it for five and a half years, almost 300 episodes. And then the pandemic summer, July 2020, you stopped and you stopped for five and a half years. You did it for five and a half, stopped for five and a half. And the reason you gave it publicly at the time, it was that, you know, there's no particular reason.
You just couldn't do it forever. And, you know, we have some quotes from you at that time that, you know, if you felt that the field was moving away from you, the show had started. when people didn't even know what the realm of the possible was in data science. And by 2020, data science had become as much of a management responsibility and scales about the algorithms. And, you know, you said the content, you know, kind of stopped pulling at you. Then nearly six years later, you named two causes you hadn't said before, a pandemic burnout, a grind of production. And you kind of alluded to that now here today, where that grind has been alleviated so much by tools like, you just said it before we started recording. What's the name of the tool for? Descript. Descript, exactly, yeah. I can't believe I didn't have that right in my brain. But yeah, an amazing tool for allowing people to edit episodes very quickly. Yeah, I don't know. I find that journey interesting, and I wonder how many people do that. But we're so delighted to have you back on air.
Well, thank you. I'm delighted to be here. And I think this is interesting for me because I haven't gone in and excavated what I was saying or thinking in 2020. That resonates. That tracks. That sounds like something I would have said. And something that I've been thinking about a lot lately, and I'm really interested to hear the seeds of it and what I think. So something I've been thinking about a lot lately, I mentioned in 2020 that data science was becoming, for me anyway, partly because of just where I was professionally, a lot more about management than necessarily hands-on keyboard. And so struggling a little bit with coming up with new content that was faithful to what I thought my audience came to us for when my day-to-day job was, like, managing people. Like, I'm not writing algorithms anymore. I'm, like, going to meetings. and in the time since then I've stayed in data science management broadly. But I think that with the advent of AI,
there's a very interesting synthesis maybe between person management as like a soft skill that you might learn because you have to do it for your job and the technical management skills that you need to be an effective user of essentially agentic AI. So the idea that my job now is context switching between different work streams that are each being carried out independently. It's about defining the task to be done and the acceptance criteria for when it's going to be complete, that there's a fuzziness or there's a lot of different ways that what I say can be misinterpreted or done incompletely or not in the way that I intended. And so I have to be checking for that and, you know, kind of a trust but verify type model. Like those are all management concepts that transfer like very very elegantly to being an effective user of contemporary AI tools So this is just an idea that I been like developing a lot because I think a lot of people
are maybe non-technical, but they have been managing people or projects or whatever for a while. That set of skills might be one that they have very developed. They're potentially being confronted with the possibility of needing to manage this new type of entity, like an AI agent for the first time, and maybe feeling like a bit out of their depth. And I would say to them, like, you know, you might actually already know more than you realize. And I think to some of the very technical people, especially software engineers who've been effective ICs in the past and are now struggling as being agent managers and are saying, like, I hate my job now. I don't like reviewing other people's content. I can't get into flow, like, 100% true. And I don't have answers to all of those problems or all of those questions as a manager. Like I struggle with flow. I struggle with context switching. I, you know, struggle to articulate what I want sometimes. But there's a lot of other people that have figured out ways to deal with that. And maybe there's some cross-pollination in the other direction.
So anyway, maybe more than you were thinking when you asked the question, but I'm interested now that we see this, you know, me back in 2020 saying like, well, I don't know if AI is like really what I do anymore because I kind of do all this management. And I'm like, oh, those are the same thing, just in different clothing. That's a really interesting answer, and I'm so glad that you got into the authentic stuff right away, because I was starting to think, as Aphra had posed this question, I was like, have we been going on about podcasting too much? Like, is this just my interest? Is the audience going to be as interested in this as I am? And I don't know, I've been, lots of advice on hosting a podcast is that you should be getting into whatever interests you, but I was still like, maybe we should be getting into the technical aspect of this. And then you did anyway. So perfect. Yes, this new world that we're in where we're doing agentic management is, you know, it's only been the past year that this is something that people are doing. You have worked at a business with tens of thousands of people where you built the agentic AI platform.
And this includes deployment, enterprise adoption, responsible AI governance. And, you know, there's a lot there to get right. I don't know to what extent you can tell us about what it's like building, you know, being the person responsible for managing a team of humans and agents to build a authentic AI platform? Yeah, that's an interesting question. I mean, it's hard. And I think it's, you know, one of the things that's very challenging right now, and I think this resonates maybe with everyone to some extent, is as much as you can build something that's compelling and maybe like a little bit future-proof and sets us up for some long-term growth and value and whatever, like when the goalposts are just moving as quickly as they are, it's really difficult. And big companies, I think, have it extra hard because they're kind of like aircraft carriers. They're just hard to turn. Once they're going in a certain direction, they can go very, very far.
But they tend to not, you know, it's just not as nimble to get 10,000 people kind of going in a particular direction. I think I do kind of wonder, you know, some of this is just reflecting where we are as a society right now. This might look very different in five years as people have had a chance to acclimate a little bit to some of the AI tools. People are maybe a little more fluent with it. Some of the norms that I think we're figuring out now might have settled in a little bit. Like it might be, you know, you feel like it's not okay to send AI slop to your coworkers or something right now. I hope it doesn't. Please stop. If you're that one guy, just stop. Yeah, it's not one guy, though. Yeah, that's the thing. It's like, you know, my AI slop is talking to your AI. I talked to Tom Davenport a few weeks ago for my podcast. He's great. He's wonderful. And he's, you know, for anyone who doesn't know Tom, he's been writing, especially in enterprise data science and analytics for decades.
And he has coined the term process slop. And I think it's an idea whose time is rapidly approaching of, like, you know, there's a whole process. I'm a job applicant, and you are the hiring manager on the other side. And it's just, like, our AI slop going back and forth. I, you know, have AI generate my resume. You have your AI that reads it, that, you know, automatically sends me some kind of, you know, reply, like whatever. Anyway, so I think that those are challenges for us in general and, you know, in particular in large organizations where you might not have direct personal relationships with the folks that you work with. You're kind of relying on the machinery of the organization and some of the processes to, you know, kind of get things to where they need to go. then injecting AI into that all of a sudden is like not necessarily, you know, fitting in exactly with how these things are working. And so there's also, I think, a really important part of the, what you might call change management or something, like just how do you get people, how do you turn that aircraft carrier? And I don't know.
I don't know how much I have to say here that's like deeply insightful or, you know, specific and insightful besides like it's just really hard work. And I think it's interesting to see in some ways as there's new companies that are popping up, they're obviously approaching how to build businesses in sometimes very fundamentally different ways. Established companies are retrofitting their operations and their technologies to varying degrees of success or maturity at this point. So it's an interesting, I guess, period of high flux for us all to be in. Managing agents is one problem. Paying for them is quite another. In episode number 1031, Dell Technologies Distinguished Engineers, Ish Shaw and Tyler Cox, explain why agentic AI burns through tokens at a rate that is catching companies out. Ish gets there by way of the anniversary present he built for his wife, a Pokemon-style video game starring their dogs, which consumed billions of tokens over a single weekend.
Why? does token consumption go up so much with Agentic AI? Like, Ish, how did you burn through billions of tokens on the weekend on a side project? And can you tell us what it is? I can. You may have to bleep out a word if I commit some sort of IP issue. So, okay, so it's my one-year anniversary this Sunday. And as my wedding gift to my wife last year, what I did was I took everybody played Pokemon as a kid. Pokemon's making a comeback. It's cool again. everything old is new again. I basically built a fan game in the art of Pokemon, where the map is my, like, area that we live in Atlanta, where my wife and I met, where we got engaged, where, and I have these little pixel art maps, and I had her caricature done as pixel art, and I replaced the Pokemon with my dogs, right? That's the project. every year, every like major life event that we have,
I build a chapter into the game. And that's like my get out of jail free card on the present part of things. And so what I've been working on is these models and their capabilities over the last couple of months have like shot through the roof, right? The artwork has gotten considerably better. The game mechanics and how much I need to supervise my little buddies as they go off and work, I can go have a cup of coffee, right? And when I come back, the chapter is built. The reason the burn was so high is because what these agents are doing in order to achieve the task, just like humans, they're divvying up the work and they're spawning sub-agents. So now you've got like an agent in charge of a bunch of other agents, right? And yes, the pie of work is finite, right? Like you have your finite pie of work. but because you've got all these sub-agents in action, like are the sub-agents doing things to like the nth level of token efficiency
that a single agent would have done or a single, it's the same thing anthropologically is when you think about like humans in a workplace, right? If one person says, everybody get out of my way, I'm going to own this path single-handedly. I'm going to do it as efficiently as possible, but I'm one person. This is like queuing theory. How much throughput do you have? Multiple sub-agents means that you go faster, means the work gets divvied up, but the pie of work might get a little bit bigger because the subagents are at liberty to do certain things, right? The point of this is best probably articulated by something that has almost nothing to do with what we've talked about, although I'm sure it'll come out. It's this organization called METR, M-E-T-R, Model of Value of Software Research. John, you're nodding, so I'm not sure if they've been on the pod or... And I talk about meter probably more than any other single thing on the podcast. And then almost every talk that I've given for a year or two now. Right. Near the beginning, I show meter charts.
Ah, so our presentation and your presentation are basically starting the same way. And then yours continue to be smart and mine kind of plateau. Meter, Model Evaluation Threat Research, and SDS listeners are going to be familiar with this at this point, has a chart, which when you land on their website, maybe we can put it in the show notes here, like it shows on one dimension time, like 2021 until now. And then on the other dimension, it shows the ability of a model to operate unsupervised to achieve a certain goal at a certain fidelity of accuracy compared to a human given the same task. Now, the reason they have this big scary name, which says threat research inside of it, their whole point was like, hey, at what point is AI going to cause harm to human beings? And we should probably be tracking that. And the heuristic they came up to track that with is this chart. How much can it do by itself?
And that chart is just like, not only is it up and to the right, it's just like, I mean, it's gone vertical, right? And at a certain point, they just kind of said, I don't know, it just keeps going up. Since the release of Mythos, they can't really track, It has gone off of the meter charts because in order to be able to benchmark the performance of models effectively on one of these charts, you have to have had humans doing these tasks and know how long it takes humans to do these tasks to be supervised. And that was easy like 2021 when you looking at GPT level capability and the tasks are only seconds long or then minutes long with GPT on average it very easy to come up with tasks that you can give humans to do, and it's not that expensive to pay them to do it and figure out how long it actually takes them on average to do it. But now that Mythos is doing, or Fable or Astra, GPT-6 from OpenAI, that class of models is now doing dozens of hours of work.
Work that would take a human dozens of hours. It could take the AI model 30 minutes or whatever to do something that takes a human 16 hours or 24 hours or 36 hours. We don't know how long those tasks, we don't know how capable these models are because we don't have any human benchmarks because how do you even, it's hard to even think of, like write a book chapter, write a book. It is quite literally off the charts, quite literally off the charts. And they accidentally invented a chart for one purpose is now like the best visual we have for like capabilities of models over time, right? But as these capabilities go up, like it's Jevin's paradox here. Like even if token costs get cheaper over time, like the base is going to move on you because people are going to realize they could do things like, you know, it took Nintendo how many years to develop a Pokemon game? Like they'd come out every two or three years when we were kids. Now it's like in a weekend, someone can sit down and build a video game to the same level of fidelity.
Like the token consumption is growing and it kind of doesn't matter how cheap you make the individual token. If the order of magnitude of usage is just constantly chain reacting on itself to get bigger and bigger and bigger. And we're rounding up a great month with episode number 1027 in which Dr. Delani Kahawala, a Harvard physics PhD who spent a decade in product leadership at Etsy, Facebook and Atlassian, tells me she has had to throw almost all of that experience away. Delani is now co-founder and CEO of Anna, an always on AI assistant for families. And she walks me through the three problems that turn out to be hardest when your product has no user interface and never stops running. So with your extensive product management background, and so to go through this to you, after doing a Harvard PhD in physics, you then went to McKinsey as an associate, Etsy as a product manager, then senior product manager at Etsy, lead product manager at Facebook, and then group product manager, head of product, head of product management at Atlassian.
And so a decade of experience in senior product leadership positions. And now you're full-time creating Anna. You're the CEO of the business, but you're surely also the head of product. That's kind of all we do. So it's been, the funny thing is we've had, I've had to pretty much throw away a decade of how we think products should be built for people. Oh, really? And that's been really fascinating. So we've had to think very first principles from when you don't have the crutch of a user interface. That's one problem to solve. The other problem is the average person doesn't really yet regularly interact with AI agents. So I'm on code every day, all day. And my mode is if I – I just ask and it will have an answer for everything.
And the mode is use ask and, like, it gives. But I don't think the average person is yet familiar with, like – there's a deterministic set of options that you can tap and drop down and things like that. You just asking an agent to do things for you is still a different, like, mental model. And so the second challenge is, like, how do you get someone to that operating model where you just ask? And then I think the third thing is how do you, you know, when you work with code code, it's never, like, it's not inaccurate, but you need to correct it, right? Like, it will give you an answer, but you will have to, like, cross-check it and be like, what did you think about this? What did you think about that? And these models are, like, optimized for coding. They don't have – they're not trained on, like, household data.
So it doesn't inherently know what to do with, like, where does this – which calendar does this go in? Like, it doesn't have that understanding. But consumers don't have that much patience for, like, your assistant getting something wrong. You don't want to be constantly correcting it. So those are the three things that I think we've had to really think about how to, like reliability, the interface, and just teaching users how to work with an agent have been the three biggest challenges. I guess something that's quite different about an agentic interface like Cloud Code and what you're building is that in Cloud Code, it is still a turn-based conversation where, yes, it goes off. It's agentic because it spins up. It figures out how to tackle a task, spins up subagents as it needs to. But ultimately, when it's done what it's doing, it just stops.
It gives you an output and then waits forever. And if you never come back to that chat, nothing ever happens again in that chat. It seems to me like with Anna, you would need to, there will be times where Anna needs to reach out, where maybe Anna has sent the last message and needs to send another one before you've responded because something has changed with your kids' football practice or an important email has come through or a reminder of an upcoming appointment or something like that. And so it seems like it's more discursive, more back and forth. It's not as linear or just turn-based back and forth conversation. Yeah, and this is, I think, the biggest change. So I think when, like, Anna is like a long-running agent, meaning that it doesn't kind of stop and wait. It is constantly working every second, every minute, working for you behind the scenes.
so when you know like the open calls of the world the herman's agents um and now see maybe like grokbot um all these agents are trying to tackle the same problem of like how do you just continuously work with someone but with like not that much success because most people like set up an open call and then they kind of give up on it after like two weeks um because what we have to have happening behind the scenes and it's constantly working for you on a set of things whether it's like checking your email or figuring out if a piece of information is noise or if I've handled this before is it already on your calendar have you already tackled this task which kid is this relevant for just constantly working in the background and you're at the same time having in conversation with it where like claude we will have a team of you know domain specific experts agents who are going and doing a bunch of things for you but anna will be like oh i picked
up that your meeting change and it's going to clash with your school pickup i need to interject and give you that message while you might be asking anna to you know book a restaurant reservation So we had to figure out how to handle that, so I queue things up in the correct way. There's a whole layer of ops that Core doesn't have to deal with yet. Yeah, that's a really interesting use case there that I hadn't even talked about in the way that I was like, oh, this must be more complex, not just having back and forth. But it is also interesting that you could be having a conversation. You're in your car talking to your car, your car phone, but your car phone is Anna, like on the other side, and you're saying you're having a conversation about scheduling some upcoming event, and then it has to actually interject and say, we're going to have to take a pause in the conversation that we're having because this important thing has come up. That is a really interesting, and yeah,
I have never experienced anything like that in any conversation with a non-human to date. So this is actually really obvious in voice mode. So when you put voice mode on, you could be, if you ask Anna to do something complex, like go find me a dentist. She has to go do some research and just like look up where you are and like who's best reviewed. That task takes sometimes like a minute or two because that's a complex task. In the meantime, you might, and voice goes pretty fast, you might have fired four or five things at her. and so she's like so we like fanned out a bunch of agents who are doing multiple things for you but it's it has to be then like queued up in the way that the conversation piece of it is understanding okay you asked me this first then you asked me this thing this thing is finished okay now I'm gonna like finish what I'm saying to you and then get back to you that was a fascinating
challenge and I don't think what's interesting is like when we started and it was only a few months ago the voice models then were not good enough to do that to even handle that like upfront conversation and it's only like three months ago that like gemini live changed substantially it could handle like the conversation piece while we have like the agentic brain behind the scenes doing all the fanning out and all right that's it for today's in case you missed it episode to be sure not to miss any of our exciting upcoming episodes. Subscribe to this podcast if you aren't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
番組の概要欄(原文)
In ICYMI Episode #1034, Jon Krohn moves from what sits underneath AI systems to what it takes to live with them every day. Hear from Aishwarya Srinivasan, Luis Serrano, Katie Malone, Ish Shah and Dilani Kahawala, discussing why the bottleneck in shipping software has moved from execution to judgment, how a physicist recast attention in transformers as words bending space, why years of managing people may be the best preparation for managing AI agents, how one weekend side project burned through billions of tokens and which three problems are hardest to solve when a product has no user interface and never stops running. Additional materials: www.superdatascience.com/1034 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (00:57) Agentic AI Skills That Matter Now (10:51) Word Gravity: How Transformers Bend Space (17:45) How AI Brought a Podcast Back From the Dead (27:39) Tokenomics: Why Your Agentic AI Bill Is Exploding (and How to Fix It) (34:28) Building an Always-On AI Agent for Busy Parents
