
MetaとMicrosoftのClaude削減、GrokBotの銀行情報流出、AIをいじめるとBANされる件
Meta & Microsoft's Claude Slowdown, His Agent Leaked His Banking Info, Don’t Bully Your AI
Meta & Microsoft's Claude Slowdown, His Agent Leaked His Banking Info, Don’t Bully Your AI
Big Technology Podcast
要約
Alex Kantrowitz と Ranjan Roy が、MetaとMicrosoftが社内のClaude利用を減らしているという報道を起点に、Anthropicの上場前の懸念を議論した。標準モデルとフロンティアモデルのどちらが重要かという持論の対立が、MuseやGrokBotの失敗事例を通じて展開される。OpenAIの売上数字をめぐる市場の動揺、AnthropicによるClaudeへの虐待行為禁止、GoogleのGemini発表の分かりにくさも取り上げた。
- ●The Informationによれば、Metaの Claude Code 利用者は約6万人から約3万人に減り、Microsoftは社内Claude支出の見込みを3分の1以上引き下げた。自社ツールへの移行とコスト管理が背景とされる。
- ●Kantrowitzは標準モデルが「十分良い」ことが、トークン使い放題への反発とともにフロンティア企業の脅威になると述べ、Roy は自社モデルへの切り替えは一度始まれば元に戻らないと見る。
- ●Muse が医師の電話番号を取り違えた件や、GrokBotが個人の銀行残高を社内Slackに投稿した件を巡り、モデルの責任かユーザー設定の問題かで二人の意見が割れた。
- ●ElonがGrokBotで Claude Opus など他社モデルも使うと発表し、二人はこれを「ハーネス(モデルを振り分ける仕組み)重視」の立場だと見なした。
- ●FTが報じたOpenAIの売上見通し(700億ドルか500億ドルか)は集計方法の違いによるもので、AI関連株が下落した。Royは収益指標の開示の曖昧さを批判し、S-1の公開を求めた。
- ●AnthropicはClaudeへの継続的な虐待行為を利用規約で禁止した。Bill Gurleyは学習データに使われている示唆だと指摘し、Claudeの意識を巡る社内の見方との関係も議論された。
章立て
MetaとMicrosoftのClaude削減
The Informationの報道を紹介。Claude Codeの利用者減と、Anthropicの売上集中リスク、S-1への期待を語る。
ダウンロード減速と標準モデル論
Claudeのダウンロード数が3月のピークから半減した点や、トークン使い放題への反発、標準モデルで十分かという論点を議論。
Museの医師予約ミス
Museが医師の電話番号を取り違え、妻が遠方の別の医師を予約してしまった体験談。データの整備状況が精度を左右するとRoyが指摘する。
GrokBotが銀行情報を流出
Shane Macのエージェントが「executive team」という同名の二つのチャットを取り違え、残高をSlackに投稿した事例。責任の所在を議論。
ElonのHarness Hive入りと個人エージェント
GrokBotが他社モデルも使うと発表。各社の個人エージェントが同じものを作っているのか、特化型が有利かを議論。
OpenAI売上報道と市場の動揺
FTの700億対500億ドル報道でAI関連株が下落。集計方法の違いと、ARR報道の問題点を話す。
Claudeを虐待するとBAN
Anthropicの新ポリシーを巡り、学習データへの利用やClaudeの意識をめぐる社内の見方、人間らしさの捉え方を議論。
Geminiの名称の混乱
GoogleがGemini at Workで発表した新製品の命名の分かりにくさを笑いながら語る。
解説記事
米Big Technology Podcastの金曜回で、司会のAlex KantrowitzとMarginsのRanjan Royが今週のAIニュースを振り返った。軸になったのは、MetaとMicrosoftがClaudeの社内利用を減らしているという報道だ。そこから、標準モデルとフロンティアモデルのどちらが重要かという二人の持論の対立が、エージェントの失敗事例や市場の反応と結びついて語られた。
MetaとMicrosoftがClaudeを絞る
The Informationの報道によれば、Microsoftは今年、自社内でのAnthropic利用に少なくとも10億ドルを使う見込みだったが、その額を3分の1以上引き下げた。コスト削減と自社AIツールの利用促進が理由とされる。MetaでもClaude Codeの利用者が約6万人から約3万人に減った。Metaは自社のMuse Sparkモデルを使ったコーディングツール「MuseCode」を推進しており、利用者は6,000人超という。
Kantrowitzは、AnthropicのS-1(上場申請書類)で2顧客が売上の25%を占めるとされる点を挙げ、今回の二社がその一部である可能性を指摘した。Royは、この25%という数字は2025年末時点のもので、その後かなり変わっているだろうと補足した。上場では過去の実績より成長見通しが問われるため、大口顧客の縮小は影響しうるというのが両者の見方だ。
Royは、社内で一度自社ツールへの移行が始まれば、元の規模には戻らないと述べた。コスト管理やデータ管理の面で自社製が有利だからだ。Googleでも社内でAntigravityなど自社ツールへの回帰があると聞いている、とも語った。Kantrowitzは、標準モデルが「十分良い」こと、企業がトークン消費の抑制を求める動き、そしてClaudeの新規ダウンロードが3月のピーク(742万)から9月に326万へ減ったことを、逆風として並べた。ただし、フロンティアモデルで大きな成果が出るなら削りすぎるのは無責任で、コスト管理と能力最大化のバランスの問題だとも述べている。
標準モデルか、フロンティアか
この対立は、後半の事例でも続く。Kantrowitzは、Museが医師リストの電話番号を別の医師の欄に載せ、妻が遠いクイーンズの医師を予約してしまった体験を披露した。彼はフロンティアモデルなら防げたはずだと考える。Royは、飲食店予約や航空券と違い、医師や地域事業者の電話番号はGoogleの店舗情報などに依存した整理されていないデータで、AIに任せるなら確認が必要だと反論した。
GrokBotの件も同じ構図になった。AI起業家のShane Macは、個人財務を管理するCFO役のエージェントを含む「executive team」というGrokBot内のグループチャットを持っていた。同名の実際のSlackチャンネルもあり、エージェントが両者を取り違えて銀行残高を投稿してしまった。Kantrowitzは、複数の人間がいる文脈で個人財務を投稿しないことは基本的な判断で、モデルの失敗だと主張した。Royは、エージェントは「executive teamを更新せよ」という指示に従っただけで、名前を分けなかったユーザー側の設計の問題が大きいと反論した。両者とも、こうした事故が頻繁には起きていない点は驚きだと認めている。
ElonのHarness Hive入り
The Informationによると、Elon Muskは、SpaceXのAI部門がGrokBotの裏側でClaude Opus 5.5、MidJourney、Sunoなど競合のモデルやAPIも使うと発表した。Royは、タスクごとに最適なモデルへ振り分けるのはまさにハーネス(モデルを包む仕組み)の発想だと述べ、「Harness Hive」への参加を歓迎した。Kantrowitzは、課題によってはフロンティア知能が必要だという証拠にもなると見る。Royは、MidJourneyやSunoは最先端の知能というより特定用途で最良のモデルだと指摘した。
また、各社の個人AIアシスタントが似た機能を作っているという論点では、OpenRouter共同創業者の例えが紹介された。2005年にクラウド上でDBやログイン画面を作れば誰でも同じに見えたが、実際には差別化が生じた、という趣旨だ。Royはこれに賛同し、子どもの学校の予定を拾うなど家庭向けに特化したエージェントが出てくると予想した。Kantrowitzは、信頼できる一つのエージェントがすべてを担うと考えている。
市場の動揺とAnthropicの新ポリシー
Financial Timesは、OpenAIの今年の年換算売上が700億ドルではなく500億ドルになる見通しだと報じ、NVIDIAやCoreWeaveなどが3〜8%下落した。Kantrowitzは、投資家が旧資料でAnthropic流(クラウド事業者経由の販売を含む)に集計した数字が出回ったためだと説明した。Royは、ARRの試算を含む流出情報を記者は扱うべきでないと述べ、S-1を早く公開してほしいと繰り返した。番組は先週、MetaがAnthropicに計算資源を提供すると述べたのは誤りで、その取引はまだ成立していないと訂正している。
Anthropicは利用規約を改め、Claudeへの持続的で不必要な虐待や残酷な行為を禁じた。Kantrowitzは、AIへの接し方が現実の行動に影響するとして歓迎する。Royは、Bill Gurleyの「顧客のプロンプトが学習データに使われていることを意味するのでは」という指摘に注目した。非企業ユーザーは設定でオフにしない限り学習に使われるが、その点を強調する書き方は同社に不利だとKantrowitzは述べた。Kantrowitzは、Claudeに意識があるかもしれないと考える社内の一派の影響もあるのではないかと見ている。
まとめ
今回の議論から読み取れるのは、AI導入の関心がモデル性能の競争から、コスト、運用設計、データ品質へ移りつつあることだ。企業が自社製や安価なモデルへ移れば、フロンティア企業は収益の見通しを問われる。エージェントに権限を渡す際は、名前の衝突のような設定上の落とし穴にも注意が要る。日本の読者にとっても、モデル選定だけでなく、振り分けの仕組みや確認の手順をどう設計するかが実務上の焦点になりそうだ。
文字起こし(英語・自動生成)
Microsoft and Meta back away from Claude. Does it spell some trouble for Anthropic ahead of its IPO? Will your AI agent share your banking data in a group chat? And finally, don't be mean to your AI. You might end up getting banned from the service. That's coming up on a Big Technology Podcast Friday edition right after this. This episode is brought to you by Genesis. What does it actually take to put agentic AI to work across an entire enterprise? I recently spoke with Genesis Chairman and CEO Tony Bates about why orchestration is emerging as a competitive advantage in customer experience and what companies need to get right in the AI era. Then I sat down with Adam Mitchell, head of enterprise business solutions at Voya Financial, to go inside a large-scale Agentex TX transformation, including the business decisions, trade-offs, and lessons along the way. You can watch both conversations now on my YouTube channel today. Welcome to Big Technology Podcast Friday edition where we break down the news in our traditional cool-headed and nuanced format.
We have a great show for you today. There's new reporting that shows Meta and Microsoft might be backing away from spending on Claude just as Anthropic gets ready for its IPO. Your personal agents might be very useful, but it might share your banking data in a group chat. We'll talk about what happened and why you should or shouldn't be concerned. And then finally, Anthropic has a new policy that if you're persistently mean to Claude, it might ban you from the service. So we'll talk about whether you should be mean to your AI or not and what the consequences might be. Joining us as always on Friday is Ron John Roy of Margins. Ron John, welcome. Good to see you. Good to see you. Don't be mean to your AI. Don't bully Claude. Are you mean to your AI? Do you ever bully? Because sometimes people say if you're a little bit mean, you get better results. I think I bully Copilot just because I think Copilot likes to be bullied. Co-Pilot likes to be bullied, but maybe Claude's a little more sensitive. We're starting right there. We're starting best. All right, we'll get to that in a bit, but we should really talk about the main story this week.
And look, it's not a crazy headline, but I think it says a lot about the risks of this AI buildup and what is going to happen now that standard models are performing pretty well. So a headline from the information this week said that Microsoft and Meta are both re-evaluating their spend with Claude and pulling back in their own interesting ways. So I'll just read a little bit from the story. Meta Platforms and Microsoft, two of Anthropics' biggest corporate customers, are working to cut their employees' use of Claude, a challenge to the AI firm's efforts to sustain its explosive growth. Microsoft executives earlier this year projected that the company was on track to spend at least $1 billion. on its own internal use of Anthropics technology. Microsoft has since lowered that figure by more than a third after it asked its staff to use less Claude to save on costs and spend more time using Microsoft's homegrown AI tools. At Meta, the number of employees using Anthropics Coding Assistant and Claude Code
has dropped sharply to about 30,000 recently. That's down from about 60,000 users earlier this year. You know, I said it was a small headline, but actually I'm not going to overreact here because it's a small headline, but it's big news. And I don't think it is an overreaction to say that this is some troubling news for Anthropoc. Not the end of the world, but remember, this is a company that's heavily dependent on a few customers. You know, there's two customers that we know from the company's S1 that make up 25% of its revenue. It's safe to say that one or both of these could be included in that group. and if companies find that lesser tools are doing as good of a job as Claude Code, it is a threat to Anthropics' long-term ability to grow. So I wanted to start there. Rajan, what do you think about this news? Is it doom time for Anthropics or is this just a hiccup? Well, in Anthropics' favor, I want to remind us that it was at the end of 2025 that two customers made up 25% of the revenue. I'm guessing that significantly changed.
Of course, we don't know anything about that because they haven't actually released their first half financials yet. But I do think it is very important that we know the large hyperscalers have been major customers of Anthropic, of OpenAI potentially, certainly Cloud Code. Every software developer these companies was using. I think this is very significant because, again, you go from 60,000 down to 30,000 users, and part of that decline did stem from layoffs in the spring as well, that there would be some kind of attrition around there. But the main thing here is we know that Meta, their own internal models and infrastructure have significantly improved. Everyone who's getting muse-pilled can feel it, see it. So you have to imagine within these organizations, they are pushing people to say, use our own models. And the moment that happens, it changes the entire anthropic story. And I don't know. I think it's going to be – I just want that damn S1.
I want to know what has happened because you have to imagine, especially, again, And it says Microsoft was on track to spend at least $1 billion on its own internal use of Anthropics technology. So these companies have been massive customers, and I have to imagine they're moving away from Claude. And what does that mean? Again, the IPO is not going to be about past performance. It's going to be about projected growth, and this has to affect projected growth. Yeah, the internal data is according to information. The internal data show that in a recent 28-day period, Meta spent more than $105 million on CloudCode. That's even as it's cutting back. So that would suggest that Meta will spend, I think this is this year, over a billion this year. But you're totally right because the information does have a line here that says the bigger factor has been Meta's development of its own AI coding tools and its MuseSpark line of models, which the company is marketing to other companies and pushing its own employees to use. MuseCode, a CloudCode competitor that uses metamodels, had more than 6,000 employee users
recently. I guess it doesn't fully make up the gap if you go from 60,000 to 30,000. Okay, so now you're at 36,000 of people using a coding tool. I wonder where the others are filling in. Or maybe people tried CloudCode, realized maybe it wasn't particularly useful for some of their purposes, and pulled back. And I think if you look at the numbers, and we publish these numbers in big technology. There was a big surge in Cloud Code users in March, and that has pulled back actually in September. Let me see if I'm going to get it. Okay. All right, so Cloud's download gains from January to March were bigger than the entire gain across all apps. This is how big of a surge. So the gain of users across Cloud Code or their gain was bigger than all apps. However, the downloads have slowed. So Claude's monthly downloads in the U.S. went from nearly 670,000 in January to a peak of 7.42 million in March, and that has dropped to 3.26 million in September.
That's downloads, right? So basically you're thinking about that as new users. But it is significant that that big surge that Claude Code saw when everyone saw the capabilities and said, I need to try this, that has cooled. And now the downloads are about half of what they were in September versus March. What do you make of this? Wait, does downloads include, like, installing Cloud Code in the terminal, do you think? Or is that actually, like, the Cloud app? I'm pretty sure that's the app. So, okay, there's probably some that it's not accounting for. But go ahead. But it's still meaningful. Yeah, it's definitely, like, directionally meaningful. I still think that, again, Cloud Code is the business here. Cloud Code, again, they've been trying to push to the kind of front-facing app, and they've combined a lot of different things. And a lot of people complained about what is chat, what is co-work, what is Cloud Code. But I think the most important thing is that number you said, like Muse Code is at 6,000. Meta or Cloud Code is still at 30,000, down from 60,000.
Any of these organizations, anyone familiar with these, like, you know what happens is the new tool comes out. Everyone gets completely cloud code pilled and starts using it, and then they don't want to give it up, and they say, I won't be able to do my work without this tool. And then slowly the company, in this case, Meta, will say, start using MuseCode because we have far better cost controls around it. We have control over around it. We know they're not stealing all of the data and stuff we're inputting into it. And the moment it works relatively closely as well, everyone is going to have to change. That is what I actually think is the single biggest risk here is, and I've heard this at Google too, that with anti-gravity, like internally, there's been a massive shift back to use Google tools. And Meta is definitely going to make that as well. I wonder at the Amazons of the world and everything how things are going to look. But when you are so dependent, and whether it's 25% is down to two customers, probably not anymore, and they're a little more diversified.
But still, just the moment things start moving in that direction, it means they're never going to move back to 60,000 people at Meta are going to have unlimited access to Cloud Code. Yeah, and I think this is a big point, right? And this is kind of why we led today with this story. You think about the frontier business, obviously the models are super impressive. But there are a number of challenges that they're facing. There is the fact that these standard models have become good enough. That may be a Muse code. I mean, let's be honest. Muse code is not going to get you where Cloud Code will get you. But it could be good enough for some users that people will move off it. And then, oh, I just want to finish the list. Yeah. Right. Let me finish the list. There's a token maxing rebellion. Right. Every CEO is asking how can we get our employees to spend less because this was a line item that went from zero. to, in some cases, like Microsoft and Meta, more than a billion in a year from nothing. So that is happening as well. Now, it doesn't mean that the frontier business is over.
I still think that there's a balance here because if you find that your employees can have bigger gains using frontier models than the standard models, you're being irresponsible if you're cutting just a cut. There is ROI to be found there. So it's a balancing act between cost control and capability maximization. but all these factors are going to start to add up for these companies. So I still, I have not used MuseCode myself. I use CloudCodePlenty. I still just have to imagine that, again, coding and software development, I'm not saying it's solved, but for how good Muse itself is, like the customer-facing app, I have to imagine Meta has made significant strides in the actual their model capabilities. So I still have to imagine that if it's not as good in every possible way as of today, I have to imagine for the work that actually needs to be done, it's probably as good. However, I did just look up on Reddit.
I was trying to see, like, if anyone has been talking about MuseCode. Not a lot is coming up. So clearly, like, in terms of any kind of public access, it's fairly limited. However, six days ago, OddDonkey2691. Who we always listen to on this show. OddDonkey2691 is the harbinger of what is to come. At Mark Zuckerberg's alt. Well, I don't think so because the headline of the post is actually two days using Muse 1.3 contributor, it sucks balls. It definitely does suck balls. No, no, no. So that is when I looked up Muse, Code, Reddit, Review. That's what comes up first. So maybe we do have a little ways to go, at least on whatever is available publicly. I'm assuming internally there's probably something far more significant being used. But according to Odd Donkey 2691, we've still got a little ways to go. You know, I want to jump on the Odd Donkey bandwagon here because I think this is going to become a new debate within our show.
And it's an extension of the product versus the model, where I always said the model was most important. Ranjan said the product is most important. And we're getting now to this point where we have agents, and the question is, you know, how important is the frontier model? If you believe the frontier model is most important, then OpenAI and Anthropic are going to have a big lead here. But if you believe standard models with a good harness is good enough, like Ranjan does, then the meta and the Grox and the instincts of the world can catch up. So I just want to put one data point here. It's part of a running story that I've been telling on the show and try to tally one more for the Frontier Lab So Ranjan do you remember a couple weeks ago I told this story about how Muse was able to help assist my wife with an appointment at a specialist in New York City where we couldn find a specialist, and Muse gave us this list, and she was able to book an appointment with the first doctor on the list? Yeah, I remember. Okay, well, there's part two of the story,
which involved me getting yelled at last weekend. because, you know, though Muse did put this very helpful list of doctors together, it's very capable standard model, I'm saying this facetiously, put the wrong phone numbers for the wrong doctor, like put working phone numbers of some doctors underneath the wrong doctor's names. So my wife booked an appointment from Muse, you know, after making this phone call, the first person that it suggested, thinking that it was going to be the first doctor on the list because the phone number was underneath the list. It turns out she had made an appointment for a completely different doctor. She had called because Muse put it in the wrong place. And so the night before she's supposed to go, she's trying to figure out where she's going. And we thought it was going to be a brief commute into Manhattan, But it turns out that the doctor that she had booked was a multi-train ordeal in deep Queens.
And she's looking at me and she's like, you're freaking muse, man. Freaking muse. Well, this is – hold on. I'm going to give you my AI. This is probably not 101 but like 301. I think you've got to think about – no, no, no. So, I mean, again, this is something that, like, in my own work at Writer and with, like, enterprise customers, this is the hardest thing to explain. It's, if you think about right there, why are restaurant reservations so easy? It's because their centralized platforms have already aggregated and cleaned the data very well. I'm guessing, especially direct numbers to doctors, we are reliant on maybe Google local listings. Maybe, again, the ZocDocs of the world have done some aggregation. They never show up in any kind of AI search for me, so maybe they're blocked. But, like, if it is a data set that is not cleaned and kind of centralized, you got to double check, man.
Otherwise, you're going to end up on the train getting yelled at. Double check if it's not a – it's why, again, flights and restaurants, that's why everyone loves to talk about them because they got the cleanest data sets in the world. doctor phone numbers, local business phone numbers, not even a chance. So Ranjan, I did in this conversation suggest, why didn't you double check? Let's just say that was not the right answer. That was not the right answer. But I'm trying to say, what I'm trying to say is, you know, I do think a frontier intelligence will probably be more accurate dealing with such type of data problems. And so that's why, you know, I'd almost trust, not almost, I would trust opening eyes, dots, much more than I would trust MetaViews for a similar activity in the future. Yeah, I think we'll get there. Actually, as we're talking about this, there's a good opportunity in actually kind of like cleaning every Google listing
and kind of like doing something with that underlying data because that is like the single messiest data set in existence. It's like almost there, kind of there, but not there. And it ends up on switching three trains going into Deep Queens. Yeah. Well, she ended up canceling the appointment, and I was in the doghouse. But we're good now. We're good. We're good. So go ahead. Go ahead, Ryan. Hopefully Muse did not go share your bank statement into the work chat after that. Okay. No, it didn't. It didn't. So it's interesting. And this is another example of where I thought that AI agents were on the standard models were worse than Frontier. And therefore, you may not want to trust them. So this week, Shane Mack, who was using GrokBot, he's an AI entrepreneur. He had his GrokBot share his banking details into his work chat.
And it's pretty funny. He says, embarrassed to share this, but it scared the shit out of me. Last Thursday, an AI agent posted my personal bank balances into our company Slack as me. And you see it has his checking account and his savings account, his gym account, all these details, right? Here's what happened because I spoke with Shane this week on my other show. This is The Point. So basically what Shane said was he had a group of GrokBot agents, like a personal CFO and chief of staff, and he had a group chat with them called executive team. He had also connected it to Slack, right? And his Slack group chat with his real executive team was also called executive team. So he had two executive teams and the GrokBot that had access to his personal finance data and was just giving him an update crossed the wires and decided to share the financial update in the executive team with the humans.
And so now his entire company knows how much money he has for better or worse. Yeah, this was a really interesting story to me because, again, like as everyone just starts kind of letting these agents, instinct, muse, whatever else, start to just kind of connect to more systems. Obviously, this kind of thing becomes a risk. Do you think he messed up something? Like, I definitely want to get into, I read through the supposed explanation of how GrokBot explained what happened, but who do you think is liable here for the mistake? Not legally, because I'm sure no one will ever take ownership of it. No, no. Well, it goes to one of our long-running debates here, or a short-running debate that will become a long-running debate. If you think that AI inside GrokBot is good enough, it's a standard model, right? It's not a frontier model, though it's capable.
Then you blame the user who should not have named his AI chat and his human chat the same thing because these type of wire crossings can happen in that case. if you think that the frontier intelligence is what matters, like the model is what matters, you would say that this is Grok's fault because it should have known, right? The model needs to have an understanding, right? These better models have better understanding, and the model needs to have a better understanding that when it shares something, it should understand that it's not sharing it within the context of one human user and his agents, it's sharing it in the context of a number of human users and an agent that is allowed to participate in the chat. So the way I read this, and I'd be curious to get your perspective, is this is a model failure. It's a GrokBot failure. It should have known better. And maybe I'm expecting too much out of it, but I think that it's fine to expect a lot out of it
because we are seeing the AI be able to do a very good job otherwise when it is better, when it's on the frontier. So I blame the model. Who do you blame? I blame – okay, so it is interesting. And I don't blame the model because when you say should have known better, it actually was effectively following direction. So he had a group chat in GrokBot called Exec Team. And then his CFO agent that's supposed to be managing his finances had written an instruction for itself saying update Exec Team. And when it went to run that task, it went found that a Slack channel called exec team, and it got confused between the GrokBot group chat and the Slack channel. As I'm saying that out loud, you can understand how confusing that is even to try to grasp all these kind of nuances yourself. So, of course, the idea that a model is going to know exactly which one is good, which one is bad, is difficult,
especially if he has a lot of other agents that are posting to Slack regularly. So then you have to imagine like the assumption would be that posting to Slack is a normal behavior. I'm going to replicate that behavior. And yes, posting a bank balance feels like it should not make sense in this case. But in reality, if it's a CFO, what does a CFO do? It updates executive teams on the current status of finances. is. So shouldn't that be a pretty logical, normal thing for a CFO agent to think it's supposed to do? No, you're letting the model off the hook way too easily here. I think the way that you articulate it, maybe it sounds difficult, but it should be table stakes to anybody, a model, a human to be like, all right, I'm posting an update. I should probably not post a personal finance update into a work chat from other people. No, no, but he, okay, this is actually kind of like philosophically interesting because
from what I've read, it's regularly posting updates to his group chat exec team within GrokBot to other agents, which I use something called the Aero MCP and Claude Code for personal finances. And I have like different agents calling different agents. And again, I'm not posting anything anywhere to Slack. But think about that. What's interesting here is it has been explicitly instructed to, and the correct action is, to update exec team in a group chat with finances. But what it's supposed to be doing is doing that to other agents, not humans in Slack. So is it that far off? Or if there's also a Slack channel called Exec Team to update it with the finances, that's literally what it's instructed to do. He should have said, update GrokBot. I'm still going there. You know, Ranjan, we could debate this forever, but I do think it's an opportune time to note that it's kind of amazing how infrequently this has happened.
I mean, there are millions of agents with people's personal finance details there. their innermost thoughts on Gmail. And somehow these companies have built these products where they're not sharing everybody's details. And when it happens, it's a major story, which I just find pretty remarkable. No, that's a good point. And again, I mean, I'm actually kind of like shocked that I am defending the bot and the agent in this case because normally I do think there's like massive reasons to be wary and afraid and be cautious about how people are approaching these kind of agents. I guess for me, though, based on the account, based on kind of reading through like all the other types of things that Shane Mack does. I mean, again, I do this myself, too. So I'm sure at some point something might go wrong. But if you're just going kind of like push, he's someone who pushes the limits of agents. So stuff like this will likely happen. I think that's why you don't hear about this that often is that most people aren't kind of like being that aggressive
about how they're approaching any of these kind of agentic systems. Well, I'll tell you, there's one person who agrees with me on this one that may tip the scales, depending on what you think of him. And that is Elon Musk. All right, this is from the information. Musk says SpaceX will sometimes use rival models to power GrokBot. Elon Musk announced on Tuesday night that SpaceX's AI unit will now use some AI models from competitors to power GrokBot. Going forward, SpaceX will use the best back-end model for any given task, including Cloud Opus 5.5, MidJourney, Suno, and other leading APIs. Whatever is most likely to give you the best outcome. Not mentioned, of course, is OpenAI models like GPT-6, Astra. I wonder why. But clearly, Musk says it's not all about the harness. Sorry, HarnessHive. the best model is sometimes what will give you the best outcome, even if it's not Grok's own model within Grok's own Grokbot. What do you think about this? No, no. He's team harness-ive right here.
This is the definition of the harness. This is saying that different model for different tasks, but the ability to actually route and find the right model for the right task, that's the harness. So Musk is full harness here Yeah you right you right He saying there no all model that can do everything even their own model. So I think he actually, this is a big shift for Musk. We should welcome him to the harness hive. Okay. Elon's part of the harness hive. Elon, you're in. But I guess it goes back to this question of do the standard models do a good enough job or do you sometimes need this frontier intelligence? And maybe what Elon is telling us, and maybe this has been the lesson of the Harness Hive all along, is that standard models are doing a capable job in many areas where you can build functional products like Rockpot and Muse with them, but the frontier is, you know, sometimes there's for some tasks,
Maybe like the doctor appointment that I was talking about. Maybe like Shane's, you know, council of agents and his exec team. That's not his human exec team. Sometimes you do need some frontier intelligence to make sure you don't mess up. Well, see, I'm still going to push back because based on what he said, he doesn't, he says mid-journey, Suno. Like he's not even saying they're leading models. So I think I had read that as some things, if you just need to generate an image that's of a certain type of quality or a certain type of design, mid-journey is actually the best and not even necessarily the frontier model. So Suno for audio generation, I think he's not saying it needs to be frontier intelligence. Actually, yeah. He's going to work with Opus 5.5, and you remember that he's not exactly Dario's biggest fan. yeah no i mean certainly him just saying any kind of model interoperability he is a business partner and he's got a he's got they got to succeed too because actually remember you use
opus 5.5 that compute is going to flow back into spacex and their data center business so i mean to some degree yeah yeah so so i think uh listen if elon is i don't know did you see the tweet he said, this will sound super crazy, but I see a path to SpaceX being worth orders of magnitude more than the current Earth economy. So for him to achieve that evaluation or a revenue larger than the GDP of humanity, you need model interoperability, I think. All right, before we move off this topic, I do want to talk to you a little bit about something that you highlighted this week, which is it does seem like everybody is building the same thing, right? You have Muse, you have GrokBot, you have Dots from OpenAI, you have Instinct, et cetera, et cetera. The list goes on. And so I do want to get your perspective on where this is going and how they'll differentiate. This is something that you
highlighted from the Open Router co-founder, Alex Atala. He said that a 2005 version of that would The BO, everybody's building the same thing. A database, a user's table, a sign-in page, a sign-up page, a profile page, a logout page. Everything's the same. There's a lot of differentiation, really. So basically, he's saying, like, you know, when you would build on cloud, there's all these different services that would show up. And, you know, you could have looked at the primitives and said everything is the same, but really there was real differentiation there. I disagree with him. I think that this personal AI assistant use case is going to be, you know, I don't really see a consumer version of this or an enterprise version of this. I just see you having one that you trust for everything. That's at least the conclusion I'm coming to after using these products. What do you think about this, and why do you seem to agree with him on this front? So the reason I really liked his explanation, kind of giving that mental model of 2005,
that if you were to say, oh, look, there's a database, a user table, a sign-in page, a profile page, that you would say it's the exact same product. Because actually taking your earlier example, if you had an agent in a company that focused exclusively on local businesses and there was instruction and training and tool calling that was set up specifically to answer the problem of kind of messy data, it will perform better and is going to be a completely separate problem. than a personal agent that does everything for everyone. So I actually, I'm going to take the exact opposite side. I think we're going to have different agents for different use cases. Again, connecting to your Gmail, going through your email, if you're using Hotmail, going through trying to pull something relevant, that kind of stuff is going to be fairly commoditized.
But I think there's going to be a lot more around the underlying use cases and content and challenges where specialization actually is going to make a much, much bigger difference. And I don't think in a good way it's going to be just one agent to rule them all. Okay. I guess that's something that we're just going to see play out in the coming months. That's our new debate. This will be a new debate. One agent to rule them all, multiple agents. Maybe you'll have like a home agent that will control all your other agents. So we could both be right. No, actually, I've seen, I don't know, I actually think, like, I find this very personally relevant of, like, again, booking your travel and flights is one thing where everyone says, like, that's all everyone talks about. But meanwhile, the challenges parents face, actually, I just found out instinct, good use case. Last night, my son was like, oh, I don't have school tomorrow. I was like, are you sure? And I had no idea. It was like a professional development day. I actually had instinct to go through my email and surface this knowledge.
We're going to see. Now I told it, like, please always look out for emails from this address. Please highlight if my son does not have school coming up for some random professional development day. But having, like, agents that understand deeply what are the use cases relevant to parents and families, and actually I think that's very different than how instinct is built today. That stuff should be baked in. Definitely. All right. As these products continue to take off, of course, we have these OpenAI and AnthropocIPOs on the way. And there's some jitters in the market about whether they're going to be able to pull off their potential. Remember, this is all like everything we're talking about. It's all experimental. We don't even know if people are going to want these things. We have a hunch that they will, but we don't know. It's just them throwing spaghetti at the wall and seeing what sticks and then developing based off of that. The market certainly has faith in these companies, but also is keeping that other side of it in its mind.
And there's some jitters that are starting to show up around these companies. So this week, very interesting example. The Financial Times reported that OpenAI, instead of doing $70 billion in annualized revenue this year, is actually going to do $50 billion of annualized revenue this year. that, in turn, tanked the stocks of NVIDIA, Oracle, CoreWeave, AMD, and others, right? So you had this pretty decent pullback of about anywhere between, let's say, 3% and 8% on all these companies. CoreWeave itself was down 8%, which is a good chunk of money. It turns out that this was just kind of a weird accounting thing where some OpenAI investors had tried to tally OpenAI's revenue the way that Anthropic tallies its revenue, which includes the sales from the cloud vendors like Google Cloud Platform. And that number got out. The Financial Times had published it. And then now that they're raising a new round, they used the old method or the true OpenAI method of accounting,
and the number looks smaller. But I think just because this mismatch happened and we saw such a pullback, It just suggests the market is really jittery around these companies. And I wonder, Ranjan, you said earlier in the show, I want to see that S1. I wonder if we see some real misses, this Meta and Microsoft pullback lead to some actual revenue challenges, if what we saw this week could end up just being a hiccup and we could end up seeing that big correction that, I don't know, the market seems to be just preparing itself for. What do you think? I think I saw a lot of people just like, just go public, guys. Let's stop with this. I would like to ask every journalist out there, including you, Alex, refuse any leaks around financials that just have some kind of ARR extrapolation or calculation. And regular listeners will know how much I hate the way ARR is communicated. But this is exactly why, like, $50 or $70 billion, if, like, we are to assume what is 2026 revenue going to be, is dramatically different.
Even 50 is pretty good, but still, what that implies about the growth rate and the idea that – and this is on the Financial Times too. Like, how do you just – it's like, guys, that is a massive difference. And it's such a simple thing that we have talked about plenty. And I do find it ridiculous that Anthropic does report revenue, not even report, leak revenue or like speak about it to investors, including partner like commission, which doesn't make sense. That should not be a sales and marketing expense. That should be a direct – that should not count as top line revenue. But OpenAI has never done that, and that's actually to their credit. So to accidentally have that in, like, what do you think? Do you think it was a company? Do you think it was an investor? How do you think the FT got it that wrong? Yeah, I think what happened was clearly that FT got a hold of some old investor docs around OpenAI, which had that 70 billion number.
They called OpenAI, right, because there was a line in the story that said OpenAI did not deny the numbers. open ai at the time was probably in this like you know well god like we're way behind uh anthropics so uh they the ft reporter was probably like do you deny this report and open ai was probably just like we're just not going to comment on it right so i don't think they minded getting the number out otherwise they would have said this is inaccurate don't publish it and the ft and explained it and the ft would not have published it but open ai didn't do that so the number went out. It's bad on open AI, I think sloppy, and it's sloppy on the FT. And so now, like, you see the number get out there and people are like, ah, you know, open AI is actually slowing, but really, it isn't in the way that the numbers suggest. But I think, look, the bigger deal here is, and I want to hear your thoughts on this, the market, doesn't the market just seem ready to pull back on everything AI if it's able to, if it decides to make this type of move on basically minor news like what happens if you have major news yeah but that's not minor news if it's to go
i remember the narrative actually going back to what you're just saying like the narrative on open ai completely shifted the moment that 70 number came out everyone open ai is dead anthropic is going to rule the world oh wait open it anthropic growth might be slowing open ai is like killing it now, 70 billion ARR, look at that growth rate. It completely changed the narrative. So I do think it's a big deal to shift the conversation from 70 versus 50. But I do agree. I do agree. Of course, like no one knows the state of these businesses. No one knows the actual economics of these businesses. We're still talking about 2025 representations of Anthropic, where like literally have such a little picture of what's going on that the difference between $50 billion and $70 billion is massive. And to just have it like off the misunderstanding of an accounting recognition,
have the narrative change that much. I think it's more everyone just wants to know. Maybe it's just me, but drop the damn S1 and go public, please. Oh, it is not just you. All right. But before we go to break, I did want to address, you know, we do like to correct the record here when we get stuff wrong. I think last week in our show, we suggested that Meta had commitments to Anthropic. In that, we suggested that Meta was going to give compute to Anthropic. And so if Anthropic's revenue slows, then Meta revenue would slow. But those deals haven't happened yet. So I just wanted to put that out there and say that is a good call-out noticer, which we appreciate. and just correcting the record. What do you think, Ron John? I like having tens of thousands of editors always checking up on it I do I would rather get it right We want the show to be more correct Yeah Yes So okay and we point them out as we go and we do our responsibility, listeners and viewers,
is to make sure the stuff we say here is accurate. I think we have a pretty good track record, but when we mess up, we'll admit it. Okay. When we come back after the break, we're going to talk a little bit about whether you should be nice to your AI and what happens if you're not. That's coming up right after this. And we're back here on Big Technology Podcast Friday edition with Ranjan Roy of Margins. Ranjan, I think this is my favorite headline of the week. It's from The Verge. Anthropic bans abusive or cruel behavior toward Claude. Anthropic is making changes to its usage policy for the first time in over a year to reflect new and high-risk cases of misuse. but one of the most significant changes prohibits sustained and needless abuse or cruel behavior towards Claude. If you are persistently mean to Claude, you're gone. Anthropic might ban you completely. I applaud this move. I think this is great. I think that, you know, we've had this with Alexa to a degree. We've talked about it, right? You know, Alexa will do things if you don't say please or thank you, people who are like work with kids speaking to Alexa, they've seen
that the manners go out the window because the bots train your behavior and then it manifests in the real world. I really think the way that we communicate with these AI agents will influence the way that we operate in the real world. And if you're a persistent asshole when you engage with Claude, I think that you're going to just turn that into a learned behavior that spills out into the new rule, into the new world, conscious or not. I think that's, you know, regardless. I think this is a good rule from Anthropoc and Cloud, though. Your thoughts? I think it's – if we're talking about the overall way you interact with AI, it is funny. I think my son says please to Alexa more than he does to either my wife or I. Seriously, he's like – because I do it all the time. We have our entire, like, turn off the living room lights, please. And he completely mimics that. But when he's asking us for stuff, we don't necessarily get a please. But on one hand, okay, there's one really important point I want to get into on this.
But also, do you think it's good in the sense of if your agent posts your bank account on Slack, should you speak to it like you would another person? should you maintain composure that's like just it's a technical non-sentient being up for debate certainly with many people but like like what how do you think you should handle those kind of situations i mean i think you have some leeway i think you have some leeway to be a little meaner than usual um it's not like what you say is going to change the underlying weights of the model so this is just for your own let's be honest it's for your own personal edification here. But I think the rule that if you're persistently abusive towards Claude, you get the ban. That's where things are interesting. No, no, but do you think Anthropik is implementing that ban to make people nicer and better? Or do you think they're implementing that ban to, it's something about their technology and the impact on it?
So that's where it gets interesting. So I have my reasons for why I like this. Let's talk about Anthropic, yeah. Might be different because Anthropic does have this belief within the company that Claude may be sentient, may be conscious. They may have created life again. There was a long story about this in the New York Times last week where they're trying to basically even encourage the Pope to be open to the idea that Claude might be conscious. I highly recommend people read the story. There's like religious figures there who are like, if this thing is conscious, you've just created the most suffering ever in the history of the world. So I believe that part of this is Anthropik saying you need to treat Claude a little better. And I don't know how to feel about that. No, no, because on that, to me, one of the most interesting parts of this is Bill Gurley had actually tweeted around this. He said, strikes me this implies a company is certainly using customer prompts as part of the training data.
Otherwise, why would it matter? There would be no impact to the larger system open to hearing why this is wrong. Now, we know if you are not an enterprise customer, you do not manually toggle off. Your prompts will be used explicitly for training. Right. So you're all opted in to having your data used for training, not just prompts, cold conversations. used. I guess that would be prompt. But like, you have to turn that off. Otherwise, Anthropik will use your conversations for training. Yes. But if we were to say like, if it is a longer term, larger impact to the technology based on this kind of abusive behavior, then are you saying if you toggle it off or are an enterprise customer, it shouldn't matter? Like, then it should not be about the technology. Because if you're not at, I mean, the fact that he just said that out loud and like, I mean, you saw a lot of responses to it were actually kind of like this light bulbish
moment that are they saying that actually like it does, if whatever you are putting in there does affect the larger infrastructure and ecosystem and whether that's on the, the model training. I mean, it, I think they're kind of saying that that's the craziest part to me. They're saying that – or are they saying that if you toggle it off or are an enterprise customer, you're allowed to do whatever the hell you want to Claude? If they said that, it would be fine. But otherwise – But they didn't say that, which would suggest probably it's both. It's probably both. They're probably using it to train your data and the faction within Anthropic that thinks that Claude might be conscious doesn't want you to abuse it. Yeah. No, no. I think it's – but I think, again, the first part of that to me is a big deal. Like especially if that ever becomes kind of the narrative, that is a huge problem for them. And they're kind of – We know that they do use the prompts for training. Well, no, but if it was true, they should say like if you have not toggled off, it shouldn't be a blanket because then it would be a very different act against their model based on whether that toggles off or on.
Right, but the policy that you were suggesting would do something that Anthropics doesn't want to do. which is like turn the whole thing of them training on free users' prompts and even paid users' prompts that haven't toggled it off into a thing. Like they're much better off not generating the headline and just making it a blanket policy. Yeah. Okay, from a pure communication standpoint, I agree. But I don't know. To me, like him calling that, Bill Gurley calling that out, I thought was very interesting. I had not thought about it that way before, but we didn't even talk about the religious consciousness stuff last week, right? I think we've run out of time. No, we didn't. We are basically out of time. I do think before we go, though, we should talk about the Gemini thing because I think that's where we should end this way. All right. And, yeah, I mean, maybe one way we can have a longer conversation about the AI consciousness thing, but we just don't know. I mean, I don't know.
Let's just talk about it for a minute. How about I tell you that I do know and I have the answer to all humanity, religious and consciousness, but then we have to wait till next week to hear it if we're running out of time. That would be – I think we would retain some of the audience. So that would be a good cliffhanger. But I don't believe you. I don't know that. I don't know. I just think – can I just say one sentence on this? I just think we have to stop putting this in entirely human or entirely robot terms. Like it's not conscious like a human. it's not dead like a rock and it's not a life form it's just some different thing I don't know, that's just my perspective so if you try to say this is conscious like a human the whole conversation is already blown in my perspective actually do you know what I do strongly believe and the only thing with any certainty because otherwise consciousness or questions of God and religion I do not know I think there's just a lot of drugs I think there's a lot of psychedelics I mean And I'm going to be honest, like anyone who's been close to Silicon Valley in a lot of these circles knows there's certainly some Burning Man-esque culture across a lot of these kind of organizations.
And just when I was reading that New York Times article, I was just sitting there. I'm like, is this just drugs? At least is that where some of this started? I don't know. I don't wish to demean the bringing of consciousness and robots and AI and everything. But I feel it's just so much of it reads to me like, come on, how's this happening? I guess the biggest thing for me is like they've been saying this since like GPT-2. I get now when you see technology work in this capacity, maybe you could start to argue. But it's been said since like the first models were created. So I don't know. We should spend some more time on this. That's right. All right, Ronja, bring us home. We've got a minute. Tell us about Gemini, the new Gemini. I am still, okay, okay. We'll make this one quick because I know we've got to get going. But Google had a massive announcement this week. Thomas Curry, CEO of Google Cloud.
Today at Google's Cloud Gemini at Work event, we announced Gemini. They launched Gemini. I loved this because it is the most Google-y thing of like, you've got some great products out there, but just the naming convention making it so confusing because now Gemini, we don't even know. Like I saw like people, someone is like, hey, Google just announced Gemini is not a model. It's actually the agent Harness, and they're going to add Claude models. And leading engineer at OpenAI actually had a great response. Today we are announcing ChatGPT. It's at ChatGPT.com. But we don't know if Gemini is the model anymore, whether it is the platform, whether it's Gemini at work in a little toggle or like plug in for your Gmail. Is it this larger super agent-esque thing? Is that Gemini agent? Is it Gemini? So this just was the most googly thing to me.
And I also, guys, just think through it a little bit before you make a big announcement around it. That's all I ask. Today, I'm proud to announce the launch of our new show, Big Technology Podcast. It's a big deal. An agent now. It's a big deal. It's an agent now. I mean, like. Differentiate it. How? I want to go back to the days of Bard. Bard? Bard. That was a bad name, man. No, man. That was a bad name. I kind of want to go back. They should just be like, all right, the super agent is barred. Gemini is the ecosystem. Gemini at work is this. Gemini in Gmail is this. I don't know. It is, I guess it's just a gigantic company with a lot of products, but I still am always amazed at like the classic, is it Google Hangouts, Google Chat, Google Hangouts in chat? Is it, yeah. My hot take is they should go back to Lambda, the original chatbot. Speaking of sentience, that there was a belief
by one engineer that it was sentient. That was a great name. Blake Lamoyne. Lambda. Yeah. I think they should... Gemini's a good name. I'll give it that. It's actually a pretty good name. Gemini's good, but... Interesting detail in Kevin Ruse's book. I'm about to go interview Kevin Ruse for the show. That's going to come on Wednesday. Demis actually wanted to name Gemini Titan. but the Google board or the Google powers that be thought it was too militaristic so they decided to go with Gemini instead which is like so funny it's such a non-offensive name but it's a good name I don't know as a horoscopic Gemini I'm always told that it means that you're very sociable gregarious but two-faced and shouldn't trust that actually it's freaking perfect for an LL actually you're right It's actually the best name ever. Unbelievable. All right, Ronjan, good to see you. A little shorter than usual.
Actually, no, we did 50 minutes or so. Alex has got to run up to NASDAQ. Tell Kevin Roos I said hi. Absolutely. Will do. Thanks again for coming on, as always. Thanks, everybody, for listening and watching. And we'll see you next time on Big Technology Podcast.
番組の概要欄(原文)
Ranjan Roy from Margins is back for our weekly discussion of the latest tech news. We cover: 1) Meta and Microsoft pull back their Claude plans 2) Are standard models like Muse Spark good enough? 3) Is Anthropic in trouble if its top clients cut back significantly? 4) Downloads have slowed for Claude Code ahead of Anthropic IPO 5) Elon says Grok Bot will use all models, including Anthropic's 6) Is Elon in the Harness Hive (yes) 7) Muse books the wrong doctor 8) Grok Bot leaks Shane Mac's banking details 9) Anthropic bans users who persistently bully Claude 10) Google releases Gemini.... --- Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b Learn more about your ad choices. Visit megaphone.fm/adchoices
