
エージェントのボンネットの中をのぞく:AIエージェントのオブザーバビリティ
Taking a look under your agent’s hood
Taking a look under your agent’s hood
The Stack Overflow Podcast
要約
Stack Overflow PodcastでRyan Donovanが、Datadog CPOのYanbing Liとエージェントの可観測性について議論した。エージェントは知的なモデルに支えられた分散システムであり、従来の指標・トレース・ログに加えて、推論やツール呼び出しなど振る舞いの観測が必要だと述べている。開発と本番の境界が曖昧になる点や、AIセキュリティ、トークン経済性(ROI)が未解決の課題として語られた。
- ●Yanbing Liは、AIエージェントは本質的に分散システムであり、メトリクス・トレース・ログという従来の柱は今も基盤だと述べている。
- ●観測対象は、プロンプトと応答だけでなく、推論、ツール呼び出し、意思決定、軌跡全体へ広がり、「ソフトウェアの健全性」から「振る舞い」の観測へ移りつつあるという。
- ●テストでは網羅的に検証できないため、本番の継続的なフィードバックと評価を開発側へ戻すフライホイールが重要になると述べた。
- ●観測データはMCP経由で提供され、SREと開発者の知識の境界が曖昧になりつつあるという。ダッシュボードは残るが、利用の場所や方法は変わる。
- ●セキュリティと可観測性は要件が重なり、トークン消費をビジネス成果にどう結びつけるかは未解決のフロンティアだと位置づけた。
章立て
ゲスト紹介と経歴
Yanbing LiはGoogleでの可観測性、自動運転トラックのAuroraを経てDatadogに至った経歴を語る。
エージェントに可観測性を適用する
エージェントは知的モデルが支える分散システムであり、従来の柱に加えモデル相互作用や振る舞いの観測が必要だと説明する。
本番と開発をつなぐ継続的フィードバック
検証し尽くせないエージェントでは、本番の評価を開発へ戻すループが重要になり、SDLCの自動化も試みられている。
境界の曖昧化とLLM評価
SREと開発者の境界が薄れ、本番と開発の評価を一体化する価値や、評価は可観測性に含まれるかを議論する。
MCPとダッシュボードの行方
MCPの提供方法や、ダッシュボードとAIエージェントによる障害調査の使い分け、SRE調査エージェントの事例を語る。
セキュリティとトークン経済性
AIセキュリティと可観測性の重なり、トークンのROI、モデルルーティングといった今後の課題を述べる。
解説記事
Stack Overflow PodcastのRyan Donovanが、Datadogの最高製品責任者(CPO)Yanbing Liを迎え、AIエージェントの内部で何が起きているかを観測する方法について話した回である。Liは、GoogleでのObservability(可観測性)の責任者や自動運転トラック企業での経験も踏まえ、エージェント時代の可観測性を語っている。
エージェントは「知的な分散システム」
Liは、AIエージェントは根本的には分散システムであり、その内部が知的なモデルで動いている点が違いだと述べている。LLMの可観測性はプロンプトと出力の理解から始まったが、現在はエージェント全体の理解へ広がっているという。推論、ツール呼び出し、応答に加え、なぜその判断をしたのか、アプリケーションや成果にどう影響するのかを把握する必要があると語った。ソフトウェアの健全性だけでなく、振る舞いそのものを観測する方向への移行である。
メトリクス・トレース・ログといった従来の柱も、依然として基盤だという。Datadog自身のエージェント可観測性製品の統計では、検出されるエラーの大きな割合がAPIのレート制限やデータアクセスに関するものだと紹介された。エージェントの速度と規模によって、従来からある問題がより顕著に現れるという見方である。
本番フィードバックを開発へ戻すループ
Liは、従来のソフトウェアではデプロイ後に観測して修正する流れが一般的だが、エージェントはテストや評価をどれだけ重ねても現実が予想外の形で驚かせるため、継続的な監視と素早い反復がより重要になると述べている。そこで、オンライン評価の結果を上流の開発へ戻す「フライホイール」が必要になる。
AIで開発が加速し、並列に多数のエージェントを動かせるようになると、ボトルネックはソフトウェアの出荷そのものに移るという。Datadogは自社のSDLCで、テスト生成の自律化やデプロイ後の段階的ロールアウト(良い信号なら拡大し、悪ければロールバック)といった自動化を試していると説明した。ローカルのエージェント向けにも、Datadog MCP経由で本番のコンテキストを使える体験を構築中だという。
開発と運用、観測と評価の境界
AIによって知識の壁が下がり、SREがコードを調査・修正したり、開発者がMCPで本番のテレメトリを参照したりするなど、観測の役割の境界が曖昧になっているとLiは述べた。AIネイティブな顧客でこの傾向が強いという。
LLM評価についても、本番ではLLM as a judgeを使うのに開発では別の方法を取る必要はなく、両者を連携させる利点が大きいという考えを示した。評価が狭義の可観測性に含まれるかという議論よりも、開発と観測を橋渡しする製品の価値を重視している。
MCPとダッシュボードの未来
MCPサーバーは単なるAPIラッパーではなく、フォーマットやエラーメッセージ、エージェント向けの賢いクエリ方法が重要だとLiは述べた。さらに、顧客がMCPでデータを取り出して使うのか、ベンダー提供のエージェントを使うのかは、今後の議論になると見ている。
ダッシュボードは「一枚の絵は千語に値する」ため残るが、Slack、CLI、IDEなど作業の場所に可観測性が出ていくとも述べた。ある顧客のSRE調査エージェントでは、約100件の調査で平均3分強で原因に到達し、人の介入もなかったという事例が紹介された。人間は運用から離れないとも強調している。
次の課題:セキュリティとトークン経済性
Liは、エージェントの挙動が意図通りかを理解するのは最初の一歩にすぎないと語る。AIはセキュリティ上、標的にも攻撃者にも攻撃経路にもなり得る。エージェントの全ステップを把握するという可観測性の要件はセキュリティと重なるが、適切な手法はまだ確立していないという。
もう一つの未解決課題がトークン経済性である。クラウドではコストが事業成長と連動していたが、AIの価値はそれと異なり、トークン消費をビジネス成果にどう結びつけるかが課題だと述べた。Datadog自身は、静的・動的なモデル選択、大小モデルや自社モデルの併用、基本的なMLモデルの利用などを実践していると語っている。
まとめ
エージェントを本番に出すには、プロンプトの記録だけでなく、軌跡全体の観測、評価、開発へのフィードバックを一体で設計する必要がある、というのがこの回の主張である。日本の開発・運用チームにとっては、従来の可観測性基盤を土台にしつつ、評価、セキュリティ、コスト対効果までを視野に入れる点が注目できる。いずれもLi自身が未解決と認める領域であり、今後の議論の行方に注目したい。
文字起こし(英語・自動生成)
Hello and welcome to the Stack Overflow podcast, a place to talk all things software and technology. I'm your host, Ryan Donovan, and today we're talking about agents. What are they building in there and how to find out? My guest for that is Yanbing Li, who is the Chief Product Officer at Datadog. So welcome to the show, Yanbing. Thank you for having me, Ryan. My pleasure, my pleasure. Before we figure out what those pesky agents are doing, can you give us a little flyover of how you got into software technology? I feel like I've been software forever, for most part of my life. I would characterize my education as I've studied until there's no more school to go to. So I have a PhD in electrical and computer engineering. Then I've been an engineer ever since. Except for a data doc, I'm far more focused on product than combinations of product and engineering.
And my introduction to observability started with my time at Google. This is seven years ago when I was leading the observability effort at Google, both for powering Google's infrastructure as services, as well as being built as products as part of the Google Cloud capability. So that was my first introduction to observability. then I had an interesting detour into everything AI where I was building self-driving trucks at a company called Aura. And Aura today is probably still not just the first but the only company that's operating driverless trucks commercially on U.S. public road. So that was kind of the experience of absolutely needing to understand everything that AI agent is doing. You know, granted, self-driving started before the chat GPT LLM days, but there is definitely a lot of parallel in thinking about what it takes to fully unleash your agent into the wild. Observability is one of those things that's really ramped up a lot since, you know, Datadog used to be, you know, metrics and dashboard company, but now fully part of the observability space.
And everybody's trying to figure out how do you apply observability to agents? How are you thinking about keeping an eye on agents when they're running? As you said, observability has existed for a long time for really complex distributed system. You know, that part remains the same. But with AI agent, there are certain things that's new and different. You know, first of all, starting with, you know, AI models playing an instrumental role in bringing reasoning and intelligence into AI agents. You know, now we deal with the non-discriminate nature of AI models. I think that tends to be the first thing that when we talk about AI observability, we started with LLM observability, really to understand what the LLM is doing. But now that is quickly expanding from just observing the AI models and the interaction with the AI models to really understand the AI agent as a whole. If we think about AI agent, it's still fundamentally a distributed system. It's just this distributed system is now powered by intelligent models underneath.
So there are certain things remain the same. We have to understand these new nuances that's introduced by AI. Yeah, the observability of LLM's is still a sort of ongoing research project, as I understand it. To get a sort of greater detail to explainability, is that right? There's certainly certain things at the research frontier, but I also think there's a lot of things that's being in production. The fact that we do rely on AI agents in our day-to-day use in some of the enterprise use cases is an indication that LLM observability or agent observability in combination with other things is giving us a level of comfort to rely on these AI agents. And I do think it's a journey because, you know, I work with a lot of the enterprises. They're still in this early phase of AI development. Can I trust my AI agent not to stay up, but also to do the right thing? I've seen a lot of successful agents, but, you know, when I speak to our customers, many more of them are still in the process of building that.
And given the rapid landscape of AI, I'm sure this how we do it is going to be evolving continuously. The LLM was you give it an input, it gives a response, and now there's multi-step, turn-taking, there's tool use. What are the pieces that you have to observe in an AI agent? I think you said it really well. Certainly, we've started with understanding prompts and LLM output, but because agents has become this distributed system in nature, so it's becoming a lot more complicated. We do understand the reasoning, the tool calls, the prompts, and the responses, and we need to understand the behavior. We need to understand why it's making certain decisions, and we need to also broadly understand what's the implication across the applications and outcome. So I think we're definitely evolving from simply observing just the software stage, software health, to truly observing software behavior.
The comparison to network software is interesting. I've been saying that AI agents are sort of speed running the microservices path. And for the network software, observability has kind of had three or four pillars, the logging, metrics, traces, and events. Do those traditional pillars apply to the AI software? I think they do still apply, but there is also more things that is definitely emerging in the realm of, you know, agent observability. We're observing additional characteristics. So the traditional pillars still is important from what I see in our customer base. We've seen AI generate tremendous demand for observability, and a lot of that is due to the red and butters of the microservices, the distributed system observability of metrics, traces, logs, that is still very, very foundational. But then you added this additional layers of uniquely understand model interactions and agent behavior. So I think it's actually the need for observability is growing in both of those.
I was actually looking at some of our own agent observability product statistics just to see what kind of errors it catches. For example, it catches a big chunk of the error in hitting rate limit of API calls or data access. This has nothing to do with agent per se. It's a problem that exists, but it gets manifested even more pronouncedly in agent given the increased speed and scale at data access. So still a lot of things remain the same, but then there are these new primitives as we've been discussing as well. Yeah, yeah, with AI agents, I feel like you can automate more failures. You know, certainly there is token limit, but it is often just a regular rate limit. And, you know, a lot of the software, they were designed for human or programmatic access. And, you know, agents are exposing a lot of rate limit. And, you know, we also constantly receive demand from our customers of raising rate limit in our system.
I think this is something that's old, but it's new in a different way. You know, when you think about observing agents outside of the standard failures, you can have incorrect tool calls, you can have PII exposures. You can have all these things that are kind of outside of most, you know, network systems that aren't AI. Are you thinking about applying observability to those sort of issues? Yes, we definitely are applying, you know, observability to those issues. You know, when we started doing agent observability, the initial focus was much more on just Aladon call itself, you know, the prompt and response. And clearly, this has now expanded into the full trajectories of, you know, the calls and the loops. And so, given the sophistications of agent has increased tremendously. And so those are all the things we observed I also think simply you know having telemetry and having you know those signals is also not sufficient because you know how that translates into the understanding of what is not working and where do I need to improve
And if I think about the contrast between the traditional system and the LLM system is the traditional cloud software system or, you know, it requires a production feedback loop, but you deploy, you trust in production fairly, you observe everything, then you go back and you fix bugs or fix issues as they arise. But, you know, for AI agent, I think there is an argument for much more continuous monitoring because a lot of the things that can be verified in a traditional system simply cannot be exhaustively verified in a AI agent system. So I think that feedback loop is much, much more important. You know, in traditional software, it tends to be bug fixing or understanding production load or sometimes understanding customer behaviors and how you improve your product. But in the world of AI, all of those things apply. But there is more of no matter how much we test, how much we evaluate during that build phase,
the real world always, you know, surprise us in many, many different ways. So I do think that continuous feedback loop and that rapid iteration is becoming much more crucial in the AI work. And this is where we see in AI, the agent observability is no longer just monitoring production, but also how that become a flywheel that help you iterate the upstream development. So you're constantly improving the outcome of your agent. And again, also this thing is also, especially now with AI coming to development, so there is the expectation of that loop becoming a much faster loop as well. So that is something that guided us to think about, you know, agent observability, not just observing the full production traces, you know, trajectories and tool calls and token usage and all of that. evaluate both through those online evaluations of my AI agent, how it's doing,
and use that to feed back upstream, going back to that build phase, so that we're constantly improving AI. Obviously, a lot of people are talking about loops and agents and feeding back production data to the build phase of the software. Like you said, are you thinking about automating that, you know, mining the signals from telemetry, giving it as a sort of direct geo ticket or whatever for the builders, and then feeding that to an agent to automatically heal the software. So certainly software development has been proven to be kind of still the dominant AI use case on the market where people are spending most of their investment. But we're definitely seeing with, you know, AI accelerating software development that you can deploy, you know, many, many parallel agents to do software development, that shipping software become the bottleneck and shipping software and iterating software. And we think about the software development cycle besides writing the PR,
then you go through testing, code review and testing and merging and pushing it to deployment and maybe in a staging environment and then roll out. I think each of these stages can be automated as well as maybe the whole loop can be automated. I think this is still an evolving landscape. We're trying to do this in our own SDLC. We're also trying to see what is the product angle. So, for example, autonomous testing. What's the new set of code? What kind of new user behavior or product behavior this may be exposing? And how can you autonomously generate test cases so that doesn't become the bottleneck? And obviously, there is also the outer loop of how you turn that SDLC, that more sequential process, into a more autonomous loop that can be paralleled in many different stages. This is definitely something interesting to us. We're trying to do this for ourselves and potentially think about how to do that for our customers as well.
And I see a lot of advanced developmental organizations is all trying to do something very similar. There's a lot of data produced by observability, right? A lot of telemetry. Are you finding ways to use AI to sort of automatically flag those unknown use cases or produce new tests or things like that? This is definitely a new AI frontier that we're invested in. And so, you know, traditionally we've always have testing that, you know, you have certain type of either test case based on previous, you know, known production use cases or synthetic use cases. And the key is, you know, how do you, let's say you're using AI to write new software, how you understand the intention and the goal you want to achieve in those new functionality. and how to turn those goals into new synthetic tests that you can run during your software development before it goes into production. So there is how much you test before you push the software, but as you push the software,
how do you bring that production feedback quickly so that you have that feedback loop working before you start a broad roll-off of the software. I think we've been doing a lot of those already in today's software development. It's just going to be we're using AI to make a lot of those things fully automatically done rather than done by human and also using AI to create bigger loops to, again, automate from a workflow point of view. The workflow is an interesting one because I think most developers encounter agents through their coding agent. Is there a case for applying observability to one zone local coding agent process? Shipping software is about code meeting production. You need to deeply understand what that implication is, and we traditionally do that in the process that defines today's SDLC. And so as I mentioned, there is this way of bringing, first of all, synthetic tests early on,
and also bringing production signals as you manage your deployment, how you autonomously apply FisherFlex to get your road out. And if you get good signals, you expand. If you don't, you roll back. So there is, you know, the autonomous loop that can be created. And we're also in the process of experimenting with a completely local agent because, you know, developers love their laptops and everything local. And we are actually building a local experience that allow the production context to be available, you know, through data dog MCPs and et cetera. The boundary for observability, I think, is shifting. You know, observability used to be squarely is the thing you hand over to whether it's the SRE team or the people carrying that production responsibility and understand. And I think with this new AI-based software development, that boundary is kind of blurry.
For example, the SREs can start to investigate and fix code because, you know, again, AI destroys that knowledge boundary. So if I'm running production with AI, the way I understand what's going on with my productions or dependencies on all the different services I don't own, The SREs have the ability to do that and potentially even make upstream service change if their organization permits that. And on the other hand, I think developers are all using production context for their development through easy access to production telemetry through MCP. And we definitely see a lot of our AI native customers operate in this way much much more pronouncedly The boundaries are blurring Like any time you start automating parts of your process or your software you need some sort of signal that it's going correctly, right? That's the domain of observability. So I think SDLC is being rewritten, and it's definitely an interesting time.
And I think this, you know, coming back to where we started with models, we definitely see the same thing. You know, when I was building self-driving trucks, So we spend a lot of money on training. We spend a lot of money on offline testing. It's essentially offline evaluation, obviously, with self-driving. This is L4 level, basically driverless, so you have uncompromised expectation for coverage, for thoroughness, or for certain behaviors absolutely needs to be guaranteed and verified. So a lot of that is done in an offline fashion, and then you deploy the software onto the trucks, you drive around, you see things, you're running in a development mode or you're running in a production mode. So this is definitely a closed-loop exercise. I think we definitely see the same with LLM development. So when we think about agent observability, you know, it's not just about observing production behavior, you know, getting all these telemetries, but, you know, constantly running evaluation
to understand is your LLM behaving correctly? Is it providing the right quality, the right levels of safety? And we also do the same, you know, bring similar technology to do so during the development phase. Development phase as you, you know, whether you're making model selections or you're building models, you're, you know, tweaking them. Whether you're understanding what's the interaction between prompts and models and tools. Whether you're thinking of I'm running certain set of experiments. You know, all of those things is we see there is the opportunity for those type of tool to be connected because there is a lot of things. I know you in your previous podcast, right? Have you talked about LLM as a judge? This is a very popular way for us to evaluate LLM behavior. But why do you do it in production this way and do it in development in a different way? So when we think about agent observability, we are definitely seeing the need of building capabilities that can monitor in production, but also can provide a similar set of capability to help you evaluate, experiment, and ship your agent product.
And I think other vendors probably are thinking the same. So it's definitely no longer just for the production business. The development side is very closely related. That's something everybody is looking at, trying to figure out where it goes, how you fit that into a system. Do you think that sort of LLM-EVAL system belongs as part of a observability suite? Certainly, from the data doc point of view, we have observability probably written on our forehead. But I do think a lot of the problems go beyond the previous narrow sense of observability. So rather than saying a particular category is evaluation observability, you know, in the narrow sense, it may not be. But from if I were an engineer, I'm an AI engineer, you know, building my next agent, do you want those things to be connected? I see there's tons of benefit for those things to be connected. It's also just like, you know, we talk about agent observability or LLM observability is also not sitting on its own island because all of those things are sitting inside your system.
It accesses your databases, it goes to your lake houses, API calls of other dependency service. It does the same thing except with this, you know, AI magic, you know, besides that. So we also believe, you know, agent observability is not isolated thing. AI agent itself is a distributed system itself and is related to other broader distributed systems. We believe agent observability is part of the broader, you know, system service observability story. And we're also, as I mentioned, whether it's in just general software development or AI agent development, this closeness between build and observe is coming together. So rather than arguing categorically, you know, is this observe view or not, we certainly think there's tremendous value of bringing products that can straddle these two different aspects of software development. You know, we talked about earlier, like, the amount of telemetry is pretty massive, and you're feeding back some of that telemetry as context in MCP servers.
How do you think about delivering the right amount of context from those servers? So we definitely think, you know, from when we were delivering our MCP service, it cannot be just an API wrapper itself. So there are certain things that we do need to become much more specific. Formatting matters, you know, what type of error messages matters, you know, how we give agent a smart way of query as opposed to just, you know, give it data. I think the more interesting thing, you know, I'm also seeing do customers build with MCP? Certainly MCP provides this universal way of accessing a lot of information from different type of AI agent or software tool. In observability, I also see this interesting battle or comparisons of where AI gets built. You know, observability is definitely going beyond just observe, meaning generate telemetry.
Even without AI, observability is no longer just generate massive amount of data. It's already about how I correlate the data, how I surface the right insight to the user. But now with AI, you know, how you take that into much more intelligent investigation and decision and remediation and action taking. You know, if I were a Datadog customer, I'm constantly thinking about where do I consume that Datadog data? Do I take it out as MCP, whether I'm doing a root causing or I'm doing development? Likely I will need a lighter set of data. Do I access it that way? Or do I access the more curated or opinionated, however you call it, the AI agent that we've built for them? Or they go to another vendor that might provide an AI agent by synthesizing the world's data from everywhere. I think it's becoming an interesting debate of what is the right approach. Do you think there's still a place for that team dashboard?
and are you thinking of new ways to consume the telemetry and results? I would expect maybe the access to the beautiful dashboard that we all love in the Datadog UX, that will remain because at a certain time, it's just the picture is worth a thousand words, but there is increasingly this new surface of I want to have Datadog everywhere I go. Maybe I want to access that through Slack. And I can even generate or view a particular dashboard in my Slack or in my CLI or in my IDE or wherever, I would say, wherever the work is happening. So we definitely see Datadog where, you know, Observability in general goes to wherever the work is happening. And what I think is that doesn't necessarily replace the need for a set of, you know, consistent, curated, well-defined dashboard because those are just the things you know. I just talked to an AI principal engineer from one of our customers, and they were saying that's true.
Sometimes, even though we have a natural language-based interface to tell me what's happening with my traces, what's wrong with my logs, they give me this dashboard, it's simply too slow versus they know exactly the few clicks you get the information they want. So obviously AI is extremely good at processing huge amount of information that beyond human ability to process in a short period of time For example in a complex incident investigation you have all the type of dashboards and telemetries at your fingertips but still today's modality is you get into a war room with tens of people or even very complex organizations, hundreds of people. And it's interesting, most people's objectives in those situations is not to find the root cause, it's actually to prove innocence so that they are not part of the problem. And with AI, all of that human-heavy process and time is short-circuited by AI agents immediately understand the context,
have all your histories, memories about what has happened in the past, can see all these telemetry, can understand what has changed in the system, and can investigate in parallel in almost infinite fashion and get to that root cause before a human can open up their computer. Because, you know, when we open up our computer, get on a Zoom, that's a few minutes. I was talking to a customer. They were running our SRE investigation agent, and they have about 100 agents in the past few weeks, and they have pretty robust statistics. and their average is three plus minutes get to that point and without any human. So I think there's certain things it's clearly, and that AI didn't open any dashboards. So it's gotten to, it would be attaching a bunch of reasoning and evidence and dashboard to explain to the human,
but the AI was able to do that in minutes without going through that. So to answer your question, I think dashboards remain as useful because sometimes they're still just, you know, the most direct way to understand and consume information in a very reliable way. But customers are using all kinds of different services to access the same information. And an AI agent doesn't necessarily read dashboard. They go to the source. And also dashboard can be done just in time. A lot of our customers are asking our AI agent to say, you know, I'm seeing some of this, but I want to compose a new dashboard that show me this particular view that's pertaining to the current investigation I'm trying to do. So it's not going away, but it's how we use it, where we use it is probably changing. But we're still obsessed with, you know, details of data, what we call data vids within, you know, our organizations, you know, how we're graphing things, dashboarding things, you know, displacing things. This is definitely details that we still obsess over because I don't think humans are going away from operations.
You know, one of my favorite stories to tell is, you know, today we have commercial trucks running at 70 miles per hour without a human. I don't expect a software system, whether the part we're building it or the part we're running it, doesn't have human involvement anytime soon because it's just far more complex in terms of the radius, the dependencies, and the type of actions you can take in a complex system. I think to me, deeply understanding agent behavior is working as you intended is really just the first step. Because just like how we build any services, making sure it works, that's the first step. But then soon you get to the point that can I secure it? Can I trust it? And I think that's also becoming an interesting problem because I don't think how we secure AI is not yet fully resolved. we are probably a few steps ahead on agent observability than, you know, what is the right
security approach. And AI security is also a multi-edged sword because AI can be your target, AI can be your attacker, AI can be the vector that you deliver attacks or so it's a very multifaceted. I think that's also something we see has close affinity to observability because security used to be built as a different silo. But with AI, what observability demands and what security demands is actually overlapping quite a bit. Basically, you understand each and every single action your agent is taking every step. And that's what we're already doing in observability. But how do you use that also to make sure we solve the security use case? I think that's an interesting area that we're also very interested in. And then there's the third question I think we haven't also solved is the tokenomics or token maxing or the ROI question.
Because when we were building large distributed cloud systems, kind of your footprint grow with your business. You have more users, you have more cloud, and you observe more. So I think the scaling factor is a lot more directly affiliated with some levels of business growth where the value out of AI is completely different. So I think this is also an interesting new frontier. It's probably started with observability of understanding, you know, tokens and consumption to call efficiency, you know, at the primitive level. But how do you link that to also business outcome? Because here is cost, but AI, we all know, is some whatever business value we deliver. So I think, you know, those are just I see the more advanced the topic, but potentially build on the same foundations of you need to deeply understand every single actions and decisions and steps that your AI agent is taking.
But, yeah, just want to make that common. You know, I don't think we're focused on some of these other topics, but I think they are important. They're related and they're still not solved. Like there could be these interesting combination use cases, right? Like even like model routing for certain tasks, you could look at the costs, look at the, eval the performance, and then automatically pick models for that. Yeah, we definitely are already doing that for ourselves. Yeah, we both static as well dynamic choice of models and we both use off the shelf models. We also build our own models. We use large models, small models. And still there are certain things you give it to a basic ML model. That's still the most reliable way. So, yeah, I think all of those things, it's not just, you know, observing, but also, you know, think about how you truly get the right value out of all the AI investments. well it is that time of the show where we shout out somebody who came on stack overflow drop some
knowledge shared some curiosity and earned themselves a badge today we're shouting out the winner of a great answer badge somebody who dropped an answer that got a hundred or more points so congrats to insert username here for answering how can i get a character array from a string If you're curious about that, we'll have the answer for you in the show notes. I'm Ryan Donovan. I edit the blog, host the podcast here at Stack Overflow. If you have questions, concerns, topics, if you want to reach out and yell at me, you can email me at podcast at stackoverflow.com. And if you want to connect with me directly, you can find me on LinkedIn. I'm Yanbin Li, Chief Product Officer at Datadog. And you can find me on LinkedIn. and certainly you can learn about Datadog on our website, our YouTube, our linking channels, etc. Or give it a try through a free trial. Well, thank you for listening, everyone, and we'll talk to you next time.
Thank you.
番組の概要欄(原文)
Ryan is joined by Yanbing Li, Chief Product Officer at Datadog, to talk about applying observability to non-deterministic AI agents, blurring the boundaries between software development and production workflows, and navigating emerging challenges in AI security and tokenomics.Episode notes: Datadog is an observability and security platform for cloud applications, providing real-time monitoring for infrastructure, distributed systems, and AI agent behavior. Connect with Yanbing on LinkedIn and learn more on the Datadog website. Congrats to Great Answer badge winner insertusernamehere for winning the badge on their answer to How can I get a character array from a string?. See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
関連エピソード

AIを形づくる5つの論争:収益とインフラ、利用層、主権、規制、データセンター
The AI Daily Brief: Artificial Intelligence News and Analysis

ゼロから NanoClaw へ:40時間の週末プロジェクトが企業向けAI事業になるまで(Changelog Interviews #686)
Changelog Master Feed

MetaとMicrosoftのClaude削減、GrokBotの銀行情報流出、AIをいじめるとBANされる件
Big Technology Podcast

1034: 2026年9月の見逃し配信(ICYMI)――判断力、ワード・グラビティ、AIエージェントのマネジメントとトークン消費
Super Data Science: ML & AI Podcast with Jon Krohn

AIがソフトウェアとゲームを解体する――すべてのソフトが事実上オープンソースになる日
AI For Humans: Weekly AI News, Tools & Trends

#259 - Dots、Sonnet、自己規制の安全協定、暴走AI
Last Week in AI