← The AI in Business Podcast

Intelligent Operations at Enterprise Scale - with Dhrubojyoti Das Deb of JPMorganChase

The AI in Business Podcast2026年10月9日26分

Intelligent Operations at Enterprise Scale - with Dhrubojyoti Das Deb of JPMorganChase

The AI in Business Podcast

0:0026:45
このエピソードの日本語要約を準備中です。
文字起こし(英語・自動生成)

Welcome, everyone, to the Emerge AI in Business podcast. Today's guest is Driba Jyoti Dashev, Senior Vice President of Global IT at JPMorgan Chase. Driba Jyoti discusses why alert overload has become one of the biggest operational drags on enterprise IT teams and how AI-driven correlation can group thousands of alerts into a handful of actionable incidents. He outlines where human accountability still sits once AI takes on more triage and root cause analysis, and what leaders can do to start moving their teams from reactive firefighting toward more intelligent operations. Today's episode is sponsored by Big Panda. A quick note for our audience, that the views and opinions expressed by Drabajati on today's program are his own and do not reflect those of JPMorgan Chase or its leadership. In this episode, we cover how AI can help teams cut through alert noise.

To go deeper on this topic and learn how to structure landing pages for higher conversion and how to use self-qualification systems to prioritize high-intent leads, download our free PDF report B2B AI lead generation guide at emerge.com.ai.g1. That's emerj.com.ai.g1 to download your copy. Now the conversation with Riva Jyoti. Thank you. conversation today because it's all about trust and alerts and noise and managing all of that.

And I think something that comes up in almost every conversation that I have around this topic is that teams are drowning in alerts and they spend half of their day trying to just figure out which of these alerts are real and need their attention and which is not really an actual problem. And you've kind of lived inside of this. So perfect person to ask, where does the friction actually show up for most teams trying to stay ahead of incidents? Yeah, that's a very interesting question. So right now, I think we have moved from less information or no information to too much information, like going down the line, like if you go back 10 or 15 years, we had systems running with very less logging, very less traceability, very less observability, very less telemetry. But going forward, like as we had evolved, so has the data, so has the logging, so has the tracing. So has the volume coming in? The business has evolved. And so the data has evolved with it. And we have created intelligent applications to log data. But I think at one

point of time, we never had too much thought given into how we are going to handle the data and how we are going to line the applications in sequence so that we know where the point is coming from, where the issue injection point is coming from. So I would say in two words, its alert overload that we are dealing with right now, where we have a complex chain, like a complex chain of systems lined up and with different applications stack, like the tech stack that we call, different applications talk to different applications. Each application has their tech stack. And when they generate the data that is in silo, that data is not getting fed to the next application and that happens and that can get cascaded down. So if you look at the whole picture, starting from top to the bottom, there are multiple applications, multiple levels of data which are kind of sitting in silo, and there is not a central point which is kind of marrying this data together. So if there is an incident, we are tracking incident at different points in the application,

the 0.1, 0.2, 0.3, 0.4, but not on a holistic level. So much of the time gets into figuring out if the issue is actually coming from the application where the alert is coming from. So let's say we have five applications which are talking to each other in a sequential manner or in a parallel manner. And we have a red dashboard on application number four, which is kind of sitting down at the bottom. Now, it so might happen that the data that got generated and fed into application four, it's impacting application four, but it's a problem with the data that got generated in application number one. So how do we know? Because no alerts came in application number one. it passed to application at two, three, and when it went to four, something's left, and then it's showing as red. So how do we know that the issue is there in application four? So once of a time, it will take operations to find out whether it's not a problem with application four, but it's going beyond and going at the top. So I would say the main frictions of ID operations right now are,

one is alert overloading, where there is too much data to process. The second is, it's a complex dependency stack, because each application, when it's talking to another application, it might be in a different stack. One is running on Java, one is running on Python, one is running on C, and then we have AWS coming in, then we have AI coming in. So all those things create a difference in the dependency stack. So we need to take care of that. The third thing that I would say is with CITD coming into picture, like continuous integration and deployment, applications are changing very fast. So we have multiple installs in a week. and multiple installs in a month. So the state of the application that was a week before, it may not be the same as a week after. And all these changes are coming with tons of data, which is going down the line. So new data is coming in. Is that getting into the traceability? Is that getting into the telemetry like the way it used to? That's another bigger question that we need to deal into. And the fourth thing that I would say is the SMEs,

the people who actually know the application, because they're engineers who know stuff on the application, and then there are engineers who know what they're doing, but they don't know how the application runs. Now, if there are people to guide those engineers to say, like, hey, the application runs like this, they have the knowledge. Like, they have the knowledge for the last 10 years or 20 years until the application is running. And this is particularly true for legacy applications. That knowledge may or may not be there because when people leave that knowledge is getting transferred over It not getting transferred over properly So I would say these are the four main friction areas right now in ID operations which kind of leads us to have difficulty in filtering the signals from the noise, as I would say. That is crazy. So I'm thinking about, especially your first and second point with the data overload also leads to application overload. and it sounds almost like a thick clutter. I imagine this very cluttered room full of things and you actually don't know where to touch and where to leave

and where to find things. And I'm thinking of two questions that I have on what you just said. First one is, on average, how many tools does a typical analyst have to open at once during a single incident? Is it between one and three or are we seeing it between three and seven, seven and ten? We normally try to keep it to a bare minimum of one and three, because if there are 10 applications and if we have a single tool for each of them, it becomes really hard for a single analyst to go through 10 tools at once, even how many monitors that person has. So we try to keep it to a bare minimum and use one tool to service as many applications as can be. So I would say like it's in between one and three. I mean, that gives you the capability of using that tool to marry the data across different applications. And that way it can find correlation in between the data that's there. Okay. I feel a bit better for the analysts than having to deal with it because I thought, oh, my goodness, if I had to use 10 tools just to diagnose a single incident or to just trace it,

I don't know if we'll have enough hours in the day. And that takes me to the second question that I have on what you just said, and that is time. Obviously, we have to investigate the incidents, decide what's real, decide what is not real. So how much time do we actually spend on that? Is it worth it to have all of these tools and all of these alerts in consideration with the time that we spend to try and figure out whether they're real? No, if you think about it, it's definitely worth it to invest in technology. I would say, like, as we grow with the data, as we grow with the business, we should have the proper tools in hand. And if you don't have, like you compare a company A and compare a company B, and let's say a company A, which is investing in technology, they're investing in cloud, they're investing in modernization, they're investing in AI, versus company B, which is not investing in any new technology. Then, you know, the rate at which company A will grow will be much higher than company B. And when I say grow, that, you know, includes everything, starting from the business, starting from, you know, sales, starting from marketing, starting from operational effectiveness, and starting from servicing the customers.

Because the end goal is to service the customer, like, no matter what business you are in. Like, I'm in the financial business, you know, for the last two decades, like, 20 plus years. And our job is to service the transaction. Like I'm into motion service processing for GFD Morgan, and we service big clients like Amazon and Walmart and Target and whatnot. And we have to be up all the time. So we have to be resilient. We have to be stable. We have to be high availability. And at the end of the day, what matters is a robust system which can give you the operational effectiveness to dive into issues real time and figure out which are the false positive and which one are the actual ones. So I would say it's very much needed for a company to have the proper insight, for the leaders to have the proper insight, to point the team into proper tools and the advancements that are going on to the system, enhance them, like harness them, and put it back into the application to get some tangible benefits in the future. Oh, what a real answer there. The next part of the conversation is where we touch on the fact that, okay, what if a team admits, listen, we are drowning, we know that we're drowning in alert noise.

The next question is then, okay, how do we fix it? And I think that is a difficult thing all on its own. And we hear a lot about AI correlation and automated alert handling and things like that. But hearing the terms and hearing what the capabilities are, but not so much about what that actually looks like on a day-to-day. So from your perspective, which AI-driven capabilities most reduce that noise and help the team to respond faster? That's, again, a great question. And right now, the issue, as we established, it's with the volume and the volume of the data and the speed with which the data is getting generated. And the only task that we need to do is, you know, segregate the signal from the noise because there's a lot of noise. 99% of the data is noise, which are kind of debug messages, which are info messages, but those are not the actual errors. So we'll have to, you know, use some tools to segregate those. And the best thing right now in the market is AI because we can leverage AI. AI has the speed. It has the mental power to do it.

However, it doesn't have a mind of its own. So what we need to give is a direction. We need to give it a purpose. We need to train the model. And that is how we can feed the data into AI and let it do wonderful things. So in terms of AI capabilities, I would say the first thing that hits my mind is, you know, use AI to do intelligent grouping of the events. So let's say there are 5,000 alerts that came in in a second or in a minute in the application. And it's not possible for, you know, to have employed 5,000 people to go through each of those alerts or even like 10,000 people to go through those alerts every minute. So feed them into AI and you give AI the proper tools, like, you know, historical data of the prior incidents that were there or maybe the deployment incidents that happened over time so that it knows what changed in the system, or even, you know, the graph in which the application increased and it enhanced. And let it do the grouping. So what it can do is it can go through those 5,000 messages and tell it like,

oh, okay, all of them are kind of variation of two incidents. So it can group it into two incidents, and that is what we need to look. So that's where 5,000 incidents are getting reduced to two, and then somebody can go and check these two. So that's where the human in the loop comes in. So AI does the smart work for us, and we take the judgment and decision. So that's the combination that you want. Now, the second thing that comes to my mind is, like, once we have the assessment and the analysis done, and we have kind of reduced the number of messages to unique groupings, then we can do a smart RCA. Because at every incident, we have to go through an RCA process, and that is where AI can come in handy very much. And when we do an RCA what we do is we look at the logs we look at the traceability we look at the different metrics from different places then we look at the deployment like whether the last thing created the issue or, you know, the issue we need to inform us at the end of time, what was the code structure at that point of time, what kind of data came in at that point of time, we look at the historical data.

So if we can have AI sit at top of all the data points, it can, and then it has access to the issue as well, it can go into each of those points and pull up specific information related to the issue. And then it can correlate each of them at different point of time saying like, hey, these are the changes that were done in each of these different points, which can lead to the issue. So it can help with a smart and intelligent RCA. Apart from that, I think it can also help with intelligent remediation in terms of like detection, diagnosis, you know, recommendation, automation, and, you know, self-review. but that's more of a future process like if we take time especially the automation and the self-healing part you cannot trust AI to solve everything for you right on the first day and so you will have to take baby steps one at a time you have to train the system like you're training a child to become a man so that's the thing that we need to do like hand hold the AI mature it into a system so that it's mature enough to take decisions for us

keep a human in the loop at every point of time and slowly leave the hand so that it can walk along its own path, it can take its own decision and think more of like a human. And if we're saying human in the loop, I'm wondering, obviously, a big part of this conversation is also going into trusting the automation that you eventually leave the hand of. So if we're keeping a human in the loop, does the responsibility still stay with the human if the AI gets it wrong during a real incident? Or if there's a real incident, AI gets it wrong, who's accountable? Yeah, I would say definitely somebody has to take the accountability, right? And we cannot say to the customer that, hey, AI did this for you. And we are taking hands off the matter and saying, like, AI did this for you. And the customer will come back. They will be really, really annoyed at that point, right? And they will not trust AI anymore. So the company will have to take accountability and responsibility, and we as humans have to take accountability and responsibility for all the decisions that AI takes for us, right? and that is why it should be not only AI.

I think AI, you can think of it as an assistant or a co-pilot and the main person driving the plane at that point of time will still have to be the human. And AI will be the co-pilot so that, you know, we can give most of the manual work to be done with AI, which the human cannot do in a short period of time because of the overwhelming data. And then take the decisions on your side and then it can be a cyclic loop as well. Like you take the decision, you send it to AI, ask AI to do specific stuff. AI does it for you. You validate it, and you send it down again, and it will refine it for you, deploy it into production, and then do the observability and traceability and matrix and everything, and feed that data again into AI. So it's kind of a tight look at different processes, and that is how we refine the system. Yeah, okay, that makes a lot of sense. And I think that's also a way to get our humans to trust it more, because they are hands-on working with it and seeing how it reacts every time that it keeps on feeding it the information And that's where I wanted to ask you if I'm going to tell a team, okay, we now have AI integrated into the workflows,

and certain workflows will be completely automated, and we need you to trust it, but you will still be held accountable if something goes wrong. How long, realistically, does it take for a team to adapt to that idea and to actually start to trust the automation? That's a pretty good question. In order for a team to trust the automation, they will have to be hands-on and be a part of the process of automating it. If I'm not a part of the team, like, so that's where the different groups come in, not only the development team, because it's mainly the development team who does the automation, then they hand it off to the product, right? And the product hands it off to the business or the sales, or it can go in front of the customer, right? But if the product does not know what the automation was or is not a part of the process, then they will not have any trust in the system. Now, as I said, like, we have to take baby steps at a time. And that goes through for the development team. So, you know, start from the beginnings. Try to figure out what are the different frictions that are there.

Like, what are the operational frictions in the system? And then you list out all the operational friction. And then you start prioritizing them by risks. Or you may prioritize by which one are easy to pick up. And then automate those pieces which are easy to pick up, which can go into production, like which can be deployed into production. and once you take those baby steps, that would give you enough confidence to go into larger steps. So let's say you have 10 operational frictions and you have already targeted two. Those are working well and then you can think about migrating the rest eight. So it will take time. It's not a thing that can be done in a matter of days or months. It will probably take years for a complex application to get automated end-to-end and take decisions amongst itself but it's a journey that we have to start with all the leaders and all the groups in it, starting from business, from product, from tech teams, so that everyone is at the same page on what the automation is happening and what AI is capable of doing and what are the risks associated with it.

And it's kind of a shared risk, I would say, but the first risk that has to be owned is not from the development team, but it's more of the product of the business because they are the ones who are selling it out to the customers. So they should be more aware of what automation is and, you know, what the risk associated can be. So it's very much a grab the team and say, listen, team, we're going on a road trip. Everyone get in the bus. And as we go down this road, we're going to get a flat tire at some point. We might run out of gas. We might get into a small little bumper bash. And we fix it as we go. And the next time the team has to go down the same road without navigation, they're already prepared for what could go wrong and how to handle it and trust the process and know what the end goal is. That's the image that we're picturing, right? Absolutely. And kind of replaying the same stuff over and over again and trying to list out all the possibilities, all the risks that are associated with it so that at one point of time you are so much familiar into the process that you can drive along the road with your eyes closed. And that's the ultimate goal.

That's exactly where we want to be. Last one before we say our goodbye today, but I know you're coming back for one more episode, so it's just goodbye for now. But last one for this episode say a leader is listening to this and they are sold on the idea that okay you know automation can be trusted There is a way to do this but they have no clue to actually start or where to start What practical steps can they take to move from reactive workflows to more intelligent operations? In that scenario, what I suggest is we have to start with knowing the system pretty well and then figuring out what the issues are with the system, ask questions on what the operational frictions are there, as I called out before, list them in order of priority, in order of risks, what things they are getting bitten by more, like whether it is the time that they're investing to figure out an issue, or whether it's that like logs are not properly evaluated, whether the same incident is getting repeated over and over again,

because that's also indicative of a failure. So those are the things that needs to be figured out and listed somewhere before, you know, the leaders take the decision of, you know, addressing things one by one. If they try to do all things at once, it will not be possible. So we need to have a game plan. We need to start with baby steps at a time. But we need to know what the system is doing, what is wrong in the system, and then have a kind of a project plan for the bigger time to have two or three things done, one thing at a time, and then integrate the whole thing from bottom up. So that would be the first thing, like figure out the operational frictions different frictions and then segregate the ones and come out with a game plan. The second thing that they can do is they will have to build the data architecture so that can be fed into AI. And AI works effectively and best when the data is structured. So if they produce, like if their applications produce unstructured data, like let's say an application is running and a submodule is running in it, and the submodule produces the data, but it doesn't have the timestamp.

It doesn't have the name of the submodule. it doesn't have the name of the application, and it doesn't have any relation to what the data is about, like what the error is about. When you feed it to AI, it will not have a freaking clue of where it's coming from, where it's coming from application A, where it's coming from application B, or even what the submodule name is. So those are the things that needs to be set at the basic ground level. And once we do that, we have to change the system to have a structured data that can be fed into AI. So that's the second thing. And the third thing I would say is we have to have guardrails and governance in it, the data governance as well as the AI governance, because AI governance is entirely different from data governance. Where data governance stops, AI governance starts from there, because AI governance hopes or rather it assumes that the data governance is already in place and that is on top of it you can do AI governance. So we need to have proper guardrails. We have to have proper boundaries set for each of the applications or each of the groups in place.

so that they know that we are not going beyond what needs to be, you know, beyond our scope of where we are going. And the last thing I would say is measurability. So anything that we put in, we'll have to measure properly, whether it's bringing any tangible benefits, whether the company is profiting or not, and what is the ROI, the KPI, and those kind of stuff. So that's the other thing. So, I mean, that is how we should operate. Like if somebody is thinking of going into it, they have a big and complex system, They don't know where to start. So start from the system. Like ask questions what the issues are. Get your SMEs involved. Get the folks involved. And come up with a game plan, a project plan. Start executing small step at a time. Maybe even do a POC. Integrate the POC and see how it's working. And then you take it into production. But make the production ready. Because many of the companies are working on legacy. They don't have production which can support AI. So in order to have a production readiness, that's also another task. So you should do it before you think of bringing AI into production. because it might be possible that you can feed the data in lower environments,

but because your production is not set, you cannot feed the data into AI. So the whole architecture needs to be looked at, and that has to be fixed before you can make AI workable. Great. I'm so comprehensive, such a comprehensive starting kit of information that you just shared with our audience. I think that's going to be extremely valuable. To wrap up our discussion today, I think what our listeners really need to understand is that alert fatigue is not solved by adding more tools. In fact, it's quite the opposite. It's kind of solved by deciding which of the noise and the alerts actually deserves human attention. And then for the second part of our conversation, I think the theme was very much AI earns trust in the operation very gradually. It's not an overnight thing. We will have small wins, but we will have to get the team involved in those wins and also involved in the failures and keep on feeding it the same information until you are completely satisfied that the return information is exactly what you need it to be for your team. And then, like you just said, for our leaders, start with the most noisy workflow almost

and not a company-wide overall because, like you just said, if you're going to try and do everything all at once, everything's going to fail all at once. So it was great having you on the show, great conversation. I'm very excited for the next time I get you to be on the show, which is going to be very soon. Thank you so much. Thank you. Thanks for having me again. Wrapping up today's episode, let's look at the three key takeaways from our conversation with Riva Joti. First, alert fatigue isn't solved by adding more to it. It's solved by deciding which signals deserve human attention in the first place. Second, AI earns trust in operations gradually through small, validated wins rather than full automation on day one. And finally, leaders should start with their noisiest workflow, not a company-wide robot, to build the data foundation and confidence automation required.

In this episode, we covered how AI can help teams cut through alert noise. To go deeper on this topic and learn how to structure landing pages for higher conversion and how to use self-qualification systems to prioritize high-intended leads, Download our free PDF report B2B AI lead generation guide at emerge.com slash AIG1. That's E-M-E-R-J dot com slash A-I-G-1 to download your copy. For further executive level analysis and to join our network of leaders delivering workflow impact with AI, visit emerge.com. On behalf of the team at Emerge, we'll see you on the next episode. Bye.

番組の概要欄(原文)

Alert volume inside IT operations has outpaced manual triage, leaving real incidents buried in noise. In this episode, Dhrubojyoti Das Deb, SVP of Global IT at JPMorgan Chase & Co., explores how AI-driven correlation and intelligent grouping can turn thousands of alerts into a handful of actionable incidents. The conversation covers where accountability sits when AI assists in triage, how teams build trust in automation incrementally, and where leaders should start when moving from reactive firefighting to intelligent operations.   This episode is sponsored by BigPanda.   Learn how to build high-converting landing pages and use self-qualification systems to identify and prioritize high-intent leads. Download our free B2B AI Lead Generation Guide at emerj.com/aig1

X でシェア