← Back
AI Engineer October 6, 2026 21m

The 6 Pillars of an Agentic Harness for Production — Varun Krovvidi, Resolve AI

Read full transcript 17 segments
  1. Can everyone hear me well Can everyone hear me well ? Perfectly. ? Perfectly. ? Perfectly. Thank you very much for coming to Thank you very much for coming to Thank you very much for coming to this session. My name is this session. My name is this session. My name is Varun. I am part of the Varun. I am part of the Varun. I am part of the Resolve AI team. We are Resolve AI team. We are Resolve AI team. We are building AI to building AI to building AI to fight fight fight fraud. Actually, fraud. Actually, fraud. Actually, we build agents we build agents we build agents to run and to run and to run and patch your patch your software. So in software. So in this session, we'll this session, we'll this session, we'll look at the six look at the six look at the six pillars of an agent pillars of an agent pillars of an agent system that you'll system that you'll system that you'll need to need to need to manage and repair manage and repair manage and repair your software. Before we your software. Before we your software. Before we begin, raise your begin, raise your begin, raise your hands so I can understand hands so I can understand hands so I can understand who is in who is in who is in the audience today. I assume the audience today. I assume the audience today. I assume most of you are most of you are engineers. How many of engineers. How many of engineers. How many of you have previously worked you have previously worked on-call? Raise your hands. on-call? Raise your hands. Perfectly. Have Perfectly. Have Perfectly. Have you tried using you tried using you tried using AI to help AI to help AI to help you on you on you on duty? Raise your duty? Raise your duty? Raise your hands. Perfectly. Those who hands. Perfectly. Those who hands. Perfectly. Those who raised their hands, raised their hands, raised their hands, can you say that can you say that can you say that 90% of your work 90% of your work 90% of your work while on duty while on duty while on duty is done by AI? Can you is done by AI? Can you say that with confidence? Perfectly. say that with confidence? Perfectly. Yes, that's exactly what we Yes, that's exactly what we Yes, that's exactly what we understood when we understood when we understood when we started the company. This is a started the company. This is a started the company. This is a very difficult very difficult very difficult task. So in this task. So in this task. So in this session, I want session, I want session, I want to share our to share our to share our journey, the evolution journey, the evolution journey, the evolution we've gone through, and we've gone through, and we've gone through, and our lessons learned in our lessons learned in our lessons learned in creating such an creating such an creating such an agent system.

  2. agent system. Before I get to Before I get to the point, I want the point, I want the point, I want to outline the context of what to outline the context of what to outline the context of what we have we have we have today. The first wave of today. The first wave of today. The first wave of AI was focused AI was focused AI was focused specifically on specifically on specifically on programming. programming. programming. Honestly, I think the Honestly, I think the way we write way we write code code code has changed radically for most of us. I also has changed radically for most of us. I also has changed radically for most of us. I also wanted to explain why this wanted to explain why this wanted to explain why this happened. First, happened. First, happened. First, the code itself is the code itself is the code itself is self-documenting. self-documenting. AI is easy to analyze AI is easy to analyze and understand to and understand to and understand to help you take the help you take the help you take the next step. next step. Second, the code is Second, the code is Second, the code is inherently very modular. inherently very modular. inherently very modular. So it's easy for AI So it's easy for AI So it's easy for AI to break it down into to break it down into to break it down into parts, move parts, move parts, move to the next stage, to the next stage, to the next stage, and help and help and help you expand and you expand and you expand and build on it build on it build on it . Third, and . Third, and . Third, and most importantly—code is a most importantly—code is a most importantly—code is a single-domain sphere. single-domain sphere. single-domain sphere. You don't need You don't need You don't need to teach AI many to teach AI many to teach AI many disciplines so that it disciplines so that it disciplines so that it can reason within can reason within can reason within one. It is easy to one. It is easy to one. It is easy to navigate the navigate the navigate the next steps, which is next steps, which is next steps, which is why we have why we have why we have such clear such clear such clear performance indicators. You performance indicators. You performance indicators. You can objectively can objectively can objectively assess how assess how assess how well AI solves this well AI solves this well AI solves this task for you. But task for you. But task for you. But in reality, as in reality, as in reality, as engineers, we engineers, we engineers, we spend 70% of our time doing something spend 70% of our time doing something spend 70% of our time doing something else: not creating else: not creating else: not creating new software, but working new software, but working new software, but working with productive with productive with productive systems—running and systems—running and systems—running and fixing programs.

  3. fixing programs. Um, and that covers a Um, and that covers a wide range of wide range of wide range of issues, right? Such issues, right? Such as operating and as operating and as operating and launching launching launching software. software. software. Generally, all of this can be Generally, all of this can be Generally, all of this can be divided into three divided into three divided into three broad categories. broad categories. broad categories. For example, think For example, think For example, think about the analogy with about the analogy with about the analogy with healthcare. This is the healthcare. This is the healthcare. This is the simplest way, which simplest way, which simplest way, which I like to I like to I like to think of too. The first is think of too. The first is think of too. The first is regular duty regular duty regular duty and and and maintenance. maintenance. maintenance. Think of it as the usual Think of it as the usual Think of it as the usual little problems that little problems that little problems that keep coming up. keep coming up. keep coming up. Maybe you need Maybe you need Maybe you need Tylenol for some Tylenol for some Tylenol for some pain. Maybe you pain. Maybe you pain. Maybe you need an Advil. So, need an Advil. So, need an Advil. So, these are typical these are typical these are typical on-duty notifications that you have to on-duty notifications that you have to work with. The second work with. The second category is category is category is incidents that incidents that incidents that require full require full require full mobilization. This is mobilization. This is mobilization. This is equivalent to a equivalent to a equivalent to a surgical operation surgical operation . Systems are shutting down . Systems are shutting down . It needs the attention of . It needs the attention of . It needs the attention of everyone everyone everyone involved or involved or involved or responsible to responsible to responsible to come in and come in and come in and fix things. And the third fix things. And the third fix things. And the third category is like your category is like your category is like your daily vitamins. daily vitamins. daily vitamins. So, you need to So, you need to So, you need to do a lot of things do a lot of things do a lot of things on an ongoing basis on an ongoing basis on an ongoing basis to keep to keep to keep your system running smoothly your system running smoothly your system running smoothly . This could . This could . This could be things like be things like infrastructure review, infrastructure review, cost analysis, or cost analysis, or cost analysis, or engineering work on a engineering work on a engineering work on a platform for platform for platform for scaling. These scaling. These scaling. These things are fundamentally things are fundamentally things are fundamentally complex because they're complex because they're complex because they're not exactly like code. They not exactly like code. They not exactly like code. They span a multitude of span a multitude of span a multitude of different domains and different domains and different domains and tools, and these are the tools, and these are the tools, and these are the consequences we consequences we consequences we are starting to see are starting to see are starting to see as the initial euphoria as the initial euphoria as the initial euphoria of AI wears off. And that's exactly what of AI wears off. And that's exactly what of AI wears off. And that's exactly what you're seeing in you're seeing in you're seeing in the news right now. There is more the news right now. There is more the news right now. There is more talk about talk about talk about the number of problems the number of problems the number of problems that that that generated AI code can create in generated AI code can create in generated AI code can create in production, and production, and production, and discussion about which discussion about which discussion about which frameworks are best

  4. frameworks are best frameworks are best suited to suited to suited to solve these solve these solve these problems. Second, the era of problems. Second, the era of problems. Second, the era of limitless AI limitless AI limitless AI has come to an end. That is, has come to an end. That is, has come to an end. That is, unlimited AI unlimited AI unlimited AI no longer exists. More and more no longer exists. More and more no longer exists. More and more often you can hear often you can hear often you can hear talk about tokens, talk about tokens, talk about tokens, token optimization, token optimization, token optimization, and and and token efficiency. How can you token efficiency. How can you token efficiency. How can you improve these improve these improve these architectures, those architectures, those architectures, those rigorous engineering rigorous engineering rigorous engineering decisions that are required decisions that are required decisions that are required to use AI effectively and to use AI effectively and to use AI effectively and accurately accurately accurately ? And the third is a ? And the third is a ? And the third is a conversation that conversation that conversation that was probably made popular was probably made popular was probably made popular by Satya by Satya by Satya Nadella, the Nadella, the CEO of Microsoft, where CEO of Microsoft, where he talked about he talked about he talked about the architecture the architecture the architecture needed on top of needed on top of needed on top of conventional models conventional models conventional models for any highly for any highly for any highly specialized specialized specialized AI to work. So there's a broad AI to work. So there's a broad AI to work. So there's a broad classification, right? classification, right? classification, right? When you look at a When you look at a When you look at a generation task, generation task, generation task, you don't have a you don't have a you don't have a specific specific specific outcome or outcome or outcome or expected outcome expected outcome . You don't want to and don't . You don't want to and don't . You don't want to and don't tell AI exactly what an tell AI exactly what an image, blog, or image, blog, or website should look like. This is where the website should look like. This is where the website should look like. This is where the euphoria comes in. But euphoria comes in. But euphoria comes in. But when you think when you think when you think about the flip side of about the flip side of about the flip side of the problem, where there is one the problem, where there is one the problem, where there is one correct answer correct answer correct answer that you want to that you want to that you want to get from the AI, that's get from the AI, that's get from the AI, that's when the when the when the frustration begins.

  5. frustration begins. frustration begins. Most people Most people Most people call this the " call this the " last mile," but last mile," but last mile," but it's the longest mile it's the longest mile it's the longest mile you'll have to you'll have to you'll have to travel with AI. That's travel with AI. That's travel with AI. That's why you need these why you need these why you need these architectures to architectures to architectures to focus on a focus on a focus on a specific problem. specific problem. In addition, In addition, the operation of the operation of the operation of production systems production systems is also a problem for is also a problem for is also a problem for many many many users and users and users and systems. You are not doing systems. You are not doing systems. You are not doing this alone. You this alone. You this alone. You need to involve need to involve need to involve experts from different experts from different experts from different teams. Over time, each teams. Over time, each teams. Over time, each of us of us of us specialized in specialized in specialized in different categories of different categories of different categories of engineering. We engineering. We engineering. We work in different work in different work in different teams to teams to teams to discuss and discuss and discuss and solve any solve any solve any problems problems problems using AI. Sorry using AI. Sorry , I mean in , I mean in , I mean in your production your production your production systems. In general, systems. In general, systems. In general, our company Resolve AI our company Resolve AI our company Resolve AI started its started its started its activities with one activities with one activities with one thesis. We truly believe thesis. We truly believe thesis. We truly believe that all of this work that all of this work that all of this work you do in you do in you do in production systems, production systems, production systems, like like like bug fixing or day-to-day bug fixing or day-to-day bug fixing or day-to-day software maintenance, software maintenance, software maintenance, should mostly be should mostly be should mostly be done by agents. That's done by agents. That's done by agents. That's one of the reasons I one of the reasons I one of the reasons I asked: can asked: can asked: can you confidently say you confidently say you confidently say that 90% or more of your work that 90% or more of your work that 90% or more of your work is done by is done by is done by agents? Engineers agents? Engineers agents? Engineers simply have to simply have to simply have to manage these manage these manage these agents, which take care of agents, which take care of agents, which take care of patching patching patching and running your and running your software. That's why software. That's why we built Resolve we built Resolve we built Resolve AI as three categories of AI as three categories of AI as three categories of agents. There are agents. There are agents. There are agents on duty to agents on duty to agents on duty to resolve resolve resolve software issues that arise software issues that arise software issues that arise daily. And there are incident agents to daily. And there are incident agents to help you navigate help you navigate incident channels and incident channels and incident channels and guide you from a guide you from a guide you from a complex incident complex incident complex incident to root cause and to root cause and to root cause and resolution. And finally, resolution. And finally, resolution. And finally, there are background there are background there are background agents that agents that agents that will help you will help you will help you with production with production with production tasks. Whether it's tasks. Whether it's deployment monitoring, or any analysis you

  6. any analysis you any analysis you need to perform, or need to perform, or need to perform, or conditions you conditions you conditions you trigger when trigger when trigger when you want to conduct a you want to conduct a you want to conduct a specific specific specific investigation. At investigation. At investigation. At the heart of all this the heart of all this the heart of all this is the same is the same Resolve AI agent architecture. This is what we will talk about today. And about what it talk about today. And about what it takes to takes to takes to get there and get there and get there and build it. So, build it. So, build it. So, let's start with a question. let's start with a question. let's start with a question. Why is this even Why is this even Why is this even necessary, right? necessary, right? necessary, right? Why can't we Why can't we Why can't we just point the just point the just point the largest model at largest model at largest model at production data? production data? production data? Models, of course, Models, of course, Models, of course, are getting better. Their are getting better. Their reasoning ability to reasoning ability to fix fix fix software problems is improving. software problems is improving. software problems is improving. It really looks It really looks It really looks impressive when you impressive when you impressive when you focus on a focus on a focus on a specific case and specific case and specific case and build your system build your system build your system on top of it, it on top of it, it on top of it, it makes for a very good makes for a very good makes for a very good demonstration. But demonstration. But demonstration. But gradually, as you gradually, as you gradually, as you start start start to scale between to scale between to scale between different different different usage scenarios and usage scenarios and usage scenarios and teams, this is where the very familiar cracks in the teams, this is where the very familiar cracks in the teams, this is where the very familiar cracks in the system start system start system start to appear to appear to appear . So . So , what are the cracks that , what are the cracks that , what are the cracks that we've seen as you we've seen as you we've seen as you start start start to scale a product to scale a product to scale a product along this path? The first thing along this path? The first thing along this path? The first thing you'll notice is that you'll notice is that you'll notice is that the models certainly the models certainly the models certainly have a binding bias have a binding bias have a binding bias . This is a very . This is a very . This is a very significant thing that you have to significant thing that you have to deal with deal with directly in directly in directly in your work your work your work tools. They are tools. They are tools. They are designed to give designed to give designed to give you a consistent you a consistent you a consistent response, and above response, and above response, and above all, as all, as all, as new models with better new models with better reasoning abilities emerge, you reasoning abilities emerge, you need to keep up with need to keep up with need to keep up with that pace as well. You also that pace as well. You also that pace as well. You also need to constantly need to constantly need to constantly change—first change—first change—first evaluating which model is evaluating which model is evaluating which model is right for right for right for the task, then the task, then the task, then updating it to one updating it to one updating it to one that better fits that better fits that better fits your scenario.

  7. your scenario. your scenario. And the second is a somewhat And the second is a somewhat And the second is a somewhat underestimated point. underestimated point. underestimated point. We always associate a We always associate a We always associate a model with a model with a model with a workflow or workflow or workflow or outcome. But outcome. But outcome. But each workflow each workflow each workflow has hundreds of different has hundreds of different has hundreds of different types of tasks. So, there are types of tasks. So, there are reasoning tasks, there are reasoning tasks, there are deterministic deterministic deterministic tasks that tasks that tasks that you go through. There are you go through. There are you go through. There are things like general things like general things like general visual reasoning visual reasoning visual reasoning or or or image-based reasoning, image-based reasoning, image-based reasoning, reasoning through SQL, reasoning through SQL, reasoning through SQL, or things like logs or things like logs . Now, all the advanced . Now, all the advanced . Now, all the advanced models that you see models that you see models that you see around, each one is around, each one is around, each one is good at a certain good at a certain good at a certain task or a certain task or a certain task or a certain kind of task. Therefore, the kind of task. Therefore, the kind of task. Therefore, the orchestration mechanism orchestration mechanism you need to you need to you need to imagine to imagine to imagine to switch between these switch between these switch between these models for a models for a models for a workflow workflow must first keep up must first keep up must first keep up with the evolution of the models, with the evolution of the models, with the evolution of the models, and secondly, must and secondly, must and secondly, must be able to select the be able to select the be able to select the best model for the best model for the best model for the best task. The best task. The second failure mode is , of course, context , of course, context , of course, context windows. I know that windows. I know that windows. I know that context windows context windows context windows are expanding, but are expanding, but are expanding, but context windows are context windows are context windows are not the only problem for not the only problem for not the only problem for solving such a solving such a solving such a complex task. complex task. The second part The second part is that is that is that if you provide if you provide if you provide too much too much too much context, the models context, the models context, the models start to " start to " over-examine".

  8. over-examine". over-examine". They will start They will start They will start making up hypotheses making up hypotheses making up hypotheses that don't even exist, that don't even exist, that don't even exist, or sometimes they may or sometimes they may or sometimes they may not even lead not even lead not even lead you to the right you to the right you to the right answer. Now, answer. Now, answer. Now, if you provide very if you provide very if you provide very limited context, limited context, limited context, of course they are " of course they are " under-researching." under-researching." under-researching." They won't be able to They won't be able to They won't be able to get to the bottom of it, get to the bottom of it, get to the bottom of it, they won't have visibility they won't have visibility they won't have visibility into what into what into what paths exist or what other paths exist or what other paths exist or what other potential hypotheses potential hypotheses potential hypotheses might exist to might exist to might exist to solve the problem. solve the problem. solve the problem. Especially for Especially for Especially for production production production incidents, this becomes a incidents, this becomes a incidents, this becomes a huge problem huge problem because the telemetry is because the telemetry is because the telemetry is literally literally literally endless. You endless. You endless. You can create can create can create as many log lines as many log lines as many log lines as you want and as you want and as you want and as many metrics as many metrics as many metrics as you need. as you need. as you need. And now, when you And now, when you And now, when you start to combine start to combine start to combine this with your code and this with your code and this with your code and infrastructure, that's when the infrastructure, that's when the real cracks start to appear in the system. Third, and my Third, and my favorite, is the favorite, is the favorite, is the definition of causal definition of causal reasoning. So, reasoning. So, models exist to models exist to models exist to please please please you. I mean, we've you. I mean, we've you. I mean, we've all had all had all had those moments those moments those moments where you push a where you push a where you push a model too model too model too hard and it hard and it hard and it starts starts starts to agree with you to agree with you to agree with you in every direction in every direction . So, models are . So, models are . So, models are specifically specifically specifically designed to designed to designed to give you coherent give you coherent give you coherent answers, not answers, not answers, not cause-and-effect cause-and-effect cause-and-effect relationships. But in relationships. But in relationships. But in industrial industrial industrial incidents, the only thing incidents, the only thing incidents, the only thing you look for is a you look for is a you look for is a causal causal causal chain of evidence, like a chain of evidence, like a chain of evidence, like a detective. What steps detective. What steps detective. What steps led to a led to a led to a specific specific specific incident so you incident so you incident so you can actually can actually can actually fix the right fix the right fix the right problem, not just problem, not just problem, not just create a patch. So, create a patch. So, create a patch. So, especially when you're especially when you're especially when you're solving a complex solving a complex solving a complex problem like a problem like a problem like a production production production justification, you justification, you justification, you need need need causal thinking causal thinking causal thinking built into the built into the built into the model and system, not model and system, not model and system, not just coherent just coherent just coherent answers. And next answers. And next , of course, are , of course, are , of course, are precautions. I won't precautions. I won't dwell too much on this dwell too much on this topic, but we've all topic, but we've all topic, but we've all read news stories about read news stories about read news stories about artificial intelligence artificial intelligence artificial intelligence deleting a certain

  9. deleting a certain deleting a certain file system or file system or file system or database. So, for database. So, for database. So, for AI, this is also a function, AI, this is also a function, AI, this is also a function, right? Not a right? Not a right? Not a mistake. The AI ​​could mistake. The AI ​​could mistake. The AI ​​could decide that decide that decide that the cleanest possible the cleanest possible the cleanest possible fix is ​​to fix is ​​to fix is ​​to simply remove simply remove simply remove that particular that particular that particular piece of code. You piece of code. You piece of code. You can't blame can't blame can't blame him for that. You just him for that. You just him for that. You just need to build need to build need to build safeguards safeguards safeguards around the system to keep around the system to keep around the system to keep it working it working it working properly. What is the properly. What is the properly. What is the minimum level of minimum level of minimum level of access required access required access required for AI to operate in for AI to operate in any given any given any given scenario? And lastly, scenario? And lastly, learning cycles. learning cycles. Any system you Any system you Any system you build should be build should be build should be scalable for scalable for scalable for your team or the your team or the your team or the entire organization at entire organization at entire organization at best. best. best. Anytime you Anytime you Anytime you use AI use AI use AI for production for production for production incidents, as I incidents, as I incidents, as I said at the beginning, it's a said at the beginning, it's a said at the beginning, it's a multi-user multi-user multi-user problem. You have problem. You have problem. You have many different teams, many different teams, many different teams, such as SRE, platform such as SRE, platform such as SRE, platform teams, backend teams, backend engineers, or engineers, or engineers, or service engineers. service engineers. service engineers. All these people must All these people must All these people must be involved. And be involved. And be involved. And most importantly, all these most importantly, all these most importantly, all these people must have the people must have the people must have the same context. same context. Now, when you Now, when you add add add temporal discussions to this, your temporal discussions to this, your temporal discussions to this, your AI system needs to AI system needs to AI system needs to have the same have the same have the same context from context from context from the previous the previous the previous investigation in the investigation in the investigation in the hundredth investigation hundredth investigation , otherwise you're , otherwise you're , otherwise you're starting from scratch every time.

  10. starting from scratch every time. starting from scratch every time. So these are the So these are the So these are the failure modes where we saw failure modes where we saw cracks appear. That's why we cracks appear. That's why we created an agent created an agent created an agent architecture that architecture that architecture that answers these very answers these very answers these very questions. So, there are questions. So, there are questions. So, there are six basic six basic six basic parts to any parts to any parts to any AI architecture that AI architecture that AI architecture that needs to work in a needs to work in a needs to work in a specific specific specific domain. The reason domain. The reason domain. The reason I mention the I mention the I mention the specific domain is specific domain is because models because models are designed for are designed for are designed for general thinking. general thinking. general thinking. Now, when you want to Now, when you want to Now, when you want to direct these models direct these models direct these models to a specific to a specific to a specific answer, you answer, you answer, you need to take care of need to take care of need to take care of six different pillars. six different pillars. six different pillars. The first is The first is The first is model orchestration. model orchestration. model orchestration. As I mentioned, As I mentioned, As I mentioned, model orchestration model orchestration model orchestration consists of two consists of two consists of two levels. First: how do you levels. First: how do you levels. First: how do you keep up with keep up with keep up with constantly updating constantly updating constantly updating models and models and models and using the using the using the latest ones? And latest ones? And latest ones? And secondly, how secondly, how secondly, how to choose the best to choose the best to choose the best model for a model for a model for a specific task specific task ? Maybe Gemini for ? Maybe Gemini for ? Maybe Gemini for image analysis? image analysis? image analysis? Or OpenAI for Or OpenAI for Or OpenAI for deterministic deterministic deterministic steps? Or maybe steps? Or maybe steps? Or maybe Claude for open Claude for open Claude for open research? You research? You research? You must constantly must constantly must constantly evaluate which model evaluate which model evaluate which model is best for the is best for the is best for the task at hand task at hand task at hand . . . So, this is a mechanism So, this is a mechanism So, this is a mechanism that we created from the that we created from the that we created from the very beginning. The second very beginning. The second is context is context is context engineering. Nowadays, engineering. Nowadays, engineering. Nowadays, context engineering is context engineering is context engineering is often reduced to just a often reduced to just a often reduced to just a database or database or execution solution. Oh, you execution solution. Oh, you 're using graph RAG? Are 're using graph RAG? Are 're using graph RAG? Are you using a you using a you using a knowledge graph? Are you knowledge graph? Are you knowledge graph? Are you using X, Y, and using X, Y, and using X, Y, and Z? These are all Z? These are all Z? These are all implementation details. But the implementation details. But the implementation details. But the main question is exactly how main question is exactly how main question is exactly how much much much context does context does context does AI need to solve this AI need to solve this AI need to solve this problem. And problem. And problem. And most often it is a most often it is a most often it is a combination of different combination of different combination of different methods. You may methods. You may methods. You may need to need to need to use a

  11. use a use a graph RAG type solution graph RAG type solution graph RAG type solution so that the AI ​​can start so that the AI ​​can start so that the AI ​​can start investigating. And when investigating. And when investigating. And when you move on to the you move on to the you move on to the next steps, you next steps, you next steps, you need to define need to define need to define very precise very precise very precise tool calls so that you don't tool calls so that you don't tool calls so that you don't waste tokens waste tokens waste tokens too quickly too quickly too quickly when executing when executing when executing queries to logs, queries to logs, queries to logs, metrics, dashboards, or metrics, dashboards, or metrics, dashboards, or code, etc. The third is code, etc. The third is code, etc. The third is cause-and-effect cause-and-effect cause-and-effect reasoning. As I reasoning. As I reasoning. As I said, this is one of the said, this is one of the said, this is one of the fundamental fundamental fundamental principles we have principles we have principles we have laid out in Resolve AI: laid out in Resolve AI: laid out in Resolve AI: root cause is always root cause is always root cause is always based on a based on a based on a chain of evidence. What chain of evidence. What chain of evidence. What exactly were the steps that led exactly were the steps that led exactly were the steps that led to this particular to this particular to this particular problem? If we problem? If we problem? If we cannot establish cannot establish cannot establish this chain, we this chain, we this chain, we should provide should provide should provide information with a low information with a low information with a low level of confidence level of confidence level of confidence and direct and direct and direct the user in another the user in another the user in another direction. This applies to direction. This applies to direction. This applies to any any any specialized AI, specialized AI, specialized AI, including Resolve AI: we including Resolve AI: we including Resolve AI: we redirect you redirect you redirect you if we don't see data if we don't see data if we don't see data after a certain point. The after a certain point. The next thing, as I next thing, as I said, is guided said, is guided said, is guided actions. It's pretty simple.

  12. actions. It's pretty simple. It depends on It depends on each team, each team, each team, each organization, each organization, each organization, access levels, access levels, access levels, restrictions, or the industry restrictions, or the industry restrictions, or the industry you work in. you work in. What specific What specific safeguards need to safeguards need to safeguards need to be identified for be identified for be identified for AI to work? Is this reading? Is this a AI to work? Is this reading? Is this a AI to work? Is this reading? Is this a recording? And if it is a record, recording? And if it is a record, recording? And if it is a record, what conditions what conditions what conditions exist for it, etc. And exist for it, etc. And exist for it, etc. And the last thing is the the last thing is the the last thing is the training system. Think training system. Think training system. Think about it this way. Every about it this way. Every about it this way. Every interaction you have with an interaction you have with an interaction you have with an AI system is an opportunity for AI system is an opportunity for AI system is an opportunity for learning. Therefore, the learning. Therefore, the learning. Therefore, the artificial artificial artificial intelligence system must intelligence system must intelligence system must learn not only from the learn not only from the investigation process itself, but also from investigation process itself, but also from how the user how the user how the user interacts with it. interacts with it. interacts with it. For example, am I giving For example, am I giving For example, am I giving you positive you positive you positive reinforcement? So, reinforcement? So, reinforcement? So, does this translate does this translate does this translate into a positive assessment? into a positive assessment? into a positive assessment? Am I giving you Am I giving you Am I giving you negative negative negative reinforcement? Am reinforcement? Am reinforcement? Am I directing you or I directing you or I directing you or directing the directing the directing the investigation in a investigation in a investigation in a certain direction? These are certain direction? These are certain direction? These are different types different types different types of evaluation. This of evaluation. This of evaluation. This brings us to the brings us to the brings us to the final part, where final part, where final part, where evaluation is also a evaluation is also a evaluation is also a very important step. very important step. very important step. This is the starting and This is the starting and This is the starting and ending point for ending point for ending point for any any any agent architecture. agent architecture. agent architecture. We look at We look at We look at assessment at five assessment at five assessment at five different levels, as I different levels, as I different levels, as I mentioned. What mentioned. What mentioned. What positive positive positive reinforcement reinforcement reinforcement can you provide? What can you provide? What can you provide? What negative negative negative reinforcement reinforcement reinforcement can you provide, and can you provide, and can you provide, and can we trace can we trace can we trace your path to how you your path to how you your path to how you arrived at the decision?

  13. arrived at the decision? arrived at the decision? Can we evaluate Can we evaluate Can we evaluate you as an engineer you as an engineer you as an engineer based on this? How would the based on this? How would the based on this? How would the best engineer best engineer best engineer conduct this conduct this conduct this investigation, and how do investigation, and how do investigation, and how do you compare to him you compare to him ? And on top of that, how do you ? And on top of that, how do you ? And on top of that, how do you calibrate yourself as an calibrate yourself as an calibrate yourself as an artificial intelligence? How do artificial intelligence? How do artificial intelligence? How do you determine this you determine this you determine this level of confidence? These are all level of confidence? These are all level of confidence? These are all systematic systematic systematic assessments that you assessments that you assessments that you build into one build into one build into one platform, and that you platform, and that you platform, and that you continually need to continually need to continually need to apply to the apply to the apply to the architecture as a architecture as a architecture as a new new new model emerges, a new model emerges, a new model emerges, a new use case, or the use case, or the architecture changes. Steeply. architecture changes. Steeply. Enough talk. Enough talk. Enough talk. Let me Let me Let me show you how this show you how this show you how this works in Resolve in works in Resolve in works in Resolve in practice. As I practice. As I practice. As I said, our platform said, our platform said, our platform is designed for exactly is designed for exactly is designed for exactly this. So, we this. So, we this. So, we work with two work with two work with two categories of agents. categories of agents. categories of agents. One is a duty One is a duty One is a duty agent, another is an agent, another is an agent, another is an incident agent, and, of course, a incident agent, and, of course, a incident agent, and, of course, a background agent. background agent. background agent. Since most Since most Since most of you here are engineers, of you here are engineers, of you here are engineers, you may be familiar with you may be familiar with you may be familiar with this system. Of course this system. Of course , most of your , most of your , most of your telemetry or telemetry or telemetry or monitoring monitoring monitoring notifications go notifications go notifications go to to to collaboration tools like Slack, collaboration tools like Slack, collaboration tools like Slack, Microsoft Teams, or whatever else Microsoft Teams, or whatever else Microsoft Teams, or whatever else you you you use. In this use. In this use. In this case, we see case, we see case, we see Grafana sending Grafana sending Grafana sending notifications to our notifications to our notifications to our shared channel in Slack.

  14. shared channel in Slack. shared channel in Slack. Resolve automatically Resolve automatically Resolve automatically picks up on these picks up on these picks up on these alerts and alerts and alerts and begins an begins an begins an investigation. You investigation. You investigation. You see the system see the system see the system reporting a reporting a reporting a log-scale error log-scale error log-scale error and starting to indicate and starting to indicate and starting to indicate the root cause of the the root cause of the the root cause of the detected detected detected alert, as well as alert, as well as alert, as well as other factors or other factors or other factors or theories it has theories it has theories it has ruled out. She ruled out. She ruled out. She provided a concise provided a concise provided a concise root cause, evidence, root cause, evidence, root cause, evidence, and recommendations. But and recommendations. But and recommendations. But let's delve let's delve let's delve into the details. As I into the details. As I into the details. As I said, in most said, in most said, in most cases it should be a cases it should be a cases it should be a collaborative effort. So, collaborative effort. So, collaborative effort. So, let's move on to the let's move on to the let's move on to the interface and interface and interface and see what's see what's see what's going on behind going on behind going on behind the scenes. The work the scenes. The work the scenes. The work begins with Resolve AI begins with Resolve AI begins with Resolve AI when it reads when it reads when it reads a notification that a notification that a notification that comes in as a separate comes in as a separate comes in as a separate request. It takes that request. It takes that request. It takes that request and launches request and launches request and launches a number of different agents, a number of different agents, a number of different agents, actually two actually two actually two categories of agents, categories of agents, categories of agents, to get the job done. The to get the job done. The to get the job done. The first category of first category of first category of agents are, in essence, agents are, in essence, agents are, in essence, investigators. They are trained investigators. They are trained investigators. They are trained to conduct to conduct to conduct investigations just as an investigations just as an engineer would. First, engineer would. First, find out what the find out what the find out what the problem is, problem is, problem is, gather more gather more gather more information from metrics, information from metrics, information from metrics, traces, etc.

  15. traces, etc. traces, etc. Identify where Identify where Identify where the problem is occurring the problem is occurring the problem is occurring using traces using traces using traces and logs, and then and logs, and then and logs, and then start correlating this start correlating this start correlating this with things like with things like with things like change events or change events or any new deployments any new deployments any new deployments in your code and in your code and in your code and infrastructure. infrastructure. infrastructure. It primarily It primarily It primarily performs this analysis performs this analysis performs this analysis or reasoning or reasoning or reasoning based on four based on four based on four types of data sources. types of data sources. types of data sources. Your code, your Your code, your Your code, your infrastructure, your infrastructure, your infrastructure, your knowledge bases, and your knowledge bases, and your observability platforms. Based on observability platforms. Based on all this, all this, all this, he begins to give out he begins to give out he begins to give out the cause or root of the the cause or root of the the cause or root of the problem we problem we problem we saw. Here is one of the saw. Here is one of the saw. Here is one of the call logs that call logs that call logs that actually has a actually has a actually has a high failure rate. high failure rate. high failure rate. He traced He traced He traced the problem back to some the problem back to some the problem back to some legacy legacy legacy integrations that were integrations that were integrations that were still active in the system. What's still active in the system. What's still active in the system. What's interesting is that it interesting is that it interesting is that it also shows some also shows some also shows some other older versions of the other older versions of the other older versions of the reasons we would reasons we would reasons we would normally start an normally start an normally start an investigation with. At the investigation with. At the investigation with. At the same time, there was a same time, there was a same time, there was a GCP outage. So, that would be a GCP outage. So, that would be a GCP outage. So, that would be a perfect perfect perfect match, right? match, right? match, right? Back to AI, Back to AI, Back to AI, which is designed to which is designed to which is designed to provide consistent provide consistent provide consistent responses. So, the responses. So, the responses. So, the GCP outage happened at the GCP outage happened at the GCP outage happened at the same time. Of course, this same time. Of course, this same time. Of course, this could be the reason for what could be the reason for what could be the reason for what we we we observe. Or are there observe. Or are there observe. Or are there other traffic spikes other traffic spikes other traffic spikes it is experiencing it is experiencing it is experiencing due to lack of due to lack of due to lack of integration. That would be integration. That would be integration. That would be another thread you would another thread you would another thread you would start pulling on.

  16. start pulling on. start pulling on. Let's say I'm a new Let's say I'm a new Let's say I'm a new engineer starting out engineer starting out engineer starting out and Resolve gave and Resolve gave and Resolve gave this as the root cause. this as the root cause. this as the root cause. Of course, we also Of course, we also Of course, we also designed the system to designed the system to operate on the operate on the principle of principle of principle of minimum trust. minimum trust. minimum trust. So, if I start So, if I start So, if I start highlighting elements, highlighting elements, highlighting elements, it will engage me it will engage me it will engage me so I can so I can so I can explore it in more detail. In this explore it in more detail. In this explore it in more detail. In this case, let's say case, let's say case, let's say we ask Resolve if we ask Resolve if we ask Resolve if this is related to a this is related to a this is related to a GCP failure? So, he GCP failure? So, he GCP failure? So, he answers it here, but there are answers it here, but there are answers it here, but there are other precise questions other precise questions other precise questions that I've been running through that I want to that I've been running through that I want to that I've been running through that I want to share with you. share with you. share with you. There were other There were other There were other deployment errors that deployment errors that deployment errors that occurred at the same occurred at the same occurred at the same time. This is what I was time. This is what I was time. This is what I was pushing Resolve towards. pushing Resolve towards. pushing Resolve towards. Back to what I Back to what I Back to what I said: if you said: if you push push any AI hard enough, it any AI hard enough, it any AI hard enough, it will start will start will start to agree with you. We to agree with you. We to agree with you. We designed the system to designed the system to counteract this. She counteract this. She will base her will base her will base her answers solely on the answers solely on the causal chain of evidence she has created. chain of evidence she has created. Because the system Because the system Because the system is able to establish this, is able to establish this, is able to establish this, it can give a very it can give a very it can give a very logical and logical and logical and consistent answer consistent answer . It also allows . It also allows . It also allows me to involve my me to involve my me to involve my colleagues so that it truly colleagues so that it truly colleagues so that it truly becomes a becomes a becomes a collaborative effort.

  17. collaborative effort. Let's say I can Let's say I can call Anvir: can call Anvir: can call Anvir: can you take a look at you take a look at you take a look at this? This allows me this? This allows me this? This allows me to add my colleague to add my colleague to add my colleague directly from Slack directly from Slack directly from Slack so we can so we can so we can collaborate in this collaborate in this collaborate in this virtual “ virtual “ situation room situation room .” It was a quick .” It was a quick .” It was a quick demonstration. As I demonstration. As I demonstration. As I said, what said, what said, what we talked about today we talked about today is just one aspect of is just one aspect of is just one aspect of architecture. But architecture. But architecture. But now there's another side of now there's another side of now there's another side of the architecture where you the architecture where you the architecture where you need to turn need to turn need to turn AI into a full-fledged AI into a full-fledged AI into a full-fledged product. If you product. If you product. If you want to present want to present want to present this to a user or this to a user or this to a user or client, you need to client, you need to client, you need to make sure everything is make sure everything is make sure everything is running on a reliable running on a reliable running on a reliable platform. These are other platform. These are other platform. These are other aspects to aspects to aspects to think about, but these are things think about, but these are things think about, but these are things we are all very we are all very we are all very familiar with. So, my time is familiar with. So, my time is familiar with. So, my time is up. If you're up. If you're up. If you're interested in learning interested in learning interested in learning more about Resolve AI, more about Resolve AI, more about Resolve AI, what customers what customers what customers are using us, and how are using us, and how are using us, and how —come visit —come visit —come visit our booth L28, I our booth L28, I our booth L28, I think. And by the way, think. And by the way, think. And by the way, before you leave, before you leave, before you leave, grab those grab those grab those cool bags that cool bags that cool bags that are right at the are right at the are right at the exit. Thank you very much.

Summary

The Resolve AI team is building agents to patch software by automating repairs, a difficult task where AI has historically been limited to programming assistance. The session will share lessons learned in creating these agent systems, highlighting that while AI excels at coding due to its self-documenting, modular, and single-domain nature, the majority of engineers' time is spent on running and fixing existing systems. The practical takeaway is that AI agent systems can be developed and improved to handle these operational tasks, analogous to a healthcare system's functions.

View original episode ↗