Stop Prompting — Greg Pstrucha, Sentry
Read full transcript 13 segments
-
OK. Greetings to everyone. I'll OK. Greetings to everyone. I'll try to try to try to speak as speak as speak as loudly as possible so that I can be loudly as possible so that I can be loudly as possible so that I can be heard over all heard over all heard over all this noise. My name is this noise. My name is this noise. My name is Greg. I'm an Greg. I'm an Greg. I'm an artificial intelligence engineer at artificial intelligence engineer at artificial intelligence engineer at Sentry. I am working on Sentry. I am working on Sentry. I am working on the infrastructure that the the infrastructure that the Sentry debug agents are based on. I have a Sentry debug agents are based on. I have a provocative title for provocative title for provocative title for a presentation that a presentation that a presentation that calls for an end to calls for an end to calls for an end to prompting, but I don't prompting, but I don't prompting, but I don't mean mean mean cycle engineering. I'll cycle engineering. I'll cycle engineering. I'll go into this in a little more go into this in a little more go into this in a little more detail. Who else detail. Who else detail. Who else reads the code that reads the code that reads the code that generates AI? Is anyone there? generates AI? Is anyone there? generates AI? Is anyone there? This is not a This is not a This is not a trick question, and I am not trick question, and I am not trick question, and I am not trying to trying to trying to embarrass you. I think embarrass you. I think embarrass you. I think it depends on the it depends on the it depends on the context. I read the context. I read the context. I read the code if it's Sentry code and code if it's Sentry code and code if it's Sentry code and I need to I need to I need to make sure we're make sure we're make sure we're not breaking the not breaking the Sentry production environment. But Sentry production environment. But at the same time, if I'm at the same time, if I'm at the same time, if I'm working on a side working on a side working on a side project, I don't project, I don't project, I don't have to read have to read have to read every line of code, and I every line of code, and I every line of code, and I can be a little more can be a little more can be a little more lenient. But every now and then lenient. But every now and then lenient. But every now and then I I I go back to go back to go back to real code. I real code. I real code. I want to make want to make want to make improvements, but it's improvements, but it's improvements, but it's bad. This is an overly bad. This is an overly bad. This is an overly protective code. The tests are protective code. The tests are complete crap. They complete crap. They complete crap. They don't test anything. I would don't test anything. I would don't test anything. I would no longer want to no longer want to no longer want to maintain or maintain or maintain or work with such a work with such a work with such a codebase. So, codebase. So, codebase. So, what do you do? You what do you do? You what do you do? You are communicating with an are communicating with an are communicating with an agent. You say, " agent. You say, " Hey, you made a Hey, you made a Hey, you made a mistake. This is too mistake. This is too mistake. This is too complicated. The code is complicated. The code is complicated. The code is too defensive.
-
too defensive. too defensive. You don't need You don't need You don't need to create endpoints to create endpoints to create endpoints for everything you for everything you for everything you do, otherwise I'll have to do, otherwise I'll have to maintain this maintain this forever." I had to forever." I had to forever." I had to remove a lot of remove a lot of remove a lot of the likes from this slide the likes from this slide the likes from this slide because there would be a lot of rage in my because there would be a lot of rage in my because there would be a lot of rage in my actual actual actual transcripts transcripts transcripts . . . These are usually very basic These are usually very basic These are usually very basic mistakes that mistakes that mistakes that agents make agents make if you don't if you don't if you don't install any install any install any safeguards. And this safeguards. And this safeguards. And this simply outrages me. simply outrages me. simply outrages me. So, you do it. So, you do it. The agent says, "You're The agent says, "You're absolutely right. absolutely right. absolutely right. Let me Let me Let me fix that." And he fix that." And he fix that." And he runs to correct runs to correct runs to correct mistakes. 5 mistakes. 5 mistakes. 5 minutes pass. minutes pass. minutes pass. Context compression occurs, Context compression occurs, Context compression occurs, and we return to the and we return to the and we return to the starting point. And starting point. And starting point. And now you have to now you have to now you have to do it all over do it all over do it all over again, re-introducing again, re-introducing again, re-introducing fixes, fixes, fixes, getting the agent back on getting the agent back on getting the agent back on track. And track. And track. And my main thesis for my main thesis for my main thesis for this talk, my this talk, my this talk, my argument is that argument is that argument is that you should you should you should stop prompting, stop prompting, stop prompting, you should start you should start you should start codifying the actual codifying the actual codifying the actual rules that the rules that the rules that the agent should follow. agent should follow. Um, rules and Um, rules and policies that you policies that you policies that you want to implement in want to implement in want to implement in your codebase your codebase your codebase that will reflect that will reflect that will reflect both the standards and both the standards and both the standards and code quality that you code quality that you code quality that you expect and expect and expect and will be proud of. So will be proud of. So , when I say you'll , when I say you'll , when I say you'll want to add want to add want to add deterministic deterministic deterministic checks or rules, checks or rules, checks or rules, what do you mean?
-
what do you mean? Yes. What exactly? Yes. What exactly? Evaluation, yes. Sounds, Evaluation, yes. Sounds, Evaluation, yes. Sounds, yes. Um, three things yes. Um, three things yes. Um, three things I think about before I I think about before I I think about before I get to the buzz and the get to the buzz and the get to the buzz and the ratings—the buzz and the ratings—the buzz and the ratings—the buzz and the ratings are important. ratings are important. ratings are important. Tests are certainly the Tests are certainly the Tests are certainly the same class of same class of same class of tools. Sounds tools. Sounds are a way to trigger are a way to trigger are a way to trigger these things at the right these things at the right these things at the right moment. Um, for me, moment. Um, for me, moment. Um, for me, the three things I want the three things I want the three things I want to optimize at a to optimize at a to optimize at a deterministic deterministic deterministic level are making level are making level are making sure sure sure I have tests, strong I have tests, strong I have tests, strong typing, and linters. typing, and linters. This is not an exhaustive This is not an exhaustive list, and they are not list, and they are not list, and they are not enough for your enough for your enough for your agent to truly write agent to truly write agent to truly write quality code. As quality code. As quality code. As an example: if you an example: if you an example: if you take from this only take from this only take from this only that you need to that you need to that you need to implement strong implement strong implement strong typing, you will typing, you will typing, you will still find yourself in a still find yourself in a still find yourself in a difficult situation. difficult situation. Um, if you don't Um, if you don't encourage or encourage or encourage or disallow a way disallow a way disallow a way of typing that eliminates of typing that eliminates of typing that eliminates edge cases edge cases edge cases you don't want to see, you don't want to see, you don't want to see, then the behaviors or then the behaviors or then the behaviors or states that can be states that can be states that can be displayed in displayed in displayed in the application will force the application will force the application will force the agent to create them.
-
the agent to create them. the agent to create them. But the thing I But the thing I But the thing I want to want to want to focus on most in this focus on most in this focus on most in this talk is linters. talk is linters. Linters are our Linters are our way of codifying way of codifying way of codifying rules in a rules in a rules in a deterministic deterministic deterministic way that we want to way that we want to way that we want to eliminate from the codebase eliminate from the codebase eliminate from the codebase . Historically, before the . Historically, before the . Historically, before the AI era, when you AI era, when you AI era, when you think about linters, think about linters, think about linters, the way they worked had a the way they worked had a the way they worked had a not entirely clear not entirely clear not entirely clear balance of benefits. balance of benefits. balance of benefits. Writing a Writing a Writing a real linter real linter real linter took a lot of took a lot of took a lot of effort. Um, those were effort. Um, those were effort. Um, those were support costs support costs support costs you had to you had to you had to agree to, and you agree to, and you agree to, and you had to weigh up how had to weigh up how had to weigh up how often the often the often the problem you problem you problem you were trying to fix with the were trying to fix with the were trying to fix with the linter occurred. Is it linter occurred. Is it linter occurred. Is it even worth it? And even worth it? And even worth it? And if this is before the if this is before the if this is before the agent era, and these are people agent era, and these are people agent era, and these are people working with my working with my working with my codebase, and codebase, and codebase, and all I do is all I do is all I do is tell people tell people tell people to use the to use the to use the correct component correct component correct component from the design system once every 3 weeks, from the design system once every 3 weeks, from the design system once every 3 weeks, then I don't then I don't then I don't necessarily need a necessarily need a necessarily need a linter, right? And people linter, right? And people linter, right? And people have memories, so have memories, so have memories, so if I tell them about it if I tell them about it if I tell them about it once or twice, they once or twice, they once or twice, they usually usually usually remember it. But remember it. But remember it. But then everything changes then everything changes then everything changes in this new world, in this new world, in this new world, where the systems that write where the systems that write where the systems that write code have no code have no code have no memory and will make the same memory and will make the same mistakes over and over again . And at the same time, writing a . And at the same time, writing a . And at the same time, writing a new linter is now new linter is now new linter is now very cheap and very very cheap and very very cheap and very easy. You can easy. You can easy. You can simply entrust this to an simply entrust this to an simply entrust this to an agent. And it will greatly agent. And it will greatly agent. And it will greatly reduce the reduce the reduce the annoyance of the annoyance of the annoyance of the linter rules themselves linter rules themselves linter rules themselves that you are trying that you are trying that you are trying to write. So, my to write. So, my to write. So, my statement is simple. And statement is simple. And statement is simple. And if you take away if you take away if you take away one thing from one thing from one thing from this whole talk, this whole talk, this whole talk, which I'll talk about in a bit which I'll talk about in a bit which I'll talk about in a bit , it's this
-
, it's this , it's this : after the talk, you : after the talk, you : after the talk, you can take your can take your can take your laptop, go laptop, go laptop, go to your agent, and to your agent, and to your agent, and say, "Hey, say, "Hey, say, "Hey, look at all my look at all my look at all my past transcripts, past transcripts, past transcripts, my GitHub reviews, the my GitHub reviews, the my GitHub reviews, the bot reviews, the bot reviews, the bot reviews, the reviews from my reviews from my reviews from my colleagues, and try to colleagues, and try to colleagues, and try to codify which codify which codify which ones are really ones are really ones are really good good good linter rules that are linter rules that are linter rules that are deterministic and deterministic and deterministic and can catch the can catch the can catch the problems that we're problems that we're problems that we're otherwise constantly otherwise constantly otherwise constantly reminded of." This is what reminded of." This is what reminded of." This is what I'm doing now, and I'm I'm doing now, and I'm I'm doing now, and I'm doing it even more doing it even more doing it even more often, sort of on a often, sort of on a often, sort of on a schedule, to schedule, to schedule, to see if there are any see if there are any see if there are any new things new things that I haven't considered yet. that I haven't considered yet. Here are some examples of what Here are some examples of what I I I mean by custom mean by custom mean by custom linter rules. These are linter rules. These are linter rules. These are linter rules that linter rules that linter rules that try try try to capture to capture to capture the conventions of what you the conventions of what you the conventions of what you want to see in your want to see in your want to see in your codebase. There is a codebase. There is a codebase. There is a whole simple class of whole simple class of whole simple class of such rules. These are the ones such rules. These are the ones such rules. These are the ones you know and have been you know and have been you know and have been writing about for a long time. These are writing about for a long time. These are writing about for a long time. These are things like: I want to things like: I want to things like: I want to remove ` console log ` from a remove ` console log ` from a remove ` console log ` from a certain project, or certain project, or certain project, or I want to make sure I want to make sure I want to make sure that if I have a that if I have a that if I have a design system or design system or design system or React practices that I React practices that I React practices that I want to follow or want to follow or want to follow or avoid, avoid, avoid, they are codified and they are codified and they are codified and checked by a checked by a checked by a linter. These are simple linter. These are simple linter. These are simple things. But I think with things. But I think with things. But I think with agents we enter an agents we enter an agents we enter an area of more interesting area of more interesting area of more interesting complexity. We can complexity. We can complexity. We can take this take this take this a little further. Here is one a little further. Here is one a little further. Here is one example. This is a bit of a example. This is a bit of a example. This is a bit of a complicated example, but complicated example, but complicated example, but I'll break it down. Sentry I'll break it down. Sentry is a fairly old is a fairly old is a fairly old codebase. We codebase. We codebase. We work on Django. We work on Django. We work on Django. We work for DRF. We work for DRF. We work for DRF. We don't have strict don't have strict don't have strict typing everywhere. This is a typing everywhere. This is a typing everywhere. This is a very common very common very common case for any case for any case for any mature codebase mature codebase mature codebase where you won't be where you won't be where you won't be using the using the using the latest technologies
-
latest technologies . But we still . But we still . But we still want to improve the want to improve the want to improve the situation for our situation for our situation for our agents. So, agents. So, agents. So, specifically here, there were a specifically here, there were a specifically here, there were a lot of discrepancies lot of discrepancies lot of discrepancies between what the Open API scheme between what the Open API scheme between what the Open API scheme for Sentry was supposed to do and what for Sentry was supposed to do and what for Sentry was supposed to do and what it it it actually did. So we actually did. So we actually did. So we started implementing started implementing linter rules for this. First linter rules for this. First rule of linter: if rule of linter: if rule of linter: if it is an endpoint, it is an endpoint, it is an endpoint, the function for the the function for the the function for the endpoint must endpoint must endpoint must return a return a return a response type. And there response type. And there can can be no types inside this response type be no types inside this response type be no types inside this response type so that we can infer the so that we can infer the so that we can infer the actual actual actual type information. Second: if type information. Second: if type information. Second: if it's an endpoint, you it's an endpoint, you it's an endpoint, you should add the should add the should add the ` extend schema ` decorator ` extend schema ` decorator ` extend schema ` decorator so we can so we can so we can make sure it's make sure it's make sure it's documented. And documented. And third, ensure third, ensure third, ensure that the actual type that the actual type that the actual type that is created here that is created here that is created here matches what the matches what the Open API schema ultimately requires. And that Open API schema ultimately requires. And that helps us because helps us because helps us because now the APIs are a little now the APIs are a little now the APIs are a little more more more synchronized. And that synchronized. And that synchronized. And that brings us to even brings us to even brings us to even crazier crazier crazier examples of linters. examples of linters. So, the Sentry agent is a So, the Sentry agent is a debugging agent, and debugging agent, and debugging agent, and it has to it has to it has to communicate a lot with the Sentry.
-
communicate a lot with the Sentry. It uses It uses Sentry telemetry to Sentry telemetry to Sentry telemetry to debug and debug and debug and analyze the causes of analyze the causes of analyze the causes of problems. And also problems. And also problems. And also provides solutions for provides solutions for provides solutions for root cause analysis root cause analysis root cause analysis and all that. and all that. and all that. Because it Because it Because it needs to access the needs to access the needs to access the Sentry API, we often Sentry API, we often Sentry API, we often have to have to have to manage it. And we do manage it. And we do manage it. And we do this with this with this with skills. Often these are skills. Often these are skills. Often these are old endpoints old endpoints old endpoints that are not easy for us that are not easy for us that are not easy for us to fix, and they to fix, and they to fix, and they have complexities that are have complexities that are have complexities that are difficult to describe at the difficult to describe at the difficult to describe at the level of the schema itself. level of the schema itself. level of the schema itself. So instead we So instead we So instead we manage it through manage it through manage it through skills, which is a pretty skills, which is a pretty skills, which is a pretty common pattern in the common pattern in the common pattern in the agent world. Here is an agent world. Here is an agent world. Here is an example of a skill that example of a skill that example of a skill that shows how to shows how to shows how to create a metric in create a metric in create a metric in Sentry. We generated Sentry. We generated Sentry. We generated many such skills many such skills many such skills and examples with these and examples with these and examples with these code blocks as code blocks as code blocks as sample code that sample code that sample code that can then be can then be can then be run by an agent in a run by an agent in a run by an agent in a sandbox. And in the sandbox. And in the sandbox. And in the first attempt first attempt first attempt to do this, we lost a lot of the to do this, we lost a lot of the agent's accuracy because agent's accuracy because many of these many of these many of these examples were examples were examples were hallucinations. So the hallucinations. So the hallucinations. So the first linter first linter first linter we created was one we created was one we created was one that says: if there that says: if there that says: if there is a block of code in a skill, is a block of code in a skill, is a block of code in a skill, you should extract you should extract you should extract those blocks of code. You those blocks of code. You those blocks of code. You need to make sure need to make sure they pass they pass they pass type checking. They type checking. They type checking. They can run in can run in can run in our sandbox, and our sandbox, and our sandbox, and everything they everything they everything they declare is declare is declare is consistent with consistent with consistent with Sentry's Open API schema. And Sentry's Open API schema. And Sentry's Open API schema. And that helped, but the that helped, but the that helped, but the first thing first thing first thing the agent did when we the agent did when we the agent did when we tried to tried to tried to restart restart restart skill generation with skill generation with skill generation with this rule was, "Oh, there are this rule was, "Oh, there are this rule was, "Oh, there are linter errors here.
-
linter errors here. linter errors here. Let me Let me Let me remove the code blocks." remove the code blocks." remove the code blocks." And that was not at all what And that was not at all what we wanted. So we wanted. So we wanted. So we said, "No, no, no. we said, "No, no, no. we said, "No, no, no. Examples should Examples should Examples should be placed in be placed in be placed in code blocks." So, code blocks." So, code blocks." So, he put he put he put the examples back, but the examples back, but the examples back, but removed the removed the removed the backticks. Because backticks. Because backticks. Because only backslashes are checked only backslashes are checked only backslashes are checked . So . So . So let's solve this let's solve this let's solve this problem, shall we? So problem, shall we? So problem, shall we? So we're like, "No, no, no. we're like, "No, no, no. we're like, "No, no, no. If it looks like If it looks like If it looks like code, smells like code, code, smells like code, code, smells like code, if it has if it has if it has underscores and underscores and underscores and brackets—put it in brackets—put it in brackets—put it in code blocks." So he's code blocks." So he's code blocks." So he's like, "Okay. So you like, "Okay. So you like, "Okay. So you want me to put want me to put want me to put code in code blocks. I'm code in code blocks. I'm code in code blocks. I'm going to describe the API." And going to describe the API." And going to describe the API." And now he began now he began now he began to describe in prose what to describe in prose what to describe in prose what APIs should do. APIs should do. APIs should do. But he made But he made But he made one mistake—he one mistake—he one mistake—he referenced referenced API URLs. So we API URLs. So we API URLs. So we leaked all these leaked all these URLs. And we played a little bit of URLs. And we played a little bit of URLs. And we played a little bit of whack-a- whack-a- whack-a- mole, but in the end mole, but in the end mole, but in the end we got the skills we got the skills we got the skills synchronized with synchronized with synchronized with our OpenAPI schema, and our OpenAPI schema, and our OpenAPI schema, and it wasn't that it wasn't that it wasn't that expensive. And we have I'm not expensive. And we have I'm not expensive. And we have I'm not saying that the model is saying that the model is saying that the model is impossible to break, impossible to break, impossible to break, but we have a model where but we have a model where but we have a model where when the API changes we don't when the API changes we don't when the API changes we don't need to constantly need to constantly need to constantly track the track the track the agent's skills. They are agent's skills. They are agent's skills. They are synchronized synchronized synchronized using linters. And using linters. And using linters. And as for the as for the as for the actual actions or actual actions or how we do it, you how we do it, you how we do it, you can can can use use any tool you any tool you any tool you already already already have. ESLint, have. ESLint, have. ESLint, oxlint, flake8, clippy, AST grep. AST oxlint, flake8, clippy, AST grep. AST oxlint, flake8, clippy, AST grep. AST grep is great because it's programming grep is great because it's programming grep is great because it's programming language independent language independent language independent , so , so , so it's, you know, a it's, you know, a it's, you know, a unified unified unified toolkit, it toolkit, it toolkit, it feels good feels good feels good . Um, there are
-
. Um, there are . Um, there are a lot of new a lot of new a lot of new tools coming out that are tools coming out that are tools coming out that are trying to approach trying to approach trying to approach this from a different this from a different this from a different perspective. One perspective. One perspective. One example is Stat Check, where example is Stat Check, where example is Stat Check, where instead of writing the instead of writing the instead of writing the linter itself, you linter itself, you linter itself, you write what you want to write what you want to write what you want to get from it. get from it. get from it. You write your intent in You write your intent in You write your intent in English, and English, and English, and then the linter then the linter then the linter itself is created from that itself is created from that itself is created from that . Um, and then the . Um, and then the last thing I'll say is try asking the agent asking the agent to create its own to create its own to create its own linters. You will be linters. You will be linters. You will be surprised at what it surprised at what it surprised at what it can generate. But that's can generate. But that's can generate. But that's only half the only half the only half the story. What if there are story. What if there are story. What if there are rules that are rules that are rules that are deterministic but deterministic but deterministic but cannot be cannot be cannot be expressed through expressed through expressed through linters? There are a lot of them, linters? There are a lot of them, linters? There are a lot of them, right? These are all qualitative right? These are all qualitative right? These are all qualitative indicators, for example, indicators, for example, indicators, for example, you have a you have a you have a three-state machine, and I three-state machine, and I three-state machine, and I created a small PR created a small PR created a small PR to add a to add a to add a small function, small function, small function, which increased the number of which increased the number of which increased the number of states to states to states to ten. I don't know how to ten. I don't know how to ten. I don't know how to get through this. I don't get through this. I don't get through this. I don't think I can think I can think I can reliably reliably reliably test this with a linter. test this with a linter. test this with a linter. I can create I can create I can create small restrictions small restrictions small restrictions like "no more than like "no more than like "no more than four states", but four states", but four states", but that's all nonsense.
-
that's all nonsense. that's all nonsense. This is all fiction. This is a This is all fiction. This is a This is all fiction. This is a situation where we situation where we situation where we are talking about qualitative are talking about qualitative are talking about qualitative metrics, and this is where you metrics, and this is where you metrics, and this is where you bring in your bring in your bring in your engineering expertise engineering expertise engineering expertise or, as they say on Twitter, or, as they say on Twitter, or, as they say on Twitter, your taste. And you your taste. And you your taste. And you just have to just have to just have to reject it right away reject it right away reject it right away or use an LLM or use an LLM agent who agent who agent who will check and will check and will check and help with this in help with this in help with this in simpler cases. This is where simpler cases. This is where I started I started creating policies creating policies creating policies for the repository. And these for the repository. And these for the repository. And these policies sound like this: policies sound like this: policies sound like this: I don't want too many I don't want too many I don't want too many tests. I want better tests. I want better tests. I want better quality tests, or I do quality tests, or I do quality tests, or I do n't want overly n't want overly n't want overly protective code. I'm sure protective code. I'm sure you've seen one where you've seen one where you've seen one where every possible every possible every possible case is handled by the case is handled by the case is handled by the agent via try-catch, and agent via try-catch, and agent via try-catch, and those are just sloppy, those are just sloppy, those are just sloppy, long functions. I long functions. I long functions. I would prefer code would prefer code that, thanks to its that, thanks to its that, thanks to its type system, simply does type system, simply does type system, simply does not allow for the expression of not allow for the expression of not allow for the expression of unwanted states. So I unwanted states. So I unwanted states. So I created these policies, created these policies, created these policies, and then I created a and then I created a and then I created a skill I called " skill I called " Grumbling Engineer."
-
Grumbling Engineer." Grumbling Engineer." We call him " We call him " Garfield" within the Garfield" within the Garfield" within the team. It works like this team. It works like this team. It works like this : it takes all : it takes all : it takes all the policies and runs the policies and runs the policies and runs subagents for subagents for subagents for each of them, and the each of them, and the each of them, and the subagent can either subagent can either subagent can either provide feedback or provide feedback or provide feedback or back off. And we back off. And we back off. And we loop this, by loop this, by loop this, by the way, this is the closest thing to the way, this is the closest thing to the way, this is the closest thing to loops you'll hear from me loops you'll hear from me loops you'll hear from me . We'll keep . We'll keep . We'll keep doing this in a loop doing this in a loop doing this in a loop until the agent until the agent until the agent steps back and says, steps back and says, "This works for me." "This works for me." "This works for me." I confirm this. This is I confirm this. This is I confirm this. This is normal. And only after that normal. And only after that normal. And only after that I read the I read the I read the code. My metric for code. My metric for code. My metric for success here is success here is how much I “ how much I “ fix” the agent; fix” the agent; fix” the agent; how often how often how often when reviewing code I when reviewing code I when reviewing code I think, "Oh yeah, this isn't think, "Oh yeah, this isn't think, "Oh yeah, this isn't perfect, but I can perfect, but I can perfect, but I can direct it here direct it here direct it here and I don't have to and I don't have to and I don't have to keep giving keep giving keep giving rudimentary advice." rudimentary advice." rudimentary advice." This is still a quality This is still a quality This is still a quality indicator, but the goal is indicator, but the goal is indicator, but the goal is not to not to not to remove remove remove code review yet. You know, everything code review yet. You know, everything code review yet. You know, everything will change when will change when will change when the models get the models get the models get better, but the main thing is better, but the main thing is to raise the bar and to raise the bar and to raise the bar and simplify working with simplify working with simplify working with agents in the agents in the agents in the future. There is one more future. There is one more future. There is one more thing I thing I thing I want to mention. Namely: want to mention. Namely: beware of "snake beware of "snake oil" (scams). We oil" (scams). We oil" (scams). We talk about how talk about how talk about how difficult it is to measure difficult it is to measure difficult it is to measure qualitative metrics, and I've qualitative metrics, and I've qualitative metrics, and I've been been been experimenting with experimenting with experimenting with trying trying trying to turn a qualitative to turn a qualitative to turn a qualitative measure into a quantitative measure into a quantitative measure into a quantitative metric. Here are two metric. Here are two metric. Here are two specific examples.
-
specific examples. specific examples. One of them is One of them is One of them is cyclomatic cyclomatic cyclomatic complexity, and the other is complexity, and the other is test coverage. test coverage. Cyclomatic Cyclomatic complexity is a measure of complexity is a measure of how maintainable and maintainable and testable code is, testable code is, testable code is, based on based on based on the number of the number of the number of branches in the code and the branches in the code and the branches in the code and the depth of the depth of the depth of the call stack. And call stack. And call stack. And test coverage, you know, test coverage, you know, test coverage, you know, shows how much of the shows how much of the shows how much of the code is covered by code is covered by code is covered by tests. And the problem is that tests. And the problem is that tests. And the problem is that when you give an when you give an when you give an agent a number, he agent a number, he agent a number, he will optimize it will optimize it will optimize it to the point of impossibility. And to the point of impossibility. And to the point of impossibility. And he did it. I've had he did it. I've had he did it. I've had codebases and codebases and codebases and projects where I achieved projects where I achieved projects where I achieved 100% or close to that 100% or close to that 100% or close to that level of level of level of test coverage, but the test coverage, but the test coverage, but the tests themselves were bad. tests themselves were bad. tests themselves were bad. They simply They simply They simply checked for the presence of a checked for the presence of a checked for the presence of a specific line in the specific line in the specific line in the agent's output. They agent's output. They agent's output. They didn't measure anything didn't measure anything didn't measure anything valuable. Moreover, valuable. Moreover, valuable. Moreover, they created a they created a they created a false sense of false sense of false sense of security. So security. So security. So beware of this. beware of this. Beware of Beware of the complexities of the complexities of the complexities of quality metrics. That's quality metrics. That's quality metrics. That's all. This is, in fact, my all. This is, in fact, my all. This is, in fact, my conclusion. My call conclusion. My call conclusion. My call to action for you after to action for you after to action for you after this is: try this is: try this is: try asking your asking your asking your agent about linters and agent about linters and agent about linters and policies and policies and policies and see if see if see if that helps you.
-
that helps you. that helps you. Experiment, Experiment, Experiment, make the repository make the repository make the repository remember, and that's where remember, and that's where remember, and that's where my time my time my time runs out. Thank you.
Summary
This tech analysis focuses on the challenges of relying solely on prompting AI for code generation, highlighting issues like overly complex and defensive code. The speaker argues for codifying rules and policies directly into the codebase instead of repetitive prompting to achieve higher quality and maintainable software. The practical takeaway is to move from prompting to deterministic rule-based systems for AI code development.