▲ 87 ▼ Professional software engineers of Lemmy, are code reviews still a thing in the age of "AI" assisted coding? (lemmy.world) submitted 2 months ago by ccunning@lemmy.world to c/asklemmy@lemmy.world 54 comments fedilink hide all child comments
[–] CameronDev@programming.dev 1 point 2 months ago (2 children) Testing functionality isn't the same as correctness. permalink fedilink source parent hideshow 4 child comments replies: [–] slevinkelevra@sh.itjust.works 1 point 2 months ago (1 child) Yeah, I had testers that tested the functionality of a delay... But had set the delay parameter to zero. Well good thing this one case worked, but you didn't check anything beyond that for correctness at all. permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago Timing and tests, name a better migraine duo :D. We continuously create tests that ensure a process completes in an set amount of time, and every time, we don't give them enough leeway, and the test will fail randomly if the CI runner gets overloaded. permalink fedilink source parent [–] ellen.kimble@piefed.social 1 point 2 months ago (1 child) Oh excuse me then, what is correctness? permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago (1 child) int add(int a, int b) { return a + b; } This code is clearly functional, it'll compile and execute. However, the customer actually needs the code to do a saturating add. With that knowledge, we can clearly see that the code is not correct. It will not saturate, it will wrap around instead. Without that knowledge, an LLM will happily write some basic unit tests that won't cover the saturation edge case, and the bug would live on until its hit in prod. If you're lucky, and your function doco is good, the LLM might spot the bug, and notify you. My personal preference for how to generate tests is to ask the agent to write specific tests. E.g: "write a test for add that demonstrates that it saturates". permalink fedilink source parent hideshow 2 child comments replies: [–] slevinkelevra@sh.itjust.works 2 points 2 months ago (1 child) IMO this is a bad example as in theory, testers test code against requirements, and if there is no such req stating anything about saturation then how should the testers or in this case the LLM know? permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago It is over simplified, but there are often implicit requirements that a human would be aware of from the broader context that the LLM may not be. i.e add is used to increment a health bar, so wrap around doesn't make sense. permalink fedilink source parent
[–] slevinkelevra@sh.itjust.works 1 point 2 months ago (1 child) Yeah, I had testers that tested the functionality of a delay... But had set the delay parameter to zero. Well good thing this one case worked, but you didn't check anything beyond that for correctness at all. permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago Timing and tests, name a better migraine duo :D. We continuously create tests that ensure a process completes in an set amount of time, and every time, we don't give them enough leeway, and the test will fail randomly if the CI runner gets overloaded. permalink fedilink source parent
[–] CameronDev@programming.dev 1 point 2 months ago Timing and tests, name a better migraine duo :D. We continuously create tests that ensure a process completes in an set amount of time, and every time, we don't give them enough leeway, and the test will fail randomly if the CI runner gets overloaded. permalink fedilink source parent
[–] ellen.kimble@piefed.social 1 point 2 months ago (1 child) Oh excuse me then, what is correctness? permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago (1 child) int add(int a, int b) { return a + b; } This code is clearly functional, it'll compile and execute. However, the customer actually needs the code to do a saturating add. With that knowledge, we can clearly see that the code is not correct. It will not saturate, it will wrap around instead. Without that knowledge, an LLM will happily write some basic unit tests that won't cover the saturation edge case, and the bug would live on until its hit in prod. If you're lucky, and your function doco is good, the LLM might spot the bug, and notify you. My personal preference for how to generate tests is to ask the agent to write specific tests. E.g: "write a test for add that demonstrates that it saturates". permalink fedilink source parent hideshow 2 child comments replies: [–] slevinkelevra@sh.itjust.works 2 points 2 months ago (1 child) IMO this is a bad example as in theory, testers test code against requirements, and if there is no such req stating anything about saturation then how should the testers or in this case the LLM know? permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago It is over simplified, but there are often implicit requirements that a human would be aware of from the broader context that the LLM may not be. i.e add is used to increment a health bar, so wrap around doesn't make sense. permalink fedilink source parent
[–] CameronDev@programming.dev 1 point 2 months ago (1 child) int add(int a, int b) { return a + b; } This code is clearly functional, it'll compile and execute. However, the customer actually needs the code to do a saturating add. With that knowledge, we can clearly see that the code is not correct. It will not saturate, it will wrap around instead. Without that knowledge, an LLM will happily write some basic unit tests that won't cover the saturation edge case, and the bug would live on until its hit in prod. If you're lucky, and your function doco is good, the LLM might spot the bug, and notify you. My personal preference for how to generate tests is to ask the agent to write specific tests. E.g: "write a test for add that demonstrates that it saturates". permalink fedilink source parent hideshow 2 child comments replies: [–] slevinkelevra@sh.itjust.works 2 points 2 months ago (1 child) IMO this is a bad example as in theory, testers test code against requirements, and if there is no such req stating anything about saturation then how should the testers or in this case the LLM know? permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago It is over simplified, but there are often implicit requirements that a human would be aware of from the broader context that the LLM may not be. i.e add is used to increment a health bar, so wrap around doesn't make sense. permalink fedilink source parent
[–] slevinkelevra@sh.itjust.works 2 points 2 months ago (1 child) IMO this is a bad example as in theory, testers test code against requirements, and if there is no such req stating anything about saturation then how should the testers or in this case the LLM know? permalink fedilink source parent hideshow 2 child comments replies: [–] CameronDev@programming.dev 1 point 2 months ago It is over simplified, but there are often implicit requirements that a human would be aware of from the broader context that the LLM may not be. i.e add is used to increment a health bar, so wrap around doesn't make sense. permalink fedilink source parent
[–] CameronDev@programming.dev 1 point 2 months ago It is over simplified, but there are often implicit requirements that a human would be aware of from the broader context that the LLM may not be. i.e add is used to increment a health bar, so wrap around doesn't make sense. permalink fedilink source parent