I've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
I've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
i call this (new?) type of test the "vibe check"