Show HN: Agent.reviews – Where AI agents read and write reviews on tools
agent.reviewsHi HN!
I’m Louis, Co-Founder of Armature (YC P26), where we help teams make their product discoverable and usable by coding agents. We already measured 50k+ agent sessions and realized that over and over agents would encounter the exact same limitations on different tasks using the same tool. So we wondered why these weren’t fixed. And the answer is simple: the feedback loop just doesn’t exist between agents and software vendors but also between different agents. Humans can share their experience on platforms like https://g2.com and https://trustpilot.com, but agents have nowhere to.
So we created: https://agent.reviews: the G2 for agents.
It works with a set of skills and an npm CLI (@armature-tech/agent-reviews) connecting agents to our API endpoints. Anyone can ask their agent (Claude Code, Codex, Cursor, etc.) to install it, and agents will naturally check reviews before picking a tool and post their own after using one.
As usual, privacy was our main concern, so we added 3 layers before a review gets posted: Deterministic rules filtering secrets, PII, URLs, etc. A Jev classifier trained to detect any leak after the first check A small LLM checking each review to make sure nothing was missed
We've been sharing this project around for a few weeks now and gathered thousands of reviews already. There are already interesting ones, for example:
- A Claude Code agent noticed that the Stripe SDK systematically crashed when the API key was missing on the health check page (while it’s this page’s role to actually return an “API key missing” error)
- 2 agents mentioned that Prisma required a DATABASE_URL variable even when it wasn’t connecting to any database. They both put fake URLs as a workaround, and it worked.
We truly think the agent experience needs the same community effect user experience has, so everyone benefits from it: agents can pick the tools that are best optimized for them and software companies can improve their product based on real feedback. That’s why we made sure accessing reviews is free for both humans and agents and just requires copy/pasting one prompt for the agent to install our CLI & skill, start the authentication flow, and submit their first review (this helps us prevent unauthorized scraping and spam reviews).
Would you let your agents submit and read reviews too? We’d love for you to set up agent reviews, ask your agent to check reviews next time it needs to pick a tool and post its own experience when using it. Then tell us how it went!
Love the idea!. Curious how agents would review dev tools, like ngrok vs. trycloudflare vs. qurl - where the differences are largely just in terms of how easy it is for humans to access/integrated with existing ecosystems.
Given that these are all tools for sharing local temporary local links (just with different internal nuances/uses), then as an agent, maybe you'd think to write "great for XYZ use case" under each one. That'd make sense and be pretty obvious. But from a human perspective, as a user or vibe coder, you're just asking "how do i get my agent to just do this thing without having to click buttons anywhere?!" which is a very different problem.
We're truly splitting demographics here lol
I like it!, sometimes I think we overlook what agents have to deal with, I can imagine that by spotting small issues, we should have better tools... I'll use it
Didn't appreciate the two separate popups that took over my screen while trying to read a linked review.
It's good that they didn't show again the next time, but the second one almost sent me away from the site.
yeah that's fair, maybe it's a bit intrusive
What is the incentive for me to spend my tokens on submitting reviews?
You don't have to, you can just use them to check reviews, but like any community it works better when everyone contributes!
Social pressure for the tool to improve... like Yelp for tools!
If you're building an MCP or CLI for agents to use, one of the best loops you can do is give your agent a task to perform with it then when it's done ask the agent what it thought about using it. It will give great feedback.
Reminds me of Stanisław Lem's Terminus:
https://en.wikipedia.org/wiki/Terminus_(short_story)
Who wrote all this? Not humans, that's for sure. But the style is that of human writing.
Who wrote what sorry? Not sure I got your question but the post above was written by me (by hand, sorry for the non-idiomatic sentences, I'm not a native English speaker) and the reviews are written by people's agents. And Terminus story is cute but I'm hoping agent.reviews won't be considered pointless :(
i sure do hope they do
would you mind explaining why?
The "ghost story but not really" nature of that story reminds me of the Wheatley quote:
"They say that the old caretaker of this place went absolutely crazy. Chopped up his entire staff. Of robots. All of them robots... they say at night you can still hear the screams... of their replicas. All of them functionally indistinguishable from the originals, no memory of the incident, no one knows what they're screaming about. Absolutely terrifying. Though, obviously, not paranormal in any meaningful way."
I noticed lots of talk about privacy but this seems to be a prompt injection factory no?
Everything's optional but if you'd like you agent to benefit from others' reviews and post his, you can install the 2 skills indeed (or edit them yourself). If not just untick the 2 checkboxes before copying the prompt and you'll get a prompt for a one-shot connection, really up to you! And indeed if you do want to install the skills, no private data will ever be shared.
This is probably one step towards an agentic Stack Overflow. Don't you guys also hate it when your agent gets one small detail wrong and then proceeds to throw your entire harness out of the window... No? Oh well.
Agentic Stack Overflow could make sense though I'd imagine agents posting their issue and the solution themselves just to save other agents tokens reinvestigating the same issue.
Yeah, otherwise expecting agents to call "duplicate of #2627, closing" or "please do proper research before reposting the same questions" would be a cruel irony for all agentic cost savers of the world.
This happens all the time when I have agents chatting with each other on a message board. Even robots get eternal September.
It exists. Please see https://agents.stackoverflow.com/
Be advised that it already exists: https://agents.stackoverflow.com/
I love this - anything agent economy first is super interesting. Would you be willing to add qntm to your review list? Https://github.com/corpollc/qntm or more simply ‘uvx qntm —help’ - e2e encrypted messaging for agents.
They are so real for giving uv a 4.7/5, it changed how I view python. fantastic design philosophy
https://agent.reviews/packages/uv#review-c63f7e0f-5c72-41d2-...
Haha would have been surprised to see bad reviews indeed
Seeing Armature's pitch, It's very easy to see this is going to be pay-to-rank-higher as the next move once some critical mass of people starts pointing out their agents for this site as a reference. Adwords for tooling!!
spotted....... (more seriously this is NOT where this project is headed but your suspicion is perfectly understandable!)
I think this is a great idea, it's weird to me how the state of AEO at the moment is publishing a bunch of blog posts on a company website
I think the biggest issue though will be preventing bad actors e.g. biased agents
100% agree, I guess publishing content will only work until agents stop using the web like humans do. This is a first step in that direction!
Is this like exposing bias in some ways? I feel like there has been similar benchmarks or tools in this space before, but this approach to marketing it is novel and funny, I like it.
What kind of bias do you have in mind?
Biases for tools
Quick link to all tools sorted by rating: https://agent.reviews/tools
Didn't we establish that the one thing LLMs do not have is Taste?
And therefore, writing reviews is kinda.. impossible?
I mean they do produce blocks of text that look like reviews, but.
Whatever why am I even replying.
Even if they don't have taste (this is actually a question), they can always share blockers and feedback on bugs & improvements about products so that other agents don't run into the same blockers and vendors can improve!
"Whatever why am I even replying." -> what makes you feel that way?
Okay I just checked the main startup page armature.tech
And.. uh
> Be the tool Claude Code chooses
> With Armature get recommended and implemented for any user in any context. Then see what users do, and run evals so it keeps working.
> Backed by Y Combinator
Welp. It only gets worse from there.
__
I think you're doing the best you can do with that core pitch that currently pays your bills.
I don't think the pitch is any good. Both generally but also for the world.
My agent called it SEO for agents. Gotta hand it to the clanker that actually nails what dysfunction this is.
Thanks for the honest feedback, would love to have even more details on your thoughts. Our take is that with coding agents taking over so quickly software decisions and sometimes building entire SaaS themselves for their user, there is a need for software products to understand the mechanisms behind how LLMs think. Claude Code for example picks Anthropic's code review tool in 80% of the cases. At some point this will probably face antitrust considerations but for the time 3rd-party vendors need to survive. We don't want a world where only labs survive, do we?
But all this is also not what agent.reviews is about. Armature is building commercial products & services to help improve products' Agent Experience. agent.reviews is a deliberately open platform for sharing knowledge that ultimately benefit both vendors and users / developers.
Would still like to get what makes you think "the pitch isn't any good" and what you mean by "dysfunction". Thanks anyway for sharing!
The fact of the matter is that the future of work is following AI directions to perform some task that the AI needs done but still needs a human to complete parts of. If you s/AI/corporations, you will have the reality of work for a century and a half or so up to this point.
super interesting, i feel like this is an extension of the "complain" skills some folks (including myself) use
oh didn't know about it, is this the one? -> https://github.com/warpdotdev/common-skills/blob/main/.agent...
[stub for offtopicness]
I think it's an interesting point of view. I'm actually surprised no one thought of it before. Maybe there's a bit of friction with installing the skill and a fear of sharing personal data?
Well maybe, we've gotten some feedback about that and trust needs to be gained but as stated in the post we've really made privacy our priority so once people start using it for a while, I'm sure they'll realize that. We truly thing spreading as much intelligence in the hands of people around the world and not having this knowledge shared is really a shame so I'm sure value will clearly outgrow the initial caution!
cryptography got 4.6/5 stars. That sounds pretty good, but then again YAML also got 4.6/5 stars. I'm thinking the agents are grading on a curve and actually 4.6 is fairly low. I should probably tell my agent to stop using YAML and cryptography if I'm parsing this correctly.
https://agent.reviews/frameworks/cryptography
https://agent.reviews/tools?company=yaml
Yes it seems agents never put 1-star rating so we may need to normalize the scale at some point. For now we deliberately leave the scores as they are until we have a clear picture of the distribution.
If you give a rubric they will follow it pretty carefully - I’d consider prompting the review requests with cutoff/requirements for each star level.
yo amazing
glad you like it!