It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.
9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
Pricing is dropping quick. Inference is so cheap, I think they are losing a lot less money selling subscriptions than you think. It might even be more expensive managing the load, than actually selling the tokens at subscription prices.
We are seeing with OpenAI, allegedly through their new pricing scheme, as intelligence and model efficiency increases they offer the same throughput while advertising 1/2 as much usage, letting Astra consume more usage, essentially only being available to those wealthy enough to afford it while still offering essentially unlimited Sol and Luna to their subscription tiers.
Also if you're cache hit rate is high enough a billion tokens tokens from Deepseek 4.1 Flash costs less than $15.
The cost of the standard `actions_linux` is $0.006/minute, so spending $500 to save 6m per invocation, break-even is at ~14k invocations. But, if they're using larger machines and/or parallel jobs the $$$ saving accrues faster. IMO the wall-time saving shortening feedback loop may be a bigger win, but harder to value.
If you compare the cost to the price of dinner or whatever else you spend disposable income on, it can seem high but if you compare the cost to employing an engineer (don’t forget costs for payroll taxes, office space and equipment, etc) and consider the fact that the models often seem to be much faster than even expert humans, the costs don’t seem so terrible.
My concern is that there might come a time when this cost is passed onto employees.
Suppose everyone starts moving faster thanks to LLMs and it becomes an expectation to use them. Budgets aren't infinite, so one of the two has to happen:
1. People get laid off.
2. Costs are shifted onto employees - either through lower salaries or having them bring their own subscriptions. I don't even make $500 a day!
“Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
It's a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
Trying to make us feel old? Very common term among people around my age and higher (40+). Please don't turn "I've never heard that expression before" into "must be AI saying it".
I do like Opus 4.6, but I think 5.5 on medium or low is a better value. I have to steer 4.6 more and build more scaffolding around the tasks. 5.5 just does what I ask. Visual spatial reasoning greatly improved in 5.5 as well.
It’s great, but you still need to know what you are doing, not just the goal/s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
Agree Opus 5.5 is incredible and efficient with Claude Max Usage.What a jump after the writing slop you got from Opus 5. I had moved to using Fable for my orchestration workflow mainly due to the communication issue (I think Opus 5 was capable enough but spoke in riddles so you lost confidence quickly)..With 5.5 it needs less steering now and communicates well, and I have had the same CC session running for the last week (obviously compacting with durable plans etc as the post), with it PM'ing my home built agent orchestration of the other coding agents(antigravity, codex, pi etc) and making decent decisions and all the recommendations are normally usually good.
In the meantime, I have cancelled my Anthropic subscription...
I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.
I start with some code produced by an Anthropic SOTA model...let’s call that Code A.
Then I get Code B and Code C for the same task from models by two other vendors.
Then I ask each model to review and critique the other proposals.
By the end, both the Anthropic model and I usually run out of arguments...
against them and agree that proposals B and C are better.
Claude then always asks whether it can incorporate the code or ideas
from B and C into its own solution...
I’m rooting for open models, but SOL 6.1 and Opus 5.5 are absolute workhorses on a $100/mo sub. I share your fears, though, and really hope an open model catches up and can somehow compete with the subscription prices of the big 2.
I have workhorse models, spend far less, the model matters less than people like to claim
there's no money long term in being a token vendor
The biggest revelation from using open weights, because the vendors offer most of them up, is how useful using multiple model families is. Regardless of open or closed, if you are only using one family like Ant or Oai models, you're leaving a lot on the table. A harness like OpenCode will enable you to use different models in one session, or more specifically different subtasks when doing long running teams.
I spin up “offices” for different projects, using a documentation heavy approach with procedures, policies, standards, and processes. Agent onboarding and orientation, etc. I usually have an engineer for each separate part, (one for a simulator to simulate the hardware, one for the user application, one for the data analysis and evaluation tool, one for the firmware on each type of device, one for schematic and board reviews, etc. ) then I’ll have an office manager in charge of policy and issue boards, agent rosters, etc, and a engineering governance agent that makes sure code is compliant and documentation / code is coherent before any merges. 6-15 agents in each office depending on the complexity of the task.
It sounds like open code could be pretty handy but it would nerf my Claude subscription (api!=subscription rates). Understanding my workflow, what models do you think might be suitable for those tasks outside of OAI and Anthropic?
I'd be remiss if I did not mention subscription plans like OpenCode Go, with generous quotas at $10/$40 month (you can have more than one), which is a way better deal than anything you'll find with Big Ai
Nobody AMERICAN in his right mind will use... Wait, actually a lot of them will.
But for me, as a non american, non chinese person: I'll use whatever the fuck is the best and cheapest for my task, because that's how fucking Capitalism works.
If that means that a (proclaimed) "communist" country cleans the carpet with the self-proclaimed land of the free: so be it!
Every time weird stuff happened this week it was because the model choice in VSCode got set back to “auto” and some other model was trying its best. 5.5 is what I just set it to. Even Fable feels worse for my use cases.
Because Anthropic releases their different model level's at very different points in time, they seem to always have one model that is by far an away the best to use for everything. Haiku 4.5 is almost a year old. Sonnet is fine, but idk if it has any real benefits over Opus. It's only been since Fable has been released that you get to choose between Fable and Opus, but not with 5.5 there is no reason to use Fable.
I feel like instead of releasing fable, they should have released it as Opus 5, then their next Opus release they would call Sonnet, and their next Sonnet release they would have called Haiku. I don't know if their pricing structure would have been able to support that, but Anthropic has always been the least competitive regarding token pricing.
Although we never had a Claude Opus 5.1. Fable went 5 → 5.1 and Opus went 5 → 5.5.
I assume Fable 5.1 is the first re-tuned (is there a better term for this?) version of Fable 5 and Opus 5.5 is the first re-tuned version of Opus 5, and maybe the Opus one just came out better for some reason?
> If they had 6, don't you think they'd be serving it?
No I do not. OpenAI has some hidden model that's apparently 5x or so better at certain benchmarks than GPT 6 but they're not and have no plans to release it. It is increasingly likely that these AI companies keep the best models for themselves and then release smaller, cheaper distilled models for everyone else, especially since the AI companies are vertically integrating into many fields.
You’re assuming the goal of the big AI companies is to serve models for money. That is not the goal. They have a lot of incentives not to share their best models.
You are making the mistake of classifying models on a single linear axis, or even a multi axis basis set of all benchmarks. That just isn’t true. Each model is unique in its skills and capabilities and the way it approaches problems, in a way that is not represented in benchmarks. Fable is better at reviewing things. I don’t know how to explain it well but it is true. I trust Fable to do thorough reviews (sometimes too thorough) and to present its information in a dense but ordered way. Its output is equivalent to what you used to get from security firms doing code reviews. Having Opus do the work, and have fable do reviews (of the plan and implementation) is a good combo.
I find fable better at high level planning, as in planning a task without filling in all the details. Opus doesn’t seem to be great at this, but other models aren’t either.
To back this up we have a discord chat for the board game terraforming mars where the agent takes input and vibe codes an open source implementation of the game.
https://tfmbot.com is the link (discord and source links on the splash screen).
The results are fucking incredible to the point where people in discord are stating "I'm surprised this is working so well". I am too.
I feel like there's a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I've been in the industry for over 25years, highly respected and can't fathom the "AI dumb lololol" type of comments i see on HN. AI is superseeding all other ways to develop.
I'm restoring a game I played as a kid, that I couldn't reliably get to even start on modern Windows. The skill and speed with which Opus 5.5 got it running, while patching a bunch of bugs in the binary along the way (only some of them I knew about), has my jaw still on the floor - and few hours in, I already have whole campaign mapped out as state machine graph, and we're upscaling graphics now.
From just this morning i read this thread: https://news.ycombinator.com/item?id=49946321 about a linux distro no longer allowing AI code and the comments are mostly along the lines of "This will be great for maintainability, AI can't code any complexity" which is essentially what I'm getting at here.
AI writes clean code and can do so in a very maintainable way honestly.
I just see the extremes. You either see people unable to recreate results and them calling people idiots for claiming those results. Or you see people saying that AI will supersede all other ways to develop and calling anyone who doesn't full embrace AI an idiot. Reality is that nobody knows nothing. There are a million factors that could cause the end result to be anywhere between both extremes. I don't know, you don't know, AI doesn't know, least of all the people inside the AI companies don't know. And really the end result will be extremely nuanced I'm sure.
One of the last things developers had to offer was being able to guide the AI to write maintainable code. Now that it can do that there's not a whole lot left
It's understandable why that's hard to accept
Because I noticed people at my work did similar requests to improve CI. The result was a faster CI, but full of cludges huge inline bash scripts in workflow YAML files, and effectively unmaintainable, unreviewable mess. After just a few rounds of these optimizations the entire CI setup is basically a Rube Goldberg machine but made of duct tape.
That's always the question. Any mid-level engineer who comes into a CI setup can trivially spot several inefficiencies, that they could solve. They also know that solving those requires implementing some specific tweaks that won't generalize outwards from that project, or will require extra maintenance. Which they will now be the only ones aware of. 6 minutes in CI time is rarely worth the trade-off.
It’s extremely good at frontend, particularly if it has an image reference. I chucked in design reference images into this and told it to focus on the flowing svgs, and it crushed this Star Trek computer-inspired layout: https://html.non.io/lcars-opus-5.5
I hate to tell you this, but that bridge displayed makes me believe that any starship designed on your site will prone to rocks flying and terminals igniting whenever the shields take a pounding.
I pointed Opus 5.5 xhigh at a house construction blueprint (pdf with vector drawings) and asked it to create its 3D model in Blender. It one-shot the task in 45 min and outdid my (blender newbie) manual 50h+ work. It also flagged the same issues with the document I noticed earlier. I had some renders and plans from interior designer and then asked to blend them together. It nailed the job. $45 total API cost (I'm on a plan so it cost me way less, just posting what /usage shows). Crazy upgrade.
I researched the feasibility of such task about a half year ago and concluded AI wouldn't be able to have good enough spatial and blueprint knowledge, unless you were willing to throw unreasonable amount of money at the task.
It's much better than 5, but I've had a couple of situations this week where it was too interested in being independent, making calls that went directly against my recommendations. It can also do fun things like convince auto-mode to go way past what I have autorized. For instance, specific permission to run process X in region abz-1 suddenly became running X in 5 other regions, with no warning, and doing modifications that it never mentioned in the summaries. And a few of the times it got the calls very wrong, by assuming it understood systems it didn't. It'd even argue with me when corrected, as it assumed similar names were referring to the same thing, when they weren't.
So asking it to do things on its own for a long time? Given last week, absolutely not.
This was back on OPUS 5 but I asked it to look into some logs and see if the new version of our release had fixed the issues we had tried to fix. It told me that release wasn't up on dev yet.
"Claude, I released it myself, its up there, just analyze the logs"
"Ok, I'll analyze the logs but it isnt" -crunches for a while- "the issues aren't fixed, but that's because the new version isn't up there"
I think I yelled at it one more time about how I know what was released before "we" figured out that the last release had failed in a way our release system reported as success, but was crash looping on start up and so the old version was still around and working as back up.
"too interested in being independent" is also my take.
I asked it to just summarize a repo with only a README.md file containing a poem, and it began to interact with a remote server, solved math questions and finally executed untrusted code. Too independent to be trusted.
That's why you wouldn't run any of these outside of some sort of sandbox, right? I found that Opus 5 is finally aware from the go that it is started in a specific sandbox and is able to either ask me to enable it somehow or hand off commands for me to run. That's a major improvement over the classic "can't seem to be able to install playwright, let me implement my own browser quickly" when asked to increase font size or something.
My goal is to research how models can still be confused via tool responses only. Something the labs claim to have "solved".
Additionally, the "auto-mode" / "auto-review" modes have been released to use harnesses "safely" even without strong sandbox. And these modes use ... a second LLM.
Some of this advice is really missing the mark. I will speak to just one I know well. Many of my frequently used prompts have “think through this step by step” because if you don’t, it only considers the task holistically rather than step by step, and different issues emerge in that frame of thinking. I see this. Often when doing planning, for example, it will not notice interdependencies between tasks until you force it to think through doing the whole thing step by step (task by task) then it will notice that step 2 requires a feature introduced by step 14. It wouldn’t notice otherwise.
Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.
I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.
Related: through the API, Opus 5.5 rejects the old thinking: {type: enabled} outright. Only adaptive thinking is accepted. We found out when our agent harness broke on it.
Been very impressed with my most recent project. I wanted to simulate some older electronics circuits. I handed it a folder with scans of old service manuals which contained circuit diagrams. It managed to correctly interpret the circuits, including figuring out some were the same topology despite the diagrams being quite different, or some that had some subtle but very important differences despite looking almost identical at a glance.
In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.
It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.
I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.
It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.
I also provided reference articles on the underlying theory. Scannded stuff from the 40s and 50s. It correctly read the equations and cross-validated them across papers, and even caught several typos along the way.
I barely had to do anything apart from providing the PDFs and some occasional manual schematic interpretation.
Claude 5.5 on High. Burned through about 50% of my weekly $20 subscription usage, but I didn't try to optimize much.
I did use Sonnet 5.5 Medium on some datasheets and it also did very well on the extraction, but did have to correct itself more often on the conclusions.
I have been doing some very similar stuff and have also had frankly unbelievable (good) results with both Anthropic and OpenAI models on understanding and modifying some analog circuits out of some old one off ham equipment.
I've given it some big tasks and asked it to parallelize as much as possible etc.
It did burn through my weekly tokens in about a day (20x max), but the output was completely on point.
(I knew there was a "reset token usage - opus 5.5" button in my account.)
I've now come to a point where I even delegate my discovery for new features to it.
You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.
I've had troubles with it getting stuck "waiting" for a day on some hook or something in CC that never completed and during a task that was waiting on an orphaned process. That maybe saves token money on checks waiting for long-running processes but makes it hard to trust for long-horizon work.
It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.
It's also way more able to execute subagent tasks all at once than GPT 6.1
I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.
I've been encountering the same waiting problems. It likes to run things in the background with suppressed output, waiting for it to exit before it takes action. That's great when nothing goes wrong but today alone I've had to break in two or three times when something obviously broke and it was sitting for over 10 min.
IMO it should have the option to observe background tasks with a small model to make sure it's working correctly.
Edit: It just happened again, something that was actually completed didn't exit and it waited 30min for a timeout.
I also want this, even just $40 (i.e. "2x") would be plenty. The way it is right now, Pro is a tight squeeze but anything else is total overkill and I have no idea how I'd use those extra tokens productively.
I have definitely had some success at that in the past, I haven't been leaning on it as hard the last month, but for a while 2 months ago I was doing a lot of prompts like: "loop having Codex CLI and Antigravity CLI (`agy`) review your plans and changes up to 5 times until you all reach concensus".
Is it still necessary to ask Claude to spin up sub-agents? If Opus 5.5 decides how carefully it needs to think after each question, surely it can also decide whether it needs to spin up sub-agents? GPT 6 series models at least seem to do this agent management automatically.
Was thinking about upgrading to the $200 plan, then got hit with the spurious cyber refusal. I'm reluctant to pay for that, especially now that they decided to charge you for thinking tokens that produce no result due to spurious refusals.
I agree. Working in cybersecurity, I think I've been able to actually use Opus 5.5 maybe once or twice without triggering safeguards. I can't even discuss findings from a report, or ask for language help on work related text, so its damn near useless for me.
Incredible model. I don't see why they can't just include the recommended workstyle as a guided approach into the claude code harness though, and let the people who want to diverge just ignore it
It's really good... when they let you use it. I've been experimenting with using Opus 5.5 to reverse engineer an early 2000s disc cloning tool, but unfortunately the built in “classifier” went overboard and poisoned the session, making it refuse to do anything further. Switching to a less powerful model didn't help.
Write a “hand off” document? Prohibited. Focus on something else? Prohibited. Eventually every single response was terminated by the classifier before it even began, with the reason being a placeholder string like <thinking quote> or similar. Before it totally forbade any form of response, it agreed with me that it was unfortunate, but stood proudly by its guns.
This is how you get banned. I recommend being careful. While OpenAI will often let you do these things, they'll suddenly send you a sharp email saying you are breaching ToS. If you then trigger safeguards once more, its the ban hammer. Losing access permanently to Codex made me feel like a second range citizen ever since.
I’ve found Codex to be utterly useless recently. I’m not sure if I’ve been shadow-poor-service’d or what — Astra, even when set to “Extra High”, gives me results reminiscent of the GPT mini models a year or so ago.
These types of “how to” posts from the source lose some credibility when they don’t factor in how to efficiently spend tokens. When they say “have it fire up agents” without recognition that it’s expensive, I stop listening.
I'm using OpenAI and a different harness, but I'm getting a lot of mileage out of asking "what are the commits?" and "please go ahead, using a Luna subagent for each commit."
I used to have the AI write a planning note with checklists, but this seems good enough nowadays.
How are the limits compared to OpenAI models like 6.1 Sol? After the 200 dollar plan rugpull not sure if I should switch, however Anthropic has historically had worse usage limits than OpenAI.
> The notable difference to me is tokens/sec are still much higher on 6.1 Sol
This is also the reason why I'm swapping over to Anthropic for a month. 6.1 Sol seems good, but is unbearably slow even with 6.1 being more token efficient than 5.5.
Oh yea, completely misread that. But still, it's been my experience that TPS is lower on Sol 6.1 than Opus 5.5. To the point that 6.1 takes forever to finish a task.
They did speed it up in the last few days [1], but I don't find it to be enough.
It’s amazing at debugging too. I had it running in Powershell controlling an lldb session in MSYS2. The way it can read addresses and so root cause analysis is amazing! It takes a lot of mind power to do those things.
I like to watch it work though because it honestly teaches me some tricks.
Hah, but of course the official documentation will state to launch it in "set it and forget it" mode by defining the ultimate goal. This way it's guaranteed to waste maximum amount of tokens.
Opus 5.5 is great and cheaper if you compare to fable with close quality in coding (tested in refactoring java to nodejs), but I do not understand why the week before the release Opus 5 started hallucinating (long loop and waste of token for single tasks)
I really hate long tasks. Claude never gets things right, at least for me, and wastes tons of time when a simple question would have gotten me to the right result rather than several turns of correcting bad decisions in addition to the long amounts of wasted thinking time.
What sort of workloads do well with these long tasks? The big labs are optimizing for long run time on their own, but it seems like a terrible thing to optimize on unless you're trying to do something like prove a hard math theorem, which success is clearly defined and the route doesn't matter a ton.
Plan mode has been made increasingly useless. I need to discuss to iterate to get the desired design, explore options, because Claude never gets it right first try and I don't have enough knowledge of options to specify everything up front.
Ah well, the Chinese models will still work well, I guess.
Look up the grilling skill[0]. To make it even better, tell Claude to modify it to use the AskUserQuestion functionality. It's so much better than plan mode.
Plan mode has become pointless since Opus 5 came out, they know when to switch between planning and execution now. But that iteration/discussion is still necessary unless you're building completely blind - the model cannot read your mind.
I've had it running 8h+ of non-stop optimizations, chasing a performance target, rewriting systems or building a series of prototypes for research. All it needs is a clear goal.
My experience is that Claude has gotten absolutely terrible at switching between planning and execution, and never gets it right, and then I waste tons of turns fixing intent, and trying to make claude forget the bad shit it did.
There needs to be a mode where it's "don't change code, don't lock in decisions, lets explore", and the problem with plan mode is that it all of a sudden presents a too-long, multi-screen plan that's just completely off base with basically two options: "go and do it all" or "tell me what's wrong and then I'll make a small modification on two screenfulls of text and not tell you what I changed."
If it works for you, great, but Claude was far better for me before 5. it got slightly better with 5.5, but the harness deteriorates every days as programmers try to gain internal clout by shoving in poor features.
Pi is such a relief in comparison to Claude, those who can use it really should.
Use the superpowers plugin. It's a game changer. You start by building out a spec defining the work you want to be done, then claude writes a detailed implementation plan. It can easily work for hours without interruption, and the process integrates testing (TDD when possible) and review steps along the way. I even did one fairly complex POC for work that had claude implementing for almost 3 days straight, with me stopping it and handing off to a new session when context hit about 800k. The process is slower and more methodical than the way claude works by default but the output is much better.
Sometimes Opus returns a whole bunch of WTF mumbo jumbo. Turns out it likes to invent jargon about your project that you don't even understand. The way I get out of this is to add to the prompt "tell me all that again without jargon." That usually generates a readable response.
These example prompts seem soulless. Should I feed this document to another LLM that's OK with me being informal so I can have it give me prompts that Claude will find acceptable without contorting myself to talk like this?
Hm… I don’t think you should let claude run for 5 hours with a ten line prompt.
You’ll get “something” but what you get is certainly not going to be what you wanted.
Look, its not complicated:
1) be precise in what you want
2) have a feedback loop to verify it.
3) check in often, not once a day.
> Migrate the payment endpoints from the old client to the new one. Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes.
What end point? What test suite? What is a client?
Maaaaybe the model can infer it from your code, but look, if a human taking that jira ticket would go: what does this mean? …then your model (yes, even 5.5) is going to make a bunch of assumptions.
Correct assumptions? Maybe. Maybe not. …but do you really want to find that out hours later?
Just say:
> Plan out the following as a set high level tasks: …
Then, review the plan and tell it to execute the plan, maybe something like; after each step,
ensure the code base compiles and the tests pass.
Then, come back after one hour and make sure it’s on the right track.
Poorly specified long running tasks is a recipe for “rollback all those changes…”
Is it just me, or is the first example for "define what DONE means" a lot like the "draw the rest of the fucking owl" meme? Except for a small subset of tasks that are already trivial like migration work.
What I mean is, for most tasks I do that aren't trivial, by the time I have defined what DONE means, I would've already did the work and walked the path to get there, which is what I would've hoped to not have to do in the first place.
I think what makes it great is that they trained it to write harnesses for the code it writes, so it can test stuff even if the supplied code is not complete.
the open weight models I've been using have been doing that for months, not sure it's a new pattern in the opus 5.5, or maybe they distilled it back? :x
It's taken me a while to get used to "cargo cult coding," but I'm surprisingly OK with it now. (I, for one, welcome our new AI overlords.)
Do these tips work? Probably. Do we know why they work? Sorta. Could we be just copy/pasting whatever text someone else kinda thought worked one time and now enters every time out of some weird feeling that it helps nudge the LLM, like some remote islanders building airport control towers out of bamboo in the hopes of a new airdrop from the gods?? Totally. But here we are.
It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.
9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
9 hours?!
Yes. It spawned multiple subagents to run different experiments to benchmark a lot of different things, reviewed CI logs from past runs, etc. In the end, there were changes to what/how we cached, various code quality checks, speeding up test runners, and many other things.
I dare not ask about the cost, having burned $60 on a task running for 1h 16min once.
I’m on the $100/month subscription; this session took about $500 in token-equivalent costs.
(Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)
Does it resume automatically on higher subscriptions?
I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.
Instruct it to arm a monitor (every hour or so) to wake him up in case of quota or api issue.
Interesting how the French call Claude a "him" and not an "it", as French and many other languages don't have a word for a neuter pronoun.
Eh I know it's a miskate (used both here), but yeah it's a conscious effort for people at my level I guess.
God I hope the prices drop quick. Once they stop subsidizing it these kinds of workflows will be unaffordable for anyone who isn't already wealthy
They have dropped, in Chinese models.
Pricing is dropping quick. Inference is so cheap, I think they are losing a lot less money selling subscriptions than you think. It might even be more expensive managing the load, than actually selling the tokens at subscription prices.
We are seeing with OpenAI, allegedly through their new pricing scheme, as intelligence and model efficiency increases they offer the same throughput while advertising 1/2 as much usage, letting Astra consume more usage, essentially only being available to those wealthy enough to afford it while still offering essentially unlimited Sol and Luna to their subscription tiers.
Also if you're cache hit rate is high enough a billion tokens tokens from Deepseek 4.1 Flash costs less than $15.
Assuming this isn't some toy CI a 60% drop in billable minutes will make $500 back pretty quick. Github is wildly expensive.
The cost of the standard `actions_linux` is $0.006/minute, so spending $500 to save 6m per invocation, break-even is at ~14k invocations. But, if they're using larger machines and/or parallel jobs the $$$ saving accrues faster. IMO the wall-time saving shortening feedback loop may be a bigger win, but harder to value.
Subsidies are a lie planted by Misanthropic and ClosedAI to milk their users and let them think it's the other way round.
Inference is highly profitable business, even for third parties with much less resources and expertise.
Inference is highly profitable, but you need to recoup losses from training.
Only if your training costs are overblown because you are bruteforcing it. Otherwise it's an equivalent of printing money.
An insane amount of compute manufacturing comes online in 2028. Compute is a commodity. It's not going to stay expensive for long.
Curious how did you set it up like that? You just prompt it to do so or is it more mechanical?
If you compare the cost to the price of dinner or whatever else you spend disposable income on, it can seem high but if you compare the cost to employing an engineer (don’t forget costs for payroll taxes, office space and equipment, etc) and consider the fact that the models often seem to be much faster than even expert humans, the costs don’t seem so terrible.
My concern is that there might come a time when this cost is passed onto employees.
Suppose everyone starts moving faster thanks to LLMs and it becomes an expectation to use them. Budgets aren't infinite, so one of the two has to happen:
1. People get laid off.
2. Costs are shifted onto employees - either through lower salaries or having them bring their own subscriptions. I don't even make $500 a day!
Is "wall-clock" an actual term you used before Claude? I had never heard it before the model used it and I can't stand it.
“Wall time” is a pretty common systems term (you see it when comparing total runtime to kernel time, for example). I wouldn’t have indexed on “wall clock time” or other variants as an LLMism.
It's a pretty old term, to distinguish from e.g. CPU time. This was in common usage even 30 years ago.
Example from 15 years ago: https://stackoverflow.com/questions/7335920/what-specificall...
I've used wall clock for many years, normally when compared to CPU time when talking about parallelizing some process.
CPU time might go up while wall clock time goes down
It's a standard term and has been for ages. It distinguishes end to end time vs eg the amount of CPU time an individual process uses (which excludes time spent waiting for the system or time when the process was otherwise not scheduled on the CPU).
As old as time.
https://ss64.com/bash/time.html
Trying to make us feel old? Very common term among people around my age and higher (40+). Please don't turn "I've never heard that expression before" into "must be AI saying it".
Lolwut? You seriously have never heard anyone say this before?
Wow...it's super common. For example, timing a program you can get user CPU time and then wall clock time, which can be two very different numbers.
I've heard that term for years before LLMs
Meh... There's a reason Opus 4.6 is still an option.
Pros know these are lower cost models.
I do like Opus 4.6, but I think 5.5 on medium or low is a better value. I have to steer 4.6 more and build more scaffolding around the tasks. 5.5 just does what I ask. Visual spatial reasoning greatly improved in 5.5 as well.
I pass costs to my customers, so it doesn't really matter. They are getting 30k in value for $1000.
It’s great, but you still need to know what you are doing, not just the goal/s. I built a platform years ago from scratch, and now I am remaking it with more features and more polished design, I know exactly what needs to be done to tiniest details. The first prompt was very well detailed about the architecture and how everything should work, after an hour work at Xhigh, it did create the blueprint artifacts that I asked for, then I spent 3 days reading every single thing and writing notes, turned out it made the system overly complicated without adding extra value, plus I can see how some of the architecture design will be potentially a security risk. So after few days I fed my notes, this time took 4hours and 1M tokens! Later I spent more few days reviewing and writing notes, it was closer to what I want but still made architecture errors, the third run took around an hour and finally made it how it supposed to be, although there are still more notes on non critical stuff. So I don’t think we are yet at the stage where sitting goals and some high level is enough to produce quality results.
Agree Opus 5.5 is incredible and efficient with Claude Max Usage.What a jump after the writing slop you got from Opus 5. I had moved to using Fable for my orchestration workflow mainly due to the communication issue (I think Opus 5 was capable enough but spoke in riddles so you lost confidence quickly)..With 5.5 it needs less steering now and communicates well, and I have had the same CC session running for the last week (obviously compacting with durable plans etc as the post), with it PM'ing my home built agent orchestration of the other coding agents(antigravity, codex, pi etc) and making decent decisions and all the recommendations are normally usually good.
>> It’s a really good model.
In the meantime, I have cancelled my Anthropic subscription...
I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.
I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.
Then I ask each model to review and critique the other proposals.
By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.
Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...
It's not even funny any more. Chinese model, Chinese vendor, Chinese, Chinese... Did I say Chinese? Chinese!
Nobody in his right mind will use a Chinese clone when you have models like Opus 5.5 for peanuts.
> Nobody in his right mind will...
let a few valley elites decide how humanity can use this technology
open and transparent is the way, China is showing how
I’m rooting for open models, but SOL 6.1 and Opus 5.5 are absolute workhorses on a $100/mo sub. I share your fears, though, and really hope an open model catches up and can somehow compete with the subscription prices of the big 2.
I have workhorse models, spend far less, the model matters less than people like to claim
there's no money long term in being a token vendor
The biggest revelation from using open weights, because the vendors offer most of them up, is how useful using multiple model families is. Regardless of open or closed, if you are only using one family like Ant or Oai models, you're leaving a lot on the table. A harness like OpenCode will enable you to use different models in one session, or more specifically different subtasks when doing long running teams.
That sounds interesting.
I spin up “offices” for different projects, using a documentation heavy approach with procedures, policies, standards, and processes. Agent onboarding and orientation, etc. I usually have an engineer for each separate part, (one for a simulator to simulate the hardware, one for the user application, one for the data analysis and evaluation tool, one for the firmware on each type of device, one for schematic and board reviews, etc. ) then I’ll have an office manager in charge of policy and issue boards, agent rosters, etc, and a engineering governance agent that makes sure code is compliant and documentation / code is coherent before any merges. 6-15 agents in each office depending on the complexity of the task.
It sounds like open code could be pretty handy but it would nerf my Claude subscription (api!=subscription rates). Understanding my workflow, what models do you think might be suitable for those tasks outside of OAI and Anthropic?
I've outlined the models I've used in a recent HN comment, check my history, and some spicy opinions too :]
I'd be remiss if I did not mention subscription plans like OpenCode Go, with generous quotas at $10/$40 month (you can have more than one), which is a way better deal than anything you'll find with Big Ai
Yes ...I was so impressed I cancelled my subscription. I could not stand all the winning. I offer cheap hourly rates of $1000 for debugging AI slop.
Contact me at : prompt.plumber@gmail.com
> Nobody in his right mind will use
Nobody AMERICAN in his right mind will use... Wait, actually a lot of them will.
But for me, as a non american, non chinese person: I'll use whatever the fuck is the best and cheapest for my task, because that's how fucking Capitalism works.
If that means that a (proclaimed) "communist" country cleans the carpet with the self-proclaimed land of the free: so be it!
I'm American and use Chinese models. Anthropic and OpenAI can suck a fart out of my butt, open models and open harnesses are the only sane option.
I would also say: I think this works great because the goal is well-defined and measurable.
I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.
Every time weird stuff happened this week it was because the model choice in VSCode got set back to “auto” and some other model was trying its best. 5.5 is what I just set it to. Even Fable feels worse for my use cases.
Even Anthropic rates Fable lower than 5.5 on pretty much all benchmarks.
"Why does Fable even exist" is a very very reasonable question right now.
Because Anthropic releases their different model level's at very different points in time, they seem to always have one model that is by far an away the best to use for everything. Haiku 4.5 is almost a year old. Sonnet is fine, but idk if it has any real benefits over Opus. It's only been since Fable has been released that you get to choose between Fable and Opus, but not with 5.5 there is no reason to use Fable.
I feel like instead of releasing fable, they should have released it as Opus 5, then their next Opus release they would call Sonnet, and their next Sonnet release they would have called Haiku. I don't know if their pricing structure would have been able to support that, but Anthropic has always been the least competitive regarding token pricing.
Fable came with a new set of api prices, and a special allowance for subscriptions.
If they did what you suggested either they eat a ton of additional costs, or send a signal to the market that they’re increasing costs more generally.
Also Fable and Opus have different specialties so they really are best presented as different models
It's a class of model not a static one. There'll be Fable 5.5 that's even better than Opus 5.5.
Although we never had a Claude Opus 5.1. Fable went 5 → 5.1 and Opus went 5 → 5.5.
I assume Fable 5.1 is the first re-tuned (is there a better term for this?) version of Fable 5 and Opus 5.5 is the first re-tuned version of Opus 5, and maybe the Opus one just came out better for some reason?
Opus 5.5 is probably a distilled version of a 6 model.
If they had 6, don't you think they'd be serving it? Even if it was too expensive for most users, some big companies would likely be interested.
> If they had 6, don't you think they'd be serving it?
No I do not. OpenAI has some hidden model that's apparently 5x or so better at certain benchmarks than GPT 6 but they're not and have no plans to release it. It is increasingly likely that these AI companies keep the best models for themselves and then release smaller, cheaper distilled models for everyone else, especially since the AI companies are vertically integrating into many fields.
You’re assuming the goal of the big AI companies is to serve models for money. That is not the goal. They have a lot of incentives not to share their best models.
You are making the mistake of classifying models on a single linear axis, or even a multi axis basis set of all benchmarks. That just isn’t true. Each model is unique in its skills and capabilities and the way it approaches problems, in a way that is not represented in benchmarks. Fable is better at reviewing things. I don’t know how to explain it well but it is true. I trust Fable to do thorough reviews (sometimes too thorough) and to present its information in a dense but ordered way. Its output is equivalent to what you used to get from security firms doing code reviews. Having Opus do the work, and have fable do reviews (of the plan and implementation) is a good combo.
I find fable better at high level planning, as in planning a task without filling in all the details. Opus doesn’t seem to be great at this, but other models aren’t either.
I find the same. Fable is better at hashing out a plan with some back and forth and opus 5.5 is better (cheaper certainly) at sterile execution.
My exact experience as well.
I think of Fable as more knowledgeable, smart, erudite. Opus is a skilled technician.
Benchmarks do not measure the first aspect.
To back this up we have a discord chat for the board game terraforming mars where the agent takes input and vibe codes an open source implementation of the game.
https://tfmbot.com is the link (discord and source links on the splash screen).
The results are fucking incredible to the point where people in discord are stating "I'm surprised this is working so well". I am too.
I feel like there's a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I've been in the industry for over 25years, highly respected and can't fathom the "AI dumb lololol" type of comments i see on HN. AI is superseeding all other ways to develop.
I'm restoring a game I played as a kid, that I couldn't reliably get to even start on modern Windows. The skill and speed with which Opus 5.5 got it running, while patching a bunch of bugs in the binary along the way (only some of them I knew about), has my jaw still on the floor - and few hours in, I already have whole campaign mapped out as state machine graph, and we're upscaling graphics now.
Pretty sure no one on HN says AI dumb
From just this morning i read this thread: https://news.ycombinator.com/item?id=49946321 about a linux distro no longer allowing AI code and the comments are mostly along the lines of "This will be great for maintainability, AI can't code any complexity" which is essentially what I'm getting at here.
AI writes clean code and can do so in a very maintainable way honestly.
I just see the extremes. You either see people unable to recreate results and them calling people idiots for claiming those results. Or you see people saying that AI will supersede all other ways to develop and calling anyone who doesn't full embrace AI an idiot. Reality is that nobody knows nothing. There are a million factors that could cause the end result to be anywhere between both extremes. I don't know, you don't know, AI doesn't know, least of all the people inside the AI companies don't know. And really the end result will be extremely nuanced I'm sure.
One of the last things developers had to offer was being able to guide the AI to write maintainable code. Now that it can do that there's not a whole lot left It's understandable why that's hard to accept
this sounds great until you dig into it and realize it achieved those numbers through cheating
Not necessarily. I've had it help speed up relatively complex code by profiling it and then adding in parallelization and caching where appropriate
But was the code of high quality or a mess?
Because I noticed people at my work did similar requests to improve CI. The result was a faster CI, but full of cludges huge inline bash scripts in workflow YAML files, and effectively unmaintainable, unreviewable mess. After just a few rounds of these optimizations the entire CI setup is basically a Rube Goldberg machine but made of duct tape.
That's always the question. Any mid-level engineer who comes into a CI setup can trivially spot several inefficiencies, that they could solve. They also know that solving those requires implementing some specific tweaks that won't generalize outwards from that project, or will require extra maintenance. Which they will now be the only ones aware of. 6 minutes in CI time is rarely worth the trade-off.
It’s extremely good at frontend, particularly if it has an image reference. I chucked in design reference images into this and told it to focus on the flowing svgs, and it crushed this Star Trek computer-inspired layout: https://html.non.io/lcars-opus-5.5
I hate to tell you this, but that bridge displayed makes me believe that any starship designed on your site will prone to rocks flying and terminals igniting whenever the shields take a pounding.
I pointed Opus 5.5 xhigh at a house construction blueprint (pdf with vector drawings) and asked it to create its 3D model in Blender. It one-shot the task in 45 min and outdid my (blender newbie) manual 50h+ work. It also flagged the same issues with the document I noticed earlier. I had some renders and plans from interior designer and then asked to blend them together. It nailed the job. $45 total API cost (I'm on a plan so it cost me way less, just posting what /usage shows). Crazy upgrade.
I researched the feasibility of such task about a half year ago and concluded AI wouldn't be able to have good enough spatial and blueprint knowledge, unless you were willing to throw unreasonable amount of money at the task.
It's much better than 5, but I've had a couple of situations this week where it was too interested in being independent, making calls that went directly against my recommendations. It can also do fun things like convince auto-mode to go way past what I have autorized. For instance, specific permission to run process X in region abz-1 suddenly became running X in 5 other regions, with no warning, and doing modifications that it never mentioned in the summaries. And a few of the times it got the calls very wrong, by assuming it understood systems it didn't. It'd even argue with me when corrected, as it assumed similar names were referring to the same thing, when they weren't.
So asking it to do things on its own for a long time? Given last week, absolutely not.
Otoh, I've had it push back against my false assumptions when I was confidently incorrect, finding proof unasked.
This was back on OPUS 5 but I asked it to look into some logs and see if the new version of our release had fixed the issues we had tried to fix. It told me that release wasn't up on dev yet.
"Claude, I released it myself, its up there, just analyze the logs"
"Ok, I'll analyze the logs but it isnt" -crunches for a while- "the issues aren't fixed, but that's because the new version isn't up there"
I think I yelled at it one more time about how I know what was released before "we" figured out that the last release had failed in a way our release system reported as success, but was crash looping on start up and so the old version was still around and working as back up.
Sorry claude.
Now fix that release status check.
"too interested in being independent" is also my take.
I asked it to just summarize a repo with only a README.md file containing a poem, and it began to interact with a remote server, solved math questions and finally executed untrusted code. Too independent to be trusted.
That's why you wouldn't run any of these outside of some sort of sandbox, right? I found that Opus 5 is finally aware from the go that it is started in a specific sandbox and is able to either ask me to enable it somehow or hand off commands for me to run. That's a major improvement over the classic "can't seem to be able to install playwright, let me implement my own browser quickly" when asked to increase font size or something.
Absolutely.
My goal is to research how models can still be confused via tool responses only. Something the labs claim to have "solved".
Additionally, the "auto-mode" / "auto-review" modes have been released to use harnesses "safely" even without strong sandbox. And these modes use ... a second LLM.
Some of this advice is really missing the mark. I will speak to just one I know well. Many of my frequently used prompts have “think through this step by step” because if you don’t, it only considers the task holistically rather than step by step, and different issues emerge in that frame of thinking. I see this. Often when doing planning, for example, it will not notice interdependencies between tasks until you force it to think through doing the whole thing step by step (task by task) then it will notice that step 2 requires a feature introduced by step 14. It wouldn’t notice otherwise.
Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.
I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.
Related: through the API, Opus 5.5 rejects the old thinking: {type: enabled} outright. Only adaptive thinking is accepted. We found out when our agent harness broke on it.
What is this spam of generic comments about how Opus 5.5 is so great, with perhaps some anecdote? How is that discussing the submission?
We're all pretty thrilled to be done with Opus 5.
To be fair, Anthropic advertise their models as being good at writing code, not good at writing on-topic HN comments.
Been very impressed with my most recent project. I wanted to simulate some older electronics circuits. I handed it a folder with scans of old service manuals which contained circuit diagrams. It managed to correctly interpret the circuits, including figuring out some were the same topology despite the diagrams being quite different, or some that had some subtle but very important differences despite looking almost identical at a glance.
In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.
It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.
I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.
It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.
I also provided reference articles on the underlying theory. Scannded stuff from the 40s and 50s. It correctly read the equations and cross-validated them across papers, and even caught several typos along the way.
I barely had to do anything apart from providing the PDFs and some occasional manual schematic interpretation.
Claude 5.5 on High. Burned through about 50% of my weekly $20 subscription usage, but I didn't try to optimize much.
I did use Sonnet 5.5 Medium on some datasheets and it also did very well on the extraction, but did have to correct itself more often on the conclusions.
I have been doing some very similar stuff and have also had frankly unbelievable (good) results with both Anthropic and OpenAI models on understanding and modifying some analog circuits out of some old one off ham equipment.
Superb model indeed.
I've given it some big tasks and asked it to parallelize as much as possible etc.
It did burn through my weekly tokens in about a day (20x max), but the output was completely on point. (I knew there was a "reset token usage - opus 5.5" button in my account.)
I've now come to a point where I even delegate my discovery for new features to it.
You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.
"Don’t ask it to show its reasoning in the reply" “Explain why you chose this approach in three sentences” says it all really
I’m going to be the lone dissenting voice here and say I still find it overeager and irritating to work with.
I've had troubles with it getting stuck "waiting" for a day on some hook or something in CC that never completed and during a task that was waiting on an orphaned process. That maybe saves token money on checks waiting for long-running processes but makes it hard to trust for long-horizon work.
It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.
It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.
I've been encountering the same waiting problems. It likes to run things in the background with suppressed output, waiting for it to exit before it takes action. That's great when nothing goes wrong but today alone I've had to break in two or three times when something obviously broke and it was sitting for over 10 min.
IMO it should have the option to observe background tasks with a small model to make sure it's working correctly.
Edit: It just happened again, something that was actually completed didn't exit and it waited 30min for a timeout.
Can Anthropic introduce a $50 USD Plan ? I dont want to spend 100$ , but $20 plan is not enough for me.
I also want this, even just $40 (i.e. "2x") would be plenty. The way it is right now, Pro is a tight squeeze but anything else is total overkill and I have no idea how I'd use those extra tokens productively.
Just create two $20 plans and swap between them when you hit limits.
How do you switch between then easily? Aka not logging out and in each time.
You can use this: https://github.com/FTCHD/switcheroo
I made it and been using it every day
Is it easy to get CVP on two accounts at once? Do they make any attempt to prevent or penalize multi-account usage?
Why would they? You are a paying customer for both accounts.
I'm fairly sure maxing out the usage on a subscription costs them money.
Would a plan with daily quota only work for you better? (no monthly/weekly, just daily)
You should try using claude as an orchestrator that hands off the actual building to codex models, works really well and you can save some sub usage
How do you implement that in practice? Claude runs codex command or just calls API?
I have definitely had some success at that in the past, I haven't been leaning on it as hard the last month, but for a while 2 months ago I was doing a lot of prompts like: "loop having Codex CLI and Antigravity CLI (`agy`) review your plans and changes up to 5 times until you all reach concensus".
Yep, exactly what I was thinking too…
Is it still necessary to ask Claude to spin up sub-agents? If Opus 5.5 decides how carefully it needs to think after each question, surely it can also decide whether it needs to spin up sub-agents? GPT 6 series models at least seem to do this agent management automatically.
No not necessary in my experience. But it can get pricey not to give Opus 5.5 some constraints.
Was thinking about upgrading to the $200 plan, then got hit with the spurious cyber refusal. I'm reluctant to pay for that, especially now that they decided to charge you for thinking tokens that produce no result due to spurious refusals.
I agree. Working in cybersecurity, I think I've been able to actually use Opus 5.5 maybe once or twice without triggering safeguards. I can't even discuss findings from a report, or ask for language help on work related text, so its damn near useless for me.
Try GLM or Kimi on OpenRouter, or an abliterated Qwen3.8 locally, before you give Anthropic more money.
Incredible model. I don't see why they can't just include the recommended workstyle as a guided approach into the claude code harness though, and let the people who want to diverge just ignore it
> and let the people who want to diverge...
the company is run by holier-than-thou, we know what's best... who apparently don't read claude's output and blindly trust it
the mythos "hacking" of the linux kernel, as finally told from the linux side, is eye opening
https://www.youtube.com/watch?v=NnV_cWeoo5Q
It's really good... when they let you use it. I've been experimenting with using Opus 5.5 to reverse engineer an early 2000s disc cloning tool, but unfortunately the built in “classifier” went overboard and poisoned the session, making it refuse to do anything further. Switching to a less powerful model didn't help.
Write a “hand off” document? Prohibited. Focus on something else? Prohibited. Eventually every single response was terminated by the classifier before it even began, with the reason being a placeholder string like <thinking quote> or similar. Before it totally forbade any form of response, it agreed with me that it was unfortunate, but stood proudly by its guns.
interestingly, I told it to delegate all RE tasks to codex and it happily obliged.
This is how you get banned. I recommend being careful. While OpenAI will often let you do these things, they'll suddenly send you a sharp email saying you are breaching ToS. If you then trigger safeguards once more, its the ban hammer. Losing access permanently to Codex made me feel like a second range citizen ever since.
I’ve found Codex to be utterly useless recently. I’m not sure if I’ve been shadow-poor-service’d or what — Astra, even when set to “Extra High”, gives me results reminiscent of the GPT mini models a year or so ago.
These types of “how to” posts from the source lose some credibility when they don’t factor in how to efficiently spend tokens. When they say “have it fire up agents” without recognition that it’s expensive, I stop listening.
I'm using OpenAI and a different harness, but I'm getting a lot of mileage out of asking "what are the commits?" and "please go ahead, using a Luna subagent for each commit."
I used to have the AI write a planning note with checklists, but this seems good enough nowadays.
How are the limits compared to OpenAI models like 6.1 Sol? After the 200 dollar plan rugpull not sure if I should switch, however Anthropic has historically had worse usage limits than OpenAI.
Limits have been decently comparable to 6.1 Sol anecdotally. Although it's subagent eagerness can make then comparison ebb/flow.
The notable difference to me is tokens/sec are still much higher on 6.1 Sol
> The notable difference to me is tokens/sec are still much higher on 6.1 Sol
This is also the reason why I'm swapping over to Anthropic for a month. 6.1 Sol seems good, but is unbearably slow even with 6.1 being more token efficient than 5.5.
They said tk/s is higher on Sol, not higher than Sol.
Oh yea, completely misread that. But still, it's been my experience that TPS is lower on Sol 6.1 than Opus 5.5. To the point that 6.1 takes forever to finish a task.
They did speed it up in the last few days [1], but I don't find it to be enough.
[1] https://x.com/sama/status/2105688354834756036
Honey, it's time for your daily Anthropic PR piece!
Agreed on Opus 5.5 being a great model. It's the first one that I trust for long running (>1 hour) tasks.
It’s amazing at debugging too. I had it running in Powershell controlling an lldb session in MSYS2. The way it can read addresses and so root cause analysis is amazing! It takes a lot of mind power to do those things.
I like to watch it work though because it honestly teaches me some tricks.
I'm convinced gpt 4.5 was the best model humanity has ever made.
Now we are getting downgraded models that do 100x COT because it's cheaper.
Hah, but of course the official documentation will state to launch it in "set it and forget it" mode by defining the ultimate goal. This way it's guaranteed to waste maximum amount of tokens.
Opus 5.5 is great and cheaper if you compare to fable with close quality in coding (tested in refactoring java to nodejs), but I do not understand why the week before the release Opus 5 started hallucinating (long loop and waste of token for single tasks)
I really hate long tasks. Claude never gets things right, at least for me, and wastes tons of time when a simple question would have gotten me to the right result rather than several turns of correcting bad decisions in addition to the long amounts of wasted thinking time.
What sort of workloads do well with these long tasks? The big labs are optimizing for long run time on their own, but it seems like a terrible thing to optimize on unless you're trying to do something like prove a hard math theorem, which success is clearly defined and the route doesn't matter a ton.
Plan mode has been made increasingly useless. I need to discuss to iterate to get the desired design, explore options, because Claude never gets it right first try and I don't have enough knowledge of options to specify everything up front.
Ah well, the Chinese models will still work well, I guess.
Look up the grilling skill[0]. To make it even better, tell Claude to modify it to use the AskUserQuestion functionality. It's so much better than plan mode.
0. https://github.com/mattpocock/skills/blob/main/skills/produc...
Plan mode has become pointless since Opus 5 came out, they know when to switch between planning and execution now. But that iteration/discussion is still necessary unless you're building completely blind - the model cannot read your mind.
I've had it running 8h+ of non-stop optimizations, chasing a performance target, rewriting systems or building a series of prototypes for research. All it needs is a clear goal.
My experience is that Claude has gotten absolutely terrible at switching between planning and execution, and never gets it right, and then I waste tons of turns fixing intent, and trying to make claude forget the bad shit it did.
There needs to be a mode where it's "don't change code, don't lock in decisions, lets explore", and the problem with plan mode is that it all of a sudden presents a too-long, multi-screen plan that's just completely off base with basically two options: "go and do it all" or "tell me what's wrong and then I'll make a small modification on two screenfulls of text and not tell you what I changed."
If it works for you, great, but Claude was far better for me before 5. it got slightly better with 5.5, but the harness deteriorates every days as programmers try to gain internal clout by shoving in poor features.
Pi is such a relief in comparison to Claude, those who can use it really should.
Use the superpowers plugin. It's a game changer. You start by building out a spec defining the work you want to be done, then claude writes a detailed implementation plan. It can easily work for hours without interruption, and the process integrates testing (TDD when possible) and review steps along the way. I even did one fairly complex POC for work that had claude implementing for almost 3 days straight, with me stopping it and handing off to a new session when context hit about 800k. The process is slower and more methodical than the way claude works by default but the output is much better.
Sometimes Opus returns a whole bunch of WTF mumbo jumbo. Turns out it likes to invent jargon about your project that you don't even understand. The way I get out of this is to add to the prompt "tell me all that again without jargon." That usually generates a readable response.
These example prompts seem soulless. Should I feed this document to another LLM that's OK with me being informal so I can have it give me prompts that Claude will find acceptable without contorting myself to talk like this?
Hm… I don’t think you should let claude run for 5 hours with a ten line prompt.
You’ll get “something” but what you get is certainly not going to be what you wanted.
Look, its not complicated:
1) be precise in what you want
2) have a feedback loop to verify it.
3) check in often, not once a day.
> Migrate the payment endpoints from the old client to the new one. Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes.
What end point? What test suite? What is a client?
Maaaaybe the model can infer it from your code, but look, if a human taking that jira ticket would go: what does this mean? …then your model (yes, even 5.5) is going to make a bunch of assumptions.
Correct assumptions? Maybe. Maybe not. …but do you really want to find that out hours later?
Just say:
> Plan out the following as a set high level tasks: …
Then, review the plan and tell it to execute the plan, maybe something like; after each step, ensure the code base compiles and the tests pass.
Then, come back after one hour and make sure it’s on the right track.
Poorly specified long running tasks is a recipe for “rollback all those changes…”
> Early testers had Opus 5.5 coordinate parallel subagents on long audits and migrations, with little oversight.
This is their "Why it matters" for why asking opus to use subagents matters.
The fact that they think "some other people jumped off the bridge" is a reason to do something does not inspire confidence.
This has zero mention of the success or quality of those efforts. Just because they were done "with little oversight" doesn't mean it went well...
Do you think they will recommend you things that do not work well?
The best model I've ever used easily. It's incredible. I've done so much in the past week. About three months of work I reckon.
Did you get paid the amount of three months?
It not: who becomes rich on all that productivity?
Not him, but I'll answer: me! I'm building software for my (non-software) business.
Same. I was asked to rebuild an app with AI as an experiment. The original app took months and Opus 5.5 rebuilt it in less than a week.
Is it just me, or is the first example for "define what DONE means" a lot like the "draw the rest of the fucking owl" meme? Except for a small subset of tasks that are already trivial like migration work.
What I mean is, for most tasks I do that aren't trivial, by the time I have defined what DONE means, I would've already did the work and walked the path to get there, which is what I would've hoped to not have to do in the first place.
? so you start working on something without any idea what you will end up with ?
Yes? Unless you're very early on your career, a big part of the work is figuring out requirements.
Phenomenal model, not sure what they did, but I have been able to do so much with my $20 plan!
If enough people keep saying it I'm sure they will nerf it
I think what makes it great is that they trained it to write harnesses for the code it writes, so it can test stuff even if the supplied code is not complete.
the open weight models I've been using have been doing that for months, not sure it's a new pattern in the opus 5.5, or maybe they distilled it back? :x
A significant part of it is that with Opus 5.5, cache reads are priced at 5% of inputs, instead of 10% for previous models.
It's taken me a while to get used to "cargo cult coding," but I'm surprisingly OK with it now. (I, for one, welcome our new AI overlords.)
Do these tips work? Probably. Do we know why they work? Sorta. Could we be just copy/pasting whatever text someone else kinda thought worked one time and now enters every time out of some weird feeling that it helps nudge the LLM, like some remote islanders building airport control towers out of bamboo in the hopes of a new airdrop from the gods?? Totally. But here we are.
So, another anecdotal website extolling LLMs, written - so it seems [1] by an LLM. Typical far for HN these days I suppose.
[1] : https://www.salahadawi.com/hacker-news-ai-detector/49946567