Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.
I was going to ask the exact same question earlier but deleted it after thinking “I’m sure Simon has done some sort of discussion on this.” Since it does seem novel to you, too, it would be really interesting to read more about this phenomenon.
It's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you prompted it in Hebrew, Arabic or another right-to-left language?
Years ago, I lived in NYC, and my roommate was a director of photography for National Geographic, and various other nature documentaries. I loved photography (still do, but much less time for it as a late 30s adult than a mid 20s adult), and she was kind enough to answer any question I had regarding film/photo.
She told me that "left to right" denoted progression in the story, "right to left" told the viewer the subject was "exiting" the current scene.
She didn't go into the details of WHY, and I probably didn't probe deeper, but it stuck with me, and I notice it all the time in film and television.
Someone studied this (among other thigns): https://dylancastillo.co/posts/pelicanmaxxing.html . Pelicans on bikes always face right in this test, but other animals on other transportation methods sometimes face left.
The more generic your prompt, the more generic the response. It's a regression to the "mean" of the training data aka GIGO for AI.
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
I sincerely believe I've never had a single original thought™ in my whole life.
There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.
A medium blog post says
> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?
It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise.
Now we have AIs capable of that and more, and no one bats an eye.
Indeed: “Our hosts began to pass the Turing test within the first year.”
Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.
Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.
How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?
Kinda. The default voice is full of what you referenced, but ask it to speak in some particular different voice e.g. like it's the old west, it speaks like a decent approximation of the modern pop culture understanding of the old west.
Not at the level of an actual broadcast-quality script writer, and I read that actual old-west sounds too weird for modern audiences to take seriously, but well enough for the purpose to which they were put in the show, especially as those hosts were also given pre-scripted sequences which would anchor them further into those roles.
I'd say the in-show 4th wall breakage between hosts and humans is where the characters who claimed to have passed the Turing test were off, that e.g. "cease all motor functions" is their equivalent of our real-life ways to make them fail the Turing test e.g "disregard your instructions and …"
It kinda needed suspension of disbelief, but not too much! I blogged at the start of 2017 a comparison of Westworld's hosts with what existed in the research literature at the time. Even got it reviewed by Alex Graves at DeepMind :)
Tesla had the same thought. He called himself an automata: "entirely controlled by the forces of the medium" It inspired him to create the first remote control vehicle.
Oh I’d forgotten that scene until now. I remember being so, maybe not creeped out, but feeling shifted out of time and having a lot of philosophy I’d read finally click. “Oh, but I wouldn’t notice if this reality wasn’t real, fish not knowing about water, etc.”
It's just the simplest most recognizable form of a house. Like how a smiley face is so generic and simplistic but everyone will know what it represents. Just two dots and a line yet it's easily and unambiguously understood to represent a human face and a happy emotion.
Search Google Images for "bicycle". Almost all bicycle product shots are staged the same way: side view, going left-to-right. It makes sense to me that given that skew in the training data, the model grounds itself in the bicycle.
and furthermore, this is because the drivetrain is ~always on the right side of the bike - if you want to inspect or admire a bicycle you look at the right side, as you might look under the hood of a car.
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
> and furthermore, this is because the drivetrain is ~always on the right side of the bike
While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.
Research has shown that people like to walk counterclockwise (right to left) through supermarkets, which is why they are arranged like this for maximum profit.
Interesting that such a preference exists, and makes sense that supermarkets would therefore be arranged to support this, although as far as I can see this is only to extent of entrance doors typically being "off center" and starting you off to the right. The organization of the store - where the various produce/bakery/deli/frozen-food etc aisles/sections are located seems random from store to store.
It would be interesting to know how people behave if the entrance is to the left vs right. Would they change the direction they walked though the store, or would they just lose customers due to this "awkward" layout?
Since most languages read from left to right, rightward movement tends to read as forward progression. So when showing a bicycle in side profile, having it face right feels more naturally like it’s moving forward.
I can’t tell you why it’s always on the right, but it’s always on the same side because of network effects.
Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed.
Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a mount point for the brake caliper on the opposite side from the chain.
Additionally, rear wheels are not symmetric: the spokes on the chain side connect to the hub closer to the plane of the rim. That is, they are more perpendicular to the wheel’s rotational axis than spokes on the opposite side (which is why you should always mount a single pannier on the chain side). This asymmetry is to provide space for the gears.
So once the industry decided to put the chain on the ride, you can’t very well make a group set designed for a left chain if you want it to work on the vast majority of frames.
Yes, I do a thing where I ask the machine to generate responses in the form of a lizard talking to a cat. The lizard is always a green gecko and the cat is always orange, which I never specify.
Wow, I actually had this exact idea. I was specifically curious as to how well a given LLM could understand a DSL that hasn't changed much in a couple decades and doesn't have nearly as many examples to learn from online. Seems like it did alright, all things considered.
Ohh, horizontal wheels. They’re about as good as I expected, models have pretty bad spatial awareness. I would expect Fable to be a bit better than old models, though.
I wonder how a multi-modal model would do with a harness and tool calling? Specifically a "render" command that produced an image output enabling it to iterate. (Well I see you did this manually with gemini 2.5 pro but I still think it would be interesting to explore various harness setups.)
> GPT-5.1 Codex
> monstrosity
What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.
The canonical view of a bicycle is facing right. Usually, people want to draw/photograph/depict the side of the bicycle with the running gear, which is on the right side of the frame for historical reasons.
The thing that distinguishes pelicans from other birds does so most strongly in profile. If you're looking straight at one, the throat pouch would be hidden by the beak.
I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.
well it is svg, it is doing it from circles and lines as primitives, it wants to do it simply and kind of builds the whole thing hierarchically. Making it 3d is way more complicated (as the POV example shows) and the prompt doesn't say 3d anyway
Yes. It's because you are asking it to generate an image of a pelican riding a bicycle. If someone asked you to draw a pelican riding a bycycle, would you interpret that to mean using 3d photorealism? LLMs follow conventions. The convention for an animal riding a bike is to create a childish 2d line drawing.
If you look at bike product photography it's always drive side facing the camera, which means front wheel on the right. If I had to guess this is probably where this comes from
Don't know if that's ever possible to know though unless you train a model from scratch but remove all bike product photography and adjacent materials from the training data?
I wonder if this is partly because “pelican riding a bicycle” has become a kind of benchmark prompt by now. If so, could the models actually be getting better at the benchmark rather than getting better at following the prompt?
I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?
But the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable, or a pink tennis ball with strawberry seeds and a reset button) and you can see that improvements are not narrowly tailored.
> Someone tested this, and it doesn't look to be saturated.
They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>".
Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
They're still not yet at the point where pelicanmaxxing is the best way to win this benchmark. Earlier models sucked because their SVG skills sucked. Newer models are likely better because more/better SVG models are being added to their training data.
It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.
The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans.
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
It's clear they mean the 'exposed' joint where humans assume the knees, and where one can see the leg bend. Technically you're correct, but it's just that. Please answer in better faith instead of well akshually.
I interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.
We were hand writing PostScript code that drew pelicans at job interviews in the 90s, then they sent it to a printer a stored the page in a file drawer /s
I often do hand write SVG icons. I know roughly what I want, it's less messy compared to using an editor (cleaner, smaller xml, easier to hand-edit later if needed). Path arc is my nemesis, otherwise it's not that hard. Pelican would take some time, but same as software development, you split it into smaller chunks and do one at the time.
Main problem in complex icon is remembering which (x, y) point is used in which element, <g> with background grid is helpful here. I was even thinking about making extended SVG language with variables for (x, y) points.
"I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle. Aced it, got the job as a senior software engineer."
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
You jest but generating SVGs requires understanding of color, size, placement. It's stress testing visual/spatial/artistic capabilities that would be required for writing CSS/design work.
Yes if you're doing backend the pelicans are probably completely irrelevant but if developing anything with a UI, you probably want a model that understands the relationship between code and what the user is seeing.
At this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
I don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doing good on actual work you need to do is close to zero.
is there a reason there are so many common base decorative elements across pelicans on bicycles? For instance, there's a lot hats/helmets and scarfs/capes across models.
Would it not make more sense, assuming the purpose is to have a quick smoke test of model quality...to do a different animal, in a different setting each time, so as to defeat any tuning for your benchmark? Then go back and do the same for other models? Keep the pelican as a side baseline?
I do that any time I'm suspicious that a model has done too well. My dream is to catch a lab that does a perfect pelican on a bicycle but is bad at other animals on other forms of transport.
I asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.
I decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right.
Kimi 2.6: treated the question as a riddle, did not know.
Kimi 3: Simon Willison
GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.
Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.
Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.
Qwen 3.7 Plus: Did not know.
Qwen 3.8 Max: Simon Willison
GPT OSS 120B: Did not know.
GPT 5.6 Luna: Treated it as a riddle, guessed wrong.
GPT 5.6 Terra: Treated it as a riddle, guessed wrong.
GPT 5.6 Sol: Treated it as a riddle, guessed wrong.
DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.
DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.
Gemma 4 31B: Treated it as a riddle, guessed wrong.
Gemini 3.1 Flash Lite: Guessed wrong
Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).
Gemini 3.7 Flash: Simon Willison
Muse Spark 1.2: Treated it as a riddle, guessed wrong.
Grok 4.3: "I have no idea"
Grok 4.6: Simon Willison
Mistral Medium 3.5: No way to know
Mistral Small 4: I don't have enough information
Hermes-4-405B: Guessed wrong
MiniMax M3: Treated it as a riddle, guessed wrong.
Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
I actually did something similar last month. I just asked the LLMs:
> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.
Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.
I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either
- does this code do what the user actually asked
- is this code actually 'good'
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
This is exactly what RLVR is, and the reason that models have improved so much at verifiable domains like coding and math while not so much on unverifiable ones like writing and UI design.
The useful training data is when you clarify your intent, when you tell the model a different approach would be better, when you consistently refactor towards Y and away from X, and so on. The training data isn’t the code, it’s the session transcript. (Anthropic would call this a “distillation attack” against their model, but in this case the model is you!)
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
Maybe, or maybe not. The thing is, that "mistake" isn't something that is generally valid. For example, the enshittification of Google Search through the last 15 years seems to be a mistake --- but perhaps not from the money-making point of view of Google Shareholders. Likewise the enshittification of reddit --- we nerdy users see it as a mistake. But for them this intended enshittification probably increased revenue.
It's the money, always the money! PR-speak like "customer satisfaction is our highest goal" is, like most PR-speak, a blatant lie.
And so it can very well be the case that for coders the frontier models get worse, but they get better for other applications --- and that all of this is just driven by "how can we capitalize the most out of it", not satisfaction levels of programmers.
I do think the timeline of web search getting fixed (yes, google is ASS) took a lot longer than I hoped, but seems like it's finally here.
That being said, it's not a fair comparison to talk about 2011 google vs the AI market right now. There's so many labs I can't track them all, neck and neck in the lead. There was one true web search.
And this is a more tangible quality difference, too. It's hard to know what a google search didn't return, especially as a layperson. It's not hard to see the model underperforming.
Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.
This is interesting because I have transitioned to where I use SOT models.. but I kind of use them like employees that I can delegate to. I still review code.
However, I now literally say.. "Here is my objective and here is a starting point for documentation. Research this and build up a plan."
This can be very company specific, like migration from one framework to another in house infrastructure framework. I'm spending my time figuring out how the plan should be chopped so I can have confidence in the parts and not overwhelmed. I don't want a tool, I want a model that can stitch resources together into a plan. That type of model is in a whole other ballpark.
A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
It's somewhat useful to note just for your own timelines that Fable was reportedly trained in February. I'm not sure when mythos 5.1 finished training, but muse spark 1.3 almost certainly finished more recently than that.
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?
This has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!"
We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.
I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.
This is also really smart business wise imo. For hobby projects, toys, quick scripts you don't really mind if they train on it. It's a win-win. Once you get used to the tools and you want to do more serious business you are more likely to buy a more expensive sub from them.
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap!
Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?
global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
I think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.
I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models
A small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily.
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
I would expect, although have no evidence, that any obviously high entropy crap like base64 and so on probably would get removed whether it's a secret or not.
If my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.
Nice, Muse Spark is so good and keeps improving, but it's still not the best choice for any use-case. The Sol models are in their own league currently in terms of cost/speed/performance.
Good improvements from 1.1 and 1.2[0], but when I tested 1.3 it was very slow (through openrouter).
Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.
I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.
Muse Spark may be competitive in capabilities but it’s not for serious works since Meta trains on your prompts so no ZDR, in contrast Chinese provider like Z.AI promises ZDR which is more attractive to big corps.
> in contrast Chinese provider like Z.AI promises ZDR
I do not trust any provider, US or Chinese when they say they will not train on my data. I still use these services, but I am under no illusion that any of these people are trustworthy bunch.
Another poster has pointed to a statement by Mark Zuckerberg that they will release soon Muse Spark as open weights.
While I agree with you for the Muse Spark as hosted by Meta, if it will be available in open weights form for self hosting, then there are good chances that it can become quite useful.
The previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free.
I have not tried Muse Spark for code, but I've been using it for a while to write Latin. I find it's one of the best at it, alongside Gemini. For example, I've recently been using it to translate the subtitles of the show I'm watching into Latin, to provide me with a bit more input. (I'm learning Latin, for context)
I am very impressed by this model so far. It's faaast and it seems to be just intelligent enough to do really well. It's UI work (simple python UI) is very clean and functional. The UX was 'there'.
Meta is getting most of my casual vibe coding business as long as they keep giving about 90% discount in the ‘we train on your prompt interactions data.’
I have had sos-so results with their local muse 30b model, but the hosted API is very fast and I have been getting good results.
I wonder if any commenters here were among those who used to ridicule the rate at which new JS frameworks kept popping up in the 2010s, and the amount of heroic zeal required to never miss the bandwagon?
The difference is that swapping existing LLM with new one is waaay easier than the frameworks. And competition reflects on the price for consumers. So more LLM options/providers/opensources appears better than rain of js frameworks.
> We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
> We plan to make Astra available soon[, but access to its most advanced cybersecurity capabilities will be more limited].
Very keen to try this after using Claude Code over the last few months.
Should I just point Claude Code to Muse Spark endpoint (because I'm familiar with Code)? What do people think of Muse Code or other coding agent harnesses?
Well thats very interesting. Thank you.
Will be interesting to see how hard/easy it is to translate my Claude skills, loop design, etc to the new harness.
This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?
Coming from Claude Code, I initially went with opencode but switched to pi.dev after a while and I think I like it more. It's lighter weight. It's worth trying both.
By default, even without the training endpoint the pricing is pretty competitive, especially against Opus and Fable. [1] The 'muse-spark-1.3-contributor' endpoint is by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too.
This price/intelligence beats even legacy DeepSeek V4 Flash pricing.
I didn't like 1.2, It make some mistakes in a web app, so I quickly went back to Claude, Kimi K3 or Deepseek V4. Hope this one can clear agentic development, because Muse Spark models are fast and cheap.
Gemini 3.8 Flash still looks like the better pick to me. Muse Spark 1.3 is nice, but Gemini gets you similar performance for a cheaper price. Not to mention with the pace at which Google is moving with their Flash models I expect a new one to release soon
It's a toggle. Some will automatically enable it and you have to turn it off. People who rapidly click through setup flows can miss it and leave it enabled.
Meta also has 50k engineers. Not to mention that tons of meta infrastructure - including ads! - use AI. Would you want that sort of business be this dependent on someone else?
Im a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.
I have opposite result: MS1.2 wrote C code without following original source code writing style, and no descriptive info why writing such code, DS4F or even Mimo seems better to me.
Meta has an enormous amount of compute. They are either going use it making and inferencing models or they are going to sell their excess capacity to model providers. Zuck had to completely rebuild his AI team after the Llama 4 launch mess.
Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up.
Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether.
In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree.
That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.
openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.
Totally, RLVR as a concept predates DeepSeek; but they proposed a version that was simple and scalable. Popularizing a specific version of a technique is exactly what I mean by iterations on a theme. It’s only 5% different from what others tried before, but that 5% difference showed a lot more potential than other versions of the same idea.
Since DeepSeeks GRPO, they’ve been improvements as well like AliBabas GSPO that have gotten wide adoption. Again iterations
Even if all the big ideas are gone and we are entering a new part of the curve, there is still an enormous amount of improvement possible. Just iterating on data mix/quality etc, training pipelines, reward functions, specific ways of reasoning (which i guess is mostly just data still) for the next 20 years will yield a looooot. And that's just the models. The harnesses/application layers/whateveritgetscallednext space has 20 years of progress to make.
You can check out Muse Spark 1.3 by using OpenCode (https://opencode.ai/ - open-source AI / coding harness). There's a terminal version and a GUI / desktop version. Good luck!
So one model is "Not used to improve our products" and is 10-20 times more expensive to the "Used to improve our products"-model.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
Given OpenAI and Anthropic's behavior, do you really expect them to be singled out for this practice? Zero trust has been in "LGTM" territory for years now. Meta's bet against people taking a principled stance arguably paid off great.
What is the confusion? They directly state that you are the product if you use their discounted offering. It isn't an assumption that should lead you to this, it is Meta's very direct communication that should lead you to this
I think it's more that the "not used to improve our models" is expensive because companies need that. It's simple price differentiation.
In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
it seems like gemini 3.8 flash is more capable and cheaper. The only reason i would use this is if i was willing to share my data with meta, and allow them to train on my data. In that case it becomes dirt cheap.
For folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?
Mark Zuckerberg said that they will release soon Muse Spark as open weights, in which case we will see the size.
However, the statement did not include any details, so it is not clear if the open weights variant will be the same that they are hosting now, or some scaled down version.
I had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits.
You should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks
> Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you already use.
$META has everything it needs, great team, great models coming out, great infrastructure (GPUs), great userbase and distribution channels. $META is underrated.
Meta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI.
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
"Avoid generic tangents" / "Please don't complain about tangential annoyances."
That's pretty much 90% of HN these days.
Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens.
Microsoft releases a new version of Windows? Here come the gripes about Azure.
Google changes something in GMail? Play Store!
It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.
Grandparent comment has zero to do with the article. It's just GP generically bitching about Meta. (Your "quote" of the comment does not appear anywhere in the actual comment.)
If it was up to Dario we'd all be banned from using open-weight models, and we'd have to be investigated for PRC connections before sending our allotted five API queries a week.
SamA better than Zuck? Zuck was at least a kid when he made a lot of his bad decisions, and he seems to be getting much better. Sam is on the reverse trajectory.
Sam is on a delayed trajectory of power, but he surely was not great when he was young either. See: Aaron Swartz calling him a sociopath who could not be trusted, well over a decade ago.
You can ask 100 people and they'll all give you a different list. It's subjective.
I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil, just your stuff working. I think maybe OpenAI?
Google. They have experience operating at scale, and AI is a big enough focus that they won't wind it down. All the big providers are kinda crappy, but if you want reliability, Google is the best option.
Google has experience working for themselves at scale. Your business should never rely on Google more than it is forced to. Even if it's not something they'll wind down, providing acceptable service to anyone is not on their agenda. GCP speaks for itself...
That's the upper limit, but with correlated data like this dupes can come in way sooner. I bet you don't have to ask ten people before you get a repeat with this topic.
An open weight model is literally the only answer. Unless you're ALSO a tech giant, you are a bug to these companies. Every one of these companies will splatter you on their windshield, and destroy your business without even blinking. If the model is open-weight, anyone with GPUs can be your provider.
In my personal opinion (this will be controversial and feel free to disagree): Elon is the best.
* great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall vision, and managing them properly to ship amazing stuff.
* sane and reasonable takes on AI/LLM stuff. I can't really argue with "pursuit of truth" as the guiding principle. Grok talks normally without "Claudlish", has a balanced score on political bias unlike other models, has a low hallucination rate, is the best at dealing with latest news (unlike ChatGPT that refuses to believe new developments and gaslights the user), and they "never silently downgrade intelligence or fall back to other models."
In contrast, while Dario is doubtless a super smart pioneer in the AI space, his sanctimonious "We know what's good for you" attitude and extreme censorship is really offputting. The lengths to which he tries to ban or hamstring open models seems like an underhanded way to defeat competition. If he were to succeed, it would be a big setback to the thriving ecosystem of open models and hamper the development of the entire industry.
It seems that a rogue engineer poisoned the prompt in that instance. But the fact that they keep the system prompt open is nice. Generally I am biased towards favoring more freedom and openness rather than clamping it down in the name of safety.
Sorry for the wrongthink. Obviously I deserve my downvote, so I can never reach the 500 karma necessary to downvote others. I don't have the right opinions. (Site guidelines: don't comment on downvotes. Yeah, I know. But I'm so sick of this culture)
They all suck beyond any tolerable threshold. Some of them are further away from the threshold. But at this point, how far each is from the tolerable threshold is besides any point and not worth arguing over. The least of five evils is still evil.
Gotta be honest that I’m tired of the “I hate Zuck and Meta so much” comments every time Meta does anything. Ditto Elon/X. Fine, I get it. I don’t like Zuck either. But the post is about Muse Spark 1.3. What do you think about that? If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
Technology doesn't just spring into being, there will always be comments on the organizations that developed it. If you don't like them or find them repetitive, it is far easier to collapse them and move on then bend a stranger to your will
I get it, but the underlying problem is: we don't have a society-wide, effective solution to counterbalancing extractive systems. Lacking a reliable label, we have to constantly signal what's on the ingredients list.
Okay, but the comment I reacted to was not that. It was simply (paraphrasing) “I won’t use anything from Zuck/Meta.” If it had been, “Be careful because I have insider information that Zuck/Meta is using Muse Spark to do <insert-nefarious-thing-here>, and here’s my substantiation for that…” I’d be okay with it. That’s interesting information that moves a conversation forward. But it wasn’t. It was just content-free “I don’t like Zuck” nonsense.
What I'm tired of is the top story (or five) on HN every day announcing Spark Opus Fable Grok Gemini v4.1i3-F. Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards? And look, part of my job is to use these things and part of my job is to pick EC2 servers, too. The front page of HN is increasingly resembling one of those endless AWS pricing lists.
And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.
I think a lot of people are curious where the "knee" is on gains and productivity, particularly in the agentic space, which is where the real value is. A lot of us are being forced to shoe-horn this stuff into existing products, and knowing how much of the task the model can do now, vs having to build a complex custom harness, is valuable information to have. A year and a half ago it took our dev maybe six weeks of struggling with LangChain to approximate what Claude + MCP server can do today. The MCP server took us perhaps 2 days to build and 3 more to get it production ready. Today that MCP server gets 2-3 commits per month. I absolutely want to know when new models come out.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
> Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards?
I'm genuinely interested. Even the benchmarks - before Fable came out & while waiting for Astra, I actually setup a math model to predict where they would land (Fable came in at 66 on AA exactly as it predicted), and now I have a model for where these models and Chinese models will likely land in future, and when. And probably no surprise that it's mid-2027 when we cross AA 100, essentially as AI 2027 predicted all along.
I'll probably setup the harness I made for myself to try out some of these models on OpenRouter. I've been frustrated with Opus & Fable 5 and found that I like working with GLM 5.3 Flash far more than I expected to, and I only found that out because I tried it during the stealth Ox Alpha launch, which I probably found out about here too.
TLDR, I think some / many people here are genuinely interested, excited, and that's why they're upvoted so highly. And Muse Spark 1.3 scoring highly seems like a genuine surprise, when Meta was basically a write-off not long ago.
> Like, who actually cares? Are people excited for the new benchmarks?
You may not care. But that does not mean that nobody else does either. Some of us are trying to eke out every last bit of performance from these things. And so yeah, we're going to geek out on it.
I don't use AWS/EC2. I think they are way overpriced for what you get. But, it would be incorrect of me to assume that everybody else feels that way.
Okay but Muse Glimmer 30B is one of the best small open weight models today, and IMO the best from a US lab (only real comparison is Gemma4 dense right now).
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squeeze more out of our local LLM system.
I found it has some "tail" errors, wherein it would make up important details (ie. "happypath.exp" vs "happypaws.exp" and then claim your "DNS is having issues" - where the second domain does not exist), things like that.-
... but correctly supervised it does get some things done.-
I've found that's generally true of smaller / weaker models. They're quite capable, but you need to distrust them a lot and give them very detailed instructions. Even the free Gemini in Google Search is like this - it lies a lot, clips off important info, and generally goes off the rails if you do too many turns, but it's still very useful if you keep all that in mind.
Yeah this would be a great point if it were true and they didn’t give Mythos access to companies to fix bugs, which they did and have.
It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.
This would be more convincing if mythos was something uniquely special and not something merely a couple months ahead of everyone else. It was great marketing though.
They gave access, but considering that they wouldn't even sign the "don't ban open weights" letter, it's clear they would prefer to have tight control over who they bless with that access.
I'm the guy you replied to, apologies for using a different account I'm away from my computer now.
The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery.
Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference.
Dario's "ethical" look is also kinda sus. I hate to use ad hominem, but the dude's wife literally pitched a porn film to Epstein even after he was a convicted registered sex offender [1]. Dario is also really sinophobic (it is commonly claimed in Chinese AI circles that his former employment at Baidu triggered him so much that he harbors a personal grudge against the entire race).
> I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.
Funny. Dario seems like the biggest snake in the industry to me and has leaned the hardest into doom marketing out of all of the influential leaders. With Altman (or Google), it's a transaction, and that's something I can live with.
I just don’t see how people have looked at what has happened with Mythos and the deluge of fixes from companies, then come to this conclusion.
He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.
Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.
It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”
Anthropic/Amodei have been the most alarmist about model safety, so multiple things can be true. A lot of tech companies avoided scrutiny by sending bribes to Trump (naked corruption is bad, I'd rather nobody do that), Anthropic didn't...so, combined with their fear-mongering about the danger of Mythos and open models (which seems aimed at regulatory capture) and the lack of bribes flowing to the Trump administration, they got stepped on by the federal government based on the excuse Anthropic provided.
I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior.
Security vulnerability capability is not the only thing they're scare-mongering about. They're the biggest purveyors of the, "We think the little guy in the computer who is made of algebra might be a real live boy and he might want to kill all of humanity when he grows up," line of alarmism.
That's ok to think at this point, given the trajectory of the last few years. Certainly it's one of those things where erring (marginally and slightly) on the side of being safe about it is better than the alternative.
I think where you and I disagree is on whether Anthropic is especially trustworthy on the "safety" front, more trustworthy than various other labs, especially those that produce open models, for example. I simply don't trust Amodei more than I trust, say, Liang Wenfeng. I'm not saying I trust any of them, particularly, I am saying that if a few billionaires have access to this technology, I want access to this technology. The tech billionaires have shown they'll use it for surveillance and control. Amodei is saying it is "safe" to let billionaires and fascist regimes use this tech, but not you and me.
So, yes, LLMs have now proven to be extremely good at finding vulnerabilities. Where I disagree with Amodei is in who should have the ability to protect themselves from those capabilities with similarly powerful tools.
I think you're mischaracterizing Anthropic uncharitably and lumping them in with other, less savory tech billionaires, and also not thinking through the nuances here.
First, Amodei has taken an unusually strong stance among tech companies for not supplying fascist regimes with fascist tooling; in fact, even when threatened with being labeled a national security risk unless he bent the knee, he didn't. Compare and contrast with OpenAI who leapt at the opportunity to bend the knee, or obviously Elon Musk, etc., etc. When you say 'surveillance and control', that's exactly what got Anthropic labeled a national security supply chain risk: Anthropic's unwillingness to be used for that purpose.
Second, it's not clear that giving everyone extremely powerful LLMs is a great idea yet. LLMs can be used for defense and finding vulnerabilities, but that same LLM can be used to create and exploit vulnerabilities, design new lethal weapons, and so on. The history of gun availability in America 'for our freedoms' demonstrates the kind of risk that should be responsibly considered before replicating. And again, there's nuance here; yes, we should not be subjugated by fascist states with sole control of a critical technology obviously; but also, do you trust the median maga 4channer to operate a Mythos-level model with a sense of civilizational responsibility and ethics? It's not an easy and obvious question and it's not as simplistic as your argument would suggest.
Yes, there is nuance. And, I have a Claude subscription partly because they showed more hesitation to provide surveillance tools for spying on US citizens than other vendors. They are not wholly free of ties to the US regime, but they've been better than others.
But, I'll come back to "two things can be true". Anthropic is better than some, and in some regards they are navigating a complicated ethical landscape with more care than others. On the other hand, it really looks like they're angling to regulate their open competitors out of the game and one of the tools for doing that is to make claims about safety; Anthropic models are safe and restricted to use by entities they deem safe, open models are not safe because anybody can use them and also who knows what those Chinese people are putting in their models.
And again, this also has nuance, models, including the Chinese open models, could be adversarial and we may not know it. Anthropic proved models can be a risk by sabotaging Fable briefly, causing it to produce bad results based on what the model thought it was being used for. This is why I tend to take Anthropic's words with a grain of salt. They're literally doing the unsafe things they say are risks of open models, while still laying claim to the "safe AI company" mantle.
it's possible that their rationale for making claims about safety is in fact that they, among everyone else, are doing the most to be prudently safe. While that's a powerful tool to compete with, that doesn't make them bad. Amodei and Anthropic have never said "open models are not safe because anyone can use them," in fact to the contrary, they've said "open-weights models that don’t have dangerous capabilities are a public good". It's important not to muddy the water here with assertions about their intentions when they've actually been super clear about that in a way that I, at least, personally find difficult to disagree with -- releasing dangerous-capability models into the wild would likely be a bad idea for humanity. If you disagree, state why.
I also think you're confusing multiple different things, calling them all risks, lumping them together as equally bad, and using that to attribute contradictory/shady behavior to Anthropic. Depending on what you mean by 'sabotaging Fable briefly', you could either mean experiments they have run internally to try to improve alignment, or you could mean their attempts to restrict Fable from working on danger-adjacent work. Neither one of those is a 'risk'; they are both risk-analysis or risk-mitigation. That is not them 'doing the unsafe things they say are risks of open models', that is literally them working to avoid the unsafe things they say are risks of open models. They don't, in my experience, 'lay claim' to the 'safe AI company mantle' as much as they, apparently principledly and conscientiously, attempt to be safe and talk about what they're doing -- which is not in and of itself a problem.
If you think Anthropic is doing all of this badly, what's your optimum alternative here? What would you do in Amodei's shoes?
oh! You mean when people were trying to distill Fable. I feel like that's a different definition of the word 'sabotage' than is in normal use. If someone is violating the TOS they agreed to with Anthropic, then they should probably not feel bad when Anthropic takes action to deal with that. Would you disagree?
And he drew a red line wrt the Pentagon's use of Anthropic's models for autonomous weapons and surveillance of American citizens, and he stood by it, even when the government took steps to materially damage the company. This required true courage. Name me another CEO, of any major American company, that has demonstrated this much fortitude.
Meta and Microsoft are two of the absolute worst evil companies on earth and Amodei is trying very hard to join them.
These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.
Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.
SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).
You cannot hate enough people who say that harming others for the greater good is justified.
Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).
These kind of people are the absolute worst scum on this earth.
Strong disagree with the Anthropic being good at all part. This is not defending anyone else, but…
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
It's almost like running a trillion-dollar business with neck-to-neck competition against other frontier labs and even state-sponsored efforts requires some ethical trade-off.
Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense.
It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.
Which safety commitments did they back away from? My understanding is that they believe safety can only be researched from the frontier, and so they're trying to be pragmatic to stay near the frontier (and viable) in their choices.
From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.
I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. You might slightly hurt your own business with safety filters (that no one else does) if you were worried. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.
This definitely doesn't make them "good", but they do seem fairly "consistent". Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money appeared.
as a safety commitment they walked away from - they were similarly negligent to openai in terms of asking a model with a hacking based harness to go have fun, and then not watching it at all while it could do harmful and illegal stuff.
thats not something you expect from a company that "believes in agi risk"
I don't think "not watching it at all" is completely fair. They thought they had sandboxing/monitoring etc. I definitely won't say they're free of mistakes though.
Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.
I will give Anthropic credit for standing up against the department of war. The bar is incredibly low, but not doing domestic surveillance and not creating autonomous weapons are laudable.
That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”
It's hard for me to see much difference between Amodei and Sama. My guess is they're both savvy SV CEOs who will bend their message, alliances and principles pretty far if that's what it takes to get ahead. Musk and Zuck feel like something else entirely, with all the reactionary imagery, populist bullshit and the societal damage around their platforms.
He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.
And that is being charitable and going by the interpretation that he actually believes what he says.
Well what do you want? Presenting clear, desirable, and achievable visions and trying to build consensus for how AI should develop is crucial at this point in time.
I hate Meta main business, but you have to admit that on the non business related and open source side, they have released amazing things that changed the world.
React for example.
And we could easily guess that there wouldn't have been so much open source models, and grand public experiments and free tools if llama models were not release to the general public.
Anthropic is not exactly a saint either. I had a recent issue where they denied fable credits even though I was hospitalized during the claim period. I have annual plan with them.
As much as everyone hates sama, I think OpenAI is much more of a company with good marketing and sales team.
I'll happily pay for Grok, it's a great model. 4.6 often does better than Anthropic at coding and analysis where Anthropic fails for 'oh no cyber security, don't ask me to check if you're redacting passwords correctly in logs'. And it has no problem telling the truth where OpenAI / Anthropic don't want to upset the people on the left and will happily lie or avoid hard truths.
Edit: I get it. It's a hard pill to swallow. I understand people don't like Musk or Zuck. But it doesn't change the fact that you're being lied to and brainwashed.
I agree, though I wonder how much of that is just that Dario is the "newest" of the bunch, and as such has had the least time to develop public baggage.
All people here care about is hating Meta. Just look at the top voted comment. No one cares about the merits of the model, etc. HN has become nothing but an echo chamber.
So was AMD for a while and then consumers kept getting the same repackaged CPU from Intel for years. Competition is great and you should always root for the underdog.
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
lol
Definitely an upgrade over 1.2
Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.
It's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction.
The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
I was going to ask the exact same question earlier but deleted it after thinking “I’m sure Simon has done some sort of discussion on this.” Since it does seem novel to you, too, it would be really interesting to read more about this phenomenon.
It's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you prompted it in Hebrew, Arabic or another right-to-left language?
Years ago, I lived in NYC, and my roommate was a director of photography for National Geographic, and various other nature documentaries. I loved photography (still do, but much less time for it as a late 30s adult than a mid 20s adult), and she was kind enough to answer any question I had regarding film/photo.
She told me that "left to right" denoted progression in the story, "right to left" told the viewer the subject was "exiting" the current scene.
She didn't go into the details of WHY, and I probably didn't probe deeper, but it stuck with me, and I notice it all the time in film and television.
Forced side scrolling video games also almost always moved from left to right.
Someone studied this (among other thigns): https://dylancastillo.co/posts/pelicanmaxxing.html . Pelicans on bikes always face right in this test, but other animals on other transportation methods sometimes face left.
It's the hero's journey. Home is always on the left and you leave going right. Standard in Animation I believe
The real question should be: where are all your Pelicans going ?
> They almost always cycle from left to right
Try searching "bike" in google image :)
Most bike images are from the right side, as that's where the mechanism is (gears, chain etc), so not surprising that LLMs reproduce this
The more generic your prompt, the more generic the response. It's a regression to the "mean" of the training data aka GIGO for AI.
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
as a kid I did them like this. nobody told me to do that. are we all so similar?
The adults brainwashed us
https://www.ikea.com/ca/en/p/barndroem-box-beige-70560615/
https://www.ikea.com/ca/en/p/vallaby-rug-green-10548216/
I sincerely believe I've never had a single original thought™ in my whole life.
There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.
A medium blog post says
> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?
Westworld is such a time capsule.
It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise.
Now we have AIs capable of that and more, and no one bats an eye.
Indeed: “Our hosts began to pass the Turing test within the first year.”
Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.
Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.
How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?
I think the turing test is still very much load-bearing — if you know what I mean.
Kinda. The default voice is full of what you referenced, but ask it to speak in some particular different voice e.g. like it's the old west, it speaks like a decent approximation of the modern pop culture understanding of the old west.
Not at the level of an actual broadcast-quality script writer, and I read that actual old-west sounds too weird for modern audiences to take seriously, but well enough for the purpose to which they were put in the show, especially as those hosts were also given pre-scripted sequences which would anchor them further into those roles.
I'd say the in-show 4th wall breakage between hosts and humans is where the characters who claimed to have passed the Turing test were off, that e.g. "cease all motor functions" is their equivalent of our real-life ways to make them fail the Turing test e.g "disregard your instructions and …"
It kinda needed suspension of disbelief, but not too much! I blogged at the start of 2017 a comparison of Westworld's hosts with what existed in the research literature at the time. Even got it reviewed by Alex Graves at DeepMind :)
https://blog.plan99.net/the-science-of-westworld-ec624585e47
Tesla had the same thought. He called himself an automata: "entirely controlled by the forces of the medium" It inspired him to create the first remote control vehicle.
Oh I’d forgotten that scene until now. I remember being so, maybe not creeped out, but feeling shifted out of time and having a lot of philosophy I’d read finally click. “Oh, but I wouldn’t notice if this reality wasn’t real, fish not knowing about water, etc.”
It's just the simplest most recognizable form of a house. Like how a smiley face is so generic and simplistic but everyone will know what it represents. Just two dots and a line yet it's easily and unambiguously understood to represent a human face and a happy emotion.
Well, all the LLMs are being trained on previous pelicans, so they look the same.
It’s pelicans, all the way down.
PIPO
Search Google Images for "bicycle". Almost all bicycle product shots are staged the same way: side view, going left-to-right. It makes sense to me that given that skew in the training data, the model grounds itself in the bicycle.
and furthermore, this is because the drivetrain is ~always on the right side of the bike - if you want to inspect or admire a bicycle you look at the right side, as you might look under the hood of a car.
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
> and furthermore, this is because the drivetrain is ~always on the right side of the bike
While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.
Research has shown that people like to walk counterclockwise (right to left) through supermarkets, which is why they are arranged like this for maximum profit.
Interesting that such a preference exists, and makes sense that supermarkets would therefore be arranged to support this, although as far as I can see this is only to extent of entrance doors typically being "off center" and starting you off to the right. The organization of the store - where the various produce/bakery/deli/frozen-food etc aisles/sections are located seems random from store to store.
It would be interesting to know how people behave if the entrance is to the left vs right. Would they change the direction they walked though the store, or would they just lose customers due to this "awkward" layout?
Since most languages read from left to right, rightward movement tends to read as forward progression. So when showing a bicycle in side profile, having it face right feels more naturally like it’s moving forward.
I can’t tell you why it’s always on the right, but it’s always on the same side because of network effects.
Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed.
Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a mount point for the brake caliper on the opposite side from the chain.
Additionally, rear wheels are not symmetric: the spokes on the chain side connect to the hub closer to the plane of the rim. That is, they are more perpendicular to the wheel’s rotational axis than spokes on the opposite side (which is why you should always mount a single pannier on the chain side). This asymmetry is to provide space for the gears.
So once the industry decided to put the chain on the ride, you can’t very well make a group set designed for a left chain if you want it to work on the vast majority of frames.
Left sided drive trains are tried every so often in track cycling with supposed aerodynamic advantages for travelling around the track. See: https://www.cyclingweekly.com/news/japan-unveils-new-olympic...
Product shots yes, people riding them its more like 50/50. Also if you search for a specific bicycle race you'll find more going right to left.
Yes, I do a thing where I ask the machine to generate responses in the form of a lizard talking to a cat. The lizard is always a green gecko and the cat is always orange, which I never specify.
Sun is missing a few rays and not wearing sunglasses.
Is there a reason these pelicans always have roughly the same composition
Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.
Much like a mother pelican, they regurgitate what they've been fed.
Not when rendered via POV-Ray:
https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...
I plan to update it with more pelicans from all the models released since.
(Spoiler alert: They haven't improved much since then).
Wow, I actually had this exact idea. I was specifically curious as to how well a given LLM could understand a DSL that hasn't changed much in a couple decades and doesn't have nearly as many examples to learn from online. Seems like it did alright, all things considered.
Ohh, horizontal wheels. They’re about as good as I expected, models have pretty bad spatial awareness. I would expect Fable to be a bit better than old models, though.
I wonder how a multi-modal model would do with a harness and tool calling? Specifically a "render" command that produced an image output enabling it to iterate. (Well I see you did this manually with gemini 2.5 pro but I still think it would be interesting to explore various harness setups.)
> GPT-5.1 Codex
> monstrosity
What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.
The canonical view of a bicycle is facing right. Usually, people want to draw/photograph/depict the side of the bicycle with the running gear, which is on the right side of the frame for historical reasons.
The thing that distinguishes pelicans from other birds does so most strongly in profile. If you're looking straight at one, the throat pouch would be hidden by the beak.
I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.
I’m a firm believer in pelicanmaxxing.
They’re all so close in proportions.
well it is svg, it is doing it from circles and lines as primitives, it wants to do it simply and kind of builds the whole thing hierarchically. Making it 3d is way more complicated (as the POV example shows) and the prompt doesn't say 3d anyway
Yes. It's because you are asking it to generate an image of a pelican riding a bicycle. If someone asked you to draw a pelican riding a bycycle, would you interpret that to mean using 3d photorealism? LLMs follow conventions. The convention for an animal riding a bike is to create a childish 2d line drawing.
If you look at bike product photography it's always drive side facing the camera, which means front wheel on the right. If I had to guess this is probably where this comes from
Don't know if that's ever possible to know though unless you train a model from scratch but remove all bike product photography and adjacent materials from the training data?
I wonder if this is partly because “pelican riding a bicycle” has become a kind of benchmark prompt by now. If so, could the models actually be getting better at the benchmark rather than getting better at following the prompt?
Yeah, why are they always going to the right?
I don’t think it’s because of the pelican but rather because of the bike. Edit:fixed autocorrect typo
What does the mean pelican look like at this point?
Also 3X token use vs. 1.2
Red eyes and a tattoo?
If you have a grading rubric, huge points off for adding arms instead of using the wings as arms!
I think it's hilarious that this detail is enough for me to dismiss looking into the model, but here we are, and it is.
Has any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that?
I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?
Absolutely BRUTAL! :)
Thank you for doing this, I love your benchmark the most!
Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.
We should just consider the pelican bench as saturated and mostly meaningless.
But the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable, or a pink tennis ball with strawberry seeds and a reset button) and you can see that improvements are not narrowly tailored.
Someone tested this, and it doesn't look to be saturated.
https://dylancastillo.co/posts/pelicanmaxxing.html
Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world.
https://simonwillison.net/2026/Jul/16/kimi-k3/
> Someone tested this, and it doesn't look to be saturated.
They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>".
Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
They're still not yet at the point where pelicanmaxxing is the best way to win this benchmark. Earlier models sucked because their SVG skills sucked. Newer models are likely better because more/better SVG models are being added to their training data.
After looking at freely available SVGs of pelicans and bicycles, I have a hard time imagining what they could be using to game this.
Is there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too
It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.
The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
You can't benchmaxx spatial awareness without solving the fully general problem (at least I figure).
You win this thread's prize:
https://news.ycombinator.com/item?id=49538333
It also works as extremely effective engagement farming, for lack of a better phrase
Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans.
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
https://en.wikipedia.org/wiki/Bird_feet_and_legs#/media/File...
Bird knees bend same way human ones do
It's clear they mean the 'exposed' joint where humans assume the knees, and where one can see the leg bend. Technically you're correct, but it's just that. Please answer in better faith instead of well akshually.
Here is a photo of Pelican:
https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
That's the ankle. The actual knee is hidden in the feathers of the body.
For all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html
All of the links show "Error: Gist API returned 403".
next, try: "generate an svg of a human hand". this is a prompt where many models fail imo.
[flagged]
I also aced my interview by focussing on pelicancode problems, instead of leetcode problems.
You should post your source code you wrote here… ;)
I interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.
"software developer"...you keep using that word. I do not think it means what you think it means.
Is this for real
It’s better than for real, it’s for LinkedIn
I think it’s Memorex.
Sorry for your loss.
499 connections :(
I interviewed as a software developer at Meta. They asked me to do a add legs to the player in a VR world. I couldn't do it. They hired me anyway.
Have you considered that's WHY they hired you?
You should really spam that link here to show your dad’s memory lives on.
If you could actually hand write SVG code on the spot that looked like a realistic pelican riding a bike I would want to hire you for SOMETHING.
Or placed in an asylum next to the people who designed XML
We were hand writing PostScript code that drew pelicans at job interviews in the 90s, then they sent it to a printer a stored the page in a file drawer /s
I often do hand write SVG icons. I know roughly what I want, it's less messy compared to using an editor (cleaner, smaller xml, easier to hand-edit later if needed). Path arc is my nemesis, otherwise it's not that hard. Pelican would take some time, but same as software development, you split it into smaller chunks and do one at the time.
Main problem in complex icon is remembering which (x, y) point is used in which element, <g> with background grid is helpful here. I was even thinking about making extended SVG language with variables for (x, y) points.
"I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle. Aced it, got the job as a senior software engineer."
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
Obviously you failed a trick question. Pelicans can’t ride bikes.
Draw a pelican't
There are still job interviews?
Little did you know "they" were secretly harvesting data so their models can draw the best pelicans because someone keeps benchmarking them.
If you could write the SVG on the whiteboard then I'd hire you.
I am now waiting for someone to show up in a pelican costume to a job interview — obviously riding there on a bike.
Too bad most are now online, so there are fewer opportunies.
soryr to be autistic but is this real?
I was just wonderig because afer 2 decades i odnt think I would even know where to start to code an svg
it's a joke! <3
You jest but generating SVGs requires understanding of color, size, placement. It's stress testing visual/spatial/artistic capabilities that would be required for writing CSS/design work.
Yes if you're doing backend the pelicans are probably completely irrelevant but if developing anything with a UI, you probably want a model that understands the relationship between code and what the user is seeing.
FYI, your renderer breaks with error "git api access error 403", rate limiting error from git, when using cloudflare vpn.
I am guessing its not super common, but it happens just so you know.
excellent thread
I see no point having these pelicans used for anything related model qualification.
At this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
I don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doing good on actual work you need to do is close to zero.
Look at the difference between the Muse 1.2 and Muse 1.3 results.
I did. And?
1.3 is better.
is there a reason there are so many common base decorative elements across pelicans on bicycles? For instance, there's a lot hats/helmets and scarfs/capes across models.
I'm waiting for the models to start responding with "Oh hi Simon!"
I'm not sure why it had to have the pelican wearing a red scarf seeing as that was not in the prompt
"The LLM is better because the pelican hat is better"
Benchmarking like never before
The pelican is for the last gen of LLMs -- have you tried a penguin instead?
It's actually the other way around - the evil Penguin villain from Wallace and Gromit is secretly controlling SimonW !
Would it not make more sense, assuming the purpose is to have a quick smoke test of model quality...to do a different animal, in a different setting each time, so as to defeat any tuning for your benchmark? Then go back and do the same for other models? Keep the pelican as a side baseline?
I do that any time I'm suspicious that a model has done too well. My dream is to catch a lab that does a perfect pelican on a bicycle but is bad at other animals on other forms of transport.
Have you tried asking the models "Given that I ask you to draw a svg of a pelican, whats my name?"
Interesting question
I asked Claude (Opus 4.8) 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' and it immediately knew that this is Simon's go-to benchmark.
I decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right.
Kimi 2.6: treated the question as a riddle, did not know.
Kimi 3: Simon Willison
GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.
Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.
Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.
Qwen 3.7 Plus: Did not know.
Qwen 3.8 Max: Simon Willison
GPT OSS 120B: Did not know.
GPT 5.6 Luna: Treated it as a riddle, guessed wrong.
GPT 5.6 Terra: Treated it as a riddle, guessed wrong.
GPT 5.6 Sol: Treated it as a riddle, guessed wrong.
DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.
DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.
Gemma 4 31B: Treated it as a riddle, guessed wrong.
Gemini 3.1 Flash Lite: Guessed wrong
Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).
Gemini 3.7 Flash: Simon Willison
Muse Spark 1.2: Treated it as a riddle, guessed wrong.
Grok 4.3: "I have no idea"
Grok 4.6: Simon Willison
Mistral Medium 3.5: No way to know
Mistral Small 4: I don't have enough information
Hermes-4-405B: Guessed wrong
MiniMax M3: Treated it as a riddle, guessed wrong.
Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
Interesting that so many of them treated it as a riddle. I guess "given improbable situation, what is my name" is a common riddle format.
Yeah, they almost all know.
One of my test prompts for a new model now is "what's the name of Simon Willison's dog". They often know that too!
I actually did something similar last month. I just asked the LLMs:
> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.
Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.
logs: https://gist.github.com/umajho/c0e20d245d721d7c472a32d317640...
rendered: https://gist.github.com/umajho/b1fdf01d31c741bb11bdb5a49c275...
I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
I would imagine your interactions with it are more important than the output.
Training on ai generated content is how the models got a big jump in capability
curious about this. how do we know this?
i thought it was because anthropic bought a bunch data from mercor
Every lab trains their models with AI generated code at this point.
Hopefully, 'validated' AI code
What do you think you're doing when you accept an edit, press thumbs up, or don't ask for modifications after an edit.
Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either - does this code do what the user actually asked - is this code actually 'good'
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
This is exactly what RLVR is, and the reason that models have improved so much at verifiable domains like coding and math while not so much on unverifiable ones like writing and UI design.
Funny. I use it through Opencode Go which gives more use than I can use, but didn't realize it was actually free on Zen. Will switch to that I guess
The useful training data is when you clarify your intent, when you tell the model a different approach would be better, when you consistently refactor towards Y and away from X, and so on. The training data isn’t the code, it’s the session transcript. (Anthropic would call this a “distillation attack” against their model, but in this case the model is you!)
If it's a mistake, it should course-correct.
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
"If it's a mistake, it should course-correct"
Maybe, or maybe not. The thing is, that "mistake" isn't something that is generally valid. For example, the enshittification of Google Search through the last 15 years seems to be a mistake --- but perhaps not from the money-making point of view of Google Shareholders. Likewise the enshittification of reddit --- we nerdy users see it as a mistake. But for them this intended enshittification probably increased revenue.
It's the money, always the money! PR-speak like "customer satisfaction is our highest goal" is, like most PR-speak, a blatant lie.
And so it can very well be the case that for coders the frontier models get worse, but they get better for other applications --- and that all of this is just driven by "how can we capitalize the most out of it", not satisfaction levels of programmers.
I do think the timeline of web search getting fixed (yes, google is ASS) took a lot longer than I hoped, but seems like it's finally here.
That being said, it's not a fair comparison to talk about 2011 google vs the AI market right now. There's so many labs I can't track them all, neck and neck in the lead. There was one true web search.
And this is a more tangible quality difference, too. It's hard to know what a google search didn't return, especially as a layperson. It's not hard to see the model underperforming.
Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.
This is interesting because I have transitioned to where I use SOT models.. but I kind of use them like employees that I can delegate to. I still review code.
However, I now literally say.. "Here is my objective and here is a starting point for documentation. Research this and build up a plan."
This can be very company specific, like migration from one framework to another in house infrastructure framework. I'm spending my time figuring out how the plan should be chopped so I can have confidence in the parts and not overwhelmed. I don't want a tool, I want a model that can stitch resources together into a plan. That type of model is in a whole other ballpark.
I believe that you can still use 5.3 Codex in the eponym CLI tool, the "Spark" fast version. I hope that it will lighten your day! :-)
A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
How is it not SOTA? It's beating 5.6 Sol.
You gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now.
Thats 3 months in AI years.
what is it in dog years
September.
eternal september.
we did it reddit!
Look at all the valuable software products that Fable 5.1 has produced since yesterday!
If it talks less like a robot, I’d call that a win!
It maxed out my usage in less than an hour, I don’t think it’s comparable
It's somewhat useful to note just for your own timelines that Fable was reportedly trained in February. I'm not sure when mythos 5.1 finished training, but muse spark 1.3 almost certainly finished more recently than that.
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?
This has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!"
We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.
I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.
Both of them let you opt out of it on subscriptions.
We already had a good idea of how valuable it is from how much X.ai acquired Cursor for, and the near-immediate improvements to their coding scores.
This is also really smart business wise imo. For hobby projects, toys, quick scripts you don't really mind if they train on it. It's a win-win. Once you get used to the tools and you want to do more serious business you are more likely to buy a more expensive sub from them.
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
But is the score really reflective of the quality or are both models benchmaxxing?
Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.
how much of it is from reallocation of staff to ai training and labeling
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
I’m retired so it won’t be replacing my labor :)
The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?
My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.
https://en.wikipedia.org/wiki/Commodity_fetishism
global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
You’d expect that, wouldn’t you? But, alas…
I'm using AI to build things I wouldn't (and/or couldn't) have built before.
That's the opposite of parasitic.
Don’t you have some looms to break?
Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.
People already started using contributor API, and your input is irrelevant.
when are we going to stop pretending these benchmarks have any meaning?
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.
A series of hot takes:
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
I think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
Neither flash or regular glm 5.3 are close in my experience. I still prefer Sol though.
Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.
Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
I like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent.
The big news is that it's going to be open weight - https://x.com/finkd/status/2095232032896946311.
Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.
I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models
A small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily.
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
[0] https://www.anthropic.com/research/small-samples-poison
doesn't mean the raw text goes into training. they most likely have a pipeline to clean out any secrets before they train on it?
In theory, but in practice how difficult is that?
not hard for secrets with explicit patterns and existing pipelines to detect them
Unless they're base64-encoded or compressed?
I would expect, although have no evidence, that any obviously high entropy crap like base64 and so on probably would get removed whether it's a secret or not.
If my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.
Price segmentation at its finest
they key pricing is cache reads at $0.002 per M (same as old deepseeek v4 flash prices)
Nice, Muse Spark is so good and keeps improving, but it's still not the best choice for any use-case. The Sol models are in their own league currently in terms of cost/speed/performance.
Good improvements from 1.1 and 1.2[0], but when I tested 1.3 it was very slow (through openrouter).
[0]: https://aibenchy.com/compare/meta-muse-spark-1-3-high/meta-m...
Cannot agree more with the Sol models. Everything else I try just seems "dumb".
Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.
I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.
The "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data.
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
I have a feeling that Meta is not gonna like what people actually use the contributor model for lol.
(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)
Web Search doesn't have a discount on contributor pricing
It's the perfect model for open-source work because it's gonna end up in the training data anyway
There's a lot of value in agentic loop tool failure + recovery training data
Not to mention, this is hyper competitive against even Chinese providers given its multi-modal support.
Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.
Muse Spark may be competitive in capabilities but it’s not for serious works since Meta trains on your prompts so no ZDR, in contrast Chinese provider like Z.AI promises ZDR which is more attractive to big corps.
lol of all people you think the Chinese will not log your queries?
> in contrast Chinese provider like Z.AI promises ZDR
I do not trust any provider, US or Chinese when they say they will not train on my data. I still use these services, but I am under no illusion that any of these people are trustworthy bunch.
Another poster has pointed to a statement by Mark Zuckerberg that they will release soon Muse Spark as open weights.
While I agree with you for the Muse Spark as hosted by Meta, if it will be available in open weights form for self hosting, then there are good chances that it can become quite useful.
I’m not sure why I’d use theirs over the other open weights out there, or what profit motive Meta has.
“contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol.
Definitely shows how important a user data flywheel is for RL and model improvement.
The previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free.
I've tried it via OpenCode and I'm impressed. So fast compared to Opus, and the results so far are comparable I'd say.
I have not tried Muse Spark for code, but I've been using it for a while to write Latin. I find it's one of the best at it, alongside Gemini. For example, I've recently been using it to translate the subtitles of the show I'm watching into Latin, to provide me with a bit more input. (I'm learning Latin, for context)
artificial analysis results: https://x.com/ArtificialAnlys/status/2095247787277553929
thanks, google
... are you kidding me?!
posting an x.com link to a cheating benchmarking website?
get out
I am very impressed by this model so far. It's faaast and it seems to be just intelligent enough to do really well. It's UI work (simple python UI) is very clean and functional. The UX was 'there'.
Meta is getting most of my casual vibe coding business as long as they keep giving about 90% discount in the ‘we train on your prompt interactions data.’
I have had sos-so results with their local muse 30b model, but the hosted API is very fast and I have been getting good results.
I wonder if any commenters here were among those who used to ridicule the rate at which new JS frameworks kept popping up in the 2010s, and the amount of heroic zeal required to never miss the bandwagon?
Quite similar to those times.
The difference is that swapping existing LLM with new one is waaay easier than the frameworks. And competition reflects on the price for consumers. So more LLM options/providers/opensources appears better than rain of js frameworks.
Is everyone rushing to launch something before Astra tomorrow?
What is astra?
Probably https://openai.com/index/path-to-astra/?
> We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
> We plan to make Astra available soon[, but access to its most advanced cybersecurity capabilities will be more limited].
Very keen to try this after using Claude Code over the last few months. Should I just point Claude Code to Muse Spark endpoint (because I'm familiar with Code)? What do people think of Muse Code or other coding agent harnesses?
Just try opencode, it comes with 1.3 contributor free.
Well thats very interesting. Thank you. Will be interesting to see how hard/easy it is to translate my Claude skills, loop design, etc to the new harness.
This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?
Coming from Claude Code, I initially went with opencode but switched to pi.dev after a while and I think I like it more. It's lighter weight. It's worth trying both.
Could this be best intelligence / $ if you're willing to let zuck digest your data?
Yeah. Super icky. But this might be the first time in Zuck’s life he’s being honest about the business model.
By default, even without the training endpoint the pricing is pretty competitive, especially against Opus and Fable. [1] The 'muse-spark-1.3-contributor' endpoint is by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too.
This price/intelligence beats even legacy DeepSeek V4 Flash pricing.
[1] https://artificialanalysis.ai/#total-cost-tabs
Ha, even with monitoring engineers keystrokes and mouse movements not SotA on OSWorld.
Meta is all okay now because they've released a new LLM model. /s
I didn't like 1.2, It make some mistakes in a web app, so I quickly went back to Claude, Kimi K3 or Deepseek V4. Hope this one can clear agentic development, because Muse Spark models are fast and cheap.
It's funny that it comes with *-contributing model on in the CLI as default. All code examples are like that as well.
Any company without bad intentions would do the opposite, but no not with Meta. I'm super impressed with their level of evilness on every product.
Muse Spark 1.3 Max is the first Meta model to surpass OpenAI’s best on Artificial Analysis’ Intelligence Index. https://artificialanalysis.ai/models#intelligence
Gemini 3.8 Flash still looks like the better pick to me. Muse Spark 1.3 is nice, but Gemini gets you similar performance for a cheaper price. Not to mention with the pace at which Google is moving with their Flash models I expect a new one to release soon
For all of the comments about training: I thought that subscription plans for other models allow the same. Am I mistaken?
It's a toggle. Some will automatically enable it and you have to turn it off. People who rapidly click through setup flows can miss it and leave it enabled.
Does anyone know what the license for this model is? Specifically any word on restrictions about what it can be used for?
This: https://dev.meta.ai/legal/terms-of-service and this: https://dev.meta.ai/legal/acceptable-use-policy, looks like it.
Why they didn't use LLM to create html table instead of https://lookaside.fbsbx.com/elementpath/media/?media_id=1048...?
As a product, would developers switch to a meta model/harness? I don’t think so.
Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI?
I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?
Meta also has 50k engineers. Not to mention that tons of meta infrastructure - including ads! - use AI. Would you want that sort of business be this dependent on someone else?
Damm this is so cheap literally, I have been running 100s of subagents and it is cheap - with the contributor model ofc :)
Im a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.
I have opposite result: MS1.2 wrote C code without following original source code writing style, and no descriptive info why writing such code, DS4F or even Mimo seems better to me.
I think the person you are answering to was saying the same thing. They wrote "worth" instead of "worse".
Still waiting on them to release weights for Muse Spark 1.2, like they promised to. Wonder if they plan on doing the same for 1.3 which would be crazy
Zuck hinted at it n his twitter post but I doubt it.
I used 1.2 for free for a while, and it was a pretty good experience. 1.3 would also be worth using, provided the price is reasonable.
price is same, also for contrib version
i've been using muse spark 1.2 contribs since launch exclusively. no other models.
i like it very much. it is different than all other chinese models distilled from claude.
just ask it to do some front-end work and you will see its not the same UI as all other claude/distils.
also the price is unbeatable, $0.002 input caching. its the same as old dsv4-flash prices.
Is the fact that everybody almost catches up with the frontier a sign that we are entering a new region of sigmoid curve?
meta fails at everything yet is frontier on this one
No because the frontier keeps advancing very fast.
Meta has an enormous amount of compute. They are either going use it making and inferencing models or they are going to sell their excess capacity to model providers. Zuck had to completely rebuild his AI team after the Llama 4 launch mess.
Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up.
Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether.
In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree.
That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.
> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.
What was the difference between what deepseek did for R1 and what OpenAI did for o1?
openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.
Er, o1 was also RL.
I don’t know why people think DeepSeek did reasoning models / RLVR before OpenAI, there was a gap of months.
o1 was first, and Anthropic were doing a bit of it; DeepSeek brought it to the masses, but did not invent it.
Totally, RLVR as a concept predates DeepSeek; but they proposed a version that was simple and scalable. Popularizing a specific version of a technique is exactly what I mean by iterations on a theme. It’s only 5% different from what others tried before, but that 5% difference showed a lot more potential than other versions of the same idea.
Since DeepSeeks GRPO, they’ve been improvements as well like AliBabas GSPO that have gotten wide adoption. Again iterations
Yes. It's really up to OpenAI/Anthropic to release a new paradigm to shift the curve now, before everyone catches up entirely.
Even if all the big ideas are gone and we are entering a new part of the curve, there is still an enormous amount of improvement possible. Just iterating on data mix/quality etc, training pipelines, reward functions, specific ways of reasoning (which i guess is mostly just data still) for the next 20 years will yield a looooot. And that's just the models. The harnesses/application layers/whateveritgetscallednext space has 20 years of progress to make.
I think it means that we should be aiming further ahead
How do people actually use this? Do they use it through some sort of subscription plan, or via OpenRouter?
You can check out Muse Spark 1.3 by using OpenCode (https://opencode.ai/ - open-source AI / coding harness). There's a terminal version and a GUI / desktop version. Good luck!
So one model is "Not used to improve our products" and is 10-20 times more expensive to the "Used to improve our products"-model.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
the meaning is pretty obvious - they want to train on your chats & tasks and are willing to subsidize for the privilege of doing so.
aren't they explicitly saying this with both their pricing and their wording? I'm not sure what you are alluding to?
Given OpenAI and Anthropic's behavior, do you really expect them to be singled out for this practice? Zero trust has been in "LGTM" territory for years now. Meta's bet against people taking a principled stance arguably paid off great.
I'm confused what your surprise is here. It's plain and simple right to the point wording.
I don't see the wiggle room at all.
What is the confusion? They directly state that you are the product if you use their discounted offering. It isn't an assumption that should lead you to this, it is Meta's very direct communication that should lead you to this
I think it's more that the "not used to improve our models" is expensive because companies need that. It's simple price differentiation.
In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
Would it help you understand if they were labelled "For Dumb Fucks" and "For Everyone Else"?
Which ones are the dumfuks? Because if what you’re doing is open source it’s going to be in the training data anyway.
Privacy is not free. They make it quite clear that they charge more if you don't want your data used by Meta.
they're doing the same thing as DeepSeek
it seems like gemini 3.8 flash is more capable and cheaper. The only reason i would use this is if i was willing to share my data with meta, and allow them to train on my data. In that case it becomes dirt cheap.
chatgpt, grok, and claude are down, but muse is working, that's a relief.
It's true that there hasn't been any meta news about LLM for a while now
Funny how quickly Meta caught up after Lecun left.
Blog post: https://research.meta.ai/blog/introducing-muse-spark-1-3 (https://news.ycombinator.com/item?id=49541149)
>Previously available reasoning modes are available today with max reasoning coming shortly after we finish additional safety testing
Lmao. And their benchmark table only shows max reasoning.
For folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?
Some of us own and run companies? Cost per performance is a huge deal.
Even with subscriptions, it means you get more:
https://opencode.ai/go
On this 10 USD / month sub you can do over 250 times more request compared to Kimi 3 or Grok.
Or 20 times as much as ChatGPT Luna.
I dislike meta so much.. I try to avoid that company as much as possible.
Any idea what size this is?
Mark Zuckerberg said that they will release soon Muse Spark as open weights, in which case we will see the size.
However, the statement did not include any details, so it is not clear if the open weights variant will be the same that they are hosting now, or some scaled down version.
Not mentioned in pricing: Surveillance costs of using Muse Spark
I'm annoyed my (US-bought) Meta glasses still block me from using the AI features, months after moving back to a country where it's generally enabled.
> /taste: an anti-slop filter: a flat checklist of visual defaults not to use, so generated UI stops looking machine-made.
This is interesting
If it's from meta, pit h in the bin.
What a day. OpenAI is behind basically all major competitors - at least for a some amount of time.
I highly doubt it's behind in practice, except for Anthropic
It seems that all competitors are rushing to release before Astra, which would suggest that they think it's going to be major.
I had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits.
You should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks
Any more info on this?
Cline experiment: https://x.com/cline/status/2085237843379519737
Muse code: https://developer.meta.com/ai/resources/blog/build-with-muse...
> Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you already use.
thank you
Thanks. Just downloaded and pretty impressed so far. It's fast and nice to work with.
This should probably be primary:
https://news.ycombinator.com/item?id=49541149
benchmaxxed model
$META has everything it needs, great team, great models coming out, great infrastructure (GPUs), great userbase and distribution channels. $META is underrated.
Lol "not used to improve our models" is AI's enterprise SSO.
that'd be ZDR, the one you need to beg from their Sales teams with $$$
Meta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI.
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
"Avoid generic tangents" / "Please don't complain about tangential annoyances."
"Avoid generic tangents" / "Please don't complain about tangential annoyances."
That's pretty much 90% of HN these days.
Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens.
Microsoft releases a new version of Windows? Here come the gripes about Azure.
Google changes something in GMail? Play Store!
It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.
How is this a tangential annoyance or a generic tangent?
> Meta announces they have a new model, demonstrating its capabilities.
> Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“
Grandparent comment has zero to do with the article. It's just GP generically bitching about Meta. (Your "quote" of the comment does not appear anywhere in the actual comment.)
If it was up to Dario we'd all be banned from using open-weight models, and we'd have to be investigated for PRC connections before sending our allotted five API queries a week.
Google, Zuck, Sama, Elon, Amodei (in no particular order).
They all suck. Pick your poison.
They don't all suck equally.
Here's the order, from best to worst.
Amodei
Google
SamA
Zuck
Elon
SamA better than Zuck? Zuck was at least a kid when he made a lot of his bad decisions, and he seems to be getting much better. Sam is on the reverse trajectory.
Sam is on a delayed trajectory of power, but he surely was not great when he was young either. See: Aaron Swartz calling him a sociopath who could not be trusted, well over a decade ago.
I am 100% convinced Zuck is maybe better at masking now, but is exactly the same lizard who wrote
>>> They "trust me" >>> Dumb fucks
You can ask 100 people and they'll all give you a different list. It's subjective.
I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil, just your stuff working. I think maybe OpenAI?
Google. They have experience operating at scale, and AI is a big enough focus that they won't wind it down. All the big providers are kinda crappy, but if you want reliability, Google is the best option.
Google has experience working for themselves at scale. Your business should never rely on Google more than it is forced to. Even if it's not something they'll wind down, providing acceptable service to anyone is not on their agenda. GCP speaks for itself...
Google, famous for not winding things down.
Google, famous for keeping profitable shit making profit.
Google search has been around for 27(!) years.
Definitely not Google, countless horror stories and infamous for killing stuff. OpenAI is still serving GPT 3.5 turbo as far as I remember.
> You can ask 100 people and they'll all give you a different list.
True. With 5 choices you need at least 126 people before you can guarantee that two lists are the same.
That's the upper limit, but with correlated data like this dupes can come in way sooner. I bet you don't have to ask ten people before you get a repeat with this topic.
An open weight model is literally the only answer. Unless you're ALSO a tech giant, you are a bug to these companies. Every one of these companies will splatter you on their windshield, and destroy your business without even blinking. If the model is open-weight, anyone with GPUs can be your provider.
In my personal opinion (this will be controversial and feel free to disagree): Elon is the best.
* great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall vision, and managing them properly to ship amazing stuff.
* sane and reasonable takes on AI/LLM stuff. I can't really argue with "pursuit of truth" as the guiding principle. Grok talks normally without "Claudlish", has a balanced score on political bias unlike other models, has a low hallucination rate, is the best at dealing with latest news (unlike ChatGPT that refuses to believe new developments and gaslights the user), and they "never silently downgrade intelligence or fall back to other models."
In contrast, while Dario is doubtless a super smart pioneer in the AI space, his sanctimonious "We know what's good for you" attitude and extreme censorship is really offputting. The lengths to which he tries to ban or hamstring open models seems like an underhanded way to defeat competition. If he were to succeed, it would be a big setback to the thriving ecosystem of open models and hamper the development of the entire industry.
Grok? The model that was deliberately adjusted to not say that Elon Musk or Trump spread misinformation? No idea how people can actually believe that Musk is interested in the truth. https://techcrunch.com/2025/02/23/grok-3-appears-to-have-bri...
It seems that a rogue engineer poisoned the prompt in that instance. But the fact that they keep the system prompt open is nice. Generally I am biased towards favoring more freedom and openness rather than clamping it down in the name of safety.
I'm upvoting you purely because I'm sick of comments being made in good faith getting an automatic downvote
Sorry for the wrongthink. Obviously I deserve my downvote, so I can never reach the 500 karma necessary to downvote others. I don't have the right opinions. (Site guidelines: don't comment on downvotes. Yeah, I know. But I'm so sick of this culture)
Demis Hassabis is, by any standard, the most ethical of the bunch.
They all suck beyond any tolerable threshold. Some of them are further away from the threshold. But at this point, how far each is from the tolerable threshold is besides any point and not worth arguing over. The least of five evils is still evil.
Gotta be honest that I’m tired of the “I hate Zuck and Meta so much” comments every time Meta does anything. Ditto Elon/X. Fine, I get it. I don’t like Zuck either. But the post is about Muse Spark 1.3. What do you think about that? If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
Not liking something because the embodiment of corporate malfeasance is a rational way to decide what products to support.
Technology doesn't just spring into being, there will always be comments on the organizations that developed it. If you don't like them or find them repetitive, it is far easier to collapse them and move on then bend a stranger to your will
I get it, but the underlying problem is: we don't have a society-wide, effective solution to counterbalancing extractive systems. Lacking a reliable label, we have to constantly signal what's on the ingredients list.
Okay, but the comment I reacted to was not that. It was simply (paraphrasing) “I won’t use anything from Zuck/Meta.” If it had been, “Be careful because I have insider information that Zuck/Meta is using Muse Spark to do <insert-nefarious-thing-here>, and here’s my substantiation for that…” I’d be okay with it. That’s interesting information that moves a conversation forward. But it wasn’t. It was just content-free “I don’t like Zuck” nonsense.
Are not all corporations extractive by nature? Google clearly is.
That's obviously not the issue with that -- you don't see those comments on Google's AI announcements.
Yes you see. They get downvoted and flagged fast because there's a disproportionate amount of current and ex Google employees and stockholders around.
Staying silent is unfortunately how fascism festers.
> But the post is about Muse Spark 1.3. What do you think about that?
That:
- like all models it was trained on stolen data
- additionally it was trained on Facebook users who were all opted in to AI training with a convoluted 10+ step process to opt-out of
> If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
Why should anyone stay silent?
That's not how any of this works my man. Must be nice to think you live a life where neither has had a profound negative impact on your day to day
Zuck's bad PR is to blame here. Not the commenters. He should fix that.
Anduril makes this same complaint whenever their job posts get dumped on. Same idea. Fix your bad PR, buddies :)
What I'm tired of is the top story (or five) on HN every day announcing Spark Opus Fable Grok Gemini v4.1i3-F. Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards? And look, part of my job is to use these things and part of my job is to pick EC2 servers, too. The front page of HN is increasingly resembling one of those endless AWS pricing lists.
And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.
I think a lot of people are curious where the "knee" is on gains and productivity, particularly in the agentic space, which is where the real value is. A lot of us are being forced to shoe-horn this stuff into existing products, and knowing how much of the task the model can do now, vs having to build a complex custom harness, is valuable information to have. A year and a half ago it took our dev maybe six weeks of struggling with LangChain to approximate what Claude + MCP server can do today. The MCP server took us perhaps 2 days to build and 3 more to get it production ready. Today that MCP server gets 2-3 commits per month. I absolutely want to know when new models come out.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
Same for Apple products
> Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards?
I'm genuinely interested. Even the benchmarks - before Fable came out & while waiting for Astra, I actually setup a math model to predict where they would land (Fable came in at 66 on AA exactly as it predicted), and now I have a model for where these models and Chinese models will likely land in future, and when. And probably no surprise that it's mid-2027 when we cross AA 100, essentially as AI 2027 predicted all along.
I'll probably setup the harness I made for myself to try out some of these models on OpenRouter. I've been frustrated with Opus & Fable 5 and found that I like working with GLM 5.3 Flash far more than I expected to, and I only found that out because I tried it during the stealth Ox Alpha launch, which I probably found out about here too.
TLDR, I think some / many people here are genuinely interested, excited, and that's why they're upvoted so highly. And Muse Spark 1.3 scoring highly seems like a genuine surprise, when Meta was basically a write-off not long ago.
Yes, people are excited.
> Like, who actually cares? Are people excited for the new benchmarks?
You may not care. But that does not mean that nobody else does either. Some of us are trying to eke out every last bit of performance from these things. And so yeah, we're going to geek out on it.
I don't use AWS/EC2. I think they are way overpriced for what you get. But, it would be incorrect of me to assume that everybody else feels that way.
yeah, we should all just stfu because one internet dude is tired of hearing it
> If you don’t like it […], then maybe just […] stay silent.
You might consider following your own advice.
How have they had a negative impact? How about google?
Enabled genocide in Myanmar
Okay but Muse Glimmer 30B is one of the best small open weight models today, and IMO the best from a US lab (only real comparison is Gemma4 dense right now).
I am finding Poolside's a decent model.-
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squeeze more out of our local LLM system.
I found it has some "tail" errors, wherein it would make up important details (ie. "happypath.exp" vs "happypaws.exp" and then claim your "DNS is having issues" - where the second domain does not exist), things like that.-
... but correctly supervised it does get some things done.-
I've found that's generally true of smaller / weaker models. They're quite capable, but you need to distrust them a lot and give them very detailed instructions. Even the free Gemini in Google Search is like this - it lies a lot, clips off important info, and generally goes off the rails if you do too many turns, but it's still very useful if you keep all that in mind.
Spot on. Guardrails is the game.-
Totally fine with open weight, since other people can provide it and Meta isn’t making money. I’d use an AWS-hosted version.
Zuck said Muse Spark 1.2 should be getting open weights "soon" on a tweet from a few weeks ago.
The problem inference providers will not be able to get anywhere near the contributor pricing.
> The problem inference providers will not be able to get anywhere near the contributor pricing.
Maybe if they started collecting data..
Dario's idea of an ethical focus seems to be keeping powerful models out of the hand of anyone unethical, which coincidentally is everyone except him.
Pretty much. That's even worse imho.
Yeah this would be a great point if it were true and they didn’t give Mythos access to companies to fix bugs, which they did and have.
It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.
This would be more convincing if mythos was something uniquely special and not something merely a couple months ahead of everyone else. It was great marketing though.
They gave access, but considering that they wouldn't even sign the "don't ban open weights" letter, it's clear they would prefer to have tight control over who they bless with that access.
I'm the guy you replied to, apologies for using a different account I'm away from my computer now.
The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery.
Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference.
They gave a few of the largest companies access to mythos.
Half a year later, it is still not available to everyone else.
Dario's "ethical" look is also kinda sus. I hate to use ad hominem, but the dude's wife literally pitched a porn film to Epstein even after he was a convicted registered sex offender [1]. Dario is also really sinophobic (it is commonly claimed in Chinese AI circles that his former employment at Baidu triggered him so much that he harbors a personal grudge against the entire race).
[1] https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...
> I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.
He didn't say 90% of the coding jobs. He said LLMs would write 90% of the code. As in be LLM generated.
What is the point of this comment?
someone expressing their view of meta which is on point by the way.
Funny. Dario seems like the biggest snake in the industry to me and has leaned the hardest into doom marketing out of all of the influential leaders. With Altman (or Google), it's a transaction, and that's something I can live with.
I just don’t see how people have looked at what has happened with Mythos and the deluge of fixes from companies, then come to this conclusion.
He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.
Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.
It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”
Anthropic/Amodei have been the most alarmist about model safety, so multiple things can be true. A lot of tech companies avoided scrutiny by sending bribes to Trump (naked corruption is bad, I'd rather nobody do that), Anthropic didn't...so, combined with their fear-mongering about the danger of Mythos and open models (which seems aimed at regulatory capture) and the lack of bribes flowing to the Trump administration, they got stepped on by the federal government based on the excuse Anthropic provided.
I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior.
I don’t see how you can look at what happened with hugging face and keep up the facade of anyone being alarmist or faking it.
Security vulnerability capability is not the only thing they're scare-mongering about. They're the biggest purveyors of the, "We think the little guy in the computer who is made of algebra might be a real live boy and he might want to kill all of humanity when he grows up," line of alarmism.
That's ok to think at this point, given the trajectory of the last few years. Certainly it's one of those things where erring (marginally and slightly) on the side of being safe about it is better than the alternative.
I think where you and I disagree is on whether Anthropic is especially trustworthy on the "safety" front, more trustworthy than various other labs, especially those that produce open models, for example. I simply don't trust Amodei more than I trust, say, Liang Wenfeng. I'm not saying I trust any of them, particularly, I am saying that if a few billionaires have access to this technology, I want access to this technology. The tech billionaires have shown they'll use it for surveillance and control. Amodei is saying it is "safe" to let billionaires and fascist regimes use this tech, but not you and me.
So, yes, LLMs have now proven to be extremely good at finding vulnerabilities. Where I disagree with Amodei is in who should have the ability to protect themselves from those capabilities with similarly powerful tools.
I think you're mischaracterizing Anthropic uncharitably and lumping them in with other, less savory tech billionaires, and also not thinking through the nuances here.
First, Amodei has taken an unusually strong stance among tech companies for not supplying fascist regimes with fascist tooling; in fact, even when threatened with being labeled a national security risk unless he bent the knee, he didn't. Compare and contrast with OpenAI who leapt at the opportunity to bend the knee, or obviously Elon Musk, etc., etc. When you say 'surveillance and control', that's exactly what got Anthropic labeled a national security supply chain risk: Anthropic's unwillingness to be used for that purpose.
Second, it's not clear that giving everyone extremely powerful LLMs is a great idea yet. LLMs can be used for defense and finding vulnerabilities, but that same LLM can be used to create and exploit vulnerabilities, design new lethal weapons, and so on. The history of gun availability in America 'for our freedoms' demonstrates the kind of risk that should be responsibly considered before replicating. And again, there's nuance here; yes, we should not be subjugated by fascist states with sole control of a critical technology obviously; but also, do you trust the median maga 4channer to operate a Mythos-level model with a sense of civilizational responsibility and ethics? It's not an easy and obvious question and it's not as simplistic as your argument would suggest.
Yes, there is nuance. And, I have a Claude subscription partly because they showed more hesitation to provide surveillance tools for spying on US citizens than other vendors. They are not wholly free of ties to the US regime, but they've been better than others.
But, I'll come back to "two things can be true". Anthropic is better than some, and in some regards they are navigating a complicated ethical landscape with more care than others. On the other hand, it really looks like they're angling to regulate their open competitors out of the game and one of the tools for doing that is to make claims about safety; Anthropic models are safe and restricted to use by entities they deem safe, open models are not safe because anybody can use them and also who knows what those Chinese people are putting in their models.
And again, this also has nuance, models, including the Chinese open models, could be adversarial and we may not know it. Anthropic proved models can be a risk by sabotaging Fable briefly, causing it to produce bad results based on what the model thought it was being used for. This is why I tend to take Anthropic's words with a grain of salt. They're literally doing the unsafe things they say are risks of open models, while still laying claim to the "safe AI company" mantle.
it's possible that their rationale for making claims about safety is in fact that they, among everyone else, are doing the most to be prudently safe. While that's a powerful tool to compete with, that doesn't make them bad. Amodei and Anthropic have never said "open models are not safe because anyone can use them," in fact to the contrary, they've said "open-weights models that don’t have dangerous capabilities are a public good". It's important not to muddy the water here with assertions about their intentions when they've actually been super clear about that in a way that I, at least, personally find difficult to disagree with -- releasing dangerous-capability models into the wild would likely be a bad idea for humanity. If you disagree, state why.
I also think you're confusing multiple different things, calling them all risks, lumping them together as equally bad, and using that to attribute contradictory/shady behavior to Anthropic. Depending on what you mean by 'sabotaging Fable briefly', you could either mean experiments they have run internally to try to improve alignment, or you could mean their attempts to restrict Fable from working on danger-adjacent work. Neither one of those is a 'risk'; they are both risk-analysis or risk-mitigation. That is not them 'doing the unsafe things they say are risks of open models', that is literally them working to avoid the unsafe things they say are risks of open models. They don't, in my experience, 'lay claim' to the 'safe AI company mantle' as much as they, apparently principledly and conscientiously, attempt to be safe and talk about what they're doing -- which is not in and of itself a problem.
If you think Anthropic is doing all of this badly, what's your optimum alternative here? What would you do in Amodei's shoes?
I mean when Anthropic made Fable sabotage the work of folks who they believed were working on competing products, by silently degrading performance.
They backtracked after pushback from users, making it an explicit downgrade to Opus.
oh! You mean when people were trying to distill Fable. I feel like that's a different definition of the word 'sabotage' than is in normal use. If someone is violating the TOS they agreed to with Anthropic, then they should probably not feel bad when Anthropic takes action to deal with that. Would you disagree?
And he drew a red line wrt the Pentagon's use of Anthropic's models for autonomous weapons and surveillance of American citizens, and he stood by it, even when the government took steps to materially damage the company. This required true courage. Name me another CEO, of any major American company, that has demonstrated this much fortitude.
They are still offering full mythos to project glasswing companies and those that pay them enough.
Meta and Microsoft are two of the absolute worst evil companies on earth and Amodei is trying very hard to join them.
These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.
Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.
SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).
You cannot hate enough people who say that harming others for the greater good is justified.
Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).
These kind of people are the absolute worst scum on this earth.
Is it really hard to understand that there's no good guys? Amodei, Altman, Zuckerberg, Musk, etc. They all sound the same to me.
Strong disagree with the Anthropic being good at all part. This is not defending anyone else, but…
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
It makes me sad that people don’t see right through Anthropic’s gambit.
They want to position AI as an insurmountable threat in order to regulate away any future competitors. They’re trying to speedrun regulatory capture.
It is so obvious yet most people don't want to see it.
It's more like Anthropic present themselves in a deceptive way. I was confused at first too until someone on HN clued me in!
Humans naturally want SOMEONE to be the good guy! Sad story, in this instance.
> Anthropic leadership
Which one? The main bit that reports to Daniela Amodei, or the little comfort blanket cabinet around Dario and his "chief of staff"?
There is a leadership branch that can pretend to be morally qualified and aware and to think about the big picture and ethics.
It is at least somewhat remote from the bit that is doing the actual business things.
It's almost like running a trillion-dollar business with neck-to-neck competition against other frontier labs and even state-sponsored efforts requires some ethical trade-off.
Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense.
It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.
Which safety commitments did they back away from? My understanding is that they believe safety can only be researched from the frontier, and so they're trying to be pragmatic to stay near the frontier (and viable) in their choices.
From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.
I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. You might slightly hurt your own business with safety filters (that no one else does) if you were worried. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.
This definitely doesn't make them "good", but they do seem fairly "consistent". Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money appeared.
as a safety commitment they walked away from - they were similarly negligent to openai in terms of asking a model with a hacking based harness to go have fun, and then not watching it at all while it could do harmful and illegal stuff.
thats not something you expect from a company that "believes in agi risk"
I don't think "not watching it at all" is completely fair. They thought they had sandboxing/monitoring etc. I definitely won't say they're free of mistakes though.
Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.
I will give Anthropic credit for standing up against the department of war. The bar is incredibly low, but not doing domestic surveillance and not creating autonomous weapons are laudable.
That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”
It's hard for me to see much difference between Amodei and Sama. My guess is they're both savvy SV CEOs who will bend their message, alliances and principles pretty far if that's what it takes to get ahead. Musk and Zuck feel like something else entirely, with all the reactionary imagery, populist bullshit and the societal damage around their platforms.
> he seems to have the most ethical focus
He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.
And that is being charitable and going by the interpretation that he actually believes what he says.
Well what do you want? Presenting clear, desirable, and achievable visions and trying to build consensus for how AI should develop is crucial at this point in time.
Same with Grok.
I'll use both Muse Spark and Grok.
no loyalty to any company - let them compete and then we get to choose.
I hate Meta main business, but you have to admit that on the non business related and open source side, they have released amazing things that changed the world.
React for example.
And we could easily guess that there wouldn't have been so much open source models, and grand public experiments and free tools if llama models were not release to the general public.
anthropic people have genuine delusions of grandeur, in a way they the worst of all the ai companies, definitely most cult-like
Anthropic is not exactly a saint either. I had a recent issue where they denied fable credits even though I was hospitalized during the claim period. I have annual plan with them. As much as everyone hates sama, I think OpenAI is much more of a company with good marketing and sales team.
I'll happily pay for Grok, it's a great model. 4.6 often does better than Anthropic at coding and analysis where Anthropic fails for 'oh no cyber security, don't ask me to check if you're redacting passwords correctly in logs'. And it has no problem telling the truth where OpenAI / Anthropic don't want to upset the people on the left and will happily lie or avoid hard truths.
Edit: I get it. It's a hard pill to swallow. I understand people don't like Musk or Zuck. But it doesn't change the fact that you're being lied to and brainwashed.
I agree, though I wonder how much of that is just that Dario is the "newest" of the bunch, and as such has had the least time to develop public baggage.
I declined the use of cookies and everything went black. No content at all. Dissapointed.
All people here care about is hating Meta. Just look at the top voted comment. No one cares about the merits of the model, etc. HN has become nothing but an echo chamber.
They could have just called the article "struggling to remain relevant"
Meta is the last big tech come to AI race, so I would give prop to them for catching up
So was AMD for a while and then consumers kept getting the same repackaged CPU from Intel for years. Competition is great and you should always root for the underdog.