simonw 11 hours ago

I posted this in the other Astra thread but it's just fallen off the homepage, so...

Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.

Astra uses less tokens overall too, for better results.

Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

  • Xunjin 10 hours ago

    Do you have other ideas of "combinations" for this kind of benchmark?

    I'm wondering if this is being trained on by the models today.

    • r_lee 9 hours ago

      I have a feeling they've been doing that for a while now, even if unintentional, as it's such a well known benchmark

  • BrokenCogs 10 hours ago

    Interesting that all of the bikes are turquoise colored, except for the medium effort

  • lspears 10 hours ago

    Luna's price seems off

    • droidjj 8 hours ago

      The price is probably coming from when Luna was first released. OpenAI slashed the price by 80% at the end of July.

      • simonw 8 hours ago

        Yes! Good catch, thanks - I'll fix that.

        • simonw 8 hours ago

          I had the prices wrong on Sol and Terra as well - they've all had price drops:

          https://openai.com/index/advancing-the-price-performance-fro...

          Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog

            Luna — costs in cents
            +--------+--------+---------+
            | Effort | Before | After   |
            +--------+--------+---------+
            | max    |   7.83 |    1.57 |
            | xhigh  |   4.24 |    0.85 |
            | high   |   2.46 |    0.49 |
            | medium |   1.26 |    0.25 |
            | low    |   0.76 |    0.15 |
            | none   |   0.71 |    0.14 |
            +--------+--------+---------+
            Per million tokens:
            Before: $1 input / $6 output
            After:  $0.20 input / $1.20 output
          
            Sol — costs in cents
            +--------+--------+---------+
            | Effort | Before | After   |
            +--------+--------+---------+
            | max    |  48.55 |   32.37 |
            | xhigh  |  24.11 |   16.08 |
            | high   |  10.38 |    6.92 |
            | medium |  10.55 |    7.03 |
            | low    |   8.33 |    5.55 |
            | none   |   5.90 |    3.93 |
            +--------+--------+---------+
            Per million tokens:
            Before: $5 input / $30 output
            After:  $4 input / $20 output
          
            Terra — costs in cents
            +--------+--------+---------+
            | Effort | Before | After   |
            +--------+--------+---------+
            | max    |  32.09 |   25.67 |
            | xhigh  |  14.67 |   11.74 |
            | high   |   3.74 |    2.99 |
            | medium |   3.46 |    2.77 |
            | low    |   3.47 |    2.78 |
            | none   |   2.60 |    2.08 |
            +--------+--------+---------+
            Per million tokens:
            Before: $2.50 input / $15 output
            After:  $2 input / $12 output
  • andai 9 hours ago

    Oh my goodness. I was not prepared for Luna on "none".

    Reminded me of https://clocks.brianmoore.com/

    • tyre 9 hours ago

      Haiku the GOAT

    • eterm 42 minutes ago

      It's so meme-worthy next to the astra-max for any situation of "What you were promised" vs "What you received".

  • vb-8448 9 hours ago

    I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.

    • thimabi 9 hours ago

      This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.

    • pizza234 9 hours ago

      Astra is the only model that correctly depicts occlusion of crank and leg, although interestingly, at max and medium levels (not in between).

      • vessenes 8 hours ago

        Looked to me like it missed the chain though.

  • jsdalton 9 hours ago

    Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?

    You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.

    • simonw 9 hours ago

      For the moment the best place to see them all is to browse the tag on my blog - 140 posts now! https://simonwillison.net/tags/pelican-riding-a-bicycle/

      I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.

      • bahmboo 6 hours ago

        Would be great to spend some money (other people's money) on paying some illustrators and artists to do the same for comparison. It could be straight images converted to SVG or direct vector graphics.

        • Gander5739 3 hours ago

          Doesn't that somewhat miss the point? The models can't see and check their work, for instance.

          • bahmboo 2 hours ago

            I'm not clear what you are saying. I am curious about what humans can do in the same context. It's becoming a standard I think for some tasks. E.g. how does a Gen AI stack up against a human in quality and "cost". Perhaps I'm missing what you are asking.

  • fHr 9 hours ago

    Luna is my go to daily model it's great value

    • ghthor 8 hours ago

      Mine as well, it’s fast and keeps me in flow; and cheap!

    • tuo-lei 8 hours ago

      Me toooo, for all my personal projects I need to pay for the tokens~ At workplace I use sol because I don't need to pay

  • bnorton 8 hours ago

    You’re closer to this than I am but do you find it odd that helmets are almost never included? Biking is almost always accompanied by helmets

    • maxlapdev 8 hours ago

      But pelicans are almost never accompanied by helmets, so it cancels out.

    • neutronicus 8 hours ago

      At least one reasoning trace I saw considered a helmet and discarded the idea because it was worried about obscuring some detail

    • sumedh 8 hours ago

      > Biking is almost always accompanied by helmets

      Isnt Netherlands the leader in bike riders and they dont wear helmets.

  • petilon 8 hours ago

    Astra pelican looks amazing: it looks like it was made by a professional artist. All the others look like they were made by kindergartners.

  • steve-atx-7600 8 hours ago

    You would not expect the developers of the model to optimize for a well known benchmark?

    • y1n0 7 hours ago

      I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.

    • Kranar 7 hours ago

      What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?

      • awakeasleep 7 hours ago

        You gotta look up how RLHF works before you ask a demanding question like this.

      • mudkipdev 7 hours ago

        Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.

      • dgellow 2 hours ago

        > Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles?

        Yes

    • mi_lk 7 hours ago

      It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it

      Treat it like a bit as is

      • benatkin 6 hours ago

        It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.

    • simonw 5 hours ago

      Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

      Quote from the thinking trace:

      > I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.

      It's pretty solid - face is a little wonky but excellent tail and scooter.

      • CamperBob2 5 hours ago

        If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.

        • ulrikrasmussen 1 hour ago

          I remember that early image models couldn't generate a cyclops no matter how you prompted it, it would at best put a third eye in the forehead.

      • benatkin 4 hours ago

        Hmm, the face makes it not work as a one shot artifact.

      • pilaf 3 hours ago

        Interesting how the background is almost identical to two of the Astra pelicans'.

  • PacificSpecific 6 hours ago

    Fallen off the homepage so what? I genuinely don't understand

    • yreg 4 hours ago

      So he is posting it again because he believes it is crucial information for us.

      Apparently the crowd agrees because they keep upvoting these.

    • dolebirchwood 4 hours ago

      Gotta double-dip on those internet points.

  • threatripper 6 hours ago

    It seems to start understanding how bicycles work. Look at the front fork. There are reasons why it's shaped the way it is in real bicycles. If you do physical simulations in 3D with reinforcement learning you will start understanding that most of the arrangements that the other agents use just won't work.

    • epihelix 5 hours ago

      Yes, the curve of the front fork on the max pelican was what I noticed also. It's impressive (if not an accident), and a remarkably accurate bike overall. The pelican just needs to raise their seat a little.

  • CamperBob2 5 hours ago

    Crazy how much the Astra pelicans resemble GLM 5.3's (https://crimson-jeri-74.tiiny.site/), down to the color of the bike and the scarf.

    The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.

  • dector 5 hours ago

    Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 so I'm pretty sure that they spent some small amount of resources to train the model to produce good svg version as well. :D

  • mkagenius 4 hours ago

    Do you ever randomize on a particular model - like try and get 3 outputs and pick one at random?

    Coz who knows if astra low will produce max like output if tried once more.

  • alastairr 4 hours ago

    It's mildly interesting that the bikes always seem to be front wheel facing the right.

  • jeffybefffy519 3 hours ago

    Shouldnt you give the AI a different task each time? otherwise the model companies just optimise for this benchmark because its in their best interest to...

jjcm 6 hours ago

It's ability to handle non-90 degree cutouts and shapes for web dev is one of the best I've seen. The vision model on this is VERY capable.

Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...

And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...

Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...

One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.

  • copperx 4 hours ago

    > That site build cost $24 - extremely non-trivial for a simple frontend.

    I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.

    • JumpCrisscross 1 hour ago

      > truth is that the cost doesn't leave much room for error or experimentation

      Compared to what?

  • mydreamof 3 hours ago

    I don't get it. For me it seems Opus was more accurate in terms of for example this small building in the right down corner

    • CapsAdmin 1 hour ago

      I would say opus was in some ways more accurate, but missed the higher level curvature feel of the site that astra picked up on.

      It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature

  • esikich 3 hours ago

    They look pretty similar to me, I'm not sure what you're seeing that is so much better.

    • eurg 3 hours ago

      Check the source oft the laser beams.

    • w4yai 3 hours ago

      Also the curves and background lines that separate the right image from the left content

XCSme 11 hours ago

That's some crazy SVG generation:

https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...

It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.

  • embedding-shape 10 hours ago

    At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.

    • XCSme 10 hours ago

      Yeah, that's confusing, the score is for the entire benchmark, not for SVG generation only.

      Good point about the mouths, I just noticed, lol

      Imo, it's still better than most models, I personally like the stylized perspective.

      You can view here all generations for all models: https://aibenchy.com/showcase/

    • XCSme 10 hours ago

      Do you prefer the fable one?

      It's more "correct" but looks a lot worse in my opinion:

      https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...

    • XCSme 10 hours ago

      I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!

      • embedding-shape 39 minutes ago

        Now I'm wondering why it's ranked #4 instead, not sure this reduced confusion :P

        What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all

        • XCSme 8 minutes ago

          Yeah, ranking is for entire model, not only SVG generation. Maybe I remove it entirely and add generation date instead.

          I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.

          I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).

    • kzrdude 1 hour ago

      And not to mention, the hamster is standing at the wrong end of the table.

      • embedding-shape 38 minutes ago

        Incredible, I've played table-tennis since I was like 14, and I didn't notice that egregious error yet I noticed all the other small ones! Thanks for pointing that out, should have been very obvious.

      • sehugg 17 minutes ago

        And also the hamster has two mouths

  • satvikpendem 10 hours ago

    That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.

    • kulahan 9 hours ago

      That photo looks pretty hilariously stupid, so this appears to be more of a first toe dip rather than some indication we can one-shot a previously difficult process.

      • appplication 9 hours ago

        Sure, it is a bit cartoonish but it’s relatively impressive. I do wonder what you would get if you asked for photorealism

        • redox99 5 hours ago

          The problem is not photorealism. The SVG is outright dumb, its on the wrong side of the table, the table has fucked up geometry (its tilted) and many more minor flaws.

          • XCSme 2 hours ago

            That's true, if it really "knows" stuff, a bit weird to have basic flaws like 2 mouths, and sitting on the side...

    • dprkh 7 hours ago

      I thought all the designers have been using vector graphics for a long time now.

      • mceachen 7 hours ago

        Aldus Freehand 1.0 was released in 1988. Adobe Illustrator was first released in 1987.

      • embedding-shape 36 minutes ago

        "Designer" is a very broad term and can mean very different things to different people, not everyone is doing stuff that can be represented nicely by vectors.

  • readams 9 hours ago

    The Astra one looks pretty good except it's standing on the wrong side of the table

  • viccis 3 hours ago

    >AGI

    >Playing ping pong from the side of the table

    I think marketing might be getting a bit absurd at this point

    • mawadev 1 hour ago

      I must live in a different world or it feels like I'm reading LLM generated comments, but who exactly is falling for this marketing?

      I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"

1saadcodes 2 hours ago

The higher price seems less important if it actually gets the job done with fewer tokens. I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all

MisterMunchkin 4 hours ago

$10/$50 is incredibly expensive compared to Chinese models which are cents.

I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.

  • KptMarchewa 3 hours ago

    the only thing that matters is cost per task. Astra seems to be massively efficient.

  • gentlewater 2 hours ago

    Not really comparable IMO. Astra and Fable are not the every day workhorse you reach for to do basic tasks (unless your company has fuck you-money), they’re the tool you break out when you need the absolute strongest performance. There are plenty of tasks where finding and fixing one or two extra edge cases saves the business a lot of money, even if the cost is high. The best example would be scanning for vulnerabilities, if these models weren’t kneecapped in that area.

    • ghosty141 1 hour ago

      We have ChatGPT Pro at work and I usually use Terra medium/high and only bring out Sol High when the big or feature actually requires "thinking"/complex behavior. This has worked pretty well for me and it's very token efficient

kingstnap 11 hours ago

Its also available finally to Pro users! Just took 24 hours.

  • InsideOutSanta 11 hours ago

    They gave out bankable resets for every day people on pro plans didn't get Astra. Given that, I wish they'd waited a few more days before activating it on my account :-D

    • wincy 11 hours ago

      They haven’t activated Astra for me yet, I have two resets now. I’ve been using the opportunity to test out how good 5.6 Sol is at computer use asking it to generate stuff in Blender which has been… interesting

      Edit: nevermind it JUST gave me a notification to use it!

    • paxys 11 hours ago

      That's a pretty genius internal incentive to move fast.

    • embedding-shape 11 hours ago

      > They gave out bankable resets for every day people on pro plans didn't get Astra.

      Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.

d2p 1 hour ago

Odd that the tool call failure rate is so high (5%) for the OpenAI provider than Azure (0.2-0.5%).

sumedh 8 hours ago

Just got access to it on Plus plan in Australia. 2 Banked resets as well.

upcoming-sesame 2 hours ago

Any tips on using Astra as orchestrator with Luna workers efficiently in codex?

  • logged4upvoting 1 hour ago

    (This applies to Sol but probably works for Astra too)

    I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.

    The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.

    Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.

    So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).

algoth1 11 hours ago

Just got it. European plus user here. Only codex, no chatgpt

friendlypenguin 4 hours ago

I was really hoping for Astra to be less expensive then Opus...

  • redox99 3 hours ago

    It is (uses way less tokens)

    • simianwords 2 hours ago

      No it isn't cheaper, any source for task vs price comparison to Sol?

      Edit:

      GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens

      GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens

      So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive.

      Why is the burden of proof on me tho!?

      • wickedsight 2 hours ago

        This is the second comment I see where you write "no it isn't", without providing a source for your statement. Then you follow it by asking for a source. So is there a source you can provide to back up your statement?

        • dgellow 2 hours ago

          It’s supposed to work the other way, if you make a positive claim that it is cheaper, where is your proof?

vb-8448 9 hours ago

Played in codex app a couple of hours today: it feels much faster than SOL, even if the TPS is half of it.

jaesonaras 9 hours ago

Anyone had success using Astra as a Foundry model via Github Copilot? The error I get is that tooling is not available if reasoning has a value.

gavinray 11 hours ago

I have GPT-6 access in Codex and OpenAI API now

I'm a Business plan user with Cyber verification enabled, FWIW.

  • embedding-shape 11 hours ago

    Same just got access literally this minute, Pro user here, no Cyber verification but have passed my ID over to them back in 2024 or something, maybe at the ChatGPT 3 API launch or something?

    Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.

forrestthewoods 2 hours ago

Threw $10 at this to help me prepare for my league’s fantasy auction this weekend. It spend $3.50 and then said “this action would cause you to go above your spending limit”.

Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.

Sure seems like Astra is expensive AF.

r_lee 11 hours ago

is Azure for this actually ZDR?

  • jiggawatts 3 hours ago

    It cracks me up that you have to ask in a public forum because the vendors purposefully obscure this critical information.

    The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.

    There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.

    If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!

starik36 11 hours ago

What is the actual utility of using this model on Azure? It's twice as expensive, according to the link.

Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?

  • itsjustkev 10 hours ago

    Compared to OpenAI flex? I'm pretty sure that is their batch processing endpoint, which is naturally cheaper.

  • claiir 10 hours ago

    They’re ZDR and the OAI ones aren’t

    • olalonde 9 hours ago

      WTHIT?

      • spdustin 8 hours ago

        ZDR = Zero Data Retention — they don't store your inputs/outputs.

        • viccis 3 hours ago

          AKA extortion

    • nibbleyou 5 hours ago

      Does anyone know if it is not ZDR on the codex app usage too. I am assuming it isn't

  • hhh 9 hours ago

    It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.

  • Peanuts99 2 hours ago

    If you've got access to an Azure subscription and you don't pay the bill personally.

cute_boi 11 hours ago

I got it, but sadly no resets...