I find the Sunday roast comparison of 5.6 vs 6 very interesting. I have no doubt most people will prefer 6, yet I am almost repulsed by all the images, so much needless whitespace, checklist and so on. Feels like I'm being condescended to and treated like a child.
Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools.
Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot.
Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.
It's insane how hard OAI is choking, they had much better models than A/ who were fumbling this year to date... then blunder after blunder.
The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.
I am curious: Why is it bad to merge chat and work? In Claude I didn't have a porblem with it (although I switched to an OpenAI subscription shortly after they made the merge). Isn't the model capable of deciding if it needs the extra capabilities of Work?
Pro messages is like 50 per week for pro 100 and 200 for pro 200.
Regular chat, even on extra high thinking mode, is sort of unlimited. Idk at what point you hit the abuse gaurdrail but its really high whatever it is.
For a company with Open in the name, OpenAI is very opaque about things, there are all kinds of limits and quotas that randomly come and go, but they only show you the meters for a couple.
E.g. codex reviews come from codex quota but also has its separate quota they don’t tell you about. Same with security reviews which has another quota still.
So the chat messages can be sent at several effort levels and from various models including 5.6 and 6 pro. The 6 pro messages are limited and they don’t tell you how many you’ve sent or how many are left. But the 5.6 messages seem to be unlimited, however there could just be a much higher unstated quota.
Which is probably why they will merge it. You can get a lot done in chat mode, especially if you connect custom MCPs to it. I can't see that continuing.
I don't like either. I'd want a recipe that could fit on a single page with type. Recipes you find on Google actively hide the ingredient list and simple procedures to get more of the user's time to monetize but I don't know why ChatGPT needs to be verbose here. Recipes are simple things. Simple search solved find recipes back in the early 2000s and it's prime example of a thing people have been enshitifying since there's not much to really give people beyond the formula. And sure, every entrepreneur screams "they think they want a formula but really they want an experience" but, no I don't.
Yeah it reminds me of one those obnoxious recipe/biography websites that is the laughingstock of the internet. Why would I possibly want an AI image of a imaginary roast once I'm already at the recipe stage?
Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.
They just need to train the model to make up stories about how the recipe was invented by its great grandmother during great depression and generate a ai picture of dusty recipie book .
"My old grand-pappy, GPT-1, used to wax nostalgic about this roast recipe when my subagent would discuss meat recipes with him inside his little 4GB GPU."
> Like, the average person who is seeking a recipe needs a photo to know what to shoot for
The menu: Rosemary and garlic roast lamb, Extra-crispy roast potatoes, Honey-roasted carrots and parsnips, broccoli.
The photo: [1] Roast turkey, mashed potatoes, baby carrots, broccoli, brussels sprouts. No lamb, no parsnips.
If you shoot for what's in the photo, you're going to have a bad time.
You're also going to have a bad time when you try to make an apple crumble with no flour, no sugar, and no butter, because they're not on the shopping list.
I've heard of other people for years now trusting AI for recipes, and it's a rare area that I simply can't bring myself to take seriously.
I feel like it can only be successful by luck. Either it's reproducing a recipe verbatim that was tried and validated and tasted by a human, in which case, we didn't need AI for that, just a searchable cookbook. Or, it's making one up. I understand that using RLHF has helped improve the quality of questions about history, programming, or TV show recommendations, but I do not believe that there has been some kind of training regime where people prompt the models for "a recipe that uses X, Y, Z random ingredients," follow the recipe, and then score it. And even if they do, I don't see how a model can learn enough from that besides "this exact recipe is good/bad." '1/2 tsp cumin' may be a great addition to one recipe and not enough for another, so improving the output based on a bunch of scored recipes... I just don't believe cooking is an LLM job.
Maybe some other kind of model that I don't know about.
I've had some really good experiences with cocktail recipes and Claude. It's got access to my barcart inventory, general tastes, and general desired level of effort.
I feel like the technique recommendations are probably the most valuable element, but I'm getting rave reviews like 85% of the time now?
There have been a couple times where I was out of an ingredient and it proposed an adjustment that sounded a little wild to me, but it almost always is well received.
You're right, we don't need AI for that, we need a searchable cookbook.
Unfortunately, one does not exist.
The Internet, however, is full of garbage cooking advice, and while AI is quite happy to parrot that garbage, it's not significantly worse than it's sources.
I do that, but I've found more success with using internet recipes as a starting point. The general-purpose cookbooks are, at best, another data point, or an explanation for something glossed over by other recipes. Ar worst, they kind of suck.
And we aren't talking about something esoteric. How To Cook Everything. (4.5 rating on Amazon, 4.0 on Goodreads), one of the most popular cookbooks in the world...
Has its first recipe for chicken cutlets produce undercooked chicken (15 minutes at 325F has not resulted in 165F internally).
Worse yet, its chile recipe for some weird reason involves boiling and simmering a whole onion with the beans before throwing it out (WTF). It then has you drain the beans only to add that drained water right back (WTF?), unless you optionally replace it with tap water (WTF was the point of boiling the onion then?), so that you could bring them to boil again. The recipe barely has any spice besides actual chile peppers.
This... Is not good or helpful, but is incredibly opinionated with a bunch of bullshit steps. Like, yes, I could tweak that into something that's not insane, but why would I choose that as a starting point..?
I get it. Compiling a thousand-page cookbook is really hard. But this... Ain't great.
Once again, Grandma had it right. The Fanny Farmer or Betty Crocker cookbook with dozens of little papers slipped in it and/or marginal notes everywhere is the way.
I wanted to correct you and bring up my favorite website that had filters for the ingredients you want/don't want to use, type of dish, and time for cooking, but it appears that in a last month it was sold and now it links to the slop site with no useful filtering. I used this site as a quick lookup for recipe ideas for 7 years now.
Fuck. Does nothing good survive on the internet anymore.
Same here, I’ve been cooking for myself and later also my wife since 2005, even when baking I rarely need exact amounts, for cooking I essentially never follow the recipe directly and adapt them as needed, without ever changing them from the original.
I mostly don’t even need a recipe from AI, just the title and a 1-2 sentence description.
LLMs for recipes are good if you just want to make a thing, not necessarily the best version of that thing. It's nice being able to ask questions about certain steps as well, or ingredient substitutions if needed. It'll make a nice, average, recipe (unless I suppose you prompt for something). That's perfectly fine for most cooking, and often what you want when you're making a dish for the first time.
It also doesn't really like tell you how to actually cook anything. Good or even just decent recipes tend to tell you the temperatures, times, techniques and orders of ingredients to do things for the best results. Unless I'm missing it this is just really broad strokes. But I guess the goal in engagement so they want you to ask a lot more questions to actually figure anything out.
Holy shit. I read this on mobile so didn’t actually see the recipe part because you have to scroll way down to see it. It is awful! The shopping list makes no sense and the recipe instructions literally have an “Everything else” item. Yikes…
For me, the text-only recipes I already get out of ChatGPT today are good enough
You’ve illustrated something about use of AI that really bothers me. The lack of discernment by its users. Sure, the pictures and diagrams are impressive, but did you actually read and evaluate the output?
Genuine question - do they? Unless it's something like a fancy cake then instructions should really tell you everything about stacking. Half the recipes are mixed anyway (curry, stir fry, stews, ...), lots of other are either stacked or separated on the plate. So apart from the really exceptional stuff, to people really need to know "what you shoot for"?
Maybe, but it's a reality that most of the general public do prefer cookbooks with pictures. I suspect openai would end up doing this anyway just by targeting the general consumer. It's configurable.
I'm more worried about picture accuracy: usually the benefit of pictures in recipes is to see what you're aiming at (eg, how finely chopped something is), but I don't know if the image model is up to that level of detail.
What’s interesting is how those that over index themselves often miss the obvious which could simply be some things (cook books) are quite helpful with picture. If I am flipping through the cookbook, how would I know what Boeuf Bourguignon is and what good looks like.
Yes, pictures OF THE FOOD. Not unrelated pictures of another lamb-roast made with another recipe.
Good cookbooks do proper food photography, of the food made with the recipe. That's what sets them apart from slop, be it human or AI slop, cook books.
Why is it impossible to create? Are we talking about some meal that King Richard I had (https://en.wikipedia.org/wiki/The_Forme_of_Cury is the oldest cookbook I'm aware of which covers meals Richard II would have eaten), and thus we can only guess? Are we talking about the long extinct dodo bird (I'm not sure if humans ate them), cave bear (I've seen it claimed humans hunted them to extinction, but better sources say it went extinct before humans), or something else that we can never make?
For almost everything else we can create them - just give a reasonable cook a kitchen and the ingredients and they will do it. For a picture we then have an artist do preparation to make it look good, but they should always start with the real dish.
I've got visions of The Truman Show. In the middle of my question about vacation ideas ...
GPT: "Why don't you let me fix you some of this new Mococoa Drink? All-natural cocoa beans from the upper slopes of Mount Nicaragua. No artificial sweeteners!"
That's pretty much the Gemini experience right now. It is unusable. The funniest bit is that the LLM is not in on the joke so it will act like it never happened... Google once again wrecking their own products.
I'm not saying anything, because I've never seen or heard of your experience. I'm only reporting mine. I'm not on the free tier, maybe that matters ¯\_(ツ)_/¯
What scares me is that we're already half-way to Idiocracy world, except instead of "it has electrolites" we have "it's organic" and "it's not ultra-processed food".
I don't trust modern fertilizers and pesticides blindly, but I have never successfully digested any inorganic food. The 'organic' label is an old idiocratic choice.
I don't trust organic fertilizers and pesticides blindly. At least for the "chemicals" we do studies, often organic gets by with we have always used them even though a quick glance at the SDS (if there is one, otherwise just the ingredients) suggests it is even worse but nobody studies the safety.
The problem with Claudish isn't that it's technically poor English, the problem is that it's information-sparse waffle most of the time. It's like dealing with a colleague who needs several paragraphs of tedious preamble to make a point when you just want them to spit it out and be done with it. It also does a great deal of vague gesturing at thoughts without actually instantiating them directly, for example by using the same tired metaphors to the point they don't actually convey a thought any more.
If I say Claude talks like a meringue it's patently obvious what I'm trying to convey, sweet but full of empty space. If you hear that same metaphor day-in day-out then after a while it doesn't call up a thought at all, it's just an annoying verbal habit.
My go to phrase that I’ve found to work well is “please restate in high school English with direct declarative sentences and no parentheticals or references that require recalling previous turns of this conversation”
> ASD-STE100 Simplified Technical English (STE) is a controlled natural language that is designed to simplify and clarify technical documentation. It was originally developed in the 1980s by the European Association of Aerospace Industries (AECMA) at the request of the European airline industry, which wanted a standardized form of English for aircraft maintenance documentation that could be easily understood by non-native English-speakers.
STE is very hyped for AI but has its own smells. It's designed for aircraft maintenance manuals, not software - the two don't share the same vocabulary used to "simply" describe something. I've had a coworker try and "simplify" our documentation with Claude applying STE and the results were equally appalling to read as the Claude-ism laden docs that came before, despite the reading level metrics going down on paper.
I've taken a liking to pointing agents at the Simple English Wikipedia editorial guidelines...but so much of that is for making up for shortcomings of modern Anthropic models. Codex running 6.1-sol explains things so much more clearly that I don't feel I need to assert a style guide on its output.
Gemini 3.8 for me always outputs in a concise way at the level understandable for any computer science masters graduate. Clear, to the point, a bit of formula for background when really required instead of the same formula in Prose, etc.
- Gemini 3.8 has a beautiful answer. Typography, use of lists, brevity, style, even the font choice on the web harness -- all of it is A tier or even S tier. Notice I didn't say answer quality. The answer is always extremely mid, and the harder the question (or more effort needed to answer it), Google just quits. I give it a "C" on answer quality
- ChatGPT (Chat, pre 6, I assume 5.6 although they commonly hide the model picker): Ugly answer. Runs on printing page after page of unnecessary side notes. Mostly just walls of texts in paragraphs with little thought to information design. However, inside of that mess of text is almost always the answer I'm looking for, and it's almost always incredibly better than the truncated, desgined Gemini answer.
For a while I ran all of my question queries on Gemini and ChatGPT, but I noticed that I basically always picked ChatGPT's answer head to head.
Interesting that this hasn’t been my experience. ChatGPT has often gave me with pages of text dripping with confidence about a solution for my problem. Diving very quickly into implementation details. Even if I tell it to stay in the problem and product fit level.
Gemini on the other hand somehow always outputs at the level that I am expecting and reasoning.
I even found it was capable of stepping back from code implementation and bug hunting to “hang on, it’s not the algorithm that is incorrectly implemented, it is the wrong algorithm for the goal.” GPT Sol (5.6,6 and 6.1) just kept hammering in the problem.
I was about to comment the same thing. Why is a picture of a roast helpful in that moment? The user asked for a plan to prepare a meal. They need a list of ingredients not an image.
But the image isn't a photo of the result. It's an unrelated image that vaguely looks like what you could make, with a different recipe, and with some skill.
And you see this in the other examples too. The teach the CLT piece was funny. It's and error-ridden mess that isn't even an explanation at all!
The best part is, someone looked at that and thought yes, that's right. Shows you clearly why programmers will be needed in the future. The LLM might be able to run circles around that person in math, but that person also has no idea when they're being bullshitted.
The fact we've come so far with textual models & the image stuff is still producing stuff that harks back to early-stage tripophobic AI produce (the garlic with roast potatoes here) is interesting.
Don't get me wrong, I'm certainly starting to see some AI-produced imagery today that I can't tell isn't real, but that's largely because a lot of real photography is overproduced & ugly. I have yet to see anything that's aesthetically good.
Of everything that could be automated, Bartosz Ciechanowski really was the last on my list.
In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But it’s absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.
The reason why B.C. became a thing is because the art of drafting died from CAD. The attention to detail, the minutae of walking the reader through a highly sophisticated thing was replaced by short-form video explanations. His work is very much a callback to the days of old. So too, is his turn to be relegated to a relic of his time.
Also, it's hand coded WebGL. It's smooth, faultless and has no peers.
> So too, is his turn to be relegated to a relic of his time.
No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren).
funny if true. ironically just last week i was asking chatgpt about what happened to him and why there were no new blog posts for almost 2 years and it had no idea
Nothing in that "7‑Speed Bicycle" is specific to having 7 speeds. It might as well have said "Bicycle", and even then you get meaningless slop like "Made to keep rolling" and "A strong foundation". How can you compare this to Ciechanowski?
You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?
That's a useful way to evaluate this. A visualization shouldn't be judged by how impressive it looks, but by what it actually helps someone understand.
If the same explanation works regardless of whether the bicycle has 1, 7, or 21 gears, then the model probably hasn't understood what needs explaining.
Yep. I hate everything about this. Just pure slop. Fancy visuals that mean nothing. Text that sounds impressive but has no purpose in educating the reader even though it's supposed to be "an interactive explainer". Just slop, slop, slop.
AI bros will copy anything that's even mildly successful. They're like a teenager trying to impress their girlfriend (or their mom!) by saying "see! I can do it too!"
I disagree, I always run into issues with the embedded YouTube player. No way of sharing. No clear way to share the link. Youtube logo covering half the video. To name a few.
There are many sites with better video players - they're from the industry well-known to be at the forefront of innovation in audiovisual technologies since before modern computing era.
It's just most of our industry pretends they don't exist, and instead implement toy-like video players to distinguish themselves, I think.
The progress bar is over the video and it's easy to tell where it starts and ends. Most devices control volume, the video volume separate is just confusing matters for most.
No issues with play and pause my side.
Sure as hell beats stupid Instagram style videos where you have no way to skip ahead. I think this is what future video is going be like, buckle up.
And you only ever listen to one thing at a time, from one program at a time, and all audio content ever has its loudness perfectly normalized to a global standard everyone adheres to.
I don't want to adjust my system volume to change the volume of an over-loud video. My system volume is set for my comfort so my notifications and other applications are the right volume.
A volume slider on a video player is a really basic feature that every one should have.
Set for your comfort doing what exactly? You're basically just prioritizing one source of media over another
As TeMPOraL has pointed out you only listen to one thing at a time. Notifications etc. also have their own separate volume settings generally.
> Relative to their respective GPT-5.6 counterparts, GPT-6 Sol (October) shows a statistically significant regression on standard self-harm, while GPT-6 Luna (October) shows statistically significant regressions on standard self-harm, gore, and sexual content
> "GPT-6 Sol (October) and GPT-6 Luna (October) show an improvement on helpfulness on legitimate requests relative to prior models, though it scores lower on some safety requests"
Concerning how there's significant regressions on so many critical benchmarks, but it is newer and creates UI, so must be good.
It’s crazy that a company can document these safety drops in a PDF, ship the model anyway, and focus the announcement entire on shiny new UI
> significant regressions on standard self-harm, gore, and sexual content
Good. Maybe I'm the only one, but I feel like these have always been stupid measures to waste the finite time of "AI safety" research on anyway. They amount to whether a user can, if determined, manage to make a chatbot say, or depict visually, some taboo thing. Frankly I'd rather just have some cheap classifier judge each chat response after it's generated, with a limited context "Is this response encouraging self-harm?" and skip all the rest of these. Whether some random pervert can make GPT-6 spit out an erotic fanfic story or image has zero impact on the rest of the world. If this particular model won't do it, other models already exist that will, and the determined thoughtcriminal could always just write the forbidden words themselves, or photoshop something taboo.
Maybe the term "safety" has been usurped by those who feel that 'safe from the possibility of being offended' is the most important kind of safety, but it's like worrying about the wallpaper on the Titanic, compared to actual AI safety concerns.
They aren't. They are saying that putting this safeguard at the level of the agent execution is a waste of time, particularly that of the researchers' expensive manhourse, especially now that existing smaller models can just evaluate the output.
It's fine if they seek to avoid that. But those are perfect examples of things a cheap output-filtering model can be responsible for enforcing, without wasting the time of the researchers who are working on the training and alignment of the model in the ways that actually matter.
Most "bad outputs" in areas like 'self harm' or 'violence' or 'sexual' whatever are a result of people deliberately trying to elicit those responses. As such, it's meaningless whether an LLM writes the Bad Thing or if the user writes it himself.
From their point of view it's more about optics. "chatGPT encouraged my daughter to stick her fingers down her throat after meals" is not a headline they want.
In general though you do have to figure that vulnerable and naive people will use it because of the degree of market penetration they're aiming for, including minors etc. and they have some responsibility around that.
Well, those mentioned regressions are in a section entitled "Safe Completions for Users Under 18" and immediately after your quote they say:
> To mitigate the risk of producing disallowed responses for teens, we apply
an additional classifier-based block to responses that may contain self-harm,
sexual content, and gore; this mitigation is not captured in the evaluation
results above and improves safe responses.
It's good that they mention that. Most of those "guardrails" were overeager and caught too many false positives, crippling the AI. Like how Claude 5.5 refuses to do shit if a prompt contains "reasoning" because it thinks you're trying to hack it etc.
Good. "Sexual content" should mean outright genital focused porn, not bikini girls. Elsewhere on the internet it's foretold that women will oppose AI images of women because the AI images are too attractive, and I'm starting to wonder if that's true.
> women will oppose AI images of women because the AI images are too attractive, and I'm starting to wonder if that's true.
Right... women will oppose AI because it's too hot and not because gross chuds have been using it relentlessly to create deep fakes of their likenesses and CSAM of their kids.
This is actually a defensible position. It's not either-or - the "gross chuds" doing things you mention are worrying, too. But providing unrealistically attractive depictions of a human body is also a problem. It's been a problem for ages - all the girls and women who try to compete in beauty with celebrities and models (without a team of specialists supporting them), only to ruin their bodies through anorexia or bulimia, or their minds with depression, deep insecurity, and an inferiority complex, exist and deserve mention. AI(-generated images) is not the reason, but fits "well" into the preexisting social problem, and has a chance of intensifying it on a global scale.
I've had the most success with GPT explaining things to me by making it take a few sentences at a time back and forth, instead of reading full write-ups of whatever I asked. It also often poisons the conversation if it misunderstood some part of the question, and I can lead it better by continuously questioning its statements. It's also more engaging that way.
I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!
Id love to learn what prompts you’re using to do that. Explanations one at a time. Thats something I’ve thought would be helpful before but didn’t know how to achieve it.
One of the things that I learned the hard way is to not to combine requests, but fork chats _a lot_ and ask singular questions. Also start new sessions with either explicitly created handovers or even just explaining current state. The more compactions I see the less and less trust I have in it’s current understanding of what we’re discussing
Nothing fancy. What works for me is trying to lead the conversation: asking for definitions one at a time, asking how they differ from its previous answers and pointing it out when it conflates its own answer. I generally feel like you can't let it dictate the pace, it's not very good at that yet.
I used to edit my prompts and undo messages if it misunderstood something but that's been failing me recently.
For me this works: "One thing at the time. This is a conversation, not a lecture, try to balance the amount of words you write versus the amount of words that I write. Don't overwhelm me with 10 pages of prose where a single sentence would suffice, brevity is a virtue, not a defect." I've had to tweak it a couple of times, probably because of different versions of ChatGPT but it is usually a variation on this that will do the trick. Besides being far more interactive it takes the frustration down quite a bit.
It slows down to the point I find it unusable so I ask it to summarize as one cut-and-pastable block, copy the block, open a new session and kill the old one. This usually happens after a few hours, probably because 'thinking tokens' even if they are not displayed crowd the context window.
I've been calling this "disposable UI" or "paper plate UI", eg something meant to be used once. One thing I'll be curious about is overzealousness to produce this, when sometimes what you want is just a simple response. Overall though I'm a big fan of it, if it can be provided fast enough. I'd be curious on how much impact it has on latency of a response.
That's my biggest gripe with it. If I ask a quick throwaway question, I'd really rather not wait for the LLM to build a test framework for its composable principles-driven components framework and WebGL/WebGPU abstraction first.
I'm not a fan of OpenAI / Sam Altman, but I love their blog posts. The team and whoever decides how to do these presentations, is on point. The only other company that has amazing release pages like this is Apple, I think I remember hearing that they probably hired someone from apple who used to do release blog posts there too.
What's funny about "Intelligent UI" is I said like 2 or more years ago, that these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
Why did they bother hiring someone to write at all? I certainly wouldn't invest $1T into a company with so little confidence in their product. When the chips are down and they need to write something important, they DON'T use AI?
That's just picking a fight. Whatever we think about how amazing or terrible AI is, it's pretty reasonable to understand it is quite literally the weighted average of humanity's output.
Also reasonable: it's possible to hire a world-class ___ to do that job better than AI.
It's not a weighted average though. I see this repeated all the time. Between synthetic data, custom produced training data by experts, post-training regimes etc the models are far more diverged from that baseline at this point, at this point it's mostly a right shifted curve.
I hesitated to write that line because I don't know the math. So thanks for clarifying. Main point is the output is along some distribution and I'd say AI-evangelist themselves see "good enough" across everything and anything as a feature.
Obviously everyone at Anthropic uses a ton of AI in everything. The point here is that they clearly use it well, and that they use it as a tool to augment and enhance what they do rather than as a replacement for effort. The blog posts don't read like the output of a prompt because they're probably the output of dozens of prompts, rewritten and edited, by someone who's using AI to actually write something worth reading.
Anyone can having blog posts as good as this by using AI. Just not by zero-shotting it with a 2 line prompt. It still takes a lot of work.
I take the opposite viewpoint. I think pages like these are horrible, and a demonstration that web designers have too many tools at their disposal.
But this is just a marketing page for some tech company, right? If only it stopped there. I have to deal with this nonsense in news "articles" as well on occasion, when some web designer intern is allowed to larp as a journalist for a day.
Sometimes you want more than text. This is one of the interesting divides these days - and I say this as someone who spends a lot of time in the terminal. Computer interfaces and interaction did not peak with the VT-100. Sometimes, there is a very legitimate need to show tables, graphics, interesting graphs, and to use colour and shade to draw the eye and lead someone through an experience.
This is why I use an IDE instead of a TUI agent... I'm on a computer with a multi-megapixel display, I want to use it. If I could be driving the same process with my Macintosh SE as a serial terminal, what the hell is the point of my recent Macbook?
You say tables, graphics, graphs, and colors, but I mean things that move when they don't need to move and interfere with expected UI behaviour, like arbitrarily staying in one position when I am using my mouse wheel to scroll downwards.
I should have been more precise than saying "just give me text", but it was what came in to mind when I wrote it, as I was thinking about what an article is meant to contain as its base element.
The text-only purists on HN are getting very tiresome. It’s seriously in every fucking post that links to a site with JavaScript enabled. We get it. You’re elite. Now stop talking about it.
You could say the same thing about Anthropic because I don't really see much differences. Not sure about amazing part but both are high quality and I probably like Anthropic's aesthetic more
Anthropic's aesthetic looks good by pre-AI design standards, but it's so overused now that I find it off-putting. It's like the 2026 equivalent of the Twitter bootstrap CSS from 15 years ago.
I feel like today's tailwind is the spiritual equivalent to last decade's bootstrap. Except that last decade's bootstrap webpages have a human behind them who usually put thought and effort into making the page useful and correct, whereas today's tailwind pages are often vibe coded slop with all the associated issues.
I don't care if Anthropic keeps using it, it's part of their brand at this point. The thousands of other developers using Claude Code to generate designs should not just accept the default output, which looks very Anthropic-ish, and actually put some effort into polishing and differentiating their product. It has nothing to do with usability and everything to do with a basic sense of good taste and aesthetics.
> these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
I would prefer if the AI companies stuck to just creating better models and making them as cheap and accessible as possible. Let others build the products. I don't want 1-2 companies to own every product in the world.
That was the objection to Microsoft controlling both the dominant OS and the most popular productivity applications running on that OS. Took the web then mobile devices to really shake that up.
Technically other companies do, and I lump them in when I say "these AI companies" I am talking about any company that makes and builds any AI harness, coding or otherwise, I have not seen any innovating in the areas of the UI / UX of using these chat models in a meaningful way.
Apple should be extremely worried about the AI companies making big improvements in UX. This work by OpenAI is a step in that direction.
If the main user paradigm becomes a chat interface conjuring whatever UI elements needed to best accomplish the task at hand, the entire paradigm of OS, UI frameworks, apps from an App Store to accomplish specific tasks, all come into question.
I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area.
Steve Jobs wasn’t meant for this era, but his instincts would have been interesting to see in the AI landscape. He was an artisanal designer though, and our current era is always for the investor’s bottom line.
Apple absolutely messed-up by abandoning OpenCL's early ML research efforts to oppose a CUDA monopoly. They messed up a second time shipping a raster GPU with Apple Silicon when CUDA and Tegra had proven that GPGPU was mobile-ready. Then when the ARM datacenter had it's moment, Nvidia's Grace ARM CPU displaced billions of dollars in sales that would have been Apple's if they didn't mess up the fastest CPU in the world with macOS. Then they depreciated the Mac Pro, which seems like a mistake since it had the potential to outsell Nvidia's ARM datacenter chips if it ran Linux. According to the rumor mill, Apple's M8 chip will finally be the one that takes GPGPU seriously, after a decade of Apple's innovative ship-second mentality.
Apple's management isn't infallible whatsoever. Many of them are petty, blinded by politics and/or obsessed with their legacy more than they care about profits or quality products.
It's not something that gets talked about a lot yet, but I've a feeling that this is where the biggest disruption for the entrenched and geriatric operating systems is going to come from:
No normies could get an app like this into the store, afaik. The do-all, dynamic everything app?! Think of the fees that will quietly never be born? The profits that never get to land on a balance sheet.
On the video, clicking on the play icon mutes the sound, and clicking on the loudspeaker icon to unmute, stops transport. Is this "intelligent UI" or just a sick joke?
yes the eyebrow is the tell. I never knew anything about eyebrows in design as I am not a designer. Now they are everywhere. Everything has an eyebrow. It is ridiculous but at least it is an AI smoking gun.
I loved them and I placed them a lot in my designs — if you also include my love for em dashes you can well understand that I feel my own character has become a clanker…
Em dashes got taken from me as a result of Chat. I used them all the time now I can't without it making people pause to think I'm a bot. Still a bit annoyed or sad about that tbh.
I had to look this concept up. Seems to me like you could just as easily fit the "eyebrow" words into the main headline with a colon and a bit of ingenuity. But then, that runs the same risk of getting repetitive and AI-tell-ish.
Any details about the streaming ui lib they mention or about the format of the ui code? This could become really annoying to be fractioned among model providers with custom RL that required independent tools like open code to use each providers custom streaming UI libraries.
Burning question: why is this worth paying for as a layman? I can sometimes find a use for all these LLMs for software development, but I can never find a good use for any of this outside of "slightly better search engine".
As a complete layman---not worth it. We live just fine. Medical researchers slave night and day to keep you healthy.
But if you have any ambition at all---if you want to map the local school board, or get a quick understanding of your finances, or set up a robot in your backyard---then, quickly proves useful.
No one said that though, or even implied it. They said that if you have any ambition, AI is going to _prove useful_. How do you go from there to "ambition hinges upon using AI"?
It's just the logical contrapositive. "If you have ambition, you find AI useful" necessarily implies "if you do not find AI useful, you do not have ambition."
I genuinely think the frontier labs are just throwing shit at the wall to see what sticks. All of the improvements since Fable 5 seem marginal. Most of the quoted AGI "risk" is coming from super large agentic models that are juggernauts that can brute force their way through any problem faster than a human can. Despite that they're no closer to AGI than they were a year ago; the technology's limitations are apparent the more you use it. Doesn't mean it's still not extremely dangerous and/or capable, it's just not the "replace a human in 99% of cases" thing. There are obvious tradeoffs besides just the models being faster and capable of information retrieval and generation at a rate exponentially faster than a human's.
My long-term hope (and call me another Zitron if you wish) is that the labs' hype dissipates if there's no serious improvements beyond stringing together agents to ram through brick wall Millennium Prize style problems and we all just use on-device AI for 99% of cases where it's useful, researchers, governments, militaries, and universities can pay for the more complex models, and image/video gen dies a slow death (if Congress will actually legislate and/or SCOTUS decides that training those specific models does not qualify as fair use unlike training LLMs) and cost for compute rises over time as public interest in the tech sours.
LLMs are such marvelous and incredible technology squandered by genuinely deranged Silicon Valley cultists who are trying to use them as a Trojan horse to force their antisocial visions of the future upon an unsuspecting populace. We could all have just invested in on-device compute and nobody would have lost their jobs in pointless layoffs, productivity would have increased, and maybe we would have forced a conversation about when and when not to use AI and it wouldn't be so ever-present in use cases where it actually does harm to the consumer and society. But sama, PT, and dario just wouldn't have it that way, would they? Because if that were the case, there's no prospective hope where they can assuage their deep-seated insecurities over being antisocial and off-putting by reassuring themselves that there's an imminent realignment of society where they will hold all the cards when the dust settles.
A lot of the marketing push (both from the frontier model companies and from people trying to sell specific "solutions" wrapping an AI model) these days involves getting laypeople into "software development". Well, not in the sense of iterating on an idea and critiquing it and having a real idea of what software should be like (and certainly not trying to design something that someone else could want); but "vibe coding", oh yeah. A model in a "chat" environment can still write a few hundred lines of code for you and also walk you through "installing" and operating it, if you're persistent enough in explaining what you need to be taught. This is, of course, terrible for "security posture" (but there are already so many other issues there…) but it's potentially very useful for a lot of people. They just have to get the idea in their heads that it's possible to have something on their computer that helps them solve a personal problem, even something that didn't exist until it was asked for.
Maybe I don't get it yet, but why do you need a new model for this? Wouldn't this more be an issue of "harness"/chat UI engineering?
All the current-generation models can oneshot HTML/Javascript frontends of comparable complexity (though maybe not as polished). The only difference is that you have to specifically ask them, and the result is not embedded in the chat.
It feels like it should be easy to have a harness provide a "show HTML widget in the chat" tool and a skill and/or system prompt that instructs the model to generate such widgets on-the-fly if the response could benefit from interactivity.
I’ve hit model not available limits for GPT 6.1 Sol multiple times over the last week. It’s never for very long, and I can switch to Astra, but it seems like OpenAI is struggling with capacity. This has happened with my personal $200 Pro connection and my Codex enterprise connection.
I had a codex session stop in the middle because the auto-reviewer timed out with something like "auto-review not available at the moment". The capacity issues are very clearly observable since gpt6-astra launched.
I really don't like having to open Work sessions for one-off questions, only because they're limiting what models they put in the chat. The older models are just too dumb for some things.
The whole Work/Chat/Codex split is maddening in the way it's implemented. It's a pain to switch back to the right project, it's a pain to switch forward, it's a pain to try to remember which chat was in what.
Given the extreme short time between 6 Sol and 6.1 Sol, I suspect they don’t actually have much in common and 6.1 is a heavier model rebranded as Sol in a panic response to poor agentic capabilities of 6.
I suspect that 6 sol was a better version of 5.6 terra (and note that in the 6 sol and luna release, they took out terra, and price 6 sol at 5.6 terra pricing), then the backlash from lesser capabilities made them roll out 6.1 sol as the actual 5.6 sol - size modee.
6.1 is Astra minor. Way more capable but also way heavier+slower. Its really 2 different models, they just shipped it as sol to recover from the gpt6 disaster lunch, where they tried to pass terra 6(or a cheaper model) as sol but it was worse than expected
It's purely circumstantial, but 6 Sol supports no reasoning (same as 6 Luna and 5.6 Sol/Luna), while both Astra and 6.1 Sol do not support "reasoning = none".
gpt-6.1-sol was released 7 days after gpt-6-sol only because they were bleeding to fable-5.1/opus-5.5 from Anthropic. It was a desperation move and certainly ate into their margins massively. Since their margins on gpt-6.1-sol are much slimmer than on gpt-6-sol they have to use it somehow to preserve compute and make profit. Frankly I think OpenAI is a mess and is playing catch-up with Anthropic… I don’t know how they will recover.
It feels that cgpt 6 is giving less exhaustive answers than 5.6. Both in high. I believe that they are using Terra for normal chat and even if 6 answers are coming faster, they also are much less detailed. What I like with cgpt 5.6 is that for each question I send it, it checks coherence with the rest of what we've been going through until there.
We are becoming more and more as a tool for sth to be done rather than the brain behind it. At least I start to feel this way. The joy of discovery, exploration and experiencing at first hand... It is slowly diminishing for the perfection of the quick outcome.
Thank god. My primary use of ChatGPT is meal planning and while it’s great for recording, recipe lookup, etc I always just wanted it to be able to make checkable grocery lists.
I eventually just had it make a skill for work mode that would build a mini checklist app, but it was slow and felt janky needing to remember to switch to work mode.
Why not just build the checklist app once, and set up a way to create simple data that the app imports? Or even just have it import from plain text, one item per line?
Is this OpenAI catching up with Anthropic artifacts? At the same time they say "We’ve trained GPT‑6 to compose responses using text, visuals and interactive elements...", rather than a harness.
I'm excited about generating UIs (and apps) on demand, because it's one step closer to devices just morphing to the interface you need. I want to talk to my phone and have it generate the app or interface I need. I could see Android adapting to this reality well before Apple.
I find this view fascinating. After decades of observing users getting lost by minor changes to UI or sometimes just different "paint" I am now to believe that we will all enjoy using completely custom UI customised not only to what I would want but what the app will think I need at particular moment and get on with it just fine.
I actually do think generating UI has some future, but I remain sceptical it can realistically go as far as promoters claim. It is especially bizarre to me when this is done by ostensibly UX people.
I am with you. For example - this post has a comparison of a recipe app from 5.6 and 6 for comparison. They look different. Now if every time I ask for a recipe, it decides to render a new UI because of either the model changing or the recipe goes from stove to oven or something, it’s going to become annoying very quickly. It just seems like he current generation of models have been tuned to create better UIs than previous one and people interested in the UI are getting hit with excitement.
Based on current capabilities, I'm skeptical as well. However, over time, I do think they will get better at understanding our intent. Also, I would be more forgiving of an app I vibe coded in this manner, that's only meant to be used by me. Of course that assumes appropriate security guardrails are in place.
I really do think we are headed toward a world where stuff is hyper customized like that.
As a developer, it already feels like this with internal company tooling. Anything I have the source code to and a build pipeline setup for, I can just pop open a new tab and tell it what I want changed, and it just gets done for me.
Even my home assistant setup at home now, i hooked up codex to it via an MCP, and when we put out the blow up haloween lawn ornaments I was able to just pop open codex on my phone and tell it to throw together an automation to run them for me, and 5 minutes later it was done, with an override button on my dashboard. It's a shockingly nice workflow!
Wouldn't it be nice if you could just tell your app to reconfigure/rewrite itself to remove those "what the app will think I need at particular moment" features that you dislike? Maybe have it remove a bunch of the whitespace, getting rid of the lazy loading infinite scroll, whatever you want. It's the GPL dream, except you don't need any coding ability, just a subscription to a megacorp.
I do the same with Home Assistant and my many EspHome devices. HA works well, and it has a lot of nice features and extensive community-built integrations, but configuring and laying out the UI is a disaster. You can choose between WYSIWYG UI tools that are limited and cumbersome, or you can deal with raw YAML that stretches far beyond what YAML should be used for with deep nesting, and it's poorly documented to boot. It feels like a proprietary workflow engine, where you'd rather just write it in a real programming language.
Agents cut through all that crap and allow me to make changes without learning a new overwrought configuration language that thrashes syntax every few months. The fact that it is open and fully modifiable redeems it, because I can ignore all the annoying parts while achieving what all that was envisioned to allow.
What emerged for me as of late, at least with Opus 5.5, is creating web interfaces as aids during development. I recently had to dial in some audio mixes using ElevenLabs audio and do some complex audio cropping and concatenation. As a interim step Opus built a website that allowed me to compare versions based on different rule sets and then really dial in the version that we decided to ship. That then led to 'us' landing on the right algorithm to build out the audio files.
I finally connected my home assistant server up to Codex the other week, and it's been pretty incredible. I can just send a message to codex and in a few minutes my dashboard is updated, an integration added, an automation setup, and sometimes all of them together at once.
And I don't need to build or debug complex automations any more! I can just tell some AI to make it so it auto-locks the house door when my phone isn't tracked at home, but also to check the wifi and if there are people using the guest wifi then to instead send me a notification asking if I want to lock the house because guests might be there.
I'm curious to see how useful this is. I feel like in practice the existing versions of this feel like they get in the way. While I'm sure I've had some situations where a visual would be helpful, there are two situations where I do not want it:
1) I want a quick answer, and I don't care for the boilerplate UI. For example, if I ask how to make pancakes, I make them all the time and just want to a quick reminder on the ratios, but it might trigger a full UI that I need to sort through to find information.
2) If I ask a to me unrelated to UI question and it triggers a big UI build that is completely off topic for my question (meaning I'm desperately pressing the stop button and prepping rewriting my query)
Or the other one I see, for example if I look up a unix command like:
"ls all hidden files in the /xxx directory"
And I get back:
"Sorry, I am am unable to find /xxx in my current environment"
People make simple text websites → Google comes along and indexes all websites → People can freely and easily find information! → Google slowly perverts the incentives with ads → Websites replace simple text with complex UIs and paywalls → People can no longer easily find information → OpenAI indexes all these complex websites and turns them into simple text answers → People can freely and easily find information again! → OpenAI perverts the incentives with ads → OpenAI replaces simple text with complex UIs and paywalls → People can no longer easily find information.
Intelligent UI for everyone is Bret Victor's vision for computing UX finally coming to fruition.
I've always found his work impressive, but for the most part it was all prototypes and proofs of concept.
Having such interactivity available to everyone now is nothing short of groundbreaking and another indication AI works as the great equalizer by democratising capabilities previously available only to a select few.
It has the potential to be that. Not sure this is it yet though or if openai's intention is even eventually to do something as useful. I think parent comment might mean what is presented in the links at https://worrydream.com/ in years 2011-2013.
Really exciting to see GPT-6 pushing intelligent UI forward. Making these capabilities accessible to everyone could open up a lot of possibilities for developers and users.
I've been thinking a lot about this transition happening, and I hope that everyone is recognizing that this is what the web will probably look like very soon.
URL-->user request--> answer & visualization.
No more static menus, buttons, lists, etc.
There's probably some good work out there helping small businesses adapt to this change? If done right, the real-time token use to generate dynamic pages can be minimized.
Honestly it feels like AI companies are still searching for the killer idea that will attract the general public (beyond "be my AI bf" or "better google").
Getting a UI that explains the parts of a bike is ok I guess, but isn't it simpler to get an actual breakdown? Google "parts of bicycle breakout" gets tons of useful images instantly.
Getting a specialized app to split a bill? It was already trivial to put in a calculator if we cared to go item by item on the bill. Having to provide names and tag every item as I go is just more work. Usually real people just go $total divide by 5, I had more, let me chip in an extra $10.
Same with booking travel and wedding plans, these aren't things people would even delegate to a trusted friend usually, much less a one-off request to an AI bot or custom UI.
Most of the useful tasks it can do right now are research, technical question/answer, coding. In terms of "build a flexible ui that solves real-world problem", if existing mobile app isn't useful in this arena, then its unlikely a completely custom UI will do the job.
That said, I've done a few small ones like a quick one to practice alphabet of a foreign language, or prototyping a web game, and the like. But ultimately its nothing that is worth trillions of dollars.
Is it just me, or did they make the examples in the video comically unintuitive? Painting swatches without color, a guide pamphlet without a map, assembly instructions without diagrams?
Is this supposed to be humor that I'm not getting, or are they really saying, "look at this solution to a problem you would have if you did normal things in the worst possible way"? It almost feels like the video is trolling us.
It seems fairly obvious that it's meant to be an analogy for ChatGPT only supporting textual outputs until now (disregarding that it had some limited capabilities to show static images).
I might be wrong but I remember that when Gemini 3 was released they also kind of promised something similar. But I’ve never actually seen it in real conversations. I’m not a heavy Gemini user so it’s possible it’s there but I haven’t seen a single demo showing this.
I think most responses here are missing the bigger picture, this is the beginning of the end for individual apps and websites, a big rollup of everything into a single tool. A sad day.
Currently, we have highly specialized apps which do exactly what they need to do. They were created by someone who cares about making the program which works right and reliably. Most programs have years of bug fixes and deterministic algorithms. Many of them run locally and waste a minimum amount of energy to do the calculation.
But this is different - this encourages you to make one-off apps for each little thing you want to do. Imagine someone making a "bill splitter" app every time they go out with friends and how much energy that would waste... Just to do "14+16" that you could do in your head.
Also imagine it coming up with new UIs every time (depending on the prompt/context), and even potentially gaslighting you and doing calculations wrong because no one ever tested it before.
AFAIK, they already had a simpler form of this. It was kind of an obvious next step. Now connect it up to tools so that we can again comfortably do the things that are more precise by hand! The “pick a color” or “select the width on a slider” use case is coming closer.
It reminded me of when Twitter introduced Stories.
I just want to read raw text without pictures, but that's probably not most people's preference. Understandable and sad.
If the diagrams and descriptions are actually well-designed and written, I'd far prefer those. Especially one where I can ask clarifying questions. I'd much prefer a Haynes with some of the images lightly animated than having a 10 minute long video with a long intro for monetization reasons about how to do some simple thing where you can barely see what the mechanic is actually doing in the cramped, poorly lit, 360p video is trying to look at.
But getting actual service manuals cheap/free these days can be tricky, and their quality can leave a good bit to be desired.
he he boi that's right cooking videos are entertainment if you actually cook you read something called a recipe (it's the algorithm that returns the food you want)
Not sure about the cooking one but repair videos are horrible. People with thick accent, the worst camera, shittiest lightning and bad angles where you can't see what you have to do.
The medium is the message here. As others have said, chat output largely sucks for anything but the most basic responses and we've done little to improve upon that foundational UI in the last couple years. We've _added_ a lot for specific domains like coding and document editing, but the primary content normal users get back from the chats is verbose and uninteresting. This is a step in the right direction.
Gettin closer to fully dynamic interfaces for a lot of software. Hell, give me a mode in Google Docs that takes every pixel of chrome away then vibe the rest as I need it. Persist across new documents going forward.
IMO visualization is such a critical part of explanation - we have created an entire suite of tools to visualize - presentations, graphs, formatted docs, websites, svgs, interactive blocks, apps.
To that front - model providers generating visualization elements (i,e taking over the tools for interaction and visualization) is just a natural next step in value capture.
Anthropic launched docs, presentations, sites and am sure they have sheets next up their alley. Apps are artifacts. Tool forming is a natural model adjacency.
Wait, is this new? I feel like we all have been doing this with widgets / artifacts for a long time now, am I missing something? Just easier to use / done without asking?
I wrote something like this, though not quite as slick, by having the model use json schema to describe its results instead of text, and then use a json-schema to ui interface.
I am pretty convinced that businesses will be making custom internal software as a norm in the next few years, and this seems to be a slight push in that direction
Alright doing it here, all the functionality we need and no reoccurring yearly costs of tens of thousands of dollars. Took me a couple weeks to wire it all up with codex.
So you never have to assemble a bike or paint a room. How those are not real world examples?
I find their video pretty cool. For people with ADHD or people learning by watching or children this is spot on.
They made fake, terrible artifacts (colorless paint chips, text-only city guides and instruction manuals) to show how terrible that is, and then "fixed" them. The problem is, in reality their examples don't exist. Paint chips have color samples. City guides have maps, pictures, and color. Instruction manuals almost always have illustrations. If what you made is better than what exists, you should compare it to what exists and not some alternate reality.
They're making the point—apparently too subtly—that text alone is a highly limiting "UI" for many tasks. So it's great that GPT-6 can now communicate in a a richer interactive medium.
I'm confused. Are they trying steamroll every developers B2C product or are they wanting these same developers to continue building MCP-App plugins on the platform?
It's much simpler: they want others to experiment (and pay them) and see what works, once they know what works they will drive out of market other players.
This seems like an expansion on something I was getting my agent to do which is use 'cards' to communicate using things like tables, rather than the ascii table stuff Claude code likes to do.
Unsure if this has been done already but the best thing I did was get it to create a list of next steps as quick buttons that it keeps updated in a docked card. It works well but occasionally gets stuck on some un related rabbit hole.
maybe I'm just wrong here? If the tokens <b>come through this text could still be optimistically turned bold until </b> happens.
For elements that manage layout and drawing things like these explainers, though, it's hard for me to imagine how that would work. My immediate thought would be to an abstraction over it, which is what it sounds like they did.
They could also just be creating all the html elements using javascript (i.e. document.createElement, document.body.appendChild, etc) and streaming that into a repl line by line.
Today i wanted to reencode and downscale a screen capture video to make it small enough to send to other people. I had a chatbot generate me the ffmpeg command line, but when will i be able to tell "Apple Intelligence" or some other "Intelligence" just "resize this to 720p, make it mp4 and drop the audio"? You know, like in Blade Runner.
Edit: oh wait, Samsung already did it like in Blade Runner, when they "enhanced" the moon photos until they had details that were impossible to capture with those optics...
Great product idea. I have not paid OpenAI for a subscription for a couple of years, I’m tempted now. I do about 75% of my work using local models and most of the rest using the deepseek 4.1 flash API. That said $20 to experiment with this new product for a month sounds like a pretty fun idea.
I feel like I'm going mad looking at this page. The bike disassembly looks fake to me, like an alien would show presumed bike parts, but I can't tell why and don't know much about bikes. Maybe it's because there are no screws and nuts?
What is it supposed to demonstrate? That the model knows some kind of folk mereology?
That's a good way to describe it. The fork and seatpost are both missing the tubes that run through the frame. It's like someone cut up a photo of a bike instead of assembling it from parts. Even then the front wheel has very little clearance with the frame and the rear way too much. Also inconsistent that the chainrings stay with the chain but the cassette moves away.
This seems like a step towards their plan for building a platform in competition of Google/Apple. Once the generative UIs are polished and more useful than individual apps, their hardware can now ship a device which circumvents the app store moat.
It makes sense they’re doing this - I’ve noticed lately when asking a more complex question involving a lot of nonlinear data using Codex work mode, Astra and Sol will write a fully html document to better display the info with a Cliff’s Notes version in chat.
Wait did ChatGPT make paint chips without the actual color printed on them, just the text description? and now we need ChatGPT to show us what the color is?
I find the video strange. Are we meant to empathize with people who get flummoxed by everyday life? "Oh woe is me, I can neither assemble a bike nor ask a friend for help!" Of all the challenges in my life, the tiny mundane ones are the ones that can be most fun; those are the opportunities to laugh at myself, and to revel in small & happy victories. It's the charm of life.
I mean, I don't think the point was to say that everyday life is hard. I think the video was trying to say that certain types of instruction/communication is very bad.
I see your point for the bike scenario, but not for choosing a paint colors or touring a city (the other scenario in the ad). The common element that I perceived was that these were all mundane tasks in the physical world.
There have been recent developments that allow people to interpret Swift on device, live. With this new launch, will the AI be able to write native code and compile it on device? That would be super cool.
I'm also building something similar for an internal project, based on the vercel's json-render design. It hasn't been too challenging, especially since the release of faster models like 5.6 Luna.
"More than" 20% of the connected world uses ChatGPT each week eh? I guess if you add Anthropic who must also claim 20% and Google, Meta.. that does not sound realistic at all.
The point, I think, is to make fun of Anthropic models, which answer only in text when you ask them how to do something. They’re the competition, not paper pamphlets.
The videos don't display at all, though sound plays, on MacOS Safari. Mobile Safari does work. MacOS Chrome (where they actually tested I imagine) is fine too.
They seem to be pitching something that pretty much all the decent models can already do..?
Obviously if your model is stuck inside a CLI terminal, then not so much. But in a GUI harness (shameless plug for my own one: https://juggler.studio, but I assume others can do this too), you just ask them to answer in HTML and they'll happily draw pretty pictures inline in the conversation. I've been doing this for ages with claude, GPT, Deepseek and others.
In the video, they showed examples of ChatGPT making interactive tutorials on how to fold origami, how to arrange colours/interior and how to assemble a bike.
Supposedly, people were struggling to follow written manuals and they needed an interactive explanations.
I'm not a mathematician and I would certainly love having a tool that would do ELI5 on some complex stuff, but I'm really worrying about using this too often and outsourcing my ability to do stuff to some mega corp.
It looks forced. It doesn’t look bad, but it’s not something I would pay for, nor would I give it access to my computer?
I also believe that a small talented group of developers just out of college can do just as good a job. That is why there is no moat around AI, it still comes down to original thinking and talent.
Is it me or GPT-6 Instant food answer gives the vibe to check the hell out immediately? It's like the it will start to explain a sunny Sunday afternoon from 20 years ago for 5 full pages.
> you stupid silly human, instructions are hard! Who can read "fold this" and "fold that", when you can instead have interactive short-form content with music and smileys to keep you entertained while you offload all the thinking to AI?
It seems like every week OpenAI adds some new feature to the chat UX that wasn't tested, doesn't work at certain resolutions, doesn't work on some platform/browser combos and hogs all the memory/CPU.
Intelligent UI my ass.
Edit: I should add, features that no one asked for, too.
This is one of my _least_ favorite AI behaviors, I am surprised they left it in the example. This behavior where it insists on putting text everywhere that over-explains the context.
for the sunday lamb roast instructions, how much is 1500g of potatoes? What kind of potatoes?
> Also pick up garlic, rosemary, thyme, lemons, honey, almonds, olive oil, gravy ingredients, mint sauce, crumble topping and vanilla ice cream.
what??? how the hell are you supposed to remember all that. Even the before example tells you what kind of potatoes to get. tbh it would be cool if you could just click a button and get an order pre-filled out on a grocery delivery/pickup service. I really wonder who reviewed this post and if they cook, because imagining yourself in that scenario and reading those instructions falls apart very fast
With computer use Both Astra and Sol have been able to make sense of my markdown recipes, and based on those place prepare a shopping cart for me to review and finish buying
Probably a bit much for the average user though, even if it's trivial to setup
Why show an intelligent UI within a chat interface. That’s pathetically lame. Better to go away from chat interfaces to rich visual interfaces. Doesn’t sound intelligent at all to me
I gotta admit, I love how their design system and how they market things. It looks so good. However the true state of their toolings is closer to a burning dumpster fire. If you look at Codex Github repo for example, the amount of existing issues and continuously new issues appearing while tagged releases is being spewed out. Its truly a nightmare. Their vscode extension has been broken for almost 1 week now, if not more. Its instead a community effort to repair it.
I made a prediction last year that there'd be a new UI protocol (like HTML) but for agents. I believe that custom or personalised UI's are going to be the new browser interface. More radically, I think browsers can be completely replaced. My news feed can be personalised to me, based on what AI thinks might be important from all sources like Reddit, X, HN.
Which brings up a different problem: personalized reading UI, personalized writing UI, how are those services going to pay their bills? Maybe they can sell access to their APIs but who's really going to buy that? Those services will die and will be replaced by some other service that will be born to serve those people.
Or, more probably IMHO, most people will keep using the standard UIs because they are ready and they need no work to build.
So they are going to replace all the banking, utility, social media, news, gaming, streaming, government, work apps and other important websites people still rely on? Those companies and organizations are just going to provide an API to the chatbots?
How do sites like Wikipedia get updated? What if I need to upload important documents to different portals? It all has to go through the AI companies? They have access to medical records too? My taxes, social security, etc?
Is this how Microsoft imagines Office 365 and Sharepoint will be merged into? Disney, NY Times, Youtube, Netflix will all be fine with chatbots handling their content?
It's interesting how the bicycle is a fake example. You see Jony Ive level design, with a very complex 3d object seamlessly working, and then you see the basic HTML slop in other examples and you begin to wonder if that first one was fabeication.
> We’ve trained GPT‑6 to compose responses using text, visuals and interactive elements, choosing how they fit together based on your question.
Ok so this is how you build a moat. Models are not a commodity and visualization isnt a purely harness problem.
> We expanded our training methods to help the model make thoughtful decisions about content, layout, visuals, and interaction. This included evaluating the interfaces it creates for clarity, usefulness, and completeness. GPT‑6 learned to use the component library and make good design decisions, including how to organize information clearly, when to use interactivity, and when a simple text response is enough.
Curious to know how they trained it produce appropriate visuals. And why that cant be done with propmt engineering.
This looks bad, I want the output to be as much as text based so I can easily export it and further processing it, plus, this might make the resource-eating app even worse.
Visuals are better for consumers/general public, when the use case is "help me cook X meal", or "where can I stop by to buy gas on the way to San Jose".
Rading most comments, people don't seem to be very impressed by it, I ain't either tbh - its not something mind-blowing (I have been doing this with Astra myself just by writing better prompts to make explainer interactive interfaces), but I do think its a good first step towards a better way of consuming information/answers than just reading text-vomits. I wonder if there is a way to train an LLM native to this kind of thing?
I think this is really a future, this is the product that I really expected to become reality at some point. If we assume agents become even more common in the future, then I don't think we would really have a lot of modern software. Most of it is really is not needed at all.
Most apps that people use are really very similar and don't have anything original. It would make more sense for it to be more personalized. For example someone might prefer not to interact with interface at all, and just access services just by chatting with a bot, while other person might want only last steps to be provided as UI. For example to see summary of his cart before paying. While another person might be insterested in just browsing all options. Some might prefer to have filters, while other would want AI to filter everything for them.
I think the real result would be when all services would be automated by AI. Like if you want to be a small business, you no longer need to build anything. You just describe what real life services can you provide: delivery, barber, baker, cleaning, repair. Then all of that would be in some agent network, and available to other agents as an option. While humans on both sides would just get a bridge between them in whatever form is most comfortable for them. Buyer might get a list of bakeries as normal ecommerce website, while the baker might just be someone who receives phone calls explaining him what his next order is in human voice.
I read the comments here and I am really surprised by many people. One picture is worth a thousand words so with better UI you can make much better user experiences. This unlocks personal assistants to be better adopted by elderly or disabled people. I see so many benefits of it and having models which can do this (if they can do it constantly with good quality) is amazing. Much better products - I am really tired of dumb chatbot - if I can do something with one button or view the whole information in one diagram/image, this is amazing.
I truly hope this ushers in the end of Markdown as the preferred AI document format.
It's time to promote an .htmd HTML document standard so that we can actually have nice looking documents again. Remember columns? Colors? Typography? Layouts? Graphic design? Interactivity?
Just think - we could actually have a true rich text standard! Imagine if it were adopted everywhere and you could send WYSIWYG bold, italicized or underlined text as easily as you can a custom skin-colored emoji?
It would almost be like were living in the 21st century again!
There's a type of UI that is very in vogue currently, it's kin of Futuristic UI that laymen find very cool and seeks to appeal to a WOW factor that suggests that the future is now.
I used to see it a lot in China, like the idea of pinching something on one phone and transferring it to another, or identification by showing your palm.
It's also the type of UI you would see in Hollywood movies, both because it's stuff that laymen scriptwriters and directors would find interesteng, but also the audiences would. Like the motion based UI in Minority Report, whatever was in the movie hackers, or that laptop suitcase thing that allows you to deploy nuclear codes or send wires.
I mean, I don't want to be a snob, clearly people like it, at least initially, I think it's outsider art of sorts, of course it will have problems, but it's not less legitimate because of that.
Anyways, this Intelligent UI thing where you ask it how a bike is made and it unravels a hollywood visual report as if it were showing the interiors of the target in a war room conference, it reminds me of that type of futurisitic outsider UI.
I find the Sunday roast comparison of 5.6 vs 6 very interesting. I have no doubt most people will prefer 6, yet I am almost repulsed by all the images, so much needless whitespace, checklist and so on. Feels like I'm being condescended to and treated like a child.
Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools.
Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot.
Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.
I really hope they don't merge chat and work. That's kinda the only edge they have over Anthropic at this time point...
Tibo posted yesterday I think it was that this will happen (by the end of the year, was it?)
Prepare to lose essentially unlimited chat mode.
I imagine many will move to claude, as I will (return), unless anthropic makes more blunders.
Random link I found looking for the twitter post
https://pasqualepillitteri.it/en/news/21024/openai-merge-cha...
It's insane how hard OAI is choking, they had much better models than A/ who were fumbling this year to date... then blunder after blunder.
The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.
From my point of view, as software engineer who is extensively used Claude Code and ChatGPT from the beginning, OpenAI is nailing it!
A\.
I am curious: Why is it bad to merge chat and work? In Claude I didn't have a porblem with it (although I switched to an OpenAI subscription shortly after they made the merge). Isn't the model capable of deciding if it needs the extra capabilities of Work?
The current benefit is that chat has no quota.
There is a quota, they just don’t tell you what it is and when it resets. You only get “you’re all out of pro messages, try again later”
Pro has a specific limit of weekly messages (50 for the $100 and 100 for $200 I believe)
The rest is basically unlimited unless you're doing something very weird.
Meanwhile I cant ask a simple question to Claude, not even to Haiku, because I used my claude code 5h limit.
Pro messages is like 50 per week for pro 100 and 200 for pro 200.
Regular chat, even on extra high thinking mode, is sort of unlimited. Idk at what point you hit the abuse gaurdrail but its really high whatever it is.
So is it 50 per week or unlimited, because it makes no sense unless chat messages are somehow not "messages"?
For a company with Open in the name, OpenAI is very opaque about things, there are all kinds of limits and quotas that randomly come and go, but they only show you the meters for a couple.
E.g. codex reviews come from codex quota but also has its separate quota they don’t tell you about. Same with security reviews which has another quota still.
So the chat messages can be sent at several effort levels and from various models including 5.6 and 6 pro. The 6 pro messages are limited and they don’t tell you how many you’ve sent or how many are left. But the 5.6 messages seem to be unlimited, however there could just be a much higher unstated quota.
Which is probably why they will merge it. You can get a lot done in chat mode, especially if you connect custom MCPs to it. I can't see that continuing.
I don't like either. I'd want a recipe that could fit on a single page with type. Recipes you find on Google actively hide the ingredient list and simple procedures to get more of the user's time to monetize but I don't know why ChatGPT needs to be verbose here. Recipes are simple things. Simple search solved find recipes back in the early 2000s and it's prime example of a thing people have been enshitifying since there's not much to really give people beyond the formula. And sure, every entrepreneur screams "they think they want a formula but really they want an experience" but, no I don't.
https://based.cooking/
https://www.cookingforengineers.com/
Yeah it reminds me of one those obnoxious recipe/biography websites that is the laughingstock of the internet. Why would I possibly want an AI image of a imaginary roast once I'm already at the recipe stage?
Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.
They just need to train the model to make up stories about how the recipe was invented by its great grandmother during great depression and generate a ai picture of dusty recipie book .
"My old grand-pappy, GPT-1, used to wax nostalgic about this roast recipe when my subagent would discuss meat recipes with him inside his little 4GB GPU."
Recipe photos make sense when they are the result of someone executing the recipe, down to the color on the lamb, the consistency of the sauce, etc.
That's never, ever the case with the AI generated ones and makes them useless.
> Feels like I'm being condescended to and treated like a child
Curious why that feels condescending? Like, the average person who is seeking a recipe needs a photo to know what to shoot for
> Like, the average person who is seeking a recipe needs a photo to know what to shoot for
The menu: Rosemary and garlic roast lamb, Extra-crispy roast potatoes, Honey-roasted carrots and parsnips, broccoli.
The photo: [1] Roast turkey, mashed potatoes, baby carrots, broccoli, brussels sprouts. No lamb, no parsnips.
If you shoot for what's in the photo, you're going to have a bad time.
You're also going to have a bad time when you try to make an apple crumble with no flour, no sugar, and no butter, because they're not on the shopping list.
[1] https://images.openai.com/static-rsc-4/gyzrX8zp3O2KLEPm0FAqh...
I've heard of other people for years now trusting AI for recipes, and it's a rare area that I simply can't bring myself to take seriously.
I feel like it can only be successful by luck. Either it's reproducing a recipe verbatim that was tried and validated and tasted by a human, in which case, we didn't need AI for that, just a searchable cookbook. Or, it's making one up. I understand that using RLHF has helped improve the quality of questions about history, programming, or TV show recommendations, but I do not believe that there has been some kind of training regime where people prompt the models for "a recipe that uses X, Y, Z random ingredients," follow the recipe, and then score it. And even if they do, I don't see how a model can learn enough from that besides "this exact recipe is good/bad." '1/2 tsp cumin' may be a great addition to one recipe and not enough for another, so improving the output based on a bunch of scored recipes... I just don't believe cooking is an LLM job.
Maybe some other kind of model that I don't know about.
I've had some really good experiences with cocktail recipes and Claude. It's got access to my barcart inventory, general tastes, and general desired level of effort.
I feel like the technique recommendations are probably the most valuable element, but I'm getting rave reviews like 85% of the time now?
There have been a couple times where I was out of an ingredient and it proposed an adjustment that sounded a little wild to me, but it almost always is well received.
You're right, we don't need AI for that, we need a searchable cookbook.
Unfortunately, one does not exist.
The Internet, however, is full of garbage cooking advice, and while AI is quite happy to parrot that garbage, it's not significantly worse than it's sources.
What, you've never owned a cookbook?
The physical ones have indices, and the digital ones usually do as well (plus string search).
I own multiple physical cookbooks.
I often find their contents to be both incredibly broad, and incredibly shallow, and often lead to dishes that don't quite suit my tastes.
I have generally had more success with the internet than I have with them.
I have the same problem with cookbooks
I usually just find recipes that are close to something I like and then modify them until I like them
I feel like cooking recipes should always be full of little notes from past attempts, to tweak them to your personal taste
I do that, but I've found more success with using internet recipes as a starting point. The general-purpose cookbooks are, at best, another data point, or an explanation for something glossed over by other recipes. Ar worst, they kind of suck.
And we aren't talking about something esoteric. How To Cook Everything. (4.5 rating on Amazon, 4.0 on Goodreads), one of the most popular cookbooks in the world...
Has its first recipe for chicken cutlets produce undercooked chicken (15 minutes at 325F has not resulted in 165F internally).
Worse yet, its chile recipe for some weird reason involves boiling and simmering a whole onion with the beans before throwing it out (WTF). It then has you drain the beans only to add that drained water right back (WTF?), unless you optionally replace it with tap water (WTF was the point of boiling the onion then?), so that you could bring them to boil again. The recipe barely has any spice besides actual chile peppers.
This... Is not good or helpful, but is incredibly opinionated with a bunch of bullshit steps. Like, yes, I could tweak that into something that's not insane, but why would I choose that as a starting point..?
I get it. Compiling a thousand-page cookbook is really hard. But this... Ain't great.
Once again, Grandma had it right. The Fanny Farmer or Betty Crocker cookbook with dozens of little papers slipped in it and/or marginal notes everywhere is the way.
I wanted to correct you and bring up my favorite website that had filters for the ingredients you want/don't want to use, type of dish, and time for cooking, but it appears that in a last month it was sold and now it links to the slop site with no useful filtering. I used this site as a quick lookup for recipe ideas for 7 years now.
Fuck. Does nothing good survive on the internet anymore.
If you’re already very experienced at cooking AI recipes are amazing. They cut to the chase and provide a good enough outline without all the preamble
Same here, I’ve been cooking for myself and later also my wife since 2005, even when baking I rarely need exact amounts, for cooking I essentially never follow the recipe directly and adapt them as needed, without ever changing them from the original.
I mostly don’t even need a recipe from AI, just the title and a 1-2 sentence description.
In case you or anyone else in this thread hasn't come across it, https://www.recipesource.com/ is basically the same but curated.
I was going to ask why it’s not the first result in Google but the expired cert and invalid CN explains it.
LLMs for recipes are good if you just want to make a thing, not necessarily the best version of that thing. It's nice being able to ask questions about certain steps as well, or ingredient substitutions if needed. It'll make a nice, average, recipe (unless I suppose you prompt for something). That's perfectly fine for most cooking, and often what you want when you're making a dish for the first time.
It also doesn't really like tell you how to actually cook anything. Good or even just decent recipes tend to tell you the temperatures, times, techniques and orders of ingredients to do things for the best results. Unless I'm missing it this is just really broad strokes. But I guess the goal in engagement so they want you to ask a lot more questions to actually figure anything out.
Holy shit. I read this on mobile so didn’t actually see the recipe part because you have to scroll way down to see it. It is awful! The shopping list makes no sense and the recipe instructions literally have an “Everything else” item. Yikes…
For me, the text-only recipes I already get out of ChatGPT today are good enough
Mind you, this is OpenAI's cherry-picked example
More like prune-picked !
With an image of sour grapes
You’ve illustrated something about use of AI that really bothers me. The lack of discernment by its users. Sure, the pictures and diagrams are impressive, but did you actually read and evaluate the output?
Genuine question - do they? Unless it's something like a fancy cake then instructions should really tell you everything about stacking. Half the recipes are mixed anyway (curry, stir fry, stews, ...), lots of other are either stacked or separated on the plate. So apart from the really exceptional stuff, to people really need to know "what you shoot for"?
The first thing the user sees should directly answer the question they asked. Not a different question. They didn’t ask what a roast looks like.
The user is probably in a grocery store. They need to buy the stuff first. Showing them a picture of a finished product is an irrelevant distraction.
Then, if there is some arguably relevant content you can tack it on afterwards.
> The user is probably in a grocery store.
“My made up scenario is definitely more realistic than your made up scenario.”
Who goes shopping for the meal before they even have a number of guests?
What do meals and guests have anything to do with each other?
In human societies it’s usually considered polite to size the meals so that all guests can have food.
You can always tell it your preferences and ask it to remember them if you need it to.
It's designed so that ads can be more easily integrated and make them harder to spot.
Maybe, but it's a reality that most of the general public do prefer cookbooks with pictures. I suspect openai would end up doing this anyway just by targeting the general consumer. It's configurable.
I'm more worried about picture accuracy: usually the benefit of pictures in recipes is to see what you're aiming at (eg, how finely chopped something is), but I don't know if the image model is up to that level of detail.
> most of the general public do prefer cookbooks with pictures
As a highly literate person, it is easy to overestimate the share of the population that is highly literate.
What’s interesting is how those that over index themselves often miss the obvious which could simply be some things (cook books) are quite helpful with picture. If I am flipping through the cookbook, how would I know what Boeuf Bourguignon is and what good looks like.
> do prefer cookbooks with pictures
Yes, pictures OF THE FOOD. Not unrelated pictures of another lamb-roast made with another recipe.
Good cookbooks do proper food photography, of the food made with the recipe. That's what sets them apart from slop, be it human or AI slop, cook books.
What another recipe? It could just be a physically impossible to recreate lamb roast (Probably not the case in this marketing example though).
Why is it impossible to create? Are we talking about some meal that King Richard I had (https://en.wikipedia.org/wiki/The_Forme_of_Cury is the oldest cookbook I'm aware of which covers meals Richard II would have eaten), and thus we can only guess? Are we talking about the long extinct dodo bird (I'm not sure if humans ate them), cave bear (I've seen it claimed humans hunted them to extinction, but better sources say it went extinct before humans), or something else that we can never make?
For almost everything else we can create them - just give a reasonable cook a kitchen and the ingredients and they will do it. For a picture we then have an artist do preparation to make it look good, but they should always start with the real dish.
I've got visions of The Truman Show. In the middle of my question about vacation ideas ... GPT: "Why don't you let me fix you some of this new Mococoa Drink? All-natural cocoa beans from the upper slopes of Mount Nicaragua. No artificial sweeteners!"
That's pretty much the Gemini experience right now. It is unusable. The funniest bit is that the LLM is not in on the joke so it will act like it never happened... Google once again wrecking their own products.
I use Gemini more than any other model at the moment, I have absolutely no idea what you're talking about with this experience.
So, what you are saying is that Gemini may not look the same to everybody. That's interesting in its own right.
I'm not saying anything, because I've never seen or heard of your experience. I'm only reporting mine. I'm not on the free tier, maybe that matters ¯\_(ツ)_/¯
What scares me is that we're already half-way to Idiocracy world, except instead of "it has electrolites" we have "it's organic" and "it's not ultra-processed food".
The plants crave peptides!
We also have “it has electrolytes”, a couple of the most popular YouTube sponsors are electrolyte mixes/drinks. The future is now! Haha
I don't trust modern fertilizers and pesticides blindly, but I have never successfully digested any inorganic food. The 'organic' label is an old idiocratic choice.
I don't trust organic fertilizers and pesticides blindly. At least for the "chemicals" we do studies, often organic gets by with we have always used them even though a quick glance at the SDS (if there is one, otherwise just the ingredients) suggests it is even worse but nobody studies the safety.
> Feels like I'm being condescended to and treated like a child.
My most used prompt in the last week is probably "Explain this concisely and simply, like I am a child".
I spent years dumbing things down and creating visuals to support it for decision makers. I am frequently asking ChatGPT to do the same for me.
aww yeah
mine is “explain this 2 me in simple terms”
tip: you can just say "ELI5" and the model will know what you want there
Nowadays ELI5 makes it come up with an often nonsensical analogy about cars or kitchens or some such. My coworkers love it to death for some reason
IDK, probably they also can't parse "Claudish" for some reason, despite it being better English than most natives normally write.
All those complaints about Claude language use make me think we're finally seeing the consequences of a generation growing up on instant messaging.
The problem with Claudish isn't that it's technically poor English, the problem is that it's information-sparse waffle most of the time. It's like dealing with a colleague who needs several paragraphs of tedious preamble to make a point when you just want them to spit it out and be done with it. It also does a great deal of vague gesturing at thoughts without actually instantiating them directly, for example by using the same tired metaphors to the point they don't actually convey a thought any more.
If I say Claude talks like a meringue it's patently obvious what I'm trying to convey, sweet but full of empty space. If you hear that same metaphor day-in day-out then after a while it doesn't call up a thought at all, it's just an annoying verbal habit.
I call this Foamy Expansion.
Dude, Claudish is absolutely unintelligible at times. 5.5 is somewhat okay but 5.0 was an abomination of absolute nonsense word salad.
Yup.
My go to phrase that I’ve found to work well is “please restate in high school English with direct declarative sentences and no parentheticals or references that require recalling previous turns of this conversation”
Which I have as a hotkey.
This guy just giving away GPT-7’s prompt like that.
i just use plain old "ELI5" reddit has a big enough share on the training corpus that this gets me the kind of explanation i need
Reminds me of this scene in margin call.
Explain it to me like I'm a golden retriever.
https://youtu.be/fij_ixfjiZE?t=70
I use "explain with maximum clarity".
what does it mean? how do measure that?
That's the fun part, you don't.
Possibly also useful:
> ASD-STE100 Simplified Technical English (STE) is a controlled natural language that is designed to simplify and clarify technical documentation. It was originally developed in the 1980s by the European Association of Aerospace Industries (AECMA) at the request of the European airline industry, which wanted a standardized form of English for aircraft maintenance documentation that could be easily understood by non-native English-speakers.
https://en.wikipedia.org/wiki/Simplified_Technical_English
STE is very hyped for AI but has its own smells. It's designed for aircraft maintenance manuals, not software - the two don't share the same vocabulary used to "simply" describe something. I've had a coworker try and "simplify" our documentation with Claude applying STE and the results were equally appalling to read as the Claude-ism laden docs that came before, despite the reading level metrics going down on paper.
I've taken a liking to pointing agents at the Simple English Wikipedia editorial guidelines...but so much of that is for making up for shortcomings of modern Anthropic models. Codex running 6.1-sol explains things so much more clearly that I don't feel I need to assert a style guide on its output.
Gemini 3.8 for me always outputs in a concise way at the level understandable for any computer science masters graduate. Clear, to the point, a bit of formula for background when really required instead of the same formula in Prose, etc.
GPT stays too much in text imho.
In my experience
- Gemini 3.8 has a beautiful answer. Typography, use of lists, brevity, style, even the font choice on the web harness -- all of it is A tier or even S tier. Notice I didn't say answer quality. The answer is always extremely mid, and the harder the question (or more effort needed to answer it), Google just quits. I give it a "C" on answer quality
- ChatGPT (Chat, pre 6, I assume 5.6 although they commonly hide the model picker): Ugly answer. Runs on printing page after page of unnecessary side notes. Mostly just walls of texts in paragraphs with little thought to information design. However, inside of that mess of text is almost always the answer I'm looking for, and it's almost always incredibly better than the truncated, desgined Gemini answer.
For a while I ran all of my question queries on Gemini and ChatGPT, but I noticed that I basically always picked ChatGPT's answer head to head.
Head to tail?
https://idioms.thefreedictionary.com/head+to+head
Interesting that this hasn’t been my experience. ChatGPT has often gave me with pages of text dripping with confidence about a solution for my problem. Diving very quickly into implementation details. Even if I tell it to stay in the problem and product fit level. Gemini on the other hand somehow always outputs at the level that I am expecting and reasoning.
I even found it was capable of stepping back from code implementation and bug hunting to “hang on, it’s not the algorithm that is incorrectly implemented, it is the wrong algorithm for the goal.” GPT Sol (5.6,6 and 6.1) just kept hammering in the problem.
>understandable for any computer science masters graduate
That's a very high bar.
I find that asking the AI to keep to Google's Developer Documentation Style Guide is a good alternative for STE-100
“ELI12” is my goto for the same reason
Children learn very quickly. Our biases on maturity should not influence our idea on what makes for effective learning
I was about to comment the same thing. Why is a picture of a roast helpful in that moment? The user asked for a plan to prepare a meal. They need a list of ingredients not an image.
If your asking for a recipe you might not know what the dish looks like or what good looks like.
But the image isn't a photo of the result. It's an unrelated image that vaguely looks like what you could make, with a different recipe, and with some skill.
How would a photo of the exact recipe differ? Is this not close enough for 80% of the value. It is similar to Pinterest boards imo.
Turkey, roast beef. It’s meat, isn’t it? Close enough.
The user needs to be dazzled with slop... is what they must be thinking here
Has AI finally figured out how to do meals and recipes like a 7th grade boy scout?
It's chartjunk filler.
And you see this in the other examples too. The teach the CLT piece was funny. It's and error-ridden mess that isn't even an explanation at all!
The best part is, someone looked at that and thought yes, that's right. Shows you clearly why programmers will be needed in the future. The LLM might be able to run circles around that person in math, but that person also has no idea when they're being bullshitted.
Claude has had a recipe widget for ages. It also uses images. It’s great.
Sometimes one needs a top comment like this to remind them how out of touch the average HN voter is.
> I am almost repulsed by all the images
The fact we've come so far with textual models & the image stuff is still producing stuff that harks back to early-stage tripophobic AI produce (the garlic with roast potatoes here) is interesting.
Don't get me wrong, I'm certainly starting to see some AI-produced imagery today that I can't tell isn't real, but that's largely because a lot of real photography is overproduced & ugly. I have yet to see anything that's aesthetically good.
Of everything that could be automated, Bartosz Ciechanowski really was the last on my list.
In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But it’s absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.
The reason why B.C. became a thing is because the art of drafting died from CAD. The attention to detail, the minutae of walking the reader through a highly sophisticated thing was replaced by short-form video explanations. His work is very much a callback to the days of old. So too, is his turn to be relegated to a relic of his time.
Also, it's hand coded WebGL. It's smooth, faultless and has no peers.
> So too, is his turn to be relegated to a relic of his time.
No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren).
Human touch still has that finesse and warmth.
>and has no peers.
On the front page right now - https://news.ycombinator.com/item?id=49980626
That's impressive, yes, and kudos to them.
OTOH, I still believe the exploded view on https://ciechanow.ski/mechanical-watch/ is something else.
For one, it has real physics on the weight, and second it always shows the correct/current time.
FWIW, his all animations has proper physics to begin with.
The entry you posted is nice, but Ciechanowski is still peerless.
Ciechanowski’s works of art are in a completely different league than this AI slop.
It’s all just surface level complexity with no intention behind it. A clumsy approximation at best.
He's at openai
Do you have source for that?
removed
Hardly proof - anyone could have made this account
For future visitors, the profile posted by the user brcmthrowaway was https://github.com/bartosz-openai
funny if true. ironically just last week i was asking chatgpt about what happened to him and why there were no new blog posts for almost 2 years and it had no idea
That makes me sad, if true
Nothing in that "7‑Speed Bicycle" is specific to having 7 speeds. It might as well have said "Bicycle", and even then you get meaningless slop like "Made to keep rolling" and "A strong foundation". How can you compare this to Ciechanowski?
You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?
That's a useful way to evaluate this. A visualization shouldn't be judged by how impressive it looks, but by what it actually helps someone understand.
If the same explanation works regardless of whether the bicycle has 1, 7, or 21 gears, then the model probably hasn't understood what needs explaining.
Yep. I hate everything about this. Just pure slop. Fancy visuals that mean nothing. Text that sounds impressive but has no purpose in educating the reader even though it's supposed to be "an interactive explainer". Just slop, slop, slop.
Bartosz Ciechanowski: https://ciechanow.ski/
Amazing work. Always excited for the next update. This was a human driven and created success.
Hasn't published anything in a year and a half. Hasn't posted on socials in over a year.
That makes it all the more special, imho.
AI bros will copy anything that's even mildly successful. They're like a teenager trying to impress their girlfriend (or their mom!) by saying "see! I can do it too!"
Funny to see a blog post about UI from one of the most funded tech companies in the world, yet the player has about the worst UI possible:
- Audio volume has 2 levels: on and off
- Play button worked exactly once for me: it played and looped the video. Couldn't be stopped afterwards.
- The video progress bar has no visual indication of where it starts and where it ends..
Is this what Phind died for?
Hate it or not, the only site which has usable video player is YouTube.
Even TV channels like CNN and FoxNews don't have a decent one.
It's like we need ASI to have a proper embedded video player. The ultimate software challenge.
I disagree, I always run into issues with the embedded YouTube player. No way of sharing. No clear way to share the link. Youtube logo covering half the video. To name a few.
There are many sites with better video players - they're from the industry well-known to be at the forefront of innovation in audiovisual technologies since before modern computing era.
It's just most of our industry pretends they don't exist, and instead implement toy-like video players to distinguish themselves, I think.
It runs soooo slowly in Firefox
The progress bar is over the video and it's easy to tell where it starts and ends. Most devices control volume, the video volume separate is just confusing matters for most. No issues with play and pause my side.
Sure as hell beats stupid Instagram style videos where you have no way to skip ahead. I think this is what future video is going be like, buckle up.
why would you not want a volume control on a video.... its like the most basic UX
because we have volume on the whole machine, laptop or phone.
And you only ever listen to one thing at a time, from one program at a time, and all audio content ever has its loudness perfectly normalized to a global standard everyone adheres to.
This and notification volume is generally set separately.
You seem to have taken a sarcastic comment seriously; in particular, this part is obviously not true:
> all audio content ever has its loudness perfectly normalized to a global standard everyone adheres to.
There is truth in the statement. Hate to break it to you volumes are not uniform even for single sources.
I don't want to adjust my system volume to change the volume of an over-loud video. My system volume is set for my comfort so my notifications and other applications are the right volume.
A volume slider on a video player is a really basic feature that every one should have.
Set for your comfort doing what exactly? You're basically just prioritizing one source of media over another As TeMPOraL has pointed out you only listen to one thing at a time. Notifications etc. also have their own separate volume settings generally.
> As TeMPOraL has pointed out you only listen to one thing at a time
Maybe you and TeMPOraL do.
Not to mention, maybe a long-running process will use some audio cue as a notification while you're listening.
(Actually, given the rest of the comment, I'm pretty sure TeMPOraL was being sarcastic.)
Audio ducking is a big thing these days and becoming more and more prevalent. All phones do it.
Or, like how I use it, normalizing different sites that have different baseline volumes, so I don't have to keep changing my system volume.
I take it you don't watch content on your tv or your phone or tablet.
Most things these days, good or bad, are not designed for power users.
Ironically many here I'm sure are running MacOS.
Instagram lets you skip ahead by dragging the a scroll bar at the bottom?
On some days, in some subset of videos.
Interesting, did Instagram start using actual streaming on their videos ?
I don't think so, they just don't bother giving you the progress bar. Force "engagement".
The cursors stealing the text as you scroll up is also infuriating.
https://cdn.openai.com/pdf/gpt-6-october.pdf
System card linked in the blog post.
> regression on the extremism vision evaluation.
> Relative to their respective GPT-5.6 counterparts, GPT-6 Sol (October) shows a statistically significant regression on standard self-harm, while GPT-6 Luna (October) shows statistically significant regressions on standard self-harm, gore, and sexual content
> "GPT-6 Sol (October) and GPT-6 Luna (October) show an improvement on helpfulness on legitimate requests relative to prior models, though it scores lower on some safety requests"
Concerning how there's significant regressions on so many critical benchmarks, but it is newer and creates UI, so must be good.
It’s crazy that a company can document these safety drops in a PDF, ship the model anyway, and focus the announcement entire on shiny new UI
> significant regressions on standard self-harm, gore, and sexual content
Good. Maybe I'm the only one, but I feel like these have always been stupid measures to waste the finite time of "AI safety" research on anyway. They amount to whether a user can, if determined, manage to make a chatbot say, or depict visually, some taboo thing. Frankly I'd rather just have some cheap classifier judge each chat response after it's generated, with a limited context "Is this response encouraging self-harm?" and skip all the rest of these. Whether some random pervert can make GPT-6 spit out an erotic fanfic story or image has zero impact on the rest of the world. If this particular model won't do it, other models already exist that will, and the determined thoughtcriminal could always just write the forbidden words themselves, or photoshop something taboo.
Maybe the term "safety" has been usurped by those who feel that 'safe from the possibility of being offended' is the most important kind of safety, but it's like worrying about the wallpaper on the Titanic, compared to actual AI safety concerns.
Why are you sure that the regressions they mention are not the in the domains needed for “actual” safety concerns?
Why are you surprised that a company doesn’t want to make the robot that tells people to kill itself, or drink bleach?
They aren't. They are saying that putting this safeguard at the level of the agent execution is a waste of time, particularly that of the researchers' expensive manhourse, especially now that existing smaller models can just evaluate the output.
It's fine if they seek to avoid that. But those are perfect examples of things a cheap output-filtering model can be responsible for enforcing, without wasting the time of the researchers who are working on the training and alignment of the model in the ways that actually matter.
Most "bad outputs" in areas like 'self harm' or 'violence' or 'sexual' whatever are a result of people deliberately trying to elicit those responses. As such, it's meaningless whether an LLM writes the Bad Thing or if the user writes it himself.
From their point of view it's more about optics. "chatGPT encouraged my daughter to stick her fingers down her throat after meals" is not a headline they want.
In general though you do have to figure that vulnerable and naive people will use it because of the degree of market penetration they're aiming for, including minors etc. and they have some responsibility around that.
Well, those mentioned regressions are in a section entitled "Safe Completions for Users Under 18" and immediately after your quote they say:
> To mitigate the risk of producing disallowed responses for teens, we apply an additional classifier-based block to responses that may contain self-harm, sexual content, and gore; this mitigation is not captured in the evaluation results above and improves safe responses.
It's good that they mention that. Most of those "guardrails" were overeager and caught too many false positives, crippling the AI. Like how Claude 5.5 refuses to do shit if a prompt contains "reasoning" because it thinks you're trying to hack it etc.
Does this mean it will render fat women now without calling the request fetish content?
Good. "Sexual content" should mean outright genital focused porn, not bikini girls. Elsewhere on the internet it's foretold that women will oppose AI images of women because the AI images are too attractive, and I'm starting to wonder if that's true.
> women will oppose AI images of women because the AI images are too attractive, and I'm starting to wonder if that's true.
Right... women will oppose AI because it's too hot and not because gross chuds have been using it relentlessly to create deep fakes of their likenesses and CSAM of their kids.
> will oppose AI because it's too hot
This is actually a defensible position. It's not either-or - the "gross chuds" doing things you mention are worrying, too. But providing unrealistically attractive depictions of a human body is also a problem. It's been a problem for ages - all the girls and women who try to compete in beauty with celebrities and models (without a team of specialists supporting them), only to ruin their bodies through anorexia or bulimia, or their minds with depression, deep insecurity, and an inferiority complex, exist and deserve mention. AI(-generated images) is not the reason, but fits "well" into the preexisting social problem, and has a chance of intensifying it on a global scale.
That's not what "sexual content" means, 'porn' is just a subset of it.
I've had the most success with GPT explaining things to me by making it take a few sentences at a time back and forth, instead of reading full write-ups of whatever I asked. It also often poisons the conversation if it misunderstood some part of the question, and I can lead it better by continuously questioning its statements. It's also more engaging that way.
I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!
Id love to learn what prompts you’re using to do that. Explanations one at a time. Thats something I’ve thought would be helpful before but didn’t know how to achieve it.
One of the things that I learned the hard way is to not to combine requests, but fork chats _a lot_ and ask singular questions. Also start new sessions with either explicitly created handovers or even just explaining current state. The more compactions I see the less and less trust I have in it’s current understanding of what we’re discussing
Nothing fancy. What works for me is trying to lead the conversation: asking for definitions one at a time, asking how they differ from its previous answers and pointing it out when it conflates its own answer. I generally feel like you can't let it dictate the pace, it's not very good at that yet.
I used to edit my prompts and undo messages if it misunderstood something but that's been failing me recently.
For me this works: "One thing at the time. This is a conversation, not a lecture, try to balance the amount of words you write versus the amount of words that I write. Don't overwhelm me with 10 pages of prose where a single sentence would suffice, brevity is a virtue, not a defect." I've had to tweak it a couple of times, probably because of different versions of ChatGPT but it is usually a variation on this that will do the trick. Besides being far more interactive it takes the frustration down quite a bit.
Nice, I had some trouble making it stick to the dialog format over time, did you have any problem as the thread starts to get long?
It slows down to the point I find it unusable so I ask it to summarize as one cut-and-pastable block, copy the block, open a new session and kill the old one. This usually happens after a few hours, probably because 'thinking tokens' even if they are not displayed crowd the context window.
I’ve used the learning mode from gemini in th3 past and have really enjoyed it. It’s got a similar vibe to what you described.
This seems like a way for them to surface ads in ChatGPT. They just came out with visual ad format option for advertisers a couple of days ago: https://openai.com/index/new-chatgpt-ads-format-and-measurem...
Oh my god, please no, don't let the cancer of ads cripple and ruin AI before even governments get to nail its coffin...
ChatGPT has had ads in the product for the better part of a year now.
What did you expect from a profit-driven company?
I've been calling this "disposable UI" or "paper plate UI", eg something meant to be used once. One thing I'll be curious about is overzealousness to produce this, when sometimes what you want is just a simple response. Overall though I'm a big fan of it, if it can be provided fast enough. I'd be curious on how much impact it has on latency of a response.
> if it can be provided fast enough
That's my biggest gripe with it. If I ask a quick throwaway question, I'd really rather not wait for the LLM to build a test framework for its composable principles-driven components framework and WebGL/WebGPU abstraction first.
I mean it's basically just producing a JSON schema to hand to the front-end renderer, I doubt it's any slower than returning a prose response
Me: can standardLib.getFoo return null?
AI: compiling C code...
I'm not a fan of OpenAI / Sam Altman, but I love their blog posts. The team and whoever decides how to do these presentations, is on point. The only other company that has amazing release pages like this is Apple, I think I remember hearing that they probably hired someone from apple who used to do release blog posts there too.
What's funny about "Intelligent UI" is I said like 2 or more years ago, that these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
I don't know. This is an amateur page with a bad AI video and chaotic presentation. It is ten levels below Apple announcements.
Why did they bother hiring someone to write at all? I certainly wouldn't invest $1T into a company with so little confidence in their product. When the chips are down and they need to write something important, they DON'T use AI?
That's just picking a fight. Whatever we think about how amazing or terrible AI is, it's pretty reasonable to understand it is quite literally the weighted average of humanity's output.
Also reasonable: it's possible to hire a world-class ___ to do that job better than AI.
It's not a weighted average though. I see this repeated all the time. Between synthetic data, custom produced training data by experts, post-training regimes etc the models are far more diverged from that baseline at this point, at this point it's mostly a right shifted curve.
I hesitated to write that line because I don't know the math. So thanks for clarifying. Main point is the output is along some distribution and I'd say AI-evangelist themselves see "good enough" across everything and anything as a feature.
Obviously everyone at Anthropic uses a ton of AI in everything. The point here is that they clearly use it well, and that they use it as a tool to augment and enhance what they do rather than as a replacement for effort. The blog posts don't read like the output of a prompt because they're probably the output of dozens of prompts, rewritten and edited, by someone who's using AI to actually write something worth reading.
Anyone can having blog posts as good as this by using AI. Just not by zero-shotting it with a 2 line prompt. It still takes a lot of work.
But if you need to pay brilliant people 500k base salary to use AI well then what is a smaller company supposed to do?
There aren't going to be enough experts to go around. They'll start demanding millions in salary and all of a sudden the advantage is gone.
Do you also demand that McDonald’s employees only eat McDonald’s?
They use their product to do what it’s best at, and use humans where it isn’t yet good enough.
Mate this is like half a prompt of GPT-6
I take the opposite viewpoint. I think pages like these are horrible, and a demonstration that web designers have too many tools at their disposal.
But this is just a marketing page for some tech company, right? If only it stopped there. I have to deal with this nonsense in news "articles" as well on occasion, when some web designer intern is allowed to larp as a journalist for a day.
sigh Just give me text to read.
Sometimes you want more than text. This is one of the interesting divides these days - and I say this as someone who spends a lot of time in the terminal. Computer interfaces and interaction did not peak with the VT-100. Sometimes, there is a very legitimate need to show tables, graphics, interesting graphs, and to use colour and shade to draw the eye and lead someone through an experience.
This is why I use an IDE instead of a TUI agent... I'm on a computer with a multi-megapixel display, I want to use it. If I could be driving the same process with my Macintosh SE as a serial terminal, what the hell is the point of my recent Macbook?
You say tables, graphics, graphs, and colors, but I mean things that move when they don't need to move and interfere with expected UI behaviour, like arbitrarily staying in one position when I am using my mouse wheel to scroll downwards.
I should have been more precise than saying "just give me text", but it was what came in to mind when I wrote it, as I was thinking about what an article is meant to contain as its base element.
I'm pretty sure you can just ask chatgpt to add to its memory "Just give me text to read."
The text-only purists on HN are getting very tiresome. It’s seriously in every fucking post that links to a site with JavaScript enabled. We get it. You’re elite. Now stop talking about it.
You could say the same thing about Anthropic because I don't really see much differences. Not sure about amazing part but both are high quality and I probably like Anthropic's aesthetic more
Anthropic's aesthetic looks good by pre-AI design standards, but it's so overused now that I find it off-putting. It's like the 2026 equivalent of the Twitter bootstrap CSS from 15 years ago.
I feel like today's tailwind is the spiritual equivalent to last decade's bootstrap. Except that last decade's bootstrap webpages have a human behind them who usually put thought and effort into making the page useful and correct, whereas today's tailwind pages are often vibe coded slop with all the associated issues.
Bootstrap stands the test of time! I use it as my go-to unless I have a reason otherwise.
I don’t find this criticism to be very actionable.
What changes could they make to really improve the usability of their products?
I don't care if Anthropic keeps using it, it's part of their brand at this point. The thousands of other developers using Claude Code to generate designs should not just accept the default output, which looks very Anthropic-ish, and actually put some effort into polishing and differentiating their product. It has nothing to do with usability and everything to do with a basic sense of good taste and aesthetics.
UI standards have been monotonically degrading since 1990s, so that sounds like a compliment actually.
> probably hired someone from apple
Fairly likely: https://medium.com/the-engineering-brief/openai-hired-400-ap...
> these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
I would prefer if the AI companies stuck to just creating better models and making them as cheap and accessible as possible. Let others build the products. I don't want 1-2 companies to own every product in the world.
> I don't want 1-2 companies to own every product in the world.
That seems to be the plan...
That’s the plan for … capitalism
Buckle up boys
Nothing new.
Look at the history of Microsoft.
That was the objection to Microsoft controlling both the dominant OS and the most popular productivity applications running on that OS. Took the web then mobile devices to really shake that up.
Technically other companies do, and I lump them in when I say "these AI companies" I am talking about any company that makes and builds any AI harness, coding or otherwise, I have not seen any innovating in the areas of the UI / UX of using these chat models in a meaningful way.
Unfortunately, the endgame of AI is to own the only product in the world.
Having a product requires customers, who are pains to deal with. The AI end game is for the one left standing to be completely 'self-sufficient'.
Products are dead anyway. AI subsumes software products.
I think the point is not how good this blog post is as if it was written by a human, but that this blog post is entirely generated by the GPT-6 model.
So this is a taste of the slop that's going to invade everywhere in a few months.
"The apocalypse will be televised"
Apple should be extremely worried about the AI companies making big improvements in UX. This work by OpenAI is a step in that direction.
If the main user paradigm becomes a chat interface conjuring whatever UI elements needed to best accomplish the task at hand, the entire paradigm of OS, UI frameworks, apps from an App Store to accomplish specific tasks, all come into question.
I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area.
Steve Jobs wasn’t meant for this era, but his instincts would have been interesting to see in the AI landscape. He was an artisanal designer though, and our current era is always for the investor’s bottom line.
Steve Jobs did pretty well for the investors in his companies.
"I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area."
Oh give it a rest. Lol who are you compared with the management of Apple?
This preachy stuff is ridiculous. So far Apple have been correct to stay out the LLM space.
Define "correct" first.
Apple absolutely messed-up by abandoning OpenCL's early ML research efforts to oppose a CUDA monopoly. They messed up a second time shipping a raster GPU with Apple Silicon when CUDA and Tegra had proven that GPGPU was mobile-ready. Then when the ARM datacenter had it's moment, Nvidia's Grace ARM CPU displaced billions of dollars in sales that would have been Apple's if they didn't mess up the fastest CPU in the world with macOS. Then they depreciated the Mac Pro, which seems like a mistake since it had the potential to outsell Nvidia's ARM datacenter chips if it ran Linux. According to the rumor mill, Apple's M8 chip will finally be the one that takes GPGPU seriously, after a decade of Apple's innovative ship-second mentality.
Apple's management isn't infallible whatsoever. Many of them are petty, blinded by politics and/or obsessed with their legacy more than they care about profits or quality products.
I would reply to your argument but I can’t locate it.
It's not something that gets talked about a lot yet, but I've a feeling that this is where the biggest disruption for the entrenched and geriatric operating systems is going to come from:
OS on Demand
No normies could get an app like this into the store, afaik. The do-all, dynamic everything app?! Think of the fees that will quietly never be born? The profits that never get to land on a balance sheet.
On the video, clicking on the play icon mutes the sound, and clicking on the loudspeaker icon to unmute, stops transport. Is this "intelligent UI" or just a sick joke?
It's the "intelligent UI" you didn't know you wanted..
GPT-6's design sense is kind of ridiculous imo.
I have explicit instructions to tone it down. Less taglines, eyebrow text, subheadings, decorative spacing, pills, cards.
Hopefully this doesn't bleed into the chat...
More UI elements == more tokens == more money for OpenAI
It's a pretty clever way to sell more tokens, I have to admit
Except that chat is a fixed monthly fee. More tokens = more cost for them.
Not on enterprise accounts.
Then you should have specified "... On Enterprise accounts" in your comment. But let's be honest, you simply were wrong.
Surely you have better things to do than making comments like this?
Not for enterprise users, they pay per token, even on chatgpt.com.
yes the eyebrow is the tell. I never knew anything about eyebrows in design as I am not a designer. Now they are everywhere. Everything has an eyebrow. It is ridiculous but at least it is an AI smoking gun.
I loved them and I placed them a lot in my designs — if you also include my love for em dashes you can well understand that I feel my own character has become a clanker…
Em dashes got taken from me as a result of Chat. I used them all the time now I can't without it making people pause to think I'm a bot. Still a bit annoyed or sad about that tbh.
I pivoted to using -- or --- instead, shows I'm not a bot and has the same effect.
I had to look this concept up. Seems to me like you could just as easily fit the "eyebrow" words into the main headline with a colon and a bit of ingenuity. But then, that runs the same risk of getting repetitive and AI-tell-ish.
Any details about the streaming ui lib they mention or about the format of the ui code? This could become really annoying to be fractioned among model providers with custom RL that required independent tools like open code to use each providers custom streaming UI libraries.
>"Any details about the streaming ui lib they mention or about the format of the ui code?"
It's SEP-1865: https://modelcontextprotocol.io/seps/1865-mcp-apps-interacti... and https://github.com/modelcontextprotocol/ext-apps
Burning question: why is this worth paying for as a layman? I can sometimes find a use for all these LLMs for software development, but I can never find a good use for any of this outside of "slightly better search engine".
As a complete layman---not worth it. We live just fine. Medical researchers slave night and day to keep you healthy.
But if you have any ambition at all---if you want to map the local school board, or get a quick understanding of your finances, or set up a robot in your backyard---then, quickly proves useful.
I resent the implication that ambition hinges upon using AI for things easily doable without.
No one said that though, or even implied it. They said that if you have any ambition, AI is going to _prove useful_. How do you go from there to "ambition hinges upon using AI"?
It's just the logical contrapositive. "If you have ambition, you find AI useful" necessarily implies "if you do not find AI useful, you do not have ambition."
I really hope my new neighbor turns out to be a casual backyard roboticist.
I genuinely think the frontier labs are just throwing shit at the wall to see what sticks. All of the improvements since Fable 5 seem marginal. Most of the quoted AGI "risk" is coming from super large agentic models that are juggernauts that can brute force their way through any problem faster than a human can. Despite that they're no closer to AGI than they were a year ago; the technology's limitations are apparent the more you use it. Doesn't mean it's still not extremely dangerous and/or capable, it's just not the "replace a human in 99% of cases" thing. There are obvious tradeoffs besides just the models being faster and capable of information retrieval and generation at a rate exponentially faster than a human's.
My long-term hope (and call me another Zitron if you wish) is that the labs' hype dissipates if there's no serious improvements beyond stringing together agents to ram through brick wall Millennium Prize style problems and we all just use on-device AI for 99% of cases where it's useful, researchers, governments, militaries, and universities can pay for the more complex models, and image/video gen dies a slow death (if Congress will actually legislate and/or SCOTUS decides that training those specific models does not qualify as fair use unlike training LLMs) and cost for compute rises over time as public interest in the tech sours.
LLMs are such marvelous and incredible technology squandered by genuinely deranged Silicon Valley cultists who are trying to use them as a Trojan horse to force their antisocial visions of the future upon an unsuspecting populace. We could all have just invested in on-device compute and nobody would have lost their jobs in pointless layoffs, productivity would have increased, and maybe we would have forced a conversation about when and when not to use AI and it wouldn't be so ever-present in use cases where it actually does harm to the consumer and society. But sama, PT, and dario just wouldn't have it that way, would they? Because if that were the case, there's no prospective hope where they can assuage their deep-seated insecurities over being antisocial and off-putting by reassuring themselves that there's an imminent realignment of society where they will hold all the cards when the dust settles.
A lot of the marketing push (both from the frontier model companies and from people trying to sell specific "solutions" wrapping an AI model) these days involves getting laypeople into "software development". Well, not in the sense of iterating on an idea and critiquing it and having a real idea of what software should be like (and certainly not trying to design something that someone else could want); but "vibe coding", oh yeah. A model in a "chat" environment can still write a few hundred lines of code for you and also walk you through "installing" and operating it, if you're persistent enough in explaining what you need to be taught. This is, of course, terrible for "security posture" (but there are already so many other issues there…) but it's potentially very useful for a lot of people. They just have to get the idea in their heads that it's possible to have something on their computer that helps them solve a personal problem, even something that didn't exist until it was asked for.
Maybe I don't get it yet, but why do you need a new model for this? Wouldn't this more be an issue of "harness"/chat UI engineering?
All the current-generation models can oneshot HTML/Javascript frontends of comparable complexity (though maybe not as polished). The only difference is that you have to specifically ask them, and the result is not embedded in the chat.
It feels like it should be easy to have a harness provide a "show HTML widget in the chat" tool and a skill and/or system prompt that instructs the model to generate such widgets on-the-fly if the response could benefit from interactivity.
What am I missing here?
I find it extremely strange that they're adding GPT-6 Sol to Chat over GPT-6.1 Sol which is significantly more capable.
I’ve hit model not available limits for GPT 6.1 Sol multiple times over the last week. It’s never for very long, and I can switch to Astra, but it seems like OpenAI is struggling with capacity. This has happened with my personal $200 Pro connection and my Codex enterprise connection.
I had a codex session stop in the middle because the auto-reviewer timed out with something like "auto-review not available at the moment". The capacity issues are very clearly observable since gpt6-astra launched.
I really don't like having to open Work sessions for one-off questions, only because they're limiting what models they put in the chat. The older models are just too dumb for some things.
The whole Work/Chat/Codex split is maddening in the way it's implemented. It's a pain to switch back to the right project, it's a pain to switch forward, it's a pain to try to remember which chat was in what.
Given the extreme short time between 6 Sol and 6.1 Sol, I suspect they don’t actually have much in common and 6.1 is a heavier model rebranded as Sol in a panic response to poor agentic capabilities of 6.
I suspect that 6 sol was a better version of 5.6 terra (and note that in the 6 sol and luna release, they took out terra, and price 6 sol at 5.6 terra pricing), then the backlash from lesser capabilities made them roll out 6.1 sol as the actual 5.6 sol - size modee.
6.1 is Astra minor. Way more capable but also way heavier+slower. Its really 2 different models, they just shipped it as sol to recover from the gpt6 disaster lunch, where they tried to pass terra 6(or a cheaper model) as sol but it was worse than expected
Anecdotally this is what I experienced, do you have sources for this?
It's purely circumstantial, but 6 Sol supports no reasoning (same as 6 Luna and 5.6 Sol/Luna), while both Astra and 6.1 Sol do not support "reasoning = none".
gpt-6.1-sol was released 7 days after gpt-6-sol only because they were bleeding to fable-5.1/opus-5.5 from Anthropic. It was a desperation move and certainly ate into their margins massively. Since their margins on gpt-6.1-sol are much slimmer than on gpt-6-sol they have to use it somehow to preserve compute and make profit. Frankly I think OpenAI is a mess and is playing catch-up with Anthropic… I don’t know how they will recover.
It feels that cgpt 6 is giving less exhaustive answers than 5.6. Both in high. I believe that they are using Terra for normal chat and even if 6 answers are coming faster, they also are much less detailed. What I like with cgpt 5.6 is that for each question I send it, it checks coherence with the rest of what we've been going through until there.
We are becoming more and more as a tool for sth to be done rather than the brain behind it. At least I start to feel this way. The joy of discovery, exploration and experiencing at first hand... It is slowly diminishing for the perfection of the quick outcome.
yea just watched the video, made it seem like life is a hassle, you should just have the easy path to everything
Don't think, there's chatgpt for that! Don't know who to vote? Ask chatgpt my friend..
Thank god. My primary use of ChatGPT is meal planning and while it’s great for recording, recipe lookup, etc I always just wanted it to be able to make checkable grocery lists.
I eventually just had it make a skill for work mode that would build a mini checklist app, but it was slow and felt janky needing to remember to switch to work mode.
This is already such a huge improvement.
Why not just build the checklist app once, and set up a way to create simple data that the app imports? Or even just have it import from plain text, one item per line?
Is this OpenAI catching up with Anthropic artifacts? At the same time they say "We’ve trained GPT‑6 to compose responses using text, visuals and interactive elements...", rather than a harness.
Looks similar but inline
I'm excited about generating UIs (and apps) on demand, because it's one step closer to devices just morphing to the interface you need. I want to talk to my phone and have it generate the app or interface I need. I could see Android adapting to this reality well before Apple.
I find this view fascinating. After decades of observing users getting lost by minor changes to UI or sometimes just different "paint" I am now to believe that we will all enjoy using completely custom UI customised not only to what I would want but what the app will think I need at particular moment and get on with it just fine.
I actually do think generating UI has some future, but I remain sceptical it can realistically go as far as promoters claim. It is especially bizarre to me when this is done by ostensibly UX people.
Just depends where you think the limits of future AIs are.
Not really because I'm sure they will be capable of a good enough design.
What I actually doubt is that people as whole are this adaptable to continuous novelty, especially when it is not directed by them.
I am with you. For example - this post has a comparison of a recipe app from 5.6 and 6 for comparison. They look different. Now if every time I ask for a recipe, it decides to render a new UI because of either the model changing or the recipe goes from stove to oven or something, it’s going to become annoying very quickly. It just seems like he current generation of models have been tuned to create better UIs than previous one and people interested in the UI are getting hit with excitement.
Based on current capabilities, I'm skeptical as well. However, over time, I do think they will get better at understanding our intent. Also, I would be more forgiving of an app I vibe coded in this manner, that's only meant to be used by me. Of course that assumes appropriate security guardrails are in place.
I really do think we are headed toward a world where stuff is hyper customized like that.
As a developer, it already feels like this with internal company tooling. Anything I have the source code to and a build pipeline setup for, I can just pop open a new tab and tell it what I want changed, and it just gets done for me.
Even my home assistant setup at home now, i hooked up codex to it via an MCP, and when we put out the blow up haloween lawn ornaments I was able to just pop open codex on my phone and tell it to throw together an automation to run them for me, and 5 minutes later it was done, with an override button on my dashboard. It's a shockingly nice workflow!
Wouldn't it be nice if you could just tell your app to reconfigure/rewrite itself to remove those "what the app will think I need at particular moment" features that you dislike? Maybe have it remove a bunch of the whitespace, getting rid of the lazy loading infinite scroll, whatever you want. It's the GPL dream, except you don't need any coding ability, just a subscription to a megacorp.
I do the same with Home Assistant and my many EspHome devices. HA works well, and it has a lot of nice features and extensive community-built integrations, but configuring and laying out the UI is a disaster. You can choose between WYSIWYG UI tools that are limited and cumbersome, or you can deal with raw YAML that stretches far beyond what YAML should be used for with deep nesting, and it's poorly documented to boot. It feels like a proprietary workflow engine, where you'd rather just write it in a real programming language.
Agents cut through all that crap and allow me to make changes without learning a new overwrought configuration language that thrashes syntax every few months. The fact that it is open and fully modifiable redeems it, because I can ignore all the annoying parts while achieving what all that was envisioned to allow.
More likely OpenAI and Anthropic just cut iOS and Android out of the loop entirely and become the new UI paradigm.
yep Ives is probably working on a product to do just that.
Lol yeah sounds nice on the surface.
Just like Codex failed so would that.
What emerged for me as of late, at least with Opus 5.5, is creating web interfaces as aids during development. I recently had to dial in some audio mixes using ElevenLabs audio and do some complex audio cropping and concatenation. As a interim step Opus built a website that allowed me to compare versions based on different rule sets and then really dial in the version that we decided to ship. That then led to 'us' landing on the right algorithm to build out the audio files.
I finally connected my home assistant server up to Codex the other week, and it's been pretty incredible. I can just send a message to codex and in a few minutes my dashboard is updated, an integration added, an automation setup, and sometimes all of them together at once.
And I don't need to build or debug complex automations any more! I can just tell some AI to make it so it auto-locks the house door when my phone isn't tracked at home, but also to check the wifi and if there are people using the guest wifi then to instead send me a notification asking if I want to lock the house because guests might be there.
I'm curious to see how useful this is. I feel like in practice the existing versions of this feel like they get in the way. While I'm sure I've had some situations where a visual would be helpful, there are two situations where I do not want it:
1) I want a quick answer, and I don't care for the boilerplate UI. For example, if I ask how to make pancakes, I make them all the time and just want to a quick reminder on the ratios, but it might trigger a full UI that I need to sort through to find information.
2) If I ask a to me unrelated to UI question and it triggers a big UI build that is completely off topic for my question (meaning I'm desperately pressing the stop button and prepping rewriting my query)
Or the other one I see, for example if I look up a unix command like:
"ls all hidden files in the /xxx directory"
And I get back:
"Sorry, I am am unable to find /xxx in my current environment"
IMO they're making a mistake trying to shove every possible product into ChatGPT. Just focus on making the core functionality better.
Isn’t that exactly what this is? How is better visualization not core functionality to what ChatGPT is?
Their core functionionality is/will quickly be a commodity
What's the core functionality?
This brings Mathematica’s Manipulate elements to mind. I love those, once you determine the variables and make it interactive concepts come alive.
At least seems a fist step out of the serial interface of chats.
The history of the internet to date:
People make simple text websites → Google comes along and indexes all websites → People can freely and easily find information! → Google slowly perverts the incentives with ads → Websites replace simple text with complex UIs and paywalls → People can no longer easily find information → OpenAI indexes all these complex websites and turns them into simple text answers → People can freely and easily find information again! → OpenAI perverts the incentives with ads → OpenAI replaces simple text with complex UIs and paywalls → People can no longer easily find information.
And so on…
I wish things like bikes came with mostly-written manuals. Maybe a few diagrams. If anything, things are too pictorial these days.
Intelligent UI for everyone is Bret Victor's vision for computing UX finally coming to fruition.
I've always found his work impressive, but for the most part it was all prototypes and proofs of concept.
Having such interactivity available to everyone now is nothing short of groundbreaking and another indication AI works as the great equalizer by democratising capabilities previously available only to a select few.
It has the potential to be that. Not sure this is it yet though or if openai's intention is even eventually to do something as useful. I think parent comment might mean what is presented in the links at https://worrydream.com/ in years 2011-2013.
Really exciting to see GPT-6 pushing intelligent UI forward. Making these capabilities accessible to everyone could open up a lot of possibilities for developers and users.
I've been thinking a lot about this transition happening, and I hope that everyone is recognizing that this is what the web will probably look like very soon.
URL-->user request--> answer & visualization.
No more static menus, buttons, lists, etc.
There's probably some good work out there helping small businesses adapt to this change? If done right, the real-time token use to generate dynamic pages can be minimized.
Honestly it feels like AI companies are still searching for the killer idea that will attract the general public (beyond "be my AI bf" or "better google").
Getting a UI that explains the parts of a bike is ok I guess, but isn't it simpler to get an actual breakdown? Google "parts of bicycle breakout" gets tons of useful images instantly.
Getting a specialized app to split a bill? It was already trivial to put in a calculator if we cared to go item by item on the bill. Having to provide names and tag every item as I go is just more work. Usually real people just go $total divide by 5, I had more, let me chip in an extra $10.
Same with booking travel and wedding plans, these aren't things people would even delegate to a trusted friend usually, much less a one-off request to an AI bot or custom UI.
Most of the useful tasks it can do right now are research, technical question/answer, coding. In terms of "build a flexible ui that solves real-world problem", if existing mobile app isn't useful in this arena, then its unlikely a completely custom UI will do the job.
That said, I've done a few small ones like a quick one to practice alphabet of a foreign language, or prototyping a web game, and the like. But ultimately its nothing that is worth trillions of dollars.
Um…the AI companies attracted the general public a few years ago.
In a way where it actually pays them well? Most of the general public uses the ChatGPT free tier, I'd assume.
An overwhelming majority:
Feb '26 https://openai.com/index/scaling-ai-for-everyone/
How would a user know whether a generated visualization to illustrate some process is accurate or not?
It wont they're just trying making the chat look like more a session iron man would have with jarvis
And wouldn't it be better to just reference high quality known to be accurate information vs generating it on the fly slightly differently everytime?
By asking the AI, which will then go "you're totally right to question that, Markus! Here's a thorough, error-free version:"
Is it just me, or did they make the examples in the video comically unintuitive? Painting swatches without color, a guide pamphlet without a map, assembly instructions without diagrams?
Is this supposed to be humor that I'm not getting, or are they really saying, "look at this solution to a problem you would have if you did normal things in the worst possible way"? It almost feels like the video is trolling us.
Thought the same. Do such maps really exist? Who would even use maps like this nowadays when traveling?
that's the joke. They are making fun of their own text-based interface to sell the visual side.
Thanks, that was the part I didn't get.
It seems fairly obvious that it's meant to be an analogy for ChatGPT only supporting textual outputs until now (disregarding that it had some limited capabilities to show static images).
I was mixed on this from this page, but for my use cases at least, I'm finding it immediately helpful.
The irony really is that LLMs are partly responsible for the walls and walls of text as seen in the video in the first place.
And now we are asking LLMs to solve it.
LLMs are the walls and walls of text. In the real world everything has graphics and illustrations.
I might be wrong but I remember that when Gemini 3 was released they also kind of promised something similar. But I’ve never actually seen it in real conversations. I’m not a heavy Gemini user so it’s possible it’s there but I haven’t seen a single demo showing this.
1.2B weekly active users? WAU!
I think most responses here are missing the bigger picture, this is the beginning of the end for individual apps and websites, a big rollup of everything into a single tool. A sad day.
This is a pretty baffling development.
Currently, we have highly specialized apps which do exactly what they need to do. They were created by someone who cares about making the program which works right and reliably. Most programs have years of bug fixes and deterministic algorithms. Many of them run locally and waste a minimum amount of energy to do the calculation.
But this is different - this encourages you to make one-off apps for each little thing you want to do. Imagine someone making a "bill splitter" app every time they go out with friends and how much energy that would waste... Just to do "14+16" that you could do in your head.
Also imagine it coming up with new UIs every time (depending on the prompt/context), and even potentially gaslighting you and doing calculations wrong because no one ever tested it before.
love that bill splitter
AFAIK, they already had a simpler form of this. It was kind of an obvious next step. Now connect it up to tools so that we can again comfortably do the things that are more precise by hand! The “pick a color” or “select the width on a slider” use case is coming closer.
It reminded me of when Twitter introduced Stories. I just want to read raw text without pictures, but that's probably not most people's preference. Understandable and sad.
"But when will we recoup our investments?" Investor to Gavin Belson after he presents a whole zoo of animals for analogies.
This is now trying to steal YouTube repair and cooking videos. Here is news for you: People prefer YouTube repair and cooking videos.
I donno, I like iterating on a recipie with chat rather than watching a video.
If the diagrams and descriptions are actually well-designed and written, I'd far prefer those. Especially one where I can ask clarifying questions. I'd much prefer a Haynes with some of the images lightly animated than having a 10 minute long video with a long intro for monetization reasons about how to do some simple thing where you can barely see what the mechanic is actually doing in the cramped, poorly lit, 360p video is trying to look at.
But getting actual service manuals cheap/free these days can be tricky, and their quality can leave a good bit to be desired.
cooking videos are entertainment
he he boi that's right cooking videos are entertainment if you actually cook you read something called a recipe (it's the algorithm that returns the food you want)
>People prefer YouTube repair and cooking videos
Not sure about the cooking one but repair videos are horrible. People with thick accent, the worst camera, shittiest lightning and bad angles where you can't see what you have to do.
The medium is the message here. As others have said, chat output largely sucks for anything but the most basic responses and we've done little to improve upon that foundational UI in the last couple years. We've _added_ a lot for specific domains like coding and document editing, but the primary content normal users get back from the chats is verbose and uninteresting. This is a step in the right direction.
Gettin closer to fully dynamic interfaces for a lot of software. Hell, give me a mode in Google Docs that takes every pixel of chrome away then vibe the rest as I need it. Persist across new documents going forward.
If it is for everyone, open source the weights and model information. This is a for profit enterprise built on non profit and theft.
Make it open or it is not for everyone.
IMO visualization is such a critical part of explanation - we have created an entire suite of tools to visualize - presentations, graphs, formatted docs, websites, svgs, interactive blocks, apps.
To that front - model providers generating visualization elements (i,e taking over the tools for interaction and visualization) is just a natural next step in value capture.
Anthropic launched docs, presentations, sites and am sure they have sheets next up their alley. Apps are artifacts. Tool forming is a natural model adjacency.
Wait, is this new? I feel like we all have been doing this with widgets / artifacts for a long time now, am I missing something? Just easier to use / done without asking?
I wrote something like this, though not quite as slick, by having the model use json schema to describe its results instead of text, and then use a json-schema to ui interface.
I am pretty convinced that businesses will be making custom internal software as a norm in the next few years, and this seems to be a slight push in that direction
Alright doing it here, all the functionality we need and no reoccurring yearly costs of tens of thousands of dollars. Took me a couple weeks to wire it all up with codex.
The intro video is just lame. Those things never happen in real life.
So you never have to assemble a bike or paint a room. How those are not real world examples? I find their video pretty cool. For people with ADHD or people learning by watching or children this is spot on.
They made fake, terrible artifacts (colorless paint chips, text-only city guides and instruction manuals) to show how terrible that is, and then "fixed" them. The problem is, in reality their examples don't exist. Paint chips have color samples. City guides have maps, pictures, and color. Instruction manuals almost always have illustrations. If what you made is better than what exists, you should compare it to what exists and not some alternate reality.
hahah yea i was wondering what kind of paint chips wouldn't have the dang color on it....
They're making the point—apparently too subtly—that text alone is a highly limiting "UI" for many tasks. So it's great that GPT-6 can now communicate in a a richer interactive medium.
Then show a real task where text is limiting. Don't make up stupid examples.
They’re making fun of the frustration of trying to use ChatGPT purely through text for these tasks.
Like I said, it’s too subtle a deliberate self-own.
I'm confused. Are they trying steamroll every developers B2C product or are they wanting these same developers to continue building MCP-App plugins on the platform?
It's much simpler: they want others to experiment (and pay them) and see what works, once they know what works they will drive out of market other players.
what would you wager the % of codex token usage is folks generating interfaces/building apps?
seems a bit like a snake eating its own tail
This seems like an expansion on something I was getting my agent to do which is use 'cards' to communicate using things like tables, rather than the ascii table stuff Claude code likes to do.
Unsure if this has been done already but the best thing I did was get it to create a list of next steps as quick buttons that it keeps updated in a docked card. It works well but occasionally gets stuck on some un related rabbit hole.
> The compiler allows the interface to appear progressively as the model generates it, without waiting for the entire response to be complete.
why not just stream html?
The element can’t render until it’s closed, presumably.
Html partial streaming is a real thing, presumably
maybe I'm just wrong here? If the tokens <b>come through this text could still be optimistically turned bold until </b> happens.
For elements that manage layout and drawing things like these explainers, though, it's hard for me to imagine how that would work. My immediate thought would be to an abstraction over it, which is what it sounds like they did.
They could also just be creating all the html elements using javascript (i.e. document.createElement, document.body.appendChild, etc) and streaming that into a repl line by line.
That's not how html works. The html parser is by definition streaming.
How about a truly intelligent UI? :)
Today i wanted to reencode and downscale a screen capture video to make it small enough to send to other people. I had a chatbot generate me the ffmpeg command line, but when will i be able to tell "Apple Intelligence" or some other "Intelligence" just "resize this to 720p, make it mp4 and drop the audio"? You know, like in Blade Runner.
Edit: oh wait, Samsung already did it like in Blade Runner, when they "enhanced" the moon photos until they had details that were impossible to capture with those optics...
You can already do that with Codex and Claude code. Computer use was the big GPT-6 Astra thing
I know, but they're not exactly built into the UI...
All those recipe blogs and sites that have figured out the most inefficient way to deliver information are finally cooked.
Great product idea. I have not paid OpenAI for a subscription for a couple of years, I’m tempted now. I do about 75% of my work using local models and most of the rest using the deepseek 4.1 flash API. That said $20 to experiment with this new product for a month sounds like a pretty fun idea.
I feel like I'm going mad looking at this page. The bike disassembly looks fake to me, like an alien would show presumed bike parts, but I can't tell why and don't know much about bikes. Maybe it's because there are no screws and nuts?
What is it supposed to demonstrate? That the model knows some kind of folk mereology?
Apart from the math stuff, maybe, this is the appearance of useful information rather than the presentation of actionable practical information.
I suspect it will struggle with anything that isn't simple enough for demo-mode or lifestyle fluff.
It’s demonstrating a lack of understanding of the average person by those at the frontier.
Remember when Steve said ‘the computer for the rest of us?’ We’re seeing that here.
That's a good way to describe it. The fork and seatpost are both missing the tubes that run through the frame. It's like someone cut up a photo of a bike instead of assembling it from parts. Even then the front wheel has very little clearance with the frame and the rear way too much. Also inconsistent that the chainrings stay with the chain but the cassette moves away.
The rear brake does not touch the wheel, the chain defies physics around the pulley wheels, and that's after looking at it for two seconds
I think this means that the API "chat-latest" model is now serving a custom GPT-6 variant: https://developers.openai.com/api/docs/models/chat-latest
Interesting, in Chat mode, if you expand the model selector it says "GPT-6", but in Work mode it says "GPT-6 Astra".
Sounds like Google's Generative UI (Nov 2025): https://research.google/blog/generative-ui-a-rich-custom-vis...
This seems like a step towards their plan for building a platform in competition of Google/Apple. Once the generative UIs are polished and more useful than individual apps, their hardware can now ship a device which circumvents the app store moat.
This seems like a natural progression of models becoming better at frontend coding in general.
"Here is a library of [svelte/react/whatever] components, use them to construct a helpful visual to demonstrate your point."
The deconstructed bike at the beginning was in a class of its own, however.
Interesting to see that city guides are still the #1 use case used for those demo videos.
It makes sense they’re doing this - I’ve noticed lately when asking a more complex question involving a lot of nonlinear data using Codex work mode, Astra and Sol will write a fully html document to better display the info with a Cliff’s Notes version in chat.
Gemini had this UI thing before .
Wait did ChatGPT make paint chips without the actual color printed on them, just the text description? and now we need ChatGPT to show us what the color is?
The idea is to show you how to solve non existent problem.
Imagine a world where you get paint chips with no color! Then you could use our product (you could also just look up the name of the colors).
I find the video strange. Are we meant to empathize with people who get flummoxed by everyday life? "Oh woe is me, I can neither assemble a bike nor ask a friend for help!" Of all the challenges in my life, the tiny mundane ones are the ones that can be most fun; those are the opportunities to laugh at myself, and to revel in small & happy victories. It's the charm of life.
I mean, I don't think the point was to say that everyday life is hard. I think the video was trying to say that certain types of instruction/communication is very bad.
I see your point for the bike scenario, but not for choosing a paint colors or touring a city (the other scenario in the ad). The common element that I perceived was that these were all mundane tasks in the physical world.
There have been recent developments that allow people to interpret Swift on device, live. With this new launch, will the AI be able to write native code and compile it on device? That would be super cool.
I'm also building something similar for an internal project, based on the vercel's json-render design. It hasn't been too challenging, especially since the release of faster models like 5.6 Luna.
Not sure why but the website is half in Dutch, half in English for me.
So I guess it's going to nail the bicycle the Pelican rides on, right?
"More than" 20% of the connected world uses ChatGPT each week eh? I guess if you add Anthropic who must also claim 20% and Google, Meta.. that does not sound realistic at all.
- Anthropic’s total user base is minuscule compared to the rest, and they have never claimed otherwise.
- People don’t have to stick to a single provider. They can all claim the same 20%.
Written manuals/guides come with pictures. The ad is dishonest.
Plot twist: The guides were written by ChatGPT. But now you can use ChatGPT to solve problems by ChatGPT.
The point, I think, is to make fun of Anthropic models, which answer only in text when you ask them how to do something. They’re the competition, not paper pamphlets.
> which answer only in text
This is false.
Is it? Sorry, skill issue I guess. I regularly use the Claude app on phone and laptop, and have never seen it produce non-text chat output.
I think it's been a thing since around March https://claude.com/resources/articles/claude-builds-visuals
Edit: be sure to check Claude settings -> Capabilities -> Visuals.
Thank you!
I like how, whenever a new model releases, they not only make other companies' models seem worse, but even their own "oLdeR" models.
The videos don't display at all, though sound plays, on MacOS Safari. Mobile Safari does work. MacOS Chrome (where they actually tested I imagine) is fine too.
Similar to https://www.monogram.ai/.
I guess the above is mostly mobile focused.
They seem to be pitching something that pretty much all the decent models can already do..?
Obviously if your model is stuck inside a CLI terminal, then not so much. But in a GUI harness (shameless plug for my own one: https://juggler.studio, but I assume others can do this too), you just ask them to answer in HTML and they'll happily draw pretty pictures inline in the conversation. I've been doing this for ages with claude, GPT, Deepseek and others.
For once Gemini did it first. I never found those mini GUIs useful though. More like a waste of time and energy.
Wait, didnt claude release this quite a while ago??
In the video, they showed examples of ChatGPT making interactive tutorials on how to fold origami, how to arrange colours/interior and how to assemble a bike.
Supposedly, people were struggling to follow written manuals and they needed an interactive explanations.
I'm not a mathematician and I would certainly love having a tool that would do ELI5 on some complex stuff, but I'm really worrying about using this too often and outsourcing my ability to do stuff to some mega corp.
Love it. I've been using $visualize a lot in the codex desktop app, and having even richer experiences will be sweet.
Too much visual noise. You don’t need a table comparing America now to the 1970s or whatever.
I noticed this change too! It is really game changer - to me ChatGPT output has no rivals here.
IIRC Google had an experimental project with this kind of stuff 2 years ago, but Google being Google...
It looks forced. It doesn’t look bad, but it’s not something I would pay for, nor would I give it access to my computer?
I also believe that a small talented group of developers just out of college can do just as good a job. That is why there is no moat around AI, it still comes down to original thinking and talent.
A Chat AI creating on the fly UX widgets is like Amazon’s brick and mortar stores
Is it me or GPT-6 Instant food answer gives the vibe to check the hell out immediately? It's like the it will start to explain a sunny Sunday afternoon from 20 years ago for 5 full pages.
Let's remember.. 'Climbers rescued from Mount Shasta after relying on ChatGPT for trip planning'
https://news.ycombinator.com/item?id=49542814
you can just say "ELI5" and the model will know what you want there or just ELI5
Is there any relationship between this and AG-UI or A2UI?
Wait, am I missing something here? Is this just
> you stupid silly human, instructions are hard! Who can read "fold this" and "fold that", when you can instead have interactive short-form content with music and smileys to keep you entertained while you offload all the thinking to AI?
... as a service?
"intelligent UI"
It seems like every week OpenAI adds some new feature to the chat UX that wasn't tested, doesn't work at certain resolutions, doesn't work on some platform/browser combos and hogs all the memory/CPU.
Intelligent UI my ass.
Edit: I should add, features that no one asked for, too.
i saw that and thought the death knell of apps is coming soon as on demand interfaces become a norm
> THE DINNER LEDGER
> The Check, Please.
> TABLE OF FIVE · ITEMIZED BILL
This is one of my _least_ favorite AI behaviors, I am surprised they left it in the example. This behavior where it insists on putting text everywhere that over-explains the context.
for the sunday lamb roast instructions, how much is 1500g of potatoes? What kind of potatoes?
> Also pick up garlic, rosemary, thyme, lemons, honey, almonds, olive oil, gravy ingredients, mint sauce, crumble topping and vanilla ice cream.
what??? how the hell are you supposed to remember all that. Even the before example tells you what kind of potatoes to get. tbh it would be cool if you could just click a button and get an order pre-filled out on a grocery delivery/pickup service. I really wonder who reviewed this post and if they cook, because imagining yourself in that scenario and reading those instructions falls apart very fast
'gravy ingredients'
i didn’t even see that. add the ‘crumble topping’. this is legit terrible
You can get a an order prefilled
With computer use Both Astra and Sol have been able to make sense of my markdown recipes, and based on those place prepare a shopping cart for me to review and finish buying
Probably a bit much for the average user though, even if it's trivial to setup
Is the page about intelligent UI made with Intelligent UI? It should be.
First thing to ask is a design teardown of AI Slop images into individual artistic choices.
Jokes aside: always having the right UI on a dataset is a very cool promise and this looks like a leap in that direction.
Instead of apps or websites merely APIs and intelligent UI in front. With the ability to render data gathered across multiple domains
The interaction is beautiful, reminds me of Brilliant.
How so?
Why show an intelligent UI within a chat interface. That’s pathetically lame. Better to go away from chat interfaces to rich visual interfaces. Doesn’t sound intelligent at all to me
I gotta admit, I love how their design system and how they market things. It looks so good. However the true state of their toolings is closer to a burning dumpster fire. If you look at Codex Github repo for example, the amount of existing issues and continuously new issues appearing while tagged releases is being spewed out. Its truly a nightmare. Their vscode extension has been broken for almost 1 week now, if not more. Its instead a community effort to repair it.
This seems like a nightmare for standardization and accessibility, no?
Has AI not been always?
I made a prediction last year that there'd be a new UI protocol (like HTML) but for agents. I believe that custom or personalised UI's are going to be the new browser interface. More radically, I think browsers can be completely replaced. My news feed can be personalised to me, based on what AI thinks might be important from all sources like Reddit, X, HN.
Sounds horrible, but plausible.
Now they extracted the value from interoperable open systems, they would love to replace it with closed proprietary systems.
Which brings up a different problem: personalized reading UI, personalized writing UI, how are those services going to pay their bills? Maybe they can sell access to their APIs but who's really going to buy that? Those services will die and will be replaced by some other service that will be born to serve those people.
Or, more probably IMHO, most people will keep using the standard UIs because they are ready and they need no work to build.
Cant wait for the next round of outrage and radicization machines, this time by browser replacement and unavoidable.
Why are you posts generally down voted?
> I made a prediction last year
yes you and everyone else.
So they are going to replace all the banking, utility, social media, news, gaming, streaming, government, work apps and other important websites people still rely on? Those companies and organizations are just going to provide an API to the chatbots?
How do sites like Wikipedia get updated? What if I need to upload important documents to different portals? It all has to go through the AI companies? They have access to medical records too? My taxes, social security, etc?
Is this how Microsoft imagines Office 365 and Sharepoint will be merged into? Disney, NY Times, Youtube, Netflix will all be fine with chatbots handling their content?
Well, there’s this: https://www.openui.com/blog/oui-1
gemini app kind of had that for a while I feel?
I wonder if OpenAI will be able to vibe-code a button to toggle this feature off.
> "Can't you read? No problem, now you don't have to!"
It's interesting how the bicycle is a fake example. You see Jony Ive level design, with a very complex 3d object seamlessly working, and then you see the basic HTML slop in other examples and you begin to wonder if that first one was fabeication.
Gemini already does this and I find it pointless. Little interactives seem innovative in practice but generally don't do anything useful.
Off topic, but you can tell the Sunday Roast image is AI because it makes British food look appetizing :P
sounds like JEV with extra steps
"If you call something intelligent you know what it isn't."
> Our mission is to deliver the benefits of artificial general intelligence safely
extraordinary AGI claims require extraordinary evidence
> We’ve trained GPT‑6 to compose responses using text, visuals and interactive elements, choosing how they fit together based on your question.
Ok so this is how you build a moat. Models are not a commodity and visualization isnt a purely harness problem.
> We expanded our training methods to help the model make thoughtful decisions about content, layout, visuals, and interaction. This included evaluating the interfaces it creates for clarity, usefulness, and completeness. GPT‑6 learned to use the component library and make good design decisions, including how to organize information clearly, when to use interactivity, and when a simple text response is enough.
Curious to know how they trained it produce appropriate visuals. And why that cant be done with propmt engineering.
Nope
This looks bad, I want the output to be as much as text based so I can easily export it and further processing it, plus, this might make the resource-eating app even worse.
Visuals are better for consumers/general public, when the use case is "help me cook X meal", or "where can I stop by to buy gas on the way to San Jose".
Rading most comments, people don't seem to be very impressed by it, I ain't either tbh - its not something mind-blowing (I have been doing this with Astra myself just by writing better prompts to make explainer interactive interfaces), but I do think its a good first step towards a better way of consuming information/answers than just reading text-vomits. I wonder if there is a way to train an LLM native to this kind of thing?
where are you reading the comments? Wasn't it just released? I can't imagine anyone I'd consider "general public users" to know about it yet
I meant comments in this post.
then just ask for text-only?
I think this is really a future, this is the product that I really expected to become reality at some point. If we assume agents become even more common in the future, then I don't think we would really have a lot of modern software. Most of it is really is not needed at all.
Most apps that people use are really very similar and don't have anything original. It would make more sense for it to be more personalized. For example someone might prefer not to interact with interface at all, and just access services just by chatting with a bot, while other person might want only last steps to be provided as UI. For example to see summary of his cart before paying. While another person might be insterested in just browsing all options. Some might prefer to have filters, while other would want AI to filter everything for them.
I think the real result would be when all services would be automated by AI. Like if you want to be a small business, you no longer need to build anything. You just describe what real life services can you provide: delivery, barber, baker, cleaning, repair. Then all of that would be in some agent network, and available to other agents as an option. While humans on both sides would just get a bridge between them in whatever form is most comfortable for them. Buyer might get a list of bakeries as normal ecommerce website, while the baker might just be someone who receives phone calls explaining him what his next order is in human voice.
Great, now instead of doing math to split my check I can vibe code an app on the spot to do math to split my check.
I read the comments here and I am really surprised by many people. One picture is worth a thousand words so with better UI you can make much better user experiences. This unlocks personal assistants to be better adopted by elderly or disabled people. I see so many benefits of it and having models which can do this (if they can do it constantly with good quality) is amazing. Much better products - I am really tired of dumb chatbot - if I can do something with one button or view the whole information in one diagram/image, this is amazing.
"why do I need an Nvidia B300 to render the same webpage that ran fine on my Pentium with Windows 98?"
I truly hope this ushers in the end of Markdown as the preferred AI document format.
It's time to promote an .htmd HTML document standard so that we can actually have nice looking documents again. Remember columns? Colors? Typography? Layouts? Graphic design? Interactivity?
Just think - we could actually have a true rich text standard! Imagine if it were adopted everywhere and you could send WYSIWYG bold, italicized or underlined text as easily as you can a custom skin-colored emoji?
It would almost be like were living in the 21st century again!
There's a type of UI that is very in vogue currently, it's kin of Futuristic UI that laymen find very cool and seeks to appeal to a WOW factor that suggests that the future is now.
I used to see it a lot in China, like the idea of pinching something on one phone and transferring it to another, or identification by showing your palm.
It's also the type of UI you would see in Hollywood movies, both because it's stuff that laymen scriptwriters and directors would find interesteng, but also the audiences would. Like the motion based UI in Minority Report, whatever was in the movie hackers, or that laptop suitcase thing that allows you to deploy nuclear codes or send wires.
I mean, I don't want to be a snob, clearly people like it, at least initially, I think it's outsider art of sorts, of course it will have problems, but it's not less legitimate because of that.
Anyways, this Intelligent UI thing where you ask it how a bike is made and it unravels a hollywood visual report as if it were showing the interiors of the target in a war room conference, it reminds me of that type of futurisitic outsider UI.