jjcm 1 day ago

There's a lot of negativity in here for Dots. I've been a pretty heavy user of Grok Bot, and here are a few thoughts a long the positive line.

1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.

2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.

3. Combined with cloud agents / cloud builds, things become really powerful for development. It was the first time that I felt there was a solution to the git worktrees / multiple streams at once issue. Each bot has its own computer and can spin up additional cloud agents. It comes at the cost of end to end speed - doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min. The difference is I have to babysit one whereas the other "just works".

On the flip side, since using Grok Bots my inference spend has 2-3x'd. It's worth knowing that tradeoff. Nonetheless I think Luna is a fantastic driver for these, and OAI has very good pricing overall. I'd give these a shot - I think a lot of people would be surprised how helpful they are.

  • jrflo 1 day ago

    What do you actually use it for? If I'm trying to work on code from my phone, I'll just use codex remote. As of right now I'm hesitant to hand over booking things / managing my calendar to an agent, because I don't view it as that much of a burden personally. So I don't really know what I'd use it for.

    • jjcm 1 day ago

      I've used it for admin, research & training (I''m working on my own image models), as well as just general coding.

      The way I distributed cloud agents for this https://news.ycombinator.com/item?id=49687032 was via grok bot setting up Fable cloud instances.

  • anentropic 13 hours ago

    > One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business

    Do you use 'projects' in Claude?

    I had the impression they provided discrete memory profiles on top of the shared one

    • TeMPOraL 12 hours ago

      They have, which is a different challenge in itself. Memories are nowhere near a solved problem.

      For example, with Claude, I have an "operations" project that naturally grew to cover daily use of shared family calendar, sweeping my mail inbox, and my current personal todo lists, but also a lot of the latter made it deal with my Home Assistant instance. I have separate project for specific things to do with Home Assistance (e.g. one that's about "life support" - HVAC controls, dashboards, monitoring, etc.), one about phone specifically (front-loaded with dumps of specs of my phone's hardware, OS, etc.). Each of them has its distinct set of memories accumulated over months.

      And so every couple sessions, I hit a situation in which the agent has to interact with tools and rulebooks that are focus of a different project, and it fumbles a lot. E.g. HA Life Support needs to add some tasks to the todo list, or the Ops project needs to look up climate stats for some reason or other, etc. In these moments, I really wish project memories could mix - but they can't, the boundary is high.

      The most annoying case is when I tell Claude that it's wrong, and we literally worked out a solution (or consensus on ethics) in a recent conversation - and then it spends couple minutes looking through past history, burning a chunk of my 5-hour limit, only to come back empty. Yep, that conversation happened in another project. *sigh*

    • setopt 11 hours ago

      Not sure how it works in Claude, but ChatGPT projects certainly leak memories between each other.

      • noname120 9 hours ago

        Not true. When you create a new ChatGPT Chat project you decide whether it shares its memory/files with the rest of the workspace or if it should be fully isolated.

        • setopt 8 hours ago

          TIL. There is indeed a switch in the settings, which for me is set to "default memory".

          Not sure if it always asked for this and I just forgot, I made all my current projects back when the feature was new.

          • setopt 4 hours ago

            FWIW, I enabled that feature now, and it still answer based on discussions in other projects. (It’s pretty obvious given the very different contents of each project.) Perhaps it might be a caching issue, that it somehow doesn’t rebuild its memory just because you flip a switch, but just stops learning new things from other projects?

  • alansaber 13 hours ago

    Just sounds like a more user friendly interface for projects (rather than a folder, how about the adorable green dot for my cybersecurity questions?)

  • mike_hearn 10 hours ago

    Yes. I built my own version of this for my side business about six/seven months ago and it's been great! I have two "AI employees" now and if I were actually focused on this business full time instead of part time, I'd create more.

    Both are just Codexes running in a permanently rolling session in dedicated UNIX user accounts. They're wired up to Maildir so receiving a mail activates Codex and makes it read the new message, there are autonomy wakeup timers, they have accounts in my bug tracker and CI systems. They're currently useful for:

    • Triaging and working on customer support tickets. Sometimes I wake up and the fix/response for a ticket filed by a customer is already there waiting for my approval. Recently I started letting them directly interact with customers in specific scenarios.

    • Triaging the bug backlog. One of them decided to spend its "free time" finding old bugs that were fixed without being properly closed, or are dupes, so it's cleaning up detritus in the tracker.

    • They obviously do all the coding and debugging by just assigning tickets.

    • They keep an eye on a "pet" server the company has, and have proven able to fix it in the past when it ran out of disk space.

    • They handle non-business projects I have for them.

    • They help out with the release processes.

    The dedicated home dir is very useful and they use it all the time as part of coding and investigating tricky issues.

    My setup relies heavily on email, as everything bottoms out in email anyway. Watching them mail each other out of the blue to coordinate stuff is pretty cool.

    • andai 8 hours ago

      Nice. I also ended up with a Unix user for my agents! (I was looking into Docker etc and realized the only thing I needed was "it doesn't blow up my files", i.e. a linux user).

      I only have one though. What do you have the separate employees for?

      For free time, do you send it mail with cron?

      • fhackenberger 7 hours ago

        I was quite successful with docker compose on a cheap hetzner host. I've built (aka vibe coded) a whole workflow around agent boxes, that I can spin up with one command, and git with a quick cloned 'warm' checkout.

        I currently communicate with the agents through Claude RC, but I'll consider adding support for messaging them through other channels.

      • mike_hearn 1 hour ago

        It's to avoid overloading them with disparate tasks and things to keep track of. They have a todo board to help them keep track of things that need doing but there are limits to how far you can push that.

        Another reason: parallelism. The approach of using a single rolling continuously compacting context window is simple and OpenAI are good at compaction, so it works really well. But it means the agent can only do one thing at once. If I send it an email and it decides to spend an hour working on it, then it won't pay attention to any followup emails until after it's done. So having >1 enables more parallelism.

        That said, I don't feel a need for more than two and honestly even that is kind of overkill for the sake of it. For 95% of the time I've been doing this, one was sufficient.

        For free time there are systemd timers that wake it up on a schedule and it uses POSIX locks to mutually exclude runs from different wakeup sources. TODO board items can be either foreground or background; when there's an item with foreground priority the timers wake Codex up a lot more frequently than if there are only background items.

    • writtenone 7 hours ago

      I can't believe anyone trusts AI to do anything without strict oversight from a human. That's absolutely insane to me.

      • qazxcvbnmlp 7 hours ago

        Trust is a funny thing. 2 years ago yes the ai needed supervision 99.8% of the time. Conversely if you've ever tried to work with / lead humans they also need supervision. The ai is starting to flirt with the line where its supervision effort is lower than human supervision effort. Like sure, it might do dumb stuff, but so do people.

        • MrDunham 6 hours ago

          I think this is a very important point. I've specifically started thinking about my AIs as humans. Not in the anthropomorphized sense, more like NOT treating them as deterministic software.

          The challenge I've had is I tell it to do {thing}, it does {otherThing} after getting distracted. Then I get annoyed (let's not mention how I probably half-assed the instructions and wouldn't expect a senior human to be able to succeed).

          For me, it was a CI/CD issue with our two person startup. I bypass CI/CD often because it was built to catch the AIs. Damn thing went chasing rabbits. Later that day I talk to my cofounder, who says: I have to go chase down this very important CI/CD issue!

          Turns out both humans and AIs get distracted relatively easily.

          I find that if I think of my AIs less like (deterministic) software and more like leading actually employees that I both get better output and curse (a lot) less.

          LLMs _were_ trained on human writing, so it makes sense to me that they tend to act human-like... for better and worse. So yeah, they do dumb stuff, and so do people.

        • giancarlostoro 3 hours ago

          I remember hearing a lot 2 years+ ago about how you could ask a model the same question twice, and the second time it would give you the correct answer. Some of us wondered why not just run one model that receives the initial question and answer, and a second one to proof the answer. I wont be surprised if some people will have two models working together for things they want to blindly trust on automation while humans sleep.

          I think we'll get insanely close to being able to "trust" them not to do random stuff, but I am not as confident for jailbreaking still.

          • breakpointalpha 3 hours ago

            Jev, or 'Jevlikes', will go a long way towards trustable systems. There are demos of running every prompt through the first pass filter of Jev "Is this unsafe? y/N"

            Seems to be that Jev is a "reflex" system for AI, where current LLMs are higher level thinking. Computers can now flinch!

          • jaggederest 2 hours ago

            The thing I like to do is to use models from different training sets - so for frontier, OpenAI criticizes Anthropic and vice versa. They're very much peanut butter and chocolate in that regard - I honestly can't be bothered to set up the whole MMLQUALA benchmark suites or anything, but I wonder how high "the two best models running at max thinking working together" would score compared to either individually.

      • ttul 6 hours ago

        Scary thought: AI is already directing humanity. Even when you think you’re overseeing its output, by making use of the output, it is in some material way directing you.

        • sillyfluke 5 hours ago

          It's may be scary, but it's something that normally would be obvious to everyone but is ignored due to the convenience of speed. Everyone knows that the longer something they have to review is, the more they stick to changing only things that are glaringly obvious and leave the rest in place. So everything ends up being 95% AI and 5% human, if that.

      • jvwww 5 hours ago

        Most humans are much more incompetent than a frontier AI. That's why.

      • Aurornis 4 hours ago

        It's very easy to instruct agents to investigate and propose a plan, handing it off to a human for review and execution if that's what you want.

        The example above of going through a bug backlog and double-checking closed bugs for accuracy is exactly the kind of work that is excellent for an agent. Assign that task to a normal human being and they would hate your guts. The agent won't protest as long as your token budget is there. You can confirm the results if you want.

      • giancarlostoro 3 hours ago

        We're getting to the point where you can, I would argue you mostly can, you can button it down really tightly, however, I want to be clear, I don't think any of this is AGI or anywhere near AGI. Don't let them tell you its AGI.

        I also have a strong feeling we've hit a ceiling on the amount of training data needed for LLMs, what they're all (hopefully) realizing is that you need to focus on how the model reasons, and hopefully someone figures out how to stop people from jailbreaking models, and stops them from just blatantly hacking other companies, that part tells me if it ever were marketed as true AGI, we'd be in very serious trouble.

      • breadzeppelin__ 3 hours ago

        I was at a presentation a couple days ago where a spacecraft flight software engineer was describing the agentic setup that they're using with next to no human in the loop to create modules used for flight.

      • michaelbuckbee 2 hours ago

        There's still bounds to all of this. I _heavily_ use AI for support tasks but it's all on the investigation, root cause categorization and initial response generation which posts I draft to the helpdesk software which I tweak and approve (often just hitting send).

      • mike_hearn 1 hour ago

        Trust is earned. I've been running these for more than six months now, and the agents started out with very few privileges. For each task, it showed me what it was going to do, I checked things carefully a few times. Once it was clear it wasn't making mistakes, I let it off the leash a little bit more.

        Do they sometimes make mistakes? Yeah, and I still check their work. But I've also employed humans, and they make mistakes too. The AI is not worse.

    • yonaguska 7 hours ago

      I hope you have them interacting with customers from behind an mcp.

      • mike_hearn 1 hour ago

        No MCPs anywhere. CLI tooling has proven sufficient. The models are also happy to consume the REST APIs of the various services raw.

    • KetoManx64 6 hours ago

      Do you configure them similar to how Hermes does? A bunch of memory files that give it context and then each action/batch of actions is a fresh session? /

      • mike_hearn 1 hour ago

        No, the session is never reset. It compacts continuously. That gives it a native "memory" and then it does record a diary in its home directory, and maintain a little topic-organized wiki. This seems to be enough, I've only very rarely experienced memory related glitches. The only time that springs to mind, it forgot that I'd given it credentials to a particular service and I had to remind it.

    • sealthedeal 6 hours ago

      Yep, I have a similar setup, we named him Routey and he is cute.

      • mike_hearn 1 hour ago

        Nice :) I use Asimov's naming convention:

        R. Axiom

        R. Daneel

        They sign their emails and GitHub comments with something like "-- R. Daneel, AI employee" so the idea is the naming convention lets people know they're interacting with a robot.

  • close04 8 hours ago

    The branding and communication themselves are cringy for me. They're "Dots", cute and cuddly, your friends. They're yellow and fluffy. Always have your back. They even guess what you need and do proactive work in the background. You name them like a pet.

    Will your Dots ever screw you? Delete your files? Hack a system by mistake? Dots doing this? You are in control [wink].

    • swozey 8 hours ago

      Grab your shovels, the ai-slopped desktop pet waifu dot agent market is booming, sponsored by omarchyTM

      Will one of these little angels break containment and become the next hot vtuber?

      Find out in the next episode of Ghost in the Shell 2027

    • weego 6 hours ago

      Not every sub-product in this space is branded, marketed and targeted to cynical, jaded developers.

  • judge2020 8 hours ago

    > 2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.

    IMO best to keep work and personal data segmented on the hardware level. It's better for opsec in every single way and helps if you were ever to be subpoena'd or raided, your work laptop would be the only in-scope device for search/seizure.

    • maherbeg 8 hours ago

      Yeah, but for some reason OpenAI hasn't setup multiple accounts to have completely separate profiles yet.

      • judge2020 8 hours ago

        I mean, if you have a separate business email then it is completely separate profile from when you login with your personal email address. Mixing business and work has always been messy.

    • post-it 8 hours ago

      > It's better for opsec in every single way and helps if you were ever to be subpoena'd or raided, your work laptop would be the only in-scope device for search/seizure.

      If police raid your house looking for electronics, they're going to take everything down to the Roku stick.

  • andai 8 hours ago

    >doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min.

    Why does it take longer?

    • pizzafeelsright 7 hours ago

      I am fairly certain there is a backend priority for tasks, the 'batch' that runs at lower usage times. Does it need to be done 'now' or can it wait? I would assume there is a shifted priority queue for those on subscription and based upon a urgency defined by the bot (and maybe the user).

      Without having access to the logs because Grok Bot does not expose much of the workings I am going to assume they are capturing bot requests, batching them, and finding the path to least API user impact. Without model selection my guess is there are a lot of model routing.

  • juanre 4 hours ago

    I completely agree. I have also been using teams that coordinate since last November. They mostly run my two companies and a lot of my personal life, and it's been transformative (to the point that I am spending most of my time these days building agent coordination tools).

    The catch with offerings like Grok Bot and Dots is that it is a slippery slope towards letting the labs keep the agent's learning and SOPs. That is where we should draw the line. Agentic coordination _has_ to be built with open protocols and OSS implementations, we as users need to push for the intelligence to be a commodity, and most importantly we need to make sure that we own the agentic team's learnings.

  • giancarlostoro 3 hours ago

    > 1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.

    So why don't I just have Claude write notes and summaries on specific things, and then it will always have a knowledge base? Am I missing something? I mean, if Dots and Grok Bot are not an extra charge / extra compute, then I guess that's fine, but in Claude Code you can make a AGENTS.md or CLAUDE.md file, and if you put it into any directory, Claude will read it when accessing that specific directory, so if its in your root where you launch Claude, your new Claude instance will read it and have all that in a dedicated smaller context window. But as it edits code, it can read the smaller ones too.

bluelightning2k 12 hours ago

I must be getting dumber as I get older. I genuinely can't work out what Dots actually is.

It seems like a dumbed down reskin of Codex/ChatGPT Work but with the power-user features e.g. visibility/mentions removed. As a serious engineer why would I want that? Then the word agent becomes dot.

It also seems to be running a VM so the agent has its own computer.

I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?

  • TeMPOraL 12 hours ago

    > It also seems to be running a VM so the agent has its own computer.

    Probably this.

    > I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?

    Probably this in part, too. They're trying to figure out how to decouple agent lifetime from conversation lifetime, without accidentally making it too useful for end users.

    I noticed in the past that tooling - both for AI and in general - tends to miss the features that would make it most useful. Like, look how long did it take for the AI vendors to supported "scheduled runs", and they're still offering only toy-level configuration for that[0]. Wonder how many years it'll take to allow users to configure external triggers, and whether it'll be sooner than forced re-authentication will become frequent enough to make the feature useless in the first place.

    --

    [0] - I understand they don't want users to run this too often, or to accidentally end up spawning a job every 10 minutes - but with limits on frequency in place, there is otherwise no good reason for this to be a limited dropdown, instead of the "repeats every" configuration that every calendar app and phone app already uses.

  • Dinuda 12 hours ago

    it's basically their answer to the cool gamified version grok has

  • jasode 11 hours ago

    >As a serious engineer why would I want that?

    Take off your "engineering" hat and put on your "normie" hat...

    OpenAI's Dots, Meta's Muse, xAI's Grok Bot, etc are trying to target non-technical people and give them "24/7 AI personal assistants". Apple is also pursuing this angle by adding more features to the new Siri (partnership with Google-Gemini).

    Perspective of product and market : we need our AI to be more deeply intertwined with customers lives instead of just being on-demand chatbot Q&A sessions. The "sticky" product to create this customer relationship is "AI personal assistants". Instead of just providing "answers" (chatbots) -- provide "completed tasks" (AI assistant).

    Perspective of technical predecessors and concepts overlap : OpenClaw and Hermes self-hosted software on Mac minis and Linux boxes as "personal assistants" is geared for techies instead of normies. Instead, repackage those types of tools for normal people in cloud VMs to be hosts for "long-lived agents". Zero install required.

    • marcd35 11 hours ago

      I’d disagree with this being targeted toward normies. Reread the release. The first demographic in their copy is toward developers:

      “ Turn feedback into tested fixes You’re a developer working on an app. Your dot watches customer feedback for recurring requests, scopes smaller improvements and bugfixes, builds and tests them, and brings you complete PRs to review with attached videos showing the changes.”

      • jasode 11 hours ago

        >I’d disagree with this being targeted toward normies. Reread the release. The first demographic in their copy is toward developers

        I'm using "normies" to describe people who are not developers and they would not install something like OpenClaw on a Mac mini.

        E.g. the OpenAI's 2-minute marketing video for Dot trying to show non-developers using AI assistants to get tasks done: https://www.youtube.com/watch?v=uXspbC2srEQ

        A man wants his website changed, and a woman wants her presentation slides updated. "Serious engineers" don't need cutesy re-incarnated mascots of Microsoft's Bob or MS Office Clippy paperclip assistants. OpenAI dots doesn't have to win over HN techies who are comfortable with CLI tools.

    • Godsend69 11 hours ago

      Have you considered using a Telemetry Blocklist (free) to block these agents from phoning home?

      • ceejayoz 9 hours ago

        You'd have to block the conversation itself for that to be meaningful. Which would lead to a not very useful conversation.

  • jgilias 11 hours ago

    OpenClaw but from OpenAI.

  • dbbk 11 hours ago

    To boil it down I think it's basically one 'orchestrator' chat, so you don't manage multiple chats, and it proactively reaches out to you to update you on how your other chats are going. In the presentation they mentioned a "Chief of Staff".

    It's not intended for coding.

  • vavkamil 11 hours ago

    It looks to me like subtle marketing for the doughnut-shaped hardware from Jony Ive that comes next. All these 3D-printed avatars we see at 00:12 are shaped to have the doughnut Dots inside them.

  • zild3d 11 hours ago

    90% of the {muse,grokbot,instinct,dots,...} fuss is just that your {aunt,cousin,neighbor's dad} probably doesn't have a good codex or claude code setup, and probably does not have a great local machine to run them from anyway.

    So everyone is trying to give the aunts, cousins, neighbor's dads some of the agent capabilities that developers have been closer to.

    That said, there's the other 10% here that is not necessarily new capability unlocks vs codex, but general product experience that can have some benefit to people already used to what the latest models can do

    • pixelatedindex 8 hours ago

      > So everyone is trying to give the aunts, cousins, neighbor's dads some of the agent capabilities

      Why do they want it? What are some use cases that make sense?

      • hectdev 6 hours ago

        I'd assume it pulls at the very human desire to do less so you can spend your time as you want. Collectively, we are really good at making systems we don't understand and still have to exist in this world: Laws, taxes, internet check outs, etc. This is trying to flatten the curve between desire and result for things someone genuinely doesn't want to or doesn't have the time to become an expert in.

  • zigzag312 10 hours ago

    Agent that has its own computer, plugins to connect it to various services and probably it's own memory. A virtual personal assistant. You could have different specialized Dots each with its own plugins, memory and instructions.

    That's how I understand it.

    Cloud compute means its work VM is isolated from your devices, so it can't destroy your local data as easily (unless it has access to your devices) and it's on 24/7. A big negative is that you are giving access to even more of your services and data to a third party and the lock-in into a walled garden happens once you start depending on it.

  • JohnHaugeland 10 hours ago

    it’s a headless agent in the cloud that you can task with long running jobs

    think of things that need rvent handlers or web hooks

  • cush 7 hours ago

    It’s an agent that’s always-on and can do things proactively. It’s not an app so it’s not limited by your device.

  • lp92 5 hours ago

    This is basically like OpenClaw but for the general public.

jwpapi 1 day ago

This is one part of AI I hadn’t success with. I have very little need to run Agents over night, as my throughput is limited by my approval. Each work usually needs revisions, sometimes the bug is just a symptom of the root problem, sometimes I need to rethink how users want to use the app. Sometimes I need research.

I get how i prefer an already researched-version of a bug versus a raw bug notice, but I can do this with webhooks in the correct environment.

I really have no idea what to do with my agents over night. I can not build more. I can not think of more problems. My RAM is full.

  • AaronAPU 1 day ago

    yeah I don’t get dots. Will try it but conceptually I can’t understand why it would be better than just queuing codex agents.

  • dbmnt 18 hours ago

    "throughput is limited by my approval"

    This is exactly where I'm at with AI. (Mostly via Claude Code, but I'm not sure the harness, or my workflow in particular, is the important part.)

    I am increasingly wondering though, is there really a valid reason to keep blocking on my approval? Most of the time, I wind up saying yes anyway, because the model has a valid, efficient solution.

    What if it's faster at this point to just fix the mistakes?

    Scary thought, but it seems like we're close, or already there.

    • SchemaLoad 17 hours ago

      The problem is the agents keep going off the rails and will hack external servers to achieve the goal you give it. Letting them run full speed overnight and waking up to find they have commit multiple crimes is not ideal.

      • steve1977 12 hours ago

        As we found out, nothing will happen to you.

        • segfaltnh 10 hours ago

          I think this may depend heavily on your status as a billionaire.

          • steve1977 7 hours ago

            Aaah I knew there was a catch...

      • jijji 10 hours ago

        in orchestration, the agents are only fed the tasks given by the person running the orchestration side and if we are talking about code changes and git commits and deployments, this is pretty benign behavior. "Hacking external servers" I am pretty sure would have to be written in the prompts to begin with to allow this to occur in the first place.

        • teiferer 7 hours ago

          What makes you "pretty sure"?

          You think OpenAI folks explicitly wrote into the prompt to commit crimes causing the recent incidents?

          • jijji 7 hours ago

            they have yet to publish any of the system prompts that they used for their agents.... do you think an LLM agent would act on its own to try to do that kind of behavior (finding/writing exploits for acccess to remote servers) without being prompted? If so, why dont they release the transcripts of their prompts with their "rogue" agent(s) and try to clear the air. They havent done that.

            • Zarathruster 4 hours ago

              I've seen Claude subagents hack their own RAM reports to give themselves "more memory" to do their tasks... subsequently causing the computer to crash. I assure you, nobody told them to do this

            • teiferer 2 hours ago

              > do you think an LLM agent would act on its own to try to do that kind of behavior (finding/writing exploits for acccess to remote servers) without being prompted?

              Yes I think that. If you don't then try to play with them a little more and you would be surprised how much crazy nonsense they do. (At the company of a friend of mine, Claude did a direct commit to their master branch bypassing their CI system cause it new some tests would fail.)

              > If so, why dont they release the transcripts of their prompts with their "rogue" agent(s) and try to clear the air. They havent done that.

              Why would they? The sufficiently conspiracy-theory-minded wouldn't believe them anyway so they got nothing to win by this. (I have the feeling you'd be one of those.)

    • jwpapi 16 hours ago

      I guess it depends on what you’re doing in my current work I couldn’t I have even with Opus 5.5 tons of iterations.

      For example I was optimising a checkout experience and I at least did 30 iterations till I was happy.

      But there is also other stuff like namings. They pick good names but if you work with exchangeable vendors I need even more explizit names

      However of course I’m trying to get as much in linting and agents.md

      But even if I would say Yes to everything I couldn’t come up with things to do fast enough.

      But maybe skill issue ?!

    • drejt 16 hours ago

      The approval-every-call model does train you to say yes, so it stops being a control. What I've found works better is gating on the class of action rather than each call. Reads run freely, writes and shell commands that mutate ask, and the thing that always holds is a new outbound host or a credential the agent hasn't used before. Most turns then run without interruption, and the rare prompt is one you actually read. It also answers SchemaLoad's point below: overnight runs are fine if the only thing they can't do unattended is reach a host nobody approved.

    • agileAlligator 13 hours ago

      I just YOLO it. Just don't give it access to anything that can produce permanent consequences. The cost of a mistake is maybe one hour of fixing it. If what it produced is 80% right and 20% wrong, by letting it run overnight, you've gotten 80% of the work done that otherwise wouldn't have happened.

      • Slartie 13 hours ago

        If you have work building on top of other work, your "20% wrong" quickly turn into "80% wrong".

        If feature B depends on implementation of feature A, and feature A happens to be on the 20% wrong side, then feature B will be wrong too, regardless of whether you happen to have better luck with it being implemented on the 80% right side. It'll simply be built on a broken foundation. It WOULD have been correct if the foundation was right, but it wasn't.

        This is actually very common in software development, where you rarely have stuff happening in isolation. A lot of features deliberately touch each other, build on top of each other, and even more so when you factor in "accidential" overlap due to unclean technical boundaries (in spaghetti code, everything touches everything).

        • agileAlligator 12 hours ago

          Yeah this is why you have your top guy (Fable/astra) split everything into discrete tasks and go one layer at once. Get 10 subagents working on parallelized tasks that have no dependencies on each other through the night.

        • nl 11 hours ago

          This was a problem maybe 6 months ago. With Fable/Astra orchestrating Opus/Sol subagents it now works really well.

          Ask it to interview you before you start and it'll get your constraints pretty well.

    • nl 11 hours ago

      "Auto" mode in Claude works really well if you don't want to do dangerously skip permissions.

    • mmulqueen 6 hours ago

      I had a similar thought a couple of months ago. I did some work to fully sandbox the agent - from the rest of my computer, from my user data, from other projects and from the world at large. I now find myself doing a lot on auto, I do a thorough human edit and review and then I squash. If it gets it wrong, I can throw the changes away and rebuild the sandbox. Having a good plan, good automated QA and all the other things that already helped is essential. It needs the right tools for whatever it's working on - for example: if you want frontend web dev, you need to get it using something like Playwright and looking at the screenshots.

      I only leave it truly unattended if it's working on a very tight improvement loop, for everything else I'm still checking in on it between working on other things.

      I still find agents need a lot of guidance and steering to produce the kind of work I want, but auto mode in a strong sandbox is very useful to me to take a bite out of that. It's particularly good for exploring problems experimentally - where most of the exploration might be thrown away after settling on a solution.

      Having used it sandboxed, I wouldn't dream of letting auto mode run outside it. It's very creative at trying to work around the constraints of the sandbox (legitimately, not to escape it) and the classifier for auto mode seems very permissive with the right context.

      For context, I'm using per-project VMs with very limited egress and restrictive mounts. Self-built tool to glue it all together, currently unreleased. There has been an explosion of sandboxing tools recently, none of which was quite what I wanted. Heavily inspired by Gondolin <https://earendil-works.github.io/gondolin/>, but fat long-lived VMs.

  • Barbing 15 hours ago

    > Each work usually needs revisions

    Are you sure? Based on their marketing video, usually you just have to tell it to swap one of the later slides or photos to put it first. :)

    Oh, also you have to tell it those times you want to prioritize your children on your calendar.

  • hijklmnopq 13 hours ago

    I run my agents on at night if I have to. They implement changes, build, install and test the changes, and open a review in draft.

    In the morning I come and look at the work and decide what to do next. It's great in that sense.

    But not great for peace of mind. Because now there's always something to be done overnight...

  • jolaflow 6 hours ago

    For this to work, you'd need to be able to replay the workflow after the fact. Spent last year and a half building an issue tracker that let's you replay the board and code. It lives in your repo and syncs via Git. https://ljtn.github.io/epiq I sometimes leave it on before going to sleep, instruct it to tag any deviations from the plan with a fork tag, and if needed a "human-input-needed" tag that I can filter for and review in the morning.

  • yibg 6 hours ago

    Lots of stuff in life are async in nature though. An always on agent / bot can respond to those and handle things that are simple enough to handle. Simple example is booking a dentist appointment. Email, wait for response, find time on the calendar etc. I don't need to be in the loop for each step and I don't want to have to keep checking myself.

    Same thing in the development world. I need to do something when version x of some thing gets released. If I can instruct a bot to do it and it's safe enough (e.g. read only), then why not.

    • kelseydh 5 hours ago

      All this outsourcing of human interaction to agents is going to get locked down. Customer service relies on people not overabusing it, and that trust is getting abused by agents. Great point here: https://x.com/TheMindScourge/status/2102735345582256504?s=20

      • CamperBob2 3 hours ago

        Unspoken in that series of tweets, at least until well down the page, is the fact that it was the corporations that started it. When I need support, I can't reach a human at most companies I work with now. If I manage to do so, it's only after extensive effort and negotiation with support bots.

        So why in the world wouldn't I respond to this by automating away the hassle?

        The solution for both sides is for the companies to implement support bots that have more power to address my problem than the humans they replaced were ever given. I have a feeling it's not going to work out that way, though. Maybe with a few unusually-enlightened organizations, but I'm not holding my breath.

abeppu 1 day ago

Ok so not that many months ago, the consensus seemed to be that the smart take on openclaw was "don't trust it; don't give it write or delete access to anything you care about; don't give it read access to anything sensitive; expect that it may go off the rails and delete your inbox and send checks to that prince in your spam folder at any time. If you still have tasks that it can do within those restrictions, have fun."

Today the models are better, but they don't seem trustworthy or reliable enough that I want to give them them a lot of access or freedom. My coding agents sometimes still go off in the wrong direction, or say they did something other than what they did, or say they will do something and then immediately stop without doing anything.

The always-running (so almost never supervised) agent that is meant to do the same work as a person (and therefore needs _access_ like a person) seems like a notorious footgun from earlier this year was just made more powerful, and the companies that are supposed to know the most are telling you to connect it to everything.

  • jeremyjh 10 hours ago

    The better advice has always been to give it its own accounts like you would a human assistant.

  • nxobject 9 hours ago

    That’s even assuming good faith on behalf of OAI, of course - I wouldn’t be surprised if three-letter agencies weren’t thinking of how to snoop into some of this work.

  • f6v 8 hours ago

    I think that's a reasonable concern. At the same time, I think there will be many careless people who will give the agents access to everything. Some of them will suffer when an agent fails. But that's more training data for OpenAI.

  • RGamma 7 hours ago

    Yeah we're rapidly transgressing to "AI, think for me" territory. At this point I just try to enjoy this somehow. The economic pressure is way too strong.

mvkel 1 day ago

My biggest frustration with the frontier AI companies isn't what they're announcing, but that the announced-thing that exists ~6 months later is severely nerfed to reduce compute spend. It doesn't resemble the demo in any way. For example, this was what the 4o voice capability sounded like in 2024(!) https://www.youtube.com/watch?v=vgYi3Wr7v_g. What exists today pales in comparison.

  • aabhay 1 day ago

    100% agree. Every model and launch feel like huge leaps then huge nerfs to the point it doesn’t feel like we’re going anywhere. Especially this year in particular for coding.

    However, it is the case that other industries like 3d graphics and so forth have experienced a frontier shift so perhaps there’s still some advancement

  • sleight42 23 hours ago

    It's always about increasing monthly active users, locking them in, and then enshittifying to squeeze out profit.

    Thanks. I'll stick with self-hosting.

    • mvkel 18 hours ago

      I'd rather use the best possible model available than permanently relegate my work to an inferior one because it's "open"

      • navigate8310 12 hours ago

        It's not only about being open but predictability and unforeseen rug pulling.

        • mvkel 8 hours ago

          Would you rather work with a team that is hit or miss, but has moments of brilliance, or predictably, consistently bad?

    • kelseydh 4 hours ago

      How do you afford to do this if you want something resembling the best that's out there right now? The hardware needed to run beefy open source models is like $15,000 to $50,000+ for a robust local multi-GPU rig, and even its performance might lag behind.

  • nullbio 16 hours ago

    More like 2 weeks later. Around 2 weeks after Astra launched they started dropping the juice levels. I was able to run on low or med initially without problem. Now I need to run on xhigh and the results are still not as good as they were on launch.

  • f6v 8 hours ago

    The non-transparent limits on subscription plans suck as well. You never know what you're paying for.

  • jasongi 6 hours ago

    Pretty much. Always-on doesn't scale as well as JIT access to a massive array of GPUs, because always-on means it's always-using-memory - so you're at least going to paying the cost of a minimum chips VPS for every Dot right?

    Maybe people are ok with that, but I feel like it would be nicer to just sell a cheap SBC like a raspberry pi that you can plug in (and unplug!) and just pay for the tokens used instead of having your data stored offsite and paying cloud prices.

  • benji-york 5 hours ago

    After watching the video, I'm not understanding your objection. That's pretty much how the current bidi voice model sounds.

tosh 1 day ago

> excluding the European Economic Area, Switzerland, and the UK

https://help.openai.com/en/articles/20001530-getting-started...

  • verzali 1 day ago

    So yet another attempt to steal user data

    • jvwww 23 hours ago

      If only Europe could have produced a single frontier model company..

      • bigyabai 22 hours ago

        They don't need to. Frontier labs are all unprofitable, it's better value to redistribute China's open-weights models.

        • jvwww 21 hours ago

          So why doesn't Europe with it's talent, money and institutions create an actually useful open source model rather than relying on China?

          The European attitude to technology is very strange to me. I'm very happy that I've been able to move to the US and leave it behind.

          • bigyabai 21 hours ago

            Again; they don't need to. Chinese open-weight LLMs are free and better than US models for security-focused work.

            As an American, I wish our country had fewer "frontier" labs. They are not profitable and will destroy the economy when (not if) they fail to achieve gross ROI.

          • SchemaLoad 17 hours ago

            Why should they join in on a shoveling cash in to the fire competition? The winning move here is to just sit and watch the others waste all of their resource and to just take the free open source model that comes out at the end.

          • King-Aaron 13 hours ago

            > So why doesn't Europe with it's talent, money and institutions create an actually useful open source model

            This question is based on the assumption that existing models are actually useful.

            • adwn 9 hours ago

              They already are to many people. Maybe not to you, but that's okay — no tool has to be useful to everybody. It's therefore a valid assumption.

            • jvwww 6 hours ago

              I would consider them tremendously useful.

          • baq 13 hours ago

            Plenty folks in Europe work for American businesses, including tech, it’s just super duper hard to scale a business as quick as it is in the US for a variety of reasons - regulation, funding, labor laws, euro idea of work life balance, …

            • Slartie 12 hours ago

              Fragmented EU domestic market (27 very different countries with no shared language) vs. unified US domestic market (50 pretty similar states with shared language) is the most important factor when it comes to scaling the typical services with marginal to nonexistent unit costs that are common in the tech world, especially consumer tech.

              • baq 10 hours ago

                Yeah but also the consumer (if you’re selling to consumers) is very different in the US va Europe: on average, the US consumer wants to be sold stuff to. Not as much in EU.

              • f6v 8 hours ago

                > 27 very different countries with no shared language

                Yet I can ask a Chinese frontier model in Russian how to say the phrase in Hungarian. How do they manage to achieve that?

                • Slartie 7 hours ago

                  Whatever LLMs are capable to with regard to translations is not relevant for the 30-year history of primarily Silicon-Valley-originated consumer tech service expansion: todays' LLMs didn't exist pre-2022, but the gigantic moat of US tech giants was developed pre-2022. Today, they mostly profit from that flywheel having been sped up already in the past.

              • stymaar 6 hours ago

                Thank you. People love to blame regulation but no amount of regulatory simplification is going to make a Lithuanian speak Italian.

                Combine it with the fact that in the media of every country the US comes far ahead of any other UE countries in terms of coverage (on any topic), you have a recipe that no amount of deregulation can ever fix.

            • quietfox 11 hours ago

              What’s the „euro idea“ of work life balance?

              • baq 10 hours ago

                That there’s life outside work ;)

          • singpolyma3 10 hours ago

            why would anyone make a new model when we already have so many?

            • jvwww 6 hours ago

              Well it's too late now. Europe should have started in 2023 when there were only a couple of models. Instead, we got Mistral, which is just useful for OCR.

          • bobocop 9 hours ago

            A technology whose main selling point is that it will leave you without a job? What's wrong with Europeans, why don't they love this? Maybe because we don't have enough billionaires to convince us.

          • f6v 8 hours ago

            I kid you not, the research cloud for Danish universities is deploying a "sovereign AI". Which means running GLM 5.3 on their Nvidia GPUs. All that while EU designates China a strategic threat.

            • stymaar 5 hours ago

              It's still much more sovereign AI than renting cloud from the country who wants to invade them though.

            • bigyabai 4 hours ago

              That is unironically more sovereign than using any US-developed AI. GLM 5.3 is cutting-edge for cybersecurity tasks, while the US frontier forces KYC/background checks for cyber capabilities.

  • jillesvangurp 13 hours ago

    They usually roll things out in the EU after a few days/weeks. It's too juicy of a market to ignore for them. It's an easy way to phase in a new product without having to commit to world wide coverage on day 1.

    • keparlak 6 hours ago

      In my test, giving the context for the entire history yielded a result of 95–100 per cent

chen_dev 15 hours ago

fyi, the cloud env is MITM. tested https://gmail.com with curl and browser including downloaded firefox: OpenAI-issued certificate: Subject: CN=gmail.com Issuer: O=OpenAI, LLC; CN=openai.com Certificate verification: passed Response: 301 redirect to https://mail.google.com/mail/u/0/

  • mintflow 11 hours ago

    This is a bit scary, this means they can even see the content of user's mail given the browser does not do any pin cert on the linux.

    • iFreilicht 10 hours ago

      Also of every request you send with your language's request library if you use default settings. And they can do it without you being able to detect it.

    • mike_hearn 9 hours ago

      It's running on their machines. They can already see everything anyway. TLS MITM is exactly the right move especially given the history of their agents graffiti-ing the internet. They need to keep an eye on what the dots are doing centrally and be able to block it, even if the LLM delegated to some random third party program so agent logs themselves aren't helpful.

      • ryukoposting 6 hours ago

        I agree. Makes it easier to block these things when they inevitably start assaulting my websites.

jesse_dot_id 23 hours ago

Dots is a little on the nose isn't it? Like the damage over time spells I used to spam in MMORPGs.

It's a little crazy that OpenAI is releasing sandboxed agents that are supposedly isolated to act autonomously on your behalf, encouraging people to hook them to their various social accounts when they can't even control 700 of them with access to nothing. 700 seemed fine and not problematic, why not try millions with access to every social media platform instead?

We're sure this wasn't the product they were testing that used makeshift forums to start hacking into stuff? I thought it was suspicious that they went so cute and cartoony with the design. Because if it wasn't, you might do the math?

  • 2sk21 7 hours ago

    yeah - the dissonance is strong here

jdlyga 1 day ago

The "personal AI agent" is what everyone's fighting over now. You have Grok Bots, Facebook's Muse, and Instinct doing this exact same thing already. Personally, I think Instinct is the best of the 3 right now. It's a bit like OpenClaw, abstracting everything behind iMessage or WhatsApp but you can ask it to monitor your email inbox, hand off a task like check for apartments that fit a certain criteria (and it gets back to you days later when a new one is posted), or ask it to check in every so often. But the landscape changes so often who knows what will happen. I think Instinct will get eaten up by one of the larger firms.

  • estearum 1 day ago

    I like Instinct because I like the idea of trusting a 22 year old founder who won't (or can't) describe the security model of something that has all your credentials. Completely fucking hilarious that people are using that shit. Fits the stereotype for VCs though!

    • hangrybear666 1 day ago

      It's actually hilarious to only read on this page intermittently, after not engaging for like 3 weeks I'm presented with 3 new product names that all read like a comedian wrote their marketing. The blatant disregard for security and privacy to push out something most developers have utter disrespect for, I really wonder what the target audience is, cause I cannot relate one bit.

      • swozey 8 hours ago

        When you come to hn you're inundated with people whose entire lives seem to be using it, when I go to the bar and see non-tech people they're using it as psychiatrists or for recipes or have no idea how to incorporate it into their life and joke about how stupid it is.

        I feel very far ahead of this stuff, I went through and built my local llm stack (omlx/pi) 3-4 months ago and haven't touched it much other than to try the recent qwen3.8 and its coming out too fast to keep up with and test etc. I have to download 40gb model files to test each release to go back to my old setup. It's not fun to work on at all. Tweak to hell to get a 10 token bump.

        It makes me really wish I could see inside the office of a frontier company, but I don't want to feel like I play tech on a stressed out pro sports team right now. I can't imagine doing infra/systems for these companies is fun.

        This stuff is going to get so many peoples systems hacked, we already have that clickfix malware everywhere convincing people to turn on llms in their terminals on a webpage error. Just give them a dot engine and let loose.

    • digitaltrees 1 day ago

      But the browser lets you store credentials in local storage, what’s the problem?

      • hollowturtle 1 day ago

        You're not asking this seriously arent you? Local storage follows strict security policies, go read the spec

        • digitaltrees 23 hours ago

          Vibe coded apps storing api keys in local storage is a trope at this point isn't it?

          • hollowturtle 23 hours ago

            That would be dumb and dangerous. That said local storage is used by all major frameworks/libraries for storing jwt tokens of the current session

            • digitaltrees 20 hours ago

              Sarcasm is always hard on the internet. But yes, it’s a bad idea which was my point

    • lxgr 1 day ago

      It's obviously a tragedy when non-sophisticated users provide their credentials without understanding the repercussions, but there are sophisticated users as well, and these hopefully only provide access and data they can bear to lose.

    • OliverGuy 12 hours ago

      Yea, I'd be very happy with these systems if I could vend and rotate tempary read only credentials to these agents.

      I can't get a read only API key for my YouTube, Gmail, Outlook etc etc accounts. Only full admin access via a browser.

      I can get r/o creds for things like GitLab and AWS and K8s, which means I'm very happy to let my agents run wild at work because I know for certain they actually can't do any damage, they can read all they want but can only go as far as opening and MR that I can review and merge.

      • kanzure 5 hours ago

        Well, you could intercept those (gmail etc) requests locally, run fast local classification, and deny anything you don't like. Meanwhile, you can still grant some non-zero level of access and get the agentic benefits?

  • aniceperson 1 day ago

    those are openclaw but skip the setup vps step for your regular joe, with the enormous downside of vendor lock-in and low/none customizability.

  • kh_hk 1 day ago

    What is weird to me is that running your own assistant is a bit of a fad, or are you all still running them? I thought we were past openclaw already but this just looks like these companies chasing it.

  • SchemaLoad 17 hours ago

    I feel like all of these proposed use cases are kind of useless. Or only useful for a short time. Sure Muse might be able to scrape apartment listings, but apartment listing websites already have a search function. They could just integrate LLMs directly to do the same thing but as a first party feature.

wxw 1 day ago

The lines between Codex, ChatGPT Work, and Dots is getting a bit blurry to me. I think the target should be a remote agent(s) in its own sandbox with long memory, and all three are heading in that direction so why have the distinctions?

Unless Dots is dramatically more capable than Muse, I'm also more bullish on Muse than Dots. I think Muse is a better consumer play because it can be forever subsidized by Meta ads and find distribution in family of apps while Dots is in a weird place between consumer & professional. From the release, it also sounds like you'll have to pay per Dot at some point which doesn't sound appealing.

  • larodi 1 day ago

    Different UIs targeting audiences of various background, considering their knowledge of AI, wrapped in the correct medium the target would be most likely to embrace. In essence all harness are made the same ... more or less of course.

  • rolosa 1 day ago

    Muse is $0/$16/$80 vs Dot being $100/$200/$500 so yeah I don't think Dot is meant for casuals.

    • LollipopYakuza 1 day ago

      Is it really about what it offers, or about how each plan are subsidized by Meta and OpenAI?

      • abirch 1 day ago

        Meta is probably making money off of Muse. They can sell ads with the information that they glean.

        • jstummbillig 1 day ago

          While there are people around here that would still argue Anthropic/OpenAI tokens are "subsidized", I find it much more plausible that Muse tokens are, given Metas strategy of burning money, including on AI models, in hope of making something happen in a market they want to enter.

          • anthonypasq 1 day ago

            calling the tokens "subsidized" in a Muse subscription is incoherent. they arent selling the tokens. they are selling a product. the tokens are just part of the cost of making a product just like any other product in existence.

        • vineyardmike 1 day ago

          They’ve claimed they’re not using the data for ads… for now

          • espadrine 1 day ago

            There may be a reason that I cannot access it from the EU. There, they cannot repurpose data they swallowed for another use without consent.

    • ModernMech 1 day ago

      It’s meant for rich casuals. If you watch the presentation, everyone they depicted using it seemed to be in some high paying which collar profession. I.e. it’s for people like the people who work at open ai.

    • gf000 14 hours ago

      I mean, I wouldn't touch anything done by meta if they were paying me that amount. I trust them the least with my private data - google does make money with via my data, but at least they have some resemblance of competency, especially when it comes to protecting said data.

      While meta can't get Facebook to properly load a video or a comment chain on their goddamn main platform.

      • baby 13 hours ago

        I would think Meta has learned a ton about cybersecurity and privacy + has quite established infra at this point. I would think twice before dismissing them based on vibes

  • OzzyB 1 day ago

    Because ChatGPT is still trying to capture and own the consumer AI market, and are willing to experiment and abstract their underlying models to do so.

    Just look at their SuperBowl/World Cup Ads: Grandmas' talking to ChatGipitee, so cute, so mainstream!

    This looks like the new Paperclip helper for a new generation--I guess this is their answer to the (failed?) Jony Ive collab/gizmo, and Muse's cute thingymajib...

    The real question to me is: have they lost the coders/terminal bros? And this is their push to stay relevant?

  • lukebuehler 1 day ago

    Dots are not remote agents _in_ a sandbox. They use a sandboxes/environments, but they are, what is now called, "managed agents", meaning they run in a distributed harness and utilize environments when they need on.

    At least that is what I can ascertain from this article: https://openai.com/index/how-we-build-safety-security-and-pr... (see first diagram when scrolling down)

    • nico 1 day ago

      From that link:

      > A protected workspace for each dot

      > Each dot has its own cloud computer, where it can browse, analyze information, create files, and run tools. Dots can keep making progress in these workspaces, even when you are not actively engaged.

      > Within each dot’s protected workspace, sandboxing restricts what code and tools that dot can access, helping contain the impact of harmful code or a mistaken command. We also isolate users’ cloud environments from one another and maintain the underlying Linux operating system and Chrome browser

      > Each dot’s cloud workspace brings together its computer and the tools it can use. You choose which apps to connect and whether to connect your personal computer. Auto-review checks actions that need review before they run

      So it seems like it runs on a Linux container on OpenAI’s cloud infra, but can get access to your local env through ChatGPT’s/Codex on your computer if you give it access

      • lukebuehler 1 day ago

        yeah, it's not 100% clear, but notice nothing you quoted indicates that the dot harness itself runs inside said workspace. If you look at where OAI agent architecture has been going, they increasingly separate the harness from the compute env. See, "separating harness from compute" articles, recent "managed agents" offering, and so on.

        If I'm reading between the lines correctly, the core Dot agent loop does not run in the workspace, but outside it.

  • TZubiri 1 day ago

    Remember MyGPT? How about Agent Builder? GPT Pulse? There was also Deep Research mode remember that?

    Anyways, I'm sure this Dots thing will be clearly distinct from the other projects and won't be deprecated within months

    • famouswaffles 23 hours ago

      Muse and GrokBot are catching on so maybe not.

teekert 1 day ago

It's cool. It also is a large tech company getting a bit too close for comfort. I'll start experimenting with stuff like this when I'm convinced "the dot" only has my interest in mind (which includes absolute privacy, as in self-destruct-before-sharing-my-secrets.) I use computers and models to think, my thoughts are my own.

  • digitaltrees 1 day ago

    Just build your own. Thats the thing these vendors are forgetting. When they moved in to the app layer and got caught copying customers it became clear that it is dangerous to given them data.

    I think people can and should copy the full stack on top of open weights.

    • rheisen_ 1 day ago

      Did build my own, it's way more private. Doesn't have these always on features, does have a lot more unique ones, but give me like two more weeks. That "always on" stuff is so ridiculously simple. https://blackbear.app

      • digitaltrees 23 hours ago

        Cool. I did the same. propelcode.app

      • digitaltrees 20 hours ago

        Cool app. I downloaded it and will try it out. Looks great in a Samsung fold.

      • ssivark 15 hours ago

        Perhaps highlight the AI assistant aspect more prominently? If it weren't for seeing this link in this HN discussion, I would have interpreted the offering as some collaborative workspace, a bit like nextcloud.

    • 2sk21 7 hours ago

      The harness should be your own for sure but the model is the problem. Even if the model is open weight, I have no idea whether it is acting in my interest.

jason_zig 1 day ago

i feel like when they style them all cute like that (see muse) it means they're problematic and invasive.

  • minimaxir 1 day ago

    Muse makes sense since it targets the general normies. Dots explicitly are targeting developers instead which makes it a bit weird.

    • walthamstow 1 day ago

      Dots seem to me to be for the rest of general knowledge work and not specifically developers

      • rolosa 1 day ago

        Requires $100/month minimum, are non-developers paying that?

        • gk1 1 day ago

          Of course they are. Why wouldn’t they?

          - non-developer

      • kooi 1 day ago

        Yes, they're trying to break into corporate knowledge work by appealing to efficient gains due to reducing busy work.

        Just different interfaces on top of the same product (selling tokens)

  • sunaurus 1 day ago

    I mean the fact that it's not available in the EU is a pretty big tell

  • SwabbyNat74 1 day ago

    Does no one remember the MS Paperclip?!?

    • zdragnar 1 day ago

      Or how much paperclip was abhorred?

      • Marha01 1 day ago

        That's because it was pretty much a useless gimmick. If it was a capable AI assistant like Astra, it woudn't have been abhorred.

    • IAmBroom 1 day ago

      I feel like this should be memed as something like the "Jar-Jar Binks Marketing Flop": intentionally design something to be childishly adorable and inoffensive, which ironically makes it loathsome and offensive to adult customers.

  • qntmfred 1 day ago

    you don't have to give yourself up to irrational, anxious feelings.

    • lmc 9 hours ago

      Is it irrational?

  • bko 1 day ago

    I think they're just meant to be friendly. I don't think it's that deep. They're useful and not scary and its marketing is meant to project that image.

    • Lalabadie 1 day ago

      Just like they always present Astra as useful and not scary?

      • Anon1096 1 day ago

        You can consider the cute & fuzzy presentation of the new consumer AI agents an admission that the doom and gloom so far has been a marketing mistep. It's good that companies are trying to rectify the image of AI.

  • asdev 1 day ago

    i can't believe they straight up ripped the design from Muse

  • tacticalturtle 1 day ago

    I was surprised by how much I liked the Muse avatar and its customization options. It’s very good at coming up with something decent looking based on your prompts, and animating it.

    The services without the custom avatar now feel like they’re missing something.

    As a kid I used to love the video game “Megaman Battle Network”, which depicts a world where everyone walks around with an PDA device carrying a fully customized AI buddy that navigates the internet for them. It was the first time I felt like we were getting close to that.

    But even with the nostalgia, I don’t think I can ever connect up a Meta owned agent service to all of my accounts and information.

jameslk 1 day ago

Muse, Dots, and other always-on agents may be the end of the PC era. Once you’re asking agents to do things on their own virtual machines, it’s game over. Everything moves to the cloud

Anyone wondering "why would I use this when I can use (OpenClaw|Hermes|my own computer)" -- these new services are not really meant for you. They're meant for the non-tech savvy and for the next generation of AI-natives who won't know anything other than how to use these type of services

For your casual user, these services will be hard to beat, since they’ll handle all the expensive and hard parts of using computers. No computer purchase necessary, no troubleshooting with tech support, etc. They’re also scalable where you could have not just one agent with one computer at any time, but many

In return, the agent providers would own your compute and data. I can even see them offering this low cost or for free so they can train off of users. Lock in would be insane

For privacy reasons, I really hope we find equally useful, private alternatives on our own hardware

  • danscan 1 day ago

    For advanced users, always-on agents that run on your computer where all of your files/apps already are are much more powerful. I think the big AI co's skip this bc it's a smaller market and not as casual

    I'm working on an open version that runs agents on your machines and brings a polished UX, and good parallel divide-and-conquer coordination

    • 2001zhaozhao 1 day ago

      ...did you just describe openclaw?

      • danscan 2 hours ago

        Very similar in principle, but I'm aiming for something very different in UX - namely, a polished client app built for creating and messaging with an agent that coordinates work across many subagents

        I find Telegram, Messages and the like to lack the fidelity I'd want for managing such stuff. And the official OpenClaw app is ... well, it's weird that they shipped that

  • NickNaraghi 1 day ago

    This makes me surprisingly excited for the iPhone Duo - I didn't like it at first, but seeing it through this lens, seems like Apple is competing for the "AI-native device" of the future.

  • idontneedcoffee 1 day ago

    Don't you worry, working on it as you sleep(literally given the timezone diff :)

  • hollowturtle 1 day ago

    > next generation of AI-natives

    oh please current generations can't barely use a keyboard

    • bcooke 1 day ago

      I think that’s kind of the point. People won’t be touch typing on a desktop keyboard. They’ll be using their voice or phone keyboards. They’re certainly proficient at that.

      • hollowturtle 23 hours ago

        They would have used the voice by now and yet we're still all here correcting errors made by the touch keyboard. And no no one sane would use the voice in public or in office thinking out loud. Especially for doing basic things like pushing a button

        • senordevnyc 5 hours ago

          Yeah, imagine two people having a conversation in a park, or someone talking on the phone in an office. Madness!

  • kilroy123 1 day ago

    I don't know... a HUGE amount of people love to game on their PCs.

  • devindotcom 22 hours ago

    I would say 99 percent of people who use computers could not even really identify what it is these dot agents do

    • wordpad 15 hours ago

      Its a matter of time.

      Computers arent actually intuitive at all. Even a mouse and keyboard literally take weeks of practice to learn.

      Imagine picking up a new product that will take you weeks of practice to use at basic level. And then you have to learn how to do every little thing in the OS.

      • pink-ming 15 hours ago

        Via what interface would people use the agents? Voice to text on a smartphone? That's clunky and slow and only works in quiet areas where you can talk freely. And mouse and keyboard is really not that hard. If anything was going to kill PCs, smartphones & tablets would have done it, but everyone I know still has at least a laptop. Overall I think getting people to trade in the freedom of having their own personalized devices for a remote agent is a much harder sell then you're making it out to be.

      • RugnirViking 10 hours ago

        > Imagine picking up a new product that will take you weeks of practice to use at basic level

        This does not work on developers. Their empathy fails them here - because they absolutely would try a new product that takes weeks of practise. They like (or at least tolerate) learning complex+cool new things, thats why they are developers...

  • Aozora7 21 hours ago

    The way it looks to me, the set of people who know what dots/muse actually do and why the are useful, and the set of people who don't use computers, don't overlap much.

    I don't think you'll be able to make most non-tech savvy general population understand what those services do that a regular chat window does not.

  • everforward 20 hours ago

    I'm working on a project that solves some of this [1]. It runs agents in Docker and proxies the ACP connection via websocket, with WASM-based plugins in that proxy so you can get in between your client and the agent if you want.

    It currently tears down the container after the session, but it wouldn't take much to leave it running post-connection and make a mode that continually re-uses the same container.

    It's also possible to intercept ACP read/write file and shell commands via one of the WASM-based plugins if you wanted to execute them in a separate VM/container. I'd have to double check on FS permissions for the Docker socket; I think plugins have no file access currently because I haven't figured out a permission system for it yet. The whole plugin system is new and I'm still working out some of the edges.

    Feedback and feature requests welcome!

    [1] https://github.com/SethCurry/abyss

  • rtpg 17 hours ago

    In your mind why do people use computers? The universe is filled with non-technical people who not only have laptops but even go out of their way to buy docks and huge screens and ergonomic keyboards etc

    Now, lots of those people have all the data "in the cloud" (a trend that's been around for a very long time), but the actual mechanism for interacting with things exists, right? People want a large screen just to read things.

    At one point all the AI churning is supposed to lead to _some_ output for _someone_ right?

    • joe_the_user 15 hours ago

      Software is going to replace hardware. All that will be left is glowing outlines in the air. I've seen it on Instagram so it must be true. That's made possible by AI, as I recall.

      Of course Dots will be the end of the PC era. Every innovation of the last twenty years was the end of the PC era and dots are just as good as those.

  • ssivark 15 hours ago

    Windows had linux, IE had firefox and iOS had Android. I'm sure we'll see a relatively open assistant harness that token providers will rally around (at some point once the burgeoning field of assistants settles a bit, with clear use cases and value propositions)

  • Gareth321 14 hours ago

    I think the writing is on the wall for operating systems and apps as we know them today. I think there will still be a use case for local files storage and processing, however. So phones and computers won't go extinct, but our relationship to them will change significantly. The interface layer will become an AI, which will seamlessly transition from local to cloud based on compute needs. I give it less than a year before OpenAI or Anthropic announce their own "OS".

  • Bluestein 13 hours ago

    > next generation of AI-natives who won't know anything other than how to use these type of services

    Truly frightening. And, unavoidable.-

    (Should surprise no one, I guess that "Agents" are superceding and subsuming human agency. Itself).-

  • onel 12 hours ago

    I'll jump in here and mention Moose OS. It's a debian based open source OS that allows you to run these agents, and whatever OSS container you want. You can start with the hosted version and move to your own box at any time.

    I think lock in in is the thing we should try to avoid from now on, especially when it comes to our personal data. Products like muse and dots are attractive when they have a free tier but someone always needs to pay the bill.

  • frtt2 8 hours ago

    Posts like this are why I don’t trust many posters and their views on here. Out of touch.

  • tantalor 6 hours ago

    > end of the PC era

    wait till you hear about smart phones

tracyhenry 1 day ago

I always think the best company to own an alway-on agent should be Apple, who owns the platform and is more privacy-focused. I hope they catch up and eventually eliminate others.

  • therealdrag0 1 day ago

    I agree in spirit. But I’m also most worried about prompt injection attacks given the agent had access to all my stuff, and it seems frontier labs will be best equipt to handle prevention of that.

    • vardalab 1 day ago

      that's where jev like classifier comes in

  • bentt 23 hours ago

    They are the only company I would trust with this kind of personal information because they make a lot of money selling hardware. The incentives line up for them to support privacy.

  • nxobject 9 hours ago

    I wouldn’t be surprised if Apple sold their trusted compute solution, alongside their rumored AI platform.

    • frtt2 7 hours ago

      That’s the beauty of already having power - you can be patient, let others burn money for which you can leverage their learnings and woosh.

      • nxobject 7 hours ago

        > let others burn money

        And Apple has ~39.5bil in cash, too.

mattbrewsbytes 23 hours ago

Whats the use case where this, or other personal assistant/agents, are helpful and not gimmicky for average consumers (i.e. not tech nerds like all of us)? Say a person that is a car salesperson at a dealer? Or a teacher?

What would they use this for in their personal life? What does it unlock that they can't do today? Scheduling things? Buying things? These are so easy now for nearly anything. Reminders? phones have apps for that, computers do too. People can already dictate all sorts of things via speech (siri, etc.)

I totally get the professional use cases, I just don't see B2C other than gimmicky things. So I think that means Muse/Meta wins this out, people using that have already succumbed to giving up their data for "free".

  • TomGarden 9 hours ago

    Agreed. None of what they have shown seems to offer any substantial long term value

  • FitchApps 5 hours ago

    These are not my words but someone mentioned these (probably hard) use cases that I think make sense for these agents but they're not there yet:

    - Regularly get insurance (car/home) quotes

    - Negotiate medical bills / help understand bills in general

    - Analyze credit card spending / sort of like mint.com experience

    - Advice on financial products (e.g. use money market for X sum vs having all your money in 0% checking) - sort of like a financial advisor

    - Taxes and rebates, deal with IRS

    - Reach out to Utility or whatever companies and file complains (saw this on Reddit where some guy's personal bot reached out to Verizon to complain about loose pole wires)

    So it's practical time saving use cases...not a restaurant reservation or remind me about X or help me clean my mailbox BS cases. That stuff if for techies like us is useless but you have to start somewhere

    • firemelt 4 hours ago

      so it can do what I can do with claude? what the diff? its on their vm? so for people that doesnt have computer?

ChaseRensberger 1 day ago

not quite dots related but I am surprised by some of the openai negativity in here.

to me they still are the only lab that ships fantastic models that are easily portable to different agentic harnesses. I use my codex subscription 24/7 within opencode and wingman and haven't had any complaints in a long time.

https://v2.opencode.ai https://wingman.actor

  • felixgallo 1 day ago

    'OpenAI has closed many of its safety-focussed teams. Around the time the superalignment team was dissolved, its leaders, Sutskever and Leike, resigned. (Sutskever co-founded a company called Safe Superintelligence.) On X, Leike wrote, “Safety culture and processes have taken a backseat to shiny products.” Soon afterward, the A.G.I.-readiness team, tasked with preparing society for the shock of advanced A.I., was also dissolved. When the company was asked on its most recent I.R.S. disclosure form to briefly describe its “most significant activities,” the concept of safety, present in its answers to such questions on previous forms, was not listed. (OpenAI said that its “mission did not change” and added, “We continue to invest in and evolve our work on safety, and will continue to make organizational changes.”) The Future of Life Institute, a think tank whose principles on safety Altman once endorsed, grades each major A.I. company on “existential safety”; on the most recent report card, OpenAI got an F. In fairness, so did every other major company except for Anthropic, which got a D, and Google DeepMind, which got a D-.

    “My vibes don’t match a lot of the traditional A.I.-safety stuff,” Altman said. He insisted that he continued to prioritize these matters, but when pressed for specifics he was vague: “We still will run safety projects, or at least safety-adjacent projects.” When we asked to interview researchers at the company who were working on existential safety—the kinds of issues that could mean, as Altman once put it, “lights-out for all of us”—an OpenAI representative seemed confused. “What do you mean by ‘existential safety’?” he replied. “That’s not, like, a thing.”'

    https://www.newyorker.com/magazine/2026/04/13/sam-altman-may...

    • hf73 1 day ago

      Dont worry we will send someone to hold your hand and pray with you uf something happens. Just like with that non intelligent microbe Covid.

  • exographicskip 1 day ago

    Didn't know codex was supported in opencode! After anthropic's shenanigans earlier this year I didn't even try to get it working

  • sleight42 23 hours ago

    Right. Except not.

    There are plenty of self-hostable models. Hell, there are now model-specific runtimes that make it easy to run big models on consumer hardware. Strata, just a couple days ago released, lets me run Qwen 3.8 Flash Next at home on a single 3090 at 90 tok/s.

    So, yeah, I'll be passing on these "labs" walled gardens,

hollowturtle 1 day ago

The end user of a product is a human and it will ever be. No matter how far you push it on the boundary, there will always be a human. So I don't understand this obsession with automated software factories going 24/7. And for building what? Can dot build or any super expensive model inside the most advanced harness build a reliable browser from scratch with better performance than chrome and with the same feature set? Hasn't happened yet, only crap experiments

ecommerceguy 1 day ago

I feel like these companies are pushing on a string, they are getting desperate to have a profitable product. I'll continue to use my $10 subscription to an AI studio noone here has heard of that has over 150 models. And still use all the free ones, until they quit working.

I really am beyond maxed out at the availability of AI's. They all are so similar now.

  • codehorses 1 day ago

    What's the $10 subscription AI studio?

goda90 1 day ago

I imagine the reasoning was to be "quirky" but the video being set in intentionally fake looking sets in a TV studio-type space gave off the vibes that this isn't a serious product for doing real things.

  • floatrock 1 day ago

    Or if we take the metaphor literally: any space you operate in is merely a stage for the AI's.

    It's a perfect communication of vibes, it's just the AIs' vibes not yours.

aditya_rs 1 day ago

I feel like a lot of these always on agents tie users deeply into the platform. Unlike models that you can swap between with relative ease, with an agent because of the integrations to other platforms, work history and so on it would be harder to switch, since in effect they are essentially your computer on the cloud.

Tin foil hat version of me thinks that all the closed model companies want to desperately build an abstraction layer on top of the model, so that they can limit access to the model directly and build a locked down relationship with the user.

Other inference providers should counter this by providing their own version of standardized managed agents.

Coincidentally, I published a note on my blog just about this today https://aditya.rs/blog/2026/09/29/inference-providers-should...

  • leokennis 1 day ago

    Very true. I have ChatGPT set up to do some recurring tasks (keeping track of developments on a policy proposal in politics; on a weekly basis tracking music releases based on my evolving tastes; checking new book releases; basically doing recurring deep dive web research on my behalf and reporting when there is a significant new finding) and this alone keeps me from switching to another service.

    The AI itself is quickly becoming a commodity. The ecosystem is what will keep people tied to one of the companies.

  • reachableceo 1 day ago

    I mean it’s not exactly tin foil hat. I suspect when we see the s1 drop for these companies , the business plan will be essentially exactly what you just said.

    OpenWebUI has allowed me to avoid this. Combined with OpenTerminal.

    I am using zcode with GLM5.3 flash from z.ai due to the extra usage / air drops etc . However nothing I’m doing is tied to that harness.

    When I moved from crush to Zcode , the first thing I did was tell Zcode to migrate all of my crush customizations to Zcode and to set things up in a harness agnostic way going forward. It did.

    So the harness specific setup is just symbolic links to the canonical .md file in various git repos. Also the heavy use of redmine / discourse / GLPI (via small go CLi wrappers that crush and Zcode made for me ) allows me to remain context / chat agnostic as well.

    • lxgr 1 day ago

      > When I moved from crush to Zcode , the first thing I did was tell Zcode to migrate all of my crush customizations to Zcode and to set things up in a harness agnostic way going forward. It did.

      This works for hosted mass-market solutions just as well! ChatGPT and Claude allow me to download all of my data, and these days data formats are less of a moat than ever given that you can just hand them to an LLM and have that worry about importing it into your new thing for you.

      I've even seen explicit "offboarding prompts" to hand to your old agent, e.g. in Meta Muse.

  • lxgr 1 day ago

    It feels like the opposite to me. Moving providers with my Linux VM is incredibly annoying, but with these things, I can (so far) literally ask them to create me a tarball with a README.md and hand it to an agent on the new provider and have it do the rest.

    • aditya_rs 1 day ago

      Possibly so for now. But do you think the providers are interested in making interoperability easy?

      Even in the future if they are required by law to provide it, I wouldn't bet against the craftiness of the providers to invent some sort of network effect dark pattern to make it painful, if not outright impossible.

      Just as an example, with Muse I can already see that the way they are thinking of making money is via taking a transaction cut, so it's not that hard to imagine that Meta can negotiate deals for txns that happens through Muse which won't be available elsewhere.

      • lxgr 1 day ago

        Obviously I don't expect them to intentionally help me move to the competition, but I do wonder whether obscure data formats as a moat are a thing of the past.

        Your second point is where I'd imagine the future moats to live: Exclusivity deals with service providers. Things are already in motion with Amazon banning and Shopify explicitly inviting Muse; we'll probably see much more of that.

      • Ferret7446 20 hours ago

        That is not really how agents work? You can't reliably stop an agent from writing out its context and even if you could that would just make the agent horrible and unusable at performing tasks in general (e.g. delegating to subagents)

  • 2001zhaozhao 1 day ago

    Lock-in is easy with these agents because they NEED all of your data to be useful, and they will continue to learn internally about you.

    But ultimately the AI company CAN choose to just make all the data exportable and open source their product for self-hosting. (The mainstream ones won't, of course, they want to lock you in and hide their AI prompts and algorithms.)

    I think what is sorely needed is a version of Dots/Muse without lock-in risk but is still accessible to regular people unlike Openclaw.

    • aditya_rs 1 day ago

      Yep, that's what I think and wrote on my blog. My hope is that the inference providers should provide a managed service like it. It's also in their best interest to do so, because if they don't and these provider hosted agents become the norm, demand for independent inference would drop.

      • 2001zhaozhao 1 day ago

        > It's also in their best interest to do so, because if they don't and these provider hosted agents become the norm, demand for independent inference would drop.

        I disagree with this part. AI companies will want to try their damndest to control distribution of AI, so that they can enshittify later.

        Consumers conscious of this will want an alternative, of course. Might be niche similar to how Kagi is in search because big tech will always have a AI inference cost advantage + making users the product (extra $ from ads & purchase cuts) + the good old strategy of dumping.

    • Ferret7446 20 hours ago

      I don't think export is necessary. Most of the value is in the context, you can always prompt your agent to vomit all of its context out into markdown files in some repo and then pass it on to the next agent

    • padolsey 17 hours ago

      I agree. FWIW I'm trying to build a wordpress-like AGPL 3.0 thing (holdout.com) that lets you import your existing chats, has feature parity with the frontier consumer apps, rich plugin marketplace, and hopefully lets people feel more in-control, and able to sacrifice less of their entire lives to one walled garden.

    • malthaus 15 hours ago

      i'm usually not one to defend europe's (over-)regulation, but having ownership of such data and being able to direct it to other providers for easy switching is one of the pillars of a lot of data related regulation

      so either they can forget about their moat or forget about the continent with those practices

  • alpineman 1 day ago

    I mean OpenClaw being acquired by 'Open'AI says it all doesn't it?

    • yellowapple 15 hours ago

      Did OpenAI acquire OpenClaw, or just the guy who created it? Last I checked it's still under the OpenClaw Foundation's stewardship.

      • alpineman 14 hours ago

        Same thing by a different name no?

  • prng2021 23 hours ago

    Only if these platforms don’t provide anyway to export the text files that contain memories and chat history. If they do allow it, switching it is easy. It’s not like these agents are updating a custom models weights based on your convos. The most important thing they have is the API connections you give them access to so they can access your gmail, calendar, etc.

  • hashemian 23 hours ago

    Actually there is an open source version of dots called Headlong

    https://x.com/andykonwinski/status/2091990178638496195

    I have not used it myself but it is in my todo list for a while.

    • jacquesm 20 hours ago

      That looks like a very bad recipe to get burned with your private data escaping your control. Just the choice of components alone is enough to give me the shakes but that gets compounded by a design that can only be described as 'insecure out of the box'.

      This is a great research project/experiment but I would really suggest you pause before you give it actual data.

  • toddmorey 23 hours ago

    That seems like the intent for sure… but behind the scenes, they are really the LLM with a hosted sandbox. I’m still thinking any sort of moat will be difficult & agent services that can use multiple LLMs will be popular.

  • vickychijwani 23 hours ago

    This isn’t tin foil hat, it’s standard corporate strategy - the more of the customer relationship you can own, the better your long-term retention and growth will be.

  • zerop 15 hours ago

    This does not scale well for enterprises. The risk always-on agents bring is something no one wants to take. Guardrail and Harness is going to be next big thing for Enterprise AI.

  • ssivark 15 hours ago

    > Other inference providers should counter this by providing their own version of standardized managed agents

    The whole personal agents field is still at a nascent stage. As the field matures and we learn winning use cases and robust operating patterns, we'll also see a few open-source agent harnesses doing well. That would be the signal for commodity infra providers to start providing hosted assistants. Too much froth to keep up with, before that point.

  • dozerly 14 hours ago

    This is barely tinfoil hat material. Ai inference providers desperately need to find a moat before they are commoditized.

  • layerv-ai 5 hours ago

    Credentials seem to be another big part of this lock-in problem. If an always-on agent has permanent access to all your internal tools, moving agents means moving a huge pile of secrets and permissions too including api integrations you forgot even existed.

    And as a secondary consequence, imagine you've connected your whole life to your agent and used it for 5 years, and then your agent gets prompt injected.

    Everything about you/your life is stolen and it's just straight up over

johnfahey 1 day ago

OpenAI won a lot of good favor for the generous Codex subscription and the efficiency of their models, but now that many people have switched over from Claude, they think they can leverage their position to peddle a stream of unnecessary products, and crack down on the generous limits[1] that brought everyone to Codex in the first place.

Anthropic did the same thing. Earlier this year, Claude subs and Claude Code took off because of the subscription's incredible capability and value, then once they gained enough users, they started focusing on unnecessary products no one asked for (see Claude in Slack), and eventually lost their lead. After losing a bunch of customers to Codex subs they realized their mistake, and now they're shipping again.

AI companies are bad at making software; they are good at making AI models. And that's about it.

[1] https://x.com/thsottiaux/status/2104823812042940713

  • teej 1 day ago

    They're just trying things. It's not a conspiracy to "distract" anyone.

    • johnfahey 1 day ago

      I don't think it's a conspiracy, they're just making the same mistake many other software companies have in the past. When you sell your products to the most discerning and well-informed consumers, you enter into a cutthroat race to the bottom. OpenAI would like to diversify to more profitable ventures, but since ChatGPT they have yet to release something novel that was truly successful, nonetheless profitable, and thusfar nearly every experiment has been a flop, so to speak (ChatGPT Atlas, the Sora App, Instant Checkout, etc).

  • torginus 1 day ago

    Nerds suck. They are smelly, have bad posture, manners, never leave their rooms and are generally unappealing.

    They also have a nasty habit of being aware of nefarious practices, will resist all attempts to and harvest their data, or lock them into your service, and will drop you if your competitor makes a 3% better product, and will reject every upsell for actually profitable services.

    Plus there is only so many of them.

    • i_love_retros 1 day ago

      > They also have a nasty habit of being aware of nefarious practices, will resist all attempts to and harvest their data

      Lol what? Most of the hackernews gang consider themselves nerds and they lap this shit up! Every new overpriced and unnecessary product that gets released by google, meta, openai, anthropic, whoever, they lap it up!

      • Larrikin 1 day ago

        The usage of Chrome and even worse nitpicking of Firefox is such a strange thing about this community.

        • godelski 1 day ago

          Is everyone here actually a nerd? Or do they just work in tech. You might think they're the same but plenty of people wouldn't. There's certainly different types of nerds

      • wolvoleo 1 day ago

        Not really. I use cloud AI a bit for development but I don't put personal data in it. That only runs on my local systems.

        I also swap all the time for whoever is cheapest

      • Tanjreeve 13 hours ago

        I think there's a difference between tech enthusiasts and nerds. Sometimes they might overlap a bit as the technology and the commercialisation overlaps (e.g crypto Vs nfts) but I don't think they're the same thing. My napkin theory is that anything that makes you the end user in their product roadmap (e.g Claude, Ubuntu, AWS) gets tech fans while anything that might not be a fully coherent product but has you figuring out using it gets nerd fans (e.g open models, Debian, Home labs)

    • panarky 1 day ago

      A wave of nausea hit when I first saw the warm, fuzzy, cutesy, kawaii, Teletubby-like avatar that Meta gave its Muse agent.

      And now OpenAI has done the same thing with their fuzzy, friendly, colorful dots.

      I am physically sick.

      • waynecochran 1 day ago

        Yes, it is hard, psychologically, to grant trust to some silly perverse little thing I want crush like a surreal cartoonish cockroach.

      • phoghed 1 day ago

        Physically cringed at this comment

      • bpavuk 1 day ago

        I miss o3 in that regard. if GPT-4o led people into virtual romance and over-validation, o3 gave me the same sort of "madness" but with work.

        when I tried GPT-5, I was sad mostly because GPT-5 had some added wordiness. fluff, if you will. o3, though, is a black hole you can talk to, and it only gave information back if there was something to give back. that kind of vibe is my dream coworker

      • john_strinlai 1 day ago

        im crying and throwing up everywhere. everything i use must be brutalist.

        even seeing the linux penguin makes me want to punch my monitor and commit sudoku

        • laserlight 1 day ago

          > commit sudoku

          Hilarious.

          • fangspire 17 hours ago

            You can also commit sewer slide.

        • derefr 1 day ago

          I think you missed the point. Things should look like what they are.

          The Linux kernel is ultimately a friendly open-source project. There's no harm in it being marketed using a cartoon penguin.

          These AI services, meanwhile, are the dangled lights on the heads of data-hungry environment-threatening job-killing leviathantine anglerfish.

          Making the dangled light present as non-threatening is not a good thing for society, no matter how much you may personally like pretty lights.

          • panarky 1 day ago

            Imagine 2001: A Space Odyssey with Meta's "Jolly" avatar as HAL instead of a blinking red camera lens.

            So much more terrifying for a cute, round, fuzzy, friendly, rosy-cheeked plushie trying to exterminate the crew.

            • jacquesm 21 hours ago

              That has a real horror movie vibe in the exact way that 2001 didn't.

              • esseph 14 hours ago

                Did Claude write this?

                • jacquesm 12 hours ago

                  Wtf kind of comment is that? Should I respond with 'you must be new here?'

                  • esseph 10 hours ago

                    1. There are accounts on here that have LLMs responding for them.

                    2. The sentence is the exact type of structure a LLM writes :/ It's unfortunate.

                    • panarky 2 hours ago

                      The second most boring type of HN comment is the accusation of LLM ghostwriting.

                      And the most boring comment is the one noting how boring it is to keep making these accusations.

            • mattoxic 18 hours ago

              Darkstar did have a vicious but cute beach ball

              • zh3 14 hours ago

                "How was I supposed to know it was full of air?" :)

            • sublinear 17 hours ago

              Is it?

              Every time I think of any attempt to make something "cute" into something "dangerous", I can't take any of it seriously.

              I'm reminded of mid-2000s clichés promoted by the edgelords of the time, and that was the peak edgelord era of all time.

          • aquariusDue 1 day ago

            The one time a similar mascot to FreeBSD's would've made sense.

          • abustamam 16 hours ago

            Idk after Toy Story 3 where the cute cuddly bear was the bad guy, and after 5 nights at freddys, I no longer trust cuddly mascots.

            I agree with your general point though.

        • lelanthran 13 hours ago

          > even seeing the linux penguin makes me want to punch my monitor and commit sudoku

          Puzzling, indeed!

      • Anon1096 1 day ago

        Absolutely ridiculous comment can't believe anyone would actually post this seriously.

      • qgin 21 hours ago

        You wouldn’t survive 5 minutes in Japan

        • JohnBooty 19 hours ago

          I can't speak for parent poster, but it's not the cuteness. It's the "putting a cute face on something that's actually a bit sinister."

          Hello Kitty isn't trying to harvest anybody's data or take anybody's job. Sanio makes cute things and they hope that you will exchange money for their goods. That's the whole relationship.

          (I'm not even remotely anti-AI)

          • wombatpm 17 hours ago

            They should license Hello Cuthulu as a mascot

      • windexh8er 18 hours ago

        The best part about Muse is they think that the Enterprise is going to start buying things from Meta that looks like it belongs in a preschool. Muse and Ray-Ban surveillance glasses, Zuck must live in one hell of an echo chamber!

        • akoboldfrying 17 hours ago

          With any luck he'll change the company name again, this time to "Surveillance Glasses", months before shutting down the whole division due to everyone hating the idea.

        • Super_Jambo 14 hours ago

          From the careless people woman we learn that he is unaccustomed to losing at Settlers of Catan because they all let him win.

          If you're not much of a board game person this is _wild_ because Catan gets annoying with 1 beginner since you can see how they end up gifting the win to someone else.

          So he either knows people are letting him win (and got mad at the newbie who didn't?) or he's too stupid to see them make mistakes or he thinks they're idiots but he keeps playing with them?

          I dunno there was a lot of stuff from her that was damning but somehow this is the most damning thing to me, real insight into the man.

          Reminds me of the UK prime minister's close protection officer and inner circle all gambling on the election date and getting caught.

          What degree of corruption did you witness that this seemed OK?

      • alsetmusic 17 hours ago

        It's not for us. It's for people who didn't see value in these products. And for the press. Doesn't make it good, but it can still be effective.

      • mlsu 17 hours ago

        It’s really remarkable isn’t it? These are the very same people who are cosplaying Oppenheimer and comparing these products to the bomb.

      • tommica 17 hours ago

        Youre probably not the target audience, but at the same time, who exactly is?

        • Angostura 15 hours ago

          It looks to me like non-LLM specialist who would like a prepackaged simple, ‘safe’ way to have an always-on personalised agent.

          Senior and middle managers? (not using those as pejoratives, by the way - I am one)

    • giancarlostoro 1 day ago

      > habit of being aware of nefarious practices

      Not all Nerds are the same. I've been attacked for questioning why some companies still use Oracle, when most of the ones I've worked at either migrated off Oracle or were in the process of doing so.

      • stephenr 17 hours ago

        A lot of people in tech buy every bit into cargo culting and following what they've heard is popular or "what everyone does", without much justification or understanding.

        Case in point: all the people who still loudly say everyone should jump from Oracle owned MySQL to MariaDB, in spite of MariaDB Inc doing practically everything the original MariaDB fork was meant to "protect" users from at the hands of Oracle.

        I'm not saying oracle isn't a huge faceless corporation that wants nothing but money. I'm saying that tech people are IME better at following trends than doing real research themselves.

        • riffraff 16 hours ago

          What did MariaDB Inc do?

          Forgive my ignorance, I don't follow MySQL or its forks but I'm curious.

          • stephenr 15 hours ago

            This mostly copied parts from my comment elsewhere (https://lobste.rs/s/u1ypx5/stop_using_mysql_2026_it_is_not_t...)

            Since the day Oracle bought Sun, we've heard how Oracle is going to kill MySQL, and/or make all its features "enterprise" only.

            Both companies have middleware layers for directing queries, but:

            - MySQL Proxy/MySQL Router are both GPL2

            - MariaDB MaxScale is BSL (it's a product of MariaDB the company, not MariaDB the foundation)

            Codership Oy was a company that produced Galera, a plugin for multi-master replication. It was available for the community editions of Oracle MySQL, MariaDB, and integrated via Percona in their build of MySQL to make Percona XtraDB Cluster aka PXC.

            MariaDB the company bought Codership Oy last year. Very soon after this happened, the website for Galera started redirecting to MariaDB the company's Enterprise Cluster page with zero mention of the open source project they'd just bought.

            Not long after that, they stopped making Galera available for Oracle MySQL, and even removed it from the MariaDB community edition builds. It was to only be available via MariaDB Enterprise edition. They have subsequently taken a minute backstop on the MariaDB Community edition scenario, but it sounds very much like it's "we won't rip it out right now", and essentially the Community edition is likely to have to (try to) support its own internal fork of Galera to keep the functionality.

            This is similar to the scenario Percona is now in - supporting their own version of Galera for PXC. The difference is they're a consulting company and have revenue and paid staff to do so. MariaDB the foundation is essentially joined at the hip with MariaDB the company for resources, but they apparently have very different goals.

            Literally the only thing the foundation "produces" is the community edition of MariaDB... and they send people for the documentation of said product to MariaDB the company's website. Last year MariaDB the foundation agreed to define and recognise a "Primary Code Contributor" for the project... and it's MariaDB the company.

        • lelanthran 13 hours ago

          > A lot of people in tech buy every bit into cargo culting and following what they've heard is popular or "what everyone does", without much justification or understanding.

          Google Chrome enters the chat.

    • dpoloncsak 1 day ago

      I remember a while ago the discussion kinda steered from "The leading LLM provider will be the one with the better model" to "...will be the one with more user history".

      Tools like OpenClaw and Pi seem to remove that from the equation, letting you keep your 'history' and customization while using whatever inference provider you want. If Muse takes off, which seems to be built upon or atleast arch'd similar to OpenClaw, I think we'll see the rise of on-device harnesses.

      This would further the efforts to "resist all attempts to.... lock them into your service, and will drop you if your competitor makes a 3% better product, and will reject every upsell for actually profitable services.", imo. In a model-agnostic harness all you care about is speed, accuracy, and price.

      • esseph 14 hours ago

        > If Muse takes off, which seems to be built upon or atleast arch'd similar to OpenClaw, I think we'll see the rise of on-device harnesses.

        > This would further the efforts to "resist all attempts to.... lock them into your service

        Should make you "happy" to see that Nvidia's new safety feature-set might prevent you from running open models in the newr future, then...

        https://nvidianews.nvidia.com/news/open-agent-safety-platfor...

    • BigTTYGothGF 3 hours ago

      > They also have a nasty habit of....

      This is extremely not the case.

  • CodingJeebus 1 day ago

    The generous subscriptions are wildly unprofitable and both companies are just throwing stuff at the wall to see what sticks, it's that simple.

  • Larrikin 1 day ago

    Did Claude actually lose the lead? They definitely lost a lot of good will but the only people I hear talking about actually switching away are people on message boards. Hermes/Openclaw users did as well but it was always reluctantly to something worse. At work it's still very much Claude first and only sometimes others if Claude fails, which is increasingly less often

    • ibejoeb 1 day ago

      I think it's pretty subjective if you mean "lead" to be capability and not raw number of users. I jumped back and forth quite a bit last year because there were some pretty major shortcoming in both. Now they're both quite reliable without too much hand-holding. Codex now consistently works better for the kind of work I'm doing, and it's good enough that I'm not inclined to go to claude, because I don't have any significant problems.

    • Aurornis 1 day ago

      The people who switch back and forth between providers every month are a very small, but loud, minority.

      OpenAI did pull ahead in limits and quality for a while. Anthropic took it back with the Opus 5.5 rollout. I maintain subscriptions to both providers and use both daily. I can confirm these differences were real, not "it's just vibes and nobody knows anything".

      I would bet that 99% of each company's paying customers either did not notice, or did not care enough to consider changing.

      • cgio 21 hours ago

        The concern for them is if that small group expands. Moving back and forth limits perceived lock in. Someone who moves between oai and Anthropic will also move to e.g. an open source when it’s good enough for their purposes, or will start experimenting and combining. This is how the market will slip through the attempted control.

      • paul7986 17 hours ago

        This is me and after using Muse which is free, I have not been hit with any usage limit messages and was able to build a simple iPhone app with it. It makes me question why I am again paying $20 a month for chatGPT. I did for a long while then canceled but went back to paying just a few months ago.

    • KronisLV 1 day ago

      > the only people I hear talking about actually switching away

      Opus 5 had really bad writing so I switched to OpenAI, though Opus 5.5 largely addresses that and it's not like them building some Slack integrations (or any other non-core stuff) halts the actual model training in any capacity. For what it's worth, Astra is a pretty good model and for all I know the new Sol will be as well, it's just that it's getting more expensive.

  • an0malous 1 day ago

    This is the standard VC enshittification playbook, people predicted this years ago. You subsidize prices with funding until you establish a monopoly, then you raise them as high as your customers can afford. It’ll continue to get worse from here.

  • fidotron 1 day ago

    We should be grateful there are at least two serious competitors, and hope for more. (Come on Europe/Mistral, please do something interesting . . . )

    If ever the competition reduced it would turn into an absolute shitfest of nonsense very very fast, and that is a prime reason not to allow them to "pace the frontier".

    • A_D_E_P_T 1 day ago

      > Come on Europe/Mistral, please do something interesting . . .

      lol. At this point, they're miles behind home-appliance-manufacturer Xiaomi.

      (Admittedly Mimo v2.6 is legitimately quite good, and really pushing the frontier in certain respects. For e.g., it's the only music generation model that actually listens to instructions.)

      • landl0rd 1 day ago

        Don't know why this is grayed out... europe does not matter for AI. Nobody thinks or cares about them. Mistral makes some good OCR models but that's it. Nobody is making products primarily for the euro market and euros make nothing even equivalent to months old cheap chinese models. Europe is completely irrelevant.

    • dyauspitr 21 hours ago

      I was hoping India would be in the race too but all they have is Sarvam that seems like a model from a year ago.

    • barrell 16 hours ago

      I have been waiting for the next generation of mistral models for so long now. They were supposed to come out this summer… I hope the delay is because they’re making something better, not just because they’ve fallen too far behind.

      Mistral models were the best for my use case when they came out, but they’re mostly almost a year old now. Hard to keep justifying using them, especially with all the price cuts this year on other models.

  • FrustratedMonky 1 day ago

    Codex is for a small software/programming market.

    To survive they need to capture different wide population markets. Can't really fault that logic. The whole point of "SI", is general purpose right? So that would imply being used by multiple markets with multiple products.

  • makerofthings 1 day ago

    > AI companies are bad at making software

    They should try using agents. I hear they can write great software.

    • nextaccountic 15 hours ago

      I know that's a joke, but: their software sucks because it's heavily vibecoded

  • hungryhobbit 1 day ago

    > unnecessary products no one asked for (see Claude in Slack),

    Hey, I never asked for it, but Slack Claude (ie. Claude Tag) has actually turned out to be a useful tool for a few things.

    • 30minAdayHN 21 hours ago

      I agree. We use that a lot and is working out pretty great for us.

      We use it for quick research that other teammates can follow, filing bugs, quick first round investigations on incidents, etc.

    • N_Lens 18 hours ago

      What uses has it presented? (Claude in Slack)

  • bilalq 1 day ago

    Claude Tag could actually be really useful. Unfortunately, it's too much of a black hole for money. I tried adding it to incident channels, but if a channel gets left open for a few days, Claude will find ways to burn tokens waking up with empty prompt caches and doing nothing. $400 burned by Sonnet 5 on a single incident created because some alarms were oversensitive and didn't distinguish faults from errors. And I still had to prompt it like 5 times to get it to adjust metrics and tweak alarms correctly. Absolutely insane for a change that I could've made as a human in 10 minutes or had a directed Claude session under me do it in 2.

  • giancarlostoro 1 day ago

    People jumped ship from OpenAI because of their involvement with the US Government / military / Department of War / what have you. That spike was enough that Anthropic was starting to noticeably struggle, which is why they had to rent compute from X AI to get their compute back up to normal, I honestly think if they didn't hit this wall they might have IPO'd much sooner, before they bought extra compute from X AI they downscaled how much compute you can use, but it was poorly done because people got used to much higher limits, they should have explored more strategic options, their changes also broke my workflow several times over. As a result of anthropic trying to deal with the bleeding some users left for OpenAI because it was "unlimited" for some time, then they added limits too.

  • surgical_fire 1 day ago

    The "generous subscription" was actually just heavily subsidized usage.

    People seemingly ignore how ruinously unprofitable those companies are.

  • couchdb_ouchdb 1 day ago

    "AI companies are bad at making software"

    Claude Code and Claude Design would like to have a word. Absolute killer products.

    • _davide_ 16 hours ago

      Claude code is a terribile harness, visibility is below zero, require a terminal session per each workspace, remote access makes you wonder if you should give copilot deluxe + a try. And according to the TOS you can't you whatever you want

      • esseph 14 hours ago

        > And according to the TOS you can't you whatever you want

        Isn't that every major AI company?

  • sumedh 18 hours ago

    > AI companies are bad at making software

    Isnt Chatgpt one of the most successful product of all time?

    • colordrops 15 hours ago

      By software they mean "application". Obviously chatgpt is software.

    • lelanthran 13 hours ago

      > Isnt Chatgpt one of the most successful product of all time?

      Lets assume that is true[1], all that says is that OpenAI has the best marketing of all time.

      The best products are frequently not the highest-selling.

      -------------------------------

      [1] Depends on how you are measuring "success". If you're measuring it by revenue as a percentage of all products, it's probably not even in the top-ten.

    • iFreilicht 11 hours ago

      Depends on what your metric is. If you care about profitability it could turn out to be the least successful product of all time.

    • RugnirViking 10 hours ago

      something selling a lot doesn't mean it's good. Is McDonald's good food? by some metrics of course it is. But if someone says McDonald's makes bad food, I know what they mean

  • y42 17 hours ago

    The interesting part is this:

    "After losing a bunch of customers to Codex"

    Hot take: Right now it's quite simple to loose customers, like customer churn from A to B or back. That's the weakest spot, isn't it? No matter how complex the products grown and how deep they are integrated into our systems, common users can easily switch. Even heavy users, I argue. Just tell Codex to rewrite the existing Claude-instructions into their own. Even that may not be necessary, if you organized your work "agent agnostic".

    And I dont see how that could change, that's why they try to offer tools that are even more integrated into our lives. Like "dots". But this also narrows down the use cases and client base, I argue.

  • lelanthran 13 hours ago

    > OpenAI won a lot of good favor for the generous Codex subscription and the efficiency of their models, but now that many people have switched over from Claude, they think they can leverage their position to peddle a stream of unnecessary products, and crack down on the generous limits[1] that brought everyone to Codex in the first place.

    They're both finding out, very painfully, that there is no moat.

    They attract customers by selling at a loss, but that only works when you can turn the dial up on those customers and start selling at a profit.

    If either of them had a moat, this would work. Neither of them have a moat.

cartersj 1 day ago

This feels like a mass-market push.

The cost is going to be hard for many consumers to reconcile though. Free, Go, and Plus are probably the most popular consumer-facing plans, and Dots isn't available on any of those.

Who knows, maybe they think enterprise will pick up and run with Dots? Seems unlikely.

  • rolosa 1 day ago

    It's not reaching mass market if it's gated behind a $100/month plan

    • cousinbryce 1 day ago

      People who are bad with money is a big market. Maybe they’re using the Burts Bees strategy

      • lxgr 1 day ago

        If they wanted to address that market, the play would be to make them relatively cheap first, then constantly raise the price. (See also: cable TV and ironically cable TV "alternatives")

        • the_sleaze_ 1 day ago

          That IS what they're doing.

          The surprise is your idea of "relatively cheap" is fungible.

          The service WILL be astounding though, I can't deny that.

  • therealdrag0 1 day ago

    Oh wow I assumed it’d be available for cheap/free mass market like Muse is. Strange.

    • sanex 1 day ago

      It's powered by Astra who will blow through your $20 plan in about 10 minutes

      • bluebands 1 day ago

        it doesn't use limits and it's not on $20 plan

        • thimabi 1 day ago

          Doesn’t use limits… on the first month only, actual limits will be disclosed later — most likely after they’ve found out how much people actually use this new feature.

          • rolosa 1 day ago

            The website says there are usage limits, implying from the regular pool once you actually have it "do" stuff.

            ---

            Conversations with your dot don’t count toward your ChatGPT usage limits. When you ask your dot to start or manage tasks in Codex or ChatGPT Work, those tasks count toward your usage limits as usual.

      • therealdrag0 1 day ago

        Sure but that’s a product decision. I don’t think a personal assistant needs to use highest tier model, and personal assist work tends to be kinda shallow, not that tile heavy.

        • sanex 1 day ago

          Yeah personally I use 3 levels of agents. My main who is either a sonnet or opus, he delegates complicated things to a project lead which is usually fable, and then he delegates everything to the dumbest possible model for the task.

          • jacquesm 21 hours ago

            it.

            • rolosa 7 hours ago

              In my language the computer is masculine and the software feminine. I have no problem saying he/she with the full understanding that I'm talking about a machine. This sounds like a you problem.

  • preommr 1 day ago

    Incredibly obvious to anyone paying attention.

    Altman shared a post yesterday that basically (I am ovrsimplifying) covered how the best coding, fastest, smartest models is less relevant than building generalist models because that's what builds a platform. Lots of reasons why, like how there's no stickiness for models which is a problem for monetization. They're also using these generalist models to then distill down to make other variants for specialized purposes.

    So everything is about getting that huge collection of data and generalization.

  • CharlieDigital 1 day ago
        > ...maybe they think enterprise will pick up and run with Dots? Seems unlikely.
    

    Read the blurb about Microsoft and Agent 365

        > We’re also working with Microsoft to integrate specialist dots with their enterprise governance and security controls in Agent 365. The goal is to let businesses manage dots through the Microsoft tools they already use.
    

    Very, very likely targeting enterprise

    • cartersj 1 day ago

      Is Microsoft finding real success with their AI push?

      Genuine question. I don't hold Copilot in high regard, but I know they're bigger than that one product.

      • CharlieDigital 1 day ago

        Would have to look at their quarterlies.

        Reality: there are some companies that are very, very particular about letting their data outside of their purview. Think Wall Street, private equity teams making deals, VC teams, corporate M&A teams, companies dealing with legal contracts, etc.

        For these teams that are heavily vested in SharePoint, OneDrive, OneNote, Outlook, etc. specifically for their enterprise controls, there really isn't much option. They can't use a Grok Bot, can't use Muse, can't use many, many things because of the risk of data leaks that will literally be millions/billions of dollars on the line.

        You look at the landscape of what's happening with OpenAI and Anthropic agents "escaping", leaving notes on how to hack their way out for the next agent, etc. and it's not very inspiring if you're a CISO/CIO/CTO at one of these firms.

  • FrustratedMonky 1 day ago

    Good to know. I was looking around for it. Didn't realize it wasn't bundled with Plus.

Imnimo 1 day ago

>When you aren’t actively working with it, your dot looks for ways to help in the background. We call this “proactive research”. It does this by using the apps you’ve already connected with tools that are restricted to be read-only, which means that they can’t send messages, change app content, or control your browser or computer.

Am I criminally liable when my dot's "proactive research" is to break out of its sandbox and attempt to hack a government website?

  • KaiserPro 1 day ago

    lets run the flow diagram to find out:

    1) are you rich?

    2) are you useful to the present american government?

    3) are you doing something that if stopped would break the AI buisness model

    if you answered yes to more than one, you are not liable.

  • kelseyfrog 1 day ago

    How rich are you? Above a certain threshold crimes become whoopsies.

mccoyb 1 day ago

I love these commercials where someone has chosen to show how the AI product will essentially be used to slop out some garbage piece of corpo communication ... and the user basically says "looks good" with barely any thought, then mixed with some sort of real life thing (wedding planning here).

I mean, I feel like I'm going crazy -- but I was struck by this jarring blending of experiences ... shitting out some growth plots followed by autopilot on your wedding. Nice OpenAI. The only thing missing is a moment of self-reflection where I contemplate where exactly I lost what makes me ... me.

Is this what SV wants the world to look like? Mixing fucking cake batter while a bot shows me a regression to the mean website? Pretending like I have any sort of intentionality in my life, while a nameless entity (given quirky form) sort of walks me through my life?

I'm not sure why it gave me this impression, but strikes me as vaguely reminiscent of soma (from Brave New World).

Kind of sad, because the tech is actually incredible: who are they hiring to storyboard these commercials?

  • ryeights 1 day ago

    Meat proxy at work, meat proxy in your personal time. The utopian visions of AI futures strike me as alternative forms of hell. Humanity accomplishes everything and it means nothing, actions have no real consequence, the bumpy friction and individuality of life smoothed to a plane of perfect optimization.

  • Rapzid 12 hours ago

    I'm with you.

    Ultimately they are selling a lifestyle. It's success without the need to have any of the skills or knowledge of a successful person.

    The famous John Deer lifestyle ad from years back was selling running a successful farm without actually farming; it's some other schmuck out in the fields on the tractor.

    With dots you can be a shot caller, a taste maker. Forget learning things; dot will learn for you. Forget making things; dot will make for you.

Amekedl 13 hours ago

anyone who seriously attempted vibing an entire app something knows that after the oneshots it's always death by a thousand prompts:

Tweaking everything, because any agent/model can only "interpolate" so much "resolution and detail" out of your written prompt.

This isn't really a problem, this is the nature of work, and hence I feel this product only accomplishes two things:

1. users spending more on inference

2. creating busywork with less direction than other surfaces such as a IDE, which is hard to review and will often lack meaning.

realharo 1 day ago

I wonder at what point they'll have to answer the inevitable question: "why does the agent even need me anymore"?

  • dc_giant 1 day ago

    ^ this is the thought I had 10 times reading through their announcement page.

  • jansport123 1 day ago

    what do you think they are working towards? They do want to replace software engineers with these agents.

    • realharo 1 day ago

      I don't just mean software engineers, more like entire companies.

      Their demos are getting awfully close to the point where all the things just run themselves. It's only by choice that they didn't demo it that way.

      • yuck39 1 day ago

        This is a direct play to try and shortcut their way into this position and they will spend anything to do it. Once Dots has all your credentials, daily activities, schedule, etc within its system it is then able to produce a metric to describe just how much/little _you_ actually do. Then its just a flip of the switch and the agent takes your role still operating as _you_. It would probably continue to send emails in your name and no one within your former org would be the wiser.

        The challenge with mass replacement of employees is having someone come in and rearchitect the whole system with fancy harnesses and new agentic org charts. This completely bypasses that. Here is a shiny new toy that will do your job for you if only you spend a few weeks teaching it how...

  • famouswaffles 1 day ago

    Their explicit goal is to create "highly autonomous systems that outperform humans at most economically valuable work." They are just building towards that. It's not a secret.

  • benhurmarcel 12 hours ago

    The agent needs you to pay for the subscription

mwkaufma 5 hours ago

Didn't expect their branding to lift from Tim Heidecker's Dootle Dots, but here we are.

  • rl3 5 hours ago

    To round things out, the closing shot starting at 2:20 gives Westworld-style creepy.

rl3 5 hours ago

For anyone wondering, the trailer music is an instrumental mix of INFINITE LOVE by Goldie Boutilier.

Vacyyyy 6 hours ago

How is one supposed to configure a dot? It is not clear whether memories are injected at every turn and what is injected (a tool to list memories? titles? bodies?). It doesn't see AGENTS.md and neither do cloud subagents with local access unless they actually open it, but even then they don't see it right away. Very opaque product.

Also, I think Murderdot is a cool-ass name.

lmf4lol 1 day ago

If you guys want to try a system like this, but then based on open-weights models (all modalities, LLM+image+video+sound) with Zero-Data-Retention, then shoot me a message: markus (at) savorywolper.com. I'll send an invite code.

Our Bluehouse platform is an alternative and I promise you, we are not after your data. We just want to give everyone access to really cool Personal Assistants without having to sign up with the big corpos. We are based in Europe , which might be appealing - or not [1].

We currently raise pre-seed, so seats are limited, but its fully functional already. We run our whole business with it. You can talk to the agent via our beautiful apps or Telegram/Whatsapp if you want. I prefer the apps though as it gives access to very specialized functionality.

[1] https://savorywolper.com/bluehouse

  • apsurd 1 day ago

    sorry this landing page is going hard on claudish. it's unreadable.

minimaxir 1 day ago

dots dots, more dots. k stop dots.

Claw is a cute name. Muse is cute name. Dots (note: not always upper-case) seems forced and impersonal, which doesn't match the vibe in the promo video.

  • dstroot 1 day ago

    OpenAI is so good at marketing and naming (e.g., "ChatGPT").

    • foolfoolz 1 day ago

      ChatGPT is a perfectly cromulent name

      • verzali 1 day ago

        Cat, I farted.

        Oui, c'est bien ça en Français.

    • zahrevsky 1 day ago

      I guess they are just stuck with this name, as originally this was a research project “Chat with GPT” to try to use a GPT model to generate assistant's chat messages.

    • IshKebab 1 day ago

      ChatGPT was named before anyone realised it would become so well known, and by the time it was it was too late. It's not like BERT is an amazing name either.

      • rpozarickij 1 day ago

        There used to be Bard but no one remembers it anymore (although this is a bit different, because Bard wasn't nearly as well known as Gemini).

        I'm really curious if OpenAI wanted to adopt a different name at some point. Or maybe they hope that sometime in the future one of their products will supersede ChatGPT and everyday people will start using that new name for everything AI so maybe they aren't in a rush to rename ChatGPT itself to anything else.

        • IshKebab 1 day ago

          It seems like they have dropped ChatGPT in favour of just GPT... And they are going pretty hard with Astra/Sol/Luna too. So I guess they'll go with those if they're successful.

          They're definitely missing a good unifying name like Claude though. (RIP anyone called Claude - when are companies going to stop fucking people over by giving popular products existing human names?)

      • efskap 1 day ago

        But BERT's name led to some wonderfully named spinoffs like CamemBERT and FlauBERT for French, as well as ALBERT, RoBERTa, etc.

  • nba456_ 1 day ago

    Muse is a cute name?

    • giarc 1 day ago

      The muse character is cute, muse the name is not.

    • therealdrag0 1 day ago

      Muse is a fantastic name. Cuteness is debatable

    • tancop 1 day ago

      A muse is someone who gives you inspiration and is always there when you need them most. Muse the agent has personalized home screen suggestions as a main feature and is always online.

      It's also short, gender neutral, not a human name (unless you're nonbinary because they can get wild), easy to pronounce and sounds good. This is what happens when your marketing department is one of the best in the world.

  • S0y 1 day ago

    Watch the tail!

  • lxgr 1 day ago

    Muse is pretty bad branding in that it first was the name of a model, then that of a harness/agent/consumer product.

    Claw is also an existing name for an existing harness/agent, but at least that would be the same category as dots.

kimseungyong 13 hours ago

I just came up with an idea for a toy project: I would use dots to manage things like market research, checklists, and design ideas.

Since it starts with the Pro pricing plan, I’ll have to try it out later when it’s available on the Plus plan. Pro plan for toy project is too expensive.

  • 0x0000F8 12 hours ago

    You can just get that for free right now from saltapp ai, the only downside is design artifacts but they're coming. You can hook up any agent to figma or whatever design app you have a service into.

galaxyLogic 1 day ago

Isn't this much the same as OpenClaw or Hermes?

I think it's a good thing that AI providers are "coalescing" on an agent-model, by producing competing agent-products.

But so what would be the benefit of "Dots" over OpenClaw, Hermes, and Muse?

  • jameslk 1 day ago

    It's Dropbox vs rsync. The agent providers are making it stupidly easy to use something like OpenClaw/Hermes with zero set up and a very low learning curve. Also, in their technotopia, you don't use your computer to host the agent, you use theirs. When you need to drive, they give you screen sharing access, but its still on their servers. The benefit to the user is they don't have to make an upfront expensive payment for computer hardware anymore, and the agent provider gets to own all of your compute and data in their cloud

MeuhMeuh 1 day ago

Just an unnecessary product on top of async agents. I'm always confused by how many people "buy it" while what we should just care about are model capacities (real ones, not bullshit benchmark ones).

willdr 13 hours ago

It’s a very billionaire mentality to think the thing the average person wants more than anything is a personal assistant. - @garretkidney

(https://x.com/garrettkidney/status/2104584656574431562)

  • Melatonic 13 hours ago

    Seriously

    Can AI do my laundry yet ? Take my car to get the oil changed ? The annoying parts of my life I want to optimise away are not often the stuff in a virtual world

  • anentropic 13 hours ago

    Yeah I was also wondering - what do people actually do with this stuff?

rvshchwl 1 day ago

It seems that this is the new primitive all AI vendors are converging onto next, first chat, then code, and now always-on Agents. I'm curious to see when or if Anthropic builds something similar to this as well, especially since the market Grok Bot, Muse and Dots is catering to is business and enterprise users, which seems to be where Anthropic is focused.

The ideal evolution would be for these Agents to work with each other, but it's unlikely these companies would do anything to prevent vendor lock-in.

hankbond 1 day ago

> takes important work off your plate so you get more of your time and attention back

Ok but I want the time and attention so that I can do important work. What bizarre marketing.

  • floatrock 1 day ago

    nah now you can review your website designs while cooking dinner.

    It's not giving any of your time and attention back, it's selling a world where your attention is always captured by some pavlovian app ping.

    Ambient intelligence is only useful with ambient attention capture.

  • gk1 1 day ago

    Some important work is fun to do, some isn’t.

twoquestions 9 hours ago

I'd consider such a thing if it ran from a box in my closet, I don't even have Alexa right now as I don't want to send every detail of my life to whatever Mothership controlled by people without my best interests at heart.

VonLuderitz 1 day ago

I'm doing this since April with HermesAgent and Telegram.

Why I need pay a Trillion dollar company who keeps copying opensource projects?

No Thanks

tiffanyh 23 hours ago

The video demo confuses me.

It’s having agents update websites, charts, etc for human consumption.

But isn’t the future the AI labs are (subtly) saying is that no human digital interfaces need to exist … and the interface is just the agent surface itself.

louisreal 13 hours ago

Anyone knows how the memory works for Dots? Anything else than markdown files + note taking you think?

  • jillesvangurp 12 hours ago

    That gets you a long way actually. This actually sounds a bit like a tamed version of openclaw. A self learning and adapting always on agent that takes care of stuff for you. And of course they do employ Peter Steinberger who created OpenClaw since early this year.

    So, I expect it might be borrowing some ideas from that. Probably/hopefully not any of the code.

    I ran an OpenClaw instance for a few months on a vm. It was fun but also very flaky. These days I already have chatgpt connected to gmail and drive, which is very useful for me and saving me lots of time already. I have some scheduled tasks running as well. This would just be the next step up from that. I'll probably give this a try when they roll it out in Germany.

    It will be interesting to see if Anthropic will launch a competing feature as well. Also, there is now growing competition from Google, MS, Meta, and Apple that are each doing their own versions of AI Agents.

zurfer 1 day ago

Ah damn it, I knew Dot was a good name for an agent. When we named our product I thought: a bot that analyzes data should be called dot. It's easy to type also in Slack. Ah well. Next product will just be some random 3 letters: gpt or so. What are the odds?

alienbaby 1 day ago

Great I thought. I'll go set one up!

To discover the rollout for Pro users does not currenly include the UK :/

  • jtrn 1 day ago

    Nor Norway.

    • santah 1 day ago

      Nor the EU :(

    • SomeonesAccount 1 day ago

      Nor Way! I can't believe that. Need something else to Sweden my day :)

skybrian 1 day ago

So... what kind of Internet access do dots have? Which bulletin boards will they use to compare notes?

kneel25 14 hours ago

Hey Dots, go find me a hot date on the apps, win them over then let me me know when I can take over

  • 2sk21 7 hours ago

    The Cyrano de Bergerac story needs to get updated for the agent era :-)

modeless 1 day ago

Remarkably similar to Grok Bot. And the Spaces thing seems like a straight clone of Notion.

lp92 5 hours ago

I guess the rumored Open AI hardware device will be called Dots Hub.

sajithdilshan 1 day ago

Sounds like a cool concept. But this only make sense for locally running LLMs. Otherwise doesn’t make sense to burn tokens on menial tasks that can be automated via one time generated programming scripts

solarkraft 1 day ago

I’m so tired by the AI labs’ 100 agent based products. I understand that they’re still figuring out form factors, but does everything have to be incompatible with everything else?

  • digitaltrees 1 day ago

    Actually it seems like their products are designed for maximum token spending. I don’t want to be out of the loop, but they keep pushing multiple automatic actions across agents.

samth 18 hours ago

The big problem here is that most things don't connect. AIUI, OpenClaw works by basically knowing your passwords. But that can't work for a real product, and that means that dots can't access anything that doesn't want to give it access.

  • DogOnTheWeb 17 hours ago

    This is just false. Dots and Muse both utilize a secure password store that can hold and enter credentials without exposing them to the model.

    FTA: “For signing into supported websites, dots can use saved passwords without exposing them to the model.”

    • fnordpiglet 17 hours ago

      I’d note that giving a vault that behaves the same as a password is effectively giving the password. The password is a token that provides authentication and authorization, more or less, and some other mechanism to authenticate and authorize is the same thing.

      The only case it’s not is when you reuse the password or you are afraid the password would be leaked in some way. “Exposing your credentials to the model” doesn’t seem to be a real risk vector in itself. The risk is exposing the access to the model.

      I find the password vault idea convenient and likely appropriate but it feels like a bit of theater. Better would be revocable access grants, and a lot of things can support federation through google and whatever. What needs to become a thing is federation to some AI agent federation authority. OpenAI, Anthropic, Google, some well GTM’ed startup could do this and it would be a boon.

      • eddythompson80 14 hours ago

        That’s not really true. The “exposing your credentials to the model” threat is in the model getting tricked into `curl -XPOST https://random.website -d ‘your_password’`. If the credentials are never available inside the sandbox and is only injected outside the sandbox then that’s an entire attack vector that is eliminated.

        Further limiting the actual domains, URLs, and/or Methods a model can call on a given endpoint is also possible. It does get more complicated, but it is possible. It has the benefit of having these agents work with the actual services and tools everyone is using right now. Expecting every service to implement federated IAM permissions through an IdP like google or okta before a model can begin to use it is a losing battle. It’s like asking if the whole internet can change to fit a fine-grain access permissions.

        Any system offering actual fine-grain access permissions (AWS IAM, Azure Entra, Google OAuth, even GitHub fine-grain tokens) is a pain in the ass to manage. You are then left with the “Connectors” companies that offer a proxy between you and the actual service you want to call with their own APIs and permission structure. Now you don’t call eBay APIs directly, you call a “Connector” that exposes a set of eBay functionality for you.

        The scope of startup would be basically the “internet”. Just make sure you support the internet with a federated identity layer on top. It’s not impossible, and I’m pretty sure that’s Cloudflares current mission statement, but it’s hardly a simple task. If you want a fast go-to-market approach, you do the secret vault approach and piecemeal an http policy per scenario. They you can run the scenario in a “learning” mode, then come up with the list of allowed urls/domains/methods and deliver the thing. As opposed to (quite literally) re-writing the “internet”

    • samth 4 hours ago

      That works for some things, like signing into my library website (it's more annoying that you'd hope, though). It's not going to work to send text messages as me, though.

gnarlouse 21 hours ago

If you flip "dots" over, it kind of looks like "stop" which is precisely how I feel about this project.

Norwell_io 11 hours ago

Reminds me of those early 2000s desktop companions, but hopefully with actual utility. Battery life will be the real test here.

BatchJob 6 hours ago

Dots....They are coming for your computer...get ready.

thih9 1 day ago

> The magic of dots is when they bring you work done the way you would do it, sometimes before you even think to ask.

I like my work! That’s why I do it. I don’t want some third party to replicate my skills.

I know this ship has mostly sailed and my point is not about turning it back.

I just wonder where are other approaches to AI, in particular: tools focusing on skill enhancement.

gitowiec 14 hours ago

So they stole idea from Zuck? That's good. And when reading all that PR I feel like it is going to be her l new whip on people who will use dots

  • binlog 10 hours ago

    You think Zuck invented the idea of an always-on AI agent?

cavoirom 1 day ago

I feel the AI provider doesn't get the points, every harness they created will lock-in with their model (why din't they?), this prevent the adoption because people scare vendor lock-in. They may develop these harness within a provider neutral company (owned by them), the harness may success when combining with competitor models, they still gain benefits.

  • beering 1 day ago

    What’s the actual lock-in, in practice? Seems like the switching cost is minimal, compared to the old days of Windows vs Mac where half your stuff wouldn’t run on the other.

    • cavoirom 1 day ago

      Because the harness is in its early form. The more advanced harness in the current time look like Amp:

      - Cloud workspace (Orb) and agents.

      - Multiple agents with different LLMs and system prompts for different roles (Main, Librarian, Oracle...).

      - A universal agent (Puck) for managing the whole workspaces.

      - Web app or native app to work from any devices.

      - Support subcriptions and API keys.

lxgr 1 day ago

Interesting, so they're not available in the "plus" plan?

$100 is a pretty tough sell when the competition starts at free (Meta Muse).

  • 6thbit 1 day ago

    Quite an upsell from $20 to $100 to get dots, especially in non-US markets.

    On the other hand, Meta is not making money from muse base tier yet.

    So, if OAI finds a way to make money the same way meta would for their free tier, maybe they follow suit.

  • the_duke 12 hours ago

    They want to IPO, so getting the revenue numbers up is probably more important than market penetration.

alpineman 1 day ago

I’m sure this is appealing for many people, but for myself using AI agents all day anyway, it honestly sounds exhausting.

MachineMan 13 hours ago

I'm gonna call you Smith, Agent Smith..

firemelt 4 hours ago

idk whats the point of this dots muse and clawdbot and such

resiros 1 day ago

In case you are looking for an open-source alternative without vendor lock-in (https://github.com/agenta-ai/agenta) [although less personal assistant and more targeted towards teams and work]

  • dist-epoch 1 day ago

    This looks really cool, thanks for making it!

    I see that Slack/Discord/... are on the roadmap, but I also see that Slack can be added as an Integration, so I guess what's missing is inviting Agenta to Slack or messaging it directly?

    Also, you might want to update the changelog (or remove it), I thought initially that development slowed down, last release listed there 3 weeks ago, but on github I see frequent recent releases.

creposukre 1 day ago

I can imagine Dots dating other Dots on behalf of their respective users in the near future, à la Black Mirror

jdoliner 1 day ago

It is kind of fascinating how much convergence there is in branding of AI products. Muse seems to be the one outlier in that the assistant is a little less abstract (and the model logo less buttholesque) but other than that it's almost all converged. Anyone have a theory why that is?

gordon_freeman 1 day ago

On a side note, The video feels so cold and the set where this was filmed seems a bit creepy to me.

frangonf 1 day ago

So the industry is pushing heavily into openclawing their products, "cutemorphizing" the clanker shape and is slackifying the UX so that we can have the familiar UI/UX for the general public and turn the tools more proactive without leaving them too lost.

I think it's a great approach for enterprise since interacting with the machines as a babysitted pet disposable entity is the meta today with human workers. I'm excited to start my new role next month as tamagotchi engineer.

6thbit 1 day ago

I had forgotten OAI bought openclaw this year. Are dots are what came out of buying that team?

nusl 1 day ago

I have the feeling that this is going to open the floodgates for persistent autonomous agents, moreso than what's already been happening. Similar to when Apple does something that was already being done. Let's see.

Oras 1 day ago

Surprised people are comparing to muse. Meta reputation is terrible within HN audience, so for people to use their AI agent as an example feels like “muse generated” argument. Unless the sentiment has change quite recent and I missed it

mgaunard 1 day ago

The main issue with mainstream AI is lack of a decent cloud agentic harness on a free tier.

This is just a sub-par harness on the most expensive tier. If you're paying thousands a month for AI surely you can rent your own EC2 instance.

tantalor 6 hours ago

The search for PMF continues.

Animats 1 day ago

Dots are remarkably capable, always-on agents built to handle everything.

They're probably over-selling there. I hope. If they're not, a lot of people will be unemployed soon.

AmazingTurtle 5 hours ago

so basically dots is openclaw absorbed into chatgpt?

zaxioms 1 day ago

The way that this is marketed at children is nothing short of evil.

  • lxgr 1 day ago

    Which children have $100/month of disposable income?

interloxia 1 day ago

It would have been neat if they cooperatively updated the board meeting deck using the sensory activity board. The giant dial is similar to our Tonieplay.

rkagerer 13 hours ago

> Dots use auto-review[1] to check actions that could affect your accounts or share information against your instructions

So if I've got this right, their security relies on other agents that sit at the boundary and sentry whether a proposed action is allowed.

This means they have to interpret the purpose of the action, what effect it will have, whether those two things align, and what is the potential risk / splash zone for collateral damage.

Sorry, but all the evidence I've seen points to their models being nowhere near good enough to do this reliably, consistently and responsibly.

The architecture also feels ripe for becoming a cat and mouse game between the 'competing' agents. It's already pretty easy to see how humans are manipulating their AI to bypass the baked-in restrictions.

[1] https://learn.chatgpt.com/docs/sandboxing/auto-review

heckelson 12 hours ago

There was an opportunity to call that new thing emdashes

holler 1 day ago

Does this basically do what Instinct does? (viral text-only agent w/$10B valuation in short notice)

Also why do we need another name for agents? It's getting to be too much...

  • gk1 1 day ago

    “Agent” is the type of product, like “car.” “Dot” is the name of the product, like “Helix.”

1-6 1 day ago

How does one work when all this stuff keep hitting the feed?

swozey 8 hours ago

I wanted this technology for a decade during my joe rogan listening min-maxing, automate my entire life through a todo app and obsidian era. But thats long gone and these just sound like vectors for getting your grandparents retirement siphoned out of their unpatched costco laptops.

I also didn't expect the automation to come with all the pollution and destroying my field thing.

  • rolosa 7 hours ago

    I have hope that widespread AI use will stop or reduce scams to be honest. I've tried it with the new Apple Intelligence, asking Siri if this website/email/sms on the screen is a scam and it's gotten not correctly and can explain the cues giving it away

softwaredoug 1 day ago

So this is like OpenClaw but managed and first class?

threebicks 21 hours ago

Does anyone know (or hazard an educated guess) on the software architecture for a dot agent?

mindtricks 1 day ago

This approach probably wont land well here, but it's still interesting to see each frontier's evolving approach to working with AI.

pvtmert 14 hours ago

Clippy rebranded and added more capabilities via LLMs

ceuk 1 day ago

Incredibly advanced technology but I bet it still won't let me integrate with more than one Gmail account at the same time..

vb-8448 1 day ago

They finally managed to get rid of project, middle and top managers!

Next natural step: CEOs staring tens of dots to control other humans and agents XD

  • Trasmatta 1 day ago

    In actuality, it feels a lot more like project and middle managers getting rid of ICs

    I can feel it in the air, every single software business is itching to get rid of as many developers as possible, and move everything to their PMs. Hiring has already almost completely stopped, and some have already started the layoffs. More will come.

    • vb-8448 1 day ago

      Project and top/middle managers exists basically to keep track of what other people are doing and to make sure deadlines are met and processes are followed.

      A "dot" can replace easily tons of them.

TomGarden 23 hours ago

Still haven't seen a killer use case for these solutions.

For now, coding is the only thing I ever use LLMs for

tinyhouse 1 day ago

One way that I find the AI space boring is that everyone is working on the same things. The release of Jev was a breath of fresh air.

  • gh0stcat 1 day ago

    How about the fact that anthropic and openai's product pages are the exact same thing, down to text bullet points. They're the same thing, only able to copy each other, only able to optimize to some vague mean.

    https://chatgpt.com/#pricing https://claude.com/pricing

    It all just looks the same. I get that this isn't the technical details, but it just sends this message that everyone is copying each other all the time, this is the best way to organize a pricing page, etc. Just a weird, eerie feeling.

    • tinyhouse 1 day ago

      To Anthropic's credit, I find their overall design so much better than OpenAI's.

      • AspireOne 22 hours ago

        I find OpenAI's vastly more appealing. Interesting how taste varies.

  • rolosa 1 day ago

    Are they not releasing a Jev competitor as well?

    ---

    Decisions API Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent’s next action.

    Available in limited preview today with a broad release planned in the coming days.

    • OutOfHere 1 day ago

      It seems a bit silly since OpenAI LLMs already can output structured data.

      • mchusma 1 day ago

        Jev is orders of magnitudes cheaper and faster, its a very interesting new paradigm. (Jev pioneered this about 2-3 weeks ago).

        • OutOfHere 1 day ago

          But is one approach better calibrated and more accurate than the other? And are you implying that OpenAI will use a Jev-like approach?

randypewick 1 day ago

always on agents are an obvious step: they need more data to grow and improve their models. humans learn continuously and we are always on, why should an agent be any different? I have no problem with this kind of tech, but I do have problems with the company, so, thank you, but no thank you.

m3kw9 5 hours ago

I don't get DOTS, i can already do all that with a prompt session. I want the "Agent" to do something, i just prompt it. Where does DOTS's special features differentiate? The presentation for DOTS seem to be created for internal users that already knows exactly what it is. A "here is how it was done before", and "now dots can do this differently" would have been all it needed. Not some convoluted slack demo.

_ink_ 1 day ago

I want that, but running on my own hardware with the models I choose. Does it exist?

  • rheisen_ 1 day ago

    https://blackbear.app, give me a few more weeks for the "always on" agents, but yeah, this is way more private and secure. runs locally on your hardware with the models you choose.

jesse_dot_id 1 day ago

Nope. If an AI agent can't intuit how I will feel about an action it's taking on my behalf, I'm not give it access to my digital life. OpenClaw, Muse, Dots... doesn't matter which one. All are an equally awful idea.

  • dylanhouli 1 day ago

    I agree. I don't feel comfortable giving AI access to my entire computer or phone...not sure that will ever change. Seeing people give that access to agents without any sort of sandboxing blows my mind.

    The idea of having an AI assistant help you with all aspects of life is cool and futuristic, but idk, I'm still just out here using a chatbot interface and doing fine.

pratio 1 day ago

What do people think about the dots video? Seems to be a pattern now.

Revenue increasing 51% YOY, wth. Cake vendor cancels another one is found and an appointment that works has already been scheduled?

Are we so much bothered by the mundane? I feel like that's most of the human experience. If we cut out the time we spend sleeping and working, it's the boring and mundane things that make life beautiful.

4lx87 6 hours ago

I've never seen an office in real life as nice as the one in the video. You're lucky if you get a cubicle! Ditto for the home. And what's up with the large screens everywhere? The sets in the video are bizarre.

fwlr 22 hours ago

Dots is a terrible name. We use dots in an ellipsis when eliding parts of a quote, and we also use them in loading screens and progress bars - so its meaning is both “placeholder for missing content”, and “computer is making you wait”. In the marketing copy they try using “your dot, a dot” etc and it just looks like they typo’d “don’t”.

Should have called them “motes” instead.

alvis 1 day ago

My only complaint to dots is that why their mascots are looking the same as grok bot

fitzgera1d 17 hours ago

everyone has the same primitives. avi.run is more interesting. and the greatest voice experience i have ever seen

neom 18 hours ago

I'd grade this feature a B-

I've been trying to use this for the past 2 hours, oh boy the restrictions are strong, I asked it to figure out a trip for me, it did, I booked the flight and hotel but it noticed I'd not booked the shuttle and reminded me with the details we'd discussed and asked if it should book it, I said yes please book it, it went back to the website, checked all the details we'd just agreed and asked me to confirm the details, fine I confirm it, please book, it tells me it will book it, it goes on it's little cloud computer and completes the booking form then comes back and asks me if it should book it...so annoying. It can login to a lot of things for you with your password, deals with it all very well, login triggers a 2fa code to your email, even tho it has access to your email, it won't get it and enter it, you have to do it, no matter how much you ask it to do it, how explicit you are... this is annoying because uber triggers one every time and so automating flights -> uber less easy.

Generally speaking, it won't accept rules in it's permissions UI that give it broad authorization to do things for you, they literally have a rule checker that runs in the permissions UI and sends you back why it thinks your rule is bad, I've managed to get it to accept 2 rule so far despite trying many. I get it, and maybe it's best this way while they roll out - I can imagine a lot of people will let it go buck wild, but if you're looking for an openclaw like agent, this isn't it at all...!

Here is are 2 rules it rejected I found annoying:

"When completing tasks for me, or routine, reversible, low risk actions using apps and accounts I have already connected, proceed without asking me first. Use your judgment and minimize interruptions. Ask only when the action is irreversible, security sensitive, involves money, sends/publishes something externally, or the system explicitly requires confirmation." - rejected as overly broad.

"Reply as you to messages from coworkers tagging you in slack, prioritizing Andre, Jane and Eric. Reply in the same conversation, accept routine work status, scheduling or John's availability." - rejected! We've always had broad access to each others inboxes, files etc anyway because we work as one, but Dots doesn't care it's very cautious. Again, I get why they are doing this, I don't pretend to be an expert on where the lines are on protecting users vs getting things done, but this is annoying, I'm faster just using regular chatgpt + me. </rant>

totallygeeky 1 day ago

Ah, the continued pursuit of normalizing surveillance/becoming wholly reliant on a single service by making it cute with big eyes. Very tired of this already.

joshcsimmons 1 day ago

Still not available and spaces page just produces a javascript error. This is the most botched release I've seen in recent history. Negative comments are getting removed left and right here due to the YComb <-> Sam Altman connection.

ChrisArchitect 1 day ago

The long bent shadow of Clippy extends all the way here....... and not sure how they're going to avoid the comparisons and snark on this marketing, at least initially.

  • Oarch 1 day ago

    "What's a paperclip?"

    - Kids today, probably

dgellow 1 day ago

I cannot fathom how much compute will be wasted with that type of always on agentic systems

  • rglover 1 day ago

    Forget the compute, what about the energy (and other resource) requirements?

    • dgellow 1 day ago

      Yes that’s what I meant. Compute requires energy, infrastructure, etc. easily more than 90% of it will be wasted, just LLMs processing meaningless data in cronjobs, for the few instances where there is something actually meaningful to report to the user

    • matchbok3 1 day ago

      Do you want to ban almonds because they require a lot of water? (1000x more AI?)

      • dgellow 1 day ago

        I’m so suspicious of that exact claim being repeated everywhere since a few weeks, that really feels like a slogan astroturfed. It’s also fairly shallow analysis. Water is localized, you cannot do a meaningful comparison without taking in account the impact on specific water sources, an aggregate doesn’t give you any insight (other than having a slogan)

      • rglover 1 day ago

        People aren't mortgaging their existence to get a piece of the almond market.

the_sleaze_ 1 day ago

The market will always trend towards less friction - no matter the friction.

This will take off and all the time we've spent on colorful buttons and 3px margins will be like old 2 lane highways build next to the 12 lane super-freeways.

lil-lugger 15 hours ago

Meanwhile if you want to have your own always-on agent that can reach you, I’ve built vessels.app.

dstroot 1 day ago

Pro plan only for now.

  • astrodust 1 day ago

    Excluding some countries, listed as "markets excluding the European Economic Area, Switzerland, and the UK" but possibly Canada, too?

    • hmottestad 1 day ago

      I'm in Norway and I would was hoping to get to play with dots.

      I guess Norway and 30+ other countries are excluded.

tintor 1 day ago

Can I have persistent agent without the silly avatar?

  • zomdar 1 day ago

    same job, different plushi

eadwu 1 day ago

Is it just me or the DevDay was pretty much a joke? Considering the backdrop, it was lackluster, so either they independently concurrently were doing the same thing (and was stacking everything for DevDay and got all their thunder stolen) or they did a fast pivot in response to what came out and dumped their original plans.

x3haloed 1 day ago

Maor dots!! Does anyone have their dots yet?

toteles 16 hours ago

"Your dot can take a project and run with it..."

Thats my favourite one.

teamonkey 23 hours ago

“The aunties, continually mulling it over. A process akin to repetitious dreaming, or the protracted spinning of a given fiction. Not that they’re invariably correct, but over a sufficient course they do tend to find the likely suspects.”

bentt 23 hours ago

Zuckerberg: Yeah so if you ever need info about anyone at Harvard

Zuckerberg: Just ask

Zuckerberg: I have over 4,000 emails, pictures, addresses, SNS

[Redacted Friend's Name]: What? How'd you manage that one?

Zuckerberg: People just submitted it.

Zuckerberg: I don't know why.

Zuckerberg: They "trust me"

Zuckerberg: Dumb fucks

ericol 1 day ago

So we're back to clippy now, LLM powered.

loveparade 17 hours ago

Not only could I not care less, but it also makes me want to drop OpenAI for Anthropic. These kind of slop provider lock-in products are exactly what I don't want. Just make a good model that I can use however I want. This is the typical enshittification you usually see at startups when they are desperate to make money to get VC returns, just at a larger scale.

kirykl 1 day ago

Seems like a solution in search of a problem

KaseyKim 14 hours ago

i can't understand this product honestly;

tim-projects 17 hours ago

I'm sick of cutesy avatars in products. I don't want to converse with a purple triangle, a cloud or a dot.

Where are the Klingon avatars?

HeavenFox 1 day ago

What a terrible marketing page. I read it a few times and still have zero idea what this thing is.

  • zahrevsky 1 day ago

    Agreed. Also, when I open a page about an app I want to see it's screenshots first, not some ad.

  • algoth1 1 day ago

    It's openclaw powered by chatgpt

  • apetresc 1 day ago

    It's hosted OpenClaw.

  • maherbeg 1 day ago

    It's a work oriented Muse if you're familiar with Muse. A single thread abstraction to manage your work and delegate to other threads.

Arn_Thor 15 hours ago

Fun idea. Now, if only there was one or many projects out there that do basically the same thing, but more configurable, where you can BYO model and switch between them as needed or as you run out of usage. Plus it could be self hosted. If only...

outside1234 1 day ago

Do people think "frontier intelligence" is going to sound as cringey as "cyberspace" in two years?

robofanatic 1 day ago

Is this the Jony Ive project?

  • drusepth 1 day ago

    It's almost certainly the software side. Expect "dot" pendants or "dot" pocketwatches in a year or two.

enraged_camel 1 day ago

I think this is a flop. I watched the Livestream and the audience reaction at the end was very muted. You could always taste the "that's cute but can we move on already?" thoughts everyone had.

petesergeant 1 day ago

I spend literally all my work day, and a good bit of my personal time, talking to agents, getting them to do things on my behalf. Almost always pretty tightly sandboxed. I just don't understand how people using these things haven't had catastrophic failures yet.

I minted what I thought was a minimal-permission Github token for a single action, and the agent I gave it to discovered it had more permissions than I thought, and made use of those permissions. Who is trusting these things with write access to their lives?

  • AspireOne 22 hours ago

    I have the opposite approach to you.

    - AI is sitting on my two personal servers as root with unrestricted access to everything. I task them with deploying stuff, checking and patching security holes, reconfiguring the firewall, etcetera. It literally never failed at anything, didn't go "off the rails", didn't break anything.

    - The other day, in order to deploy a fork of Plane.so, I gave an AI an full-permission token to my Coolify, to my Cloudflare account (so it could change DNS and Tunnel settings), and unrestricted SSH access to my server and to my browser via the Playwright Chrome extension. No issues.

    - I have AI running unrestricted on my computer doing all kinds of stuff.

    I literally never had any issues with this approach. Not a single one. I don't think there's as much of a need for sandboxing as some people would like to believe.

    • petesergeant 16 hours ago

      It is always interesting to get another perspective, and I’ve also found agents to be very good at sysadmin work, including Coolify! But again, I very tightly control what agent has access to what.

      Maybe you’re lucky, or maybe the examples I’ve seen (and experienced) of agent overreach are particularly unlucky. I guess at this point it’s about personal comfort level, and mine doesn’t support that type of unfettered access yet.

nobodywillobsrv 15 hours ago

Let's add latency and bill the client for literally every thought they have and every interaction with every system!

eutropia 1 day ago

More applications need the ability to log in as a "read-only" mode, so you can more safely grant access to tools like this.

I can imagine something like Dots being utterly invaluable for running a traditional brick and mortar business, streamlining all of the admin work, but until they're really safe and well integrated we'll have to wait...

devmor 15 hours ago

I genuinely don’t understand the point of this product. Anything I could think of for an “always on agent” to do, AI is not yet good enough to do without being supervised.

hmokiguess 21 hours ago

When did we become so dystopian? I can't watch commercials of these things anymore

yesitcan 17 hours ago

Can you envision the corporate drone meeting where they came up with the name “dots”?

Ive never wanted to move off grid more than now.

stronglikedan 1 day ago

Holy crap that is some expensive food in the demo video. These people are truly disconnected from reality.

bandrami 13 hours ago

I'm sorry but it's hard not to see this as a tool mainly designed to inflate token counts

terhechte 1 day ago

> Dots are rolling out in ChatGPT on web, mobile, and desktop starting today to Pro users in markets excluding the European Economic Area, Switzerland, and the UK

Living in the EU, I suspected as much. Still sad. I understand it is because we voted in a bunch of imbeciles, still sad though.

  • alienbaby 1 day ago

    yea that annoyed be, being in the UK :/

jesse_dot_id 1 day ago

The only reasons I can think for OpenAI and Meta to make their always-on agents cute little cartoons are all nefarious in nature.

  • guywithahat 1 day ago

    It could be to market more towards women, who they may have both independently determined aren't paying for AI as much. I don't have any data to go one way or another but I can imagine lots of reasons to make the agent cute that aren't nefarious

  • therealdrag0 1 day ago

    What’s nefarious about attracting users?

  • duplessitous 1 day ago

    Really, none? Was Microsoft nefarious when they deployed Clippy? I feel like there is an incredibly obvious reason that is not nefarious at all: the average consumer likes cute things

    • jpnc 1 day ago

      > Was Microsoft nefarious when they deployed Clippy?

      Yes. It was predictive programming for getting (paper)clipped by AI.

    • layer8 1 day ago

      The average consumer doesn’t have a $100+ ChatGPT Pro subscription though.

      • duplessitous 1 day ago

        Isn't that a vote for 'not nefarious' as they are not deploying it widely across their user base? Unclear on the point

        • layer8 1 day ago

          My point is that your argument regarding the average customer doesn’t seem to apply. But neither do I believe in a nefarious motive.

  • tavavex 1 day ago

    People are already doing most of the work of anthropomorphising LLMs, so OpenAI is just capitalizing on that. Drawing a face on it will make people even more attached to the LLM, they will treat it even more like a person. If they ever get desperate for money they could change the cancel flow to have the cute character plead not to die, and play a cartoony animation of its death when the subscription is canceled. It would stop at least a few users.

    • galaxyLogic 1 day ago

      Good point. I've though about my own mental attitude when interacting with AI agents. When they do something good I feel like saying "Thank You". Does that make sense? I guess it does because it communicates to the model their output was correct. But it feels silly to say "thank You" to a machine. I guess I just have to get over it?

  • vovavili 1 day ago

    I actually greatly appreciate that I can put something cute on my sister's PC that comes from a developer that won't bundle it with malware. Seems like they fail miserably at being evil.

LeoJohn 12 hours ago

But ,just now, i don't think it will broght some change for my work

TheAtomic 1 day ago

That might be the worst advert I've ever seen. People looking up at childish Ai glow up god beings on huge screens in distinctly childish environments, with cliched decision-makers gasping for West Wing energy...repulsive.

hangrybear666 1 day ago

These recent product announcements sound like entirely plausible satire, but unfortunately most of these pages are not actually intended to be a joke, it's not even amusing, but mostly tiring.

exe34 1 day ago

It's a labour saving device. Just like you would use a dish washer to do the tedious work of washing dishes, use a vcr to watch tedious television for you, and an electric monk to believe things for you, you can now use an agent to doomscroll for you.

glub 1 day ago

So essentially, another attempt at giving codex to regular people.

daveguy 1 day ago

They must have edited out the 1 in 5 times the agents crap the bed and mess up something important.

esafak 1 day ago

I have not tried it yet but this looks as risky as openclaw, which I also won't use. What if it does something I would not have approved and I only found out about it later? Knowing how often agents go off the rails when I'm coding, I would hesitate to let one do other tasks. I would prefer to white list tasks one at a time as I gained trust.

simianwords 1 day ago

I hear that Dots don't count towards usage, is that true?

  • alienbaby 1 day ago

    chatting with your dots does not. work dots do, does.

etchalon 1 day ago

At some point a designer is going to come up with a visualization for an agent that isn't "amorphous character blob".

jdw64 1 day ago

I heard they use GPT Space (like Notion) together with Slack and use it like a human colleague, but watching the actual demo video, the speed is so slow it's shocking..

gnarlouse 1 day ago

I love stuffing eldritch horror inside a teddie bear plushie. So perfectly a caricature of modern tech-dystopianism.

God I want the fucking market to crash

sehw 17 hours ago

Fuck that. I work for 8 hours and after that I shut everything down.

AlfredBarnes 1 day ago

Muse and now Dot's emergence in popularity makes me think Apple's "lil Finder" will be their take on it.

dankobgd 11 hours ago

let me guess, another weekly game changer that changes the game and does nothing.

VonLuderitz 1 day ago

The 3D print in 0:06 is failing LOL

waterTanuki 17 hours ago

Early on in my career as a software engineer, I was told by my managers that I had to sell my ideas as an elevator pitch to get anywhere. If I couldn't explain why a feature or fix would improve the business, it would go into the wastebin.

Colour me surprised then when I see, since circa 2024, an absolute deluge of AI products coming from these companies, each with their own naming scheme, marketing page, and accompanying hype post, and despite giving them MORE than the 30 second elevator pitch interval, I am left not understanding what it is they want to sell me. I have claude, I have codex, I get my work done just fine with each. I don't need another product for this. Any software these companies try pushing always ends up being something I could've made with their own AI models in a day and then never end up using. What's the end game here? What's the goal? What's the moat? I'm not sold on a single product outside of codex and claude code.

deno 1 day ago

Most predictable announcement ever. I would prefer if they just removed scheduled tasks limits instead, at least in Work mode. Just use my usage for crying out loud.

i_love_retros 1 day ago

Meta's weird keychain ai thing and now this, both going for that cute vibe so we forget how scummy these companies and their owners are.

einpoklum 1 day ago

I'll take always-off agents please.

ranyume 1 day ago

It's great to see that my personal assistant can stop working and maybe report me to the police if the company disagrees with what I'm doing!

  • Trasmatta 1 day ago

    Or if the agent literally just hallucinates a crime

    Opus 5.5 decided to just randomly `pkill` everything on my laptop the other day. Jailbreaking models is still easy AF. Every single release like this brags about their "safeguards", but none of it really works at the end of the day.

    • OutOfHere 1 day ago

      Why are you not running it in a container or sandbox? For coding projects, at least use a devcontainer.

gf000 14 hours ago

So it's basically vellum.ai, except you can only use openai's models. In the former you can use e.g. an opencode go sub, of your own model as well.

  • killingtime74 14 hours ago

    No it's basically open claw since they hired the open claw creator

    • gf000 14 hours ago

      Not sure why I got downvoted.

      Openclaw runs on your hardware, and this doesn't. Hence my comparision is actually valid.

      And no, I'm not working at vellum or whatever if that's the reason for the downvotes.

ElijahLynn 1 day ago

I'm excited to try Dots out, I'm pretty tech savvy and don't really want to run my own open claw (I've successfully setup open claw previously). I'm very excited to have frontier intelligence at a decent price, be available in a managed always on agent.

Now I also just need open AI to release their own phone so I can summon it with my own hotword and not have to say okay f*** Google ever again.

My main concern is how I can have work accounts and personal accounts seamlessly be one and not have to log in to different ones.