pbasista 1 day ago

Tangential:

I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.

This partial prompt data might potentially be used to "pre-warm" some kind of cache.

But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.

  • ShinyLeftPad 1 day ago

    I wouldn't be surprised if their privacy policy would say "what you send is private" and then it wouldn't apply to unfinished prompts on technicality

    • jwstillwater 1 day ago

      This is my concern as well- the same wiggle-room methodology that allowed a business to claim not to “sell or share” PII, because “user data collaboration” was not part of the legal definition prior to CCPA.

      OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.

      This comment by Falserum on the mathematics research post articulates it well:

      https://news.ycombinator.com/item?id=49649992

  • kridsdale1 1 day ago

    Enough Meta executives have been hired there. Expect the same behavior over time.

  • derefr 1 day ago

    > But it may also be used to track the user's writing cadence, error correction style

    I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.

  • sroussey 22 hours ago

    You open the same convo on another device and the partially written text is there to continue. Does it not do that for you?

    • port11 11 hours ago

      “Oh, we’ll store your unfinished thoughts, privately typed into a text box, on the off-chance that you want to continue writing it on another device. And no, we won’t ask you nor give you any semblance of true privacy.”

      Sounds good.

      • broken-kebab 9 hours ago

        "Privately typed" is a bit of a stretch, frankly speaking. That text box literally exists to send your texts to a remote entity.

kdaniel_03 1 day ago

It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't. Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.

postalcoder 1 day ago

My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy.

Perplexity does this. Visiting a past perplexity search url exposes your full conversation.

  • msdz 1 day ago

    Genuinely asking: If you don’t share the UUID-based URL yourself, what makes it not privacy-friendly?

    It’s not like someone’s gonna guess that URL… right?

    • postalcoder 1 day ago

      Yes, technically, guessing a url is impossible. But browser histories are stored in cleartext and trivially accessible to sketchy actors.

      I also accidentally paste random stuff into input boxes all the time.

      • msdz 1 day ago

        Good points, thanks.

        Although I think at the point of some on-device program reading your browser history against your will, you’re gonna have bigger problems.

        • kjs3 1 day ago

          You already have bigger problems then. Most people have given any number of plugins, etc., access to peek at browser history, clipboards, etc and didn't realize it. Run an ad blocker? VPN? Check out the permissions those sorts of things have on your device.

          • lukan 1 day ago

            I want to get rid of whatsapp. Well, since years, but now the need increased.

            On android it is possible to give permissions "once" "while using the app" "always".

            But whatsapp now only accepts the full camera access. If you set "ask every time" to indeed only make a picture once and then no camera access anymore, it refuses and sends you to the permission dialoge.

            • kjs3 1 day ago

              So they aren't even pretending to not say "screw you and your privacy". Glad I don't have need, increasing or otherwise, to use Whatsapp.

  • albert_e 1 day ago

    Security by obscurity -- such an age old anti-pattern!

    I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link

    Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped

    There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines

    This is shockingly lax approach to data security and privacy by design

    • postalcoder 1 day ago

      Chat UIs are a minefield of “if you accidentally click this your data will be shared or trained without you realizing it!”

    • kevindamm 1 day ago

      The same assumptions are true about giving any human that shareable link. They could pass it on to anyone, screenshot it, paste it into their own session. This has been true since before "share with link" permissions on Docs and elsewhere.

      If you click "provide a shareable link" you should decide (and behave) as though that made it public.

      I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."

  • someonebaggy 1 day ago

    Isn't that equivalent to a password? Knowing my password exposes my full data.

    • wtetzner 1 day ago

      You don't store your password in the URL.

      • someonebaggy 1 day ago

        I store my session token in a cookie, which is even worse because it's sent with every request.

        • bsharper 1 day ago

          Not in a URL generally, and if it is the only people who can see the full URL are the receiver and the sender if HTTPS is properly enabled.

        • SahAssar 1 day ago

          It's not. The cookie only gets sent to the domains/servers you specify and is not accidentally exposed via browser history or copying a link.

    • layerv-ai 1 day ago

      not as bad since the blast radius is only your chat vs. knowing your password exposes all your data.

      however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)

  • 40four 21 hours ago

    It certainly doesn’t expose it to anyone else besides you (when you are logged in), unless you explicitly select the share option.

delis-thumbs-7e 1 day ago

In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.

We have all become Milhouse now.

  • Avicebron 1 day ago

    Inequality has eroded trust in society in ~50 years, a lot of the old models (heh) of how we see the world aren't relevant. It's hard to exist when everything around us is adversarial.

    • kevin_thibedeau 1 day ago

      Data brokers were around 50 years ago. They just had more limited sources to draw from.

      • buellerbueller 1 day ago

        That's like saying my 1980s casio calculator watch and an apple watch are the same. they're not the same.

        • walt_grata 1 day ago

          I love how some people think the scale of a problem doesnt matter. Like yeah we has data brokerw 50 years ago, way less of them and the scope they could collect data about was significantly more limited. That matters a lot

          • buellerbueller 1 day ago

            Most of YC and VC (and the internet) is built on the math of network effects, but you are correct, so many people here don't think scale differentiates things. It might not always be a distinction with a difference, but in some cases, it clearly is. That's why we have the whole darn concept of network effects lol.

          • dataflow 1 day ago

            "A difference in degree becomes a difference in kind" seems to escape too many people when it comes to negative social impact, especially those making money in the scaling business.

        • TeMPOraL 1 day ago

          Indeed. The Casio calculator watch was a reliable, ergonomic watch, with a reliable, ergonomic (for its size) calculator.

    • gmd63 1 day ago

      It's not inequality. It's who we've chosen to reward. Adversarial people have eaten a lot of the world and that's because we let them.

      People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way. The deeper the con, like Trump and Elon, the more damage to trust.

      • keybored 1 day ago

        Eating the world sounds like an, on the balance, unequal outcome.

        • gmd63 1 day ago

          Yes. So we should stop enriching people who do that.

      • coliveira 1 day ago

        Yes, the media has conditioned most of us to believe that anyone starting or investing in a tech company and at least a billion dollars in equity is "doing good" and should be emulated.

        • indoordin0saur 1 day ago

          This just means they're "doing well". Whether or not they're also "doing good" depends on what exactly that company provides to society.

        • consensus1 1 day ago

          The media has been hyperventilating about how evil the tech industry is for the last decade.

      • sssilver 1 day ago

        Isn't it more like "it's who capitalism rewards"?

        All we've chosen is free markets.

        • gmd63 1 day ago

          Capitalism is fine so long as capital is allocated on merit. It isn't, because we don't have capitalism, we have corporate socialism.

          • sssilver 1 day ago

            I guess I don't understand the term "corporate socialism". What does it mean, and how did it create the world of NVIDIA and AMD, Google and Microsoft, Anthropic and OpenAI?

            Weren't Mr. Huang, his cofounders, and his peers, rewarded based on merit? Where did socialism step in?

            • gmd63 1 day ago

              Whether they were rewarded by actual merit isn't yet known. Their recent wealth is present-dated imaginary future gains paid by speculators. If by "merit", you mean the ability to perform skillfully to serve the whims of people with a lot of money, then yes.

              Who is spending the kind of money that makes regular NVIDIA employees millionaires? You have to consider how those people attained it. And at the scale of NVIDIA's growing value downstream of a rapid investment in AI infrastructure, you can trace a good amount of it to Elon Musk, who has interfered illegally in an election to purchase favor with a party that ended several serious investigations into his companies. Now, the government is using Elon's AI which trails in several benchmarks. Elon's storied history of gaming systems to keep his companies alive only begins there.

            • buellerbueller 1 day ago

              If an individual acted in a way consistent with the "rogue AI agents" that some of your name-checked companies claim their AIs acted, then that individual would be pursued by law enforcement. (For example, Aaron Swartz.) That these companies are allowed to misbehave without repercussion is a form of letting the gains belong to a few, while the risks are socialized among the many. Another example would be "Too Big to Fail" bailouts after 2008. Those bailed out companies should belong to the American populace now, paying us dividends, so that the employees could keep their livelihoods. Alternatively, but I think this is a more callous approach because of knock on effects, they could have been allowed to fail. That they were not is clear and convincing evidence to most people of what I think the prior commenter is referring to as "corporate socialism."

    • keybored 1 day ago

      It’s just inequality full stop. Trust is a word by wannabe-Bernays polsci people who would just so dearly want the proles to trust their betters, but for some dastardly material reasons they don’t.

  • flipping_beacon 1 day ago

    This can't be real I am go gonna ask Willie

    • davsti4 1 day ago

      But, Willie isn't real?

      • water-data-dude 1 day ago

        No, but now you can talk to a model that pretends to be Willie ;)

  • consensus1 1 day ago

    But we are not really like Milhouse. In reality the teachers didn't care about us enough to make fun of us in the teachers lounge. We were just that year's batch of work soon to be forgotten when they move on to the next. This is exactly the same as the advertising companies. You are just a member of the cohort of a particular target campaign. Nobody cares about you or even knows you are part of that target cohort.

    Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.

    • port11 11 hours ago

      I would point you to the insurance companies deciding on premiums based on people’s Amazon purchases as an indicator of risk. There’s plenty of potential for long-term abuse of your profiling data. Nobody cares about you in a specific way, but the algorithms can be designed to target people “like you” and hurt YOU in a very specific manner.

j4k0bfr 1 day ago

This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!

My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.

Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).

  • alansaber 1 day ago

    AI companies rushing an implementation? Surely not :).

  • amarcheschi 1 day ago

    I'm taking an onboarding process for an Ai company helping other (much) bigger Ai companies and the amount of vibecoded platforms and documentation is staggering. Like, training process so broken that the platform just doesn't load sometimes, things that have never even been tried are published and you have to use them and they suck so much because it is apparent that no human ever touched that and probably wouldn't want to

    • Forgeties79 1 day ago

      Doesn’t sound like your new company is going to be very helpful for other AI companies lol

      • amarcheschi 1 day ago

        It's a part time job hiring contractors, I'm doing this only because they pay would be nice but I know very well I can be fired anytime or when the project is done...

        Anyway I'm a student so I'm not looking for something stable. BTW now it's getting better but if I were the client paying for those services I'd be pretty fucking furious for what's going on - guess they'll never know tho

        • Forgeties79 1 day ago

          Soak it up. Front row seat to some wild times lol good luck man

  • mrweasel 1 day ago

    So I only read the abstract, but the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

    If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).

    I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.

    • duskdozer 1 day ago

      Well, some data is deliberately provided at least. It's not a situation of other apps managing to grab the data like the facebook-localhost exploit

      >The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.

    • kjs3 1 day ago

      the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

      Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.

    • dgellow 1 day ago

      I don’t think it matters much if it is intentional or not, no? In both cases it’s a breach of privacy that should be punished. Though I myself do not believe they are selling, I don’t think they reached that level of sophistication yet, they seem to be way too hacky for that type of scheme. But that will for sure come later on

20k 1 day ago

A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you. They don't care about the agreements you've signed. You're just a pile of cash to them

  • int3trap 1 day ago

    > A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you

    Who thinks OpenAI or any big company for that matter give a crap about them? This isn't a popular sentiment at all, it's just patently false.

    • oblio 1 day ago

      All those people with AI friends, boyfriends, girlfriends.

      • jerf 1 day ago

        Yes, everyone who has an AI boyfriend or girlfriend should make sure that they use only open source models that nobody can take away. Even if a particular cloud service stops hosting them at least you can make other arrangements. And the nice thing is those other arrangements should generally get cheaper over time, and that's nice, when your significant other gets cheaper over time.

        In the modern world, it's just critically important that you exert full control over your SO. You wouldn't want them to run off and leak your most intimate secrets to others without your knowledge or control. I'm sure a lot of us could tell stories about our SOs that we didn't fully control and how they ended up leaking a lot of critical information about us. So full control is definitely mandatory in any SO relationship. It's really just prudent

        wait something's gone wrong here

  • Mistletoe 1 day ago

    That’s not fair, we are also more free training and refining of their models.

  • anotha_one 1 day ago

    > You're just a pile of cash to them

    I thought I was the lazy, irrational, inferior human filth whose job they were replacing.

    • sodapopcan 1 day ago

      You're both, which is why none of this makes sense and is yet somehow "inevitable."

  • dspillett 1 day ago

    > You're just a pile of cash to them

    You are currently a cost to them¹ - using your data as free training/refinement, and potentially selling², it is a way to offset that a little.

    ----

    [1] https://isaiprofitable.com/ - some of the green bars are creeping up a bit, but not much unless you count the shovel sellers

    [2] sorry, leaking³

    [3] Though it could of course be incompetence rather than malice, a mistake they are not actually making anything out of, as per Grey's amendment⁴ to Hanlon's razor.

    [4] Any sufficiently advanced incompetence is indistinguishable from malice.

segmondy 1 day ago

Prior to this new AI age, your data was calculated and what was inferred about you was "shallow", but as of today. The sort of profile and things that can be known about you is scary especially if you are constantly engaged with cloud AI. IMO, the number one risk of using cloud AI is loss of privacy and loss of freedom. With AI and capabilities, more controls can be placed on people and the more you put yourself out there, the more you are going to lose.

For example, we now have self driving cars, we have cameras everywhere. Based on your chat with a cloud AI, you can automatically trigger an automatic monitoring event that follows and tracks you in the real world with the fleet of cameras, cars, GPU, cell signal. Your tracking due to AI has moved into the real world and eventually, a self driving car will take you in to be "processed" against your will, not even for what you posted in a public forum, but for your private ribbing and chatting with some cloud AI.

So definitely put local AI into the mix and keep personal stuff and thoughts local only.

vivekpolavarapu 1 day ago

Ad-tech spent 20 years trying to infer intent from clickstreams. Chat apps now hand over an AI-written one-line summary of intent, labeled and keyed to a cookie. Are we sure the ads business model in AI is about ads in the chat, and not the chat as the targeting signal?

troyvit 1 day ago

This paper focuses on the use of the providers' web and phone tools, and the data sharing arrangements they built through ad networks and tracking services like Data Dog. It doesn't talk about how they handle data from API calls. For one thing that can be pretty opaque.

I have a friend/client who understandably doesn't trust the existing privacy policies of the major providers. They have the same problem many of us do: We want the most powerful models, we're willing to pay for them, but we see over and over how much of a frontier the frontier actually is. Frontiers are ugly if you don't have guns.

So for now maybe platform tools like Open WebUI and TypingMind are a good workaround since the big boys don't (apparently) train on API data (for now) or (probably) send that data to advertisers. It would be interesting to confirm that.

zug_zug 1 day ago

time for somebody to make a quick script to poison your chat history by starting 1000 fake conversations with contradictory identifying details "I worry as a lesbian woman, my tween daughter doesn't blah blah Shabbat blah blah move home to Australia"

  • nick486 1 day ago

    ask another llm to generate those prompts for you

  • Madmallard 1 day ago

    probably will lead to account punishments at some point

dhanushnehru 1 day ago

It’s not a leak if it’s the business model.

yoaviram 1 day ago

Not surprising. When privacy is a selling point on the enterprise plan regular users, even paying ones, are the product. OpenAI had this from day one, the writing was on the wall. We all knew it so let's not act surprised now.

[1](https://chatgpt.com/pricing/?type=team)

  • autoexec 1 day ago

    > When privacy is a selling point on the enterprise plan regular users, even paying ones, are the product.

    There's really no reason to think that enterprise plans are immune. While this analysis didn't test enterprise plans, and only focused on third party tracking, the people behind these companies have already demonstrated that they are willing to break the law to get what they want, are willing to lie to their users, willing to lie to the public, and even willing to lie to congress. Yet somehow people seem convinced that they'd never dare to lie to Random Corp LLC

Traster 1 day ago

I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.

First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.

  • jeltz 1 day ago

    > a whole bunch of their execs come from Meta, and all those guys learned the hard way.

    Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.

  • delis-thumbs-7e 1 day ago

    > First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

    And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!

    > But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

    Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.

    • sodapopcan 1 day ago

      > And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!

      Not only this, many will continue to help the company by bullying anyone who decides to stop using the drugs.

  • gorbachev 1 day ago

    > I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.

    The lesson they've learned is that if the profit from an activity is X, and the sanctions/reputational hit costs less than X, it's full steam forward.

  • coliveira 1 day ago

    The only thing they learned from Meta is that this does work!

Coeur 1 day ago

"multiple providers disclose sensitive conversation-derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation."

Not good at all.

  • Aboutplants 1 day ago

    “Screenshots”

    The amount of sensitive information that accidentally gets left on screenshots is pretty large. This is a pretty massive security issue

  • baggachipz 1 day ago

    I'm shocked... SHOCKED! that these companies would sell this data to advertisers and violate the privacy of their users. Who could have seen this coming??

    • davsti4 1 day ago

      No way will your future spouse be an AI entomed robot, with a percentage of your salary endowed to the corporate owner of the hardware.

lukehandcool 1 day ago

You should really assume that any information you give to a private model is going directly into a database. These companies don't care about your privacy. Local, open weight models are the only way to truly protect your data.

  • oreally 1 day ago

    you got any recommendations for local models setups? lightweight enough for a 5060 if possible

    • lukehandcool 1 day ago

      Yes I do! 5060 is tough because you have 8GB of VRAM. That means any model you load onto it will have to be smaller than that (a model that is 6GB on hugging face will take up 6 GB of VRAM just loaded onto your GPU, then need some more headroom for the actual inference).

      So, you could get away with Qwen3.8 quantized to 1 bit which is available here: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.

      Now, depending on how much RAM you have, you can offload some of the inference to the RAM, but this is a performance bottle neck (less tokens per second)

      I haven't tried the 1 bit quantization yet, but I often run the Q4 quantization on a 3090 machine I have Qwen3.8-27B-UD-Q4_K_XL.gguf which is 17.6GB and fits nicely on a 3090 with 24GB of VRAM.

      To run these models you need to download them (usually from hugging face) and run them with something like llama.cpp (I personally recommend this over ollama). If you are doing that you need a GGUF file which is available from many people on hugging face, most famously the account named "unsloth" which takes the Safetensors weights that a company publishes and turns it into these GGUF files which can be run on a desktop with llama.cpp.

      With the 5060 you are limited. If you can get a 3090 you can do really great local work with Q4 Qwen3.8-27B

      In my personal set up (which I will release fully open source soon), my qwen model can actually browse the internet as if it was me (you can actually watch it view pages, its pretty cool), which means your little local model can scrape up to date info off the internet.

      Hopefully that was helpful. Open source is the future!

DrMandalay 1 day ago

The word is "sell" not "leak". This title takes away all agency from the thieves selling private data to advertisers.

  • otabdeveloper4 1 day ago

    The personal data fell off the back of a truck, man.

jsw97 22 hours ago

I would be curious if these concerns survive for chatgpt if you switch off all the marketing-looking stuff in their settings. They let you turn off a bunch of cookie types and there are a couple of useful looking toggles under "marketing privacy".

flipflowdev 20 hours ago

The web vs. mobile comparison is interesting. I’m curious how much of the privacy risk comes from the agent itself versus the platform APIs and permissions around it.

pluc 1 day ago

You thought... they didn't?

gagan2020 1 day ago

All sells but I saw Chinese models are upfront about that most of the time.

classified 1 day ago

Is it still called a leak if it was the whole point and purpose of the deal?

Someone should have to investigate, but I suppose it's all "legal"?

  • lava_pidgeon 1 day ago

    In the US.

    In EU law it is very likely against GDPR.

r0b05 1 day ago

Why do you think they are fighting so hard to dethrone Google?

Google is an ad company.

robertclaus 1 day ago

Hanlon's Razor given that these tools are almost certainly vibe coded at this point?

  • jeltz 1 day ago

    Incompetence and carelessness is rampant in our industry abd with vibe coding it has only gotten worse.

maybewhenthesun 1 day ago

Yea no shit sherlock!

My stance used to be that the invention of the filter bubble combined with targeted advertising is the most dangerous invention for human society of the last 100 years.

AI turns that up to 11

reedf1 1 day ago

my first guess is always Gboard.

okokwhatever 1 day ago

I never thought this could happen (XD)

nwhnwh 1 day ago

I am very surprised.

charcircuit 1 day ago

The paper doesn't say when the app sends the conversion artifact.

titzer 1 day ago

Not surprised. Just wait until they bake the goddamn advertising right into the model.

bix6 1 day ago

Hard to read on mobile. Is there a tldr of how bad each provider is?

Grimeton 1 day ago

Oh no, what, what happened?

What happened? Oh no!

How terrible! That’s just, that’s just awful!

How terrible! Oh no!

folkrav 1 day ago

Insert surprised pikachu meme

Razengan 1 day ago

The ads ~~industry~~ racket is a cancer upon civilization.

This shit needs to be stamped down by law.

People take up pitchforks and torches against AI data centers,

well how much time, money, and resources have been wasted on advertisements over the centuries?

How many people, including children, have been deceived by ads?

How much privacy has been violated in the name of "sErViNg YoU rElEvAnT aDs"?

Have you ever tried browsing YouTube from a poor connection, and noticed how long it takes to serve ads? How much bandwidth has been wasted on ads so far?

Who[m] does it all benefit?

micromacrofoot 1 day ago

these companies are complete dumpster fires and we just keep giving them more to burn

damaru2 1 day ago

Entire conversations via exposed permalinks. For Grok: trackers receiving the conversation URL could access the full chat because the link lacked access controls.

Screenshots of conversations. TikTok received screenshots of Grok chats during sharing, exposing the actual visible conversation content.

Conversation-derived content tied to persistent identifiers, including prompts and automatically generated chat titles revealing sensitive facts. "Salary 85k NYC: mortgage 280–350k".

  • Traster 1 day ago

    That is a staggering level of incompetence.

    • ex1fm3ta 1 day ago

      what do you want it's the era of vibecoding.

aitoolcrux 15 hours ago

This is exactly the kind of rigorous analysis the AI industry needs. From our hands-on testing of 500+ AI tools, we've seen the same pattern: most conversational agents ship with telemetry baked in, and very few disclose what's actually being collected beyond the standard "we may use your data to improve our services" boilerplate.

The prompt-injection angle is particularly concerning — it's not just about passive tracking, it's about the agent's context window becoming an attack surface. When a customer support agent pulls in untrusted web content or email threads, every loaded prompt becomes a potential data exfiltration channel.

More tools need to follow the local-first model (like tools that run entirely on-device or via self-hosted infrastructure) as the default for sensitive workflows.

cleochan 20 hours ago

One important limitation is that the paper excludes enterprise and government tiers, so I would not automatically map these findings onto a contracted business service. But it does suggest a better procurement test than asking only whether prompts train the model. For an approved workplace assistant, document the exact client and tier being used, inspect outbound domains from both web and mobile clients, test reject-cookie and no-sharing configurations, and ask the supplier to identify subprocessors, retention periods, conversation-link controls, and material change notifications. Repeat the test after major client updates. Where verification is not possible, restrict the data classes users may enter. I would also distinguish observed third-party contact from proof that full prompt content reached every endpoint.

spncai 1 day ago

The shared links are bad, but persistent memory makes the blast radius much larger. At that point the product holds a durable record of what someone has worked out over time, not just one conversation. That needs real access controls and revocation, not a UUID.

sick_of_slop 1 day ago

Of course they do. As soon as ads entered the equation this was always going to happen. Nobody should be suprised.

drywater2 1 day ago

No, they don't "leak data", data is sold. Leaking data requires a mistake. This is intentional.

  • kjs3 1 day ago

    Thank you. I keep saying this and the spin doctors keep using words that imply "oh, my...I am so sorry we had that tiny little privacy issue...we'll get right on that in the next sprint...". And winning the message war.

    It's not a 'leak'. It's in the T&Cs you agreed to. This is the business plan.

  • dgellow 1 day ago

    I’m not sure. It could definitely be incompetence. We are talking about an industry rushing as fast as possible in every directions, their systems are very likely a complete mess. I really wouldn’t be surprised if advertisers are getting that data for free just because the AI labs didn’t bother

    • coliveira 1 day ago

      It's because of people like you that these companies have a free ride. When we're talking about billion dollar companies there's no such thing as a "mistake", because they need to be held accountable for everything that they do. Thinking otherwise let them keep the rewards while getting away with anything wrong they cause to society.

      • wtetzner 1 day ago

        Just because something is due to incompetence instead of malice doesn't mean they shouldn't be held accountable.

        • yoyohello13 1 day ago

          Blaming incompetence instead of malice is always used as a backlash diffusion mechanism. We should stop attributing incompetence to obvious malice.

          • paimapi 1 day ago

            there's always some organizational, systematized intent in the creation of incompetence. all organizations should be accountable for it's workers and their adherence to policy - it's why safety regulations and training exist, after all

            if you have an individual who has been sexually harassing others regularly the problem exists both with the individual and with there not being clear reporting standards, remediation processes, and person-by-person training opportunities (or only offering lackluster, barely enforced versions of the training)

            organizations, especially ones as large and well-funded as LLM model providers that can hire whole fleets of trainers, ethics watchdogs, etc, that then produce large-scale, systematic incompetence are and should be the primary culprit and the one targeted by accountability processes

            outside of someone just straight up embezzling funds or grossly misusing their power, organizations should also bear the brunt of the culpability especially when it relates to harms to the public good, especially if those harms benefit the organizations directly. and even in those cases the question becomes why weren't there accountability practices? checks on power? etc

          • dgellow 1 day ago

            Those companies are grossly incompetent to the point of committing what is likely to be felonies. If you think I’m deflecting you’re misguided, I’ve been advocating for months for the AI labs to be shutdown and their leadership investigated for gross negligence and potential crimes.

            I would still bet they aren’t making any money from data leaking to advertisers