dwa3592 1 day ago

I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink

A few things jump out since this is done by the government:

- the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.

- it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.

- technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.

  • bberenberg 1 day ago

    I think all of your points are very valid, but it doesn’t remove the point that the government is trying to make it free and easy for application builders to do a little bit better than they are today. This is commendable on its own, even if it’s not perfect.

    • iAMkenough 23 hours ago

      I find it interesting that the co-author of the OP violated long-standing security practices and copied a social security database into an insecure AWS instance without proper oversight.

      Is it still commendable if in order to develop this, the taxpayer-funded developer themselves published our PII to a consumer, privately controlled third-party server without following federal data security protocols? Seems like an attempt to make up for the damage they caused.

      Standing up your own AWS-equivalent internally is not much work, yet the PII data leaked to publicly available services due to this developer's personal choices and lack of oversight to prevent sensitive government data from being published externally.

      Notably, this resulting work has not been published as public domain (CC0).

    • prometheus1992 21 hours ago

      what i understood from the above comment is that the government should first do the basic policy making (set some non-negotiables for the ai companies etc) before jumping into solution building. let the market figure out the tech solutions.

      • tancop 5 hours ago

        Letting the market handle things is how you get all the problems America has today. Markets are a good way to distribute goods and services but they don't really work for public services with no money to be made. Things like roads, hospitals and open source software.

        The purpose of a government is to serve its people, not make sure every business gets the same conditions at all costs. Picking winners and losers is fine if the losers still have a chance to beat the winners when they're really better.

        If someone makes a free PII detector better than Rampart I bet the US government will either switch to it or make Rampart v2. For now it's enough. Perfect is the enemy of good.

  • applesauce3572 1 day ago

    Point is completely fair. It seems to be sort of a grey area that they skim over entirely on your point about sharing deep personal info.

  • RainedOnCat 12 hours ago

    Absolutely this. "PII" is a moving target that's heavily context-sensitive. Any entity that's large enough to care about PII almost certainly has some category of it that would be indiscernible from a handful of jumbled alphanumerics without a clear understanding of the context, both linguistic and organizational, in which the text sits. A general-purpose model would struggled to solve this in the broadest case, and will certainly be incapable of solving it in any specific case.

bob1029 1 day ago

I have presented approaches like this to banking clients and they are still not very interested. The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.

If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie. ZDR is your best mitigation if you actually want the magical AI experience to work the way the investors hope it can.

PII can often become disclosed by way of many correlated factors that are not considered PII on their own. Even a perfect AI system cannot capture all of these relationships. You could probably locate where I live within a 20 mile radius if you spent enough time analyzing my HN comments over the years. Not one of these comments on their own would trigger a PII filter.

  • Terr_ 17 hours ago

    > If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie.

    For related reasons (e.g. rule compliance) I'm trying to convince folks to shrink the LLM's role to a kind of "guess which valid choice the user meant to ask for" level, and let the data-entry parts be small and deterministic pieces.

    Alas, convincing anyone to use some symbolic GOFAI (or a strict "harness" that is 90% of the way there) seems hard these days. The allure of a promised magical silver bullet is strong.

sscaryterry 20 hours ago

I can't deal with "radaction".

nhinck2 1 day ago

98.4% is nowhere near good enough to call it PII redaction.

throw03172019 1 day ago

Plain text in a chat input is only one piece of the problem. What about files like PDFs and documents filled with PII.

stuaxo 18 hours ago

Nothing to do with the fantastic megadrive game Rampart which is a shame.

handfuloflight 1 day ago

Why did the National Design Studio see the need to put all the readable text on the right column of the page?

  • tancop 5 hours ago

    Because there's one part where you also get text on the left side so Claude thought it's okay to keep it as blank space for the rest. Or they wanted to have an active example on the left for the whole page and ran out of tokens.

    • handfuloflight 3 hours ago

      I am Jack's complete lack of concern for the end user.

swiftcoder 1 day ago

What we'd really like is a PII redaction model for video...

Onavo 1 day ago

Is this good enough for HIPAA?

  • fwip 1 day ago

    No.

    • Onavo 1 day ago

      It's the technique used by a lot of AI healthtech

      • estearum 1 day ago

        What is? This is nowhere close to what HIPAA requires.

        • Onavo 1 day ago

          A lot of clinical decision support AI tools essentially scrub the PHI data on the frontend.

          • estearum 1 day ago

            PHI is a different set of data points from PII amigo.

            • Onavo 1 day ago

              Same technique is good enough.

              • estearum 23 hours ago

                Which is why I was asking what you meant by "this technique."

                If your question is: "Is client-side de-identification sufficient for HIPAA," then sure, assuming you are de-identifying the data elements required by HIPAA to be de-identified. You'd have to be a real big idiot to rely on something like this though.

                If your question is: "Does this library, as described in the blog post, de-identify data sufficiently for HIPAA?" then no, it doesn't appear that it does.

      • fwip 23 hours ago

        If they're relying primarily on this type of tool to scrub PHI, I can tell you already they are not in compliance.

sgnelson 1 day ago

The skeptic in me really can't trust our current government to protect my information.

  • goodmythical 1 day ago

    I mean, it can be run locally, so you don't necessarily have to trust it, unless there's been any model-as-a-vector CVE.

    That hasn't happened yet, has it? Where running a model from HF directly compromises the machine as opposed to some breakout or exfil done by the model after the fact?

    That said, if you're filing your taxes or have a 'real' ID, the government already has all pertinent information required to fuck you over either deliberately or via a leak, so...

    • Ohentis 1 day ago

      The vulnerability to watch out for would be the model being trained to not redact some specific information.

  • Cider9986 1 day ago

    You could never trust the US government to protect our information since it moved online.