palmotea 13 hours ago

> The public posts and discussions being had about this subject are already informing AI companies on how to train their multimodal models to get around these obfuscations, most of which have already been broken. I'd argue every new font and tech demo is effectively a benchmark, daring AI firms come up with solutions to sidestep them. And they will be sidestepped, one way or another. If a human can see the information, that means there is a way the information can be parsed. "Ghost" fonts will become just another scraping obstacle with its own set of contingencies.

1. I don't like the sense of futility and powerlessness this advocates for.

2. I'm not sure it is so futile. I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.

That could happen if:

1. There are so many schemes out there the catalog of circumventions gets unwieldy.

2. Doubly so if the schemes allow generation of new obfuscated fonts per site or per page.

3. Then you're forcing the scrapers to pay a greater tax to get your text: spin up a Chrome instance to OCR a screenshot, or spend some a buck or two or LLM credits to reverse engineer the page in order to scrape it.

  • danudey 11 hours ago

    The main argument against these fonts is that you are removing accessibility for humans permanently in exchange for removing accessibility for scrapers temporarily.

    It might raise the cost for some scrapers, to some degree, but it raises the cost to, effectively, infinity for anyone who needs to use a screen reader, a browser's 'reader' view, or any other assistive technology.

    Things like this just remind me of EA's Spore; it released with DRM so draconian that legitimate players were getting locked out of the game during the first week, while people who pirated the game had no problem whatsoever and were playing the game without issue even before its release. The legitimate users of the thing were the only ones punished by the technology designed to stop everyone but them.

    This is going to be the same thing; a website which AI will be able to read in short order but which assistive technologies will not be able to read ever.

    • jimmaswell 11 hours ago

      I also see no virtue in stopping an AI from reading something to begin with. If anything, I find it anti-social - hide your work from the machine trying to learn from it, taking nothing away from you in exchange for benefiting all of mankind?

      • r_lee 11 hours ago

        you mean benefit a private company training their AI model which they market to the public in order to get to their trillion dollar IPO?

        • infinite_spin 10 hours ago

          That seems like an unfair equivalence. Lots of things you enjoy, including this very forum, benefit already wealthy private companies. That doesn't seem like it's a very good litmus test for whether something is overall a benefit to humanity.

          • Retric 8 hours ago

            > overall a benefit to humanity.

            By that metric a complete ban on LLM’s might be on the table, which I don’t think is something you’re advocating for here.

            • infinite_spin 6 hours ago

              I'm advocating for a litmus test that can't be easily abused, and I'm advocating against unfair equivalences. I also am not calling "overall a benefit to humanity" a metric, it's more of a conclusion we could arrive at by an application of relevant metrics.

              • Retric 5 hours ago

                > a conclusion we could arrive at by an application of relevant metrics.

                That’s dangerous ground because of how you can arbitrarily change weights of different metrics. If the goal is the “overall benefit of humanity” then that’s what’s important not metrics.

                • infinite_spin 5 hours ago

                  How do you propose assessing "overall benefit of humanity" without metrics?

                  • Retric 5 hours ago

                    Qualitative assessment etc not just metrics.

                    How do you propose to asses “overall benefit to humanity” with metrics?

                    • infinite_spin 4 hours ago

                      Number of successful outcomes reported (e.g. with disabled students, cancer patients). Labor costs. Latency of services. Scientific discoveries found. Etc.

                      I'm not trying to discount the qualitative approach, but I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.

                      • SiempreViernes 1 hour ago

                        > I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.

                        One would hope for the slightly higher ambition of metrics that are causally connected to "overall benefit of humanity", settling for merely correlated has that unpleasant risk of the association breaking as you try to optimise.

        • satvikpendem 8 hours ago

          That's why one should advocate for open weight models.

          • wesleywt 2 hours ago

            A trillion dollar IPO buys a lot of politicians to make these illegal.

        • SideQuark 1 hour ago

          So jealousy is the argument? If someone reads my work am I worse off? Why would I post it to the web? If someone indexes it so others can find it am I worse off? Should I complain they also make money helping others find it? If someone build on it in a way I did not do, never planned to do, and found really neat new tech from reading lots of works, why should I get so upset?

          Seems oddly sour grapes.

          • neuroticnews25 1 hour ago

            I don't think we should hold our relations to big corporations to the same moral standards we invented for interacting with people, like "don't be jelous" or "don't be petty", because corporations won't reciprocate. They banned my account, so fuck them, simple as.

      • JodieBenitez 6 hours ago

        > taking nothing away from you

        Hosting is not free.

        • silon42 5 hours ago

          That is the problem. I believe web is long overdue for a torrent like model where hosting is shared among all users (and ISPs instead of 'cloudflares').

          • brnt 5 hours ago

            Opera Unite. The idea was that hosting HTML should be as simple as browsing it.

        • silon42 5 hours ago

          That is the problem. I believe web is long overdue for a torrent like model where hosting is shared among all users (and ISPs).

        • neuroticnews25 2 hours ago

          Marginal cost of serving a text document is ~0.

          • pwdisswordfishq 1 hour ago

            Marginal cost of serving a text document repeatedly to hordes of reckless scrapers hiding behind residential proxies is ⋙0.

      • GlacierFox 3 hours ago

        Is this bait? Wtf haha

        • akramachamarei 2 hours ago

          If the comment baited you, maybe it's because you're holding onto a popular but nonsensical belief system, and you're experiencing painful cognitive dissonance.

          • SiempreViernes 1 hour ago

            Are you intentionally being rude, or is this just how HN raised you to think?

    • customguy 5 hours ago

      So scrapers are using those people as hostages. So let's have something like a HTML meta tag to link to an accessible version of a website and hard jail for using it for scraping.

      It's not a technical issue, it's a social one, people behave differently given the same set of possibilities and incentives, and we can and should target those who fuck it up for everyone.

    • nottorp 2 hours ago

      And even if you don't use a screen reader, how about search? Reader mode? Saving for later?

      Oh and lynx/links...

  • BrenBarn 7 hours ago

    I also don't like the fatalistic mindset, but I feel like we're better off trying to retaliate against the makers and operators of the bots, rather than getting into a technological arms race against the bots themselves. That is, the fight is one of policy, law, and morality, not of technology.

    • palmotea 4 hours ago

      Why not both? Better that than putting all eggs in one basket.

  • pwdisswordfishq 1 hour ago

    > I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.

    Did it work for non-cryptographic DRM? (Broadcast flag, Macrovision, deliberately miswritten floppy sectors, port dongles, physical manual challenge-response...)

evnp 17 hours ago

Thanks for introducing me to shieldfont.org! It's the first of these I've seen that feels designed to be more than a visual experiment, reading through their landing page is interesting. In particular, their section on accessibility seems to contradict this post's opening premise:

> Screen readers get the real words. A screen reader reading down the page is never handed scrambled text, and our NVDA test asserts exactly that. Screen review and touch exploration are untested. ShieldFont hides shielded passages from accessibility tools by default with aria-hidden="true", because a decoy read aloud is fluent, grammatical, wrong English, and that is worse than silence.

> The real words remain sealed in the same page, and a visible notice above the block carries the control that uncovers them. It is on by default and reachable by mouse, keyboard and screen reader alike. Pressing it sets the reader’s browser to solving a compute-heavy puzzle: JavaScript and a few seconds of processing, more than most mass scrapers are willing to spend. That puts the words within reach of a screen reader, a translator and copy/paste.

I'd love to hear your thoughts on that. Also just aside, love this TUI-esque blog design and color palette (maybe a bearblog theme? still worth an upvote)

  • kstenerud 17 hours ago

    When you look at their live demo (https://shieldfont.org/demo/), it says: If you use a screen reader, custom font, or translator, please uncover the text before reading.

    They also actively block copying the text, telling you to "uncover" the text first. The uncover operation is VERY expensive.

    Anyone using assistive technologies or trying to copy "protected" text is SOL.

    Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.

    • evnp 16 hours ago

      Isn't this the same sort of cost CloudFlare and anime catgirls are making us pay daily? Only enough to deter bots, or "a few seconds processing."

      I take your point that one extra click/interaction required for screen readers only is objectively more friction, but it's only by exploring these technologies instead of dismissing them that we'll arrive at UX solutions truly work for people of all stripes (and ideally, not for bots, scrapers, and the like). I think if you're applying a tool like this you already have very different priorities than fueling Google.

      • kstenerud 16 hours ago

        As soon as you need to put a "decode" button for accessibility functions to work, you're effectively posting the key along with the cipher. It's self-defeating because any tool that supports screen readers will also support scraping. It's a fool's errand.

        • HeatrayEnjoyer 14 hours ago

          Accessibility isn't optional, so what do you propose?

          • kstenerud 10 hours ago

            I propose using the internet the way it was designed: Open and readable by everything (including bots). Scraping is a legal issue. There is no technical prevention mechanism that isn't theater.

            • mschuster91 1 hour ago

              > Scraping is a legal issue.

              The problem is... it's a legal issue we cannot solve. America and China are large enough to not give a fuck about what everyone else wants.

    • 317070 13 hours ago

      > Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.

      It's only search engines? And who uses those anymore anyway? Other bots?

      It doesn't affect organic traffic, so you are not really putting yourself into a black hole. A lot of website are driven by social media and organic traffic, so they would be just fine with this approach.

      • 8cvor6j844qw_d6 12 hours ago

        Same thoughts. The black hole concerns are overstated for personal stuff when nothing is at stake. Planning to adopt one of the fonts and see how it goes.

  • csallen 16 hours ago

    It's funny, I was just thinking that the one thing I hate most about the terminal is that its monospaced fonts and overly-long lines are a nightmare for reading. So the fact that somebody designed a blog reading experience to mimic this is just… ugh for me.

    But you seem to appreciate it, and I'm sure others do too. Different strokes for different folks.

  • svara 4 hours ago

    A bit ironic that this is written in idiomatic Claudese.

  • SideQuark 1 hour ago

    > I'd love to hear your thoughts on that.

    “Claude: make the scraper mimic a screen reader.”

    And just like that, in 10 seconds, their site feeds my “screen reader” the real words.

logseman 3 hours ago

> Like it or not, all publicly available information will inevitably become accessible to anyone and anything that has permission to access it. This is the baseline scenario people need to plan around.

The entire point of the situation is that permission is not involved. They just do it. Meanwhile, if I do it to them, I am fined/sent to prison/executed. Until such a time that this baseline scenario of inequality is somehow remedied, there will be a motivation to stop them.

simonh 1 hour ago

So a rounding error of nerds, who themselves are a rounding error, will invest a lot of effort and trouble to use tech that makes their and their user's lives harder, and probably won't even work, to hide text that LLM big tech couldn't care less about anyway.

Cool.

pwdisswordfishq 1 hour ago

> I don't have anything against these specific examples [...]. This post is meant to critique the idea itself. I'm not trying to put-down anyone here!

This is self-contradictory. Either you want to critique something or you do not. Make up your mind.

condour75 18 hours ago

Are these even meant to be used though? It seems more like performance art.

  • gruez 18 hours ago

    With this kinda of stuff it's hard to tell whether the person is doing it unironically, or knows it's "performative art". A while ago there was a trend of using a tool which imperceptibly perturbs an image in a way that supposedly breaks AI training on it. Of course, artists ate it up, despite the skepticism from AI researchers. Same with people setting up their sites to be "AI scraper traps", generating gibberish content. Probably also trivial to filter out, but people do it.

    • pixl97 15 hours ago

      The problem with being dumb satirically is dumb people look up to you as a thought leader.

      • yieldcrv 11 hours ago

        there was this guy that was anti-bitcoin - or specifically against most arguments from enthusiasts - and people started looking up to him to validate their feelings

        and on news and podcasts he wound up correcting so many dumb arguments that he sounded pro-bitcoin and could never get to his own points

        "well, no, not like that, the difficulty algorithm...."

        "there are ways to use it with the power off"

        "well, no, the transaction fees supplant the block reward so ..."

  • ffsm8 14 hours ago

    "caveman speak" skill, need I say more?

    People aren't particularly bright. That's why the scientific method was developed to counteract our built-in tendency for... Unorthodox approaches

  • Animats 13 hours ago

    Agreed. There was a thing a few years ago for dazzle-painting your face to avoid face recognition. This just makes you stand out.

  • HlessClaudesman 5 hours ago

    "Go ahead, obfuscate your contribution to the repository of all human knowledge, see if that impedes our imminent invasion! Moooahahahar!!!" - Kang and Kodos

blehn 17 hours ago

The irony of championing accessibility using low-contrast simulated VGA text...

  • VCFundedGenYer 17 hours ago

    Also the notion that the sides of the pages flash as you scroll due to simulating that old Macintosh monochrome monitor dithering effect. My eyes.

    • boxed 5 hours ago

      Isn't that because you (and me!) have screens with slow response times for pixel color change though?

  • Narishma 17 hours ago

    Where do you see the low contrast?

    • smohare 17 hours ago

      The linked article has a fairly light gray background with white text. I can read it, but the low contrast is just tiresome.

  • hn_throwaway_99 10 hours ago

    When I first was reading this I thought that the author was deliberately using a shitty font to make the point that obfuscated fonts are hard to read.

  • flexagoon 52 minutes ago

    Agree. My vision is not even that bad and yet I literally had to squint and hold my phone very close to be able to read this page.

teekert 1 hour ago

My agents act on my behalf, if you make it difficult for them I (or they) will get their info elsewhere. Will this split the internet? How long is this feasible? How long until any agent can work around any barrier, just at the expense of more energy?

kazinator 7 hours ago

I can tell right away that is stupid because one of the things I've used AI for was for deciphering someone's illegible handwriting, which it did amazingly. Once it figures out for you what is written, you can't unsee it, so you know it has to be right.

Illegible fonts will only create accessibility problems for humans, even ones with normal vision, high literacy and no cognitive defects (dyslexia), while AI will blow right through the text.

If this is done in electronic documents, where the AI won't even see the glyps becaue it's reading the underlying character codes, it's even stupider.

I can't believe anyone would even try this (and then believe it is working without putting their hypotheses to the test).

frollogaston 8 hours ago

Is the video on shieldfont.org AI-generated? It sounds like that voice again.

hk1337 18 hours ago

Anti-AI fonts seems like scrambled porn on cable back in the 1980s.

wesleywt 1 hour ago

Just give up and don't even try. This is what I am getting from this post. There are already efforts to pollute musical data that seems effective.

piker 18 hours ago

There could be benefits unlocked in legal documents by retaining a machine-readable version and distributing the obfuscated version with a legend at the top. We proposed one that said:

"This document contains mitigations against review by automated systems. Recipients should ensure that they have read the contents on screen or in print. Recipients with bona fide vision impairments may be entitled to unmitigated documents upon request."

In testing, obfuscating small portions of text slipped under the radar of most (then-)frontier LLMs.

We used a font that was rendered on the fly and reported faulty or fake Unicode mappings: https://tritium.legal/blog/noroboto but others have proposed and done the same with ligatures.

  • tbalsam 18 hours ago

    There was a story once about a boy with a wheelchair who needed a ramp to get into school, and the school made him use the loading dock ramp used for garbage and other things at the back. The school argued that it was an appropriate accommodation.

    Accessibility is not accessible if you need to go through extra steps to get it.

    Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.

    Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.

    • piker 18 hours ago

      My dad caught paralytic Polio at age 2 and has had limited mobility his entire life, so I'm familiar with that issue.

      Our internal, hypothetical use-case was between contracting parties who were looking to avoid terms escaping into the wild. This shouldn't show up in standard ToS or similar. There are already really good legal reasons for that.

      • doctorpangloss 18 hours ago

        okay, i get that as a lawyer who wants to make money, every client is "heckin cute and valid." and you can hypothesize that this thing is something that clients want: "terms escaping into the wild," whatever that means - are you saying that you think copying and pasting an agreement into an LLM makes its contents escape into the wild, by some mechanism?

        Look, I understand, you don't have to explain to me the theory for how that happens, I know it already. Since I know your a smart guy, to some extent you care about that only because you imagine that clients do. But in reality, in the real world, every email you send is read by at least two people, every contract you sign has multiple parties, etc. You make some obfuscated thing or whatever, but eventually, someone has the real text of the document - it might be YOUR client, it might be the person you are negotiating with, and you rarely represent ALL the sides. You never own ALL the information and all the parties and IT systems in totality. Eventually someone will put the text into a chatbot. Or maybe they put a salient piece of the pre-final text, like some legal theory or merely a question, into the chatbot.

        So I see this font stuff, or watermarking stuff, or all this provenance and control stuff, as deeply illusory. It is the worst circlejerky kind of aesthetic experience making. When you mess with anti AI fonts you are trying to compete in the same business DocuSign is in, that is, in the business of selling holistic social experiences - a whole 7000 person company whose main competition is a fucking pen - but it's not like you're doing something creative. If you care about aesthetic experiences, write a short story! Are you getting it? The itch you are scratching with this weird thing, nobody wants.

    • stronglikedan 18 hours ago

      > Accessibility is not accessible if you need to go through extra steps to get it.

      That seems a bit entitled to me, especially in the story you shared. They had access to the school just like everyone else. Why does it have to be in the exact same spot? Surely they could be dropped off by the loading dock as easily as others could be dropped off out front, and maybe even moreso. Should the wheelchair accessible stall be the first stall in the bathroom so they don't have to go through "extra steps" to get to it? As long as the ramp had the proper gradation to satisfy the ADA, I don't see the problem.

      • shnock 18 hours ago

        > They had access to the school just like everyone else

        They literally did not. They had access to the school from a different entrance than everybody else. Specifically, one intended for cargo before people.

        • fluoridation 18 hours ago

          That's a different sense of "just like". Not "in the same manner", but "to the same extent".

        • binaryturtle 18 hours ago

          Would it have been different, if they had everyone take the cargo entrance?

          • grim_io 17 hours ago

            Yes. It's about discrimination and dehumanizing everyday cruelty.

        • gblargg 17 hours ago

          They also literally cannot have the same access like everyone else, since they're in a wheelchair. Everything will be different.

  • gizmo686 18 hours ago

    It is not sufficient to work against current AI. It needs to also work against AI that has been trained by a competent team aware of your mitigation. Or worse, a competent developer with no particular AI skills.

    Otherwise, you are relying on obscurity, and will lose as soon as you become interesting enough to matter.

    You will also break non-AI machine processing use cases. That isn't just accessibility, it is things like search.

  • aleksejs 18 hours ago

    You will surely not have a good time enforcing the terms of a legal document that explicitly spells out that it is intentionally obfuscated from the party it intends to bind.

    • piker 18 hours ago

      No, that’s not at all what is going on here.

fwlr 4 hours ago

>>We’re better off with the way the web is now

Well, bad news, you aren’t going to get to keep that either. Whether for high-minded reasons like “the knowledge that everything you create will be fed into the slop machine so it can later be regurgitated without attribution is having a deleterious effect on the morale of some contributors”, or for incredibly mundane reasons like “lacking the resources to either serve or block the crawlers”, I fear the web you love is dying of ai with or without anti-ai fonts.

hartator 18 hours ago

It also mostly don’t work.

  • bawolff 18 hours ago

    Yeah, it seems like most of these would probably be easier for an AI to read than a human once you give even a tiny bit of training to the AI.

    there is a reason nobody uses text based captchas anymore.

    • theblazehen 4 hours ago

      Even just giving the image and "do image analysis to figure out what it says" to your LLM will get it done

rappatic 10 hours ago

Ironic that an article about bad fonts would use such an ugly, garish font

Varelion 17 hours ago

Is there evidence shieldfont doesn't work?

  • gs17 17 hours ago

    It only works while it's rare. If it was more common, scrapers would switch to OCR or simply reverse the font so they can decode the ligatures.

    Their own whitepaper brings up a bigger issue: if it works, it poisons search engines as well.

    • Varelion 17 hours ago

      OCR is a more expensive option, right? I don't think these fonts need to stop ai from being trained; I think they just need to make it more expensive and difficult.

      • gs17 16 hours ago

        It's slightly more expensive, but the cost would be worth it if this became common. It's not designed to be hard to OCR, it's designed to be hard to copy out of the source code, so it doesn't require a very advanced OCR system (and all the regions that would need OCR-ing are clearly marked).

      • pixl97 14 hours ago

        More expensive and not worthwhile are two different things. Also when certain implementations become popular it's much more likely someone will write a very efficient kernel for decoding said text making it much less expensive than generic OCR.

        Also more expensive doesn't mean that something won't happen, only the dynamics of how it happens. For example if you put all your documents in images then some service might just sell the AI providers the text. That service may do underhanded things like bundle OCR in an app that does something else and use your phone to get the text out of these images all day.

charcircuit 3 hours ago

>accessibility tools will parse. Already, you've cut out many of the humans you're trying to reach.

I figure less than 1% of the people you are trying to reach would be cut out by text that some screen readers read incorrectly.

  • akramachamarei 2 hours ago

    Well, yes, that's kinda the whole question of accessibility. The proportion of people who e.g. can't distinguish colors or use the stairs might be a rounding error in some populations. We still find it valuable to strive to serve them. Some of this has been expressed and thus calcified into regulations.

anax32 13 hours ago

Love that page styling.

fluoridation 18 hours ago

Wouldn't a font that shuffled the codepoint-to-glyph assignments be more effective?

  • CodesInChaos 17 hours ago

    That wouldn't make the decoy text look plausible to the AI.

NetOpWibby 6 hours ago

Who's gonna be the one to bring back Flash?

tabarnacle 10 hours ago

Shieldfont approaches the accessibility issue mentioned by not obfuscating text on screen readers.

dombiscoff 5 hours ago

Quite frankly I don't understand why the cat and mouse argument here is meant to serve as a shutdown for anti-AI methods. Whole industries entirely exist in a cat and mouse state (cybersecurity, anti-cheat, etc) and no one in those industries imply that any solution is or can be a permanent dunk. If anything, the very fact that theres no permanent dunk is what leads to such industries developing a competitive service based industry to begin with. Why can't a theoretical anti-AI industry develop into the same thing? These fonts just seem like the infancy steps for such.

dana-s 20 hours ago

I believe the cat and rat game is already there, for multiple places, spam, captchas and now for AI content, yes, it's objective is to make it harder for AI companies to get such data, if it wastes their time, it's a win.

  • rpdillon 19 hours ago

    > it's objective is to make it harder for AI companies to get such data, if it wastes their time, it's a win.

    That's only one half of the equation, though, isn't it? What if it makes it harder for legitimate users as well? It seems there's a balance to be struck.

    • pixl97 14 hours ago

      Kind of funny how you get downvoted for a rational take, but one that's not anti-ai.

      You'd ask that same person how much they like captchas and I'm sure they'd think their a terrible idea and they've ran into all kinds of issues with them.

yieldcrv 11 hours ago

I think this is an example of just catering to the gullible solely because the market exists without pondering anything about the individuals in the market

like the "pink tax", which isn't a tax at all but just a premium on consumer gullibility as the consumer can purchase other products that do the same thing simply marketed in a different way

playing into anti AI sentiment in a useless way fits the criteria

aussieguy1234 9 hours ago

It's probably trivial for AI to work around this either now or in the near future.

1. Screenshot page

2. Parse with OCR

Then train or do whatever else the font was trying to prevent...

ares623 10 hours ago

Look at what they make us give.

grim_io 17 hours ago

DRM for your eyes == garbage idea.

waffletower 19 hours ago

I don't foresee anti-AI fonts being widely adopted. I see them largely as the symbolic saber rattling of intellectual property trolls.

  • KPGv2 18 hours ago

    Every AI font proponent I've seen has been a fanfiction writer who just doesn't like AI stealing their shit to use against them.

  • dgellow 18 hours ago

    Actually, pretty sure it was an art performance

  • wesleywt 1 hour ago

    I don't think its necessary. When everything is AI slop, people most likely will revert to trusted sources.

hellojomp 19 hours ago

We are now in a weird middle ground where we want to write things OCR algorithms have trouble transcribing which also means we write things people with accessibility issues have trouble seeing. No child left behind?

  • mister_mort 19 hours ago

    It's like the old tale about the national park bin with the smartest bear / dumbest tourist crossover, except we're now comparing capabilities of the smartest AI with disabled humans.

    • Xirdus 18 hours ago

      It already was a major issue in 2010 - home desktop-grade OCRs could easily beat an average grandma on reading heavily garbled text.

      • no-name-here 18 hours ago

        > in 2010 - home desktop-grade OCRs could easily beat an average grandma

        I’ve heard Tesseract OCR recommended, but even in 2026, no matter how I scan receipts or documents, the OCR output seems far worse than human reading?

        • Xirdus 17 hours ago

          Well, I used specialized software specifically for solving CAPTCHA on select few file sharing websites. My later experience with general purpose OCR software was just as awful as yours. But it's a product problem, not technology problem - you really don't need an advanced AI for it, you just need a good implementation of traditional shape-matching OCR that doesn't seem to exist anywhere for some reason.

        • strangecasts 17 hours ago

          I think the difficulty was specifically with CAPTCHA challenges, which had to be quick to generate but still legible - OCR on physical documents has to be robust to a different set of problems

          That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?

          [1] https://github.com/PaddlePaddle/PaddleOCR

          • no-name-here 4 hours ago

            Good points.

            >> no matter how I scan receipts or documents, the OCR output seems far worse than human

            > what kind of errors are you getting?

            Just noticeably worse character recognition, particularly where the document is faded, water-stained, or the paper document (not the scan) was low resolution to begin with, as compared to 'normal' human recognition.

            I was largely using Tesseract in conjunction with the self-hosted Paperless-NGX, and I wanted to stay free/local without yet investing in AI-focused hardware. But you're right that AI will continue to advance, including in smaller local models, and it looks like Paperless 3.0 released recently (after my testing earlier in 2026), including with AI functionality.

            > or having the documents smooshed together because the OCR can't parse the layout

            I haven't even been worrying about that yet - I'm just at the point of trying to get the OCR characters right. :-)

unethical_ban 19 hours ago

Is it part of the joke that the site is intentionally over-pixelated while the author critiques readability? (edit: I don't mind esoteric design and I play old games. I found it funny to see a blog with aesthetics that are not optimized for long reading to complain about the readability of fonts)

I don't think any anti-AI font design is more than design-as-art statements against AI. If there is evidence of them being used exclusively for business and without accessible fonts aside them, I'm willing to be wrong.

avazhi 18 hours ago

Not everything that annoys or inconveniences you is harmful, as if this needs to be said to an adult.

bradthebeaverfa 12 hours ago

This author greatly overestimates how much I care about accessibility.

If I have to block a blind person from reading my blog to block an AI from training on it, I'll make that trade every time.

  • Ohentis 12 hours ago

    Well of the fonts listed, 2 out of the 3 would also prevent any human from wanting to read it and the third becomes an ineffective counter measure if it's widely used.

  • hyperadvanced 9 hours ago

    For real. It’s such a weak cop-out of an argument to lead with that I clicked out, never to read this blog again.

  • phoghed 8 hours ago

    I don’t understand how you think a font would block a bot from reading your blog in the first place?

    It’s not even going to render the damn website. If it decides to and detects your retarded font it could just change the font trivially. You’d have to fundamentally fuck up the html text content for it to work at all, and then all you’ll do is inconvenience real people that you’d be extremely lucky to have attracted to your blog in the first place.

  • dombiscoff 4 hours ago

    I think the line of reasoning that a method should provide as much accessibility as possible to not alienate real humans is a good one. You use 'have to' as if there aren't better solutions that could be invented, but such will only occur if we challenge the imperfect solutions we have now.

yipinwong 17 hours ago

Not trynna to be funny. The author's post is anti-human, thus useless and harmful (medically).

I can't read this bad font, sizing, spacing, etc. The main offender is the color choice, and fonts that are just god aweful to read.

  • jotux 17 hours ago

    Found it awful to look at, tried to zoom out and everything on the page got larger. Seriously gross accessibility.

  • dwrodri 17 hours ago

    Accessibility is important, and I find that "Reader Mode" in most browsers is quite good. Everyone should have access to the tools to consume content. Did your browser not provide that functionality?

    • yipinwong 16 hours ago

      basically make it look hard to read for the majority for the sake of 1%? let those 1% use a diff tool view instead of 99% of us having to suffer.

      That website no-way is accessible for my 30 year old eyes.

  • dombiscoff 4 hours ago

    I think he just enjoys terminal styling, man.

kokanee 17 hours ago

I'm a bit frustrated by what seems to be a widespread strong negative reaction to anti-AI fonts. The accessibility problem is real, but I feel like that's a reason to push the investigation deeper for solutions to that problem, not a reason to abandon the effort entirely. The largest intellectual property infringement in the history of the universe is actively unfolding, and it's resulting in an existentially threatening transfer of wealth and power. That's a problem worth exploring every solution for, and solving it may entail some serious sacrifices.

  • Uberzi 17 hours ago

    It's simply that a font won't solve anything. That concept is worst than security by obscurity, as it causes more problem and add more constraints than what it solves... for a very limited time until AI bots are adjusted to decode those fonts properly.

    • waterTanuki 11 hours ago

      how can you make such a bold claim so early with 0 evidence? A successful (and much more realistic outcome) could be to make a font that is financially infeasible for LLMs to parse but easy for humans.

      • SideQuark 1 hour ago

        Yeah, that worked really well for captchas. Now even simple models far surpass humans at recognition.

        So there’s plenty of evidence.

        Now provide evidence it’s technically possible to make any font “ that is financially infeasible for LLMs to parse but easy for humans.”

  • pixl97 14 hours ago

    https://en.wikipedia.org/wiki/Generative_adversarial_network

    Every half assed means of trying to confuse an AI is just a small bit of learning away from making the AI better than you.

    Worse when you have people with disabilities, which I seem to be this week, you just make doing things a pain in the ass.

    What I don't get is people like you think there is a solution to this. There is not. The harder you try you either exclude more actual humans or you align the AI closer to how people actually see.