1. I don't like the sense of futility and powerlessness this advocates for.
2. I'm not sure it is so futile. I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.
That could happen if:
1. There are so many schemes out there the catalog of circumventions gets unwieldy.
2. Doubly so if the schemes allow generation of new obfuscated fonts per site or per page.
3. Then you're forcing the scrapers to pay a greater tax to get your text: spin up a Chrome instance to OCR a screenshot, or spend some a buck or two or LLM credits to reverse engineer the page in order to scrape it.
It might raise the cost for some scrapers, to some degree, but it raises the cost to, effectively, infinity for anyone who needs to use a screen reader, a browser's 'reader' view, or any other assistive technology.
Things like this just remind me of EA's Spore; it released with DRM so draconian that legitimate players were getting locked out of the game during the first week, while people who pirated the game had no problem whatsoever and were playing the game without issue even before its release. The legitimate users of the thing were the only ones punished by the technology designed to stop everyone but them.
This is going to be the same thing; a website which AI will be able to read in short order but which assistive technologies will not be able to read ever.
Oh and lynx/links...
It's not a technical issue, it's a social one, people behave differently given the same set of possibilities and incentives, and we can and should target those who fuck it up for everyone.
By that metric a complete ban on LLM’s might be on the table, which I don’t think is something you’re advocating for here.
That’s dangerous ground because of how you can arbitrarily change weights of different metrics. If the goal is the “overall benefit of humanity” then that’s what’s important not metrics.
How do you propose to asses “overall benefit to humanity” with metrics?
I'm not trying to discount the qualitative approach, but I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.
One would hope for the slightly higher ambition of metrics that are causally connected to "overall benefit of humanity", settling for merely correlated has that unpleasant risk of the association breaking as you try to optimise.
This matters due to the difference between a hypothetical perfect analysis covering every possibility and one than can actually take place in the real world.
Seems oddly sour grapes.
Hosting is not free.
This whole distinction is futile imho. And no, I'm not Sam Altman.
I mean training is one thing, you should honor people's licenses, but browsing and gathering information? Why force me to do it with my biological neural net?
Having principles means applying them uniformly -- even to large entities or those you hate (it's fine if the principles themselves have size bounds in them, though -- versus them being implicitly glued on -- but then you need a universal justification for why that size. Which is possible and valid.)
I think that one should use the best algorithms, the best information, the best knowledge they have access to -- period. I don't like using gimped machines, I don't like making gimped machines, and I certainly don't like being sold them.
So I'm not going to turn around and say LLMs need to be gimped via arbitrary restrictions on their training data.
We have definitions of plagiarism and copyright infringement that apply to people based on what they put out into the world, not how they trained themselves to get there[0]. Artists literally trace art and study specific examples in detail to learn, and that's not a bad thing. Humans can also accidentally plagiarize or make things identical to past works - it's easy to think a great guitar riff just came to you when it was really based on a song you heard years ago that stuck in a part of your mind but you don't even consciously remember the influence, for example.
It would certainly be nice for AI output to provide citations if it realizes it's using a significant chunk of an idea from its training that has a clear source (or many), though this is difficult in the same way it would be difficult for me to cite where I learned about the Towers of Hanoi. The AI frequently does web searches for specific resources to get ideas from these days, which are easy for it to cite.
However, it would have a chilling effect on progress as a whole if we all started jealously guarding our ideas so close to our chest that machines couldn't read them and only a select few humans who passed some kind of gate (or even paid us) were allowed to see them. We would be nowhere close to where we are if we had always had such a mindset.
---
[0] We use proven absence of viewing certain material as a legal shield against copyright infringement e.g. clean room engineering, but this isn't strictly necessary, and possibly even discouraged with modern precedent: https://reactos.org/forum/viewtopic.php?t=21740
Did it work for non-cryptographic DRM? (Broadcast flag, Macrovision, deliberately miswritten floppy sectors, port dongles, physical manual challenge-response...)
Yes?
Did those measures stop piracy completely? No. Did they raise the cost of piracy so there was less of it? Probably.
It's not that simple. What you say may be true, but the reason arms races happen is the alternative is surrender and domination of you or your people. You can't unilaterally choose to not participate without accepting those consequences.
> Screen readers get the real words. A screen reader reading down the page is never handed scrambled text, and our NVDA test asserts exactly that. Screen review and touch exploration are untested. ShieldFont hides shielded passages from accessibility tools by default with aria-hidden="true", because a decoy read aloud is fluent, grammatical, wrong English, and that is worse than silence.
> The real words remain sealed in the same page, and a visible notice above the block carries the control that uncovers them. It is on by default and reachable by mouse, keyboard and screen reader alike. Pressing it sets the reader’s browser to solving a compute-heavy puzzle: JavaScript and a few seconds of processing, more than most mass scrapers are willing to spend. That puts the words within reach of a screen reader, a translator and copy/paste.
I'd love to hear your thoughts on that. Also just aside, love this TUI-esque blog design and color palette (maybe a bearblog theme? still worth an upvote)
They also actively block copying the text, telling you to "uncover" the text first. The uncover operation is VERY expensive.
Anyone using assistive technologies or trying to copy "protected" text is SOL.
Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
It's only search engines? And who uses those anymore anyway? Other bots?
It doesn't affect organic traffic, so you are not really putting yourself into a black hole. A lot of website are driven by social media and organic traffic, so they would be just fine with this approach.
I take your point that one extra click/interaction required for screen readers only is objectively more friction, but it's only by exploring these technologies instead of dismissing them that we'll arrive at UX solutions truly work for people of all stripes (and ideally, not for bots, scrapers, and the like). I think if you're applying a tool like this you already have very different priorities than fueling Google.
The problem is... it's a legal issue we cannot solve. America and China are large enough to not give a fuck about what everyone else wants.
Not sure if that's due to ublock or one of the font settings in ff.
But you seem to appreciate it, and I'm sure others do too. Different strokes for different folks.
“Claude: make the scraper mimic a screen reader.”
And just like that, in 10 seconds, their site feeds my “screen reader” the real words.
Their stance appears to be "[t]his is not an anti-AI font. AI is in our lives. We are actually a pro-consent font." (https://github.com/isaqueseneda/shieldfont/issues/2) Might be a bit more niche market.
Setting browser. display. use_document_colors or browser.display.document_color_use to 0 fixes things.
It does break some stuff, like voting buttons on hn. Im happy with the tradeoff.
The entire point of the situation is that permission is not involved. They just do it. Meanwhile, if I do it to them, I am fined/sent to prison/executed. Until such a time that this baseline scenario of inequality is somehow remedied, there will be a motivation to stop them.
and on news and podcasts he wound up correcting so many dumb arguments that he sounded pro-bitcoin and could never get to his own points
"well, no, not like that, the difficulty algorithm...."
"there are ways to use it with the power off"
"well, no, the transaction fees supplant the block reward so ..."
People aren't particularly bright. That's why the scientific method was developed to counteract our built-in tendency for... Unorthodox approaches
Illegible fonts will only create accessibility problems for humans, even ones with normal vision, high literacy and no cognitive defects (dyslexia), while AI will blow right through the text.
If this is done in electronic documents, where the AI won't even see the glyps becaue it's reading the underlying character codes, it's even stupider.
I can't believe anyone would even try this (and then believe it is working without putting their hypotheses to the test).
"This document contains mitigations against review by automated systems. Recipients should ensure that they have read the contents on screen or in print. Recipients with bona fide vision impairments may be entitled to unmitigated documents upon request."
In testing, obfuscating small portions of text slipped under the radar of most (then-)frontier LLMs.
We used a font that was rendered on the fly and reported faulty or fake Unicode mappings: https://tritium.legal/blog/noroboto but others have proposed and done the same with ligatures.
Accessibility is not accessible if you need to go through extra steps to get it.
Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.
Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.
Our internal, hypothetical use-case was between contracting parties who were looking to avoid terms escaping into the wild. This shouldn't show up in standard ToS or similar. There are already really good legal reasons for that.
Look, I understand, you don't have to explain to me the theory for how that happens, I know it already. Since I know your a smart guy, to some extent you care about that only because you imagine that clients do. But in reality, in the real world, every email you send is read by at least two people, every contract you sign has multiple parties, etc. You make some obfuscated thing or whatever, but eventually, someone has the real text of the document - it might be YOUR client, it might be the person you are negotiating with, and you rarely represent ALL the sides. You never own ALL the information and all the parties and IT systems in totality. Eventually someone will put the text into a chatbot. Or maybe they put a salient piece of the pre-final text, like some legal theory or merely a question, into the chatbot.
So I see this font stuff, or watermarking stuff, or all this provenance and control stuff, as deeply illusory. It is the worst circlejerky kind of aesthetic experience making. When you mess with anti AI fonts you are trying to compete in the same business DocuSign is in, that is, in the business of selling holistic social experiences - a whole 7000 person company whose main competition is a fucking pen - but it's not like you're doing something creative. If you care about aesthetic experiences, write a short story! Are you getting it? The itch you are scratching with this weird thing, nobody wants.
That seems a bit entitled to me, especially in the story you shared. They had access to the school just like everyone else. Why does it have to be in the exact same spot? Surely they could be dropped off by the loading dock as easily as others could be dropped off out front, and maybe even moreso. Should the wheelchair accessible stall be the first stall in the bathroom so they don't have to go through "extra steps" to get to it? As long as the ramp had the proper gradation to satisfy the ADA, I don't see the problem.
They literally did not. They had access to the school from a different entrance than everybody else. Specifically, one intended for cargo before people.
It may not be sufficient, and we should criticise when it isn't sufficient, but we must in practical terms also accept tradeoffs in a scarcity economy. In the worst case, some of those tradeoffs accommodate one disability over another disability.
If the dean had a swanky office and the wheelchair kid must roll past the bins, then the school made the wrong choice. But if it's a ramp at the front or textbooks, then we must think a bit harder.
Otherwise, you are relying on obscurity, and will lose as soon as you become interesting enough to matter.
You will also break non-AI machine processing use cases. That isn't just accessibility, it is things like search.
Cool.
This is self-contradictory. Either you want to critique something or you do not. Make up your mind.
there is a reason nobody uses text based captchas anymore.
Their own whitepaper brings up a bigger issue: if it works, it poisons search engines as well.
Also more expensive doesn't mean that something won't happen, only the dynamics of how it happens. For example if you put all your documents in images then some service might just sell the AI providers the text. That service may do underhanded things like bundle OCR in an app that does something else and use your phone to get the text out of these images all day.
And this will be the end of the internet as we had it, and as we have it right now. With no way to block AI scrapers and no hope of AI companies acting ethically, shutting down the free internet is the only end result.
I'll remember these fun, bespoke attempts to stop AI from ruining the world once the internet is gone. That alone makes them worth existing for me.
Well, bad news, you aren’t going to get to keep that either. Whether for high-minded reasons like “the knowledge that everything you create will be fed into the slop machine so it can later be regurgitated without attribution is having a deleterious effect on the morale of some contributors”, or for incredibly mundane reasons like “lacking the resources to either serve or block the crawlers”, I fear the web you love is dying of ai with or without anti-ai fonts.
like the "pink tax", which isn't a tax at all but just a premium on consumer gullibility as the consumer can purchase other products that do the same thing simply marketed in a different way
playing into anti AI sentiment in a useless way fits the criteria
That's only one half of the equation, though, isn't it? What if it makes it harder for legitimate users as well? It seems there's a balance to be struck.
You'd ask that same person how much they like captchas and I'm sure they'd think their a terrible idea and they've ran into all kinds of issues with them.
AI labs are broadly within the law and would like to not become criminal organisations, because money laundering is expensive.
I figure less than 1% of the people you are trying to reach would be cut out by text that some screen readers read incorrectly.
1. Screenshot page
2. Parse with OCR
Then train or do whatever else the font was trying to prevent...
We will probably have to wait Europe and China to do this since the US executive , legislative and legal systems have all been bought by Ames oligarchs .
I’ve heard Tesseract OCR recommended, but even in 2026, no matter how I scan receipts or documents, the OCR output seems far worse than human reading?
That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?
>> no matter how I scan receipts or documents, the OCR output seems far worse than human
> what kind of errors are you getting?
Just noticeably worse character recognition, particularly where the document is faded, water-stained, or the paper document (not the scan) was low resolution to begin with, as compared to 'normal' human recognition.
I was largely using Tesseract in conjunction with the self-hosted Paperless-NGX, and I wanted to stay free/local without yet investing in AI-focused hardware. But you're right that AI will continue to advance, including in smaller local models, and it looks like Paperless 3.0 released recently (after my testing earlier in 2026), including with AI functionality.
> or having the documents smooshed together because the OCR can't parse the layout
I haven't even been worrying about that yet - I'm just at the point of trying to get the OCR characters right. :-)
I don't think any anti-AI font design is more than design-as-art statements against AI. If there is evidence of them being used exclusively for business and without accessible fonts aside them, I'm willing to be wrong.
I can't read this bad font, sizing, spacing, etc. The main offender is the color choice, and fonts that are just god aweful to read.
That website no-way is accessible for my 30 year old eyes.
So there’s plenty of evidence.
Now provide evidence it’s technically possible to make any font “ that is financially infeasible for LLMs to parse but easy for humans.”
I suppose Uberzi has a confidence level in their claim, but it's tedious to hedge every damn thing.
Every half assed means of trying to confuse an AI is just a small bit of learning away from making the AI better than you.
Worse when you have people with disabilities, which I seem to be this week, you just make doing things a pain in the ass.
What I don't get is people like you think there is a solution to this. There is not. The harder you try you either exclude more actual humans or you align the AI closer to how people actually see.
If I have to block a blind person from reading my blog to block an AI from training on it, I'll make that trade every time.
It’s not even going to render the damn website. If it decides to and detects your retarded font it could just change the font trivially. You’d have to fundamentally fuck up the html text content for it to work at all, and then all you’ll do is inconvenience real people that you’d be extremely lucky to have attracted to your blog in the first place.