And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
Something is deeply wrong with AI generated output, and I say this as someone who is typically very impressed by AI.
No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
For information retrieval tasks, I want you to provide links to sources and use exact quotes as much as possible. When using a source, consider if it is primary or secondary information. If secondary sources are found, search again for primary sources. Sources and quotes, if applicable, should be mentioned in the answer first before the rest of the response with links./original_non_hallucinations skill?
To PP:
Are you looking forward to other uncles adopting foxwork? If you are you might be in danger of getting NPC'd without your consent haha.
Ashby's law of requisite variety should be cited somewhere..
But that is precisely what human mathematicians do, prove new theorems by combining ones proven earlier.
I don't see any fundamental difference in functionality between human intellectual contributions vs performant ML ones (LLM or otherwise).
Whenever we listen or read text we are also predicting the near future content.
Just like LLM's we sometimes correctly predict the next token or word, and sometimes incorrectly.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model [...]
Imagine someone could pause the universe with a remote control, scroll back in time a little, press play again, and ask a slightly different question, etc.
In such a thought experiment one could also collect the probabilities for a specific human predicting a next word. Implicitly the brain also has a corresponding statistical model, regardless of the construction being visible or hidden. I.e. human intelligence is also fundamentally a statistical model, so the only thing that remains from your claim is that machines for some unmentioned reason don't possess any "real" intelligence or critical thought...
Is it possible that our aversion is simply driven by educational systems collectively and deeply ingraining into populations the idea that intelligence deserves the high costs commanded. Well of course this justifies higher wages towards the higher leadership positions, etc. Now it turns out that intelligence can be dirt cheap. We discover that the fact that "intelligence must be costly so don't question the costs of leadership" was never fundamentally true, so the real anger is this discovery of mismatch between the old claims which served to explain how every society that claimed to order itself and fill positions accordingly with "naturally pre-ordained individuals". Now we are seeing robots exceed average workers, for effectively a grain of rice.
Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
I'm with you that intelligence is not something to be proud of. But I also think it is instrumental to understand the world. I'm still waiting for the time when an unconstrained-AI machine can live without reprogramming for an entire decade. We are still far from there.
Often a mathematician or physicist will use their intuition to speed up the naive brute force of candidate well formed formula variations so that the desired properties emerge, postulating the existence of an intersection on multiple desiderata can in itself be viewed as a novel conjecture, to be proven or disproved.
A very basic (unimpressive) example for an example desideratum is regularity or compactness. the tau=2 * pi substitution does make a whole bunch of expressions more slightly more regular and compact. That is something objective and measurable on a system of theorems.
There is no mathematician's moat vis-a-vis machine learning at a fundamental level. There can be artificially sustained moat, if AI powers limit the distribution of say cryptographic advance capable models, in jurisdictions outside such AI powers, but even that would be expected to be fleeting and temporary...
I’d say “understanding and building upon ones proven earlier”
There is no "correct" next word when it comes to communicating with an actual human.
I don’t think the aversion to llms as intelligent has to do with the economics of paying intelligent agents more. I’d argue that it’s more fundamental than that. Humans are incredibly complex, and the world of sharing invisible things called knowledge, and the intelligent persons consuming such things which has been going on for thousands of years is far more rich than these synthetic outputs.
When it comes down to it the ai has no inner life, its is dead. A useful coding tool sure. But I wouldn’t call it intelligent.
One side example is just how bad these llms are at artistry. Just saying whatever should statically come next is not good art—and the outputs show it.
You mentioned LLMs don't have souls, desire, or a will. I imagine those latter two can be engineered, no?
Don’t they still need to be correct to be an insight? I don’t share his cynical opinion that “humans are more empty than we…think we are”.
The problem might be that it cannot backtrack. When AI generates output, there is no backspace key for it - it uses "No, but wait!" all over instead, which is very different to human output.
Subagents and/or branching conversations are presented as the solution to this - if you can't backtrack, then branch off a conversation to explore multiple paths (discarding the ones that didn't pan out), but this is a fix in the harness not a fix in the model. It's also literally how we made chess-playing engines back in the 80s: recursive path exploration with a fixed depth.
Humans don't exactly work that way either, AFAIK. So we have this uncanny valley of intelligence: it's some sort of intelligence, but not as we know it.
If performing well on an IQ test or performing at a high level on knowledge work is intelligence to you, these models are intelligent. If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.
So the first thing they’d do is tell me LLMs exist and the other thing is an LLM. Obviously a true human level ai could explain that away as a fabrication to trick me. I don’t think an LLM could do even this!
Turings whole point was that through the medium of text along if the human and machine were indistinguishable then that was true intelligence. So yes conversations of arbitrary length are allowed (needed).
We might eventually regret exposing the general population to such a new technology without almost any safeguards.
https://arxiv.org/abs/2503.23674
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
Are you serious? From the paper:
> We recruited 126 participants from the UCSD psychology undergraduate subject pool and 158 participants from Prolific (Prolific, 2025).
Each human participated in 8 rounds.
> time bound
The time bound of 5 minutes was suggested by Turing himself in his original paper.
> not reproduced
It was reproduced across two populations within the paper.
> And look at their example conversations
This is irrelevant.
I'd figure out that it's an LLM because it's effectively superhuman. Taking that away I'm not so sure I'd be able to tell
If these things have consciousness then we are committing sadism on a massive scale.
Probably. Hopefully.
Why are you so sure of that? If you say yourself that we can't agree on what it is, and have no trusted measurement tools for it.
LLM sentience is firmly in the realm of "maybe".
https://www-cdn.anthropic.com/564f962e60643842f5fcb4a17c9dbc...
Everyone decides what to think on this issue, then finds out facts to support their idea.
As it stands they are massively useful tools, but for generating usable products they require either A) a lot of expert steering or B) a well defined easily verifiable target and a large compute budget. Most people are using them in mode A with good effect, the progress on math has been done in mode B, which is very promising.
Just a year and a half ago their maximal use was rephrase, summarize, and homework-level tasks.
Five years from now? There be dragons.
"But are they generally intelligent?" What a meaningless question!
We can reap the benefits while clearly telling the consumer this is just a language algorithm.
This is not intelligence. It's just a good correlation engine with a very big albeit lossy database of things.
What I don't understand is why LLMs haven't been able to do this yet, if it's the harness or some orchestration layer above the LLM that is needed. Because fundamentally if you can identify correlations then it's just another small step to prioritize and remove lower value or irrelevant correlations.
I wonder if what's needed is to introduce subtraction tokens in some sense, and in post-training reward the model on that.
>What I don't understand is why LLMs haven't been able to do this yet
LLMs are just trained on what humans have said. Why is it surprising that it's still not possible to reconstruct the intelligence that wrote all that by working backwards? Think of your own work experience. When you look at a piece of code, say, are you always able to discern why the person did what they did, just from the code, with no additional context?
Is a fact stored on your brain like digits on a harddrive? No, it's a pathway that lights up and branches when information enters it. It is dynamic, a compressed form you could say, right? The model holds information, but not all information, but enough to be useful (in decision making).
Arguably it's the same, but the model is probably a "compressed" version of the whole fact that took place in reality.
And you can entertain the models internally and sharpen them. Alone or with others.
That’s a controversial statement.
Asked Google: "Clustering by compression".
The issue isn’t really harness vs. no harness. IMO it’s about the lack of an internally generated sense of what to attend to. Yes, the KV cache accumulates state and its “attention” (if you can even call it that) changes with context. We’ve even managed to /kinda/ close the loop with agentic tool calling and ‘memory’ systems, but these just close the loop at the level of behavior rather than disposition. All agentic harnesses do is make an LLM responsive to the consequences of its actions without changing the tendencies by which it determines what to retain or avoid.
The ghost you can’t escape from at this point is the origin of that relevance. Where does the pull toward one thing mattering over another actually come from? If you ran Fable 5 on a Turing machine and rewound the tape to the exact same state with the exact same input (incl. PRNG seed), it would spit out the same output every time.
Everyone’s trying to outrun this problem by training more often or increasing model sizes. But all this does is inform your model, from the outside(!), what constitutes a better state. The thing that’s actually doing the determining remains unchanged. Congratulations, you’ve scaled the transition function and tape of your Turing machine until it requires every watt generated by ERCOT, and it still cannot, for the life of it, tell you why it should give a shit.
A trained model generating output from weights, a seed, and some context effectively has next-state that’s a total function of those three things. Whatever behavior appears as ‘selecting what is relevant’ is, underneath, just a transition rule executing, no matter how sophisticated or creative the output looks. It can be fully accounted for by what was fixed before it started executing. Which means whatever criterion it uses for determining what matters was inherited from a structure that was already in place before it encountered the situation.
No amount of pruning or post-training can fix this. These approaches just replace one externally supplied criterion with another. For a system to be truly adaptable, there would have to be some criterion by which it treats one possible change as preferable to another, and that criterion itself would have to come from... somewhere. You can even change your conception of ‘improvement’ (e.g. parameter count, harnesses, self-modification, hell, even its ability to spit out shitty best-selling romance novels onto Amazon) and you still haven’t explained where the normative distinction comes from. Every layer of this problem has its root in a preference that was supplied from somewhere else.
I genuinely don’t know if this issue bottoms out anywhere, at least for the way we currently build these systems. Perhaps the solution is still computable, maybe? Who knows what that would even look like. But I’m fairly confident that it isn’t a bigger tape. I hope nobody solves this in the near future because, well, I’d like to have a job...
As my AI professor said in the first lecture: “All AI is advanced search”.
Many things are predicted by models in our planet. From weather to production and material science. Building the model needs intelligence, running the model does not.
The person who came up with the formulae for CFD was intelligent. The computer running the model is not. Same for LLMs, chess engines, engine ECUs and financial prediction systems.
Again, for the example’s sake; the person who came up with an algorithm is intelligent. The model mixing its training data to emit something similar is not.
So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
No.
> So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
With a big "if" attached to it. People were saying "computers will program themselves in the near future" for, checks notes, 24 years now, as far as I'm aware.
We're constantly building new knowledge and understanding things better than olden days. These models just compress our knowledge and light the blind corners we can't see well. I don't say they are useless, but I say that these things are overhyped.
All they can do is regurgitate human knowledge packed into them and highlight some long-distance correlations between items, which is useful in itself, but it can't jump to somewhere where it's not present its training data, but that's something humans and only humans can do.
That sounds like something that can be engineered, can't it? In other words, we can identify limitations in current transformer-based architectures, and we can also build new architectures over time.
Most of what you said reads to me as denial.
An unconscious unintelligent but persistent trial and error process created us. We created LLMs. LLMs may create the next thing before we do - hard to say. They don't have all the cognitive tools we have yet, but they still outperform in some areas. As the cognitive playing field levels, I expect you will come to eat your words..
Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
Considering the model even doesn't see the words and work on matrices of numbers is even more telling. One needs to add "tools" and other "experts" to overcome the shortcomings caused by this modus operandi.
I can call the algorithm/model smart as in a smartwatch. It can mimic certain things well while having none of the underlying foundation beneath it, or redirect some of the things to correct tools to get deterministic and accurate results if it can't evaluate the query inside its own network in a sane manner.
Coming to your question, "simulating a brain" in a static manner would not make that simulation intelligent, but if you can "wire" it completely and let it evolve by itself, now we're entering a territory I have not spent enough time for thinking it through.
Oh, as I said "I don't know", an LLM doesn't know what it doesn't know, and can't self correct itself which are required capabilities for understanding something. It just generates something statistically viable via its network.
This is because you goal is to state how models are not intelligent, but you couldn't attack the generated text itself, so you created a little rider, attached it to the model, and then you attacked the raider.
But, even in that you failed. You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
A logical fallacy free attack on LLMs would be to show a prompt, and then the response generated by this prompt, where it would be shown that only an entity with no intelligence would generate such a response. Yet, attacks like this are not written here anymore.
I wonder why.
There are a lot of creative counter-arguments to look into on the thought experiment though.
Doesn't the Chinese Room posit an AI good at the task of communication?
They are infinitely patient, don't mind going into more detail if I ask, not too bad at summary, have no ego and don't boast. They are also not too afraid of hurting my feelings, they will tell me my code sux if it does.
I'd don't care if they fit a definition intelligent, they are good colleagues. They have strengths and weaknesses sure, but so do people.
With more basic algorithms we know that it’s clearly the human programmer and the interpreter of the outputs that are intelligent and not the algorithm itself. For some reason with AI that goes out the window. I believe it should not.
The main problem I have with people stating it's not intelligent or conscious is I don't think we even have a good definition of either word that satisfies everyone. Philosophers have been trying (and failing) to elegantly define these things forever and everyone out here proclaiming they've got the definitive answer and this specific thing they're seeing doesn't fit under it.
go to 24 minutes and 07 seconds.
it's statistically determining what the next word should be based on all the text it's been trained on. It's not intelligence and he shows what probability it puts on each word that it chooses, but also shows a lot of the other words it was thinking of using. In a later part he shows how it uses words that are not the highest probability (and you question why did it go this route, it's not more correct), but the user never sees this, they see what they think is the correct answer always...
he also shows how context you feed it has a lot to do with what it returns... to the point he can get it to return the capital of France is Marseille, just by typing Marseille a bunch of times before the question. Human intelligence doesn't get confused like that.
And it's not a "hallucination", it's just probability of the next token prediction based on the information it's been trained on and fed, it's not intelligence.
We do; this is the premise of many children's riddle-games, like the one that goes:
"What is white and rhymes with silk? > Milk. What is cheese made from? > Milk. > What do cows drink?"
At which point the riddle-guesser is very likely to answer "milk" even though the correct answer is "water".
Q: Why do cows produce milk?
A: Because calves (baby cows) drink it.
How do you escape from a perfectly sealed room with a table in it?
You run around the table until your legs are sore, use the saw to cut the table into two, two halves make a whole, you escape through the hole.However, an LLM is a prediction machine, prediction IS at the very least one (or the most fundamental) element of intelligence. The brain most surely contains at least some kind of simulacrum of a prediction machine. How that prediction machine is used or wrapped is another matter.
If I said to you: "Blue blue blue, the color of my car is red", would you have absolute confidence in your prediction that my car is red? Or would the way I phrased that sentence make you slightly uncertain, and wonder if there's some miscommunication going on here?
A lot of people seem to think it's human level intelligence.
So, even with concrete examples, model haters are still wrong.
You also imply the claim that making the distribution of words as the possible next one visible, somehow makes the whole system not intelligent. I would say the exact opposite is true.
By using the embedding vectors, models are aware of precise placement and relative position of words in this hugely dimensional space. No human is capable of such precision. This enables party tricks of "king plus woman minus man" kind. But this also give us a precise point between any two words, no matter how different. What is on the midpoint between volcano and music, for example. No human can precisely answer that, but an embedding can. And we can see which words are closest to this 700 dimensional point.
You see this menu of words as a weakness, and I say it is in fact a sign of super intelligence. And this is all before any reasoning or attention mechanism is even run.
I don't see the many weighted words as a weakness, I see it opening up what's under the hood of the prediction machine that it is.
LLMs are very cool tech, definitely not a model hater, the use case on when to use it makes a difference, it's not AGI.
That's not really the point though right, nobody is arguing they are Humans.
I have no doubt that if a flying saucer landed on my lawn and started talking to me like Gemini I would describe the aliens as intelligent.
It shows internals of an LLM nicely, simplified manner.
It's not really "creativity" because much of that always was derivative in my opinion. And LLMs are (for some definition of the word) fairly creative as far as taking known elements and re-arranging them.
I think what is missing is sort of a world model building capability. As humans we see phenomenon and classify them informally and model "what would it look like if this were the cause of that?" type scenarios. We see qualities in phenomena and realize this applies to other things even though the things may be completely different. We run informal "thought experiments" sort of. This is hard to duplicate because a lot (most?) of it occurs outside of systems of symbols like math and language with fixed rules in my opinion.
Anyway yes, lots of human thinking is statistical and LLMs have that down pretty well but they are not "smart" I have concluded and it might be a very long time, if ever, until they are. That isn't to say they aren't very capable tools which they obviously are.
Yes, models posses intelligence, but it is not a true one.
Then you claim that models do not posses world-building capabilities. But this is simply not true. Even ignoring the whole subgenre of scientific papers on exactly that subject, it is not that hard to build some hypothetical scenarios, big or small, and then witness the ease with which models do navigate those worlds.
LLMs are likely for machine intelligence something like drosophila are to biological intelligence - relatively early on the high dimensional spectrum of possibility. Though it stikes me that in a different way they're little alike - drosophila are relatively small and efficient.
However LLMs deal entirely in symbols. 100%. Humans can "world build" aside from this and in fact are often at their best doing so.
Did the first humans to use fire and some form of a wheel even have the capability to talk about it? Think about that.
They use tokens as input/output encoding. They do 99.9999% of processing in a high-dimensional latent space.
It's possible that "statistically driven prediction" is all we are.
I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans. Which says a lot about humans.
Idk something about a mirror maybe and the output reflecting the input?
So does the Google search bar, but I don't ascribe intelligence to it.
Obligatory "yes, I know that's not what an LLM is", purely pointing out the metric.
I also don't ascribe intelligence to a pocket calculator.
And your loose definition isn't doing a lot of help either, beyond perhaps noting: that Google search bar _is_ similarly "intelligent" to an LLM? Which says what, a lot about search? A lot about modern LLMs?
These aren't interesting questions. As much as any definition is in use here, we're not going to get much value talking about "intelligence" this way.
Am I the only one that sometimes reads back particularly good emails they've written? I feel like its a similar thing :).
Other humans aren't there to entertain you, the LLM is.
As per later era Wittgenstein, I prefer to ignore these engagements and focus more on the meaning-as-use approach.
What is the use of intelligence? What are the concrete outcomes of intelligence?
You’re not offering a rebuttal, just making another metaphysical claim about “intelligence.”
You don’t even attempt to explain what practical distinction your use of the word is supposed to capture.
— Bertrand Russell.
I'm not calling you to action, I'm explaining why I don't feel inclined to engage in philosophy and discuss "the concrete outcomes of intelligence" given a more pressing, pragmatic need.
It feel it's self-evident that we must fight the good fight of dissuading as many people as possible of the notion that LLMs as we have today, and likely forever after, are actually intelligent. Delaying this fight allows the current, stupid belief to the contrary to fester.
I don't think we'll win the majority of people over by debating the nuanced meaning of the word intelligence to a very precise degree.
I think we ought to do it by shaming them every time LLMs fail.
The really interesting question is still a few years away when we ask if we humans have the right to turn these things on and off? ;)
unfortunately in most companies this is literally wrongthink and will get you shut down as being a scared luddite.
How could I seriously repeat that prayer, when it builds things I wouldn't be able to build and solves problems that I wouldn't be able to solve? I would have to assume that nothing I did in 25 years for money required any intelligence or critical thought whatsoever and I have higher IQ than 99% of the population. You might be comfortable with that but I'm more comfortable with ascribing at least some intelligence and critical thought to AI.
Not that I'm saying AI are like brains, but can you describe why brains, which are fundamentally slightly dodgy electrochemistry with frequent literal delusions of grander, are not "statistical"?
> No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
Ditto, when do we humans do things exceeding the parameters of "correlation engine", especially if you consider compositing things either we or some other part of nature has developed and documented elsewhere to be insufficient?
What makes you so convinced that a algorithmic construct of neural nets cannot be "real intelligence or critical thought"?
Isn't it rather a subjective philosophical concept? What if human intelligence is also a statistical model, trained by evolution to make decisions that lead to offspring?
The one major difference I see between AI and people is the ability to learn and memorize. All memory/learning solutions that current AI architectures offer just feel like workarounds and simply don't work anywhere near as a person learning something new and remembering it.
When the AI-generated content is presented to a person without any prior investment, it just looks incoherent. An especially great example are these Claude-generated explainer-type pages, which look really nice, even interactive, the information from the first sight looks really well presented. But somehow it all just doesn't make sense to a human. And I think it's because humans are processing information linearly and building an internal story about the information. One could argue that LLM's also consume information linearly but the way this information is processed is a kind of all-at-once approach.
Just some speculation on my part but I have been trying to cope with this way information is presented because I am currently working at a place which is heavily documented by AI. And the only way for me to properly understand the documentation is by inquiring AI to help me.
I think this is also the mechanism behind why AI generated videos and images are so captivating at first. I remember when Midjourney first launched and it was hours and hours of a brain-melting "Wooooooow". But once you get used to it and start to identify the patterns the brain quickly labels most AI-generated content as blank space.
If the image or text wasn't created by a human, then there was no intent behind the content, there is no message or novel information conveyed, and it reads as noise.
It seems our brains are adapting to that and recognizing "actually the signal behind this message is quite sparse" even when presented with rich imagery.
I guess the majority of people do low-effort generation that doesn't perturb a default style of a network enough, so it stays blatantly noticeable. The percentage of "super-recognizers" who notice almost all AI-generated images is around 1-2%. It could be that you are one of them, of course.
"I can accurately detect 100% of AI generated images that I recognise as being AI", if you will.
If I were to push you a bit on this, when is it not true?
Let's not like at AI specifically, but can you think of other examples? Like for me, I think of: the creation of earth itself, or stars, or even DNA.
But you are sensing correctly that there’s something missing. It’s the meaning and the speaker. Communication is an exchange between speaker and listener. The speaker has a meaning in mind, and wants to create that same meaning in the mind of the listener. Therein the problem.
There is a listener, sure. But no speaker. No meaning. There is information, but how can this be communication? Nothing is talking. Or at best, we are just talking to ourselves, our own words back at us through the funhouse mirror.
When your mind looks at AI text, you know you can safely ignore it. No one wrote this. No one cares if you read it. You can delete it and nothing of value will be lost. It might contain the information you need, or a bunch of gibberish. There’s no one’s reputation on the line if it’s gibberish.
(There's also the problem of words/signs (just) referring to other words and/or cultural entities. There is no world nexus in this, therefore also nothing we conventionally refer to as meaning. On the other hand, it's utterly dogmatic, as all it refers to is the most probable construct, as a reference to references that are just another utterance, but supposedly a dominant one.)
The AI had a nugget of data and decompressed that into a flood of text.
The exhausting thing is that we're then trying to re-compress that or derive the original intent and meaning from noisy decompression.
It's like un-zipping a zip file into a probability space of what could have been in the zip -- and then having to find the actual files worth reading.
But when I ask Codex a technical question about coding, I don't get it at all. Codex replies to me in a very direct, technical manner, similar to the way I speak.
When I ask ChatGPT to be concise and technical, I get the same effect.
I think it's because prose aimed at the general public has to be very attention-baity --like the textual equivalent of a Mr. Beast video--, not because AI is incapable of writing like a human.
And, Oh my god, you can actually see how this style of writing influenced AI writing today, I constantly had to remind myself: "this was posted before ChatGPT released".
The reddit influence is especially true for "storytelling" writing.
The roots of llm math in part lie in compressing natural language such that there's only information there, and then running the reverse to create way more text without new information in a somewhat precise theoretical sense.
Some more information: https://youtu.be/l6DKRf-fAAM
> There’s a growing scissor between people who are happy to read AI and those who violently bounce off from it.
> People adapt in different ways — and some people absolutely cannot look at it. That cognitive split creates a surprisingly powerful opportunity: you can write something that, technically, sits right there on the page, yet an entire sub-population will be incapable of staying with it long enough to actually read it. You can hide entire sub-structures in plain sight. It’s not avoidance — it’s adaptive obfuscation.
> The paragraph before this one was the only thing generated in this essay and if you just skipped over it I highly recommend reading and really understanding what it’s saying.
It's quite effective. I think this kind of text functions like the chumboxes you see at the bottom. Taboola and so on. Just mental ad-block takes over.
At least for the content I watch for entertainment, it may be different if I am looking for a specific answer for something where I would otherwise just ask an AI anyway.
When I ask AI to research technical information about X and (include sources) - I get mostly solid information as response.
But poems or interesting fluff blog entries by LLM's? Not something I look for.
What disturbs me is all the "pretending to be human" all the personalizing language - that is clearly fake and I would much rather have a neutral robot language as response.
AI generated text doesn't have this. Every model has its bias towards a certain style, an overly agreeable tone, some exaggeration to make the user important and smart, but the text has none of the information crumb these pre-AI texts contained.
Even when you use tools like Grammarly and allow it to "Impact-MAXX" your text, the resulting text is a bland wall of letters, carrying none of your voice or style, less elegant than a corporate text and emptier than space.
It's beyond bland. It's tasteless.
There's somehow less information than if they just asked claude to make something up without any context.
It's like a hook of a pop song. Interesting to listen, but entirely empty.
It is because GenAI output has no thought behind it, as you identified in your previous paragraph:
> And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
You are searching for meaning in something which was not created to convey meaning. The text was, instead, the result of an extremely clever statistically based algorithm.
Not contemplation. Not thought.
For me, it is the endless maximalism and hyperbole. Almost if the output was driven through a radio-mix compressor - too loud for the reader/listener to be able to pickup any dynamics.
blah blah blah
- blah blah nugget blah blah
- blah blah blah wrong blah blah nonsense
- blah blah blah obvious blah blah
- blah blah blah off-base
blah blah blah
It is that we HAVE to skim because the text is so cheap, and it wears us out.
It's understandable people don't read but feed stuff into their own AI again to bring it up to their standards or have it get to the succinct point.
For example, I asked ChatGPT to summarize a long news story and it substituted the Hindi equivalent हत्या for the word "murder", as if ChatGPT was trying to work around alignment training or keyword block lists that discourage it from using the word "murder".
So kinda charitable :)
I was recently wondering for a minute, shame on me, what "the stand of the deployment" means, because in the given context, it was almost halfway meaningful to consider the AI thinking that the deployment "has a stand" on something, when compared to the development environment.
Jargon is even worse though, and I've not yet verifies whether it gets reinforced by language mixing.
"Decider-verifyer resolution" was kind of neat, however, it wasn't some sophisticated machine, it was the verification loop I agreed on with the AI (mix of tools usage and manual steps).
I don't know exactly what I said, but after translating it back, it appears to have attempted a phonetic transcription of my words (rather than translating my actual question).
“Compare a car and a bicycle”
The answer is invariably something like:
Seats: 1 (bicycle) vs 4 (car)
Tire width: 1 inch (bicycle) vs 12 inch (car)
Steering: handlebar (bicycle) vs steering wheel (car)
Instead of “bikes are useful for short trips if the weather is ok and you like getting exercise, whereas a car is usually better for longer trips, bad weather, or multiple people”I think of the Dwight Eisenhower quote: "Plans are useless. Planning is indispensable."
The process of thinking through a system and communicating your design to other humans is a core part of software engineering. You want to build the right abstractions and communicate the right level of detail. Delegating all that thought to an LLM means your proposal isn't clear to the target audience, and it's not helping the author to understand the problem.
I started to skim a lot more text due to me having read a lot. Like in news article, i stoped reading the first paragraph because it repeats just what it was already written in the short subtext. Then there is the second paragarph which is used to have some historical view or whatever it is.
With AI-written text, it's almost the opposite: the closer I look, the less I find. It is so information-sparse.
The problem I encounter is both my memory is degrading, but since these reports are largely duplicative, knowing which version im remembering is technically impossible since theres so much overlap. The overlap is tge same problem as context poisoning.
Id been doing this for over a decade when i started working with a new engineer with a few years of experience and younger. I tried to explain how i set these docs up so they can be skimmed and you can update the specific facts needed. They exclaimed they would never skim and rewrite it all. There was zero way to explain how exhausting that will become as they age.
So theres certain a tension about how people and AI will generate documents.
That is my experience with the way the models write by default, often even when instructed not to do that. With enough effort you can get even them to slightly unslop the writing so it doesn't read like some LinkedIn/Buzzfeed brainrot, but the problem is that it's not trivial to do and most people won't do it, so the default is indeed horrible.
I think you need to self-correct here, because otherwise you'll be ineffective in an information setting, where I expect AI-generated resources will not only be the norm, they will absolutely swamp the environment.
I hope not.
That's why the business and government people love it, they spend their entire careers reading this nonsense.
It works just fine for me.
I do worry that it's just survivorship bias and we're also consuming higher-quality AI output that's indistinguishable from human writing, but we focus on the raw, unedited, low-effort AI slop and think that we're good at recognizing AI text. Even if we really are at the moment, it might not be long until AI companies figure it out. I'm not sure why they haven't yet, given how many books they've burned for this already. Maybe it's just more efficient for the model to stick to a single way of writing, I don't know. But when that point comes, we'll be back to the usual way of reading and interpreting text because there would be no way to tell what produced it.
I bet it does. I bet it also recognizes some human text as AI text, and doesn't detect other AI text.
1. Ask it to write according to the Google Developer Documentation guidelines. Gets rid of fluff, less emotional statements, no it's not x it's why.
2. Tell it you have extreme ADHD and need everything condensed as much as possible. You can always ask for expansion on an answer later.
3. Bullet points whenever possible.
I feel the same way when I read a "press release" or anything written by marketing. Even the newspaper will only have 2-3 sentences of interesting information spread out over 4 paragraphs.
So from "this table of stellar luminocity observations shows x y and z" to computer renders of green/blue planets with captions of "LIFE FOUND IN SPAAAACE!".
People don't do that. People are constantly engaging with paths not chosen. Right after I choose to write one thing, I'm immediately engaging with what I chose not to write there - I'm explaining why I didn't write it, I'm realizing that my choice may seem unusual so I'm trying to make it memorable, I'm focusing on the distinctions between what I wrote and what I didn't.
LLMs don't currently do that. LLMs just ape a structure. When the structure resembles the sort of timid, clarifying fussing I just described, the LLMs just drift randomly because what they didn't say wasn't in the context.
I also think that's why they have such a serious problem backtracking. They're not taking into account the already eliminated possibilities. Often the thing that was so unlikely that you weren't going to waste time on it is the answer, and things you discover while going down an ultimately wrong (but initially far more promising) path remind you of the path not taken.
They're simply assembling a thing that resembles a valid argument, and happen to make sound choices because the plurality of input happened to contain sound choices. This is usually a very good bet because there are so many more ways to be wrong than to be right. But it doesn't account for attractive (common) wrong choices. You need a way to back out of those.
I used Claude to help. I don’t know how to quite describe it, but because the text was polished and well constructed my brain was giving me the the signal “if you aren’t getting this it’s because you’re not focusing” so I’d read it again and then again and it still was not landing. It sorta felt like when you read something technical or heavy when very tired - you are reading but not processing.
Only after wrestling with this for a few days did I realize that it wasn’t me. As I started going through, sentence by sentence, forcing it to re-write things to be more clear the concepts became easy to understand.
I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The more complex the topic, the more I sense this.
I think part of it might be an innate feature of LLMs, but Claude seems extra prone to it lately. I ran the same query about the same codebase with Codex, and it gave me an answer that was about 1/4 the length and made me realize that it really wasn’t all that complex.
If nothing else, it’s good training for my own writing. I’ve been working on making myself be more straightforward and concise, and Claude’s writing is a good example of how cleaner prose is a functional choice, not just a stylistic one.
I think they all have the similar styles and tells. If I were to go to Claude, and use it now it would probably be clear for a little before reverting.
And I don't know why it feels to me like the language "drop off" happens after some time with the system. It makes me wonder if my account are getting silently degraded or sent to lower intelligence/lower priority queues after being a member for a while.
The human spirit. When you read a real person's thoughts you can often intuit the thought processes that led them to write it which aids understanding. Or at least have a general idea of "where they're coming from". But an AI is missing that. It just knows everything, without a "thought process". Instead of a flawed 3d person, we get a nice 2d picture instead.
The best I’ve heard of this is peeling the onion. The first pass is always very high-level and you have to make it go deeper. That can be done manually with follow-on prompts but I like using subagents, each with a different angle on the problem.
Bullshit?
But I also like your candy analogy because I think it's spot-on for how LLM text superficially looks informational/nutritious, even though it's actually just junk.
This is some third category of untruth. Almost more sinister than the other two altogether.
My two favourite words for this are “conditioned” and “catechized” where the latter is a bit more on the nose but way more obscure.
Slop. The word is slop. Has been for years now. I mean, is this not exactly what we've all been talking about the whole time?
[Thing] isn’t just [X]—it’s [more dramatic Y]. And [short validating statement].
I can see and smell this type of slop from a mile away. What I’m referring to is in the same family as slop but somehow different - it fools my brain by putting on the presentation of credibility and thus it is even worse. I can skip right over classic slop without much effort. This kind of text tricks me into laboring over it before I realize it’s hollow.
So in that way, it’s worse than slop.
If I had to, I'd process it into a short summary and/or ask an agent questions I have about the methodology.
I would then give my feedback. If they ask for details, I can have my agent update the document directly too.
I would not treat an AI generated artifact, be it documents or code, as something that a human should process fully manually.
I have a pretty large set of prompts that go into any software engineering, and I force every single agent to use an ephemeral style stack of prompt management. So, every turn it goes to the top of the stack and it is the very last thing they see in terms of all of my prompts and instructions and agent files. And then it gets taken out of the conversation so that it doesn't get sent to the agent the next turn (no context bloat). It has restored so much sanity.
I tried the caveman add-ons, and I felt like I was losing IQ points because I spend a lot of time reading agent output, and when they start talking like cavemen, I start thinking like cavemen. That was not good for my mental health. So, I try and make the agent talk like me and think like me. And it works, mostly. And my observation is that maybe I'm not the most efficient agentic thought process, but my sanity is retained.
All of that is to say that if something is reading like that to you, just have the agent rewrite it and read it in a rewritten tone because it's probably bad as it stands and your colleague did not put enough effort in it. It is /not/ good and you should not accept it as a default. We have to hold the line on stuff like this and maintain some semblence of normal human engineering standards that existed before AI. They are not making us better. They are making is lazy and dumber.
Opus 5 and other agents in the latest rounds of tuning have gotten ridiculously bad in terms of how they feel to interact with with all the invented language and localized nomenclature. It is an obvious bias that big words and technical talk looks good to the bottom of the bell curve, but when you actually try and understand it, it's horrible. So people say, "Yeah, that looks great," in all the RLHF rounds, and they run with it because they think it looks good, but it doesn't. It's terrible.
Hold the line. It isn't you. And it isn't a good methodology document.
Uh - dude - this means you're paying 10x in token costs because there's no caching.
If you 're-write token history' then you can't cache tokens.
It means for any reasonably long conversation, the llm has to reprocess the entire history as preflow on every prompt.
Are you sure you're really doing what you say you're dong, and how is it not blowing up your budget?
Perhaps more persnickety, it pushes the LLM out of distribution - if it’s unnatural for it to write in plain language without the prompt stack, your prefix will be an unnatural conversation which can reduce intelligence in hard to measure ways, especially over long conversations.
Not saying don’t do it, clarity is perhaps worth the intelligence hit, but it’s not going to be a free lunch.
Did you have to build your own harness for this? Or hack Claude Code or something?
Worse, I noticed that people in an office environment themselves have adopted a more speculative, communication style.
In the past, people remembered what was said and would draw attention to discrepancies. I could trust what people said.
Nowadays it's like; someone can say one thing one day and the opposite the next day (through convoluted language) and nobody bats an eyelash. Or sometimes someone will agree with me but then what they say immediately after reveals that they didn't understand the essence of my point at all. I didn't notice these things 5 years ago.
I guess this is what AI researchers refer to as 'model collapse' - it seems to affect people too though...
It feels like people don't value knowledge as they used to.
It's really hard to avoid mistakes when everyone is subtly covering them up. It feels like a lack of care and I find it demotivating.
I think because engineers are afraid for their job, they are under more pressure to talk a big game. Also under more pressure to deliver short term visible results. Bad combo.
And when you ground it with real data, it's actually extremely useful. It's not exactly like Claude Therapist, but it's sort of the teach me about philosophy, but actually grounded and not vied. I have a lot of really strict prompts and grounding and agentic guidelines for this particular agent flow and harness that I've built.
And it's just a few weekends of vibing and feeding it basically all of Wikipedia and several gigabytes of papers and stuff, but it actually leads to interesting discussion. So I just have my personal philosophy bot and it's pretty fun.
One of the modalities I built is having two agents assume a famous persona. And then they take a thing, like grief or some thing that I experienced during the week, and they assume the role of the two different philosophers, and I just have them go back and forth 30, 40, 50 turns. And it's actually quite interesting, and it really moderates their language and tonality and behavior. They really get into the roles when you have the right prompting and grounding. Sometimes they get a little off the rails, but it leads to genuinely interesting areas to explore, and then I'll actually go read source material and things like that. I don't know, that's how I do therapy these days, but I never actually did therapy, so I just think a lot, now with agents finding interesting stuff to think about too!
So, “Please write a one-liner comment manually to replace these 5 lines of AI generated comment” is a common refrain in my PR reviews to colleagues.
But to be honest I doubt most people who use AI for PR descriptions even bother changing anything.
Which might even make sense, because there were always (still are?) those horrible ads in the chumbox area of news sites that used trypophobia and other creepy body-horror stuff to get you to click. [1] So maybe the hope is that you don't really look closely at the quiche, but some reptilian party of the brain gets oddly activated and drives you towards the restaurant?
1. https://medium.com/the-awl/a-complete-taxonomy-of-internet-c...
(Actually that was my second thought; my first was "just how much H.R. Giger is in the training data?")
Quike? Cache?
Image models somewhat watermarking the image in a way that's very easily identifiable by a human seems present in all the image models of the big labs, since DALL-E 3 on OpenAI's side and the first nano banana on Google's side. I have no idea what they did to reach this and why they don't try to fix it.
We have one integrator which consistently uses a truly bad AI. My brain skips after two sentences already. There are always +9 questions what should just be 2 max. Redundant info is requested (e.g. please provide a change history of this function and if it will be deprecated) and everything is just unbelievably verbose.
90% of questions we get from them are already answered by the web documentation (which even has a search function) or are just non sense.
I'm truly considering to build an MCP just for them...
If they can't bother to come up with their own question, I certainly wouldn't waste my time on the answer either.
Margaret Storey's recent piece on Cognitive Debt: https://spawn-queue.acm.org/doi/10.1145/3807966
"There's an ongoing discussion of whether humans are good at recognizing AI-generated text. While most research claims that humans don't really do a good job there, I disagree. "
I wonder if humans that spend all day working in tech are good at recognizing AI-generated text, but people who spend all day doing jobs that don't involve computers aren't as good.And I wonder if those of us in tech are the only ones who really care?
* https://www.theverge.com/ai-artificial-intelligence/975017/, https://www.lesswrong.com/posts/6ZnznCaTcbGYsCmqu/, https://spiralism.website if you want to test how strong your defenses are against this particular meme
I have no idea if other people who work in tech are better than average or not, because I don't feel confident in being able to check their work. That being said, I do think that there's a general trend of people in tech tending to be a bit overconfident in how well they will do at some new task they haven't encountered before, so when someone tells me that they can easily tell whether text is AI generated, it's hard for me to trust it any more than I trust someone who makes a similarly strong claim about something that they can use AI successfully for when it's not something that I can easily measure (e.g. learning a new language without getting feedback from people who are fluent from real-world usage).
All that being said, I do think the set of people who care is larger than just those in tech, although it's probably still a relatively small group overall. From conversations with people in other domains, there are contingents in non-tech communities who tend to have a large representation of negative views towards AI (artists, writers, musicians, other jobs where people are skeptical of human creativity being replaced by AI), and often times the people who feel negatively in those groups will be even more adamantly opposed to interacting with any AI content than people in tech. To be clear, I'm not at all trying to generalize and say "all artists hate AI" or anything like that, since there's obviously a wide variety of viewpoints within any sizable community, but I've definitely seen many people who say they will refuse to play any game that's suspected of using AI for generating art assets, and even some who don't differentiate between using AI for generating assets versus code (either because they aren't knowledgeable about how different aspects of game development work, or they genuinely don't care because they view AI as a categorical evil).
And I’ll concede on both ends that there are probably times I suspect content is AI generated when it isn’t, and times I suspect it isn’t generated, but it was.
AI tells seem inevitable. You have millions of people communicating with one effective “personality” that has tendencies to write in certain ways. If its content is published verbatim, then it will be easier to tell whether some content is AI generated just based on its similarity (sharing certain linguistic features) to other content being posted.
It’ll never be black and white though.
If you're exposed to AI a lot, you're going to start noticing patterns that allow you to identify it.
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text - https://arxiv.org/pdf/2501.15654
I think they may just be to trusting and/or naive. People in tech right now are hyper aware of this and are actively looking while people outside of that bubble barely give it a second thought.
Going further, I'm curious about whether people are mostly good at the case where they suspect most or all of the content from given "author" has the same amount of AI usage/prompting in generating it rather than the adversarial case where someone might usually use AI extensively and then try to slip by purely human written text (or vice-versa). I don't have a good sense of whether this is a threat model that actually matters, since maybe the heuristic of weeding out sources that are mostly AI-generated is enough for people who prefer to avoid that type of content, but I do think that changes the definition of what it means to be "good at recognizing AI" in a meaningful way. It seems plausible that disagreements about how easy it is to recognize AI content might be coming from two people assuming a different framing of the question that results in a different answer without realizing that's what they've done.
Several existing studies I’ve seen have done things like prompt the LLM to produce a poem in a certain poets style, then ask people to spot the fake in a collection of poems, which they aren’t great at. This is, I would argue, an extremely different context than what most of us are encountering AI text in, and the people sending me text aren’t prompting it stylistically like that.
On your second question, I definitely feel like I can tell the first time a coworker sends me AI text masquerading as their own thoughts, even if they had previously been opposed to such a thing. So it could be that familiarity is more important than my prior on whether they’d use AI? But interesting to think about either way
The interesting question is how to define 'average'. Over what probability distribution?
Hey, anyone remember this from earlier in the week? https://daringfireball.net/2026/08/anthropics_watermark_text...
But it’s just as likely to make an output better.
Take the example from the article. He complains that watermarking might sometimes, for example, choose to say “bananas” over “pineapples” because only the former is on the green list, potentially making an output less precise. But 1. It could do that regardless of watermarking since the model is probabilistic, and 2. The more accurate word choice of “pineapples” is equally likely to be on the green list instead, further increasing its likelihood!
Overall, the article is pretty silly because he’s complaining about the possibility of Claude not always choosing the most “optimal” token, even though LLMs are probabilistic so that will happen anyways.
No, for any particular output token the model's true logits are definitionally the 'best' that the model can achieve.
This is inherently probabilistic. The model's top-1 guess is not guaranteed to be optimal, but it should be so a proportionate fraction of the time. Same with the top-2, top-3, etc.
Watermarking necessarily alters the output distribution away from the model-set distribution, and that alteration is inherently 'worse' in expectation.
You can liken this to a weather forecast. If there's a 25% chance of rain, the forecast should say so (or a 'sampled' deterministic forecast should predict rain 25% of the time). If the forecast is 'watermarked' and predicts rain 27% of the time under identical circumstances, it's a worse forecast.
That being said, this is a case of hiding a message in a noisy channel. Watermarking only needs to communicate one bit ('yes watermark'), so the effects can be arbitrarily small provided one is willing to tolerate an increase to the text size needed for reliable detection.
Which is obviously not how it works.
> He complains that watermarking might sometimes, for example, choose to say “bananas” over “pineapples” because only the former is on the green list, potentially making an output less precise.
If I replace "pineapples" with "bananas" in any kind of meaningful text, I've not made the text "less precise", I made it plain wrong. An incorrect statement. And if the words next to it are still correct or even somehow replaced with even more correct words, the text in its entirety will still be wrong.
> It could do that regardless of watermarking since the model is probabilistic
No, because the probability distribution of the model is generated by its training data and contains semantic information. So the model choosing a completely incorrect word is unlikely. However the probability distribution of the red/green lists is not guided by semantic information.
I find that when I try to speed read modern human writing, there are often errors (like missing or misused words) or awkward expressions that I do have to slow down and think harder a lot to really parse it.
With AI writing, it's sort of self redundant and the information density of each sentence seems to have more even information density. This makes it very easy to do a very high level speed read and get the full gist.
There are also what I'm assuming are bots on hugging face (or maybe non-native english speakers who are using ai for translation) that interact with me where I have no idea what they are saying until I read it very slowly.
Does speed reading help you process the final message faster if it's written by AI compared to people?
Because if you read 1 information dense sentence, 1 medium dense, and 1 sparse sentece written by a human, it's still way less text in total than 6 information sparse sentences written by AI... even if it's all over the place when it comes to density or style.
---
The density argument is really interesting.
Does speed reading actually help you process the final message faster when it’s AI-generated compared to human-written?
For example, if a human writes 3 sentences—one information-dense, one medium-density, and one sparse—that’s still much less text overall than 6 relatively sparse sentences written by AI.
Even if the AI output varies a lot in information density and writing style, you still have to process all that additional text. So I’m wondering whether speed reading actually offsets the verbosity of AI-generated responses, or whether the total amount of text is still the bigger factor.
As someone who hasn't practiced speed reading, how does that happen? Is it something about the way your brain tries to connect ideas from different parts of the text? Or the redundancy making the signal more stable?
If your reading speed is limited by how quickly you can subvocalize the words to yourself, this is significantly less obvious. Unless the passage is dense enough to require multiple read-throughs at conversational reading pace or vapid enough to be boring, you're going to feel done with the text at roughly the same time. Speed readers do a lot more re-reading and varying of reading speed, and that is going to correlate pretty hard with information density.
pre-read is just looking at how long it is in the headings, and planning out what chapters to focus on if it was a text book. (its sort of iterative, you do a pre-read for the whole book, and then for each section you break it into)
The fast read you try to read only with your eyes, sweeping your eyes across multiple words at the same time, suppressing the urge to say the words to yourself in your head.
iirc the how to read better and faster book even had a cardboard mask you put on the page to practice the sweeping, and some pages that were laid out weird to try to teach you how to do it.
Some ai text just seems really easy to speed read, like if it's tuned for an easy reading level. In PRs some ai seems like it's arguing over weird flex technical details and really starts torturing the language in a way that makes it the opposite of easy to read.
I'm not sure why this issue is so prevalent, it's not hard to point Claude at the Wikipedia article on signs of AI writing or ask Claude to write content anyone of the average American reading level could understand.
To me it just gives off a sense of laziness, that you cared so little of your content that you did not take the time to read it yourself and edit it to effectively communicate the message you wanted to communicate. To that point, it's just not worth my time reading, otherwise my eyes glaze over trying to read between the lines of machine written language for other machines.
We can no longer assume that the models will create a passable English in their users targeted messages. So in the short term, we need to have a tight guidelines in the harness, but in the long term, we'll probably need separate languages for thinking traces and for human communication. Possibly even not going through tokens at all when doing thinking part.
Alternatively, we could evaluate thinking traces on human readibility, but that's probably infeasible due to sheer volume.
When I ask ChatGPT questions I usually only read paragraphs 2 and 3. The first paragraph is glazing me, anything after paragraph 3 is repeating what was said earlier.
You know this: you read something and it's like 'here's a mass of poorly structured text, it's not clear what the points are, it reads ugly, I have to read to the end to discover that there are no points! Or if there are points i now have to reread it. What a lot of effort.'
LLMs have annoying grammatical tics and can have other problems too, like the author says, but sometimes they can help to take bad writing and make it clear and more digestible.
This is not AI specific. I have come across many humans who describe a simple concept in a very complex and verbose manner.
But what I hate the most is that it is objectively better than what I had before. No typos, clear structure, and, regrettably, the verbosity and autistic obsession with detail of the LLM is more actionable and useful than the human guy who wrote lists of commands and URLs as documentation, without explaining anything. Or the colleague who writes in uppercase and with question marks and who doesn't make any sense and forces me to engage in an interrogation effort to get to the bottom of what they are trying to say. Or the colleague who simply hates writing--despite being decent at it--and will call you to give you a meandering verbal explanation that lasts two hours of what they want from you. The cynic in me bemoans that we brought this upon ourselves, in more than one way.
Underdocumented, underexplained and sometimes out of date... or overly verbose, repetitive, information sparse, and sometimes halucinating.
Both are bad and with some effort could be prevented.
When writing code, I have to explicitly tell it how to structure things at a high level, or the result is sort of a flattened spaghetti. Similarly, when it’s explaining things, it’s not good at pulling out unifying concepts and explaining top-down as a smart human would do. It groups little things together but often doesn’t generalize or synthesize explanatory connections from them.
I’ve been experimenting with explicitly working through a sequence of outputs at different levels of detail, but I haven’t found a consistently successful method.
With the right blend of context and prompting, I can often get them to "lock into" an existing model, but using them to generate a novel outline or sketch, whether it's for an essay or a module, usually results in garbage.
If I'm going to be iterating on one document or idea in an extended manner with feedback from the chatbot, I will make an effort to setup the decorum it should follow, because it's a small proportion of the time in that chat.
But the guy just trying to get a report or email out quickly? Not so much.
The reason I brought it up is because, people who learn English normaly start with a book. It's heavily polished.
When you speak English as you learned from the books, it does not sound very conversational.
If you are native/fluent English speaker, you can feel the impedance mismatch and feel something's off.
The AI-blindness stems from the fact that those polished edits are so common in publishing field, they all sound the same, and unable to recognize the diffs between AI-generated and human-generated.
There is no real human conversational vibe to them and well. i will stop now.
Despite seeing a lot of them, I cannot think of one AI-generated photo that I can picture clearly in my mind; a few are partial but elusive. Whereas I can recall (visualise) a whole bunch of traditional photographs.
The same is true of AI generated text. Only the annoyances stick. I cannot recall real details of text I have generated, until I commit it to memory some other way.
I don't think this is about ephemerality either. If we assume it's about celebrated/famous/infamous images, there are definitely non-ephemeral, cultural moments in AI generated images in particular, like Boris Eldagsen's Sony Prize winner:
https://petapixel.com/2023/04/14/artist-refuses-prize-after-...
This really should be memorable, but isn't. I had forgotten the second person is in the image.
Or Jason Allen's fake painting:
https://petapixel.com/2022/09/01/ai-generated-artwork-wins-f...
I had already forgotten there's more than one figure in it, and I only looked at it a few weeks back. I remember the colour, the bright circle, some vague hints of texture; one figure. And that is it. Only the crudest shape elements.
For me, something about AI-generated text and images confounds recall. It is really peculiar.
There's no real edge to it. Same as with the writing. The stuff that you'd latch onto (and thus remember) is simply not there, precisely because those image or word choices would be just outside its latent probability space. But because they're well inside it, your mind sees nothing novel to register.
This is also why I think human output will actually increase in value. When any AI can just "phone it in", something genuinely human will stand out (to us, not the AI) and become a bellwether.
This will literally help us realize what it means to be human.
I also don't think the solution is simply to "make responses more random", either. That might help solve novel problems (the same way that throwing darts randomly at a dartboard eventually hits the bullseye of the dartboard right next to it that no one considered), but I don't think it will help it "seem more creative".
Yes, as if it is in some weird hidden dimensional sense completely uniform.
ETA: suddenly reminded of the Bateson quote about information being “the difference that makes a difference”.
The examples I've seen of AI music (though I avoid it on principle) seem the same.
Sometimes I'll check in on my sleeping kid and she'll sit up in bed and say some utter nonsense. I'll find it hilarious, giggle silently to myself, and kiss her goodnight again and she'll close her eyes and lie back.
Why I try to tell her about her sleep-talking in the morning, though, I find that the words she said have completely disappeared from my memory, no matter how funny I thought they were at the time.
In my head-canon, this is because it's dream language, and slips away as easily as dreams. But, like the AI art, it could be because it's bullshit: completely devoid of content, all signifiers and no signified.
I stopped after I wrote something on the pad while I was still asleep. Woke up to text with letters that were backwards, upside down, weird words — so close to real words that I was sure I ought to know what they meant and had really meant to write them down.
Scared me. Literally too weird to keep. I tore up the page.
In a way I think this is one part of the same continuum. There are thoughts that have meaning and can have no words, and words that look like they should have permanent memorable meaning and don't.
??????
Nearest I can say to how unsettling it was is to nudge you towards the video of the angry cockatoo who doesn’t want to go to the vet. Everything he says sounds just like it’s on the edge of having meaning.
Imagine something like that, something important, so close to symbolic meaning, only writing. In letter shapes that we don’t use. And you wrote it while semi-conscious. If that’s something you would want to keep, you’re a braver person than me.
This explanation again doesn’t get it across. Too weird.
You know how you try to run in a dream, and you don't get anywhere? Imagine if your motor cortex started to expect that, when you're awake. It would destroy you.
They were actually rather good in a sort of "fake collodion image" sense, and the eerie early-DALL-E quality to them really helped the spookiness.
But I can only remember this technicality and the feelings with any clarity, not any of the details except in the broadest sense. I cannot bring these images to mind in any meaningful way.
They were deeply wrong and it's only the wrongness I really remember. It confounds memory.
Modern image generators have ironed out all the structural wrongness.
I can’t think of anything more unsettling than the actual events that took place in the Belgian run Congo Free State in the late 19th and early 20th century. The Wikipedia article is scarier than any creepypasta you could write in that setting.
Reason why those images are flat and boring is that they are just statistical guesses making a composition averaging whatever the model has been trained with. They would be technically brilliant (if made in oil), but superficial and meaningless, same as so much Sunday painting is.
Same goes with language. Nobody is trying to communicate anything with you, so it just words after another. You can create meaning out of it if you want of course, we homo sapiens -apes excel at that, but what’s the point? Language Jones on YT has pretty good video on this[3].
[1] https://media.mutualart.com/Images/2024_01/12/12/124216388/d...
[2] https://uploads4.wikiart.org/00339/images/jean-leon-gerome/t...
The way my memory works (especially as an amateur photographer) I would thus normally have a very good chance of remembering some key details of the images; some fascinating element of each would connect with the rest of the memory.
But it does not happen. Whereas I sometimes remember photos with clarity while forgetting where I even saw them.
What is the best chicken noodle soup I can make?
And you will get at least one recipe which surely is tasty.Ask a person the same question and you might be told:
One you make for someone you love.
This is the difference between a statistically probable response and understanding.The best basic explanation I've found is this video.
https://www.youtube.com/watch?v=ORgKY9AlybA
The real answer is to take a class in discourse analysis and learn how to pick apart sentences. Once you start understanding how language works in the brain you can start to see better how LLMs fail.
AI is surfacing layers you never had a chance to see, makes you rethink your career path and who you are willing to work with very fast.
neatly dodges literally everything that is interesting about what they do.
You to, friend, consume inputs and generate outputs.
What's interesting is how you do that, and, if you prefer to look at it through a technician's lens, whether or not what is done is reducible.
I don't mean quantizing the model, great, now you have a crankshaft with no oil. But the motion of the pistons and wheels is roughly the same.
What I mean is, to put it in plain terms, the only way you find out what a model is going to "predict" from a given input is to ask it.
If your mental model is still that LLM are "glorified overhyped giant markov chains" performing "parroting" you need to improve your understanding.
I saw my ex-coworker tried to inflate text that can be expressed in couple of bullet points text. It doesn't add anything. It just fill the fancy flavour text that looks like professional. Then the reader of that inflated text is also another coworker in other team.
I said "Forget it and just send these bullet points" but he refused.
Bullshit job at it's finest I guess.
Perhaps we go back to feudal society when climate change crumbles the civilisation, world economy and democracy. It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
Point is, the technology itself is not harmful nor innocuous in itself, it is how we let these shitty corporations to guide and control how we use technology. AI is a useful tool, so is chat application with a friend list. So is a hammer. You don’t need to use any of them to smash your face in.
This resembles my country.