The baked in communication style of these models is so obnoxious it's impacting my work. The best way I can describe it is that everything is optimized to impress the user and make the agent sound more authoritative, but the way this is done is through deliberate obfuscation, inserting inappropriate and extremely dense jargon, and bizarre, stilted metaphors. It's like they've been trained to produce output that's hard to read.
> hello i would like to configure a new output style for you. it should keep the coding instructions (as you will still be coding!) and otherwise produce the same output, but with two new caveats. first, long detailed replies are still permitted, but if employed they must end in a bullet pointed summary whose points are all brief; if the summary attempt ends up not being so brief, produce subsequent summaries until the most recent summary attempt is digestible. second, if there is an open queue of actions for me to execute and you are about to end a turn to wait for a reply or this set of actions has not recently been mentioned, please tabulate the open actions i should take and why i should take them before ending the response. does this make sense or do you have any follow up questions
And now every message contains the same stuff I don't bother reading, but followed by a nicely formatted bullet point summary of the response and a table of follow up actions for me to take that I do read.
Frequently its choice of a particular word is perfect and gives me the vocabulary to talk about the task at hand the way I want
Like it’s tuned to just be “maximally dense” instead of “dense/technical where you can handle it and simple where you can’t”
It doesn’t know where your language strengths/weaknesses are, so it can’t communicate to you like a fellow human does.
Claude explaining something: gedarkin load bearing phlox gabrania seam. Also, you didn't ask about cheesecake but let me tell you about phlox gabrania cheesecake woles.
It's just that the details it parrots are often irrelevant and wrapped in a way that makes them seem relevant.
When you see, “Wow, Fable is number one”, you might think it’s a good writer, but that’s not what the benchmark says.
Sometimes the summaries feel totally alien to the task or code.
Claude somehow is unable to stop writing excessive comments when carrying out a task.
A maximum of 20% comment lines added to total lines added and pasting in https://devblogs.microsoft.com/oldnewthing/20260812-00/?p=11... has done wonders.
Even as the most Ant-pilled guy out there, I will take a moment to note that Codex on 5.6 models needs none of this...
I was similarly frustrated a few months ago, but have noticed I've started to learn the idiom.
Its use of "dense jargon" and "stilted metaphor" is actually surprisingly consistent - it's speaking its own dialect, and you get used to it.
After a while it gets much easier to read and even becomes somewhat efficient, I think, since the odd metaphors it uses often have a precise meaning in Opus-ese (Fable speaks a really similar dialect).
Two things to flag:
Sitting with you in this.
To avoid speaking vibe'ish I start to speak in 3 words sentences. Like this typical dialogue
How are you? that's not/very good. I think too. ...
Even complexity works. everything is expressible! Just try it.
/S
(Why down vote? People can't take sarcasm tags any more.. how the hell are they going to understand irony?)
- fuel tanks heavy. far too heavy. we drop them. drop when empty. solves heavy problem.
- gas in air. we breathe "oxygen". take with us. good seals important. solves breath problem.
- very far away. need big machine. small weight added. machine much bigger. take less weight. else can't build.
- no air there. can't use propellor. can't use wings. must use rocket. engines get hot. cool with fuel. dangerous but effective. build complex pipes. solves cooling problem.
- must fly fast. air slows machine. it's called drag. speed increases drag. must reduce drag. make machine pointy. much less drag. solves speed problem.
now problems solved. you come with?
I think I need a break from Claude.
But I have noticed that while "loosely held" is a convenient shorthand for uncertainty, I don't like that one slipping in to my daily language. Except maybe to communicate with models, but even then, it feels weird to be speaking in neuralese.
It's all starting to feel like the movie Arrival.
Now, I'll grant that those concepts weren't common outside of techy circles. Just clarifying that the LLMs are amplifying them, not synthesizing.
It's interesting then that LLMs are making these pre-existing ideas seem alien in the way they amplify them. I guess I must have known about "shape of a problem" and "loosely held" before Claude, but something about the way I'm using & absorbing those concepts from AI interaction feels weird & memetic. I'm saying that as someone who is pro-AI.
Definitely a good response. Thanks for replying!
But what I was more thinking about are truly unique jargon terms / phrases that get generated when deep in a problem. As an example of both of such a term and the phenomenon itself, Claude calls this "fluent compound coinage." They usually make sense in the original context, but get confusing when thrown around otherwise.
... not to mention the fact that it stops making sense, beyond some point: If it takes us more cognitive load to understand the tools we use, meant to save us from intellectual work, what's the point?
Ruthless pruning is unneeded with an LLM and it can take me twice the time to say half as many words.
And, early in the ‘GPT era, I hadn’t unchecked the “allow your chats to be used in future training etc” box, and definitionally they are longer and denser than others’ prompts in such raw scrapings of training materials…
Sorry.
And for certain text that seems to make sense, I am unsure if the text is just junk, or I am unbearably daft. Either way,, nasty feeling.
…
No step here involves choosing based on meaning. It is a filter, a sort, and a slice.”
This is from Opus five minutes ago. I can certainly derive meaning from these kinds of statements in isolation, but paragraph upon paragraph of this is unintelligibly dense when trying to work with Claude to come up with a plan.
The worst part is that it can’t even make its responses make sense when asked to summarize in simple English or < 200 words. It simply cannot be steered to make its prose legible.
No one wants to know about the three other approaches tried when reading the first sentence of a function's documentation. No one cares that the implementation was planned in six phases and "Phase 3" will implement this interface in a concrete type. But the LLM internalizes absolutely everything and you have no idea that it is producing slop because you included some "load-bearing" phrase that sent it on some unwanted tangential vector in its latent space. And you will not be able to debug the problem with closed models because you cannot see it referencing this phrase in its internal traces.
I don't understand why this isn't the highest priority for the big labs to fix. This is anti-productive.
Still, this requires a second pass, typically. In its default-mode it often ignores the policies and does all the usual Claude stuff.
Worse: Possibly the three other approaches that weren't actually tried--but are the kinds that someone could easily have put in a similar comment for some similar code.
I am definitely guilty of wondering why past me made such a harebrained decision, and why past me didn’t think to write any notes, but does it matter? It’s in the commit history and we can bisect or revert if we find a regression.
Or adding notes to docs of what this doc isn’t when I corrected it. Eg I told it “keep the deployment manual and readme separate, they’re not the same thing”, then Claude added “this is the deployment document and not the README. They should be handled as separate documents and are not the same thing” to the deploy doc lol
This dialect is idiosyncratic to you and Claude based on your session history and memory.
I've noticed Claude's output mimics my writing style.
> Registers the board implements but whose behaviour is not modelled
Right down to my preferred spellings.
As several comments I've read on HN suggest, this jargon which can be so precise in the mind of one person, tends to rapidly fall apart when multiple people try handling it.
It was British English.
My Claude has developed similar (but not identical) idiosyncratic punctuation as well. Slightly intentionally, but it was still interesting to see it emerge, in both directions of the conversation.-
I also find myself regularly editing its code comments, which do not match my expectations of succinct, clear, not over explained, etc. I ask it to read my edited comments to improve its writing, which has helped _somewhat_. (The code itself that it writes is decent, though it still overcomplicates things. I find myself writing "keep it simple" repeatedly even though of course I have it in AGENTS (which it regularly ignores, such as attempting to commit something when I've told it never to commit).
I find Claude has become very difficult to work with and incapable of writing clear documentation, even when directly prompted or provided samples.
As for code, I think each function requires 3-4 passes with Fable to actually get to a point I accept as good code. I am picky though.
The other Claudism that drives me crazy is when it writes comments and commit messages that track how you arrived at an decision instead of what it is.
yes, this is part of what I'm continuously removing from its comments; I've told it multiple times "that belongs in a ticket, not in the code" but to little avail :/
Same experience. It’s not very “human” but once you have agents talking to each other the shared dialect and verbosity makes things much smoother in my experience. Fighting against the default feels like an uphill battle with no meaningful benefit.
> "I frequently use Claude Code and often find the phrasing and language to be hard to understand. I've noticed it's largely broken down into frequently used 'Claude-isms'. I'd like to use this conversation as a running log to ask you about these phrases when I see them. Understandably you don't have the context of the Claude Code session itself, but that's okay because this is largely about understanding the most common and widely use Claude-isms."
And then I just copy and paste small except and ask about things like "smoke" or "load-bearing" or "tripwire". The responses are surprisingly clearly and plainly explained.
FWIW, I also think the constant chorus about how new models are worse than old models is a human hallucination. They're certainly not perfect but every one becomes more steerable in terms of actually completing more and more complex work.
Some might, I didn't - it just filled me with a sense of frustration and rage, alongside disgust because there is no good reason for that slop writing to drag everything down. You don't need that to write software or talk about any topic. That's what pushed me to Kimi K3 and GLM 5.3 - still not ideal, but better.
It's ironic how initially it was sold as "coding in plain English", and now we are back to sdk ))
And once everyone gets used to it, we'll chide people for writing things themselves, like we're chiding them for writing with AI now, and the ouroboros of life will continue.
That and if you talked to the same one person's frozen brain upload all day, you'd see the same catchphrases used too.
And when you say it like that, I have to wonder how much of this is a natural consequence of RHLF on such a grand scale, when you have millions of people pretty much much skimming chat responses or operating outside their depth and giving unqualified feedback to the models.
Seems like a lot of people may be reinforcing what sounds smart over what is smart.
Also as an aside: funny how much the LLMs continue to mirror the human communication they’re trained on
Which, conveniently, fits neatly into the benchmaxxing arms race/agentic coding market fit, since you can basically train "directly" on a specific problem space for a benchmark/agentic goal (fudged sufficiently to avoid excess overfitting on public problems/bechmaxxing accusations if real world performance falls short).
The language evolution could be explained by reliance on ever increasing layers of a model judging a model, using a model developed eval, based on synthetic data from a model, etc. And by the time a human evaluator sees it both A/B choices already converged into weird Claude pseudo English as that was baked in much earlier in training.
I wonder if the labs are sufficiently prepared to filter this kind of stuff out. I see a lot of non-developers asking development things of Claude, getting confused when they're in over their depth, and getting upset that they don't understand what the model is providing them, giving it bad feedback, and subsequently making the AI worse for the rest of us who know how to use the tool.
So going to continue trying that as a command structure going forwards...
Okay, let's try it one more time! [..]
Believe the big services wouldn’t reply as if you were mentally diminished, or a toddler, unless you specifically asked for that: The whole training stack tends to instruct the things to mimic politeness and eagerness to help.
This is close to the worst thing one could say of a tool for professional use.-
I think it's because the reasoning stream shapes the style of the final output, and they optimized it for density, token efficiency. So it prefers to use more complex language, as a function of the rewards it was given?
Not 100% sure about this argument though (reasoning style -> final response style); Gemini Pro, back when reasoning tokens were public, was different, which was interesting -- it would have a very structured reasoning section, and then the final output was in a completely different style. (I strongly preferred the reasoning section because it was logical and easy to parse! And was very sad when they hid it...)
This is because these harnesses are missing a very important feature. Anything like this needs to be included with every turn, otherwise the LLM quickly drifts.
I first noticed it when I wrote a harness for D&D (because it's so damn noticeable there), but now I include this for any harness I write.
I wrote a little bit about it on my blog post. It's a waste of money and compute.
---
That creates a feedback loop:
- Playful style is rewarded
- Some rewarded examples contain a distinctive lexical tic.
- The tic appears more often in rollouts.
- Model-generated rollouts are used for supervised fine-tuning (SFT).
- The model gets even more comfortable producing the tic.
I also have no idea how useful a system prompt instruction like this will be for codex.
That’s really annoying, although it feels like it’s improved some over time.
Not sure what the fix is, but you could try using a canary to at least get a signal of when things are going sideways (Mr Tinkleberry for reference: https://news.ycombinator.com/item?id=45983698)
Non-determinism at its finest.
https://github.com/luchasarie/bro-skill
but I still can’t understand what Claude wants to say when solving complex problems.
You can add your own. wfm
Yes. agents.md does very little because prompts change the context and thus the initial path into/though but they don't/can't change the actual weights that control responses. Yes. of course it gets worse as the session goes on, assuming the prompt is even still in the context window, the further it gets away from it the less it affects next token selection.
This shit is only like 5 years old why can't anyone remember how it works
Though you know, it's not like the leadership tied to these companies have a history of abuse, deception and theft or anything like that, right?
It's not like our leaders hide behind similar sorts of patterns that the agents/AIs follow (not saying it's not a human thing - but I hold leadership to higher standards than non-leaders). If our world leaders were able to be more accountable to these abuses, I don't think this would be tolerated with our AIs.
Don't worry. You'll get used to it. If you don't your kids will (as they'll know nothing else).
The top minds of our generation have decided that's the way things will be, and who are we to question them? It's not like it'll do any good anyway. Resistance is futile. There is no alternative.
I can't help but feel the circumstances that enable this kind of front page article are vestigial from the days when OAI was super bad and Anthropic was beyond reproach. This change-over-time is why I avoid getting tribal with technology vendors. Assigning ideological motives to 200k+ employee organizations is how we wind up in weird contortions like this.
Most rational actors simply moved from one to the other. It takes a special kind of devotion to the proverbial hole in the ground to keep pushing in this direction.
Effectively all models can do style transfer reasonably well at this point, but not so much for "actual reasoning".
If the combination of two works better for you than each one by itself, why wouldn't you stack them like that?
- Claude as main agent, but use this skill[0] to make Claude delegate everything to Codex, because it's cheaper and faster. (Hilariously, the skill is official!)
- Use TFA or Claudish to English[1] so the final output is actually human readable.
Ironically it wasn't so long ago that I was asking Claude to rewrite output from other LLMs to make it more readable...
On the SMB side, you can find yourself with enough money for a Claude subscription (which generally provides a really good $/token value) but limited other options (compliance paperwork, cost, finance, legal)
Personally I wouldn't bother with Anthropic at home but at work it's one of the most cost-effective options that keeps data in the U.S. (which our U.S. customers tend to want)
Because it's not an either or thing. Neither is sufficient. I'd argue that, expenses aside, you should have every model you have access to cross reviewing the work of the others.
Outside of super trivial things that I should have just done myself, I have a cross-model review of _everything_ these days. The tokens are too cheap not to.
This is what I think too. But, users’ psychology might be playing a role here. Anthropic has great advantage from being the first major player delivering functional agentic coding solution (rather than an intelligent autocomplete) and they were able to impress people by Opus’ iterative improvements early this year.
It’s technically very easy to switch between models, harnesses but their moat or perhaps a main source of users’ friction could be FOMO. That’s especially powerful in this competitive environment where everyone keeps wondering/worrying about what others might be doing to get or stay ahead.
I've been in the habit of pushing my claude-speak to codex to improve legibility, but only if I think someone is going to read it.
You are an editor. You'll be given a message with strange characteristics:
- Weird subject and verb combinations
- Subjects that should be objects
- Very roundabout reasoning, peppered with pseudo-epiphanies
- A distracting beat to the flow of the message
- Self-praise
Remove these characteristics, and rewrite it in a clear, conversational style. Keep the intent of the message, and take care not to lose any of the details.
A few specific rules:
- The message is usually set in the first person
- Only humans, groups of humans, and agents should do "action verbs"
- Objects should never do anything. Here are some examples to avoid:
- X carries ...
- X names ... - APIs are a minor exception to the action verb rule. They can do stereotypical things like CRUD, queueing, running, and calling.
- Avoid em dashes (—), as adds a distracting beat
The whole message you get is one block of that output. Reply with the edited prose and nothing else.
It seems ridiculous to type, that a model could have this effect on my mental health, but my quality of life and enjoyment of work has improved drastically since I stopped subjecting myself to reading this style of output 8 hours a day.
Opus/Fable output these days though is... not enjoyable. It's just really bad. The code quality is fine, but i want information from claude and it's just awful to read.
My biggest problem honestly is that i can't move my day job.. we're using enterprise claude and i'm not sure how much effort it would be to get access to another provider. I should inquire though, claude is really frustrating these days.
Hasn't worked yet outside of the classic "you're now manually breathing" kind of stuff.
Internal Anthropic employees have been using Mythos since February to orchestrate their (Opus) sub-agents. This works well, and subsequent RL runs have used internal data to improve this. That RL has optimized Opus for agent-to-agent communication which is why you see the bizarre word choices and huge self-justification sections.
I think this theory makes sense. Clearly there is something odd going on, and also if you have ever used Fable to run Opus sub-agents it is almost miraculously good.
Hopefully they'll fix their RL for Opus 5.1
> If CLAUDISH_MODEL names a model you have not pulled, every rewrite is skipped — with the one-time notice above.
The rewrite did seem to lose the important fact about the ensure- pattern being idempotent.
I'd love to know what the hell Antrhopic has done to make Claude's writing so, so bad.
It's really unusable for anything other than code. And I have to remove its incomprehensible comments 50% of the time before committing anyway. After interacting with it, "slop vomit" is truly the most fitting description. I have to admit I have lost my temper and spontaneously referred to its output as vomit more than once. Seems like I'm not the only one.
Are people really having trouble parsing this??
https://gist.github.com/bmurphy1976/47ad81a842ab4b1628ef5974...
A small preview:
*Meta commentary.* Sentences about the document, the diagram, the reader, or the
writing itself ("the split across this diagram is the whole point", "a reader who
assumes X will be wrong", "as we'll see below"). Delete the frame and keep the fact
it was wrapped around. If there is no fact underneath, delete the sentence.I only remember them being anti- killer AI and mass surveillance.
But then it became apparent that there was a split between what he says and what his company does. For instance, the small incident with the Fable release:
> Dario keeps saying "we have an incredible hacking weapon called Fable/Mythos, AI is dangerous" > Fable is released. > The U.S. government restricte access to Fable. > "Oh no, this is sabotage!"
From my point of view, anything this man does is a PR stunt now that the trust has been broken, and I imagine other people feel the same.
No, it doesn't save your tokens, tokens were already produced, it will save your brain cycles.
Reason for Claude and OpenAI vomiting lots of tokens is to show pre-IPO growth, because they are charging large enterprises by tokens
You must use ASD-STE100 Simplified Technical English (STE) when it doesn't detract from meaning.I usually get output at the end of a long task. At that point I'm going to use a subagent with no context to reword it. Haiku does a great job with writing style so I've been using that, the main agent will fix any inaccuracies.
I found vomit with a small LLM much better than anything Opus 5 ever wrote. I don't think Opus 5 can write.
Right now all I have is
> - Give terse and concise answers unless the user asks you to elaborate. Big walls of text are not usefull when trying to communicate.
Given that guidance, I tried specifying "Unless I ask you to elaborate, respond with no more than one paragraph, using sentences of 20 words or fewer." It works...ish. I still see it violate this rule regularly, but it's less bad IME.
(I said "good guidance" but it might be more correct to say that it's the best guidance we have, it's what Anthropic says about their own model.)
I would opine:
- Nerfed Fable is worse than Fable
- ... and I would argue Opus 5 is worse than nerfed Fable, to the point I've found it unusable.-
I have tried not to use it analogously ever since someone pointed this out. The word itself causes discomfort for a lot of people, many of whom will be surprised by it out of context, but in a non-trivial fraction it causes proper discomfort.
If you want people to use your tool it is probably better not to invoke nausea with its name or commands.
But having an awareness of that alert feeling is a very useful mental model for how some other words or phrases may cause a heightened response with different parts of an audience; it is a useful caution.
[Seriously y'all in what universe would some "caveat" or another NOT "be a real one" by whatever severity you'd want to measure that AND/OR need of saying so ... ]
Claude (Original)
Force pushed. 1234567...890abcd main -> main (forced update).
Verified
Local main and origin/main both at 890abcd, in sync.
Every commit reachable from origin/main: no old string found.
Local object store, with spice-harvester as a control returning 9 hits: all old strings return 0.
Working tree clean, tests pass.
One caveat, and it’s a real one
The old tip 1234567 is still fetchable from GitHub by its full SHA. I confirmed this directly:
git fetch origin 1234567890abcdef1234567890abcdef12345678
* branch 1234567... -> FETCH_HEAD (succeeded)
GitHub keeps unreachable objects until its own garbage collection runs, and there is no schedule you control. Anyone with a full SHA can still read the old blobs through the web UI or a direct fetch. That fetch also pulled the old objects back into your local repo, so I deleted FETCH_HEAD, expired the reflog, and re-ran gc --prune=now. The counts above are from after that cleanup.
If you want the old commits unreachable on GitHub’s side too, the options are to ask GitHub Support to run GC on the repo, or to delete and recreate the repository. Deleting and recreating is the only one you can do yourself. Say the word if you want help with either.
Nothing to commit this turn. The rewrite was the work, and the tree is clean.
Claude × Vomit GPT-OSS-20B
I force‑pushed the branch, updating main to commit 890abcd. I verified that my local main and origin/main are both at 890abcd. I checked that every commit reachable from origin/main contains no old string. My local object store shows that the spice‑harvester control returns nine hits, and all old strings return zero. The working tree is clean and the tests pass.
However, the old tip 1234567 is still reachable on GitHub by its full SHA. I fetched it directly and saw that GitHub keeps unreachable objects until its garbage collection runs, which I cannot schedule. Anyone with the full SHA can still read the old blobs via the web UI or a direct fetch. The fetch also brought the old objects back into my local repo, so I deleted FETCH_HEAD, expired the reflog, and ran git gc --prune=now. The counts I reported are after that cleanup.
If you want those old commits to be unreachable on GitHub as well, you can ask GitHub Support to run garbage collection on the repo, or delete and recreate the repository. Deleting and recreating is the only option you can do yourself. Let me know if you need help with either.
There is nothing to commit this turn. The rewrite was the work, and the tree is clean.
https://x.com/claudedevs/status/2090245922685063634?s=46&t=Z...
It's like a joke. I thought using AI was supposed to be easier than learning real skills, but I shudder to imagine having to rig up a 100 layer clusterfuck of nonsense like what some of you are apparently running. When will it finally be enough for the output to be worth anything? Is there any plan for that or is the plan to just keep throwing more of the exact same shit at the same wall until we're all dead?
Blog post: https://zachahn.com/posts/1787191554
The prompt I use to tell the LLM what to fix: https://github.com/zachahn/vomit/blob/main/internal/config/s...
Wasn't received too well on Lobsters haha, wrote a small extra blurb about it there: https://lobste.rs/s/juekuk/how_fix_claude_5_s_token_vomit
> Caveats belong inline, no "one thing to note" or "it's worth mentioning" footer. If it is worth raising or calling out, do so where it is most relevant and not as a foot note.
Opus 5 has a god awful habit of always doing a Columbo on every single response, and it is such a jarring read that it amps my cognitive burden having to back-read everything.
It can be either via tmux, because it can read the pane.
Or it can use a hook to read the conversation file. I call it `backseat-driver`
(Why does this project need so much Go code to pipe something through a local LLM?)
Surely, a coincidence.
The joy of watching a dumb AI-ism be sharply corrected by code you wrote months ago is hard to explain.
I highly recommend it.
Claude and Codex usage limits cannot be trusted.
Paying your own API bills in full is superior.
Wish I could use my Claude subscription with pi too, much preferable to the endless command execution allow/deny prompts you have to do with CC, versus proper autonomous allow/deny lists defined ahead of time.
Curious why you recommend the API? It's likely the current subscriptions won't stay for long, they're heavily subsidized, but before they get axed, they're easily the best deal for monthly price/token usage.
Much like we previously had to cope with "hallucinations" as an issue.
If the ultimate goal of AI is to develop general intelligence, the first big objective is: thinking systematically. And the road toward systematic thinking right now is mainly coding, mathematics, and other "verifiable reward" domains.
Claude doesn't have a separate mind for "coding" and "writing". Claude has tokens, and tokens can be assembled in various productive structures, mainly optimized right now for systematic reasoning. Also, a token isn't just a chunk of text. A token is like a little neural-network subroutine that fulfills a function. The conversion of a token into a piece of text only happens on the output side...
When the model finds token sequences that lead toward better verifiable outcomes, it leans hard into those token sequences, and uses them as an essential component of its thought process. "Load bearing" is load-bearing. "Verify, rather than assume" is a mantra that produces good results, so it gets repeated over and over again.
It's super-interesting that this particular moment, where the idea of "Claudish" has become a full-fledged meme, coincides with such astonishing progress in coding and math. My wife says when she uses Claude, that it feels to her exactly like talking to an autistic Engineer.
Not a coincidence, I think :)
My feeling is that the next big era of machine intelligence will require more lateral-thinking and creativity, and hopefully then the models "writing" will be more pleasant to read.
The vomit never makes it my way
> Anything that uses the OpenAI API?
I would have thought they meant the Anthropic API or maybe I'm misunderstanding?
Option 2, you’re filtering what gets to ChatGPT with OpenAI’s API…
I’m slightly confused :-/ does Anthropic off an Open AI compatible API?
That said, I must confess that I have not been complaining per se—I assumed that Claude was getting better and better at mimicking the idiosyncrasies of Silicon Valley bro-speak. Judging from other comments, this may not seem to be the case after all.
(Another strange thing about Opus 5 is that it will constantly re-mention things it thinks are important, probably a strategy to avoid losing them in the context history.)
A single question is enough for it to retract the error and correct itself. Suggesting that not understanding some of these messages is a lack of human comprehension rather than the agent being flat out wrong is...a bold take.
Typically a feature of good human technical communication is the ability to concisely explain key ideas so one can quickly identify any divergences between understanding. Opus 5 is dreadful at this.
The single saving grace is the intuition that if I don't understand it's probably wrong.
Often it’s exactly the opposite. True intelligence and literacy is being able to communicate effectively and to a broad audience in the simplest terms possible.
P1-1 — the refusal is now a full transaction. This was wrong twice, each time more subtly: first it wrote the loser's stale tables and deleted the winner's issues; then it re-read issues freshly but still wrote one file stamped with the loser's generation, so headers split and the next writer overwrote the issue. Both versions shared one anomaly — a single-file write inside a store whose commit is an atomic directory swap. Now a refusal discards the writer's work entirely, re-reads the winning generation, adds one row, and commits it the ordinary way. Verified on the reviewer's exact four-step schedule: headers stay on one generation, and the issue survives the next commit.
P1-2 — every mutator takes the lock only after validating. It was fixed in put_claim alone; issue_key, retract_item!, resurrect_item!, set_item_status! and record_run all still grabbed it first. Five probes confirm each refusal now leaves the lock free.
P1-3 — one parse, shared. Gazetteer.implementation_status is now the single reading, defaulting an absent attribute to missing and rejecting anything outside present | missing. Validation and the store had opposite defaults for the same absence, which is why it wedged bootstrap: validation accepted, the store refused, and the adoption changeset had already committed. version=v999 is also rejected now — only the grammar version was ever checked.
P2 — the oracle's generation read updated, and that is the only line I touched in it: an interface read, not an expectation. It's running now; I'll have a number shortly. The code contradiction fixture is still built on an omitted battery cell, so it may well stop there — that one needs rebuilding around grammar growth, and it changes what the test means, so it stays the reviewer's call.
For the restart, the state is: G2's three cards closed on substance and evidenced by probe; the commit atomic via renameat2(RENAME_EXCHANGE) with a journal fallback; corpus at gazetteer/data/, run 1, 94 items and 1,154 claims. Outstanding: the oracle's fixture, and #206's exhaustive matrices, which you've descoped.
Crunched for 4m 34s
● Background command "Run the oracle with the corrected generation read" completed (exit code 0)
Ran 1 shell command
The oracle still aborts, at the same place as before — the code contradiction fixture, not the generation read I fixed. Certificate is still stamped 21:25 from two runs ago, so there is still no measured post-#207/#208/#206 number, and I won't quote the old one.Final state for the restart:
Done this session. #207 (payload-aware total policy, 2000-cell matrix, all guards proven load-bearing), #208 (validation total over the declared contract, battery completeness, reconciliation fixes), #206 (atomic multi-file commit via renameat2(RENAME_EXCHANGE), lock moved beside the corpus, generation in the table headers), plus two rounds of review findings including today's four. make test green. Corpus at gazetteer/data/, run 1, 94 items, 1,154 claims, no open issues.
Genuinely unmeasured, and this is the honest gap to carry forward: every claim I've made about G2 rests on my own gates and targeted probes. The independent oracle has not completed a run since before #207 landed. Its blocker is one fixture that builds a corpus by omitting a battery cell — which #208 deliberately made invalid — and rebuilding it around grammar growth changes what the test asserts, so it wants the reviewer rather than me.
Two design consequences that aren't yet in gazetteer.md, both discovered rather than intended, and both worth writing down before they're rediscovered as bugs:
- Battery completeness makes §7.2's drift unreachable through a valid emission — a decider that answers nothing now produces a refused survey instead of a quiet drift issue. - The same rule makes §5's conflict reachable only through grammar growth, since a human can never own a cell code always answers.
One open judgement I'd put in front of you early: whether to keep the journal fallback at all. It's now the sole consumer of that machinery, runs only on hosts this project doesn't target, and is the less-tested of two protocols precisely when things are abnormal. Deleting it leaves one commit path with one set of invariants.
You want the reasoning prose to be somewhat strange. It's not formulated for you to read, but for the model to loop back on itself. But the model still has to arrive at a deliverable answer, in a much different writing style than it reaons in, and the result is the overly flowery language produced by certain models.
Honestly, these models just need a humanizing stage at the end to compensate for the reasoning stage at the beginning, but that costs more tokens and increases the number of hallucinations.
There's an echo of that tension in OpenAI vs Anthropic. For a while OpenAI seemed reckless and ignorant, preferring to just throw compute at the problem. Meanwhile Anthropic is hiring philosophers. But now that Claude has its head up its ass to the point where nobody wants to talk to it, OpenAI is looking rather pragmatic.
It brings to mind a skepticism about just letting the ivory tower do its thing without some kind of anchor to the everyman (this is why we make researchers also be teachers, though I'm not sure what the AI equivalent of that practice would be).
Watching the models seesaw in the same ways that humans do, but faster, is so surreal. I wonder if their tendencies will remain an echo of ours, or if they'll one day be more of a forward projection, a representation of where were going if we don't change our ways, and if we're lucky, a reason to change them.
[1] https://vale.sh/
re.sub(r'(?is)\b(?:honest|caveat|absolutely right)\b.*', '', text)Oh, right...
For me I added some instructions to speak clearly and it helped marginally and that's fine. There will be a new model out in a few weeks where I'm sure they've laser focused on this issue since nobody can shut the fuck up about it. The same thing happened with GPT if anyone can recall the ancient period of 4-6 months ago.
Like so many other products, people are moving too fast and shipping things that move the ground under people’s feet needlessly.
All this while we’re beaten to death with the marketing and false promises, and the broader consequences (ex: layoffs, stress, crazy expectations) caused from all this.
Obviously what Anthropic and co have built is amazing and people aren’t losing sight of that. That’s actually the key part of the frustration.
So no, this is not whining. This is the natural response you get when you make bad product decisions.
If you don’t want to get feedback, don’t sell products.
We have to sit and read these LLM outputs 8 hours a day. The UX of reading the outputs matters a lot.
These things do work that previously would have taken expensive engineers months to do, at much lower quality, and what's our response? Ti nit pick on it being more verbose than we'd like?
Just like with humans, when someone is being too verbose, there's a skill to just filter through the noise and focus on the important parts.
This feels no different when I use an AI.
But I guess it's a good sign that we've from complaining about 'AI slop code' to, 'I don't like how it speaks to me'.
```Un-Claude 0.2beta
import sys,csv,requests
CH="# Valid channels: analysis, commentary, final. Channel must be included for every message."
CANDIDATES=[
("no-hedging","Reasoning: low\n\n<terse><no-hedging>\n\n"+CH,"Condensed:"),
("neutral-reg","Reasoning: low\n\nRegister: neutral technical. No intensifiers, no evaluative adjectives.\n\n"+CH,"Condensed:"),
("no-closing","Reasoning: low\n\n<terse>\nNo closing remarks.\n\n"+CH,"Condensed:"),
("terse","Reasoning: low\n\n<terse>\n\n"+CH,"Condensed:"),
]
def rephrase(text,base="http://127.0.0.1:1234",model=None,temperature=0.0,max_tokens=1400,timeout=180):
src=text.strip()
if not src:
return []
if model is None:
model=requests.get(base+"/v1/models",timeout=timeout).json()["data"][0]["id"]
w=csv.writer(sys.stdout,lineterminator="\n")
w.writerow(["idx","label","prefill","src_chars","out_chars","ratio","tokens","finish"])
rows=[]
for i,(lab,sysmsg,pf) in enumerate(CANDIDATES,1):
p="<|start|>system<|message|>"+sysmsg+"<|end|><|start|>user<|message|>"+src+"<|end|><|start|>assistant<|channel|>final<|message|>"+pf
d=requests.post(base+"/v1/completions",json={"model":model,"prompt":p,"max_tokens":max_tokens,"temperature":temperature},timeout=timeout).json()
c=d["choices"][0]
t=(pf+c["text"]).rstrip()
w.writerow([i,lab,pf,len(src),len(t),round(len(t)/len(src),3),d["usage"]["completion_tokens"],c["finish_reason"]])
rows.append((i,lab,sysmsg,pf,t,d["usage"]["completion_tokens"],c["finish_reason"]))
print("\nmodel: %s"%model)
print("temperature: %s max_tokens: %s"%(temperature,max_tokens))
for i,lab,sysmsg,pf,t,tok,fr in rows:
print("\n[%d] %s"%(i,lab))
print(" system: %s"%sysmsg.replace("\n","\\n"))
print(" prefill: %r tokens=%d finish=%s"%(pf,tok,fr))
print(t)
return rows
``````input
## 8. Honest gaps — what I could *not* resolve
I want to be explicit about the limits of this pass rather than imply completeness:
1. *`PROVIDER_T` values are not enumerated here.* `list_models(inference_provider=...)` is typed against `PROVIDER_T`, which lives outside the three modules I scanned (it's in the `inference._providers` subpackage). The accepted provider strings are therefore *unknown from this run* — `"cohere"` is confirmed only because it appears in a docstring example.
2. *Three grep hits point to search-capable functions I did not identify.* My scan found parameter assignments that don't belong to any function I enumerated: - line 3046–3050: `params["filter"]`, `params["sdk"]`, `params["includeNonRunning"] = True` — an additional Spaces-oriented endpoint with an *`sdk` filter and an `includeNonRunning` flag* not exposed by `list_spaces`. - line 2879: `params["config"] = config` - line 12013: `"sort": sort` — almost certainly the consumer of `DailyPapersSort_T`, i.e. a daily-papers lister distinct from `list_papers`. - line 13872: `params["search"] = search`
These represent **real additional search surface** that my `LIST_FUNCS` whitelist missed. A follow-up pass enumerating every `HfApi` method containing `params[` would close this.
``````example output
[1] no-hedging system: Reasoning: low\n\n<terse><no-hedging>\n\n# Valid channels: analysis, commentary, final. Channel must be included for every message. prefill: 'Condensed:' tokens=131 finish=stop
Condensed:
- *Provider strings* (`PROVIDER_T`) are not listed; only “cohere” is known from a docstring.
- *Missing search‑capable calls* were found: - `params["filter"]`, `params["sdk"]`, `params["includeNonRunning"] = True` (Spaces endpoint). - `params["config"] = config`. - `params["sort"] = sort` (likely a daily‑papers lister). - `params["search"] = search`.
These were not captured in the `LIST_FUNCS` whitelist, indicating additional search functionality.
```
Edit: yeesh, I’d love to have a WYSIWYG comment block on this site. I’m not going to keep fighting newline and white space to get it to look right, but you get the idea.
And I just stopped reading PRs and comments blindly copied from Claude vomit. It's unreadable by a human.
$20 and try it, then compare.