Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.
And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.
(Makes sense to me, just curious about this additional piece.)
Looking at reserved capacity cost for PTUs on azure, which I think they’d probably not subsidize but can’t be sure, I’m inclined to not agree with the vast undercharging for tokens hypothesis.
And this one is easy to calculate: take your monthly API spend to K3, then rent a stack of 8xB300 for a month and see how much it costs. I would say you're about to save 10-15k dollars per month if you have enough traffic compared to pay per token pricing.
It's not very complex math, and the hardware of course is cheaper if you bought it last year and if you have extra GPUs waiting in your warehouse (depending on if you can produce enough energy cheaply).
But the biggest discount people see is subscriptions. You get a few thousand dollars of work from a couple hundred dollars.
Use only the deepseek provider, you can configure openrouter to do that(but they create some obstacles, go to configurations and allow all providers)
Right now i'm testing deepseek harness, don't wait to test it. The plugin architecture and self awareness of workflows really make you think about "what is a tool vs what is a project".
It has been an amazing experience
Well, DeepSeek just raised prices.
I assume it's highly use case dependent, though?
Even before the price cut seems like Sol was price competitive with Kimi
https://artificialanalysis.ai/models?models=gpt-5-6-sol-xhig...
And now it should be considerably cheaper
You cannot just look at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.
If there's independent data showing this feel free to share a link, I haven't seen it. DeepSWE has been most closely matching what I see in my own use.
It's not always Chinese models. For example GLM 5.2 just did not work for us at all. And Gemini is still the best cheap model for non-text agents.
If you don't have a good eval set and if you don't check the models weekly, you are missing on things. And Opus 4.8 is still the absolute quality king for agentic tasks. Too bad it's so expensive.
And the clearest thing here is that Fable, Opus, and Sol are all too expensive. I'd say a healthy 75% cut to token prices and they are back in competition.
Surely you don't want them to be the reason the bubble bursts?
Fable seems to be generally more impressive at outputting one-shot web apps. I'm not really saying that to try to downplay what Fable can do, it's just that if I compare the two, this is one of the few definitely noticeable areas that you can easily demonstrate. Obviously, one-shotting programs is much better as a demonstration of a model's capabilities than it is practically useful (not that it is useless, but hopefully my point is understood).
However, whatever Fable truly is better at, one thing I really like about GPT 5.6 Sol is even harder to quantify: taste. GPT 5.6 Sol outputs are still LLM outputs and they contain many things that people would probably consider "Claude-isms" for better or worse, but overall I really prefer the GPT 5.6 Sol output. I find it to be generally more tasteful. Hard to quantify, but when talking to people I've had enough people seemingly agree with me to convince me that it really is true.
I was very pleasantly surprised to find Sol wasn’t obstructive over what was clearly a very grey area endeavour.
Some of the topics it’s flagged have been hard for me to understand what it seeing that can be remotely concerning in my requests.
Sol is great and has never blocked a request, and generally gives great answers. Happily switched over to it now.
Which is definitely protecting their turf, but also probably a little bit hiding their “RSI” abilities for competitive reasons. My theory is that a lot of “safety blocking” is actually WIP training of new business directions. Anthropic has started hiring biologists and has opened a preview of a “Claude code for bioinformatics”. I’m guessing they’re tweaking their bioinformatics market play, and block “bio safety” requests so competitors can’t learn about their training.
That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.
Competition is competition and it’s two sides of the same coin.
There's people that have tried to contact Jeff McBride and follow the IP trail but the IP is currently owned by a company that went defunct. Not sold, but no one is even bothering to register its LLC any more, it's simply dead.
Compared to what I am doing at home experimentally, I feel like day-to-day work is absolutely nothing. Not only am I also working with existing codebases in my experimental prototyping, but I am also doing things vastly more complex with vastly harder constraints.
But I wouldn't trust lower tier models for end to end solutions.
Sol over engineers now and then (hey please don't factor that function out into its own file....) but it doesn't do the same level of stupid terra does.
That said, plan with Sol, implement with terra, have Sol fix all the mistakes, then I go over the code and make recommendations for the architecture to fix Sol's foolishness.
Getting the AI to output code that you like is difficult.
As an example, let's say in React you have a "useLocale()" hook.
The AI will happily pass down locale as a prop to 5 child components instead of just calling the hook in the component.
A review from another model did not flag such stylistic issues either.
I believe that the latest models are very good at functionally achieving the goal, but still have poor taste for UX or code quality.
The most productive use of AI for software development happens in an environment where you do not review the code but test the UX end to end.
I think it sometimes worked, for example for testing preferences, but sometimes it did not.
Could be a problem with the harness also.
In any case, I feel that it's a bit playing whac-a-mole with explicit rules for things that a more intelligent model should do by default.
MAI also offers a ultra cheap version that's competitive with Luna.
So much so that the models look like they were designed by a product manager explicitly to eat away OpenAI's market share.
Vscode even pushed them quite hard onto users with the latest release, going to the extent of putting up a modal to convince users to try them out.
It's just vibes
When people figure out any reliable strategies to test and benchmark them, that's insane, and in the positive sense. This very same issue has been a thing for humans as well forever, and remains only very questionably solved (IQ, academic tests). This is not easy.
Just consider your own example. Do you think a less or more "rebellious look" is not something designers can actually ellicit? Less so in software design, sure, but in character design for example? Or general product design? Do you think e.g. Monster energy drinks are branded the way they are completely due to happenstance or something?
Except people don't usually put numbers to it, because they understand that that's hard to defend. You're the one who's describing such an idea, and wants such a thing to happen, classifying anything else as just vibes (that's the point!) and unhingedness. You're handwaving the difficulty and fundamentally limited nature of that, assuming that it is some laziness or mental delusion that's preventing it instead. You're also pretending as if it was somehow not real as a result. What I'm telling you is that you're wrong about that. Any kind of qualitative analysis that's actually defensible with these is genuinely difficult and limited in nature. See also all the opining about benchmaxxing. It also doesn't mean they're useless though, see also benchmarking.
The guy above didn't put numbers to his vibe assessment, they just drew a comparison, exactly because they know that there's not much else they can earnestly offer. You're sulking at them not lying to you by overstating their rigor, and you're flipping the arrow as if this limitation was some sort of mistake, not a necessary and intrinsic property, which it absolutely is. Natural language is an inherently subjective medium.
[0] In fancier and more mathematical terms: https://abeljansma.nl/2026/07/10/truth-is-not-a-direction.ht...
Where you’re wrong is pretending there is any intellectual rigor to the discussion which justifies promoting from the domain of vibes to actual reasoned debate
To give you a practical example, small, self-hostable models are very popular on HN. But my personal experience with them has been absolute dogwater, so this difference informs me that this is not the place where I should shop for a signal on whether a specific model like that is worth trying, or on whether it's game over for large models and remote models yet. I can "take the temperature" and make use of that without it having to be any rigorous, high assurance or mechanistic thing. It's suboptimal, but not useless.
Conversely, it also inspires people to try and substantiate these issues, so that it can eventually be more rigorous, higher assurance, and mechanistic. This is why I brought up benchmarks and benchmaxxing. Hard to know what points of consideration are salient when there is no abstract grievance to investigate, but Goodhart's law does also keep looming.
If people start moving away from Claude to GPT because e.g. Claude's output is too hard to work with and parse, that's relevant for the respective model providers, because it's a revenue share shift.
If I really like using Claude and strongly prefer its output stylistically, then claims otherwise will infuriate me, and will drive me to substantiate. If for no other reason, then because it's likely that Claude will have its language tuned in response to people's feedbacks and the revenue shifting, which may not be to my liking, and so I better prepare to call it and contend it.
It's literally like any other real world thing ever. Think tuning video codecs. Or tuning user journeys in frontend development.
But human coders can have bad taste too. There is code where there is nothing obviously objectively wrong, yet the choices feel like they were made by someone who just doesn't value or put emphasis on the right things, yet spends a lot of effort on trivialities. It comes in many forms.
To be honest, in this case, I wasn't even personifying them, because I in fact didn't say that an array of GPUs has "taste", or in fact even that model weights did. I was saying I preferred GPT 5.6 Sol's outputs as a matter of taste. I can see why someone would confuse the two statements since in this case they're pretty much the same thing, but if you re-read what I said I was actually more careful than you're giving me credit.
It's... really just vibes?
Always has been.
The only obvious objection I can think of to this line of thinking (at least from a technical perspective) is "how does someone build up the knowledge to be able to use a tool effectively in that way if not by doing things by hand at first?" The honest answer that is "I don't know, but that's also pretty much exactly the type of thing my employers have never been paying me to solve in the first place". Even just a decade into my career, there have already been plenty of times in my career I've struggle to convince people that we should do stuff in a way that won't bite us in the ass a month or two down the line, and in the times I've managed to succeed, it's usually only by putting in more of my own time and effort to make the initial investment seem more palatable. Luckily right now I'm not in one of those times when I'm having to go full throttle to keep the lights on a few months from now, but I don't have enough fuel in reserves to work on a plan for when we need to build a new rocket in another ten years. Maybe ask me next month.
After extensively using both on Max 20x plans, I've concluded that Fable is better for problem solving and coding, whereas Sol 5.6 Ultra shines in debugging specific issues: tackle a problem with Fable then leverage Sol to clean up, double check, or fix specific issues.
Fable (imo) had the edge on the $200 plan, but after this 50% reduction I'd say Codex is better value by far and there's no contest.
---
Using Fable as the orchestrator and delegating tasks to Sol 5.6 Ultra via the codex plugin in Claude Code yielded good results, but still there was a lot more over-engineering (thus time and tokens spent) than Fable by itself would've done.
Both models suffer from doing-too-much. But both models are fundamentally really smart and knowledgeable. I think it's really close and pricing cuts really spice things up for us consumers! Sol is a clear winner in the value department and the $100 plan is enticing!
---
*Claude Code usage is reducing by 33% in 2 days, Wednesday August 19... cmon anthropic: clau.de/cc-50-promo
However, I think these are very different models in terms of orchestration. Long-horizon tasks are way more predictable with Fable. It just doesn't lose track of details. Thus I ended up building a small wrapper around Pi (where I run Sol) so that CC can delegate via background tasks, automatically wait for completion, and do what was one of the most effective parts - steer Sol toward simplicity, getting Sol out of code-review infinite loops (Pi calls for Codex review to ship better, but generally gets stuck on P2 and results in vastly overengineered work).
One of the worst experiments was enforcing coverage at 100%. Only Sol, with an enormous amount of code and significant pushback (on architecture decisions) to Fable, was able to reach it. It made me think this is somehow related to overengineering in general, so that instructions on acceptance criteria in claude.md plus proper DX (e.g., Lefthook) actually led to okay results. It mostly helped that responsibilities were clearly split: Fable designs architecture, Sol handles coding and debugging.
Remove the one thing that differentiates these 2 SOTA models (long-horizon task scaling) and you're left with assessing the raw intelligence of both models.
Both models are really fucking smart. We should be intentional when discussing effort levels when it comes so SOTA models because currently, effort levels are the essential lever to evaluate task scaling.
I've never seen that leaderboard link but I think my sentiments reflect exactly the findings: Sol 5.6 is smart, Fable is smart, but Sol is more value for the end user (even more so when you lower effort levels because that doesn't degrade the model's raw intelligence/knowledge). Not to mention Codex resets, that's just the cherry on top!
But smart != capable and this is evident once you start assessing both models on long-horizon tasks with higher effort levels. While Fable is (imo) at least a little bit better, both are still very good. If you want to test the raw intelligence of said models, you should lower the effort level.. if you want to test the model's capabilities fully, you should increase the effort level.
Economics aside, Fable is the better model (imo), but there's no need for a binary stance here. Both models are very good yet there is a clear winner on the value front.
Is it more about just avoiding any mistakes? Seems like that would be costly when medium or high would work fine?
For small tasks, you can just use something like low or medium effort and it can usually avoid mistakes; after all, the model will test the code anyways and can do some baseline level of iterating.
In regards to cost, we need to acknowledge how generous OpenAI was in the last couple months with Codex usage credits (no weekly limits) and usage resets. It afforded me many a dive with Codex! Yes it uses more tokens, but sometimes it's worth it -- just depends on what you're working on.
Finally, Ultra(code) isn't that bad when it comes to cached tokens. I think folks overstate the general token usage of ultra effort on both providers.
---
Both models are great at green-fielding a project when given detailed specs.
Both models overthink too liberally (imo) during these larger multi-shots. Sol overthinks more than Fable.
Both models are really smart and perform great for general knowledge and regular coding tasks.
Given the 50% discount on Sol and how smart it is, yeah it's unprecedented value. If you only want to use low effort, there's a clear winner here on value and it's not even close!
I'm in Australia, and Fable downgrades to Opus when testing for bugs in memory in a legacy C code base. If Fable starts taking initiative and writes a test case that involves writing to a null pointer, that's the end of the conversation.
What on earth are you asking it?
Also Opus 5 is fine if your codebase is simple.
> even after completing the verification program
Was it easy to complete it?
I ended up in some weird state where I can't even attempt the verification at all. Opened the Persona tab once, closed it and then it never opened ever again. It says a verification precheck failed.
Even without TAC, Sol doesn't seem to get blocked very often. Fable would downgrade to Opus if I looked at it wrong.
I hear you on the downgrades, I'm 13/13 on downgrades, and last downgraded me to Sonnet for asking for reasoning chain.
Its a crutch that is no longer competitive
It the first model to actually make me pissed off to use AI. I absolutely hate the model so much.
I don't even want to see the codebases this model is fucking up.
It might just be good at finding bugs that about it. That all I would ever use it for just because it works harder than Claude models.
I will not be told what I can and can't do by AI and I will no longer be supporting American companies run by despicable people. GPT only gets my money right now because its so fast and cheap but I'll be back to Chinese models in no time.
I just can't get Opus (Opus 5 is dumb as a rock, to be fair) or Sol to do that, so I exclusively use Fable for personal work. When I reach my weekly limit, usually on the last day close to the reset, I just go back to coding by hand ¯\_(ツ)_/¯
Heck, it even does security reviews and fixes, as long as I don't ask it to "attack" the codebase. I'm planning on using Kimi or GLM for that part.
Sol also doesnt _really_ work but it sort of tricks me into thinking it does more convincingly :p.
cancelled my subscriptions few days ago. (was on 100$ ones, not sure if there is diff in quality for higher tiers or not.. there might be that too).
what i hate the most is that they will make any obvious mistake you do not tell them to avoid. then on the next plan to fix it, your token limit is hit at step 4/5 -_-. Both models seem incredibly good at that mostly...
for tasks outside of coding and program design i do find them quite useful. like devops crap. maybe because i hate that, i like their help there more.
They STILL don't have an option to "Sign in with Apple" on the website, but they do for Google??!? (and on iPhone of course)
Screw that asinine UX
(and no it wasn't better than Codex at this particular task)
Can somebody at Anthropic tag claude in slack or whatever goofy shit you do and ask it to add Apple OAuth to your website? Clearly humans aren't testing it.
and sure enough, I was right to do so: They don't even let you remove your payment method afterwards. Every other store, Steam etc., lets you.
No way I have enough trust to install their desktop app after that, so I just want to try it through their website..
Can Sign In with Google, but not with Apple
so you gotta open the Passwords app, copy your random email, paste into the website, then copy the OTP from your email..
It's been that way for at least a year
The desktop app was clunky too the last couple times I tried it a few months ago
and the AI itself hasn't been that hot compared to ChatGPT/Codex either: https://i.imgur.com/jYawPDY.png
So all the Claude hype posted on HN seems like a case of the emperor with no clothes to me
(P.S. The thing I just now tried to do on Claude hit the weekly usage limit after 2 minutes)
I have witnessed 5.6 Sol Ultra edit line after line of literally empty lines ... for hours.
I wasn't literally watching it, I came back to a goal (that it started for itself without my approval!) that had done nothing but that for some reason.
It couldn't explain why it had started.
Maybe they want to see how much market they can grab with Sol?
This might help but there are already cheaper models with Sol's intelligence more or less, the most notable being Grok 4.6 at $6/m which makes it a tougher sell
It's now my open weight inference provider of choice, since on top of the privacy/security characteristics it's also reasonably cheap.
tinfoil asks more than 10x the output cost ... $1.90 per 1M tokens instead of $0.18 per 1M tokens for my favorite model (Deepseek V4 Flash 0731) on my favorite provider (DeepInfra) currently, for example.
Tinfoil OTOH claims hardware attestation and confidential computing, which is a pretty strong and (in theory) verifiable promise.
You absolutely would not recommend tinfoil.sh to someone indifferent to privacy trying to save money.
It's a very different service than OpenRouter.
Batch API is a bit harder, not many models/providers support Batch. It primarily is only the Gemini/ChatGPT/Claude models that do. DigitalOcean does support a 50% discount on Batch API via them directly, not listed on OpenRouter.
[1] https://docs.digitalocean.com/products/inference/how-to/use-...
I am quite happy with them so far.
For Mythos and even Fable they require prompt retention on their end.
edit: or more precisely if you want to access Mythos/Fable ZDR does not apply, and depending on config the exclusion can affect other models.
But if their employer is bringing millions of dollars of potential spend to the table, can tell you from years of experience that turns a lot of 'no's' to 'yes'.
You do not have ZDR with fable/mythos, nor will you ever have it again. The bedrock customers who thought they had ZDR learned rather rudely that this was not the case after the government decision that also banned non US citizens from using these models.
I'd be surprised if they closed out over 100B tokens today (they did 101B on sol 5.6 yesterday).
If Sol isn't the best model, it is up there...
You don't cut the price of the best model for no reason...
Always has been. My prediction is that both OpenAI and Claude will go bust unless they deliver a killer product. And unlike scrappy startups, they have a pretty serious deadline because creditors will come a-knockin'.
There's little to no functional difference between Kimi, Qwen, Sol, Opus, etc. All flagship models are within like 1-5% of each other and the real moat will be what's always been the hard part: making a good product.
Don't know about that.
I'm using code review of my lone lisp project as a benchmark. It's a massive parallel code review where a coordinator cuts up the codebase into sections and dispatches agents to consider each part from different perspectives like quality, maintainability, consistency, correctness, rigor, etc.
Ran a complete Fable/max code review. Took over a month on a subscription. Now I've switched to OpenAI and am repeating the exact same review with Sol/max.
It's still not done yet but preliminary findings suggest Sol can only reproduce 70-90% of Fable's findings. So I think these models aren't as close as we've been led to believe.
My methodology consists of launching a 242 cell parallel code review matrix and committing all Fable/Sol max effort agent prompts and their full reports to a private orphan branch on the repository. This is the part that is taking me months to complete. This thing can kill my $100 subscription in about 12 hours.
When done, these raw findings will be semantically deduplicated and merged into a list of findings per model. This list will then be audited for hallucinated or otherwise made up findings. This will refine the list, and hallucination rate is its own data point. I'm also counting things like cybersecurity refusals and downgrades.
When all this is done, I'll analyse the final results and publish them on my website.
Some preliminary analysis:
Which code review lenses were the most valuable, where value is defined as number of serious issues identified? Rigor, followed by tests, robustness, correctness, and so on. I was able to create a tier list of reviewer personas using evidence! I can now run focused code reviews using the highest value lenses.
What's the most expensive code review? Correctness and rigor, of the lone lisp machine specifically.
API costs per finding? $0.91 to $3.77. API costs per serious finding? $8.94 to $28.51. All Fable.
How long did it take? 28.2 calendar days, 66.3 agent-hours.
Is it worth it to run the code review multiple times? If a review matrix's defect capture probability is 57%, then a second run captures 81% of the estimated/extrapolated defect population, a third run captures 92%, a fourth run captures 97%, and further runs yield severely diminishing returns. Probably worth it to code review important stuff three times.
What's the impact and cost of the safety classifier? Out of the 52 Fable review cells that triggered the safety classifier, 35 died without producing any output whatsoever, so 67.3% of the cells were a complete waste of tokens. 25% produced at least some output.
Does the safety classifier trigger most often on the important code that actually needs SOTA models? For the most part, yes. Fable was most often barred from reviewing the most important and complex files in the codebase, such as the virtual machine, the parser and I/O layer. These files also have the most CRITICAL+HIGH severity findings. Only a couple outliers broke this pattern.
https://www.cnbc.com/2026/08/15/anthropic-revenue-jumps-to-o...
The Chinese models are cheap because no one is using them. But they can't actually afford (or have capacity) to serve enough people to kill the giants. This is evidenced by them all recently hiking prices or limiting usage.
It's possible that they build out in China at an unreal pace, China doesn't have concept of "community input" to drag down state projects, but then you are left giving your IP to China. Just ask western hardware businesses how well that goes.
If compute is the moat, that doesn’t really make their position any less precarious.
They're guaranteed to get over whatever hump you think they're in unironically. Uber/Tesla have been in far worse situations and despite Elon being an idiot/liar you see how they performed when even the most bullish of investors called for their heads
I just think the whole “moat” discourse is silly. OpenAI and Anthropic’s success depends on the same thing every business’s success depends on: their customer base, and their continued delivery of services their customers want to pay for. Not their tech or their compute or anything else. They have no moat because moats aren’t a thing.
Claude already has a killer product (claude.ai/chat is a Swiss army knife) but just relying on people typing stuff into chat is not enough to sustain the company.
The other strategy is entrenching yourself as the LLM of choice into existing products (like ChatGPT is on Apple products).
Depends on your use case. the Chinese models are not there yet.
This is basically undercutting KimiK3 and Grok 4.6 where previously utilised gad soke advantages but was a step more expensive
That shows a bunch of models, including Sol, with a discount. None of them say how long it’s for, but I’d assume in all their cases it’s for a limited time as the banner said, and only on OpenRouter.
I posted elsewhere, but the Azure uptime & performance for Sol is truly dire. OpenAI is offering 5x faster latency, 4x faster tokens generation, and vastly better uptime (Azure US has only 87% uptime), all for 50% of the price now. I assume the pricing is to compete with other shiny new models (Grok, Qwen etc), but it might also be to cut-off a truly poorly performing Microsoft hosting experience.
OpenRouter attributes this promotion to OpenAI https://x.com/OpenRouter/status/2089416739398254662
What's the incentive here?
Open Responses API doesn't appear to support state management (yet)
So yes, presumably a very small share of their total traffic.
Their api pricing is absurdly expensive.
I assume at this point that it subsidizes subscriptions.
I've gotten more work done on a second chatgpt pro $100/mo subscription than I did with ~$150 of paying for usage through the app.
Privacy, experimenting with ML and "unorthodox" needs are currently the only acceptable reasons to do local.
Experimentation and privacy are definitely advantages, but it's also quite a lot of fun.
Also if you don’t specify, most end up being the same as the parent model which is pretty wasteful.
I engineered a skill that spins up Terra High agents for most sub-agents, resorting to Sol Medium for technical research and Luna High for code/in-project research tasks.
On a slightly different topic, Luna Max is incredibly capable and doesn’t use as much quota (Luna tokens are dirt cheap).
No one can judge the enjoyment, learning, and hobby aspects. Just wondering if there is an end goal for that much overall expenditure (time, money, energy, etc.)
I think on average AI energy usage is not as big a deal as everyone is panicking about, but your usage is truly absurd and I don't know how you can live with that. It's immoral.
As for co2, it depends on the provider, it could be way lower as well.
As for ethics, you don’t know what he works on, and how effectively - he might be saving 10x that much of co2 for the planet.
My (only somewhat facetious) opinion is that physicist access to programming languages should be controlled like doctors' access to opiates.
1. Why did you take from my OP that I tell codex "write a climate simulation code, make no mistakes" and go suntanning on a beach in the tropics for the rest of the semester?
2. Perhaps you have a different experience from me in writing HPC codes, but my experience is that > 90% of the code is boilerplate. I find GPT 5.6 can be prone to overengineering, but with a little steering and good judgement it generates very nice interfaces and high level code. I just have to think about the solver structure or metastructure.
3. Even with core numerics - pre-AI, it was a bunch of iteration going back and forth between code and optreports. Now codex will just do it. I suppose this may seem grim to you if you loved decorating every variable with !DIR$ ASSUME_ALIGNED, and manually batching array operations or whatever, but I didn't and I'm glad I no longer need to.
4. I'm now highly motivated to write tests, and AI makes it way easier to write the immense boilerplate around good tests (sorry not sorry, my {FUNDING_AGENCY} program manager doesn't give a flying fuck what my test coverage is, and my next grant won't depend on that in the slightest, so pre-AI I did the bare minimum. You can argue that the results will be worse, yadda yadda, but the incentive structure that {FUNDING_AGENCY} has in place don't promote good software standards, and my career never suffered for it)
5. I can generate docstrings with high accuracy (see the above)
I’m guessing that Wh/token estimate is several orders of magnitude too high.
Leaked financial documents from 2025 show the company reported an operating loss of approximately $20.9 billion against $13.1 billion in revenue.
Is any amount of tokenmaxxing moral?
Do some people still deny you can do a shit ton of work with AI?
I ended up borrowing my gf's phone number just so I could get access for work. Ridiculous
```
This changes how Claude Code communicates with you
1. Default Claude completes coding tasks efficiently and provides concise responses
2. Proactive Claude executes immediately, minimizes interruptions, and prefers action over planning
3. Explanatory Claude explains its implementation choices and codebase patterns
4. Learning Claude pauses and asks you to write small pieces of code for hands-on practice
```You can set up a custom one under `vi ~/.claude/output-styles/eli5.md` and `eli5` would show up in the list above:
```
---
name: ELI5
description: keep it simple please
keep-coding-instructions: true
---
I can't process lengthy text, talk to me like I'm 5.
Small words, short sentences, short paragraphs. If you have to use a big word, explain it right after. Only return what's actually necessary. Just tell me what you did, did it work, what do I do now.
If I have to decide something: show 3 options max, the context I need to pick fast, and which one you'd go with.
Keep paths and commands exact. I have no brain cells left for the rest.
```
I'm not sure if it's a good thing that there's a config option for this, it kinda means they know there's issues with verbosity.
God knows I've tried. I've got a variant of the ASD-STE100 trick which does the job, mostly, at the start … but get to about 100k of context and it goes out the window.
The model's personality is too strong for simple suggestion, alas.
Codex has always beaten claude in coding benchmarks, hasn't it?
Do you include research and training costs? Of all models or only the ones being served? What percent of the R&D budget do you allocate to inference? What about data center capacity? Do you count future commitments? All the circular financing deals? Do you count employee equity grants as costs? At what valuation?
We also have another solution for "whatever accounting decides": generally accepted accounting practices. It's far from perfect, but GAAP figures are what you should be looking at; not "adjusted GAAP" or whatever invention.
OpenAI's docs still show non-discounted pricing https://developers.openai.com/api/docs/models/gpt-5.6-sol
https://vercel.com/changelog/gpt-5-6-sol-is-50-off-on-ai-gat...
You're literally encouraging someone else to come in and steal your customer base,
Multiple times a year, retailers here in Australia have co-ordinated sales on Apple products. Apple.com or their retail stores don't have these sales.
But they're clearly Apple-funded when competing retailers launch the same sales on the same days; and the margins aren't enough for retailers to take a loss.
It also seems to be providing a vastly better user experience - Azure has less than 99% uptime (Azure USA only has 87% uptime), latency of 20 - 30 seconds, and a mere 8 tokens per second. OpenAI is offering 32 tokens per second (4x faster), 4 seconds latency (5x faster), and all for half the price of what Microsoft is charging for a vastly inferior experience.
Data taken from this page:
Has OpenAI struck a deal with openrouter and that's why we're seeing preferred pricing?
Is openrouter taking a loss on sol API calls to grow adoption?
How temporary is the reduction in price?
Overgrown datacenters or mounds of GPUs dumped into the harbour next ?
This is what happened after the great crypto GPU dumping.
Indeed, there is a point where the cost of disposal is higher than the expected market value of components. Thus, the asset turns into a liability if held too long.
I have seen factory liquidations, and everything goes... right down to the bolts in the floors. =3
Hopefully we can look forward to all that useless datacenter AI crap gets repurposed in a similar manner into something actually useful for users.
https://www.youtube.com/watch?v=rE75WvOtcu8
The Shrek movie market correction correlation may be due again in July 2027. =3
One person can use as many GPUs as they want.
That’s why Chinese models are gaining traction and it’ll be the only way for OpenAI or Anthropic to keep up.
(there's probably going to be a reply about 'but how can you trust them'; I'm just stating what they say)
If this nudges Anthropic to give me more Fable usage, that's even better.
Fwiw, you could do this with any small or medium model, and it's easier with the aws-docs mcp. AWS is pretty stable, well documented, and programmatic, so most AI can figure out what it needs pretty quick
What I can say is that over the course of ~1 week I was able to review, plan, implement, test and release a significant change to a production system using Code Sol max. Any other claim about how any other model might have completed the same task is outside of my experience.
I don't get this thread.... Really. Is it full of bots?
However it's ignorant to think that there aren't bots on HN and especially for the very many motives people have for swaying public opinion.
Paper: https://arxiv.org/abs/2603.07267
FWIW, there’s not that much value protected here anyway IMHO, and even raw thinking text can lie (as shown by Anthropic’s amazing research), so for legitimate interpretability research it’s limited.
Scaling frontier performance hasn’t been SFT-bounded for a while now; it’s now basically how much you can scale RL rollouts.
Here are all the providers giving discounts: https://openrouter.ai/collections/discounted-models
Another thing some people don't notice is flex pricing, which is way lower than default pricing, for slightly worse latency and reliability. Depends on the provider and model
They're reacting instead of leading, basically.
Cutting API prices 50% while millions of your paying subscribers have had their limits slashed and are all literally looking at the salivatingly-cheap chinese API prices availalbe on openrouter...
Not only did OpenAI and all of their cash somehow MISS the opportunity to purchase OpenRouter ...
Now they're giving a discount on an API that nobody even uses (get real, nobody's paying API prices to OpenAI ...
I calculated a 5.6 sol coding session the other day ... $680+ USD ... and it actually destroyed the codebase it was working on during that session).
Needless to say, I will not be spending another dime with Codex or OpenAI.
This entire Codex reset limit fiasco has taught me they are not to be trusted.
Deepseek, here I come.
> I calculated a 5.6 sol coding session the other day ... $680+ USD ... and it actually destroyed the codebase it was working on during that session).
Yeah, sorry, skill issue. If you let AI run wild (if one session is $680 yeah it ran pretty wild) don't complain how it messed up your codebase.
Not the OP, but I consider myself a skilled and heavy AI user, with multiple subscriptions in both platforms, plus OpenRouter.
A couple of weeks ago I gave 5.6 Sol a small/medium sized ticket to simplify part of the auth system. The ticket had a lot of details another Sol agent had collected during an exploratory session, and it was all vetted by Opus 5. I thought to myself that the implementation agent should have everything it needs. I still had it write a plan just in case, read the plan, made sure it matched the ticket, then clicked Approve and walked away.
I came back later that afternoon to a horror show. The agent had written 25,000+ LoC in the worktree. After 15 minutes of skimming through it, I realized that it had made the specced change, then convinced itself that it needed stronger verification, and over a series of compaction cycles ended up writing a static analysis harness so that it could prove that the change would be safe. Total bonkers.
Except, according to another Sol agent I showed the worktree to, the harness didn't actually do what the original agent claimed. The review agent said 98% of the worktree's code should be thrown away, and only the fix and its relevant unit and integration tests should be retained. I also asked Opus 5, and it theorized that 5.6 Sol must have gone through too many compaction cycles and lost track of its original goal.
This never happens to me with Claude models. Yes they write a lot of code and verbose comments, but I've never had a situation where a ticket that should take several hundred LoCs ended up with tens of thousands. When Claude overengineers something, I catch it during the planning phase, and it implements plans faithfully.
5.6 Sol is simply unreliable. It's too relentless and doesn't know when to stop. That's probably what caused the OP's $680 incident. I find it fascinating that people like it so much.
With models a commodity at this point there isn’t much leverage for the big labs to keep their pricing anywhere near where it’s at. And that’s at the worst possible time as they need to be dramatically raising prices to have a viable business model.
Expect pricing to rapidly fall towards the underlying cost of compute and as players get really desperate we’ll likely see inference at less than the cost of compute as the market starts to rationalize and squeeze out weaker players who’s only play left will be to be the cheapest option in town.
The AI bubble is just waiting for the first player to scream mercy and cut capex as they simply can’t afford to throw more cash on the burning pile. That will be the trigger that implodes this bubble.
Deepseek v4 and Kimi k3 have tightened the screws on the frontier models. With open-weights, anyone can host these for cost of compute. So, there is zero leverage left for the frontier labs.
What implodes it in my view is that enterprise adoption will stall. Enterprises are struggling to actually use these things in real workflows outside of coding and support.
People actually have to select and want to use Sol 5.6 in their routing.
I’d have thought that even today people would validate a number of models for certain tasks and on a daily basis go for the cheapest provider when they run that task.
Maybe I'm wrong, but "reasoning as a service" is looking more and more like a... commodity.
Their token usage on 5.6 Sol isn't even expected to double today.
I asked it to write a user todo and it turned out a four page essay. I gave the same task to 5.4 and got the small list of checkboxes I expected.
Then I switch models (to luna) before implementation. I find this combo nearly always does what I want.
I also use a skill called ponytail, its goal is to keep things terse and edits small. It may have contributed to the successes above.
I like that skills are easy to try out, too.
I agree Luna is great for task execution, either as a sub-agent with Sol planning and coordinating or if the task is well defined and straightforward, but there are lots of models now that you can say that about.
And yes, I sub in other coding models like DeepSeek. Mostly with good results.
I find that I get exactly the effort that I asked for, which is pretty nice. The other side of that coin is that these are the least lazy models I’ve used so far. They will go on elaborate tangents to complete the task when I want them to.
But otherwise I don’t use Claude anymore.
At this point, I'm considering going back to cursor over codex due to the ability to get more control over what model I use since there is clearly a heap of user preference and having frontier providers constantly shift the goal post with "State of the Art" is complete non-sense.
The TypeScript code which was transpiled into Rust (and is compatible with most hugo templates) runs faster than the original hugo.
[1]: https://github.com/tsoniclang/tsonic-examples/tree/main/rust...
The transpiler is still WIP, but the fact that it can do this says a lot of about how far LLMs have come.
Its like rehiring an employee every few months then training them up. Its honestly tiring and cant stay like this.
Opus has the same problem too…
This is the csharp target for tsonic (a TypeScript to C#/Rust/Python/Triton transpiler). It has a bunch of tests here: https://github.com/tsoniclang/tsonic-csharp/tree/main/test
More comprehensive e2e proving grounds are at
1: https://github.com/tsoniclang/proof-is-in-the-pudding
2: https://github.com/tsoniclang/tsumo/
They were built specifically for testing the C# target. There are several other large projects we built specifically for e2e testing.
But more interesting would be the tooling built to support this. For example, our current TypeScript parser [1] is a file-by-file port of Microsoft's TypeScript V7 compiler written in golang. The challenge here is that every time Microsoft changes code, we'll have to fix our code and tests. It's doable, but a fair amount of work.
So we decided to write tooling to transpile Microsoft's v7 compiler from golang, and autogenerate our compiler. That tool is called gotots [2] - and it already produces a fully working TypeScript compiler. It's 3x slower than TypeScript v6 compiler, but we hope to get to rough performance parity in a week or so. Everytime Microsoft makes an update, we run gotots and our parser gets updated as well.
[1]: The old parser - https://github.com/tsoniclang/tsts-legacy
[2]: Golang to TypeScript transpiler - https://github.com/tsoniclang/gotots
My general point is that tests and tooling is tremendous value, and they are guardrails for LLMs to converge. I could have, for example, chosen not to write the go-to-ts transpiler, and live with porting Microsoft's parser line by line. But making such tools is something LLMs are good at, so it's a tradeoff well worth making. And the upside is that you don't have to use LLMs to port Microsoft's parser/compiler (a large and complex project) line by line.
if you want to solve basic problem then use Luna
I’ve used Claude exclusively for the past few months
Was excited when Sol came out a few weeks ago and loaded it up
I made the mistake of treating it as if it were Claude - I’d assumed they were close enough in ability and treated them that way
Well, turns out my instruction sets for Claude are 100% too complicated for Sol
Sol made the stupidest assumptions, constantly did things that it wasn’t asked to do and always approached code in what I considered a weird way - I had redo a lot of my prompts to get it anywhere close
Now, did it do good work?
Yes, on occasion. But with LLMs and coding, consistency is the name of the game. Constantly having to correct the LLM and constantly feeling paranoid that it won’t listen makes for an exhausting session
Maybe if you “came up” in the codex world you’re more fluent with it, but sticking with Claude for now
Didn’t mean to make you upset sorry
Kudos to you though for being your authentic self so publicly
Is the HN community just too online and sucked in to the musk mind manipulation vortex? Or what is going on? Why does nobody seem to care?
Anything else they don't save it. Even if they tell you the model provider saves your data for training.
I’d bet that explains this move!
If you remember programming language discussions, they are exactly like this.
Software development is still in the leeches and bloodlettings phase.
It's a bunch of bread bakers talking about wheat suppliers.
Fwiw I love K3 and use it as a daily driver. I haven't tried Sol, as I dislike OpenAI.
So this isn’t really a price cut. As to why, lots of possible reasons. Perhaps an agreement with OpenAI to help them drive up more diverse traffic priorities.