PR open counts, issues created , ceos/founders spending more time on linear don't automatically lead to better outcomes (in my experience they are often negatively correlated:-) )
Even that's hard. There aren't enough signals to attribute changes directly to AI, so these apps seem to correlate the signal that the user was interacting with AI to the changes they made e.g "Bob used AI at that time, and they opened a PR at a similar time, so Bob probably used AI to make that PR."
Until the tooling for gathering data on AI usage improves the data will be fairly interesting because correlations often point to something related, but won't be a source of truth.
My guess is some will jump up (where the team has managed to make themselves move effective and responsive) and many will plummet (doesn't need explaining).
Then there might be something to look at.
Very few businesses can accurately attribute customer value to the work they do, especially once they're passed start-up scale. A mature company makes lots of small changes and they're rarely measurable.
Skill/Hook usage rates, budget spend, auto-approval rate, focus area heatmaps, MTTR, MTTD are all there now
Also, somewhat amusingly, the “founder” consistently being at the top of the use curve might just have something to do with everyone else also using it, but that more implies people are using it because the chief at the top wants them to use it… not because it’s actually useful. A pattern that’s typical of bad AI deployments.
This sounds possible but how can we know? It could just as well correlate. Effort of all kinds correlates with success even when it's not obvious or not... linear.
I'm sure someone is going to reply to and say "actually it's not clear and the tech debt is about to explode everything real soon now" or something to that effect, but it's hard for me to explain this theoretically when it is so obvious in practice in my day to day.
ROI is about how much its costs vs the value of the “help.” That’s where the present crisis is. Is the ROI there to pay for the trillions in commitments that have been bet on that ROI? That’s a clear no at this point hence the growing panic.
That's only a potential crisis for those who have made concrete investments.
On the use side, the 'cost' of AI spans more than two orders of magnitude. Looking at recent models (<6mo) with reasonable performance (intelligence index >= 45) on OpenRouter, the output cost ranges from $50/MTok (Fable) to $0.153/MTok (DeepSeek Flash 0731).
From the perspective of a user of LLM/agent assistance, there's very likely a range where the benefits outweigh the costs.
If the ROI for the model developers isn't there, then that just impairs the future trajectory of the field. Current models are just bits that aren't going anywhere, and as long as they can be served (in inference) above their marginal cost they will continue to be so-delivered.
All other things being equal, increasing the speed of some part of the development process will increase the overall pace of development. However, By Amdahl's law that increase will be sublinear, and that is why we should take "pull requests" as an imperfect metric.
We also don't get to pick the form that 'better technology products' take. While we'd probably like to keep cost(/effort) and complexity constant and increase robustness and performance, the market equilibrium might be 'worse is better' and reward whiz-bang features and lower effort.
Right, but we didn't invest a trillion dollars of capital into any of those things in the span of a couple years, thus forcing them to capture value and show a return on such a massive investment.
Also, large companies are still experimenting on how to integrate LLMs into their workflows. Due to the fast cadence of releases, people forget that LLMs became robust (regardless of the capability level/parameter count/data size) enough to use semi-reliably in company-specific ways only 1 year ago. And bigger the org, the slower the process. I don't expect it to settle and get productive used across the majority of very large companies for another year atleast.
High prices are mostly a result of DC capacity. As more and more DCs get built out, prices will drop. At a unit level the inference business is extremely sound regardless.
I recently set up a local llm to see what the fuss is about and other than the few ringer solutions my experience is as you described. 2min promping, 5min waiting, 3hrs debugging or just doing it myself.
I am very likely doing it wrong, and it does speed up some aspects, but I wouldn't say I trust llm code any more than my own. Until it runs and throws an error, the llm is 100% confident that it has written perfect code.
So I asked Claude to first figure out a way to instrument GPU particle code so it could itself check his solutions. After that I left the agent running for 3 hours and it came back with a solution (continuous collision detection) and a bug fix (particles get their velocity applied an extra frame after colliding). I'm sure someone would find something to complain about the code (which is why I haven't upstreamed it) but it looks visually perfect and I had plan to fork the engine for my project anyway
You will get a solution that works with a proper workflow, but you won't get one that scales or would be truly maintainable. Which is also what you get with random midwit drive-by contributors, but faster. I'll give it that.
But more importantly: isn't what you call "intent" just a series of optimization goals? You want your code to satisfy the constraint of being correct™, while also maximizing various other goals like being maintainable, easy to understand, having few lines of code, as little tight coupling as possible, etc. Goals that often conflict, but when given two implementations you could likely tell which hits the better tradeoff (in your engineering experience)
Those are all things that theoretically - with a tight enough specification and enough compute - a constraint solver could solve. No human intent necessary.
The issue is more that we can't fully specify all those side goals, and even if we could the LLM would struggle following them. A classic paperclip maximizer problem (where nobody told the paperclip maximizer to keep the planet inhabitable and all the other side conditions we implicitly assume)
The good thing is that you don't have to agree with that, as the fundamental technical reality does it for you. LLMs do work like that. They are just statistics and probabilities.
And where those constraints prevent us from matching the shape there is no guarantee that we match anything like the mean. Though maybe we can agree that that's where the optimizer would tend to steer towards when it can't do any better
So I wonder whether, in your experience, the results you've seen, could have improved by providing sufficient context? - or what context was given.
I.e. if you have the agent that same context, as one of your colleagues would have/require to solve a problem.
You've missed my point. I didn't dismiss agents. I did dismiss the industry.
I don't need to add more context to a statement that operates on a layer above where context injection would influence it. It is a conceptual impossibility. Not a technical roadblock.
If can put properly engineered intent in the prompt that is verifiable, it works wonders. Anything that can defaults to the llm doing its way, you're right, it just can't converge to good, not with proper constraints.
Epiq solves this with an architecture that supports workflow auditing, allowing you to time-travel state in a filtered view to reconstruct what happened, when, by who, and where intent started drifting, while also allowing you to correlate the evolution of the board with the corresponding commits.
The LLM produce so much code that the best I can is skim and look for obvious flaws, also pass it trough adversarial one.
who is winning here? lol
It reminds me of matt levine's reframing of insider trading where it's not about fairness it's about theft. You're supposed to get secret insights and use them to get an edge. What you can't do is get an edge for yourself with secret insights that your employer got.
So it roughly comes down to "that data is valuable and rightfully belongs to the originating company." Which then makes this a contract diligence type situation.
I don't believe this very strongly. Again it's a contract issue not a moral one. But the metaphor still holds I think.
Double the PRs doesn't mean double the output.
Butt-head: Huh huh huh. Hey Beavis, he said "delve". Huh huh huh.
Beavis: Yeah, I bet he'll use an em-dash next. Heh heh heh.
- The prose before the data appears to be AI generated. Not a big deal but it makes the reader work harder to figure out what's actually being said.
- Linear didn't control for their platforms AI changes over the past year. The platform has become much more AI integrated, so a lot of numbers will move. I'm not sure how you do it but this is only useful signal with a control.
I think the measurement for this may be flawed. We do mostly use AI to decide how to build. But what we build is influenced by AI-driven research into a problem or task. That's largely done in coding and desktop AI tools, not Linear Asks/AI.
I'm working on accelerating my team's work by implementing AI-driven code pipelines with guardrails to eliminate as much unnecessary review time as possible. Also making a chatbot for turning repetitive tasks & PRs into buttons, and an "architectural guidance" chatbot that gives advice tailored to our business, software/system architecture, cloud, standards, etc. This puts AI and automated jobs in the center of both how (automated task) and what (architecture guidance).
But this has a not-so-great implication for Linear. With my tools, a human never has to touch a ticket, so we could use any ticketing system with an API or CLI. Linear is a great product because they made a great interface. What happens when I replace their interface with a chat bot?
Recent studies by e.g. Google show a much broader range of adoption.
Tim (author) if you're there: it'd be amazing to see the split of which agents people are using, if you have that data.
I use LLMs to write my code, but this does not show up in this data.
If it's an attempt at rebutting their claims etc, it'd be easier to interact if you provided some data, or concrete observations :)
It will save people time reading the 5-10 other identical vacuous comments