> You'd essentially need to write a single test suite that works for both projects in each pair to compare fairly.
That's pretty much what we did. We start with a repository, and
- (vibeclean pipeline) make changes to clean analyzer issues (some of them involve moving big code chunks around to reduce cognitive complexity), and tweak the existing test sets when needed.
- (slopify pipeline) inverse, add analyzer issues and increase complexity, add dead code etc.
In both cases, we ensure that the test coverage remains the same, and the tests are passing, before we start our experiments.
To be sure, we had hidden tests that validate whether agent implemented the task appropriately. And this is something we paid a lot of attention to (pass rate, in our paper.)
What we didn't do (stupid oversight on my part) was to ensure that there are no regressions in the remaining tests in the repo (unrelated to the current task).
In practice, when using Sonnet 4.6 IRL, I don't see a lot of regressions because often the agent runs the test before calling it done. But it could have gone either way. We don't know.
Our initial suggestion was to do something along the lines of your proposal. But we found that when we ask an agent to implement features for `t-n`^th PR, even when we are overly specific makes the code rather divergent, to the point that sometimes the `t-n+1`th task description doesn't really make a lot of sense.
In fact, there are some papers (that we cited) which create a set of tasks doing exactly this, and it is non-trivial [1].
First author here. Please let me offer a clarification. Our notion of "clean" isn't to just ask the agent to write better code. Rather, we give it a list of 50-100s static analyzer rule violations (and code LOC), and ask to remove them. We then check if the rule violations are resolved.
Using LLMs to rewrite code to remove these violations is a rather accepted practice. Sonar's existing one-shot LLM based approach [1] (in production since 1+ year), and a recent agentic approach [2] to do the same work rather well to do this.
If you think HN crowd is Anti AI, try to talk to a random person on the street. HN crowd is generally more measured on the topic. Current AI is not the perfect solution, and comes with many many soci economic problems and general HN statement reflects that alongside the mindblowing strides made
I guess at some level it is a matter of incentives. In their city, we have electricity 20-22 hours per day (used to be 12-18 when i was growing up) and we can’t rely on the state to provide us electricity consistently.
But also, due to infrastructure. Everyone who could afford it has had a battery and inverter in our homes since forever. Hooking up some solar panels to it is relatively straightforward.
I think there are also some state sponsored subsidies involved although I couldn’t tell you how much.
My 65yo parents installed Solar panels on the roof of their house in a Tier 2 city in the poor parts of India. So did pretty much most of their neighbours.
So i would have to disagree. We are significantly far ahead from the initial “idea”.
As a smoker who transitioned to vaping, I see immense health benefits.
My home country (India), and others (Singapore, others?) have outright banned all electronic cigarettes which is a regulation I hate. I acknowledge that vapes reduce barriers to entry to kids. This is partly solvable in countries with strong governance.
I grow tired of the might makes right world we inhabit again. If you are not a citizen of a hegemon, or their allies, all the best envisioning a stable environment to thrive for your children when you know that the price of sovereignty and nationalised natural resources is a US invasion.
You're not wrong. But, I recently did the mistake of upgrading my iPad to version 26 (the liquid glass version). I had a relatively smooth experience on my 6 year old tablet which now runs painfully slowly. Even scrolling through different parts of home-screen lags.
My point being, with time performance might go up. But instead of that making my device faster/long-lasting, developers use that extra performance to cram in more stuff, at the end of which I come out only slightly better if not worse (as is in my case)
Based on my limited understanding, making a rendering engine (what mozilla folks are doing) and making a browser using existing rendering engines (what firefox is, what probably-but-im-not-really-sure Orion is) are different in terms of complexity by orders of magnitude.
Location: France
Remote: yes
Willing to relocate: no
Technology: pytorch; numpy; pandas; transformers; LLM inference; sparql; RAG; OpenAI, anthropic APIs etc; static analysis.
About me: Doing NLP (research, then industry, then failed PhD attempt ~3yr, then industry) since 2015. Worked extensively with transformers (and RNNs, pre-transformers) including BERT* models, and thereafter with LLMs. Worked extensively on query generation from free text. Currently happily employed in Switzerland, but want to move to France to be with my SO.
I can appreciate that but as an Indian, the thought of subjecting myself and my devices to search for “problematic” material to attend a scientific conference is not something I am willing to do. To me, the USA is the USA.
Also, while there are a lot of people unhappy with your state, I wouldnt say the same for your citizens.
In 2015 a PhD scholar attending a security conference was sent back citing national security concerns. This was absurd as she was an Indian, studying in Montreal and has no past involvement in any untoward thing.
In 2017, a friend doing his PhD in artifical intelligence in Germany was made to undergo a thorough interview at the border to determine if he is a threat on account of his work. Again, this was absurd to say the least.
In this March, my SO (French) chose to not attend a tier 1 conference in AI where she was going to present her work. She, having the brains for both of us, was prescient enough to cancel her trip in Feb-March, a bit before the current border policies came into full force and europeans were detained.
I have never gone, nor will I ever go to the United States. Not for scientific purposes or leisure. For over a decade I have been voicing concerns about hosting conferences in a country which is inaccessible or hostile to a vast section of the scientific community. I am glad to see this shift.
I am sure your opinion is formed based on some experiences you have had in life.
I would like to disagree. Unions are a tool for the poor, the people who don’t have a lot of rights, and protects them from the whim of the rich. If you are working a minimum wage job, and you are being made to work excessive hours, what is your recourse? What is your bargaining power?
Okay, one answer may be to quit and try somewhere else since there isn’t anything to lose here.
Well, I can tell you a very real scenario. My mother was working as a bank clerk in India. Has been her whole life. In the same bank (branches changed but she never changed the bank). When she was 50, there was a fraud. There was a transaction from a local businessman to someone, worth 3x her annual salary. She approved the txn. Once discovered, the businessman took the bank to court who in turn put the blame on my mom of will full ignorance. Businessman offered to settle out of court but we couldnt afford it, of course.
At this point, the union came to mum’s help. They pressured upper management to get their house in order, not shift the blame to the tellers, and do not even think about firing her.