Hi Dex,
you write " There are no good benchmarks for a model's ability to maintain codebase quality. "
This might be interesting for you:
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via CI
https://arxiv.org/html/2603.03823v3
I'm half through reading " Why Software Factories Fail " and so far it absolutely aligns with my experience.
Thanks for sharing your knowledge!
Hi Dex,
you write " There are no good benchmarks for a model's ability to maintain codebase quality. "
This might be interesting for you:
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via CI
https://arxiv.org/html/2603.03823v3
I'm half through reading " Why Software Factories Fail " and so far it absolutely aligns with my experience.
Thanks for sharing your knowledge!