CONTRIBUTING.md: establish initial automation/AI/LLM policy - #514587
Conversation
|
|
||
| All covered use of automated tooling for a contribution must be disclosed as part of that contribution. | ||
|
|
||
| In the case of LLM‐based AI tooling used for commits, this **must** be in the form of a Git commit trailer in the format `Assisted-by: AGENT_NAME:MODEL_VERSION`. |
There was a problem hiding this comment.
I feel like "agent name" is a little too vague. "Agent name" could be a harness that was used (opencode, Codex etc.), it could be some sort of nickname or codename the author uses for their agent (a combination of harness, language model or model family, system prompt and perhaps a memory system in case of autonomous persistent agents).
Additionally, Assisted-by could (and, in my belief, should, though for purposes of this policy I'd put in an RFC2119 "MAY") be used for other tools, not just language model-based coding assistance agents.
I seem to recognize that the format was copied from the Linux kernel coding assistant guidelines. I think we should just copy that section in full (with explicit acknowledgment and backlink):
Assisted-by: AGENT_NAME:MODEL_VERSION [TOOL1] [TOOL2]Where:
AGENT_NAMEis the name of the AI tool or frameworkMODEL_VERSIONis the specific model version used[TOOL1] [TOOL2]are optional specialized analysis tools used (e.g., coccinelle, sparse, smatch, clang-tidy)Basic development tools (git, gcc, make, editors) should not be listed.
There was a problem hiding this comment.
Also I think it would be nice to clarify behavior if several harnesses/models were used. One Assisted-by trailer for each model/harness pair?
There was a problem hiding this comment.
Some examples of AGENT_NAME:MODEL_VERSION could be useful, as there is still ambiguity.
For example, what would be the best form amongst the following?
opencode/unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL- @codex
/Qwen3.6-27B github.com/openai/codex/Qwen3.6-27Bgithub.com/openai/codex/unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL(and we are lost in the/).
There was a problem hiding this comment.
For example, what would be the best form amongst the following?
@hoh all of the examples you listed below have a / where a : should be per the policy described above, and the 1st and 4th example also use a colon for quantization specifiers (the only place where a colon is used in model releases anyway).
All of this definitely requires clarification, and official examples would be useful.
I imagine it like:
opencode:qwen3.6-35b-a3b(downsides: no quantization specifier)opencode:qwen3.6-35b-a3b:q4_k_m(downsides: none, but watch out for the second colon if parsing via automation)opencode:unsloth/qwen3.6-35b-a3b-gguf:q4_k_m(downsides: a little more verbose, though this form can be useful if using obscure finetunes/quants/variants; implied huggingface link)opencode:hf.co/unsloth/qwen3.6-35b-a3b-gguf:q4_k_m(downsides: a lot more verbose, but extremely specific; I could paste this into llama.cpp and get the exact same model you're running, save for the sampler parameters, which seem to be a matter of taste anyway, it seems)github.com/simonw/llm:<any of the above>(downsides: extremely long, but also has a direct link to the harness. One could even do a specific commit, or point to a fork if harness is customized; still doesn't fully capture the workflow, and can't hope to, because some harnesses allow plugins for a total conversion into something else entirely; also, excludes non-OSS harnesses, which to me feels like a good thing, but some proprietary/unpublished harness users will rightfully object, and soon I may be among the latter)
There was a problem hiding this comment.
When reflecting on this topic, I am wondering what the real value is for mentioning the specific tools and models used.
A user may also be using a collection of tools, such as:
- Cursor with Opus 4.7 for creating a plan
- OpenCode with Qwen 3.6-35b-a3b for implementing it
- Codex with GPT 5.5 for reviewing the changes
- GitHub Copilot for a second code review
- Claude code with Haiku for fixing typos
There are also web versions of Codex/Claude/...
This would end up with a very verbose list of advertising for the companies behind these tools, but I don't see the value above something like Assisted-by: Coding agent.
There was a problem hiding this comment.
seconding @hoh here - the prevalence of e.g. claude's suggested opusplan default (have the bigger model handle planning mode, the mid-sized one do implementation) already makes this format awkward.
this further raises the question of what exact problem we were trying to solve with that detail. it's not like model details make things reproducible, while i imagine the reason LLM vendors had the Co-Authored-By tags include any info on this was marketing, basically.
like, should we go along with that and make nix contributors do marketing for LLM vendors just because big AI wanted us to?
There was a problem hiding this comment.
We might end up collecting a bunch of «not even the top models are reliable» examples of mess-ups, and maybe a few guidelines of type «if you use this, check this and document the check or get the PR auto-closed for carelessness». And yeah, maybe we might end up needing to emergency revert as much as possible of commits touched by a specific setup.
There was a problem hiding this comment.
We might end up collecting a bunch of «not even the top models are reliable» examples of mess-ups
tbh LLMs are 'garbage in garbage out' w.r.t. prompts as well as to prior context, and either way should be challenged by someone with some domain knowledge on suggested solutions, if not further to conversely have the machine challenge the user on ideas and implementation as well.
to an extent then, i would be inclined to regard any use of a tool settling for a shitty output as using it wrong, such that i would argue having some users publish shitty output should not have to imply another wouldn't be able to use it better.
There was a problem hiding this comment.
(Joining threads: I’ve responded to some of the points raised here in #514587 (comment).)
There was a problem hiding this comment.
to an extent then, i would be inclined to regard any use of a tool settling for a shitty output as using it wrong, such that i would argue having some users publish shitty output should not have to imply another wouldn't be able to use it better.
On the one hand, sure, «get 70% there and just edit the last part by hand» (possibly with another round of fully-hand-handled LLM review) is solid advice so one could say that anything that gives worse results than that is holding-it-wrong. On the other hand, some data on how bad things can or cannot get with an allegedely vibecoding-safe setup if one doesn't pay attention could still be useful.
|
This is one of the more well thought-out policies. While it would not be one that I would choose personally (and I have opted for a different approach in the past), I believe this is the key part that reasonably justifies picking this particular approach:
|
|
edit: moved accessibility topic to #514587 (comment) |
Many speech-to-text tools nowadays definitely use language models to reduce error rates, and some are architecturally close to what is most centrally called LLMs, and possibly some of them use or will soon use unambiguously instruct-LLMs to allow specifying what kind of text one is dictating. If one includes machine translation into the accessibility umbrella, and includes code comments into the notion of code, then there too one can have LLM-generated code due to accessibility tools. |
|
|
||
| # Automation/AI policy | ||
|
|
||
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and Matrix, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. |
There was a problem hiding this comment.
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and Matrix, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. | |
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub, Discourse, and Matrix, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. |
There was a problem hiding this comment.
Do we consider Discourse a nixpkgs development venue? (honest question; I see it more as a support venue most of the time)
There was a problem hiding this comment.
Category: Development
Development discussion for Nix, NixOS, Nixpkgs, NixOps, RFCs, etc.
There was a problem hiding this comment.
Not in entirety, but the specific threads/topics/spaces. Same with Matrix, only development chats (#dev, #ci, #hydra and so on). I'd suggest specifying that moment:
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and Matrix, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. | |
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and in development-related spaces on Matrix and Discourse, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. |
There was a problem hiding this comment.
The reason Discourse wasn’t included here originally is that it covers a range of topics outside our jurisdiction, like Nix itself. On GitHub and Matrix, the categorization is distinct enough that it’s fairly unambiguous which venues are related to Nixpkgs and therefore in scope; on Discourse, a lot of topics are just in the “Development” category with no further categorization.
I think it would be a good idea in itself to include Nixpkgs‐related topics on Discourse, but the discontinuity between standards in one topic to the next depending on its subject (and the ambiguity when that subject shifts) may be surprising, compared to GitHub and Matrix where the subject of a venue is more static. Happy to hear people’s thoughts about this.
@acid-bong I think your suggestion is redundant: “Every contribution to Nixpkgs and related development venues, including … communication on GitHub and Matrix”? Note that even for GitHub, related development venues does not include e.g. the Nix and Hydra repositories, because they’re outside of the Nixpkgs core team’s authority.
(Of course, we’d be happy to see this work grow into a more widely‐applicable policy beyond Nixpkgs in future!)
|
|
||
| Any use of automated tools to generate non‐trivial amounts of output as part of a contribution, in whole or in part, verbatim or edited, is covered by this policy, except as listed in the Exemptions section. | ||
| Both LLM‐based AI tools and hand‐written automation are covered. | ||
| Contributions include code and documentation in commits, commit messages, pull request summaries and reviews, issue and vulnerability reports, GitHub comments, and Matrix messages. |
There was a problem hiding this comment.
| Contributions include code and documentation in commits, commit messages, pull request summaries and reviews, issue and vulnerability reports, GitHub comments, and Matrix messages. | |
| Contributions include code and documentation in commits, commit messages, pull request summaries and reviews, issue and vulnerability reports, GitHub comments, and messages in other official community spaces. |
There was a problem hiding this comment.
(I think this is the same case as #514587 (comment); let me know if there’s nuance I’m missing.)
This comment was marked as duplicate.
This comment was marked as duplicate.
welp, there's a difference:
the former should go into its own category, like "LLM-assisted (accessibility)", the latter — into "LLM-generated" |
| If you think someone is continuing to break the policy after this, please escalate to the [Nixpkgs core team](https://nixos.org/community/teams/nixpkgs-core/) rather than fighting over it. | ||
|
|
||
| Deliberate violations of this policy are considered to break the [Code of Conduct](https://github.com/NixOS/.github/blob/master/CODE_OF_CONDUCT.md) clause against “Wasting other people’s time with low quality contributions, including but not limited to LLM and bot spam”. | ||
| If a contribution clearly violates the policy – e.g. the contributor admits it is in violation without working to fix it, or there are AI tool attributions that do not meet our required format – it is acceptable for anyone with technical permissions to close or hide it, pointing to this policy if necessary. |
There was a problem hiding this comment.
To clarify, is closing a PR in case of "AI tool attributions that do not meet our required format" exempt from the above "assume good faith"? More specifically, are such PRs that use Co-Authored-By for AI disclosure always to be closed directly without giving a chance of correction?
There was a problem hiding this comment.
I hope not, especially for first-time contributors who otherwise seem to be acting in good faith -- granted, I might be exactly the kind of contributor that you want to drive away with this as someone who has admittedly been a bit too liberal with coding agents in some of my PRs, and in that case I'll take my lumps and step away from the discussion, but I really do think that continuing to be a welcoming place for new contributors is worthwhile, even in a world where those new contributors will increasingly make at least some use of AI-based tools and a lot of people in the community have strong feelings about that.
There was a problem hiding this comment.
To me this paragraph reads like the idea is to grant people who inadvertently violated this policy an opportunity to fix their mistake.
There was a problem hiding this comment.
I worry a bit that this paragraph could be used to justify unwarranted behavior on the part of maintainers. I'd favor a phrasing that keeps the temperature low, and we could adjust the policy if we end up becoming more of a target for slop somehow, and even then, the CoC already covers this, albeit with fewer specifics.
Furthermore if this scares away LLM users in the first place, that would be bad.
- Below the generated code are usually real issues that are worth solving
- LLM users contribute real-world testing
- Improving open source together should reduce LLM usage and its negative externalities (i.e. avoid duplicated LLM work between many users)
I'd consider dropping this paragraph and only reintroducing it when proven necessary.
There was a problem hiding this comment.
The assumption of good faith is not meant to be optional here, of course!
By the definition of extractive contributions, however, we need a way to handle them that requires sufficiently low friction on the part of the person doing triage; see LLVM and Rust for concrete precedent on a “quick rejection” flow for PRs that don’t meet standards set by related policies. I believe that we cannot omit some clear indication that contributions that don’t meet the requirements aren’t accepted, and that it’s fine to deal with them quickly and move on rather than debating the policy every time.
We already have a significant triage problem; a dramatically lowered cost of submitting plausible‐looking contributions exacerbates it, which is one of the primary motivations for this policy. In general, I think Nixpkgs is quite bad at saying “no” to things; contributions tend to stall out rather than being closed, which I think is no less frustrating for contributors.
I agree, though, that this shouldn’t come at the expense of being polite, or giving the contributor a chance to address the problem. The ideal flow I imagine looks something like this:
-
We have a label for contributions that don’t meet the policy. CI detects contributions that can be unambiguously determined to violate the policy (e.g.
Co-authored-by:withoutAssisted-by:) and applies it; contributors can also manually apply it where relevant. -
CI detects the label, marks the PR as a draft, and posts a comment explaining the need to meet the policy and pointing the contributor to it.
-
When the contributor has adjusted the contribution to meet the policy, they can check the PR checklist item, and the PR will be undrafted and the label removed.
-
If the contribution isn’t adjusted within a week, CI closes it automatically.
I think this should allow us to minimize the burden posed by contributions that are in violation of the policy, while being kind to contributors. We’ll take a look at the specific wording here, though, as we don’t want to encourage reviewers to cause conflict, especially unilaterally in ambiguous cases.
|
I think this is a reasonable policy, as a starting point, and if we find shortcomings in its usage we can iterate and learn. (moved feedback to comment) Great work y'all; thanks for getting this out there! :) |
|
|
||
| * Use of standard deterministic editor/IDE/formatter/text transformation tooling to produce changes that the author manually reviews and understands is exempt, including inline “auto‐completion” (even if LLM‐based) of short, rote snippets of text that do not contribute anything beyond boilerplate the author would have written anyway. | ||
|
|
||
| * Use of standard community automation is exempt, such as `nix-update`, the official Nixpkgs CI bots, the @r-ryantm update bot, and the Nixpkgs security tracker bot. |
There was a problem hiding this comment.
How does this affect custom update bots as some maintainers have used them in the past (including me)? Are these still allowed to create PRs on their own without the operator's review before undrafting, or are they explicitly addressed by this policy (i.e, operators are meant to review the automated PR before it is undrafted)?
Context for my use of this (currently retired) bot: zed-editor has a rather short release interval, and for a while there were no other contributors willing to take care of update PRs (which is not the case anymore), so I tried to automate the process in the meanwhile. Also r-ryantm does not notify me of issues in the update process, which my own bot did. I personally think requiring such automated PRs to be drafts until the operator has verified their output is reasonable though.
There was a problem hiding this comment.
Perhaps include updateScript and tooling that's created or approved by the maintainers of the individual package?
We can spread the decision making here.
| * Use of standard community automation is exempt, such as `nix-update`, the official Nixpkgs CI bots, the @r-ryantm update bot, and the Nixpkgs security tracker bot. | |
| * Use of standard community automation is exempt, such as `nix-update`, the official Nixpkgs CI bots, the @r-ryantm update bot, and the Nixpkgs security tracker bot. | |
| * Use of automation approved by the maintainers of the affected set of files, provided that any automated communication such as PR descriptions and automated commit messages are clearly identifiable as the result of that particular automation as laid out in [Transparency](#Transparency). |
(without enforcing that agreement to be formal, so I think it should be fine for a solo maintainer could just go ahead)
There was a problem hiding this comment.
Custom update bots seem like a reasonable use case, and the fact that they’re serving a specific maintainer who has opted in to seeing the PRs certainly helps. However, it seems tricky to delineate the boundaries here. I am not sure we would want to give the green light to, say, LLM‐based bots submitting automated changes en masse, even with maintainer consent, even as drafts, which it seems like @roberth’s amendment would do.
I think that “standard community automation” ought to include update scripts, but update scripts themselves don’t open PRs, which runs into “It is not permitted to submit automated contributions without any manual review or intervention, outside of standard community automation”. Maybe we could have a specific carve‐out here by simply generalizing “the @r-ryantm update bot” here to e.g. “bots that run in‐tree update scripts”?
It’s of course also an option to have these bots present these changes to the relevant maintainers outside of Nixpkgs (say in a fork), and then to migrate them to Nixpkgs once the maintainer has approved, which would meet the policy requirements. But I realize that complicates the workflow, so it would be good to find a way to permit this as‐is.
|
I think it would be good to also add a But however I have nothing else to say about this, I think it's a great idea to be transparent about AI in nixpkgs, it might also be good to something for other repos in nixos github as well (like the nixcpp repo). |
The whole idea of this change is to enforce keeping humans in the loop. Adding dedicated hands off LLM hooks seems like a step in the wrong direction. We want humans to follow the rules explicitly. |
|
Hi everyone, and thank you for the lovely feedback so far! I'd like to draw attention to the following request from Emily's PR body though:
Posting top-level comments that can't be threaded can make discussions like this quite difficult to read when they start filling up, because different conversations get mixed up together. |
| Contributors are expected to be able to answer questions about their contribution and respond to feedback appropriately. | ||
| A contributor submitting a contribution intended for inclusion in Nixpkgs is also responsible for ensuring that it is [appropriately licensed](https://github.com/NixOS/nixpkgs/blob/master/COPYING) and credited, and not encumbered by any incompatible copyright. | ||
|
|
||
| When output from automated tooling is used in contributions, a contributor must establish confidence in that output. |
There was a problem hiding this comment.
One thing I would like to request, though I'm not sure where it'd go here, is a sort of rule that:
Contributions from tools should at the very least reduce toil for future maintainers.
The idea there being that if you contribute via LLM or similar for packages or services, we expect to see:
- Addition of tests (if they were missing)
- Addition of update scripts
- Addition of (as much as possible) well-documented and well-typed program and service options (instead of, say, the not-uncommon practice of "here's a config string we'll turn into a file, good luck").
(nix-facts is a tool you can use to try and see how many packages are missing tests or update scripts, though it's in "beta")
Basically, if we use LLMs, we should be using them not only for our own convenience but to achieve code quality and ease-of-maintenance that is currently not there.
There was a problem hiding this comment.
I think that this is something valuable to keep in mind as a principle, but I think it's a little bit too subject to interpretation to make a hard and fast guideline.
There was a problem hiding this comment.
I think a hard-and-fast rule of "if you're adding a package/service using an LLM, we better see tests and update scripts" is a pretty objective measure--either the tests and update scripts are there, or they aren't.
But, I'm not sure that belongs in this doc, as such.
There was a problem hiding this comment.
"leave the codebase better than you found it" is a good guideline, but would be indeed subjective to be a hard rule.
May be worth a mention somewhere in the contribution guidelines though. Explicitly, because even though it's common-sense for many long-time maintainers, I feel like newcomers may benefit from a reminder. And I feel the positive sentiment as in "try to strive for this ideal" would offset the "don't do this, don't do that" of hard-and-fast rules aimed at low-effort contributions mentally, encouraging good contributions instead of just discouraging bad ones.
There was a problem hiding this comment.
We would definitely like to encourage contributors to raise our technical standards! But I agree with the other replies that this is a more general principle, and would best go in other contributing documentation, as it seems particularly important for this policy to stay focused; any subjectivity is more of a risk here than elsewhere, and I think it already risks being too verbose in the effort to avoid that.
| Please don't blow up situations where progress is happening but is merely not going fast enough for your tastes. | ||
| Honking in a traffic jam will not make you go any faster. | ||
|
|
||
| # Automation/AI policy |
There was a problem hiding this comment.
Hey all,
I fear that this policy is still too permissive. I see the effort put into guardrails to guide LLM users to use their tools responsibly, but am scared that this will still encourage a flood of low effort contributions.
As you say, we are fundamentally bottlenecked on manual review. I am not enthusiastic about reviewing code that could be AI generated. I do not want to spend my time explaining why something fails, and how it can be fixed, to someone who will just pass it along to their chatbot. I do not think that allowing LLM generated content, even clearly labled, will help the burden of manual reviewers.
I believe automated, deterministic, auditable tools fundamentally differ from LLMs. While I agree that holding them to at least the same standards is a good starting point, I believe they should be held to a much stricter one - if allowed at all.
That said, I appreciate that a stance is being taken. Thank you for your transparency and request for feedback.
There was a problem hiding this comment.
I, and many other maintainers have had multiple workflow disrupting interactions with people using AI as a substitute for thinking. Many contributors will not review, advise, or otherwise knowingly interact with any contribution that has been generated by an LLM. That will not change, regardless of what policy is implemented.
I understand that this is not the case for everyone here, and have had many discussions with people on both sides of the fence. One of the arguments I hear the most is that banning AI all-together is unenforceable. Or that a full ban will alienate new contributors. These are not reasons to implement a more permissive policy. We want new contributors to write their own code, be corrected, and learn from the experience. We want our energy to benefit the community.
Enforcing an AI ban will not be impossible. As mentioned here, banning AI is doable. In most situations it is immediately obvious when someone has used an LLM. If they are dishonest about their use, this can easily be escalated and verified. Multiple other large, community governed projects have chosen for a stricter policy, if not an outright ban (see Wikipedia, Gentoo).
Something similar to the proposed rustlang AI policy is much more in line with our community values. Our policy should aim to lighten the load of reviewers, and encourage human learning.
It is imperative that we implement a policy, and these discussions are prone to stall. I urge the core team to close the discussion after the two week mark.
There was a problem hiding this comment.
We absolutely agree that contributions and reviews being passed unthinkingly back and forth to LLMs is not acceptable, and this policy is very much intended to forbid that; I’ve said more about this and other matters you’ve raised at #514587 (comment), so I’ll point there to avoid duplicating responses :)
| Please don't blow up situations where progress is happening but is merely not going fast enough for your tastes. | ||
| Honking in a traffic jam will not make you go any faster. | ||
|
|
||
| # Automation/AI policy |
There was a problem hiding this comment.
I know I'll get push-back on this, but I really think it's worth including at least a minimal AGENTS.md in the repository that just highlights these standards, if only because then we can enlist the agents themselves to help inform users about the expectations and responsibilities they have to adhere to as contributors.
There was a problem hiding this comment.
They will read CONTRIBUTING.md at times, and I have a global instruction on my system for my agent to do so, but AGENTS.md is directly injected into the context window as part of the system prompt (or at least automatically injected into the context window at the start of all conversations) so they don't even have to read it.
There was a problem hiding this comment.
I'll note that figuring out this sort of thing was the point of one of the GSoC proposals that didn't make the cut. It might be worth revisiting, or even suggesting a stub AGENTS.md that has the shape you're considering--a strawman would go a long way. :)
There was a problem hiding this comment.
might adding such a file like this not defeat the spirit of the stated brown M&M test though?
|
|
||
| # Automation/AI policy | ||
|
|
||
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and Matrix, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. |
There was a problem hiding this comment.
Not in entirety, but the specific threads/topics/spaces. Same with Matrix, only development chats (#dev, #ci, #hydra and so on). I'd suggest specifying that moment:
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and Matrix, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. | |
| Every contribution to Nixpkgs and related development venues, including code, documentation, and communication on GitHub and in development-related spaces on Matrix and Discourse, must have a **responsible person in the loop** who is accountable for that contribution and reviews it before submission, and must **transparently disclose** any non‐trivial use of automation to produce it, including but not limited to LLM‐based AI tools. |
There was a problem hiding this comment.
(here, because it's a reply to Emily's header message, not this file's contents)
we would not want to ban the use of accessibility tools that often use LLMs like machine translation, speech to text, text to speech, and OCR
accessibility tools don't generate code, which is what this policy should regulate, they only interpret existing media, such as reading the given text or type from the person's voice. a contribution from a visually impaired person, who reads with LLM-powered tools, but codes everything by hand, isn't LLM-generated. it is certainly possible to restrict LLM usage and not restrict accessibility tools
original replies:
There was a problem hiding this comment.
A coding agent can itself also be an accessibility tool, if someone cannot type for any reason (e.g. RSI).
There was a problem hiding this comment.
RSI meaning repetitive strain injury?
anyway, once again, see my points in linked messages: there's a difference between voice-typing word by word (and knowing what you do) and "hey, claude, make this thing for me"
|
Thank you everybody for the thoughtful feedback and discussion. The Nixpkgs core team has decided to move forward with this policy, in its current state as amended thanks to community feedback. As we've already said, we do not intend for this policy to be set in stone. We'll monitor how it goes and consider potential amendments based on our own observations and community feedback. |
|
This pull request has been mentioned on NixOS Discourse. There might be relevant details there: https://discourse.nixos.org/t/call-to-ban-ai-commits/78262/36 |
|
This pull request has been mentioned on NixOS Discourse. There might be relevant details there: |
|
This pull request has been mentioned on NixOS Discourse. There might be relevant details there: https://discourse.nixos.org/t/the-nixpkgs-core-team-has-disbanded/79413/1 |
|
Here, I believe, is an interesting step somewhere in the right direction: hawkw/mycelium@b77e854. It might not seem that way, especially in terms of managing people's feelings and emotions, but only at the first glance. TLDR:
Call it a "proof is in the pudding" policy. |
|
That only works, I think, in a solo-run project, and not in a project where 300 people have merge access and a portion are pro-slop. Especially since a fair bit of undisclosed slop has already made its way into nixpkgs including core tools. |
Even in before LLM applicability opinions became relevant, conflicting feedback in reviews was a long-standing fact of life in Nixpkgs, and mixing differently controversial issues together sounds unfortunate to me. |
|
This pull request has been mentioned on NixOS Discourse. There might be relevant details there: https://discourse.nixos.org/t/should-we-add-ai-tool-co-authored-by-linting-to-ci/80177/2 |
The Nixpkgs core team feels it is overdue to establish an official policy on the use of automation for Nixpkgs contributions. The Code of Conduct has a clause against “Wasting other people’s time with low quality contributions, including but not limited to LLM and bot spam”, but does not define this in detail, and we have seen a large increase in automated contributions over the past months, often without disclosure.
After discussion with the community team bootstrap group and within the Nixpkgs core team, we’ve agreed on submitting this proposed policy for community review.
This is an area where people understandably have strong views, and which goes beyond technical matters to touch on legal and ethical concerns. It’s unlikely we can achieve complete consensus among contributors on the topic, and we’re willing to make judgement calls as necessary for the benefit of Nixpkgs, but we want to start out with a baseline policy that we think can gain strong consensus.
Therefore, we’ve focused on formalizing existing norms around automated contributions and applying them generally to include LLM‐based AI tools, ruling out what we think contributors will widely agree are clearly unacceptable cases: undisclosed use of complex automation, and automated contributions submitted without any manual review or understanding. The hope is that this will also give us more visibility with which to iterate further as necessary.
The core of the proposed policy is:
See the rendered version of the full policy for easier reading.
This policy takes inspiration from similar policies in LLVM, Mesa, Fedora, and the Linux kernel, along with a proposal by the author of Anubis. We’re also following the Rust project’s work on a more elaborate, stricter policy.
We’ll leave this pull request open for a public feedback period of at least two weeks. We’re happy to hear concerns or suggestions, either on this pull request or in private.
To ensure the discussion stays on track and civil, we’ll be liberally hiding comments that aren’t focused specifically on what Nixpkgs policy should look like, or that restate previously‐raised points. Please leave each piece of feedback as a separate review comment on specific lines of the diff so that threads of discussion can be kept separate and resolved as needed; use the header line if you have a comment about the policy as a whole.
Why cover all automation rather than just LLM‐based AI tools?
We have established norms around how we expect people to disclose, review, and verify automated contributions like treewide refactors. In the past, there have been cases where such changes have been merged without adequate verification and had to be reverted.
We think that formalizing those norms makes sense on its own merits, and that they serve as a good starting point for an AI policy: our expectations for the use of LLM‐based AI tools should clearly be at least as strict as those for more deterministic, reviewable tooling.
Why not ban LLM‐based tools entirely?
There’s clearly a divide in the community on these tools, with some prominent contributors forgoing them entirely and some using them extensively. As we said in the summary, we’re willing to make judgement calls here for the benefit of Nixpkgs, but only after doing our best to facilitate consensus. We believe that we need to establish a baseline of transparency and accountability before considering whether further steps may be appropriate, and that formalizing existing norms will help with this.
We believe that a maximally hardline policy of banning all use of LLM‐based tools at any stage of any contribution would likely be untenable; we would not want to ban the use of accessibility tools that often use LLMs like machine translation, speech to text, text to speech, and OCR, and declining genuine vulnerability reports solely on the grounds of LLM use during research would be shooting ourselves in the foot. We are not concerned with designing norms around dishonest contributors, but believe any policy will need to strike some kind of balance to be reasonable and enforceable.
We care deeply about the technical quality of Nixpkgs and don’t want a race to the bottom on standards, but believe it should be possible to use these tools without sacrificing rigour. While it may be financially inadvisable, I have seen Claude use to produce trivial version and hash bumps, and clearly in that case there is no meaningful risk to Nixpkgs compared to doing it by hand or with
nix-update. For a less trivial example, it’s not uncommon for initiatives to be blocked on treewide changes that are tedious to write by hand, but comparatively easy to programatically verify (e.g. checking whether a refactor avoids causing rebuilds); in the past, these have sometimes been accomplished with a combination of naive scripting to handle common cases, and manual busywork to handle the rest. LLM‐based tools show potential to reduce the manual toil for these changes in combination with deterministic verification code.In terms of copyright, we generally trust that submitted code is unencumbered and licensed appropriately. Cases where there is clear cause for concern are raised and handled as people notice them. LLM output is one of many areas of legal ambiguity around software copyright; too many, sadly, for us to strictly avoid all of them. This shouldn’t be taken as expressing any official project view on the copyright status of LLM output in general – there is clearly LLM output that does not pose a copyright issue (e.g. the aforementioned trivial version bump), and clearly LLM output that is a much greater concern (e.g. eliciting verbatim reproduction of existing code snippets, or using them to rephrase a codebase in an attempt to “launder” its licensing). If anyone spots plagiarism or copyright violation in any contribution, they can point it out and it should be handled appropriately. Of course, the appropriate copyright policy for LLM output may change as regulation and precedent develop.
We understand, however, that many community members have objections to the training and use of LLMs that go beyond matters of technical quality or legal concerns, and don’t intend to dismiss them out of hand. We would like this policy to be seen as establishing consistent standards based on pre‐existing norms around automation and ruling out the most problematic cases, not as an endorsement by the project, nor necessarily as the final word on the matter. We hope that transparency requirements will give people choice about which contributions they spend time on in the immediate term, and help facilitate a broader discussion on these topics to inform future work.
Why require disclosure?
We believe that transparency is the foundation of any effective policy. Establishing a baseline expectation of disclosure is required to enforce any further requirements or prohibitions, and to evaluate whether a policy is working.
We already have an established expectation that people who use scripts to generate large treewide changes disclose this in their contributions, and preferably make efforts to verify and understand the results.
We believe that it’s especially important in the case of LLM‐based tools: they are particularly good at producing confident and convincing output even when it’s not correct. While contributors are of course fallible in general, and code review inherently involves applying appropriate scrutiny rather than taking things on faith, we believe that this decoupling between the appearance of confident expert research effort and the correctness of the changes is especially pronounced for LLM output, and has a significant effect on review dynamics. It also decreases the extent to which reviewers can rely on pre‐existing trust in the nominal author.
The fundamental character of interacting with LLM‐based tools is also different – providing extensive feedback on changes from new contributors can usually be expected to help them learn and improve over time, which is not the case for an LLM. Contributors may wish to avoid spending time on this, or adjust their reviews accordingly.
Why require
Assisted-by:overCo-authored-by:?Since LLM‐based AI tools use
Co-authored-by:attribution by default, we think that using a different format will serve as a brown M&M test and assist with triage. A commit that contains such a tool inCo-authored-by:but notAssisted-by:is, by definition, one that did not follow the policy and can be closed, potentially automatically. It also matches the format standardized on by the Linux kernel and Fedora, and will make it easier to review and query for use of these tools.Why require a person in the loop?
We are fundamentally bottlenecked on manual review; there is no benefit to Nixpkgs to the submission of changes en masse that we cannot hope to triage. It’s the expected norm that contributors check their contributions before submitting them. A contributor sending automated changes without any review or validation takes up limited triage resources while not contributing anything that someone else with access to the same tooling couldn’t. Since any kind of automation allows changes to be produced much more quickly, requiring oversight ensures that contributions happen at a sustainable pace, and that every contribution is already vouched for by its contributor.
See the LLVM policy’s remarks on extractive contributions for thinking that resonates strongly with ours.
We also place trust in maintainers and committers in our own automation, e.g. through the merge bot. Unsupervised automation of changes and reviews subverts the principles behind these mechanisms.
This also mitigates the issues around feedback, by ensuring that there is always someone on the other end who can learn from it.
Why not cover upstream practices for packaged software?
We have been following the discussion about this, but believe that establishing a policy for Nixpkgs itself is the priority. Defining criteria and maintaining data for the entire package set would be a huge undertaking that goes beyond the scope of this proposal. A blanket policy covering packages with any LLM output would be untenable, as we cannot avoid packaging core components like the Linux kernel, LLVM, systemd, and HarfBuzz, and it would likely be very difficult to produce a usable, up‐to‐date system while entirely avoiding all such components.
Things done
passthru.tests.nixpkgs-reviewon this PR. See nixpkgs-review usage../result/bin/.