Conversation
|
|
||
| ## Attribution | ||
|
|
||
| When AI tools are used, proper attribution helps track the evolving role of AI in the development process. |
There was a problem hiding this comment.
if AI tools are prohibited everywhere is there any point to this section?
There was a problem hiding this comment.
Thanks. Good question.
My understanding is that unsupervised LLM generated output is not accepted.
If you use LLM tools for pre-commit code review or to run some sort of validation, spell cheking... this is allowed
There was a problem hiding this comment.
The word attribution stands out to me. This seems to contradict the TLDR?
Listing AI tooling as a co-author, co-signing commits using an AI tool, or using the assisted-by, co-developed or similar commit trailer is not allowed.
I've seen this format in other projects:
Assisted-by: AGENT_NAME:MODEL_VERSION [TOOL1] [TOOL2]
And usually it is for signing commits, which this policy says earlier is not allowed.
There was a problem hiding this comment.
This is what is most fresh in my memory. No wonder I thought I had seen this before. :)
https://docs.zephyrproject.org/latest/contribute/guidelines.html#ai-coding-assistants
There was a problem hiding this comment.
The word
attributionstands out to me. This seems to contradict the TLDR?
Given that we should be banning the inclusion of any outputs, I don't think that a section on labeling the outputs is useful?
There was a problem hiding this comment.
If you use LLM tools for pre-commit code review or to run some sort of validation, spell cheking... this is allowed
If you run a spell checker I don't think that you need to attribute the output to it?
The crucial thing here is that the LLM itself, or any LLM tooling, cannot be allowed to write to the code, whether you "review" that or not, because human review of LLM outputs tends to decay over time, and so we have to set a hard line to prevent that gradual slide. You cannot copy and paste the output of the LLM into the code.
If you were to run an LLM to spellcheck your docstrings, and it gave you a list of errors, you could manually correct those errors after reviewing the list, because that's not "including LLM output". However, the manual process of typing in each correction yourself is necessary to avoid the vigilance decrement and the slippery slope that comes along with agent harness tooling.
There was a problem hiding this comment.
Thanks for the feedback. I added this section mostly to get more feedback.
I will remote it and replace it with a summary regarding attribution.
There was a problem hiding this comment.
I moved the labeling/attribution into a separate section dedicated to Issues/Security reports... and keept it short to just a simple informal disclosure
| ## Summary | ||
|
|
||
| In practice, this means: | ||
|
|
There was a problem hiding this comment.
can you re-iterate the issues and bugs cannot be investigated by LLM eg #12538 (comment) would have been prohibited under this new policy
There was a problem hiding this comment.
Good question. Thanks for the feedback.
I would say that LLM research/investigation is allowed ... it's just that the fix for the issue should not be LLM generated.
But I will see what others are saying :)
I have created this PR so that we can define what exaclty is allowed and what is not.
There was a problem hiding this comment.
Note that the current text of the PR is not the final versions... this is the initial draft and we will update it based on the feedback from the current developers
There was a problem hiding this comment.
I did specifically say that I thought this use was allowable given that the outputs were out-of-tree.
I personally don't like it but there has to be room for people who personally disagree with me about things to contribute.
There was a problem hiding this comment.
I will try to update the text to make it clear.
One thing that I think that we should expand is using LLM tools for security reports.
One idea is to move those reports to public issue tracker.
The reasoning being that if a public tool can discover that defect, anyone else can discover it and there is little point in hiding it and doing all the CVE/private fix stuff... and with GitHub tools we can't run CI on those fixes.
|
@hynek thanks for creating the AI policy text. Note that this version is 99% your version... with a few changes copied from kubernetes. |
|
Thanks for your feedback. I pushed some changes based on the feedback. I tried to clarify that you can use LLM/AI tools for research or generate reports, but just make sure output is not copied to the Twisted source code. Maybe we should be more explicit that "vibe coding" is not welcomed. Please take another look. needs-review |
|
There is the godot article about AI contribution. There is this tool which I guess uses LLM to detect LLM https://agentscan.tools/ |
|
@graingert thanks for the review. Can you please check the changes and let me know if you still think that this needs more work? Thanks |
|
I would like to have it merged to serve as base. My biggest issue with LLM interaction is that it tried to hide the fact that the comment/finding/code was LLM generated see an example here twisted/pydoctor#799 (comment) where the comment and code was 100% LLM agent, the GitHub user is a human, yet the posted comment contains things like This is why in my initial draft I added the info about disclosure, to know when you are dealing with a human. |
|
Another thing related to LLM generated code in Twisted. twisted/twisted repo is kind of the source of truth for Twisted releated code and LLM agents are trained on the twisted/twisted code. If we let LLM agents generated code in twisted/twisted and then have LLM agents train based on their own generated code, what do we get ? So maybe this should be the important and pragmatical reason for rejecting LLM contributions, and I hope that people will understand. At the same time, some mechanical refactoring do benefit from some LLM work . LLM will generate 5% garbage , but maybe for the other 95% it is worth using it, since you are more producting ... so I don't know.... it's complicated I was not using LLM tools before. I start testing LLM tools since I have created this PR, to get a better understanding of what is this all about and to help detect LLM generated output. I think that with a competent human developer, an LLM tool can help productivity. |
The problem is that this conclusion is based on your subjective impression, and LLMs are extremely good at manipulating subjective impressions. |
I think that this is really wordy and somewhat misleading. It says stuff about Let me see if I can get a new, simpler policy authored that we can use as a more straightforward basis. |
:) Thanks for the info.
Thanks. I know that the document is not ideal. I have used LLM as security scanners, but without code generation. This is why I delayed merging it. But I feel that we need some point of reference for LLM contributions. So having some soft of policy should be high priority for the project. Thanks again. |
|
this needs-changes |
|
Sorry @gudnimg for now merging on your approved review... but I was not 100% happy with the result |
|
@adiroiban while I went in a different direction with my rewrite, I really appreciate the work that you put into this; without some of the discussion here I would not have known what to write. |
glyph
left a comment
There was a problem hiding this comment.
(I really should have submitted my previous comment as a needs-changes review, sorry for just doing it as a random comment.)
I'd really prefer to merge my simpler policy #12835 which skips a lot of the controversial sections here, that just says the two main things:
- Including the outputs is not allowed; a human must generate whatever text we're actually using.
- Using LLMs for examining the code (research, both security and otherwise) is not meaningfully possible to forbid in a world where so many tools are so tightly integrating it, and so much information is LLM-derived.
so I would probably propose that we just close this.
I've also included it as a regular policy doc, but I'm not sure if .github/AI_POLICY.md has some kind of significance. Recent research indicates that these types of files might just be ignored anyway so if this is a convention for agents, I think we can skip it. But if it shows up in the UI somewhere then I'll add one to my branch.
|
I don't think that AI_POLICY.md . LLM agents are kind of rude . GitHub/MS is kind of force feeding us AI tools, so I don't expect to see any UI that will demote the AI usage. This should be closed and we have #12835 to merge |
This comment was marked as abuse.
This comment was marked as abuse.
|
@fly1d Just to clarify since something was taken out of context.
I used In general, the LLM model are designed to "look like a human" and to "blend" with human conversations. I think that this is one big design flaw of these products... I understand that LLM companies are greedy and they wanted money and success at any cost. It would have been much better if the LLMs product were designed to be honest. For example, in GitHub there is support to discose that a comment is a bot. Another reason why I am +1 rejecting LLM contributors. As a reviewer, it's ugly to spend time reading and suggestion various fixes... but then the I am on the fence with LLM usage, but I definetly feel the current way they are designed, they are toxic for open source projects and I fully understand Glyphs total rejection. |
Scope and purpose
Fixes #12676
This creates the initial policy for LLM contribution.
This is based on attrs and kubernetes and Linux.
The idea is that we reject LLM output... but we can't prevent people from using AI.
You can preview the pages here:
Not sure if we should include this as part of Development of Twisted
I left it as Markdown, to leave it as close as possible to the
attrsversion.Once this is finalized, we can reformat it to RST.
Contributor Checklist:
This process applies to all pull requests - no matter how small.
Have a look at our developer documentation before submitting your Pull Request.
Below is a non-exhaustive list (as a reminder):
please review.Our bot will trigger the review process, by applying the pending review label
and requesting a review from the Twisted dev team.