Comment on KDE developers' attempts at creating LLM guidelines are not going well
MoogleMaestro@lemmy.zip 1 week ago
The issue with KDE accepting AI generated code is that it’s arguably not the right of the programmer to “submit” code produced by an AI as the copyright of the training material is questionable. It’s like the inverse of white-room engineering, there’s actually no way of knowing whether the code coming out of the black box is “authentic” or not.
They need to establish rules grounded with the understanding that AI generated code cannot be copyrighted by the user who produced it, which means it cannot be licensed under the KDE licenses as is.
That’s my 2 cents. I do think them having a discussion around it is important going forward.
9tr6gyp3@lemmy.world 1 week ago
Im trying to understand.
If I learned python from python’s own manuals, then wrote some proprietary python functions in a closed source proprietary project, but then also wrote that same function in an open source project, is it free and open source code, or is it proprietary?
MoogleMaestro@lemmy.zip 1 week ago
You’re fundamentally misunderstanding my argument.
If you use AI to, for example, learn how to decompress or recompress a file, how can you be sure that the outputting function doesn’t match one-for-one an (open or closed) source implementation? You literally cannot, as the AI system generally operates on a black box and cannot “learn” how to attribute and, cynically, the people who make the AI tools do not want to learn how to attribute code from the original developers.
If you take Godot’s renderer code (MIT project) and copy and paste it into a project that is also MIT, that doesn’t mean you can get rid of the Godot attribution. You must preserve the attribution of the code – this is important to learn for all open source developers.
The “learning” argument you’re proposing, which is buying into AI anthropomorphisation, still doesn’t protect you from copyright. If you have photographic memory and worked at adobe, for example, you can’t simply “learn” the code and output the same code to an open source competitor. This is generally why people who work on closed source try to avoid working in same-field open source – the risk of violation is way too high unless you intentionally break design decisions.
Individual code snippets might not be covered but it depends on the line count – but that also changes from territory to territory. IANAL, so my general advice for people who are “learning” from closed source to contribute to alike open source: just don’t. There are other open source projects that can use your skills. And if you’re looking at other open source projects and taking whole functions, you must attribute the project in question (this is pretty easy but depends also on the license of the code you’re looking at.) AI doesn’t have the capacity to understand any of this nuance, and I’d argue it doesn’t understand this by design: hurting copyright hurts open source more than it hurts big exploitative businesses.
9tr6gyp3@lemmy.world 1 week ago
Honestly, thanks for writing all that out. I think I understand where you are coming from. You’re basically saying that this is an attribution issue and the LLMs cant determine that.
But at the same time, now you arent able to use properly attributed software either. Any open-source project that allows any user to contribute has a chance to be using unattributed code. Upstream code that your operating system uses to function could also be using unattributed code that your code runs through and could output tainted results.
The genie is out of the lamp, and it has already made the wishes of the user come true and that wish has landed in your repository. I just don’t see how its feasible to expect any open-source codebase to be fully compliant with full attribution. It never was to begin with, whether LLMs were used or not, but it was expected for people to abide by it.
Now that its easier to generate code off of the training data that went into the model, you have to be real selective on what code is used.
My suggestion for open-source projects is to recommend an open-weight model that has had its training data thoroughly audited so its not a black box. Have it cite code responsibly. Maybe only allow pull requests from users who use that particular model if they must use LLMs. That doesnt mean users are going to abide by that recommendation, but we should start guiding people towards a solution. People are going to use LLMs, thats not going away.