There’s a big difference between training From (their LLM and generative-ai offering) on it, and training a moderation tool
Comment on Elon Musk’s xAI used child porn to train Grok models, lawsuit says
NaibofTabr@infosec.pub 5 days ago
I just want to point out that there is a valid use case for such a model. The massive amount of images uploaded to the internet every hour has created an exponentially growing need for content moderation, and unfortunately that means some people have to spend time reviewing material flagged as CSAM. This leads to a lot of mental trauma: researchgate.net/…/373947227_The_psychological_im… (not surprisingly).
This is a fucking awful job, and there’s way too much for human moderators to handle. Most people who do this job burn out within a year or two. Supporting their work with automated recognition takes some of the strain off of them. In an ideal world we wouldn’t need humans to do this work at all, but that’s unlikely.
There is a real benefit to training some models to automatically identify CSAM, but such a model should be specialized and only used internally for its intended purpose. It should never be publicly accessible. I suspect that what happened with Grok is an example of gross negligence/incompetence, where a model intended for CSAM identification was included in the overall Grok system.
If that’s not the explanation, then this is way more disturbing because anything else would have been done with intent.
apotheotic@beehaw.org 5 days ago
NaibofTabr@infosec.pub 4 days ago
Yes… I believe I made that point in my comment.
apotheotic@beehaw.org 4 days ago
Yea, wasn’t disagreeing with you, just echoing that particular sentiment
pulsewidth@lemmy.world 5 days ago
Occam’s Razor, I suspect X AI simply downloaded/scraped all porn they possibly could access on the internet to build their training data, and this included a large amount of child porn.
Likely Elon-style factors: