Why Would Meta Download So Much Porn?

hebelehubele1 pts1 comments

Why Would Meta Download So Much Porn? - The Atlantic

Listen−1.0x+<br>Seek<br>0:009:43

Editor’s note: This work is part of AI Watchdog, The Atlantic’s ongoing investigation into the generative-AI industry.<br>Amid tech companies’ ongoing efforts to vacuum up as much data as possible to train AI models, Meta has taken public posts from its own platforms, scraped massive amounts of content from the rest of the internet, and pirated millions of books to train its AI models. Now the company is being accused of downloading much more illicit material: huge volumes of porn, nonconsensual celebrity nudes, blueprints for 3-D-printable handguns, and millions of passwords acquired by hackers.<br>Strike 3, the parent company of Vixen Media Group, a prolific producer of adult films, is suing Meta for allegedly downloading 2,973 of its copyrighted videos. But in its efforts to collect evidence about these purported downloads, Strike 3 also captured information relating to many other files that Meta may have acquired over the course of two years: They include images from “Celebgate,” a 2014 hack that resulted in the leak of private photos belonging to Jennifer Lawrence, Kirsten Dunst, and others, plus several collections of deepfake celebrity porn featuring the faces of Natalie Portman, Scarlett Johansson, Elizabeth Olsen, Gal Gadot, and others.<br>Strike 3’s file-transfer lists include ordinary movies and TV shows, alongside dozens of videos from GirlsDoPorn, which was shut down after six people associated with the site were charged with sex trafficking. Meta allegedly downloaded several individual episodes from the website and two “GirlsDoPorn MegaPack” archives. Strike 3 also presents evidence that Meta made some content publicly available in addition to downloading it.<br>When I reached out to Meta, a spokesperson told me via email that “these claims are bogus.” Meta also noted in a court filing that Strike 3, which regularly files lawsuits against people who download its work illegally, “has been labeled by some as a ‘copyright troll.’”<br>This is a complicated case. Piracy can be hard to track with precision, and both parties are clearly conflicted: Strike 3, in seeking damages, wants Meta to look bad, and Meta would prefer not to be associated with this kind of content. Yet if Strike 3’s data are accurate, they would show that participating in mass piracy is a by-product of modern tech development—whether or not any of the material collected is explicitly used to engineer AI models or other programs.<br>Read: The hypocrisy at the heart of the AI industry<br>Although Meta acknowledged the possibility that employees, contractors, or even visitors to Meta’s offices may have used the company’s networks to download the porn in question, its defense in court was that the material was accessed for “personal consumption” rather than as part of a company project. But neither the sheer volume nor the pattern of downloads looks like personal consumption. The file-transfer logs produced by Strike 3 show sequences of files that are unrelated except for a certain keyword or concept. For example, two files that were transferred consecutively on June 2, 2023—“Bajillion Dollar Properties S01E07” and “First Time Home Buyer Anal Fantasy”—are both loosely themed around real estate, but the similarities likely end there. Judge Eumi Lee, who denied Meta’s motion to dismiss the case, cited other such juxtapositions, such as consecutive downloads of “Teenage Mutant Ninja Turtles (1987-1996)” and a video labeled “Teen Sex Sessions 2 (2012).”<br>Strike 3 has argued that the content could be used for AI training. A Meta spokesperson told me, “We don’t want this type of content, and we take deliberate steps to avoid training on this kind of material.” In its legal defense, Meta pointed out that its terms of service prohibit users from using its AI products to generate images containing pornography. Even if Meta isn’t building a porn-generating bot, however, tech companies can use “unsafe” content to test guardrails for their systems—effectively telling the software what not to generate. (Lee noted that Meta’s terms of service are irrelevant to the question of whether it wants to acquire porn.)<br>It’s also possible that Meta is compiling an archive of material for no precise purpose, just as Anthropic acquired millions of books that it says it had no intention of using for AI training but wanted to keep for its “research library.” As AI-training techniques advance, there is no telling what kind of content a company might find useful in the future.<br>Nineteen of the allegedly downloaded files contain models of functional handguns that can be produced with 3-D printers. A file labeled “5.7 million passwords list” purportedly contains 5,718,107 passwords from accounts that were hacked from 2015 to 2019. The downloads also include password-cracking tools that can be used by hackers to break into people’s accounts. Other files include the movies BlacKkKlansman, The Banshees of...

meta strike from porn content files

Related Articles