Fed up with Big Tech, communities turn to data collectives for control

thm1 pts0 comments

How data collectives are helping communities fight Big Tech AI extraction - Rest of World

Skip to content

Rest of World/iStock

By Rina Chandran

23 July 2026

Marginalized communities and creators are forming data cooperatives to claw back control from Big Tech’s scraping practices.<br>These grassroots collectives ensure that local communities — not Silicon Valley — dictate how their data are used.<br>By leveraging collective bargaining, these groups are successfully forcing tech giants to negotiate for data rather than exploit it.

A growing pushback against big tech companies, and greater awareness of the value of data, is spurring interest in data collectives and cooperatives, which give communities control over the collection, management, and distribution of their data. This alternative allows creators to benefit from data sets that may otherwise be ignored or misused.

A handful of tech companies dominate the generative artificial intelligence industry, with American and Chinese frontier models controlling the lion’s share of the market. Some countries are building their own large language models because their language and culture are not adequately represented in GPT, Gemini, Claude, or Qwen.

The big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now.”<br>Raffi Krikorian, chief technology officer, Mozilla

Communities that possess smaller or unusual data sets can gain from having control over them, Raffi Krikorian, chief technology officer at Mozilla, told Rest of World. The nonprofit last year set up Mozilla Data Collective to provide a platform for such data sets from communities, organizations, and individuals around the world.

“More people are starting to feel like they’re sitting on unique information, not generally available on the internet, and they want to turn the tables on what governance looks like for that data,” Krikorian said.

“The anti-Big AI, anti-Big Tech push is a convenient bedfellow. The big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now,” he said.

A blueprint for responsible use

Companies including Meta, Open AI, Google, and Anthropic have scraped nearly all available data from the internet to train their AI models, and have been accused of using copyrighted material without consent. Tech companies have said it qualifies as fair use, which allows the use of such material for research and other purposes. Some countries are trying to balance the need for good data with the need to compensate creators.

Data collectives, or cooperatives, offer a blueprint for the responsible use of data, and ensure the value goes to those generating the data, Astha Kapoor, co-founder and director of Aapti Institute, a tech research firm in India, told Rest of World.

“Beyond consent and compensation, collective action around data gives communities the opportunity to direct data towards issues they may care about,” she said. “Communities can negotiate the terms on which their data is used at every stage of the AI lifecycle [with] mechanisms for accountability and redressal, in case their terms are breached.”

Workers, producers, consumers, and others have been establishing cooperatives and other community-led associations to pool resources, share benefits, and address socioeconomic challenges for centuries. The United Nations marked 2025 as the year of cooperatives, positioning them as “essential solutions to today’s global problems,” kindling renewed interest in data collectives and cooperatives.

They cover a wide range: The Kerala Food Platform enables about 2,500 farmers in the southern Indian state to trace and market their produce, including rice, fish, fruits, and vegetables. Mexico-based PescaData helps small-scale fishers in Latin America and the Caribbean to manage and benefit from their catch records, while the Native BioData Consortium is a repository of the genetic and environmental data of Indigenous people.

Increasingly, data sets are being created for AI-related purposes, including in low-resource language communities, whose data, including voice data, is valuable for training small models and creating speech recognition tools. For these communities, a data collective is a more practical solution, Krikorian said.

“An OpenAI or Anthropic is not going to prioritize a language that’s only spoken by 1 million people,” he said. “But if it existed, a chatbot that can communicate in their language is hugely beneficial to that community.”

"Linguistic identity crisis"

Long before the launch of ChatGPT, Meesum Alam realized that dozens of languages were dying in his native Pakistan. He belonged to the Baloch community but could not speak Balochi, the language of his forefathers, and experienced a “linguistic identity crisis,” he told Rest of World.

Every other day, I get a text or a voice note from someone who is able to...

data communities tech from collectives world

Related Articles