Communities Seek Control Over Personal Data - data collectives
Communities Seek Control Over Personal Data

Communities are turning to data collectives as a way to regain control over the information they generate, a movement gaining momentum as criticism of big‑tech data practices intensifies.

Why data collectives are emerging

Large AI firms such as Meta, OpenAI, Google and Anthropic have built their models on massive data scraped from the internet, often without explicit consent from creators. The practice has drawn accusations of copyright infringement, while the companies argue that it falls under fair‑use provisions. In response, groups are forming cooperatives that let them dictate how their data is used, ensuring that any benefits flow back to the original contributors.

Raffi Krikorian, chief technology officer at Mozilla, told Rest of World that “the anti‑Big AI, anti‑Big Tech push is a convenient bedfellow.” He added that “big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now.” Mozilla’s own Data Collective, launched last year, provides a platform for diverse data sets from communities worldwide.

Related: AI Boom Raises Concerns Over Human Job Loss

Examples from the field

Initiatives such as the Kerala Food Platform help about 2,500 farmers track and market produce, while Mexico’s PescaData assists small‑scale fishers in managing catch records. The Native BioData Consortium preserves genetic and environmental data of Indigenous peoples, illustrating the breadth of applications for collective data stewardship.

How collectives differ from other models

Data trusts place a trustee in charge of managing data for a group, while data unions aggregate individual contributions to negotiate with buyers. Data commons like Wikimedia and OpenStreetMap operate under distinct governance structures, and data donation schemes such as the Personal Genome Project rely on voluntary contributions for public benefit. Collectives combine aspects of these models but emphasize community‑level decision‑making and accountability.

Astha Kapoor, co‑founder of the Aapti Institute, highlighted that collectives allow “communities to negotiate the terms on which their data is used at every stage of the AI lifecycle… with mechanisms for accountability and redressal, in case their terms are breached.” This approach aims to balance consent, compensation and the ability to direct data toward issues that matter locally.

Related: AMD Readies Ryzen AI MAX PRO Launch

While the concept is appealing, governance and scaling remain challenges. Kapoor warned that “building data cooperatives solely to steward data is not feasible because sustainability becomes an issue, and the only viable pathway becomes monetization of the data the cooperative is meant to safeguard, which is problematic.” Recent steps by Mozilla to offer paid commercial licensing for community‑generated data sets illustrate attempts to address financial viability while preserving community interests.

Compared with earlier attempts at data pooling, today’s collectives benefit from clearer legal frameworks and growing public awareness of data rights, making them more resilient against corporate overreach. This shift mirrors past movements where workers formed co‑ops to secure fair wages and shared ownership, suggesting a broader trend toward decentralized control of digital assets.

As AI adoption expands, such stories show why data collectives may become a cornerstone of equitable digital participation.