LONDON, July 30, 2026 — Mozilla Data Collective today announced the launch of Compensated Datasets. First previewed earlier this month, Compensated Datasets enables organizations to list datasets for paid licensing through the platform for the first time.
As organizations increasingly build AI products for global audiences, demand continues to grow for high-quality, multilingual datasets that are transparently sourced and responsibly licensed. AI labs, enterprises, startups and researchers need data that accurately reflects the cultures and contexts they want to serve.
At the same time, many of the organizations and communities creating and stewarding that data have had few transparent ways to participate in the AI economy while retaining agency over how their data is licensed and used. Compensated Datasets helps bridge that gap by enabling organizations to make datasets available through Mozilla Data Collective while retaining control over pricing and licensing, helping expand access to multilingual, multicultural and multimodal data.
With Compensated Datasets now available, uploaders can set their own pricing for dataset access. Downloaders pay for a license to use the data, while uploaders receive 100 percent of the license fee directly. Mozilla Data Collective charges downloaders a separate 5 percent platform fee to cover infrastructure and support costs, while uploaders pay nothing to use the platform.
Compensated Datasets is initially available to verified data providers in the United Kingdom, France, Japan, the Netherlands, Singapore, Spain, and the United States, with additional regions to follow.
“The future of AI depends on more representative data, but it also depends on moving beyond extractive models for how that data is sourced,” said E.M. Lewis-Jong, Founder and CEO of Mozilla Data Collective. “Compensated Datasets is how we begin putting a different model into practice. It’s a step towards an AI ecosystem built on human agency and fair value exchange, where builders have access to better data and the people and organizations behind that data participate more directly in the value they create.”
Compensated Datasets launches with contributions from participating organizations, including TAUS, Pangeanic, Karya, Spotlite, ContentX Labs, and YUX Design, offering responsibly sourced voice, text, image and video datasets for AI builders.
“We see Mozilla Data Collective as more than another distribution channel for datasets. It’s helping build the trusted infrastructure needed to connect data creators and AI builders through transparent licensing, fair compensation and responsible data sharing,” said Manuel Herranz, CEO at Pangeanic. “We’re excited to contribute to an ecosystem that makes high-quality, multilingual datasets more accessible while recognizing the people and organizations behind them.”
“For years, we’ve believed there should be a better way for organizations to share and monetize high-quality language data, so it’s exciting to see Mozilla Data Collective bring that vision to life,” said Jaap van der Meer, Founder & CEO at TAUS. “Compensated Datasets creates a sustainable path for organizations like ours to reinvest in new datasets and AI innovation, while helping developers build more accurate, multilingual models with professionally curated data that reflects languages and communities often overlooked by today’s AI systems.”
Organisations interested in licensing datasets or making their own datasets available through Mozilla Data Collective can learn more here.
About Mozilla Data Collective
Mozilla Data Collective is a mission-locked British social enterprise, backed and incubated by Mozilla Foundation, building the data platform for human agency and fair value exchange. Mozilla Data Collective enables communities, organisations, and individuals to share global cultural datasets on their own terms, while helping downloaders build more representative and culturally grounded technologies with data they cannot find anywhere else. Built by the team behind Mozilla’s Common Voice, the world’s largest open, public-participation speech dataset, Mozilla Data Collective already supports more than 190 organisations sharing over 600 datasets across more than 300 languages. Learn more at mozilladatacollective.com.
Source: Mozilla Data Collective
No related posts found.
The new Fire phone that Amazon launched this week looks like your ordinary black smartphone,…
We know the Internet of Things forecasts: 50 billion connected devices by 2020. Apparently, there’s…
We’ve seen tremendous technological innovation in the data analytics space over the past 10 years….
LIVE from GTC12 — The flock of birds that weaves seamlessly through the sky, propelled…
In the quest to achieve data-driven insight, Hadoop running on Intel X86-based processors has emerged…
Dialog and networking were on the Datanami agenda this week as we kicked off our…
BigDATAwire recently spoke with Cindi Howson, Chief Data and AI Strategy Officer (CAIO) at ThoughtSpot,…
AI coding assistants have become part of everyday software development. They write functions, generate tests,…
By now, most of us have heard about the “wonder drug” Ozempic. Many experts believe…
Databricks and Microsoft have deepened their partnership, which now extends into the 2030s. The two…
At BigDATAwire, we recently sat down with Saurabh Gupta, CEO of The Modern Data Company,…
NetApp has acquired DataPelago in a move aimed at making enterprise data ready for AI…
Alex Woodie Editorial Director +
HPCwire Managing Editor

Ali Azhar BigDATAwire Managing Editor
Jaime Hampton AIwire Managing Editor
Drew Jolly QCwire Managing Editor
Position: High Performance Computing Engineer
Location: Switzerland
Started in 1989, KDD is the oldest and largest data mining conference worldwide, covering technologies like deep learning, differential privacy, and ethical machine learning.
PEARC26
Black Hat
Ai4 2026

© 2025 BigDATAwire. A TCI Media Publication (Formerly Tabor Communications). All Rights Reserved.
Use of this site is governed by our Terms of Use and Privacy Policy.