Mozilla Data Collective Raises $5M to Scale a More Equitable Data Ecosystem for AI
Mozilla Data Collective Raises $5M to Scale a More Equitable Data Ecosystem for AI
Funding follows rapid growth for the data-sharing platform, with datasets spanning more than 450 languages and the company exceeding its annualised revenue milestone
LONDON--(BUSINESS WIRE)--Mozilla Data Collective, the data-sharing platform redefining how AI data is created, shared and governed, today announced it has raised $5 million from Mozilla. The funding marks the company’s next stage of growth following significant progress since becoming a standalone entity earlier this year.
The idea that we have to choose between giving AI builders the data they need and giving people agency and fair value is a false choice.
Share
AI is being built for a global population, but the training data available today still fails to reflect much of it. As demand for diverse, high-quality data grows, Mozilla Data Collective is working to close that gap by bringing more multilingual, multicultural, and multimodal datasets to AI builders, while creating better ways for the people and organisations behind the datasets to participate in the AI economy.
Since launching publicly, Mozilla Data Collective has built a carefully curated offering that serves both AI builders and the people and organisations behind the datasets. Every contributing organisation is vetted and every dataset reviewed before it reaches the platform, resulting in intentionally gathered, consentful and culturally grounded datasets with clear provenance and licensing. Today, those datasets span more than 450 languages and a wide range of AI and NLP use cases, with 350 organisations approved to contribute data.
“AI is moving incredibly quickly, and we have a window right now to move beyond the extractive models that have defined how data is sourced and make sure what comes next works better for everyone,” said E.M. Lewis-Jong, Founder and CEO of Mozilla Data Collective. “When we started Mozilla Data Collective, we had a pretty big ambition: to prove that we could give AI builders access to better, more representative data while also doing right by the people behind it. In less than a year, we’re seeing demand from both sides and proving that this model works. The idea that we have to choose between giving AI builders the data they need and giving people agency and fair value is a false choice. We can do both, and this funding gives us the opportunity to prove that at a much bigger scale.”
That early traction is translating into meaningful adoption and commercial growth. Major AI labs, thousands of AI startups and scale-ups, and dozens of unicorns are already using datasets from Mozilla Data Collective, while the company’s annualised revenue run rate has reached 9x the milestone set for this stage of growth. At the same time, new regulation, including the EU AI Act, is raising expectations around transparency into AI training data, putting greater importance on the provenance, licensing and consent that have been central to Mozilla Data Collective’s approach from the start.
The platform has also continued to evolve with the launch of Compensated Datasets, new tools to help builders find and request the data they need, and initiatives such as Lost in Transcription, a competition challenging developers to improve speech recognition for underserved, code-switching language communities.
Mozilla Data Collective was launched in November 2025 as the first social enterprise incubated by Mozilla Foundation, before becoming a standalone, mission-locked British company in 2026. The $5 million investment from Mozilla marks the next chapter in that relationship and reflects a shared commitment to a more equitable data ecosystem for AI.
“We thought AI needed a different data economy, and that we had a window to build it before extractive models became the default. So we helped build Mozilla Data Collective. Less than a year in, the market is validating that bet faster than we expected. We’re thrilled to have proof that human agency is an excellent starting point for innovation, not a constraint,” said Nabiha Syed, Executive Director of Mozilla Foundation.
The new funding will support Mozilla Data Collective’s expansion into multimodal cultural video datasets and larger-scale text corpora across EU, African and South Asian languages, alongside new licensing and affordable subscription options designed for startups and scale-ups. The company will also bring new security and data-improvement capabilities from its R&D Lab to the platform, helping organisations share large datasets with greater control and making complex archives easier for AI builders to use. As the company continues to grow, Mozilla Data Collective will look to work with mission-aligned investors who share its vision for a data ecosystem built around human agency and fair value exchange.
About Mozilla Data Collective
Mozilla Data Collective is a mission-locked British social enterprise, backed and incubated by Mozilla Foundation, building the data platform for human agency and fair value exchange. Mozilla Data Collective enables communities, organisations, and individuals to share global cultural datasets on their own terms, while helping downloaders build more representative and culturally grounded technologies with data they cannot find anywhere else. Built by the team behind Mozilla’s Common Voice, the world’s largest open, public-participation speech dataset, Mozilla Data Collective already supports more than 350 organisations sharing over 1,700 datasets across more than 450 languages. Learn more at mozilladatacollective.com.
Contacts
Media Contact
Max Borges Agency for Mozilla Data Collective
mdc@maxborgesagency.com
