Why does representative data matter?

Representative data is data whose composition reflects the full population a system is meant to serve, rather than the subset of that population most easily measured. For gender equity, this means gender-disaggregated data collected with the same rigor applied to every other category in the dataset, not appended as an afterthought.

A model can only be as fair as the data it learns from. Where women and girls have been undercounted, misclassified, or absent from historical records, whether in health research, financial systems, or labour statistics, any system trained on that record inherits the omission and repeats it as prediction. Representative data is the precondition for every other claim of fairness that follows.

The A+ Alliance approach: The A+ Declaration commits to producing open, gender-disaggregated datasets, investing in human-in-the-loop verification of data collection, and prioritizing the quality of inclusive datasets over sheer quantity. It further calls for a set of internationally agreed metrics for digital inclusiveness, measured with sex-disaggregated data across institutions including the UN, the IMF, and the World Bank.

 

Related projects

A+ Declaration, Feminist AI Research network, Gender & AI Innovation Collective

 

Related resources

Feminist AI Network Papers, Feminist AI PubPub, AI4D Knowledge Synthesis Webinars