arXiv:2107.08928 Abstract | arXiv Analytics

arXiv:2107.08928 [cs.LG]Abstract References Reviews Resources

Introducing a Family of Synthetic Datasets for Research on Bias in Machine Learning

William Blanzeisky, Pádraig Cunningham, Kenneth Kennedy

Published 2021-07-19Version 1

A significant impediment to progress in research on bias in machine learning (ML) is the availability of relevant datasets. This situation is unlikely to change much given the sensitivity of such data. For this reason, there is a role for synthetic data in this research. In this short paper, we present one such family of synthetic data sets. We provide an overview of the data, describe how the level of bias can be varied, and present a simple example of an experiment on the data.

Categories: cs.LG, cs.CR, stat.ML

Keywords: machine learning, synthetic datasets, synthetic data sets, significant impediment, introducing

Related articles: Most relevant | Search more

arXiv:1811.07216 [cs.LG] (Published 2018-11-17, updated 2018-11-24)

Machine Learning for Health (ML4H) Workshop at NeurIPS 2018

Natalia Antropova et al.

arXiv:1811.06128 [cs.LG] (Published 2018-11-15)

Machine Learning for Combinatorial Optimization: a Methodological Tour d'Horizon

Yoshua Bengio, Andrea Lodi, Antoine Prouvost

arXiv:1701.07179 [cs.LG] (Published 2017-01-25)

Malicious URL Detection using Machine Learning: A Survey

Doyen Sahoo, Chenghao Liu, Steven C. H. Hoi

arXiv Analytics

arXiv:2107.08928 [cs.LG]Abstract References Reviews Resources

Introducing a Family of Synthetic Datasets for Research on Bias in Machine Learning

Links

Toolbox

arXiv:2107.08928 [cs.LG]AbstractReferencesReviewsResources

Introducing a Family of Synthetic Datasets for Research on Bias in Machine Learning

Links

Toolbox

arXiv:2107.08928 [cs.LG]Abstract References Reviews Resources