Measuring the quality of Synthetic data for use in competitions

Jordon, James; Yoon, Jinsung; van der Schaar, Mihaela

Computer Science > Machine Learning

arXiv:1806.11345 (cs)

[Submitted on 29 Jun 2018]

Title:Measuring the quality of Synthetic data for use in competitions

Authors:James Jordon, Jinsung Yoon, Mihaela van der Schaar

View PDF

Abstract:Machine learning has the potential to assist many communities in using the large datasets that are becoming more and more available. Unfortunately, much of that potential is not being realized because it would require sharing data in a way that compromises privacy. In order to overcome this hurdle, several methods have been proposed that generate synthetic data while preserving the privacy of the real data. In this paper we consider a key characteristic that synthetic data should have in order to be useful for machine learning researchers - the relative performance of two algorithms (trained and tested) on the synthetic dataset should be the same as their relative performance (when trained and tested) on the original dataset.

Comments:	3 pages, 1 figure, 2018 KDD Workshop on Machine Learning for Medicine and Healthcare
Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1806.11345 [cs.LG]
	(or arXiv:1806.11345v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1806.11345

Submission history

From: Jinsung Yoon [view email]
[v1] Fri, 29 Jun 2018 10:39:59 UTC (17 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2018-06

Change to browse by:

cs
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

James Jordon
Jinsung Yoon
Mihaela van der Schaar

export BibTeX citation

Computer Science > Machine Learning

Title:Measuring the quality of Synthetic data for use in competitions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Measuring the quality of Synthetic data for use in competitions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators