Mnist Reborn, Restored and Expanded: Additional 50K Training Samples

Mnist Reborn, Restored and Expanded: Additional 50K Training Samples

Increasingly impressive published performance on the MNIST raised researchers’ concerns that models were overfitting to the small test set, throwing the MNIST itself into question: Why trust any new conclusions drawn from this dataset? In the paper Cold Case: The Lost MNIST Digits, researchers reconstruct the MNIST dataset by tracing each MNIST digit to its original NIST source and metadata; and augment the test set with 50,000 additional samples. LeCun believes it may be time for researchers to update their character recognition models: “If you used the original MNIST test set more than a few times, chances are your models overfit the test set.

Source: medium.com