Evaluation Sets That Lie to You
A held-out set is only held out if it was separated before you started making decisions. Usually it was not.
Placeholder article. Leakage rarely looks like leakage. It looks like a model that performs suspiciously well.
The common shapes
Random splits over data with temporal structure. Near-duplicates across the split. Normalisation statistics computed before splitting. Hyperparameters tuned against the set you later report on.
The fix is procedural
Split first, by the axis that matters — time, patient, site, document. Then do everything else. Touch the test set once.
Replace this placeholder article with your own writing.
Leave a comment