Evaluation Sets That Lie to You

A held-out set is only held out if it was separated before you started making decisions. Usually it was not.

Placeholder article. Leakage rarely looks like leakage. It looks like a model that performs suspiciously well.

The common shapes

Random splits over data with temporal structure. Near-duplicates across the split. Normalisation statistics computed before splitting. Hyperparameters tuned against the set you later report on.

The fix is procedural

Split first, by the axis that matters — time, patient, site, document. Then do everything else. Touch the test set once.

Replace this placeholder article with your own writing.

Read next

17 Sep 2026

Your Model Is Not the Product

The model is perhaps a tenth of what you are shipping. The other nine tenths decide whether anyone keeps using it.

6 min read · 11 views

1 Sep 2026

Retrieval Is Mostly Chunking

Swapping embedding models buys you a little. Getting chunking right buys you a lot. The attention goes the wrong way round.

7 min read · 6 views

0 comments

No comments yet. Be the first.

Leave a comment

Never published.

Plain text only. Comments from visitors are reviewed before they appear.