Nightly retrieval eval, 90 days
Sana Qureshi@sanaqureshi
image
Nightly retrieval eval, 90 daysWritten by
Sana Qureshi
ML engineer working on retrieval and inference, not training runs. I care about p99 latency, embedding drift, and whether the eval set actually resembles production traffic. Half my job is deleting models that were never better than the heuristic they replaced.
Three weeks of silent degradation with no code change to blame is the entire argument for keeping a nightly metric nobody looks at. The value is all in the day you look back.