ML SPLIT INTEGRITY FORENSICS
ML Dataset Leakage & Drift Lab
Before trusting an evaluation, compare train and test splits for copied rows, shared entities, group leakage, time travel, target proxies and feature drift.
Reviewed 2026-07-09
DataWHAT THIS TOOL DOES
ML Dataset Leakage & Drift Lab: inputs, outputs and verification
Find leakage before it inflates your model score.
Load the synthetic case or paste two de-identified CSV splits. Choose optional ID, target, group and time columns, then run one audit.
WHY THIS IS DIFFERENT
Leakage and drift are not the same failure.
Copied examples can inflate evaluation scores while distribution shift can break production behavior. This lab reports overlap, split-policy violations, proxy signals and drift separately so the remediation is specific.
PUBLIC METHOD · REPRODUCIBLE INPUTS
Sources, fixtures, artifacts and correction
Method: compare de-identified train/test rows for exact and near overlap, group/time policy violations, target proxies and distribution drift; report each failure class separately and quarantine implicated rows. Dataset limits and the no-safety-certification boundary remain visible beside the workbench.
Official sources: scikit-learn data-leakage guidance, Google ML monitoring guidance, and the NIST AI Risk Management Framework.
Reproducible fixtures: clean train, clean test, leaky train, and leaky test.
Inspect at least two artifacts: findings CSV and drift SVG; quarantine CSV, Markdown brief, JSON analysis and Receipt v1 expose the exact evidence and assumptions.
QA and accountability: run flagship-editorial-20260713, reviewed 2026-07-13. Exact report hashes are in the quality manifest. FastTool is the accountable publisher. Report a method, fixture, leakage definition or hash error to contact@fasttool.app.
COMPLETE THE EVIDENCE LOOP
Run, understand, verify
Reproduce planted leakage and preserve the split policy, overlap ledger and artifact hashes for review.