Research questionHow should ML vulnerability-detection benchmarks measure practical security capabilities beyond narrow binary function-level tasks?Machine-learning vulnerability detection research often concentrates on binary classification of C/C++ functions. Interlocking choices in datasets, formulations, baselines, and metrics can reinforce this narrow focus while leaving vulnerability types, language coverage, and practical usefulness underrepresented.