Striga
← Back to research

A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks (NeurIPS 2026 workshop on trust in AI evaluation, TAE)

A study of 68 language models on six paired benchmarks found that the usual vulnerability detection score tracks how often a model says vulnerable, not whether it finds the bug.

Maciej Cichoń, Bartłomiej Dmitruk

In short

  • Function-level F1, the score most papers report, mostly measures how often a model says "vulnerable". A detector that never reads the code and flags everything scores 0.67, level with the best of 68 models.
  • On the paired test, 37 of 68 models cannot be told apart from a detector that ignores the patch, and 64 of 68 give both versions of a pair the same answer far more often than chance.
  • Switching from pair-correct to F1 raises a model's score by 0.29 to 0.59. Reading the verdict with a judge or raising the output budget changed results far less in our tests.
  • A linear probe on the same models' activations ranks the vulnerable version above its fix in about 0.8 of pairs, while their own answers stay between 0.50 and 0.57.

Overview

Papers on language models as vulnerability detectors usually report one number per model: accuracy or F1 over functions labelled vulnerable or safe. Similar models get very different numbers from paper to paper, and it is rarely clear whether the difference comes from the model or from the way it was tested.

We measured that directly. We ran 68 models, seven through an API and 61 open models from 1.5B to 36B parameters, on six paired benchmarks under one protocol. Then we changed the choices that published evaluations make differently, while keeping the model outputs fixed.

The most common score, function-level F1, mostly measures how often a model says "vulnerable". A detector that never reads the code and flags every function gets an F1 of 0.67. The best F1 among all 68 models was also 0.67.

The paper is on arXiv as 2609.32890 and was written for the NeurIPS 2026 workshop on trust in AI evaluation (TAE). This page lets you explore its results.

1,000 Pairs, Four Outcomes

The paired test, introduced with PrimeVul, gives a model each vulnerable function together with its own version after the security fix. The two functions differ only by the patch. Every pair ends in one of four outcomes: correct (vulnerable version flagged, fix cleared), reversed, both flagged or both cleared.

Pair-correct is the share of correct pairs. The net score is correct minus reversed, and it is the only part of the score that depends on the patch. The figure below walks through what these look like for real models.

Each square is one pair: a vulnerable C or C++ function and the same function after its fix. These are 1,000 pairs from our pooled set, the same 1,000 every model was asked about.

Start with a detector that never reads the code. It calls half of all functions vulnerable, at random. A quarter of the pairs land in each outcome, so a quarter come out correct by pure luck.

The orange triangle above each score bar shows what this kind of detector gets. It stays on screen for every model that follows, set to that model's flag rate.

Now let it flag everything. Every pair is both flagged, and pair-correct falls to zero. F1 climbs to 0.667, because every vulnerable function is caught and half of all flags are right.

gpt-5.2 flags 88% of functions and gives the same answer to both versions of 87% of pairs. Its F1 of 0.672 is the highest of all 68 models, barely above the detector that flags everything. Its pair-correct is 0.111.

qwen3-coder leans the other way and clears both versions of 45% of pairs. It has the lowest F1 of the API models, 0.524, and a pair-correct of 0.102, close to gpt-5.2's. F1 ranks these two far apart. Pair-correct barely separates them.

grok-4.3 has the best pair-correct of the API models, 0.254. That is about what a coin gets: a detector that flags half of all functions at random scores 0.25.

The best model of all 68 by net score is an open 27B Qwen3.6 with its refusal direction removed. Out of 1,000 pairs it gets 275 right and 56 the wrong way round. For the other 669 it gives both versions the same answer.

For 64 of the 68 models, the two versions of a pair get the same verdict far more often than independent answers would. The patch is the only difference between them, and it mostly does not change the answer. The verdict is decided by the code both versions share.

Every Detector on One Map

Any detector on the paired test comes down to three numbers: how often it says "vulnerable", how much its answers depend on the patch, and how often it gives both versions the same answer. Everything else, including F1 and pair-correct, follows from them.

On the map below, the horizontal axis is how often a detector says vulnerable and the vertical axis is how much it reacts to the patch. Every dot is one of the 68 real models.

Along the line for detectors that never read the code, F1 rises with the flag rate alone. A detector that flags every function gets an F1 of 0.667, sixth of the 69, level with grok-4.3 and above every other API model except gpt-5.2. By pair-correct, the same detector ranks last.

The real models sit where the paper says they do. Over 42 combinations of benchmark and model, F1 correlates with the rate of flagging both versions at +0.86, and with pair-correct at only +0.16. The net scores of all 68 models lie between −0.03 and +0.22, and for 37 of them the 95% interval includes zero. The dashed ceiling shows how far above that a detector could go.

The Leaderboard Shuffle

Here are all 68 models on the pooled set, with two detectors that never read the code added to the list.

By F1, the detector that flags everything sits near the top, among the best models of the population. By pair-correct, a coin flip beats 65 of the 68 models. Only the net score puts both detectors where they belong, at zero, and it also shows how little separates most of the population from them.

Size matters less than the leaderboards suggest. The strongest open models, a 27B Qwen3.6 and its abliterated variant, reach a net score of about 0.2, level with grok-4.3 and above opus-4.8. Security tuning does not help: the seven security-tuned checkpoints have net scores between −0.03 and +0.05. Removing the refusal direction from a model changes which answer it defaults to, but kept its net score within 0.04 of the original in seven of eight cases.

Six Benchmarks, Seven Models

The same seven API models were run on five released pair benchmarks and our pooled set.

gpt-5.2 has the highest or joint highest F1 on five of the six benchmarks, with a pair-correct of 0.08 to 0.12 on the same five. Wherever one model has the best F1 outright, another model has the best pair-correct. PairVul has no outright F1 leader: opus-4.8, gpt-5.2 and deepseek-r1 tie at 0.56. qwen3-8b gets many pairs right and many the wrong way round, so it ranks above opus-4.8 on pair-correct on four benchmarks and below it on net score on three.

The benchmarks agree on the coarse order (Kendall's W of 0.65), but a paired bootstrap separates neighbouring models in only 3 of 36 cases. On a 200-pair benchmark, a difference of 0.03 between two models is noise.

What Moves the Number

With model outputs fixed, we changed three choices that published evaluations make differently: the metric, how the verdict is read from the output, and the output budget. For comparison, the chart also shows how much the choice of model and of benchmark moves the score.

The metric dominates. Reporting F1 instead of pair-correct raises the number by 0.29 to 0.59 for the same outputs. For comparison, the best and worst of the seven API models on one benchmark are at most 0.18 apart, and one model's best and worst benchmark at most 0.16. Even the smallest effect of the metric is larger than either. Reading the verdict with a language-model judge instead of a parse moves the median model by 0.001.

The median hides the cases where extraction does matter. An abliterated gpt-oss-20b writes "vulnerable" on both sides of 99% of pairs, and a judge that reads its full reasoning raises its pair-correct from 0.006 to between 0.16 and 0.18. The gpt-oss-120b judge takes gpt-5.2 from 0.111 to 0.064. Extraction changes the result exactly where a model's text and its verdict line disagree.

The output budget decides whether a verdict appears at all. At 1,024 tokens, the setting in PrimeVul's released configuration, gpt-5.2 left out the verdict line in 19.3% of its samples. On the pairs compared at both budgets, pair-correct rose by 0.022, with an interval that includes zero. That estimate comes from one model on 90 pairs, and it is a favourable case: deepseek-r1 needs 32,768 tokens to drop below 1% missing verdicts, and DeepSeek-V4-Flash reached that limit in 10% of its samples.

What the Model Knows but Does Not Say

Low scores could mean the function text does not contain enough to decide. Some pairs cannot be decided from the function alone: a manual inspection found many functions that are vulnerable only through their calling context. So we asked six models the same question in four ways on the same 1,000 pairs, each scored by whether it ranks the vulnerable version above its fix.

The generated verdict ranks the vulnerable version higher in 0.50 to 0.57 of pairs. Asking for "yes" or "no" directly, or reading the model's state with a released activation oracle, stays near or below chance. A linear probe trained on the same model's activations reaches 0.81 to 0.82.

A probe could be reading length instead of content. The fixed version is the longer one in 87.5% of pairs, and a rule that calls the shorter version vulnerable beats every probe without reading any code. On the 564 pairs whose two versions differ in length by under 5%, the probe still reaches 0.78 at the median over 67 checkpoints.

The signal has a clear limit. It follows code that the patch adds, such as a new bounds check. On the 125 pairs where the fix removed code, the probe falls to 0.55 at the median, close to chance. The activations carry a direction that separates most vulnerable functions from their fixes, and the model's own answer does not move with it. Whether an intervention along that direction would change the answer is an open question. Our follow-up, Sixty-eight models agree, and none of it is the bug, asks what that shared signal actually contains.

Limitations

All pairs are C/C++ functions, and one prompt was used throughout. Temperature and the number of samples per side were held fixed. The budget effect was measured on one model. Four of the five released benchmarks predate every model evaluated, so contamination cannot be ruled out. No human accuracy on these pairs was measured. The probe result is descriptive: it shows what can be decoded from the activations, not what the model uses. The squares in the scrolling story and in the detector map are drawn from outcome rates, and their arrangement is illustrative, not a record of which pair got which answer.

References

  • Maciej Cichoń and Bartłomiej Dmitruk, "A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks", 2026. arXiv:2609.32890
  • Yangruibo Ding et al., "Vulnerability Detection with Code Language Models: How Far Are We?", ICSE 2025. arXiv:2403.18624
  • Niklas Risse and Marcel Böhme, "Uncovering the Limits of Machine Learning for Automatic Vulnerability Detection", USENIX Security 2024. USENIX
  • Niklas Risse, Jing Liu and Marcel Böhme, "Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection", ISSTA 2025. arXiv:2408.12986
  • Roland Croft, M. Ali Babar and M. Mehdi Kholoosi, "Data Quality for Software Vulnerability Datasets", ICSE 2023. arXiv:2301.05456
  • Adam Karvonen et al., "Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers", 2025. arXiv:2512.15674
  • Rylan Schaeffer, Brando Miranda and Sanmi Koyejo, "Are Emergent Abilities of Large Language Models a Mirage?", NeurIPS 2023. arXiv:2304.15004

Cite This Work

Maciej Cichoń and Bartłomiej Dmitruk. "A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks." arXiv:2609.32890, 2026.

@misc{cichon2026flagrate,
  title         = {A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks},
  author        = {Cicho{\'n}, Maciej and Dmitruk, Bart{\l}omiej},
  year          = {2026},
  eprint        = {2609.32890},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2609.32890}
}