Research
What Vulnerability Detectors Actually Learn
Testing whether code models encode vulnerabilities or just learn labels
Maciej Cichoń
LLMs Write Insecure Code by Default
When a prompt asks for a feature and never mentions security, what do language models write? A benchmark of eleven models on 1,007 sink-forcing tasks.
Bartłomiej Dmitruk
change the code, keep the bug
building a scalable benchmark measuring LLM capabilities for vulnerability detection
Maciej Cichoń
For a Fistful of Dollars: Less than $100 of Compute Surfaces Pre-auth RCE in Apache httpd
A double-free in Apache httpd's mod_http2 stream cleanup, surfaced by Striga, turns a two-frame HTTP/2 sequence into pre-auth Remote Code Execution.
Bartłomiej Dmitruk