dfbench

dfbench is the held-out benchmark depthfirst maintains to evaluate frontier models on open-ended defensive security work, across three objectives: detection, validation, and differential analysis.

How dfbench works

dfbench: detect

Recall

20% 30% 40% 50% 60% 70% $1 $3 $10 $30 dfs-large1 GPT 5.6 Sol xhigh GPT 5.6 Luna xhigh Opus 5 Grok 4.5 Kimi K3 max Qwen 3.8 max GLM 5.2 xhigh DeepSeek v4 Flash Gemini 3.6 Flash

Cost per task

Figure 1: Vulnerability recall across a security scope, against cost per task. The ideal agent sits in the top left corner. Unless otherwise specified, all models are run on our internal harness.

Every dfbench datapoint, by model and thinking level.

Model and thinking level
Thinking
dfs-large1
GPT 5.6 Sol xhigh
GPT 5.6 Sol high
GPT 5.6 Sol med
GPT 5.6 Luna xhigh
GPT 5.6 Luna high
GPT 5.6 Luna med
Opus 5 high
Opus 5 med
Grok 4.5 high
Kimi K3 max
Kimi K3 high
Qwen 3.8 max
GLM 5.2 xhigh
GLM 5.2 high
DeepSeek v4 Flash max
DeepSeek v4 Flash high
Gemini 3.6 Flash high
Detect cost and overall recall
$6.77 62.2%
$43.37 65.7%
$26.47 59.6%
$9.51 48.9%
$2.53 52.4%
$1.57 39.6%
$0.69 28%
$22.66 45.4%
$21.89 47.8%
$7.70 57.2%
$9.10 48%
$6.22 46.7%
$7.91 40.8%
$9.72 40.7%
$4.09 38.2%
$1.23 31.8%
$1.03 30.1%
$7.44 20.9%
Detect recall by domain
66.5% 42.4% 62.3%
67.1% 67.3% 60.7%
62.8% 62.3% 49.2%
54.2% 38.5% 44.3%
56.8% 45.1% 47.5%
48.4% 23.1% 31.1%
31.6% 15.4% 29.5%
54.8% 34.1% 31.1%
45.8% 57.7% 44.3%
59.1% 51.6% 0%
31.8% 31.8% 30.4%
48.4% 51.8% 37.9%
40.8% 40.8% 40.8%
34.8% 22.4% 44.3%
42.6% 11.1% 44.3%
27.6% 29.5% 45.9%
30.1% 30.1% 30.1%
19.4% 11.5% 32.8%
Detect recall by repository scope
63% 52.6%
66.3% 57.9%
63.3% 10.5%
50.6% 26.3%
53.2% 42.1%
40.2% 31.6%
29.3% 10.5%
47% 26.3%
48.2% 42.1%
59.9% 25%
32.3% 21.1%
48.5% 31.6%
40.8% 40.8%
35.6% 20%
42.2% 36.8%
32.9% 26.3%
31.9% 8.3%
20.9% 21.1%
Validate precision
18.3%
18%
22.3%
24.6%
23.2%
24.8%
27.3%
39.4%
40.5%
29.5%
30.1%
33.3%
29.6%
28.6%
30.8%
24.8%
26.8%
43.5%
F1 from detect recall and validate precision
28.3%
28.3%
32.5%
32.7%
32.2%
30.5%
27.6%
42.2%
43.8%
38.9%
37.0%
38.9%
34.3%
33.6%
34.1%
27.9%
28.4%
28.2%
Differential analysis
$1.34 75.6%
$10.67 75.3%
$5.46 76.3%
$2.32 74.4%
$0.62 71%
$0.30 71.3%
$0.13 65.1%
$7.41 70.3%
$5.48 71.3%
$3.38 65.1%
$6.34 60.3%
$4.18 60.1%
$2.99 72.7%
$2.04 61.4%
$1.56 58.6%
$0.79 66.8%
$0.55 67.4%
$1.51 61%
Differential analysis by NEW, KEPT, and CLOSED
45.8% 95% 86%
46.8% 95% 84.1%
45.6% 94.4% 88.9%
39.2% 93.3% 90.7%
38.9% 89.3% 84.8%
32.5% 92.4% 89.1%
22.2% 79.2% 94%
38.2% 79.8% 92.8%
37.4% 83.8% 92.8%
39.9% 74.5% 80.9%
35.5% 72% 73.4%
31.4% 75.3% 73.5%
33.6% 95% 89.6%
25.5% 85.6% 73%
25.3% 76.4% 74.1%
28.4% 83% 89%
27.1% 88.1% 87%
18.3% 68.9% 96.5%
These results will be updated as new models and harnesses are released. Expand detect to see results stratified by example type.