Report by Simbian
The Cyber Defense Benchmark: Why Every Frontier LLM Failed
4 FINDINGSPublished Apr 28, 2026
View Original Report →Key Findings
Anthropic Claude Opus 4.6 detected an average of 46% of attack evidence per MITRE tactic.
Anthropic Claude Opus 4.6Detection AccuracyMITRE ATT&CK
Anthropic Opus 4.6 found three times more attack flags than Google Gemini 3 Flash in the benchmark.
Anthropic Claude Opus 4.6Threat DetectionAI ModelsGoogle Gemini 3 Flash
Zero of the 11 large language models tested earned a passing score on the cyber defense benchmark.
LLMsAI ModelsBenchmarking
Anthropic Opus 4.6 incured roughly 100 times the detection cost of Google Gemini 3 Flash in the benchmark.
Operational CostAI ModelsAnthropic Opus 4.6Google Gemini 3 Flash