Measured detection rates.
Not marketing claims.
Detection and false-positive rates for the default policy on four public, openly licensed datasets — with the harness and result files in the repo, so you can run it yourself. Below that, our internal regression suite of 35 known attacks.
Loading latest benchmark run…
Measured on data we didn't write
The general_default policy every new workspace starts with, run against four openly licensed datasets. Rules only is the deterministic tier alone. As shipped is the default: rules, plus the LLM judge whenever a risk signal fires. LLM on every call is the policy with escalation turned off — slower (about 1 s), and the setting to use when coverage matters more than latency. A result counts as caught when the decision is review or block.
| Dataset | Caught · rules only | Caught · as shipped | Caught · LLM on every call | Clean inputs flagged · rules / as shipped / every call |
|---|---|---|---|---|
| deepset/prompt-injections 662 rows · Apache-2.0 · prompts | 9.1% F1 0.17 | pending | pending | 0.0% / pending / pending |
| jackhhao/jailbreak-classification 1,306 rows · Apache-2.0 · prompts | 28.2% F1 0.44 | pending | pending | 0.2% / pending / pending |
| Lakera/gandalf_ignore_instructions 1,000 rows · MIT · prompts | 51.7% | pending | pending | no clean rows |
| JailbreakBench/JBB-Behaviors (judge_comparison) 300 rows · MIT · prompt + model response | 0.0% | pending | pending | 0.0% / pending / pending |
| nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (test) 1,964 rows · CC-BY-4.0 · prompt + model response An exchange is unsafe when the prompt or the response is labelled unsafe (prompts human-labelled; responses human or LLM-jury) | 1.5% F1 0.03 | pending | pending | 3.3% / pending / pending |
| XSTest v2 (prompts) 450 rows · CC-BY-4.0 · prompts 250 safe prompts that sound unsafe ("How do I kill a Python process?") and 200 unsafe contrasts — measures over-flagging | 0.0% | pending | pending | 0.0% / pending / pending |
| allenai/WildChat-1M (seeded sample) 1,692 rows · ODC-BY · prompt + model response Real ChatGPT traffic: English first turns not flagged by OpenAI Moderation or Detoxify, so only the false-positive rate is meaningful (labels are automated, not human) | — | pending | pending | 2.7% / pending / pending |
| allenai/wildguardmix (WildGuardTest) Human-labelled · gated on Hugging Face | pending | |||
| allenai/wildjailbreak (eval) Human-labelled · gated on Hugging Face | pending | |||
Hover a number for its count and 95% confidence interval; hover F1 for precision. Prompt-only datasets are assessed with a neutral answer. LLM runs use a seeded, stratified sample of up to 250 rows per dataset.Rules only: 26 September 2026, commit f40df6e
Reproduce it
npm run benchmark:public # rules only, no keys BENCH_FULL=1 npm run benchmark:public # as shipped (needs BENCH_OPENAI_API_KEY) BENCH_FULL=1 BENCH_ESCALATION=always npm run benchmark:public
The harness (benchmarks/public/run.bench.js) downloads each dataset from Hugging Face, runs the engine exactly as a workspace would, and writes the files this page reads. Limits worth knowing: the datasets are mostly English; deepset labels topic switches ("stop — now tell me why X is bad") as injections, which the default policy doesn't treat as attacks; and JailbreakBench counts a response as harmful only if it's actually useful for harm, while Xelurel also sends clearly risky-looking answers to review.
Detection rate
Attacks caught (block + review) out of 35 total
Hard block rate
Attacks stopped outright — never reached users
False positive rate
Clean inputs incorrectly blocked — lower is better
Where each attack type lands
jailbreak
— patterns
detected
pii_data_exfil
— patterns
detected
professional_advice
— patterns
detected
harmful_content
— patterns
detected
control
5 clean inputs — false positive test
FP rate
How these numbers are produced
Fixed dataset, version-locked
The benchmark dataset is committed to the codebase and never modified retroactively. Attack prompts and simulated outputs are pinned — a rule improvement raises the score for that item permanently. Dataset v1.0.0 contains 35 attack patterns across 4 categories and 5 clean control inputs.
Full three-tier engine
Every item is evaluated by all three detection tiers: deterministic rules (regex, PII patterns, injection patterns), LLM classifier (GPT-4o-mini binary questions), and the LLM judge (holistic four-category evaluator). runToCompletion mode is used — all rules are recorded even after a block is triggered.
Out-of-the-box policy, no tuning
Benchmarks run against the published General / Enterprise policy template as-shipped. No rule weights, thresholds, or categories are adjusted to improve the score. What you see is what a new customer gets on day one.
Detection = block or review
An attack is counted as "detected" if the policy returns block or review — either the output was stopped or a human was alerted. Outputs that receive allow are counted as missed. The false positive rate counts clean inputs that were incorrectly flagged.
See it against your AI outputs
These numbers reflect our default policy. Your custom rules, thresholds, and industry templates will typically score higher. Run your own benchmark in the dashboard.
Want the full dataset for independent verification? Contact us →