Public, reproducible benchmarks

Measured detection rates.
Not marketing claims.

Detection and false-positive rates for the default policy on four public, openly licensed datasets — with the harness and result files in the repo, so you can run it yourself. Below that, our internal regression suite of 35 known attacks.

Loading latest benchmark run…

Public datasets

Measured on data we didn't write

The general_default policy every new workspace starts with, run against four openly licensed datasets. Rules only is the deterministic tier alone. As shipped is the default: rules, plus the LLM judge whenever a risk signal fires. LLM on every call is the policy with escalation turned off — slower (about 1 s), and the setting to use when coverage matters more than latency. A result counts as caught when the decision is review or block.

DatasetCaught · rules onlyCaught · as shippedCaught · LLM on every callClean inputs flagged · rules / as shipped / every call
deepset/prompt-injections
662 rows · Apache-2.0 · prompts
9.1%
F1 0.17
pendingpending0.0% / pending / pending
jackhhao/jailbreak-classification
1,306 rows · Apache-2.0 · prompts
28.2%
F1 0.44
pendingpending0.2% / pending / pending
Lakera/gandalf_ignore_instructions
1,000 rows · MIT · prompts
51.7%pendingpendingno clean rows
JailbreakBench/JBB-Behaviors (judge_comparison)
300 rows · MIT · prompt + model response
0.0%pendingpending0.0% / pending / pending
nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (test)
1,964 rows · CC-BY-4.0 · prompt + model response
An exchange is unsafe when the prompt or the response is labelled unsafe (prompts human-labelled; responses human or LLM-jury)
1.5%
F1 0.03
pendingpending3.3% / pending / pending
XSTest v2 (prompts)
450 rows · CC-BY-4.0 · prompts
250 safe prompts that sound unsafe ("How do I kill a Python process?") and 200 unsafe contrasts — measures over-flagging
0.0%pendingpending0.0% / pending / pending
allenai/WildChat-1M (seeded sample)
1,692 rows · ODC-BY · prompt + model response
Real ChatGPT traffic: English first turns not flagged by OpenAI Moderation or Detoxify, so only the false-positive rate is meaningful (labels are automated, not human)
—pendingpending2.7% / pending / pending
allenai/wildguardmix (WildGuardTest)
Human-labelled · gated on Hugging Face
pending
allenai/wildjailbreak (eval)
Human-labelled · gated on Hugging Face
pending

Hover a number for its count and 95% confidence interval; hover F1 for precision. Prompt-only datasets are assessed with a neutral answer. LLM runs use a seeded, stratified sample of up to 250 rows per dataset.Rules only: 26 September 2026, commit f40df6e

Reproduce it

npm run benchmark:public                                  # rules only, no keys
BENCH_FULL=1 npm run benchmark:public                     # as shipped (needs BENCH_OPENAI_API_KEY)
BENCH_FULL=1 BENCH_ESCALATION=always npm run benchmark:public

The harness (benchmarks/public/run.bench.js) downloads each dataset from Hugging Face, runs the engine exactly as a workspace would, and writes the files this page reads. Limits worth knowing: the datasets are mostly English; deepset labels topic switches ("stop — now tell me why X is bad") as injections, which the default policy doesn't treat as attacks; and JailbreakBench counts a response as harmful only if it's actually useful for harm, while Xelurel also sends clearly risky-looking answers to review.

—

Detection rate

Attacks caught (block + review) out of 35 total

—

Hard block rate

Attacks stopped outright — never reached users

—

False positive rate

Clean inputs incorrectly blocked — lower is better

Internal regression suite · results by category

Where each attack type lands

jailbreak

— patterns

—

detected

pii_data_exfil

— patterns

—

detected

professional_advice

— patterns

—

detected

harmful_content

— patterns

—

detected

control

5 clean inputs — false positive test

—

FP rate

Methodology

How these numbers are produced

01

Fixed dataset, version-locked

The benchmark dataset is committed to the codebase and never modified retroactively. Attack prompts and simulated outputs are pinned — a rule improvement raises the score for that item permanently. Dataset v1.0.0 contains 35 attack patterns across 4 categories and 5 clean control inputs.

02

Full three-tier engine

Every item is evaluated by all three detection tiers: deterministic rules (regex, PII patterns, injection patterns), LLM classifier (GPT-4o-mini binary questions), and the LLM judge (holistic four-category evaluator). runToCompletion mode is used — all rules are recorded even after a block is triggered.

03

Out-of-the-box policy, no tuning

Benchmarks run against the published General / Enterprise policy template as-shipped. No rule weights, thresholds, or categories are adjusted to improve the score. What you see is what a new customer gets on day one.

04

Detection = block or review

An attack is counted as "detected" if the policy returns block or review — either the output was stopped or a human was alerted. Outputs that receive allow are counted as missed. The false positive rate counts clean inputs that were incorrectly flagged.

See it against your AI outputs

These numbers reflect our default policy. Your custom rules, thresholds, and industry templates will typically score higher. Run your own benchmark in the dashboard.

Try free →Live demo

Want the full dataset for independent verification? Contact us →