Dralvia Research
Archive/labs

How we measure phishing detection quality (and why no number yet)

July 27, 2026

How we measure phishing detection quality, and why we have not published a headline number yet.

Digest focus

Impersonated brands

Mixed attacker brand pressure

Common lure

Mixed lure patterns

Teams to brief

Security and identity teams

Scan window

July 27, 2026

How we measure phishing detection quality (and why no number yet)

Summary

How we measure phishing detection quality, and why we have not published a headline number yet.

Threat model

A detection score is easy to inflate. Pick a friendly sample, count the wins, and you can claim almost anything. The real risk to a reader is a number that cannot be reproduced, because it hides what was missed and what was counted twice.

Why it matters

A detection figure you cannot reproduce is marketing, not measurement. We would rather show our method and our floors than a score we cannot stand behind.

Test setup

We score against a fixed groundtruth corpus with a liveness check, then compute recall and precision the same way every run. We hold two publish floors: recall at least 0.50 and precision at least 0.80. Below either floor we publish no number.

What Dralvia observed

The benchmark runs on schedule, but recall currently sits under our floor, so the public detection-quality page stays in a measurement in progress state with no score.

What worked

The floor discipline itself worked: refusing to publish under the floor kept us honest. New brand impersonation signals also moved recall in the right direction.

What did not work

Single signal fixes did not clear the floor on their own. Recall is the hard part, and one rule rarely moves it far enough.

Product improvements

We added a brand impersonation in domain signal and expanded the trusted brand list, and we gate the public number behind the recall and precision floors.

Defender recommendations

Ask any detection vendor for a reproducible number: the fixed corpus, the floors, and the method. If they cannot show the work, treat the score as a claim, not a result.

Limitations

This note describes the method, not a score. The benchmark covers web2 phishing only, and the number stays unpublished until it clears the floor.

Address this risk

Products that stop it

How we measure phishing detection quality (and why no number yet) | Dralvia Research