Every program gets the same question eventually. It arrives from a CFO, a board member, or a CISO three days out from a quarterly review, and it sounds deceptively simple: how do we know this is working?
The reflex answer is a number. We took down 5,000 malicious sites last quarter. It fills the slide, it trends up and to the right, and it survives roughly one follow-up question.
Because the honest response to “we took down 5,000 sites” is “compared to what?” Volume tells you how busy your team was. It says nothing about whether customers were safer, whether threats were caught early enough to matter, or whether the people running those campaigns had a worse quarter than you did.
Why takedown volume flatters the wrong things
Counting takedowns is easy. Measuring protection is not. That gap explains how volume became the default metric, and it also explains why the metric quietly misleads the people reading it.
It says nothing about who was at risk
Picture two programs. Organization A removes 5,000 phishing sites, most of which drew a handful of visits before anyone found them. Organization B removes 500, and each one sat inside an active campaign that reached thousands of customers with a working credential harvest.
On a volume chart, A wins by a factor of ten. In practice, B almost certainly prevented more fraud losses, more account takeovers, and more calls into the support queue. The goal was never to remove the most. It’s to protect the most.
It rewards speed over judgment
Not every flagged domain or profile is malicious. Before anything gets submitted, someone or something has to establish whether the content actually impersonates the brand, whether credentials or payment data are being collected, whether malware is in play, and whether the evidence will hold up with a registrar or platform.
Weak validation shows up later as false positives, wasted analyst hours, requests that go nowhere, and strained relationships with the abuse desks you depend on. Burn credibility with a hosting provider this month and your next legitimate request moves slower. Bolster AI’s detection models carry a 99.999% accuracy rate precisely because the alternative is a program that pays for its volume in trust.
It treats a campaign like a pile of one-offs
Threat actors don’t launch a phishing site. They launch a campaign: dozens or hundreds of registered domains, several hosting providers, paid ads buying traffic, social profiles manufacturing credibility, a Telegram channel for coordination, and sometimes a mobile app to close the loop.
Pull a single domain out of that structure and everything else keeps running. What matters isn’t how many assets you removed, it’s how much of the campaign you mapped before you started removing them. That principle runs through domain monitoring, social media monitoring, app store monitoring, and fraudulent ads monitoring: measure the campaign, not the channel.
It ignores repeat offenders
Attackers reuse what works. Phishing kits, naming conventions, registrars, bulletproof hosts, email addresses, wallet addresses, and page templates all resurface, often with barely any modification. If the same operator reappears every month on recognizable infrastructure, you’re winning individual engagements and losing the relationship.
It can’t see prevention
The best outcomes generate no takedown at all. A malicious ad rejected before it runs, a typosquat caught the day it’s registered, a fake app pulled before its hundredth download, phishing infrastructure spotted while it’s still staged and unpublished: none of that moves the volume number, and all of it lowers risk.
There’s one more problem worth naming, and it’s structural. Generative tooling has collapsed the cost of standing up a convincing lookalike, so raw volume climbs for reasons that have nothing to do with how well your program performs. A metric that rises when attackers get cheaper isn’t measuring your defense.
Six metrics that hold up under scrutiny
None of these are exotic. They’re the numbers that answer the question an executive is actually asking, which is some version of: were our customers exposed, and for how long?
1. Mean Time to Detect (MTTD)
The average time between malicious content going live and your program finding it. Every hour in that gap is an hour attackers use to harvest credentials, distribute malware, run payment fraud, and erode the trust you spent years building.
MTTD improves with breadth of coverage and the quality of your intelligence sources. Watching newly registered domains, active and passive DNS, WHOIS records, and certificate transparency logs tends to move the number more than anything else. For context on where phishing volume is heading, the APWG Phishing Activity Trends Reports remain the most useful quarterly baseline available.
2. Mean Time to Validate (MTTV)
How long it takes to confirm that detected content is genuinely malicious. Slow validation stalls the takedown; sloppy validation floods the queue with false positives and trains your team to distrust its own alerts.
Three things drive MTTV: false positive rate, the share of verdicts that can be automated, and how well your playbooks are defined. Bolster AI resolves 95% of takedowns without manual intervention, with human analysts reserved for the edge cases and complex campaigns where judgment genuinely changes the outcome. That’s the split worth aiming for, whoever you use.
3. Mean Time to Takedown (MTTT)
The time from confirmed verdict to content removal. Customers stay exposed for every minute of it.
MTTT is heavily influenced by counterparties. Some hosting providers act within hours and others take days, registrar policies vary by jurisdiction, and social platforms, app stores, ad networks, and marketplaces each run their own enforcement clock. Tracking response time by provider is how you find the bottleneck instead of guessing at it. Without automation, the industry average for a fraudulent removal sits at 10 to 12 days. With automated takedowns, Bolster AI’s mean time to response is roughly 60 seconds, and confirmed malicious URLs reach global blocklists in about 6.5 seconds, so browsers start warning users before the site is fully gone.
4. Customer Exposure Window
Detection, validation, and takedown stacked end to end: the total time malicious content was reachable by a real person. This is the metric to lead with in an executive readout.
It’s also the one that exposes uncomfortable truths. A program can report an impressive MTTT and still leave customers exposed for four days, because detection took ninety-six hours and nobody was measuring that leg of the trip. Shrinking the window is the objective. Everything else is diagnostics.
5. Detection Coverage
How much of the attack surface you’re actually watching. Attackers move fluidly between platforms, so coverage gaps become the path of least resistance almost immediately. A mature program monitors lookalike domains, social impersonation, mobile apps, paid advertising, commercial marketplaces, messaging platforms, search results, and the dark web. Bolster AI scans over 3 million sites daily across more than 1,500 top-level domains, and customer-reported abuse mailboxes often surface the threats no crawler reached first.
6. Campaign Recurrence Rate
How often the same actors or campaigns come back after disruption. Removal is not the same as deterrence, and recurrence is the cleanest signal of which one you achieved.
Track whether returning activity reuses phishing kits, domain naming patterns, hosting infrastructure, email addresses, wallet addresses, or page templates. A falling recurrence rate means you’re raising the cost of doing business for the adversary. A flat one means you’re providing a cleanup service.
What good looks like
Benchmarks are useful even when your program is nowhere near them, because they tell you which direction is up. These are the numbers Bolster AI publishes for its own platform:
- Mean time to response: roughly 60 seconds
- Manual intervention: required in only 5% of takedowns
- Blocklist submission: 6.5 seconds from confirmed verdict
- Detection accuracy: 99.999%, which keeps validation cheap and abuse-desk credibility intact
- Post-takedown monitoring: continuous, because rebuild attempts are the norm rather than the exception
If your current provider can’t produce equivalents for these, that absence is itself a finding. The Bolster AI difference page breaks down how the detection and enforcement stack produces them.
Where these metrics are heading
Measurement is shifting from counting what happened toward anticipating what’s about to. Four changes are already underway.
Risk-based prioritization. Instead of treating every suspicious asset identically, platforms now score threats on brand similarity, infrastructure reputation, observed traffic, and phishing behavior, then route analyst attention accordingly. The queue stops being first in, first out.
Predictive detection. Historical campaign data reveals behavioral patterns: which registrars an actor favors, how they name domains, how long they stage infrastructure before going live. Those patterns support detection before the page is published rather than after a customer reports it.
Provider-level transparency. Median mitigation time, fastest and slowest hosting providers, and the most common abuse types are now dashboard metrics rather than quarterly guesswork. Takedown Insights inside the Bolster AI Web Dashboard exists for exactly this reason: enforcement should be measurable, not a black box.
Reporting in business language. Security leaders need output an executive can act on, expressed as reduced exposure, faster response, and lower fraud loss. The Verizon Data Breach Investigations Report has spent years making that translation for breach data, and brand protection reporting is finally catching up.
Rebuild the report around exposure
Takedown volume isn’t worthless. It’s a workload indicator, and workload indicators belong in operational reviews, not board decks. The trouble starts when a program lets an activity count stand in for an outcome.
So change the first slide. Lead with Customer Exposure Window, support it with detection, validation, and takedown times, show coverage against the platforms where your customers actually get targeted, and close with recurrence to demonstrate whether adversaries are being deterred or merely inconvenienced.The strongest programs aren’t the ones removing the most assets. They’re the ones that shrink the window between a threat appearing and a customer being safe from it. If you want to see what that measurement looks like against your own brand, request a demo.