Threat Scoring Methodology
The API combines indexed indicator data with available live enrichment. Review the sources, observation dates and missing fields in each response before using a verdict in an investigation.
How the Scoring Pipeline Works
Multi-Source Ingestion
Scheduled jobs ingest feeds into the indexed dataset. Lookups read available indicator records and can request live enrichment. A configured feed, a source contributing to a record and a live enrichment provider are different things; their counts do not describe the number of checks performed for every query.
Configured source weights
Source weights are scoring rules stored in configuration, not measured accuracy or false-positive rates. For example, Google Safe Browsing is weighted 0.95 and StopForumSpam 0.70. URL-based listings count at half weight when scoring the host itself.
Cross-Provider Correlation
When cross-provider signals are available, they contribute a separate confidence factor. Agreement and contradiction help interpret a result, but shared upstream data means multiple listings are not necessarily independent evidence.
Confidence Score Calculation
The confidence calculation uses the base weights below and renormalizes over the factors present. Scanner consensus is included only when engines return verdicts; the cross-provider factor is included when correlation data exists. Confidence is a scoring output, not a measured probability that the indicator is malicious.
Enrichment & Context
Available enrichment depends on the indicator type, upstream response and cache state. It can include geolocation, ASN, WHOIS, DNS, certificates or OTX. CVE records use a separate vulnerability endpoint. Missing enrichment is not evidence that an indicator is safe.
Confidence Score Factors
| Factor | Base weight | Description |
|---|---|---|
| Source Agreement | 33 % | Distinct publishers among threat listings; duplicate listings and tracker-only entries do not add independent votes |
| Source Quality | 22 % | Configured weights of the flagging sources, not measured accuracy |
| Scanner Consensus | 18 % | Included only when scanner engines returned verdicts |
| Cross-Provider Agreement | 12 % | Included when cross-correlation data is available |
| Data Completeness | 10 % | Presence of expected facets for the indicator type |
| Data Freshness | 5 % | Age from available observation or update dates; dates can be missing |
Examples of configured source weights
Examples read from the scoring configuration, reviewed 2026-10-08. This is not an inventory of active feeds or providers queried for every request.
| Source | Category | Configured weight |
|---|---|---|
| Google Safe Browsing | Phishing & malware | 0.95A |
| PhishTank | Phishing | 0.90A |
| OpenPhish | Phishing | 0.90A |
| URLhaus (abuse.ch) | Malware distribution | 0.90A |
| Abuse.ch | C2 & malware | 0.90A |
| Spamhaus | Spam & abuse | 0.85B |
| SURBL | Spam URL blocklist | 0.85B |
| Feodo Tracker | Banking trojan C2 | 0.85B |
| AlienVault OTX | Threat intelligence | 0.80B |
| Malware Domains | Malware hosting | 0.80B |
| Tor Project | Anonymization network | 0.80B |
| Malware Domain List | Malware hosting | 0.75B |
| Blocklist.de | Brute-force & abuse | 0.75B |
| ThreatCrowd | Threat intelligence | 0.75B |
| AbuseIPDB | IP abuse reporting | 0.70C |
| StopForumSpam | Forum spam | 0.70C |
Data Freshness
Feed ingestion, collection reconstruction, client polling and upstream publication are separate clocks. Polling every hour does not guarantee that every indicator was observed within that hour.
Use the returned source dates, firstSeen, lastSeen, lastUpdated and dataTrust fields when present. An update timestamp is not necessarily the time the upstream first observed the threat.
Schedules, caches and upstream failures can delay updates. A missing date or incomplete source coverage must remain visible in the evaluation; there is no universal 24-hour freshness guarantee.
False Positive Handling
When returned, contradictory signals and source details help an analyst review a verdict. They do not establish a measured false-positive rate or guarantee that a result is correct.
Shared infrastructure and repeated upstream listings require context. Use the returned evidence alongside your logs and existing intelligence rather than treating a score as an automatic blocking decision.
Report a suspected false positive through the contact form, including the indicator, supporting sources and why the result appears incorrect.
Inspect a real API response
Compare the fields actually returned with the sources and context already available in your tools. Check timestamps, missing data and contradictory signals as part of the evaluation.