Google Threat Intelligence Indicator Score
Our threat scoring system is designed to help SecOps teams prioritize the most significant security threats. We calculate threat scores for various entities, including files, domains, IP addresses, and URLs. This document explains the logic behind our scoring system and provides guidance on interpreting those scores.
Overview: Demystifying the Threat Score "Black Box"
Security operations teams are constantly inundated with security alerts, forcing analysts to make rapid, high-stakes decisions. To help your team prioritize effectively, Google Threat Intelligence features a native Threat Score Explainability framework.
This methodology is the direct codification of expert Mandiant threat analyst triage logic. By automating the analytical processes of seasoned threat hunters, Google TI evaluates files, IP addresses, domains, and URLs using a rich, continuously expanding library of over 100+ unique threat telemetry signals and behavioral heuristics. This pipeline undergoes continuous manual and automated evaluation to optimize precision (minimizing false positives) and recall (ensuring true threats are caught).
Rather than outputting an arbitrary, opaque number, GTI provides complete transparency by revealing the core contributing factors, signal age-offs, and analyst overrides behind every score.
To prevent ambiguity for analysts, a clear distinction must be made between two complementary metrics:
| Metric | Purpose | Primary Consumer | Operational Usage |
|---|---|---|---|
| Threat Score | Holistic Risk & Triage Priority (0–100) | SecOps Analysts & Automation Rules | Primary decision metric for severity classification, automated alerting, and incident triage. |
| GTI Confidence Score | Machine-Learning Malicious Probability (0–100) | Threat Intelligence Analysts & Detection Engineers | Supporting evidence / contributing factor providing ML telemetry to corroborate the final Verdict. |
Threat Score is HolisticThe Threat Score represents the overarching risk assessment that takes into account all available information from the IoC—including Verdict, Severity, Mandiant analyst intelligence, sandbox detonations, and the automated GTI Confidence Score.
How the Threat Score is Derived
The Threat Score is a function of the Google TI Verdict and Severity, and leverages additional internal factors to generate the final score.
Core Concepts: Verdict vs. Severity
The Threat Score is synthesized directly from two distinct evaluation dimensions:
- Verdict (Maliciousness): How confident are we that this indicator is harmful?
- Severity (Impact): Categorizes the impact of malicious activities possible.
What is a Verdict? (Confidence of Maliciousness)
The Verdict represents our analytical confidence regarding whether an indicator is associated with malicious activity. GTI classifies indicators into five core verdicts:
- Malicious: Definitive evidence and high confidence that the indicator is part of active threat infrastructure, malware campaigns, or unauthorized actor activity.
- Suspicious: Anomalous indicators or suspicious behaviors are observed (e.g., suspicious sandbox behavior or Safe Browsing reports), but they lack definitive proof of targeted malicious intent. This category specifically includes:
- Inactive Historical Threats: Malicious indicators suspected to be currently inactive. While the infrastructure may be dormant, historical threats can still carry high residual impact, making them critical for enterprise investigation.
- Dual-Use Infrastructure & Binaries: Legitimate tools or administrative utilities frequently co-opted by adversaries. To minimize disruptive false positives, these are not classified as malicious outright, but they remain vital hunting leads for security teams.
- Undetected: No suspicious or malicious signals have been observed in our telemetry.
- Benign: Confirmed safe infrastructure, globally recognized trusted software, signed binaries, or clean domains/IPs.
- Unknown: Data unavailable; default verdict.
What is Severity? (Potential Impact)
Severity measures the potential impact or damage an indicator could cause to an organization. In other words, should you care about this indicator? It allows your analysts to ignore low-impact nuisances and focus on high-impact operational dangers:
- High: Associated with severe, business-disrupting threats that typically operate later in the attack lifecycle (e.g., the exploitation or actions-on-objectives phases). Examples include ransomware, active command-and-control (C2) servers, critical backdoor access, or credential stealers.
- Medium: Associated with infrastructure or tools that typically appear in the earlier stages of an intrusion (e.g., the distribution or delivery phases). Because it is technically unknown at this stage whether the distribution will succeed or what the ultimate payload will be, these are capped at medium. Examples include downloaders, droppers, exploits, keyloggers, obfuscation networks, or bulletproof hosts (which may ultimately traffic either low or high-severity threats).
- Low: Low-impact threats or nuisances that should have a limited blast radius if an organization's baseline security posture is strong. Examples include cryptocurrency miners, adware, spam, or vulnerability scanners. While opportunistic vulnerability scanning should be mitigated by standard patch management, these remain important to address if an organization has known underlying weaknesses.
- None: No risk identified (default for Undetected or Benign indicators).
How the GTI Threat Score is Calculated
GTI combines the Verdict and Severity dimensions into a single numeric Threat Score ranging from 0 (completely benign) to 100 (highly malicious and severe).
The score maps directly to specific logic combinations to provide granular risk levels:
| Verdict | Severity | Threat Score | Scoring Logic & Category Modifiers |
|---|---|---|---|
| Benign | None | 0 | Confirmed safe or explicitly allowlisted infrastructure (e.g., non-routable private IPs, known safe domains). |
| Undetected | None | 1 | Default clean state. No malicious markers observed. |
| Suspicious | Low | 21 | Mild anomalies with low potential impact (e.g., minor grayware or suspicious low-impact PUA (Potentially Unwanted Application) behavior). |
| Suspicious | Medium | 25 | Suspicious behaviors tied to delivery-stage activity, obfuscation networks, bulletproof hosts, or dual-use utilities lacking confirmed threat actor pairing. |
| Suspicious | High | 29 | High-Impact, Low-Confidence Safety Cap: Active severe threat signature (e.g., an enterprise-grade backdoor or command-and-control behavior) that has been structurally mitigated, muted, or degraded. Examples include: A known high-severity C2 domain that has been flagged as hijacked or dormant for over 30 days. A high-severity malware file that matches legitimate allow-list override rules. This cap allows analysts to remain aware of the high potential risk without triggering disruptive automated blockings. |
| Malicious | Low | 30 – 59 | Confirmed threats with minor organizational impact: Examples include adware, generic spam. Category Modifiers Include: Cryptocurrency miners and banker trojans. |
| Malicious | Medium | 60 – 79 | Confirmed threats capable of initial intrusion or lateral movement: Examples include standard downloaders/droppers. Category Modifiers Include: Generic hacktools and being identified in verified finished frontline threat intelligence. |
| Malicious | High | 80 – 100 | Severe, critical threats requiring immediate incident response: Examples include active C2 servers, ransomware. Category Modifiers: Associations with known, active threat actors, campaign or ransomware. |
How the GTI Confidence Score is Calculated
The GTI Confidence Score is an automated machine-learning model score calculated across indicators. It serves as one of the key contributing factors driving the indicator's Verdict, Severity, and final Threat Score.
The score maps directly to the following machine-learning confidence scale:
| GTI Confidence Score | Category | Interpretation & ML Telemetry Guidance |
|---|---|---|
| > 95 | Critical Malicious Confidence | Overwhelming ML telemetry indicating active malicious infrastructure or confirmed malware tooling. |
| 80 – 95 | High Malicious Confidence | Strong corroborating evidence of malicious behavior, threat actor infrastructure, or campaign associations. |
| 60 – 79 | Suspicious / Anomalous | Anomalous behaviors, suspicious redirects, or infrastructure overlaps detected; recommended as proactive hunting leads. |
| 20 – 59 | Undetected / Neutral | Insufficient telemetry or neutral markers observed to determine maliciousness. |
| < 20 | Benign / Safe Infrastructure | High ML confidence of legitimate, clean infrastructure or known safe entities. |
Operational Best PracticeBase automated blocking rules and SIEM/SOAR playbooks on the holistic Threat Score and Verdict, using the GTI Confidence Score as supporting diagnostic evidence.
Explainability Cards Taxonomy
To provide context and answer the "Why" behind any given score, the GTI user interface (and API) displays diagnostic explainability cards. These cards translate raw backend scoring attributes into clear, logical groupings.
The 10 Card Categories
- Highest Impact Factor: Highlights the primary indicator or signal that most heavily influenced the verdict and final score.
- Human Verified Ground Truth: Shows determinations, manual overrides, and direct attributions made by threat analysts (e.g., Mandiant incident responders).
- Mitigating Factors: Identifies attributes that reduce the severity of the threat or suggest a benign nature (e.g., trusted metadata, allowlists, or high global popularity).
- Threat Intelligence: Collects attributions, campaign mappings, and reports from premium threat intelligence sources.
- Technical Evidence: Displays automated observations, such as sandbox runs, configuration extractions, and multi-engine detection agreements.
- Tool Abuse: Identifies instances where legitimate administrative, security, or dual-use tools are suspected of being abused in a threat context.
- Propagated Factors: Captures threat indicators that are inherited or associated via related assets (e.g., a URL hosted on known malicious infrastructure).
- Collective Detections: Groups collective signal feeds, blocklists, community YARA rules, and shared reputation indices.
- Machine Learning Models: Reflects determinations made by specialized ML classifiers.
- Historical Context: References past historical activity and signals that have aged off.
UI Classification Examples
Below are a few examples of the explainability cards to their respective explainability classifications:
| Card Title / Signal | Category | Description |
|---|---|---|
| Verified Malicious by Mandiant Analyst | Human Verified | Manual analyst verification during an active investigation. |
| High Sandbox Detonation Malicious | Technical Evidence | Automated detonation observed high-severity malicious actions. |
| Matched Mandiant YARA Rule | Technical Evidence | Matches high-fidelity Mandiant-authored signatures. |
| ML Malware Code Similarity Detection | Machine Learning | ML model matched code structure to known malware families. |
| Abused Legitimate Tool / Service | Tool Abuse | Warns analysts that a safe administrative tool is being leveraged maliciously. |
| Redirects to Malicious Destination | Propagated Factors | Traffic automatically routes users to an active infection node. |
| Attributed to Mandiant Threat Actor | Threat Intelligence | Direct mapping to a tracked advanced persistent threat (APT) or UNC group. |
| Aged-Off Threat Signal | Historical Context | Shows that a historical signal has expired and no longer affects the score. |
| High Global Traffic Popularity | Mitigating Factors | Identified as a high traffic domain which reduces the likelihood of targeted maliciousness. |
Dynamic Card Association Limit
To maintain a clean, actionable UI layout and prevent "alert fatigue" or dashboard clutter, the explainability section limits the display of associated entities.
The UI enforces a maximum of three of each association type (such as associated threat actors, malware families, or rules/YARAs). If more than three associations exist, the UI displays the three highest-impact items and provides an option to view the complete list in the threat details drawer.
Signals Age-Off & Score Decay
Threat infrastructure is highly volatile. A compromised IP address hosting a malicious loader today might be cleaned and reassigned to a legitimate business next week. To prevent security teams from chasing stale indicators, GTI employs a dynamic Signal Age-Off model.
How Age-Off Works
- Expiration Deadlines: Every contributing threat signal (e.g., a sandbox execution run, an AV signature detection, or third-party feeds) is assigned a strict expiration timestamp.
- Age-Off Factors: When a threat signal reaches its expiration deadline without any new corroborating malicious activity, it "ages off." The system retires these signals from current calculations but keeps them logged in the indicator's history as an age-off factor.
- Automatic Score Decay: If the primary signal that drove a malicious score ages off, the threat score is recalculated. This dynamically drops the indicator’s score back down to Suspicious or Undetected protecting your team from false positive alerts.
Customer Guidance
Security teams should prioritize entities with "Malicious" verdicts and "High" severity scores. We recommend further investigation of "Suspicious" verdicts and careful monitoring of all entities with severity scores greater than "None."
We recommend aligning your SIEM/SOAR playbooks with the following score ranges (keeping in mind that your situational context has a big impact on the actions you take):
- 80 to 100 (High Severity): Consider blocking the traffic and initiate high-priority investigation (Active C2, Ransomware, Targeted Attacks).
- 60 to 79 (Medium Severity): Alert security analysts for active investigation (Intrusion tools, Droppers). Initiate proactive hunting sweeps across the network, as these indicators signify the enterprise is being targeted by an early-stage distribution campaign. Search logs for successful distribution.
- 30 to 59 (Low Severity): Log the event and cross-reference your controls to ensure these commodity threats are fully covered by existing passive defenses. If a defensive gap is found that allows the threat to succeed, escalate the incident to a higher priority for immediate remediation.
- Under 30 (Suspicious / Anomalous): Treat as proactive hunting leads. Do not auto-block to avoid false positives. Instead, leverage local environmental context to definitively classify the activity as malicious or benign; this step is especially critical for evaluating dual-use utilities. Hunting Tip: If an internal asset interacts with a Score 25 (Dual-Use utility) followed shortly by a Score 29 (Dormant/Aged-off C2), escalate the investigation immediately, as this pattern potentially suggests a revived historical threat.
Threat Score Explainability: User Experience FAQ
Q: Why did this specific indicator receive this score?
A: Every indicator page features a Threat Score Explainer card. This card functions as an "open ledger of evidence," showing the active contributing factors—such as specific YARA rule matches, connected Mandiant Finished Intelligence Reports, behavioral sandbox alerts, or reputation lists—that drove the verdict and severity calculations.
Q: I suspect an indicator has been incorrectly flagged (False Positive). How can I verify this?
A: Review the Contributing Factors listed on the Explainer card. If a manual correction is made by our teams (e.g., a Mandiant analyst marking an indicator as a false positive or using an bypass list), the system's "analyst mistake" or override logic will trigger. The Threat Score will immediately decay to its correct level, and the Explainer card will highlight the analyst intervention.
Q: How should our team prioritize incidents using the Threat Score?
A: See the score range recommendations in the Customer Guidance section above.
Key points
- The GTI Threat Score is designed to reflect the potential impact of a threat. Higher scores indicate a greater risk to your environment, but risk can only be assessed by using customer derived factors.
- The verdict and severity classifications are the primary drivers of the GTI Threat Score.
- Specific categories, threat intelligence sources, and analyst expertise all contribute to the final score.
- The score is dynamic and is updated as new factors affecting the scoring are observed or as active signals age off.
Updated 7 days ago