Building Vulnerability Aging and SLA Tracking
Overview
With over 30,000 new vulnerabilities identified in 2024 (a 17% increase from the prior year), organizations must track how long vulnerabilities remain unpatched and whether remediation occurs within defined Service Level Agreements (SLAs). Vulnerability aging measures the time between discovery and remediation, while SLA tracking enforces severity-based deadlines. Industry benchmarks indicate standard SLAs of 14 days for critical, 30 days for high, 60 days for medium, and 90 days for low vulnerabilities, though more aggressive timelines (24-48 hours for actively exploited critical CVEs) are increasingly common. This skill covers designing SLA policies, building aging dashboards, implementing automated escalations, and generating compliance metrics.
When to Use
- When deploying or configuring building vulnerability aging and sla tracking capabilities in your environment
- When establishing security controls aligned to compliance requirements
- When building or improving security architecture for this domain
- When conducting security assessments that require this implementation
Prerequisites
- Vulnerability management platform with historical scan data
- Asset inventory with criticality ratings
- ITSM/ticketing system for remediation tracking
- Reporting platform (Splunk, Elastic, Power BI, Grafana)
- Stakeholder agreement on SLA timelines and escalation procedures
Core Concepts
Standard Vulnerability SLA Framework
| Severity |
CVSS Range |
Standard SLA |
Aggressive SLA |
CISA KEV SLA |
| Critical |
9.0-10.0 |
14 days |
48 hours |
BOD 22-01 due date |
| High |
7.0-8.9 |
30 days |
7 days |
14 days |
| Medium |
4.0-6.9 |
60 days |
30 days |
N/A |
| Low |
0.1-3.9 |
90 days |
60 days |
N/A |
| Informational |
0.0 |
Best effort |
Best effort |
N/A |
Adaptive SLA Modifiers
| Factor |
Modifier |
Rationale |
| Internet-facing asset |
-50% SLA |
Higher exposure risk |
| CISA KEV listed |
Override to 48h |
Active exploitation confirmed |
| EPSS > 0.7 |
-50% SLA |
High exploitation probability |
| Tier 1 (crown jewel) asset |
-25% SLA |
Maximum business impact |
| Compensating control in place |
+25% SLA |
Risk partially mitigated |
| Vendor patch unavailable |
Exception with review date |
Cannot remediate yet |
Key Performance Indicators (KPIs)
| KPI |
Formula |
Target |
| Mean Time to Remediate (MTTR) |
Avg(remediation_date - discovery_date) |
< 30 days overall |
| SLA Compliance Rate |
(Vulns remediated within SLA / Total vulns) * 100 |
>= 90% |
| Overdue Vulnerability Count |
Count where age > SLA |
Trending downward |
| Vulnerability Aging Distribution |
Count by age bucket (0-14d, 15-30d, 31-60d, 60+d) |
Majority in 0-30d |
| Remediation Velocity |
Vulns closed per week |
Trending upward |
| Exception Rate |
(Exceptions / Total vulns) * 100 |
< 5% |
Workflow
Step 1: Define SLA Policy Document
Vulnerability Remediation SLA Policy v1.0
1. Scope: All information systems and applications
2. Severity Classification: Based on CVSS v4.0/v3.1 base score
3. SLA Timelines: See Standard SLA Framework table
4. Adaptive Modifiers: Applied based on asset context
5. Exception Process:
- Must be documented with business justification
- Requires compensating control description
- Maximum extension: 90 days (one renewal)
- CISO approval required for Critical/High exceptions
6. Escalation Path:
- 50% SLA elapsed: Automated reminder to asset owner
- 75% SLA elapsed: Escalation to manager
- 100% SLA elapsed (overdue): CISO notification
- 120% SLA elapsed: VP/CTO escalation
7. Metrics Reporting: Monthly to security committee
Step 2: Build the Aging Calculation Engine
import pandas as pd
from datetime import datetime, timedelta
class VulnerabilityAgingTracker:
"""Track vulnerability aging and SLA compliance."""
SLA_DAYS = {
"Critical": 14,
"High": 30,
"Medium": 60,
"Low": 90,
}
def __init__(self, sla_overrides=None):
if sla_overrides:
self.SLA_DAYS.update(sla_overrides)
def calculate_aging(self, vulns_df):
"""Calculate aging metrics for each vulnerability."""
today = datetime.now()
vulns_df["discovery_date"] = pd.to_datetime(vulns_df["discovery_date"])
vulns_df["remediation_date"] = pd.to_datetime(
vulns_df["remediation_date"], errors="coerce"
)
vulns_df["age_days"] = vulns_df.apply(
lambda row: (row["remediation_date"] - row["discovery_date"]).days
if pd.notna(row["remediation_date"])
else (today - row["discovery_date"]).days,
axis=1
)
vulns_df["sla_days"] = vulns_df["severity"].map(self.SLA_DAYS)
vulns_df["sla_deadline"] = vulns_df["discovery_date"] + \
pd.to_timedelta(vulns_df["sla_days"], unit="D")
vulns_df["is_overdue"] = vulns_df.apply(
lambda row: row["age_days"] > row["sla_days"]
if pd.isna(row["remediation_date"]) else False,
axis=1
)
vulns_df["sla_compliance"] = vulns_df.apply(
lambda row: row["age_days"] <= row["sla_days"]
if pd.notna(row["remediation_date"]) else None,
axis=1
)
vulns_df["days_overdue"] = vulns_df.apply(
lambda row: max(0, row["age_days"] - row["sla_days"])
if row["is_overdue"] else 0,
axis=1
)
vulns_df["sla_pct_elapsed"] = (
vulns_df["age_days"] / vulns_df["sla_days"] * 100
).round(1)
return vulns_df
def generate_kpis(self, vulns_df):
"""Generate KPI summary from aging data."""
open_vulns = vulns_df[vulns_df["remediation_date"].isna()]
closed_vulns = vulns_df[vulns_df["remediation_date"].notna()]
kpis = {
"total_vulnerabilities": len(vulns_df),
"open_vulnerabilities": len(open_vulns),
"closed_vulnerabilities": len(closed_vulns),
"overdue_count": open_vulns["is_overdue"].sum(),
"mttr_days": closed_vulns["age_days"].mean() if len(closed_vulns) > 0 else 0,
"sla_compliance_rate": (
closed_vulns["sla_compliance"].mean() * 100
if len(closed_vulns) > 0 else 0
),
}
kpis["overdue_by_severity"] = (
open_vulns[open_vulns["is_overdue"]]
.groupby("severity")
.size()
.to_dict()
)
return kpis
def get_escalation_list(self, vulns_df):
"""Get vulnerabilities requiring escalation."""
open_vulns = vulns_df[vulns_df["remediation_date"].isna()].copy()
escalations = []
for _, vuln in open_vulns.iterrows():
pct = vuln["sla_pct_elapsed"]
if pct >= 120:
level = "VP/CTO Escalation"
elif pct >= 100:
level = "CISO Notification"
elif pct >= 75:
level = "Manager Escalation"
elif pct >= 50:
level = "Owner Reminder"
else:
continue
escalations.append({
"cve_id": vuln.get("cve_id", ""),
"severity": vuln["severity"],
"age_days": vuln["age_days"],
"sla_days": vuln["sla_days"],
"days_overdue": vuln["days_overdue"],
"sla_pct": pct,
"escalation_level": level,
"asset": vuln.get("asset", ""),
"owner": vuln.get("owner", ""),
})
return pd.DataFrame(escalations)
Step 3: Dashboard Visualization
# Grafana/Kibana query examples for vulnerability aging
# Age distribution histogram (Elasticsearch)
age_distribution_query = {
"aggs": {
"age_buckets": {
"range": {
"field": "age_days",
"ranges": [
{"key": "0-7 days", "to": 8},
{"key": "8-14 days", "from": 8, "to": 15},
{"key": "15-30 days", "from": 15, "to": 31},
{"key": "31-60 days", "from": 31, "to": 61},
{"key": "61-90 days", "from": 61, "to": 91},
{"key": "90+ days", "from": 91},
]
}
}
}
}
# SLA compliance trend (monthly)
sla_trend_query = {
"aggs": {
"monthly": {
"date_histogram": {"field": "remediation_date", "interval": "month"},
"aggs": {
"within_sla": {
"filter": {"script": {
"source": "doc['age_days'].value <= doc['sla_days'].value"
}}
}
}
}
}
}
Best Practices
- Start with achievable SLA targets and tighten them as processes mature
- Adapt SLAs based on asset criticality and threat context, not just CVSS scores
- Automate escalation notifications to reduce manual tracking overhead
- Track MTTR trends month-over-month to demonstrate improvement
- Build exception workflows that require documented compensating controls
- Report SLA compliance to executive leadership monthly for accountability
- Include aging metrics in security committee and board-level reporting
- Integrate SLA tracking with ITSM ticketing for end-to-end remediation visibility
Common Pitfalls
- Setting unrealistic SLA targets that teams cannot meet, causing SLA fatigue
- Not adapting SLAs for asset criticality, treating all systems equally
- Lacking exception processes, forcing teams to either ignore SLAs or request blanket waivers
- Measuring only open vulnerability count without considering age and SLA compliance
- Not tracking the SLA clock from discovery date (using report date instead)
- Failing to re-baseline SLAs as team maturity improves
Related Skills
- implementing-vulnerability-remediation-sla
- building-executive-vulnerability-risk-report
- implementing-security-metrics-and-kpis
- performing-remediation-validation-scanning
1---2name: building-vulnerability-aging-and-sla-tracking3description: Implement a vulnerability aging dashboard and SLA tracking system that measures time-to-remediation against severity-based deadlines (e.g. 14 days critical, 30 days high, 60 days medium, 90 days low), with automated escalations and compliance metrics reporting. Use when designing SLA policies, building aging/remediation dashboards, or proving compliance with remediation timelines.4license: Apache-2.05---6# Building Vulnerability Aging and SLA Tracking
7
8## Overview
9With over 30,000 new vulnerabilities identified in 2024 (a 17% increase from the prior year), organizations must track how long vulnerabilities remain unpatched and whether remediation occurs within defined Service Level Agreements (SLAs). Vulnerability aging measures the time between discovery and remediation, while SLA tracking enforces severity-based deadlines. Industry benchmarks indicate standard SLAs of 14 days for critical, 30 days for high, 60 days for medium, and 90 days for low vulnerabilities, though more aggressive timelines (24-48 hours for actively exploited critical CVEs) are increasingly common. This skill covers designing SLA policies, building aging dashboards, implementing automated escalations, and generating compliance metrics.
10
11
12## When to Use
13
14- When deploying or configuring building vulnerability aging and sla tracking capabilities in your environment
15- When establishing security controls aligned to compliance requirements
16- When building or improving security architecture for this domain
17- When conducting security assessments that require this implementation
18
19## Prerequisites
20- Vulnerability management platform with historical scan data
21- Asset inventory with criticality ratings
22- ITSM/ticketing system for remediation tracking
23- Reporting platform (Splunk, Elastic, Power BI, Grafana)
24- Stakeholder agreement on SLA timelines and escalation procedures
25
26## Core Concepts
27
28### Standard Vulnerability SLA Framework
29
30| Severity | CVSS Range | Standard SLA | Aggressive SLA | CISA KEV SLA |
31|----------|-----------|-------------|----------------|-------------|
32| Critical | 9.0-10.0 | 14 days | 48 hours | BOD 22-01 due date |
33| High | 7.0-8.9 | 30 days | 7 days | 14 days |
34| Medium | 4.0-6.9 | 60 days | 30 days | N/A |
35| Low | 0.1-3.9 | 90 days | 60 days | N/A |
36| Informational | 0.0 | Best effort | Best effort | N/A |
37
38### Adaptive SLA Modifiers
39
40| Factor | Modifier | Rationale |
41|--------|----------|-----------|
42| Internet-facing asset | -50% SLA | Higher exposure risk |
43| CISA KEV listed | Override to 48h | Active exploitation confirmed |
44| EPSS > 0.7 | -50% SLA | High exploitation probability |
45| Tier 1 (crown jewel) asset | -25% SLA | Maximum business impact |
46| Compensating control in place | +25% SLA | Risk partially mitigated |
47| Vendor patch unavailable | Exception with review date | Cannot remediate yet |
48
49### Key Performance Indicators (KPIs)
50
51| KPI | Formula | Target |
52|-----|---------|--------|
53| Mean Time to Remediate (MTTR) | Avg(remediation_date - discovery_date) | < 30 days overall |
54| SLA Compliance Rate | (Vulns remediated within SLA / Total vulns) * 100 | >= 90% |
55| Overdue Vulnerability Count | Count where age > SLA | Trending downward |
56| Vulnerability Aging Distribution | Count by age bucket (0-14d, 15-30d, 31-60d, 60+d) | Majority in 0-30d |
57| Remediation Velocity | Vulns closed per week | Trending upward |
58| Exception Rate | (Exceptions / Total vulns) * 100 | < 5% |
59
60## Workflow
61
62### Step 1: Define SLA Policy Document
63
64```
65Vulnerability Remediation SLA Policy v1.0
66
671. Scope: All information systems and applications
682. Severity Classification: Based on CVSS v4.0/v3.1 base score
693. SLA Timelines: See Standard SLA Framework table
704. Adaptive Modifiers: Applied based on asset context
715. Exception Process:
72 - Must be documented with business justification
73 - Requires compensating control description
74 - Maximum extension: 90 days (one renewal)
75 - CISO approval required for Critical/High exceptions
766. Escalation Path:
77 - 50% SLA elapsed: Automated reminder to asset owner
78 - 75% SLA elapsed: Escalation to manager
79 - 100% SLA elapsed (overdue): CISO notification
80 - 120% SLA elapsed: VP/CTO escalation
817. Metrics Reporting: Monthly to security committee
82```
83
84### Step 2: Build the Aging Calculation Engine
85
86```python
87import pandas as pd
88from datetime import datetime, timedelta
89
90class VulnerabilityAgingTracker:
91 """Track vulnerability aging and SLA compliance."""
92
93 SLA_DAYS = {
94 "Critical": 14,
95 "High": 30,
96 "Medium": 60,
97 "Low": 90,
98 }
99
100 def __init__(self, sla_overrides=None):
101 if sla_overrides:
102 self.SLA_DAYS.update(sla_overrides)
103
104 def calculate_aging(self, vulns_df):
105 """Calculate aging metrics for each vulnerability."""
106 today = datetime.now()
107
108 vulns_df["discovery_date"] = pd.to_datetime(vulns_df["discovery_date"])
109 vulns_df["remediation_date"] = pd.to_datetime(
110 vulns_df["remediation_date"], errors="coerce"
111 )
112
113 vulns_df["age_days"] = vulns_df.apply(
114 lambda row: (row["remediation_date"] - row["discovery_date"]).days
115 if pd.notna(row["remediation_date"])
116 else (today - row["discovery_date"]).days,
117 axis=1
118 )
119
120 vulns_df["sla_days"] = vulns_df["severity"].map(self.SLA_DAYS)
121 vulns_df["sla_deadline"] = vulns_df["discovery_date"] + \
122 pd.to_timedelta(vulns_df["sla_days"], unit="D")
123
124 vulns_df["is_overdue"] = vulns_df.apply(
125 lambda row: row["age_days"] > row["sla_days"]
126 if pd.isna(row["remediation_date"]) else False,
127 axis=1
128 )
129
130 vulns_df["sla_compliance"] = vulns_df.apply(
131 lambda row: row["age_days"] <= row["sla_days"]
132 if pd.notna(row["remediation_date"]) else None,
133 axis=1
134 )
135
136 vulns_df["days_overdue"] = vulns_df.apply(
137 lambda row: max(0, row["age_days"] - row["sla_days"])
138 if row["is_overdue"] else 0,
139 axis=1
140 )
141
142 vulns_df["sla_pct_elapsed"] = (
143 vulns_df["age_days"] / vulns_df["sla_days"] * 100
144 ).round(1)
145
146 return vulns_df
147
148 def generate_kpis(self, vulns_df):
149 """Generate KPI summary from aging data."""
150 open_vulns = vulns_df[vulns_df["remediation_date"].isna()]
151 closed_vulns = vulns_df[vulns_df["remediation_date"].notna()]
152
153 kpis = {
154 "total_vulnerabilities": len(vulns_df),
155 "open_vulnerabilities": len(open_vulns),
156 "closed_vulnerabilities": len(closed_vulns),
157 "overdue_count": open_vulns["is_overdue"].sum(),
158 "mttr_days": closed_vulns["age_days"].mean() if len(closed_vulns) > 0 else 0,
159 "sla_compliance_rate": (
160 closed_vulns["sla_compliance"].mean() * 100
161 if len(closed_vulns) > 0 else 0
162 ),
163 }
164
165 kpis["overdue_by_severity"] = (
166 open_vulns[open_vulns["is_overdue"]]
167 .groupby("severity")
168 .size()
169 .to_dict()
170 )
171
172 return kpis
173
174 def get_escalation_list(self, vulns_df):
175 """Get vulnerabilities requiring escalation."""
176 open_vulns = vulns_df[vulns_df["remediation_date"].isna()].copy()
177
178 escalations = []
179 for _, vuln in open_vulns.iterrows():
180 pct = vuln["sla_pct_elapsed"]
181 if pct >= 120:
182 level = "VP/CTO Escalation"
183 elif pct >= 100:
184 level = "CISO Notification"
185 elif pct >= 75:
186 level = "Manager Escalation"
187 elif pct >= 50:
188 level = "Owner Reminder"
189 else:
190 continue
191
192 escalations.append({
193 "cve_id": vuln.get("cve_id", ""),
194 "severity": vuln["severity"],
195 "age_days": vuln["age_days"],
196 "sla_days": vuln["sla_days"],
197 "days_overdue": vuln["days_overdue"],
198 "sla_pct": pct,
199 "escalation_level": level,
200 "asset": vuln.get("asset", ""),
201 "owner": vuln.get("owner", ""),
202 })
203
204 return pd.DataFrame(escalations)
205```
206
207### Step 3: Dashboard Visualization
208
209```python
210# Grafana/Kibana query examples for vulnerability aging
211
212# Age distribution histogram (Elasticsearch)
213age_distribution_query = {
214 "aggs": {
215 "age_buckets": {
216 "range": {
217 "field": "age_days",
218 "ranges": [
219 {"key": "0-7 days", "to": 8},
220 {"key": "8-14 days", "from": 8, "to": 15},
221 {"key": "15-30 days", "from": 15, "to": 31},
222 {"key": "31-60 days", "from": 31, "to": 61},
223 {"key": "61-90 days", "from": 61, "to": 91},
224 {"key": "90+ days", "from": 91},
225 ]
226 }
227 }
228 }
229}
230
231# SLA compliance trend (monthly)
232sla_trend_query = {
233 "aggs": {
234 "monthly": {
235 "date_histogram": {"field": "remediation_date", "interval": "month"},
236 "aggs": {
237 "within_sla": {
238 "filter": {"script": {
239 "source": "doc['age_days'].value <= doc['sla_days'].value"
240 }}
241 }
242 }
243 }
244 }
245}
246```
247
248## Best Practices
2491. Start with achievable SLA targets and tighten them as processes mature
2502. Adapt SLAs based on asset criticality and threat context, not just CVSS scores
2513. Automate escalation notifications to reduce manual tracking overhead
2524. Track MTTR trends month-over-month to demonstrate improvement
2535. Build exception workflows that require documented compensating controls
2546. Report SLA compliance to executive leadership monthly for accountability
2557. Include aging metrics in security committee and board-level reporting
2568. Integrate SLA tracking with ITSM ticketing for end-to-end remediation visibility
257
258## Common Pitfalls
259- Setting unrealistic SLA targets that teams cannot meet, causing SLA fatigue
260- Not adapting SLAs for asset criticality, treating all systems equally
261- Lacking exception processes, forcing teams to either ignore SLAs or request blanket waivers
262- Measuring only open vulnerability count without considering age and SLA compliance
263- Not tracking the SLA clock from discovery date (using report date instead)
264- Failing to re-baseline SLAs as team maturity improves
265
266## Related Skills
267- implementing-vulnerability-remediation-sla
268- building-executive-vulnerability-risk-report
269- implementing-security-metrics-and-kpis
270- performing-remediation-validation-scanning