Ferritaas Incident Intelligence
Defect Intelligence Report Bug 0899120 Severity: High (Route Instability)

SD-WAN SLA Jitter Threshold Flapping Triggers Route Instability Storm

Sub-millisecond probe sensitivity and absence of state hysteresis in FortiOS SD-WAN Performance SLA health checks cause continuous link failover loops, desynchronizing enterprise VoIP calls and active database connections.

Technical Root Cause Analysis

FortiOS SD-WAN uses active synthetic probe packets (ICMP, HTTP, TWAMP) dispatched every 500ms to calculate link latency, jitter, and packet loss metrics against configured SLA thresholds.

When administrators configure stringent jitter targets (e.g. 5ms) without adequate dampening intervals, standard internet jitter variance causes the calculated moving metric to breach the threshold for a single probe cycle. The link steering engine instantly evicts the primary WAN member from the SLA group and shifts routes to the secondary link. Moments later, the primary link recovers, triggering an immediate switchback. This rapid "ping-pong" route oscillation destroys TCP throughput and causes session drops.

Affected Firmware & Blast Radius Matrix

FortiOS Branch Vulnerable Builds Confirmed Clean Build Status & Workaround
FortiOS 7.2 7.2.0 – 7.2.6 7.2.7+ Apply failtime 5 & recoverytime 10
FortiOS 7.4 7.4.0 – 7.4.2 7.4.3+ Exponential moving average smoothing
FortiOS 7.0 Moderate Impact 7.0.14+ Increase probe interval to 1000ms

Platform Impact: Branch and campus SD-WAN deployments steering voice, video, and mission-critical cloud ERP traffic across hybrid MPLS/Broadband/5G circuits.

Step 01: Free Verification CLI (Safe Read-Only)

Execute these commands to inspect active SLA probe measurements, link flapping logs, and SD-WAN steering state:

Diagnostic Commands

# 1. Check live SLA probe latency, jitter, and packet loss stats
diagnose sys sdwan health-check status

# 2. Dump recent SLA transition logs to detect flapping frequency
diagnose sys sdwan sla-log

# 3. Inspect active SD-WAN steering service rules
diagnose sys sdwan service

# 4. Verify current kernel routing table entries
get router info routing-table all

Remediation & Workaround Steps (Teaser Preview)

Access the complete SD-WAN SLA stabilization and hysteresis tuning runbook in the Ferrite interactive platform:

Step 02: Enforce Probe Failtime and Recoverytime Hysteresis

Configure failtime 5 and recoverytime 10 to require sustained failure before link failover.

🔒 Interactive CLI Available in Ferrite Runbook #13

Step 03: Apply Link Flap Hold-Down Timer

Configure link-cost-factor and hold-down intervals to eliminate rapid route oscillating.

🔒 Interactive CLI Available in Ferrite Runbook #13

Step 04: Session Preservation on Primary WAN Member

Bind existing TCP sessions to primary path until natural termination during minor jitter spikes.

🔒 Interactive CLI Available in Ferrite Runbook #13
⚡ Ferrite Platform Superpowers

Execute Runbook #13 with Live Browser Automation

Connect your FortiGate via browser console (Web Serial) or local SSH bridge, verify each command in real-time, generate ready-to-run Tera Term scripts, and export sanitized TAC dossiers.

Live Browser Automation Direct terminal connection with live step checkoff.
📟
1-Click Tera Term (.ttl) Generate scripts for air-gapped jumpboxes.
🛡️
Zero-Trust Scrubber Scrub serials and credentials in local browser RAM.
📄
TAC P1 Escalation Dossier Standardized evidence export with SHA-256 seal.

Frequently Asked Questions

What causes Bug 0899120?

Lack of hysteresis dampening in the SD-WAN health-check probe calculation causes brief internet jitter spikes to trigger instant link failovers and rapid switchbacks.

What is the recommended failtime and recoverytime?

Setting failtime to 5 (5 consecutive dropped/out-of-spec probes) and recoverytime to 10 prevents spurious flapping on commercial broadband.

Which FortiOS releases fix Bug 0899120?

FortiOS 7.2.7 and 7.4.3 introduced moving-average exponential metric smoothing to filter out isolated packet anomalies.