The Security Tax in WSL3: Benchmarking Kernel Mitigations Against WSL 2.x
In our earlier look at The Security Tax in WSL 2, we showed how kernel-level mitigations for CPU speculative execution vulnerabilities (Spectre, Meltdown, Retbleed) imposed a steep penalty on developer tasks. On a real-world Go backend compilation workload, turning off mitigations slashed kernel system time by 45.7% and delivered a 31.7% drop in wall-clock compile time.
With the rollout of WSL v3.0.2 (running Linux kernel 6.18.40.1-microsoft-standard-WSL2), how does the security tax hold up today? Does modern hypervisor paravirtualization eliminate the overhead, or are developers still paying an invisible performance toll on syscall-heavy workloads?
We ran microbenchmarks and multi-process scheduler stress tests comparing default WSL3 (mitigations=on) against tuned WSL3 (mitigations=off), and contrasted the findings with our earlier WSL 2.x benchmark data.
The WSL 2.x Baseline: Go Compilation & IPC Sched Pipe
In our earlier WSL 2 benchmark, we measured both the macro Go backend compilation and synthetic inter-process communication using perf bench sched pipe -l 100000.
WSL 2.x Go Compilation (go build)
| WSL 2.x Metric | Mitigations ON (Default) | Mitigations OFF | Difference |
|---|---|---|---|
| User Time | 81.87s | 82.15s | Negligible (+0.3%) |
| System Time | 24.89s | 13.51s | -45.7% ▼ |
| Total (Wall) Time | 26.98s | 18.31s | -31.7% ▼ |
| CPU Efficiency | 395% | 522% | +127% ▲ |
WSL 2.x Synthetic IPC (perf bench sched pipe -l 100000)
| WSL 2.x Metric | Mitigations ON (Default) | Mitigations OFF | Difference |
|---|---|---|---|
| Throughput (ops/sec) | 22,493 ops/s | 26,659 ops/s | +18.52% ▲ |
| Latency (usecs/op) | 44.45 µs | 37.51 µs | -15.61% ▼ |
| Total Execution Time | 4.445s | 3.751s | -15.61% ▼ |
In WSL 2.x, disabling mitigations yielded a respectable ~18.5% gain in IPC throughput (jumping from 22.5k to 26.7k ops/sec).
Apples-to-Apples: WSL 2 vs WSL 3 IPC Performance (perf bench sched pipe)
Because both benchmarks ran the exact same command (perf bench sched pipe -l 100000), we can line up the raw operations per second across environments:
| Environment | Mitigations ON | Mitigations OFF | Mitigation Penalty / Gain |
|---|---|---|---|
| WSL 2.x | 22,493 ops/s (44.45 µs) | 26,659 ops/s (37.51 µs) | +18.5% ops/sec (+4,166 ops/s) |
| WSL v3.0.2 | 123,159 ops/s (8.12 µs) | 217,736 ops/s (4.59 µs) | +76.8% ops/sec (+94,577 ops/s) |
| Generational Uplift | 5.48x faster | 8.17x faster | — |
Key Observations from the Comparison:
- Dramatic Baseline Architectural Uplift: WSL 3.0.2 with Linux
6.18.xstarts off 5.5x faster than WSL 2 even with mitigations active (123k ops/s vs 22.5k ops/s), reflecting modern hypervisor optimizations and kernel IPC efficiency. - The Mitigation Tax Has Widened: While disabling mitigations in WSL 2 granted an 18.5% throughput bump (+4,166 ops/s), in WSL 3 it unlocks a massive +76.8% jump (+94,577 ops/s). As the rest of the virtualized virtualization and IPC stack became drastically faster, speculation barriers and context-switch fences became the dominant remaining bottleneck.
The WSL v3.0.2 Testbed & Workload
To pinpoint exactly where the overhead originates in the new architecture, we benchmarked WSL v3.0.2 across entry/exit latency, IPC pipe bandwidth, and multi-process scheduler contention using native perf bench suites.
Environment & Setup
- Host Platform: Windows 11 / WSL v3.0.2 (Hyper-V VM)
- Kernel:
6.18.40.1-microsoft-standard-WSL2 - Allocation: 16 vCPUs
Kernel Command Line Comparison
- Mitigations ON (Default):
initrd=\initrd.img WSL_ROOT_INIT=1 panic=-1 nr_cpus=16 hv_utils.timesync_implicit=1 printk.devkmsg=on page_reporting.page_reporting_order=5 swiotlb=force hv_pci_swiotlb=67108864 hv_storvsc.storvsc_max_hw_queues=4 console=hvc0 debug pty.legacy_count=0 WSL_ENABLE_CRASH_DUMP=1 - Mitigations OFF (Tuned):
initrd=\initrd.img WSL_ROOT_INIT=1 panic=-1 nr_cpus=16 hv_utils.timesync_implicit=1 printk.devkmsg=on page_reporting.page_reporting_order=5 swiotlb=force hv_pci_swiotlb=67108864 hv_storvsc.storvsc_max_hw_queues=4 console=hvc0 debug pty.legacy_count=0 WSL_ENABLE_CRASH_DUMP=1 quiet loglevel=0 audit=0 mitigations=off vsyscall=none
Executive Summary: WSL v3.0.2 Benchmark Results
| Benchmark Metric | Mitigations ON | Mitigations OFF | Delta (OFF vs ON) | Improvement Ratio |
|---|---|---|---|---|
Go Build (go build -a ./...) Total |
41.558 s | 39.842 s | -4.13% ▼ | 1.04x faster |
| Go Build System Time | 71.41 s | 55.52 s | -22.25% ▼ | 1.29x faster (kernel overhead) |
Syscall Basic Throughput (getppid) |
5,353,135 ops/s | 14,798,678 ops/s | +176.45% ▲ | 2.76x faster |
| Syscall Basic Latency | 0.1868 µs/op | 0.0676 µs/op | -63.83% ▼ | 2.76x lower latency |
| Sched Pipe Throughput (IPC ping-pong) | 123,159 ops/s | 217,736 ops/s | +76.79% ▲ | 1.77x faster |
| Sched Pipe Latency | 8.1196 µs/op | 4.5927 µs/op | -43.44% ▼ | 1.77x lower latency |
Sched Messaging (hackbench 400 proc) |
1.682 s | 0.997 s | -40.73% ▼ | 1.69x faster |
Detailed Analysis: Where WSL3 Spends Its Cycles
1. Raw Syscall Latency & Throughput (perf bench syscall basic -l 10000000)
- Mitigations ON: 5,353,135 ops/s (0.1868 µs/op)
- Mitigations OFF: 14,798,678 ops/s (0.0676 µs/op)
- Impact: +176.45% throughput gain (2.76x speedup).
Kernel Page Table Isolation (KPTI) page-table swapping and speculation barriers on guest kernel entry/exit remain the single heaviest tax inside virtualized Linux. Removing these barriers drops per-syscall latency from 186.8 ns down to 67.6 ns.
2. IPC Pipe Ping-Pong (perf bench sched pipe -l 100000)
- Mitigations ON: 123,159 ops/s (8.12 µs/op)
- Mitigations OFF: 217,736 ops/s (4.59 µs/op)
- Impact: +76.79% throughput gain (1.77x speedup).
Inter-process communication via pipes stresses context switching alongside indirect branch speculation barriers (IBPB and STIBP). By disabling mitigations, ping-pong latency drops from 8.12 µs to 4.59 µs.
3. Multi-Process Contention (perf bench sched messaging -g 20 -l 1000)
- Mitigations ON: 1.682 seconds total runtime
- Mitigations OFF: 0.997 seconds total runtime
- Impact: -40.73% wall-clock reduction (1.69x faster).
hackbench simulates 400 concurrent processes exchanging messages across sockets and pipes. This mirrors the behavior of parallel toolchains: compiler processes, Node.js/Jest test suites, and package manager indexing. Disabling mitigations eliminates more than 40% of runtime under heavy multi-process contention.
4. Real-World Macro Benchmark: Clean Go Compilation (go build -a ./...)
We tested a full clean rebuild of the codebase with go build -a ./...:
| Metric | Mitigations ON | Mitigations OFF | Delta (OFF vs ON) | Improvement Ratio |
|---|---|---|---|---|
| User Time | 341.21s | 343.06s | +0.54% | Baseline compute parity |
| System Time | 71.41s | 55.52s | -22.25% ▼ | 1.29x faster (kernel overhead) |
| Total (Wall) Time | 41.558s | 39.842s | -4.13% ▼ | 1.04x faster |
| CPU Utilization | 992% | 1000% | +8% | Higher multi-core saturation |
While pure user-space compilation cycles stayed virtually identical (~341s vs ~343s), disabling mitigations saved nearly 16 seconds of kernel system time (-22.25%), directly translating to a lower wall-clock compilation time and allowing the compiler toolchain to maintain a solid 1000% CPU utilization across threads.
WSL 2.x vs WSL 3.x: The Evolution of the Tax
Comparing our earlier findings with the current WSL v3.0.2 data yields clear patterns:
| Aspect | WSL 2.x Series | WSL 3.x Series | Takeaway |
|---|---|---|---|
| Go Compile Sys Time Savings | -45.7% (24.89s -> 13.51s) | -22.3% (71.41s -> 55.52s) | Kernel syscall tax remains substantial during heavy parallel compilation. |
| Macro Workload Impact | -31.7% total build time (Go compiler) | -40.7% total messaging time (hackbench), -4.1% Go build total |
High concurrency workloads still experience severe contention taxes. |
| Kernel / System Time Penalty | 45.7% of kernel time wasted on mitigations | 63.8% syscall latency penalty (0.1868 µs vs 0.0676 µs) | Raw ring-transition overhead remains severe inside Hyper-V VMs. |
| IPC & Thread Scheduling | High context-switch latency | Pipe throughput throttled by 43.4% when mitigations active | Modern 6.18 kernels still incur heavy branch-prediction barrier overhead. |
While WSL 3.x provides a modernized kernel base (6.18.x) and smoother Hyper-V integration, the fundamental architectural cost of speculative execution barriers inside a virtual machine has not decreased. If anything, as CPU core counts increase (16 vCPUs in our testbed), multi-threaded speculation fences compound across active cores.
CPU Vulnerability Diffs in WSL v3.0.2
Checking /sys/devices/system/cpu/vulnerabilities/* reveals the exact protections toggled:
| Vulnerability | Mitigations ON (Default) | Mitigations OFF (Tuned) |
|---|---|---|
spectre_v1 |
Mitigation: usercopy/swapgs barriers and __user pointer sanitization | Vulnerable: __user pointer sanitization and usercopy barriers only; no swapgs barriers |
spectre_v2 |
Mitigation: Retpolines; IBPB: conditional; IBRS_FW; STIBP: always-on; RSB filling; PBRSB-eIBRS: Not affected; BHI: Not affected | Vulnerable; IBPB: disabled; STIBP: disabled; PBRSB-eIBRS: Not affected; BHI: Not affected |
spec_store_bypass |
Mitigation: Speculative Store Bypass disabled via prctl | Vulnerable |
spec_rstack_overflow |
Vulnerable: Safe RET, no microcode | Vulnerable |
tsa |
Vulnerable: No microcode | Vulnerable |
Configuring .wslconfig for Maximum Throughput
For dedicated local development machines where untrusted multi-tenant code is not being executed inside the same WSL guest, you can disable the mitigation tax via %UserProfile%\.wslconfig:
[wsl2]
# Disable speculative execution barriers and audit overhead
kernelCommandLine=mitigations=off quiet loglevel=0 audit=0 vsyscall=none
After modifying the file, restart the WSL VM from PowerShell:
wsl --shutdown
Once rebooted, verify the status inside WSL:
grep -H . /sys/devices/system/cpu/vulnerabilities/*