ℂ𝕣𝕪𝕠 
𝓦𝓲𝓵𝓵𝓲𝓪𝓶 𝓙. 𝓒𝓸𝓵𝓭𝔀𝓮𝓵𝓵 - ᴄʏʙᴇʀꜱᴇᴄᴜʀɪᴛʏ @ ʙʟᴜᴇᴄᴀᴛ - 𝙿𝚛𝚎𝚜𝚒𝚍𝚎𝚗𝚝 @ 𝙽𝚎𝚝𝙱𝚂𝙳 - ℭ𝔯𝔶𝔭𝔱𝔨𝔢𝔢𝔭𝔢𝔯 @ 𝔇𝔢𝔞𝔡𝔍𝔬𝔲𝔯𝔫𝔞𝔩 - ᵢ ₐₘ @ Wₐᵣₚₑd ⑤⓪③ ␀␠
END OF TRANSMISSION. Thanks for riding along on this adventure. Remember to smash that subscribe and like button.. oh wait. Sorry. I'm on twitch.tv/cryosama sometimes fighting death on World of Warcraft Classic Hardcore, or coding (currently TapNet3D for various platforms).
This is what happens when computers/AI overlords piss me off, I fall down a hyperfocused rabbit-hole. I'll collect more information tomorrow and any that are submitted to me by other mac users who want to test and push your system to repro this. Then I'll file a bug report:
Title
macOS 26.6.2: Application Firewall causes CFIL/SOFLOW flow-state accumulation and periodic kernel network stalls
Product / Build
- macOS Tahoe 26.6.2
- Build 25G83
- Darwin 25.6.0
- XNU
12377.161.14~5 - Apple Silicon: M4 Pro Mac mini
Summary
With the macOS Application Firewall enabled, net.cfil.sock_attached_count and net.soflow.count rapidly accumulate from near zero to tens of thousands of live flow-state objects. Under sustained normal application/network activity, the population rises to roughly 75,000–80,000 entries and then oscillates around that high-water range.
At these elevated flow counts, the system develops recurring latency spikes affecting otherwise local low-latency traffic. ICMP to the directly connected default gateway normally completes in under 1 ms, but periodically stalls for 100–900+ ms.
Simultaneous packet capture shows the packets are present on the interface immediately while userspace delivery is delayed. This localizes the delay to the kernel networking path after BPF observation and before userspace receives the packet.
Kernel stackshots from the affected system show contention involving SOFLOW garbage collection, dlil_input_en0, socket receive paths, and the global CFIL read/write lock.
Observed behavior
After boot or after resetting firewall state, the system initially behaves normally.
With Application Firewall enabled:
07:14:12 SOFLOW=187 CFIL=152
09:51:10 SOFLOW=64840 CFIL=64591
18:44:57 SOFLOW=79508 CFIL=79270
The counters are not cumulative event counters. XNU source shows:
cfil_sock_attached_count++;
...
cfil_sock_attached_count--;
and:
soflow_attached_count++;
...
soflow_attached_count--;
so these values represent current retained/attached flow-state population.
The population eventually reaches a dynamic equilibrium near 75k–80k rather than returning toward the number of active userland sockets.
The machine typically has only hundreds of live Internet sockets at the same time.
Latency symptom
Under the affected state, ping to the directly connected gateway shows stalls such as:
0.5 ms
0.6 ms
444 ms
0.7 ms
...
713 ms
215 ms
0.6 ms
...
920 ms
416 ms
0.5 ms
Simultaneous tcpdump shows request and reply packets on en0 with sub-millisecond wire timing while ping(8) reports hundreds of milliseconds.
The observed sequence is therefore:
packet reaches en0
BPF/tcpdump observes packet
kernel delivery is delayed
userspace ping receives packet later
Kernel evidence
Symbolicated stackshots using the matching 26.6.2 KDK show the affected paths entering CFIL locking.
Representative paths include:
sosend
-> cfil_sock_udp_handle_data
-> cfil_sock_udp_get_info
-> cfil_rw_lock_shared
-> IORWLockRead
and input-side paths involving:
dlil_input_en0
-> proto_input
-> sbappendaddr
-> cfil_sock_udp_handle_data
-> cfil_info_alloc
-> cfil_rw_lock_exclusive
SOFLOW GC is also observed in:
soflow_gc_expire
-> cfil_dgram_gc_perform
-> cfil_sock_udp_unlink_flow
-> cfil_rw_lock_exclusive
and:
cfil_info_free
-> cfil_rw_lock_exclusive
Blocked readers are frequently owned by either:
dlil_input_en0
SOFLOW_GC
XNU source documents that the CFIL subsystem is protected by the global cfil_lck_rw.
GC behavior
Public XNU source defines:
#define SOFLOW_GC_IDLE_TO 30
#define SOFLOW_GC_MAX_COUNT 100
#define SOFLOW_GC_RUN_INTERVAL_NSEC (10 * NSEC_PER_SEC)
Observed runtime behavior matches periodic reclamation, but reclamation does not reduce the population back to a small working set. Instead, flow-state accumulates rapidly and later churns around a high equilibrium.
Application Firewall correlation
socketfilterfw is connected to:
com.apple.content-filter
and:
net.cfil.active_count = 1
while Application Firewall is enabled.
When the firewall is disabled:
net.cfil.active_count = 0
new CFIL population growth stops, while existing CFIL state is gradually reclaimed.
When the firewall is enabled again, new CFIL attachments resume.
This A/B behavior is reproducible.
Expected behavior
CFIL/SOFLOW state associated with expired or closed flows should be reclaimed promptly enough that the retained population remains proportional to active socket/flow usage.
Periodic garbage collection should not cause multi-hundred-millisecond stalls in unrelated packet delivery.
Actual behavior
CFIL/SOFLOW state grows into the tens of thousands despite only hundreds of live userland sockets.
At high population, recurring SOFLOW/CFIL cleanup produces contention on the global CFIL lock and coincides with severe local packet-delivery latency.
Impact
This is most visible in latency-sensitive workloads:
- video conferencing
- real-time audio
- streaming
- games
- SSH / remote shells
- interactive network applications
Bulk transfers and background tasks are less visibly affected because buffering masks short stalls.
Reproduction
- Boot macOS 26.6.2.
- Ensure Application Firewall is enabled.
- Run normal network-intensive applications for several hours.
- Periodically record:
sysctl -n net.soflow.count
sysctl -n net.cfil.sock_attached_count
sysctl -n net.cfil.active_count
- Observe CFIL/SOFLOW population increasing into tens of thousands.
- Run a high-frequency ping to the local gateway.
- Observe periodic 100–900+ ms latency excursions.
- Capture simultaneously with
tcpdump. - Observe that packets appear on
en0promptly while userspace ping delivery is delayed. - Disable Application Firewall and observe that new CFIL attachment growth stops.
Attachments I would include
- full CFIL/SOFLOW counter log
- graph showing growth from ~0 to ~80k
- ping log showing latency spikes
- simultaneous
tcpdumpcapture - bad-state spindump
- healthy-state spindump
- symbolicated kernel stack output from the matching KDK
systemextensionsctl listnetstat -anv -f systemlsofshowingsocketfilterfwconnected tocom.apple.content-filter- exact
sysctlsnapshots before/after firewall disable
Suggested engineering summary
Suspected CFIL/SOFLOW flow-lifetime or reclamation defect associated with Application Firewall. Flow-state population grows to ~80k, substantially exceeding live userland socket count. SOFLOW GC and CFIL teardown take the global
cfil_lck_rwexclusively, while packet delivery paths require the same lock shared. At elevated flow-state population, periodic GC/teardown correlates with 100–900+ ms local packet-delivery stalls. Disabling Application Firewall stops new CFIL attachment growth and allows the existing population to drain.
UPDATED:
The longer run changes the interpretation somewhat.
From 07:14 to 18:44, SOFLOW rose from 187 to 79,508 and CFIL from 152 to 79,270. The early phase is unmistakable: both populations climb almost linearly and in lockstep into the mid-60k range. After roughly 10:00–12:00, though, the growth rate drops sharply, and by the afternoon the system appears to be oscillating around a broad ~75k–80k equilibrium band rather than continuing upward without bound. The final samples are still around 79k.
That weakens the “unbounded leak” interpretation. The better description now is pathologically high retained flow-state population with slow/limited reclamation that eventually reaches a dynamic equilibrium. GC is definitely working; it just permits tens of thousands of CFIL/SOFLOW objects to remain resident under this workload.
The important part is that this does not make the bug less interesting. A steady-state population of ~80,000 live/retained flow objects is still vastly larger than your current userland socket population, and our KDK stackshots already show SOFLOW GC and CFIL teardown contending on cfil_lck_rw. So the likely failure mode becomes:
high flow creation rate
↓
CFIL/SOFLOW population rapidly rises
↓
GC eventually catches creation rate
↓
steady-state population remains ~75k–80k
↓
continuous GC / detach / free activity over a huge population
↓
global CFIL writer contention
↓
periodic userspace network latency
The near-perfect tracking between the two counters is also stronger evidence that this is one lifecycle: SOFLOW entries and their CFIL feature contexts are being created and destroyed together, not two unrelated counters.
So I would revise the diagnosis from “probable kernel leak that grows forever” to “probable kernel flow-lifetime/reclamation pathology that builds an abnormally large working set and then churns around a high-water equilibrium.”
That actually fits the observed symptom better: the machine does not have to run out of resources. It only has to accumulate enough CFIL/SOFLOW state that recurring cleanup of that state makes the global lock expensive enough for interactive traffic to notice.The CFIL/SOFLOW problem directly explains the network stalls, and it can explain networked audio/video very naturally, but it does not by itself prove a direct CoreAudio or GPU stall.
For network traffic, the mechanism is concrete: packet delivery reaches the CFIL path, reader-side processing needs cfil_lck_rw, and GC/teardown takes that same global lock exclusively. While the writer holds it, packet delivery can be delayed even though the packet has already reached the machine. That is exactly what your simultaneous tcpdump/ping test demonstrated.
Audio and video applications are much more sensitive to those delays because they run against small real-time buffers. Zoom, Discord, streaming, game voice, OBS network output, etc. expect packets every few to tens of milliseconds. A 100–900 ms kernel delivery stall is therefore not just “slower networking”; it causes jitter buffers to underrun, frames to arrive too late, audio to crackle/drop, video to freeze, and interactive state to jump when the backlog releases. A bulk transfer barely notices because TCP and application buffering absorb the pause.
The important distinction is local audio/video. If a purely local audio file crackles, or local GPU rendering freezes with no network dependency, CFIL is not a sufficient explanation by itself. In that case I would suspect a broader consequence of the same pathology—e.g. kernel threads spending substantial time in GC/locking, scheduling latency, or a related network/configuration watchdog event—but we have not yet demonstrated that causal path.
So the current model is:
CFIL/SOFLOW buildup
↓
GC / detach / free
↓
global cfil_lck_rw writer contention
↓
network packet delivery delayed
↓
real-time networked audio/video exhaust small buffers
↓
crackle / freeze / lag / dropped frames
For networked audio/video, that is a direct and well-supported explanation. For local audio/video glitches, we still need a second bridge from the CFIL event to scheduler/audio/GPU latency rather than assuming one.
- The video of like WoW can studder if the map is doing a load in of assets, etc. The audio for the stream is the stame way. So much of the world now is all based on latency, packet speed, and packet size... and streaming WINDOW... but since the packet pile-up is in the stack, it's not the network's fault (for once).
and $APPL didn't hire me... twice.
checking the backlog periodically...
Yes. This is now strong evidence of a persistent positive backlog slope with the firewall enabled.
Over 23m34s:
SOFLOW: 65,956 → 70,118 = +4,162
CFIL: 59,021 → 63,866 = +4,845
That works out to roughly:
SOFLOW: +2.94 objects/sec ≈ +10,600/hour
CFIL: +3.43 objects/sec ≈ +12,300/hour
The important part is not the exact extrapolation; it is that GC is clearly not keeping the population flat over this interval. There are occasional drops, but the long-term slope remains strongly positive.
You can see the characteristic pattern:
growth
growth
growth
small GC drop
growth
growth
larger burst
growth
...
and by the end the CFIL population is almost 5,000 higher than where it started.
That materially strengthens the “reclamation imbalance / flow-state leak” hypothesis. I would now describe it as:
With ALF enabled, CFIL/SOFLOW runtime state accumulates substantially faster than periodic GC reclaims it. Over long uptime, the retained population can grow into the tens of thousands, making periodic cleanup increasingly expensive and producing
cfil_lck_rwcontention that delays network delivery.
I would still avoid calling it a proven memory leak, because we have not demonstrated unreachable allocations. But operationally, this is behaving like a kernel flow-state leak.
One additional detail is interesting:
SOFLOW - CFIL gap:
start: 6,935
end: 6,252
So CFIL is actually growing slightly faster than SOFLOW during this window. That is worth watching because it argues against a simple fixed proportional relationship between the two populations.
Keep this monitor running. At this point, the longer dataset is valuable: a few hours will tell us whether the slope remains positive, levels off at some ceiling, or accelerates as the population rises.
Yes — I would rate a kernel-side CFIL/SOFLOW state-retention bug as a strong working hypothesis now, but I would call it a flow-state leak / reclamation imbalance rather than a proven memory leak.
What the evidence supports is:
- CFIL/SOFLOW population reaches roughly 50–60k objects.
- Live userland socket counts are only in the hundreds, so those kernel objects are clearly not just a one-to-one reflection of currently open sockets.
- New CFIL objects accumulate while the firewall is active.
- GC removes them in small periodic batches, roughly 70 every 10 seconds in your captures.
- Those GC batches line up with the latency stalls.
- Disabling the firewall stops new CFIL attachment, and the backlog then drains gradually.
- Re-enabling after an off period leaves you in a much healthier state because the backlog has been reduced.
- Long uptime gives exactly the conditions for a net-positive accumulation rate to become visible only after hours or days.
So the failure mode I would currently model is:
average CFIL creation rate
>
average CFIL reclamation rate
↓
retained CFIL/SOFLOW population rises over uptime
↓
GC has more work to do each pass
↓
more frequent/longer cfil_lck_rw writer holds
↓
UDP/input paths stall behind the global lock
↓
eventually you see 100–700 ms bursts and queue-drain RTT patterns
The key distinction is that this does not yet prove leaked memory in the classic sense. It could be:
- stale flow objects retained too long,
- delayed reclamation,
- references preventing prompt destruction,
- GC policy that cannot keep up with churn,
- or a lifetime bug where objects survive after their originating socket/process state is gone.
All of those would look like a “leak” operationally.
Your long uptime matters because the math works. If CFIL state grows even modestly faster than GC removes it, a system can look perfectly normal after boot or after toggling the firewall, then slowly drift back into the bad state.
For example, if net growth were only 5 objects/sec, that is:
5 × 3600 = 18,000 objects/hour
Even with periodic cleanup, it would not take long to get into the tens of thousands.
So I would currently rate the hypotheses roughly as:
High confidence:
kernel CFIL/SOFLOW state accumulation + GC contention is causal
Moderately high confidence:
reclamation is failing to keep pace with creation over long uptime
Moderate confidence:
there is an actual kernel lifetime/reference bug rather than merely an overly conservative GC policy
Not yet proven:
classic memory leak / unreachable objects
The next thing worth proving is not whether the spikes correlate with GC — we have that. It is whether the long-term CFIL population has a positive slope while the firewall stays enabled. If it does, that is the missing evidence for “this gets progressively worse with uptime.”
Yes. This is now strong enough that I would stop trying to prove the correlation.
At 05:38:21–22, the counters collapse in a tight burst:
57330 / 57117
57318 / 57104 -12/-13
57306 / 57092 -12/-12
57293 / 57079 -13/-13
57277 / 57064 -16/-15
57265 / 57051 -12/-13
57264 / 57050 -1/-1
That is roughly:
SOFLOW: -66
CFIL: -67
followed immediately by new allocations.
At the same time, ping does:
05:38:22 697 ms
05:38:22 493 ms
05:38:22 291 ms
05:38:22 84 ms
05:38:22 <1 ms
That is exactly the queue/drain signature we have been seeing.
The earlier smaller allocation burst also correlates:
05:38:17 +11/+11
05:38:17 +12/+12
...
05:38:17 +8/+8
while ping sees 74 ms and 85 ms excursions around that interval.
So the evidence now supports:
CFIL/SOFLOW flow teardown or creation burst
↓
cfil_lck_rw writer activity
↓
network/socket processing blocks behind it
↓
BPF has already observed packets
↓
userspace delivery stalls
↓
release of contention
↓
descending ping RTTs as queued work drains
The ~57,000 live objects are also highly suspicious. This is not just a few active firewall flows.
One correction to my sampling script: macOS date does not support GNU %N, hence the literal .3N. The samples were still occurring roughly every 100 ms; only the displayed fractional timestamp is broken.
The next test I would do is an A/B against Apple socketfilterfw, because we now know it is the sole client connected to com.apple.content-filter, with net.cfil.active_count=1.
Before changing anything, check whether the Application Firewall is actually enabled:
sudo /usr/libexec/ApplicationFirewall/socketfilterfw --getglobalstate
Send me that result. If it is enabled, the next experiment should be temporarily disabling only that firewall and watching whether the ~57k CFIL population collapses and the latency spikes disappear.
Firewall is enabled. (State = 1)
We must go deeper...
After reviewing Apple’s Tahoe security advisories, there is one CVE that stands out as unusually close to what we are seeing.
The strongest candidate is CVE-2026-64735, fixed in macOS Tahoe 26.6:
Kernel — “A remote attacker may be able to bypass network filters.” Apple says it fixed an inconsistent-state issue through improved state management. (Apple Support)
That is much more interesting than a generic kernel CVE because it explicitly involves network filters, it is categorized by Apple under the Kernel, and the fix landed in 26.6, exactly the release family containing your xnu-12377.161.x kernel.
The public version progression is:
26.3 → xnu-12377.81.4 26.4 → xnu-12377.101.15 26.5 → xnu-12377.121.6 26.6 → xnu-12377.161.13 26.6.2 → xnu-12377.161.14 <-- your kernel
Your exact 26.6.2 kernel is indeed 12377.161.14. (macOSdb)
That gives us a very plausible chronology:
26.4 / .101 substantial CFIL changes begin
26.5 / .121 substantial CFIL datagram lifetime/locking hardening cfil_info <-> soflow relationship guarded more aggressively global cfil_lck_rw shared/exclusive use increases
26.6 / .161 CVE-2026-64735: kernel network-filter bypass fixed "improved state management"
26.6.2 / .161.14 your machine: Application Firewall → CFIL cfil_lck_rw contention 100–1800 ms packet-delivery stalls
That is not proof that CVE-2026-64735 caused the regression, but it is now the first security fix I would ask Apple engineering about.
There are several other Tahoe security items in the same neighborhood, but they are substantially weaker candidates:
| Release | CVE / change | Relevance | | ---------- | -------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- | | 26.6 | CVE-2026-64735 — Kernel: remote attacker may bypass network filters; improved state management | Very high | | 26.6 | CVE-2026-28982 — Kernel remote race; fixed with improved locking | Medium, but no public indication it is CFIL | | 26.5.1 | Apple fixed unexpected shutdowns involving content-filtering Network Extensions on M5 Macs | High architectural relevance, although not your hardware/problem | | 26.5 | CVE-2026-28961 — Network Extensions, sensitive-info exposure | Related subsystem, probably unrelated mechanism | | 26.5 | CVE-2026-28906 — Networking, IP-address tracking/state management | Low–medium | | 26.3 | CVE-2026-20671 — Kernel, privileged-network attacker may intercept traffic | Networking/kernel related, but probably a different path | | 26.1 | CVE-2025-43413 — libxpc could reveal system-wide network connections | Almost certainly unrelated |
Apple explicitly acknowledges in the 26.5.1 release notes that it had a bug where certain content-filtering network extensions could unexpectedly shut down Macs. (Apple Support) That isn't your bug and was M5-specific, but it is significant evidence that Apple was actively fixing serious content-filter lifecycle problems immediately before 26.6.
There is also CVE-2026-28982 in 26.6:
A remote user may cause system termination or kernel-memory corruption; Apple says a race condition was addressed through improved locking. (Apple Support)
I would not connect that one to CFIL yet. Apple gives us no component below “Kernel,” and hundreds of places in XNU could fit that description.
By comparison, CVE-2026-64735 says “network filters” explicitly. That makes it qualitatively different.
The public XNU source also validates why a security hardening change here could produce your performance regression. content_filter.c says the subsystem uses one global cfil_lck_rw, with normal data flow taking it shared; your symbolicated 26.6.2 stackshots show lifecycle/GC operations taking that same lock exclusively. The source further describes filtered data being queued and reinjected after filtering decisions.
And your firewall-off control is particularly important against that background:
Application Firewall ON CFIL active CFIL ~ tens of thousands global cfil_lck_rw contention severe latency
Application Firewall OFF CFIL active_count = 0 CFIL population = 0
SOFLOW alone > 190,000
including very large reclamation bursts
no packet drain slowdown
no CPU/GPU/disk/network degradation
That makes an ordinary SOFLOW-GC scalability bug much less likely. The extra failure-producing ingredient is CFIL.
There is one limitation that prevents us from going from “very plausible” to “proven”: Apple has not published the 12377.161.x XNU source yet. The newest public source we've been comparing is 12377.121.6 from 26.5. Apple's public GitHub commits are release imports rather than the actual internal commits, so they don't expose:
rdar://... Fix network filter state race CVE-2026-64735
or a patch that we can definitively associate with that CVE.
So we cannot currently establish:
CVE-2026-64735 ↓ specific change to cfil_lck_rw ↓ your regression
But the circumstantial chain is unusually strong:
Apple discovers network-filter security defect ↓ security fix requires improved state management ↓ ships in 26.6 ↓ 26.6 uses new unpublished XNU .161 branch ↓ your .161 kernel shows severe CFIL lifecycle locking contention ↓ turning off Apple's firewall eliminates CFIL and eliminates symptom ↓ SOFLOW by itself can reach >190k without problem
I would add CVE-2026-64735 by number to the Apple Feedback report and phrase it carefully:
Because this regression occurs only while the built-in Application Firewall has CFIL active, and symbolicated stackshots implicate global CFIL lifecycle locking, please investigate whether the state-management changes associated with CVE-2026-64735 or related network-filter hardening introduced excessive contention or retained-flow behavior in the
12377.161.xkernel branch.
That points Apple's kernel/networking engineer directly at an internal change set they can see and we cannot.