Elektrine
Log in Register
Paige Chat Timeline Gallery Friends Email Drive DNS Private DNS Domains VPN Kairo Nerve
Remote

ℂ𝕣𝕪𝕠 :netbsd:

@Cryo@infosec.exchange
mastodon 4.8.0-alpha.3+glitch
  • Open on infosec.exchange

𝓦𝓲𝓵𝓵𝓲𝓪𝓶 𝓙. 𝓒𝓸𝓵𝓭𝔀𝓮𝓵𝓵 - ᴄʏʙᴇʀꜱᴇᴄᴜʀɪᴛʏ @ ʙʟᴜᴇᴄᴀᴛ - 𝙿𝚛𝚎𝚜𝚒𝚍𝚎𝚗𝚝 @ 𝙽𝚎𝚝𝙱𝚂𝙳 - ℭ𝔯𝔶𝔭𝔱𝔨𝔢𝔢𝔭𝔢𝔯 @ 𝔇𝔢𝔞𝔡𝔍𝔬𝔲𝔯𝔫𝔞𝔩 - ᵢ ₐₘ @ Wₐᵣₚₑd ⑤⓪③ ␀␠

0 Followers
0 Following
16 Posts
Joined October 29, 2022
:birdsite: EOL:
twitter.com/cryo
GPGP:
0xF97CC215
␀␠:
https://cryo.ws
cryo @Keybase:
https://cryo.keybase.pub/infosec-exchange.html
cryo @Github:
https://raw.githubusercontent.com/cryo/cryo.github.io/main/index.html
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to
I'm pulling all data content (backups, photos, etc) off of iClod.
1
0
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 4mo ago
Replying to
@NanoRaptor@bitbang.social Fatality! right in the age hole.
3
0
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 7mo ago
Replying to
@lukhash@mastodon.social this is kick ass.
1
0
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to

END OF TRANSMISSION. Thanks for riding along on this adventure. Remember to smash that subscribe and like button.. oh wait. Sorry. I'm on twitch.tv/cryosama sometimes fighting death on World of Warcraft Classic Hardcore, or coding (currently TapNet3D for various platforms).

This is what happens when computers/AI overlords piss me off, I fall down a hyperfocused rabbit-hole. I'll collect more information tomorrow and any that are submitted to me by other mac users who want to test and push your system to repro this. Then I'll file a bug report:


Title

macOS 26.6.2: Application Firewall causes CFIL/SOFLOW flow-state accumulation and periodic kernel network stalls

Product / Build

  • macOS Tahoe 26.6.2
  • Build 25G83
  • Darwin 25.6.0
  • XNU 12377.161.14~5
  • Apple Silicon: M4 Pro Mac mini

Summary

With the macOS Application Firewall enabled, net.cfil.sock_attached_count and net.soflow.count rapidly accumulate from near zero to tens of thousands of live flow-state objects. Under sustained normal application/network activity, the population rises to roughly 75,000–80,000 entries and then oscillates around that high-water range.

At these elevated flow counts, the system develops recurring latency spikes affecting otherwise local low-latency traffic. ICMP to the directly connected default gateway normally completes in under 1 ms, but periodically stalls for 100–900+ ms.

Simultaneous packet capture shows the packets are present on the interface immediately while userspace delivery is delayed. This localizes the delay to the kernel networking path after BPF observation and before userspace receives the packet.

Kernel stackshots from the affected system show contention involving SOFLOW garbage collection, dlil_input_en0, socket receive paths, and the global CFIL read/write lock.

Observed behavior

After boot or after resetting firewall state, the system initially behaves normally.

With Application Firewall enabled:

07:14:12  SOFLOW=187     CFIL=152
09:51:10  SOFLOW=64840   CFIL=64591
18:44:57  SOFLOW=79508   CFIL=79270

The counters are not cumulative event counters. XNU source shows:

cfil_sock_attached_count++;
...
cfil_sock_attached_count--;

and:

soflow_attached_count++;
...
soflow_attached_count--;

so these values represent current retained/attached flow-state population.

The population eventually reaches a dynamic equilibrium near 75k–80k rather than returning toward the number of active userland sockets.

The machine typically has only hundreds of live Internet sockets at the same time.

Latency symptom

Under the affected state, ping to the directly connected gateway shows stalls such as:

0.5 ms
0.6 ms
444 ms
0.7 ms
...
713 ms
215 ms
0.6 ms
...
920 ms
416 ms
0.5 ms

Simultaneous tcpdump shows request and reply packets on en0 with sub-millisecond wire timing while ping(8) reports hundreds of milliseconds.

The observed sequence is therefore:

packet reaches en0
BPF/tcpdump observes packet
kernel delivery is delayed
userspace ping receives packet later

Kernel evidence

Symbolicated stackshots using the matching 26.6.2 KDK show the affected paths entering CFIL locking.

Representative paths include:

sosend
  -> cfil_sock_udp_handle_data
  -> cfil_sock_udp_get_info
  -> cfil_rw_lock_shared
  -> IORWLockRead

and input-side paths involving:

dlil_input_en0
  -> proto_input
  -> sbappendaddr
  -> cfil_sock_udp_handle_data
  -> cfil_info_alloc
  -> cfil_rw_lock_exclusive

SOFLOW GC is also observed in:

soflow_gc_expire
  -> cfil_dgram_gc_perform
  -> cfil_sock_udp_unlink_flow
  -> cfil_rw_lock_exclusive

and:

cfil_info_free
  -> cfil_rw_lock_exclusive

Blocked readers are frequently owned by either:

dlil_input_en0
SOFLOW_GC

XNU source documents that the CFIL subsystem is protected by the global cfil_lck_rw.

GC behavior

Public XNU source defines:

#define SOFLOW_GC_IDLE_TO            30
#define SOFLOW_GC_MAX_COUNT          100
#define SOFLOW_GC_RUN_INTERVAL_NSEC  (10 * NSEC_PER_SEC)

Observed runtime behavior matches periodic reclamation, but reclamation does not reduce the population back to a small working set. Instead, flow-state accumulates rapidly and later churns around a high equilibrium.

Application Firewall correlation

socketfilterfw is connected to:

com.apple.content-filter

and:

net.cfil.active_count = 1

while Application Firewall is enabled.

When the firewall is disabled:

net.cfil.active_count = 0

new CFIL population growth stops, while existing CFIL state is gradually reclaimed.

When the firewall is enabled again, new CFIL attachments resume.

This A/B behavior is reproducible.

Expected behavior

CFIL/SOFLOW state associated with expired or closed flows should be reclaimed promptly enough that the retained population remains proportional to active socket/flow usage.

Periodic garbage collection should not cause multi-hundred-millisecond stalls in unrelated packet delivery.

Actual behavior

CFIL/SOFLOW state grows into the tens of thousands despite only hundreds of live userland sockets.

At high population, recurring SOFLOW/CFIL cleanup produces contention on the global CFIL lock and coincides with severe local packet-delivery latency.

Impact

This is most visible in latency-sensitive workloads:

  • video conferencing
  • real-time audio
  • streaming
  • games
  • SSH / remote shells
  • interactive network applications

Bulk transfers and background tasks are less visibly affected because buffering masks short stalls.

Reproduction

  1. Boot macOS 26.6.2.
  2. Ensure Application Firewall is enabled.
  3. Run normal network-intensive applications for several hours.
  4. Periodically record:
sysctl -n net.soflow.count
sysctl -n net.cfil.sock_attached_count
sysctl -n net.cfil.active_count
  1. Observe CFIL/SOFLOW population increasing into tens of thousands.
  2. Run a high-frequency ping to the local gateway.
  3. Observe periodic 100–900+ ms latency excursions.
  4. Capture simultaneously with tcpdump.
  5. Observe that packets appear on en0 promptly while userspace ping delivery is delayed.
  6. Disable Application Firewall and observe that new CFIL attachment growth stops.

Attachments I would include

  • full CFIL/SOFLOW counter log
  • graph showing growth from ~0 to ~80k
  • ping log showing latency spikes
  • simultaneous tcpdump capture
  • bad-state spindump
  • healthy-state spindump
  • symbolicated kernel stack output from the matching KDK
  • systemextensionsctl list
  • netstat -anv -f system
  • lsof showing socketfilterfw connected to com.apple.content-filter
  • exact sysctl snapshots before/after firewall disable

Suggested engineering summary

Suspected CFIL/SOFLOW flow-lifetime or reclamation defect associated with Application Firewall. Flow-state population grows to ~80k, substantially exceeding live userland socket count. SOFLOW GC and CFIL teardown take the global cfil_lck_rw exclusively, while packet delivery paths require the same lock shared. At elevated flow-state population, periodic GC/teardown correlates with 100–900+ ms local packet-delivery stalls. Disabling Application Firewall stops new CFIL attachment growth and allows the existing population to drain.

0
0
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to
macOS Tahoe users play-at-home game: Terminal 1: sudo ping -i 0.5 YOUR_ROUTER_IP | while IFS= read -r line; do printf '%s %s\n' "$(date '+%Y-%m-%d %H:%M:%S')" "$line"; done | tee ping-test.log You can do this from multiple Macs on your network, and you'll see that on those not affected, the ping times will be normally low. As the machine becomes affected it will do things like: 2026-09-12 19:05:24 64 bytes from 172.27.1.1: icmp_seq=17330 ttl=64 time=0.553 ms 2026-09-12 19:05:25 64 bytes from 172.27.1.1: icmp_seq=17331 ttl=64 time=0.786 ms 2026-09-12 19:05:25 64 bytes from 172.27.1.1: icmp_seq=17332 ttl=64 time=0.974 ms 2026-09-12 19:05:26 64 bytes from 172.27.1.1: icmp_seq=17333 ttl=64 time=0.727 ms 2026-09-12 19:05:26 64 bytes from 172.27.1.1: icmp_seq=17334 ttl=64 time=0.523 ms 2026-09-12 19:05:27 64 bytes from 172.27.1.1: icmp_seq=17335 ttl=64 time=461.501 ms 2026-09-12 19:05:27 64 bytes from 172.27.1.1: icmp_seq=17336 ttl=64 time=0.460 ms 2026-09-12 19:05:28 64 bytes from 172.27.1.1: icmp_seq=17337 ttl=64 time=0.562 ms 2026-09-12 19:05:28 64 bytes from 172.27.1.1: icmp_seq=17338 ttl=64 time=0.984 ms 2026-09-12 19:05:29 64 bytes from 172.27.1.1: icmp_seq=17339 ttl=64 time=43.713 ms 2026-09-12 19:05:29 64 bytes from 172.27.1.1: icmp_seq=17340 ttl=64 time=0.454 ms 2026-09-12 19:05:30 64 bytes from 172.27.1.1: icmp_seq=17341 ttl=64 time=0.731 ms 2026-09-12 19:05:30 64 bytes from 172.27.1.1: icmp_seq=17342 ttl=64 time=8.705 ms 2026-09-12 19:05:31 64 bytes from 172.27.1.1: icmp_seq=17343 ttl=64 time=0.693 ms 2026-09-12 19:05:31 64 bytes from 172.27.1.1: icmp_seq=17344 ttl=64 time=0.814 ms 2026-09-12 19:05:32 64 bytes from 172.27.1.1: icmp_seq=17345 ttl=64 time=320.964 ms 2026-09-12 19:05:32 64 bytes from 172.27.1.1: icmp_seq=17346 ttl=64 time=0.738 ms 2026-09-12 19:05:33 64 bytes from 172.27.1.1: icmp_seq=17347 ttl=64 time=0.637 ms 2026-09-12 19:05:33 64 bytes from 172.27.1.1: icmp_seq=17348 ttl=64 time=0.549 ms 2026-09-12 19:05:35 64 bytes from 172.27.1.1: icmp_seq=17349 ttl=64 time=1372.799 ms 2026-09-12 19:05:35 64 bytes from 172.27.1.1: icmp_seq=17350 ttl=64 time=872.529 ms 2026-09-12 19:05:35 64 bytes from 172.27.1.1: icmp_seq=17351 ttl=64 time=454.173 ms 2026-09-12 19:05:35 64 bytes from 172.27.1.1: icmp_seq=17352 ttl=64 time=0.391 ms Terminal 2: Graphable csv over time (not the system wasn't doing streaming or game at this time because I was asleep, lol): 2026-09-12 19:00:59,on,79766,79509,1,+131,+128 2026-09-12 19:01:29,on,79810,79553,1,+44,+44 2026-09-12 19:01:59,on,79940,79684,1,+130,+131 2026-09-12 19:02:29,on,79938,79685,1,-2,+1 2026-09-12 19:02:59,on,80181,79920,1,+243,+235 2026-09-12 19:03:30,on,80263,80009,1,+82,+89 2026-09-12 19:04:00,on,80306,80052,1,+43,+43 2026-09-12 19:04:30,on,80421,80167,1,+115,+115 2026-09-12 19:05:00,on,80464,80211,1,+43,+44 2026-09-12 19:05:30,on,80631,80377,1,+167,+166 2026-09-12 19:06:00,on,80683,80431,1,+52,+54 2026-09-12 19:06:30,on,80783,80530,1,+100,+99 2026-09-12 19:07:00,on,80822,80562,1,+39,+32 2026-09-12 19:07:30,on,81004,80749,1,+182,+187 2026-09-12 19:08:00,on,81102,80846,1,+98,+97 2026-09-12 19:08:30,on,81049,80795,1,-53,-51 2026-09-12 19:09:00,on,81007,80750,1,-42,-45 ---- #!/bin/sh INTERVAL=30 OUT="${1:-/tmp/cfil-backlog-$(date '+%Y%m%d-%H%M%S').log}" FW="/usr/libexec/ApplicationFirewall/socketfilterfw" prev_soflow="" prev_cfil="" printf "# Started: %s\n" "$(date)" | tee -a "$OUT" printf "# interval=%ss\n" "$INTERVAL" | tee -a "$OUT" printf "# timestamp,firewall,soflow,cfil,active,delta_soflow,delta_cfil\n" | tee -a "$OUT" while :; do ts=$(date '+%Y-%m-%d %H:%M:%S') fw=$( sudo "$FW" --getglobalstate 2>/dev/null | awk ' /State = 1/ { print "on"; found=1 } /State = 0/ { print "off"; found=1 } END { if (!found) print "unknown" } ' ) vals=$(sysctl -n \ net.soflow.count \ net.cfil.sock_attached_count \ net.cfil.active_count 2>/dev/null) soflow=$(printf '%s\n' "$vals" | sed -n '1p') cfil=$(printf '%s\n' "$vals" | sed -n '2p') active=$(printf '%s\n' "$vals" | sed -n '3p') if [ -n "$prev_soflow" ]; then dso=$((soflow - prev_soflow)) dcf=$((cfil - prev_cfil)) else dso=0 dcf=0 fi printf "%s,%s,%s,%s,%s,%+d,%+d\n" \ "$ts" "$fw" "$soflow" "$cfil" "$active" "$dso" "$dcf" | tee -a "$OUT" prev_soflow=$soflow prev_cfil=$cfil sleep "$INTERVAL" done
0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to

UPDATED:


The longer run changes the interpretation somewhat.

Updated full-run graph

From 07:14 to 18:44, SOFLOW rose from 187 to 79,508 and CFIL from 152 to 79,270. The early phase is unmistakable: both populations climb almost linearly and in lockstep into the mid-60k range. After roughly 10:00–12:00, though, the growth rate drops sharply, and by the afternoon the system appears to be oscillating around a broad ~75k–80k equilibrium band rather than continuing upward without bound. The final samples are still around 79k.

That weakens the “unbounded leak” interpretation. The better description now is pathologically high retained flow-state population with slow/limited reclamation that eventually reaches a dynamic equilibrium. GC is definitely working; it just permits tens of thousands of CFIL/SOFLOW objects to remain resident under this workload.

The important part is that this does not make the bug less interesting. A steady-state population of ~80,000 live/retained flow objects is still vastly larger than your current userland socket population, and our KDK stackshots already show SOFLOW GC and CFIL teardown contending on cfil_lck_rw. So the likely failure mode becomes:

high flow creation rate
        ↓
CFIL/SOFLOW population rapidly rises
        ↓
GC eventually catches creation rate
        ↓
steady-state population remains ~75k–80k
        ↓
continuous GC / detach / free activity over a huge population
        ↓
global CFIL writer contention
        ↓
periodic userspace network latency

The near-perfect tracking between the two counters is also stronger evidence that this is one lifecycle: SOFLOW entries and their CFIL feature contexts are being created and destroyed together, not two unrelated counters.

So I would revise the diagnosis from “probable kernel leak that grows forever” to “probable kernel flow-lifetime/reclamation pathology that builds an abnormally large working set and then churns around a high-water equilibrium.”

That actually fits the observed symptom better: the machine does not have to run out of resources. It only has to accumulate enough CFIL/SOFLOW state that recurring cleanup of that state makes the global lock expensive enough for interactive traffic to notice.

The CFIL/SOFLOW problem directly explains the network stalls, and it can explain networked audio/video very naturally, but it does not by itself prove a direct CoreAudio or GPU stall.

For network traffic, the mechanism is concrete: packet delivery reaches the CFIL path, reader-side processing needs cfil_lck_rw, and GC/teardown takes that same global lock exclusively. While the writer holds it, packet delivery can be delayed even though the packet has already reached the machine. That is exactly what your simultaneous tcpdump/ping test demonstrated.

Audio and video applications are much more sensitive to those delays because they run against small real-time buffers. Zoom, Discord, streaming, game voice, OBS network output, etc. expect packets every few to tens of milliseconds. A 100–900 ms kernel delivery stall is therefore not just “slower networking”; it causes jitter buffers to underrun, frames to arrive too late, audio to crackle/drop, video to freeze, and interactive state to jump when the backlog releases. A bulk transfer barely notices because TCP and application buffering absorb the pause.

The important distinction is local audio/video. If a purely local audio file crackles, or local GPU rendering freezes with no network dependency, CFIL is not a sufficient explanation by itself. In that case I would suspect a broader consequence of the same pathology—e.g. kernel threads spending substantial time in GC/locking, scheduling latency, or a related network/configuration watchdog event—but we have not yet demonstrated that causal path.

So the current model is:

CFIL/SOFLOW buildup
        ↓
GC / detach / free
        ↓
global cfil_lck_rw writer contention
        ↓
network packet delivery delayed
        ↓
real-time networked audio/video exhaust small buffers
        ↓
crackle / freeze / lag / dropped frames

For networked audio/video, that is a direct and well-supported explanation. For local audio/video glitches, we still need a second bridge from the CFIL event to scheduler/audio/GPU latency rather than assuming one.


  • The video of like WoW can studder if the map is doing a load in of assets, etc. The audio for the stream is the stame way. So much of the world now is all based on latency, packet speed, and packet size... and streaming WINDOW... but since the packet pile-up is in the stack, it's not the network's fault (for once).
0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to
backflow versus drain over a little time. Interactive applications notice this first because they depend on low and consistent latency: voice/video calls, remote shells, streaming, games, and other real-time workloads immediately expose a 100–700 ms stall as frozen audio, lag, jitter, delayed input, or dropped frames. Less interactive workloads such as downloads, background sync, mail retrieval, backups, and ordinary web requests are much more tolerant because TCP buffering, retransmission, application queues, and asynchronous processing hide short pauses. As the CFIL/SOFLOW backlog grows over long uptime, the same periodic kernel stalls can become progressively longer or more frequent, eventually becoming noticeable even to those less latency-sensitive applications.
0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to

and $APPL didn't hire me... twice.

checking the backlog periodically...


Yes. This is now strong evidence of a persistent positive backlog slope with the firewall enabled.

Over 23m34s:

SOFLOW: 65,956 → 70,118   = +4,162
CFIL:   59,021 → 63,866   = +4,845

That works out to roughly:

SOFLOW: +2.94 objects/sec  ≈ +10,600/hour
CFIL:   +3.43 objects/sec  ≈ +12,300/hour

The important part is not the exact extrapolation; it is that GC is clearly not keeping the population flat over this interval. There are occasional drops, but the long-term slope remains strongly positive.

You can see the characteristic pattern:

growth
growth
growth
small GC drop
growth
growth
larger burst
growth
...

and by the end the CFIL population is almost 5,000 higher than where it started.

That materially strengthens the “reclamation imbalance / flow-state leak” hypothesis. I would now describe it as:

With ALF enabled, CFIL/SOFLOW runtime state accumulates substantially faster than periodic GC reclaims it. Over long uptime, the retained population can grow into the tens of thousands, making periodic cleanup increasingly expensive and producing cfil_lck_rw contention that delays network delivery.

I would still avoid calling it a proven memory leak, because we have not demonstrated unreachable allocations. But operationally, this is behaving like a kernel flow-state leak.

One additional detail is interesting:

SOFLOW - CFIL gap:
start: 6,935
end:   6,252

So CFIL is actually growing slightly faster than SOFLOW during this window. That is worth watching because it argues against a simple fixed proportional relationship between the two populations.

Keep this monitor running. At this point, the longer dataset is valuable: a few hours will tell us whether the slope remains positive, levels off at some ceiling, or accelerates as the population rises.

0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to

Yes — I would rate a kernel-side CFIL/SOFLOW state-retention bug as a strong working hypothesis now, but I would call it a flow-state leak / reclamation imbalance rather than a proven memory leak.

What the evidence supports is:

  • CFIL/SOFLOW population reaches roughly 50–60k objects.
  • Live userland socket counts are only in the hundreds, so those kernel objects are clearly not just a one-to-one reflection of currently open sockets.
  • New CFIL objects accumulate while the firewall is active.
  • GC removes them in small periodic batches, roughly 70 every 10 seconds in your captures.
  • Those GC batches line up with the latency stalls.
  • Disabling the firewall stops new CFIL attachment, and the backlog then drains gradually.
  • Re-enabling after an off period leaves you in a much healthier state because the backlog has been reduced.
  • Long uptime gives exactly the conditions for a net-positive accumulation rate to become visible only after hours or days.

So the failure mode I would currently model is:

average CFIL creation rate
        >
average CFIL reclamation rate
        ↓
retained CFIL/SOFLOW population rises over uptime
        ↓
GC has more work to do each pass
        ↓
more frequent/longer cfil_lck_rw writer holds
        ↓
UDP/input paths stall behind the global lock
        ↓
eventually you see 100–700 ms bursts and queue-drain RTT patterns

The key distinction is that this does not yet prove leaked memory in the classic sense. It could be:

  • stale flow objects retained too long,
  • delayed reclamation,
  • references preventing prompt destruction,
  • GC policy that cannot keep up with churn,
  • or a lifetime bug where objects survive after their originating socket/process state is gone.

All of those would look like a “leak” operationally.

Your long uptime matters because the math works. If CFIL state grows even modestly faster than GC removes it, a system can look perfectly normal after boot or after toggling the firewall, then slowly drift back into the bad state.

For example, if net growth were only 5 objects/sec, that is:

5 × 3600 = 18,000 objects/hour

Even with periodic cleanup, it would not take long to get into the tens of thousands.

So I would currently rate the hypotheses roughly as:

High confidence:
kernel CFIL/SOFLOW state accumulation + GC contention is causal

Moderately high confidence:
reclamation is failing to keep pace with creation over long uptime

Moderate confidence:
there is an actual kernel lifetime/reference bug rather than merely an overly conservative GC policy

Not yet proven:
classic memory leak / unreachable objects

The next thing worth proving is not whether the spikes correlate with GC — we have that. It is whether the long-term CFIL population has a positive slope while the firewall stays enabled. If it does, that is the missing evidence for “this gets progressively worse with uptime.”

0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to
WHY WOULD TURNING THE APPLICATION FIREWALL OFF AND BACK ON MAKE IT APPEAR TO WORK (nintendo blowing on cartridge mode activated) ----- That explains why toggling it off for a while and back on made the machine better: you partially drained a pathological accumulated CFIL working set. The other likely reason it stays better after re-enabling is that existing flows probably aren't all retroactively given CFIL contexts. Re-enabling starts filtering newly encountered/newly evaluated flows, rather than instantly reconstructing the ~58,000 old CFIL associations. So you effectively get a partial reset without rebooting. This also explains the earlier “reboot/reinstall makes it good for a while, then it comes back” behavior: a reboot starts with essentially no accumulated SOFLOW/CFIL flow state; the population then builds up over time until the GC + global cfil_lck_rw contention becomes noticeable again. I would now treat CFIL state accumulation plus periodic GC of that accumulated state as the leading causal mechanism, rather than simply “the firewall occasionally pauses networking.” The quantitative match between the 10-second GC rate and the ~9-minute off-period reduction is particularly strong.
0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago
Replying to

Yes. This is now strong enough that I would stop trying to prove the correlation.

At 05:38:21–22, the counters collapse in a tight burst:

57330 / 57117
57318 / 57104   -12/-13
57306 / 57092   -12/-12
57293 / 57079   -13/-13
57277 / 57064   -16/-15
57265 / 57051   -12/-13
57264 / 57050    -1/-1

That is roughly:

SOFLOW: -66
CFIL:   -67

followed immediately by new allocations.

At the same time, ping does:

05:38:22  697 ms
05:38:22  493 ms
05:38:22  291 ms
05:38:22   84 ms
05:38:22  <1 ms

That is exactly the queue/drain signature we have been seeing.

The earlier smaller allocation burst also correlates:

05:38:17   +11/+11
05:38:17   +12/+12
...
05:38:17    +8/+8

while ping sees 74 ms and 85 ms excursions around that interval.

So the evidence now supports:

CFIL/SOFLOW flow teardown or creation burst
        ↓
cfil_lck_rw writer activity
        ↓
network/socket processing blocks behind it
        ↓
BPF has already observed packets
        ↓
userspace delivery stalls
        ↓
release of contention
        ↓
descending ping RTTs as queued work drains

The ~57,000 live objects are also highly suspicious. This is not just a few active firewall flows.

One correction to my sampling script: macOS date does not support GNU %N, hence the literal .3N. The samples were still occurring roughly every 100 ms; only the displayed fractional timestamp is broken.

The next test I would do is an A/B against Apple socketfilterfw, because we now know it is the sole client connected to com.apple.content-filter, with net.cfil.active_count=1.

Before changing anything, check whether the Application Firewall is actually enabled:

sudo /usr/libexec/ApplicationFirewall/socketfilterfw --getglobalstate

Send me that result. If it is enabled, the next experiment should be temporarily disabling only that firewall and watching whether the ~57k CFIL population collapses and the latency spikes disappear.


Firewall is enabled. (State = 1)

0
1
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3mo ago
Live on Twitch: Working on my 3D game called TapNet3D with the Playdate port. Check it out, click follow, watch me yell at LLMs as they fight me and each other to produce working code that I accept... https://twitch.tv/cryosama
twitch.tv
0
0
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3mo ago
I'm live on https://twitch.tv/cryosama doing some vibe yelling at LLMs working on TapNet3D for the Playdate handheld. WORKING ON GETTING A DEMO RELEASED! Come follow me (show interest - please clap...)
twitch.tv
0
0
0
0
Open post
ℂ𝕣𝕪𝕠 :netbsd: @Cryo@infosec.exchange
· 3w ago

We must go deeper...


After reviewing Apple’s Tahoe security advisories, there is one CVE that stands out as unusually close to what we are seeing.

The strongest candidate is CVE-2026-64735, fixed in macOS Tahoe 26.6:

Kernel — “A remote attacker may be able to bypass network filters.” Apple says it fixed an inconsistent-state issue through improved state management. (Apple Support)

That is much more interesting than a generic kernel CVE because it explicitly involves network filters, it is categorized by Apple under the Kernel, and the fix landed in 26.6, exactly the release family containing your xnu-12377.161.x kernel.

The public version progression is:

26.3 → xnu-12377.81.4 26.4 → xnu-12377.101.15 26.5 → xnu-12377.121.6 26.6 → xnu-12377.161.13 26.6.2 → xnu-12377.161.14 <-- your kernel

Your exact 26.6.2 kernel is indeed 12377.161.14. (macOSdb)

That gives us a very plausible chronology:

26.4 / .101 substantial CFIL changes begin

26.5 / .121 substantial CFIL datagram lifetime/locking hardening cfil_info <-> soflow relationship guarded more aggressively global cfil_lck_rw shared/exclusive use increases

26.6 / .161 CVE-2026-64735: kernel network-filter bypass fixed "improved state management"

26.6.2 / .161.14 your machine: Application Firewall → CFIL cfil_lck_rw contention 100–1800 ms packet-delivery stalls

That is not proof that CVE-2026-64735 caused the regression, but it is now the first security fix I would ask Apple engineering about.

There are several other Tahoe security items in the same neighborhood, but they are substantially weaker candidates:

| Release | CVE / change | Relevance | | ---------- | -------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- | | 26.6 | CVE-2026-64735 — Kernel: remote attacker may bypass network filters; improved state management | Very high | | 26.6 | CVE-2026-28982 — Kernel remote race; fixed with improved locking | Medium, but no public indication it is CFIL | | 26.5.1 | Apple fixed unexpected shutdowns involving content-filtering Network Extensions on M5 Macs | High architectural relevance, although not your hardware/problem | | 26.5 | CVE-2026-28961 — Network Extensions, sensitive-info exposure | Related subsystem, probably unrelated mechanism | | 26.5 | CVE-2026-28906 — Networking, IP-address tracking/state management | Low–medium | | 26.3 | CVE-2026-20671 — Kernel, privileged-network attacker may intercept traffic | Networking/kernel related, but probably a different path | | 26.1 | CVE-2025-43413 — libxpc could reveal system-wide network connections | Almost certainly unrelated |

Apple explicitly acknowledges in the 26.5.1 release notes that it had a bug where certain content-filtering network extensions could unexpectedly shut down Macs. (Apple Support) That isn't your bug and was M5-specific, but it is significant evidence that Apple was actively fixing serious content-filter lifecycle problems immediately before 26.6.

There is also CVE-2026-28982 in 26.6:

A remote user may cause system termination or kernel-memory corruption; Apple says a race condition was addressed through improved locking. (Apple Support)

I would not connect that one to CFIL yet. Apple gives us no component below “Kernel,” and hundreds of places in XNU could fit that description.

By comparison, CVE-2026-64735 says “network filters” explicitly. That makes it qualitatively different.

The public XNU source also validates why a security hardening change here could produce your performance regression. content_filter.c says the subsystem uses one global cfil_lck_rw, with normal data flow taking it shared; your symbolicated 26.6.2 stackshots show lifecycle/GC operations taking that same lock exclusively. The source further describes filtered data being queued and reinjected after filtering decisions.

And your firewall-off control is particularly important against that background:

Application Firewall ON CFIL active CFIL ~ tens of thousands global cfil_lck_rw contention severe latency

Application Firewall OFF CFIL active_count = 0 CFIL population = 0

SOFLOW alone > 190,000
including very large reclamation bursts

no packet drain slowdown
no CPU/GPU/disk/network degradation

That makes an ordinary SOFLOW-GC scalability bug much less likely. The extra failure-producing ingredient is CFIL.

There is one limitation that prevents us from going from “very plausible” to “proven”: Apple has not published the 12377.161.x XNU source yet. The newest public source we've been comparing is 12377.121.6 from 26.5. Apple's public GitHub commits are release imports rather than the actual internal commits, so they don't expose:

rdar://... Fix network filter state race CVE-2026-64735

or a patch that we can definitively associate with that CVE.

So we cannot currently establish:

CVE-2026-64735 ↓ specific change to cfil_lck_rw ↓ your regression

But the circumstantial chain is unusually strong:

Apple discovers network-filter security defect ↓ security fix requires improved state management ↓ ships in 26.6 ↓ 26.6 uses new unpublished XNU .161 branch ↓ your .161 kernel shows severe CFIL lifecycle locking contention ↓ turning off Apple's firewall eliminates CFIL and eliminates symptom ↓ SOFLOW by itself can reach >190k without problem

I would add CVE-2026-64735 by number to the Apple Feedback report and phrase it carefully:

Because this regression occurs only while the built-in Application Firewall has CFIL active, and symbolicated stackshots implicate global CFIL lifecycle locking, please investigate whether the state-management changes associated with CVE-2026-64735 or related network-filter hardening introduced excessive contention or retained-flow behavior in the 12377.161.x kernel branch.

That points Apple's kernel/networking engineer directly at an internal change set they can see and we cannot.

Apple Support

Sobre o conteúdo de segurança do macOS Tahoe 26.6 - Suporte da Apple (BR)

Este documento descreve o conteúdo de segurança do macOS Tahoe 26.6.

0
0
0
0
Back
313k7r1n3
Elektrine

Tor hidden service

elekhj7afj4qnrr4yd3bkzslsyo5jgfxw3orgjkhlcxifueodybyiiad.onion

I2P eepsite

j6b6cyk6gjmepjih7jjadxgxvvf3lzzujljuu2v4biemzpg3naya.b32.i2p

Platform

  • Email
  • Chat
  • Timeline
  • VPN
  • DNS

Company

  • About
  • Contact
  • FAQ
  • Lite (no JS)

Legal

  • Terms of Service
  • Privacy Policy
  • Transparency Report
  • Report Abuse
  • Warrant Canary
  • VPN Policy

Support

  • support@elektrine.com
  • Report Security Issue
Mail client setup IMAP mail.elektrine.com:993 POP3 mail.elektrine.com:995 SMTP mail.elektrine.com:465
© 2026 Elektrine. All rights reserved. Server: 19:09:47 UTC