The Architectural Bottleneck of Agentic Execution\n\nWhen scaling agentic artificial intelligence workloads on Kubernetes, the primary infrastructure constraint transitions rapidly from raw compute availability to memory density. Modern agent frameworks rely on iterative execution loops. An individual agent step typically involves fetching contextual history, constructing prompts, calling Large Language Model (LLM) inference endpoints, parsing response schemas, and invoking external tools or software APIs. While the LLM inference phase requires massive compute and GPU memory, the surrounding agent orchestration runtime—handling session state, conversational memory, workspace files, tool execution environments, and dynamic context windows—executes entirely on standard CPU nodes within volatile System DRAM.\n\nIn real-world deployment topologies, agentic loops spend a disproportionate fraction of their total lifespan waiting. An agent may pause while an external HTTP service responds, an asynchronous database query completes, an isolated code sandbox executes, or a human operator provides approval. During these extended wait states, the CPU utilization of the agent container drops to near zero. However, the process resident set size (RSS) remains static. The system memory allocated to retain the agent context, active runtime dependencies, and transient state remains pinned in physical DRAM. In high-density deployments attempting to host tens of thousands of concurrent agent sessions, reserving unshared physical DRAM for idle execution contexts causes severe resource inefficiency, inflates infrastructure spend, and triggers premature cluster node autoscaling.\n\nHistorically, cloud-native operational paradigms dictated that swap memory should be disabled entirely across Kubernetes clusters. This guidance stemmed from the coarse memory control mechanisms of cgroup v1. Under cgroup v1, swap usage was poorly isolated between co-located containers. A single memory-intensive container undergoing heavy swap operations could saturate host disk I/O, causing cascading latency degradation, uncontrollable page fault thrashing, and unpredicted noisy-neighbor impact across unrelated workloads on the same host node. Consequently, cluster operators enforced strict failSwapOn: true policies in the kubelet, treating system memory as an un-overcommittable resource.\n\nThe maturation of cgroup v2, combined with native Kubernetes swap support and enterprise Non-Volatile Memory Express (NVMe) solid-state storage, alters this trade-off fundamentally. By combining modern Linux page reclamation algorithms with low-latency PCI Express (PCIe) Gen4 and Gen5 NVMe drives, I have designed host configurations that treat high-speed NVMe storage as an extended memory tier. This pattern allows cluster nodes to safely overcommit physical DRAM by offloading cold, idle agent contexts to NVMe swap while keeping active execution contexts in high-bandwidth system memory.\n\n## Low-Level Linux Kernel Page Reclamation and Cgroup v2 Primitives\n\nTo safely overcommit memory without exposing latency-critical operations to performance collapse, you must understand how the Linux kernel manages memory reclamation under cgroup v2. The kernel divides physical memory into distinct pages, typically 4 KiB in size. These pages broadly fall into two categories: file-backed pages, which mirror data stored on a filesystem, and anonymous pages, which store process heaps, stack frames, and dynamically allocated execution memory.\n\nWhen system DRAM pressure rises, the kernel asynchronous page reclaim daemon, kswapd, scans memory zones to free pages. Under default system configurations with swap disabled, the kernel can only reclaim file-backed pages by evicting them from the page cache. Anonymous pages cannot be evicted unless a swap device is active. When system DRAM is depleted without swap, the kernel is forced into direct reclaim mode, pausing allocating processes synchronously while attempting to clean file pages. If direct reclaim fails to yield sufficient free pages, the kernel invokes the Out-Of-Memory (OOM) killer, terminating processes based on their oom_score.\n\nUnder cgroup v2, memory and swap tracking are unified within a single hierarchical control interface. Rather than treating swap as a global dumping ground, cgroup v2 provides fine-grained controls that allow you to dictate precisely how individual container cgroups interact with physical DRAM and swap media. Understanding these control knobs is critical for establishing operational boundaries:\n\nThe memory.min knob defines a hard physical DRAM protection boundary. The kernel page reclamation subsystem will never reclaim anonymous or file-backed memory from a cgroup whose total memory usage is below this threshold, unless the entire system faces an unrecoverable catastrophic memory shortage. For an agentic AI pod, setting memory.min guarantees that the base execution runtime, core language interpreter, and immediate control flow buffers remain pinned in physical DRAM, completely immune to swap eviction.\n\nThe memory.low knob establishes a soft protection threshold. If a cgroup operates below its memory.low mark, the kernel avoids reclaiming its memory pages unless memory pressure across all other system cgroups becomes acute. This creates a secondary buffer for active agent containers, minimizing unnecessary I/O overhead during transient operational pauses.\n\nThe memory.high knob acts as the primary throttling boundary and memory offloading trigger. When an agent container's memory consumption exceeds memory.high, the kernel subjects the processes inside that cgroup to aggressive background memory reclamation and I/O rate-limiting. The kernel identifies cold anonymous pages—such as inactive session histories and idle tool buffers—and writes them out to the configured NVMe swap device. The container continues executing without interruption, but its physical DRAM footprint drops, freeing memory for adjacent active pods.\n\nThe memory.max knob represents the absolute maximum physical DRAM a cgroup can occupy. If allocation demands exceed memory.max and background swapping cannot offload pages fast enough to satisfy the request, the process enters direct reclaim and face execution stalls. If memory cannot be freed, the OOM killer terminates the container processes, provided memory.swap.max is also reached or exhausted.\n\nThe memory.swap.max knob sets the upper bound on the total amount of swap space a specific cgroup is permitted to consume. This control is essential for multi-tenant isolation. By capping memory.swap.max, I prevent any single rogue agent container from consuming the entire host NVMe swap partition and starving adjacent workloads.\n\nKernel sysctl parameters further refine page scanner behavior across the entire host node. The vm.swappiness parameter, which ranges from 0 to 200 under modern kernels, dictates the relative cost weight assigned to swapping anonymous memory versus evicting filesystem cache. Setting vm.swappiness to a value between 60 and 80 instructs the kernel to reclaim cold anonymous pages aggressively when high-performance NVMe storage is present, preserving the filesystem page cache that holds shared application binaries, runtime libraries, and cached container layer files.\n\nComplementing swappiness is vm.vfs_cache_pressure, which controls the kernel tendency to reclaim directory and inode VFS structures. Keeping vm.vfs_cache_pressure around 50 ensures that file system lookup metadata remains cached in DRAM, reducing disk seek overhead for container runtimes. Finally, tuning vm.watermark_scale_factor to higher values forces kswapd to wake up earlier and work more gradually, preventing sudden memory allocation bursts from triggering synchronous direct reclaim stalls.\n\n## Storage Layer Mechanics: Host-Level NVMe Tuning and Block I/O\n\nThe viability of node-level swap for high-density agent workloads depends directly on the latency and throughput characteristics of the underlying block storage. Traditional spinning disks or slow network-attached block volumes introduce multi-millisecond page fault latencies that freeze execution threads, leading to severe application timeouts. In contrast, modern PCIe Gen4 and Gen5 enterprise NVMe drives deliver random read and write latencies in the microsecond range, combined with sequential throughput capabilities exceeding 6 to 12 GB/s per controller.\n\nWhen a swapped-out memory page is accessed by an active agent process, the CPU execution thread generates a hardware page fault exception. The operating system handles this major page fault by suspending the faulting thread, issuing a block read request to the swap device, transferring the 4 KiB page back into physical DRAM, updating the process translation lookaside buffer (TLB) and page table entries, and resuming thread execution. On enterprise NVMe storage, this complete major page fault sequence resolves in sub-millisecond timeframes, making page restoration virtually imperceptible to upper-level agent orchestration logic.\n\nTo achieve this performance profile, the host operating system must bypass unnecessary I/O queues and scheduler abstractions. Raw NVMe block devices should be partitioned specifically for swap operations rather than relying on swap files created inside existing filesystems. Mounting raw partitions eliminates filesystem journaling overhead, block allocation overhead, and metadata locking contention during high-frequency parallel write operations.\n\nThe block device configuration must align with modern hardware controller capabilities. You should configure the host kernel block I/O layer with the none I/O scheduler for high-performance NVMe drives. Modern NVMe hardware features multiple hardware submission and completion queues that map directly to physical CPU cores. Software I/O schedulers such as mq-deadline or bfq introduce unnecessary CPU lock contention and queuing latency; setting the scheduler to none allows block I/O requests to pass directly from the kernel block layer into hardware controller queues.\n\n## Declarative Configuration and Deployment Setup\n\nImplementing high-density NVMe swap requires coordinating host-level storage initialization, kernel sysctl tuning, systemd resource controls, and Kubernetes kubelet feature enablement. The automation script below demonstrates how to initialize a dedicated raw PCIe NVMe partition for high-priority kernel swap, apply low-latency memory reclamation parameters, and configure the kubelet daemon to operate under LimitedSwap mode with cgroup v2 integration.\n\n#!/usr/bin/env bash
set -euo pipefail
1. Identify raw enterprise NVMe device and partition for high-density swap
TARGET_DEV=


