Keone’s blog
← All writing

Async I/O and io_uring

performance

EVM execution is dominated by state access. Every SLOAD, balance lookup, and EXTCODESIZE call ultimately reads from the state trie, which lives on disk. An NVMe SSD serves a random read in roughly 100 microseconds -- fast by storage standards, but in that time a modern CPU could execute hundreds of thousands of instructions. If the CPU blocks on every read, it spends most of its time waiting.

Monad's I/O layer is built on Linux's io_uring interface, which allows the execution engine to submit disk reads and continue doing useful work while the kernel completes them. Combined with lightweight fiber-based concurrency, this keeps the CPU productive even when many concurrent state reads are in flight.

Why synchronous I/O is the bottleneck

In a traditional Ethereum client, a state read works like this:

  1. Execution reaches an opcode that reads state (e.g. SLOAD).

  2. The client issues a synchronous read to the database.

  3. The thread blocks until the data arrives from disk.

  4. Execution resumes with the retrieved value.

A single NVMe SSD random read takes about 100 microseconds. But a single state lookup (e.g. SLOAD) doesn't translate to a single disk read; the value lives in a trie, and traversing the trie requires reading several intermediate nodes. Each node that isn't cached is a separate disk read, so a single state lookup can require multiple 100-microsecond reads in sequence. An average transaction performs several state lookups, so the total I/O wait time per transaction can easily reach milliseconds.

In practice, Ethereum clients mitigate this by caching aggressively: keeping as much of the trie in RAM as possible so that most node lookups are memory reads rather than disk reads. Some clients use memory-mapped files, where the operating system manages a page cache and loads pages on demand. This looks like memory access in the code, but when a page isn't cached, the thread takes a page fault: the kernel suspends the thread, issues the I/O, and resumes the thread when the data arrives. The thread is blocked just as effectively as with an explicit synchronous read -- the blocking is simply hidden behind the memory mapping abstraction.

Relying on large in-memory caches works when state is small enough to fit, but it doesn't scale. Blockchain state grows continuously, and RAM is roughly 30x more expensive per byte than SSD storage. Requiring validators to provision ever-growing amounts of RAM raises hardware costs and limits decentralization. As state outgrows the cache, more trie node reads hit disk, and the synchronous I/O problem resurfaces.

The fundamental problem is that synchronous I/O (whether explicit or via mmap) ties one thread to one I/O operation. To have N concurrent reads, you need N threads, each consuming kernel resources and stack memory. The scalable solution is to keep state on disk and make disk I/O fast enough that it isn't a bottleneck.

io_uring: asynchronous I/O done right

io_uring is a Linux kernel interface (introduced in kernel 5.1) that provides truly asynchronous I/O with minimal overhead. It works through two ring buffers shared between user space and the kernel:

  • Submission queue (SQ): The application writes I/O requests here.

  • Completion queue (CQ): The kernel writes completion notifications here.

The application submits a read request by writing an entry to the SQ. It doesn't need to make a system call for each submission -- entries can be batched, and an optional kernel polling thread can drain the SQ without the application explicitly calling into the kernel at all.

When the I/O completes, the kernel writes a completion entry to the CQ. The application polls the CQ to pick up completed operations by reading from the shared ring buffer.

Both the SQ and CQ live in shared memory mapped into both user space and kernel space. The kernel reads submissions directly from the SQ and the application reads completions directly from the CQ -- there is no data copying in either direction. This design minimizes both the number of kernel transitions (system calls) and the amount of data copied per I/O operation, which is where traditional async I/O interfaces like epoll spend most of their overhead.

How Monad uses io_uring

Monad's I/O layer wraps io_uring with a fiber-aware interface. The ring wrapper manages ring setup and teardown; the buffer pool manages pre-registered I/O buffers; and the database config captures the tuning parameters for each use case.

Ring configuration

The I/O layer maintains two ring configurations tuned for different workloads:

  • Read-write database (used during block execution): 512 SQ entries, 1,024 concurrent read limit, SQ polling enabled.

  • Read-only database (used during state queries): 128 SQ entries, 600 concurrent read limit, SQ polling disabled.

When SQ polling is enabled, the kernel spawns a dedicated polling thread that continuously drains the submission queue. This thread is pinned to a specific CPU and has a 60-second idle timeout -- if no submissions arrive for a minute, the kernel idles the thread and re-creates it on the next submission. With SQ polling active, the application never needs to call io_uring_enter() to submit I/O; the kernel thread picks up new SQEs automatically. This eliminates the last remaining system call from the submission path.

Both configurations allocate 1,024 read buffers. The read-write configuration additionally allocates 4 write buffers (each 8 MB).

io_uring features used

Beyond basic submission and completion, Monad uses several io_uring features to reduce per-operation overhead (see the I/O submission implementation):

  • Registered file descriptors. All file descriptors are registered with the ring at construction via io_uring_register_files(). Every SQE sets the IOSQE_FIXED_FILE flag and references the registered index rather than the POSIX fd. This avoids a file-table lookup in the kernel on every I/O operation.

  • Registered buffers. Read and write buffer pools are registered via io_uring_register_buffers(). Reads use io_uring_prep_read_fixed() and writes use io_uring_prep_write_fixed(), both of which reference buffers by index. This avoids per-operation page pinning and get_user_pages() overhead in the kernel -- the pages are pinned once at registration time rather than on every I/O.

  • I/O priority classes. The I/O layer assigns Linux I/O priorities to individual SQEs. Critical reads (e.g. a state read on the critical execution path) can be submitted at IOPRIO_CLASS_RT (real-time, priority 7); background reads (e.g. prefetch or compaction) use IOPRIO_CLASS_IDLE. Normal reads use default priority. This lets the NVMe scheduler prioritize latency-sensitive reads over background work.

The read path

When the trie traversal reaches a node that isn't in memory:

  1. The I/O layer constructs a read request with the node's on-disk location and size (both encoded in the node's chunk_offset).

  2. The request is submitted to the io_uring SQ.

  3. The current fiber yields, allowing other fibers to run.

  4. When the kernel completes the read, a completion entry appears in the CQ.

  5. The polling loop picks up the completion and resumes the waiting fiber.

  6. The fiber continues trie traversal with the loaded node.

Buffer management

Read buffers are managed in two tiers:

  • Short reads (up to 4 KB): Drawn from a pre-allocated pool of 1,024 fixed-size buffers. Each buffer is 8 disk pages (8 * 512 = 4,096 bytes). Pool allocation avoids per-read malloc overhead for the common case.

  • Long reads (above 4 KB): Dynamically allocated with alignment to the DMA page size (64 bytes). This alignment is required for direct I/O (O_DIRECT), which bypasses the kernel page cache for predictable latency.

The short-read pool uses an intrusive free list: the next-free pointer is stored in the first 8 bytes of each free buffer, so the pool requires zero auxiliary memory. Allocation is a single pointer dereference and swap; deallocation writes the current head into the buffer being freed and updates the head pointer. The entire BufferPool struct is 8 bytes -- just the head pointer.

On the write side, 4 write buffers of 8 MB each are pre-allocated and registered, giving 32 MB of write buffer space. Both read and write buffers are backed by huge pages to reduce TLB pressure.

Write path

Writes use a separate write ring with larger buffers (8 MB). New trie nodes are written sequentially to append-only chunks, which aligns naturally with io_uring's batching capabilities: multiple node writes can be submitted as a batch and completed in a single kernel operation.

Completion handling

The I/O layer supports two completion-polling modes:

  • Normal mode (default): Process one completion queue entry (CQE) per poll call. This gives fine-grained interleaving between I/O completions and fiber execution.

  • Eager mode: Drain all available CQEs into a batch, then process the entire batch at once. This amortizes the cost of polling when many completions arrive simultaneously.

When a read returns the temporary failure EAGAIN (the kernel could not complete the operation immediately), the I/O layer retries after a 50-microsecond sleep. A retry counter tracks how often this occurs.

The I/O layer maintains an IORecord with running statistics: in-flight read and write counts, peak in-flight counts, total reads submitted, total bytes read, and a retry count. An optional latency- capture mode records per-operation timing for profiling. These statistics are used by the backpressure system and are available for operational monitoring.

Fibers: lightweight concurrency without threads

Monad uses Boost Fibers as its concurrency primitive. A fiber is a user-space coroutine: it has its own stack and can be suspended and resumed, but it runs within a regular OS thread. Each fiber allocates an 8 MB protected stack (via boost::fibers::protected_fixedsize_stack), with guard pages that catch stack overflow. Thousands of fibers can be multiplexed over a small number of threads.

When a fiber issues an async I/O request and yields, the scheduler swaps stack pointers and resumes another runnable fiber. This is a user-space operation -- on the order of tens of nanoseconds, compared to the microseconds required for an OS thread context switch. Thousands of concurrent reads can be in flight, each in its own fiber, with negligible scheduling overhead.

Fiber pool architecture

The fiber infrastructure is organized in layers. A FiberThreadPool manages N OS threads that share a work queue. Each thread installs a priority-aware scheduler that replaces Boost's default round-robin policy. On top of the thread pool, FiberGroup instances create fibers and feed them tasks through a buffered channel (capacity 1,024 pending tasks).

Priority scheduling

The scheduler uses an Intel TBB concurrent_priority_queue as its ready queue. Each fiber has an attached priority value; when a fiber becomes runnable, the scheduler inserts it into the priority queue, and each thread always picks the highest-priority runnable fiber. This lets the execution engine assign priorities to transactions so that critical-path work runs first.

When a fiber picks up a new task, it sets its priority to the task's priority and yields. The yield causes the scheduler to re-insert the fiber into the priority queue at its new position, ensuring it competes fairly with other fibers at the same priority level.

Work-stealing

Fibers can migrate across threads through the shared priority queue. Each OS thread runs a PriorityAlgorithm instance that checks the shared queue when it has no local work. Because the priority queue is thread-safe (TBB's concurrent data structure), this migration works without additional synchronization in the scheduler. The result is automatic load balancing: if one thread finishes its work while another is overloaded, the idle thread pulls fibers from the shared queue.

Backpressure

The I/O layer enforces a configurable limit on concurrent reads (default 1,024 for the read-write database, 600 for read-only). When this limit is reached, new read requests are placed in a deferred-read deque rather than submitted to io_uring. Each time a read completes, the I/O layer dequeues the next pending request and submits it. This keeps exactly the configured number of reads in flight at all times, queuing excess requests until a slot opens.

The deferred-read queue stores the operation handle, the target buffer, and the on-disk offset for each pending read. The in-flight accounting includes both the reads currently in the io_uring pipeline and the deferred reads waiting in the queue, giving the execution engine an accurate picture of total outstanding I/O.

Without backpressure, a burst of concurrent transactions could submit thousands of reads simultaneously, causing the SSD's internal queue to overflow and latency to spike. The limit keeps I/O latency predictable under load.

Originally published on X on February 12, 2026.

More writing →