Keone’s blog
← All writing

How parallel execution actually works

protocol

Ethereum executes transactions sequentially: transaction 1 completes before transaction 2 begins. Monad executes transactions optimistically in parallel, preserving Ethereum's sequential semantics while doing more work concurrently.

The primary benefit is initiating more state access in parallel. Sequential execution can only issue storage reads for one transaction at a time -- the next transaction's reads don't begin until the previous transaction finishes. With parallel execution, multiple transactions' storage reads are in flight simultaneously, hiding I/O latency across transactions the same way async I/O hides it within one. Computational work also runs in parallel across CPU cores, so EVM execution, signature recovery, and other CPU-bound work for different transactions overlap rather than queue up.

This pairs directly with MonadDB, Monad's custom storage engine. MonadDB is optimized for concurrent random reads, so the parallel access pattern that execution generates maps well onto what the storage layer can efficiently serve.

The key insight is that most transactions in a block don't conflict. Two token transfers between different accounts touch entirely separate state. A DEX swap and an NFT mint have no overlapping storage slots. Sequential execution imposes a strict ordering even when there are no data dependencies, waiting for each transaction to finish before the next one can begin.

Optimistic concurrency control

Monad's parallel execution is based on optimistic concurrency control, a well-studied technique from database systems. The idea:

  1. Execute transactions concurrently, as if they were independent.

  2. Track what each transaction read (its read set).

  3. After execution, check whether any transaction's read set overlaps with an earlier transaction's write set.

  4. If there's a conflict, re-execute the conflicting transaction with the correct state.

  5. Merge state changes in the original transaction order.

This preserves the semantics of sequential execution: the final state is identical to what you'd get by executing every transaction one at a time, in order.

Pre-execution: parallel sender recovery

Before any transaction executes, all ECDSA sender signatures are recovered in parallel. recover_senders() submits one fiber per transaction to the priority pool, each recovering the sender address via recover_sender() (a secp256k1 ecrecover call). All recoveries complete before block execution begins.

This matters because ecrecover is one of the most CPU-expensive operations per transaction. Parallelizing it across all transactions in the block means this cost is paid once, concurrently, before execution begins -- and re-execution never pays it again.

Fiber infrastructure

Each transaction runs in a fiber -- a cooperative coroutine multiplexed over a thread pool, built on Boost.Fibers. Fibers yield when waiting on I/O or on a predecessor transaction, allowing other transactions' fibers to run in the meantime.

Transactions are submitted to the fiber pool with their index i as priority, so lower-indexed (earlier) transactions are scheduled first. This reduces stall time: earlier transactions are more likely to be on the critical path, blocking later ones from merging.

Block execution loop

execute_block_transactions submits all N transactions to the fiber pool concurrently, then waits for all of them to complete and merge in order. Ordering is enforced by a chain of boost::fibers::promise objects: each transaction waits on its predecessor's promise before merging, then sets its own promise to unblock the next transaction.

State architecture

Two state layers separate the global block view from each transaction's local view.

BlockState is the single authoritative state for the block, backed by a TBB concurrent hash map that all fibers can read simultaneously. For each account and storage slot, it stores both the value at block start and the current merged value. The current value is updated as transactions merge in order.

Each transaction works against a fresh State that records two things as execution proceeds: what it read from BlockState (the read set) and what it intends to write (the write set). At merge time, can_merge() checks whether the read set is still consistent with the current BlockState; if so, the write set is applied.

Each State carries an Incarnation -- a (block number, transaction index) pair -- to prevent stale storage reads across contract lifetimes: after a SELFDESTRUCT and re-CREATE at the same address, the new account has a different incarnation and its old storage is ignored.

Within a transaction, VersionStack supports EVM call-level rollback: sub-call writes are pushed onto a stack and discarded on REVERT, restoring the pre-call state.

Transaction execution

ExecuteTransaction::operator() validates and runs the transaction against a fresh State, then waits for its predecessor to merge. Once the predecessor completes, can_merge() checks whether every account and storage slot the transaction read still holds the same value in BlockState. If so, the write set is applied and the next transaction is unblocked. If not, the transaction re-executes from scratch against the now-current state, which is guaranteed to be conflict-free.

Why conflicts are rare

In a typical block, the conflict rate is low. For example:

  • Token transfers: touching different accounts → no conflict.

  • DEX swaps on different pools: different contract storage → no conflict.

Conflicts arise when two transactions write the same storage slot (two swaps on the same DEX pool, two mints that increment the same counter). Even then, only the later transaction re-executes; the earlier one merges unconditionally.

Why re-execution is cheap

When re-execution occurs:

  • Signature recovery: already done before block execution; never repeated.

  • State caching: accounts and storage slots read during the first execution are now in BlockState's concurrent maps, so re-execution reads from memory rather than from the database.

  • JIT compilation: if the called contract was compiled to native code during the first execution pass, that compiled code is used immediately.

In practice, re-execution costs significantly less than the first execution because I/O, crypto, and compilation have already been paid.

Integration with async I/O

Fibers yield when they issue storage reads via io_uring. While one transaction's fiber is waiting for a disk read, the scheduler picks another ready fiber. Multiple transactions' I/O requests are in flight simultaneously, and the SSD handles them concurrently. This synergy between parallel execution and async I/O is what sustains high CPU and storage utilization simultaneously.

Correctness guarantee

The final block state after parallel execution is identical to sequential execution. Merges happen in transaction-index order via the promise chain, so each transaction sees the definitive state of all predecessors when it merges. Any transaction that read stale data is re-executed before merging, and that re-execution is guaranteed to succeed because all earlier state is already final.

The result is that Monad blocks have the exact same semantics as Ethereum blocks: a linearly ordered set of transactions, each seeing the state produced by all previous transactions.

Originally published on X on April 11, 2026.

More writing →