JIT compilation
The EVM is a stack-based virtual machine that executes smart contract bytecode. In a standard Ethereum client, an interpreter processes bytecode one opcode at a time: fetch the next instruction, check gas, check stack bounds, dispatch through a function table, execute, repeat. These per-instruction overheads add up, and the same contract code is re-interpreted identically on every call.
Monad includes both an optimized interpreter and a JIT (just-in-time) compiler that translates EVM bytecode to native x86-64 machine code. The compiler eliminates per-instruction overhead and enables optimizations that are impossible in an interpreter, while maintaining exact EVM semantics.
The dual-path architecture
Monad maintains two representations of contract code:
Intercode: An optimized intermediate representation used by the interpreter. This is the fallback for all contracts and is always available.
Nativecode: Compiled x86-64 machine code, produced by the JIT compiler for hot contracts. When available, it replaces the interpreter for that contract.
Both representations are cached together in a Varcode structure (cache), keyed by the contract's code hash. The Varcode also tracks cumulative gas usage under interpretation, which drives the decision of when to compile.
Adaptive compilation
Not all contracts are worth compiling. A contract deployed once and never called again doesn't benefit from compilation. A heavily-used DEX router called thousands of times per block benefits enormously.
Monad uses gas-weighted adaptive compilation. Each time a contract is executed via the interpreter, the gas consumed is accumulated in the Varcode's atomic counter. When the cumulative gas crosses a threshold proportional to the bytecode size -- roughly 32x the bytecode length -- the contract is submitted for background compilation.
This means:
Contracts are compiled in order of impact: the contracts consuming the most execution gas are compiled first.
Small, simple contracts are compiled at a lower absolute threshold than large, complex ones.
Contracts that are called rarely stay on the interpreter, avoiding wasted compilation effort.
Compilation happens asynchronously in a dedicated background thread. A concurrent queue holds pending compilation jobs, and the compiler thread drains them one at a time. While a contract is being compiled, it continues to execute on the interpreter. Once compilation completes, the native code is inserted into the Varcode cache and subsequent calls use it immediately.
Intercode: the optimized interpreter representation
Before any JIT compilation occurs, raw EVM bytecode is transformed into Intercode -- a pre-processed form that the interpreter can execute more efficiently.
One transformation is code padding: the bytecode is extended with 30 bytes before and 33 bytes after the actual instructions. This padding prevents the interpreter's lookahead logic from reading out of bounds when scanning multi-byte instructions like PUSH32 near the end of the code.
Intercode also pre-computes a JUMPDEST map -- a vector<bool> that marks valid jump destinations. When the interpreter or compiler encounters a JUMP or JUMPI, it validates the target against this map in O(1) time. The JUMPDEST scan is careful to skip over immediate data bytes that follow PUSH instructions (e.g. the 32 bytes after a PUSH32 cannot contain a valid JUMPDEST).
The compilation pipeline
The compiler translates EVM bytecode to x86-64 machine code through several stages:
1. Basic block analysis
The bytecode is divided into basic blocks: sequences of instructions with a single entry point and a single exit. Basic blocks end at control flow instructions (JUMP, JUMPI, RETURN, REVERT, STOP, SELFDESTRUCT, or invalid opcodes) and begin at jump destinations (JUMPDEST).
Each basic block has a terminator type:
| Terminator | Description |
|---|---|
| FallThrough | Block ends by falling into the next block |
| Jump | Unconditional jump to a dynamic target |
| JumpI | Conditional branch (JUMPI) |
| Return | RETURN instruction |
| Stop | STOP instruction |
| Revert | REVERT instruction |
| SelfDestruct | SELFDESTRUCT instruction |
| InvalidInstruction | Illegal opcode; always reverts |
The basic block analysis also computes stack delta for each block: the net change in stack depth (delta), plus the minimum stack depth reached during the block (min_delta) and maximum (max_delta). These deltas allow the compiler to validate that a block never underflows the stack and to know the exact stack layout at each block boundary.
2. Gas check batching
In the interpreter, gas is checked before every opcode. The compiler performs a single gas check at the beginning of each basic block, covering the total static gas cost of all instructions in the block. This replaces N gas checks (one per opcode) with one gas check per basic block.
The compiler uses a threshold of 1,000 gas for when a gas check is worth emitting. Blocks with less than this amount of static gas may have their check deferred to the next block, reducing branch overhead in tight inner loops.
Instructions with dynamic gas costs (like SLOAD, CALL, or memory expansion) still require individual gas checks at runtime, but the static portion is batched. The emitter distinguishes between "static work" blocks (where the total gas is statically known) and "unbounded" blocks (which contain dynamic-gas instructions and must check more carefully).
3. Constant folding
Sequences of constant operations are evaluated at compile time. For example:
PUSH1 0x02
PUSH1 0x03
ADD
The compiler recognizes that the result is a constant (5) and replaces the three-instruction sequence with a single literal value. The gas accounting is preserved -- the compiled code charges the correct gas for all three original instructions -- but the CPU work is eliminated.
The virtual stack (described below) represents values not just as register or memory locations but also as literals: compile-time known constants. Arithmetic on two literals produces another literal, and the constant is propagated as far as possible before it must be materialized into a register.
4. Register allocation
The EVM is a stack machine, but x86-64 has 16 general-purpose registers and 16 AVX/SSE vector registers. The compiler maps EVM stack slots to machine locations through a virtual stack -- an internal data structure that tracks where each stack element currently lives.
Each stack element (StackElem) can be in one of four states:
Literal - A compile-time constant -- no register needed
AvxReg - An AVX register (ymm0–ymm15 or zmm0–zmm15)
GeneralReg - An x86-64 general-purpose register (rax, rcx, rdx)
StackOffset - Spilled to the runtime stack at a fixed offset
The register allocator maintains:
16 AVX registers for 256-bit values (the EVM's native word size). AVX2/AVX-512 instructions can perform 256-bit operations in a single instruction, directly matching the EVM's word width.
3 general-purpose registers (rax, rcx, rdx) for values that participate in address computations, control flow, or operations that map naturally to scalar x86 instructions.
Two dedicated registers are reserved: rbx holds a pointer to the execution context, and rbp holds a pointer to the EVM stack in memory.
When more values are live than available registers, the allocator spills values to the runtime stack at fixed offsets. The stack frame layout reserves 6 slots for function arguments plus additional temp storage.
Because the virtual stack tracks the exact location of each operand, instructions are specialized to their operand locations at emit time. For example, an AND with both operands in AVX registers emits a single vpand (ymm, ymm, ymm) instruction operating directly on 256-bit registers, with no data movement. If one operand was spilled, the emitter falls back to a memory-source form instead. This location-aware emission is what makes the virtual stack worth tracking at all: the same EVM opcode can produce different (cheaper) native sequences depending on where its inputs currently live.
The virtual stack also implements a deferred comparison optimization: conditional branches (JUMPI) check a boolean condition. If that condition was just computed by a comparison instruction, the compiler can avoid materializing the 0/1 result into a register and instead emit a direct conditional jump based on the CPU's flags register.
5. Native code emission
The compiler emits x86-64 machine code directly, using the asmjit library as an assembler backend. This is a deliberate architectural choice: Monad does not use LLVM or any other general-purpose compiler framework. Direct emission provides:
Maximum control over the generated code: every instruction is chosen specifically for EVM execution patterns.
Minimal compilation latency: no intermediate optimization passes, no linker invocation, no relocation processing.
Predictable performance: the generated code is closely tied to the input, with no surprises from optimizer heuristics.
The trade-off is that the compiler must implement its own optimizations rather than relying on a framework, but the optimization surface for EVM bytecode is narrow enough that this is practical.
The emitter uses a custom EmitErrorHandler that captures asmjit errors and converts them into typed error codes (NoError, Unexpected, SizeOutOfBound), allowing the compilation pipeline to cleanly handle failures and fall back to the interpreter.
6. Code size bounds
The compiler enforces a maximum compiled code size of 32 times the original bytecode length. This bound prevents integer overflow in relative jump addressing within the generated code: if a conditional branch in the native code uses a 32-bit signed offset to reach another block, and each bytecode byte can produce at most 32 bytes of native code, the offset can never exceed the addressable range.
If compilation would exceed this bound, it fails with SizeOutOfBound and the contract falls back permanently to the interpreter.
7. Read-only data section
The compiler maintains a read-only data section (RoData) for literals (256-bit constants, addresses, etc.) referenced by the generated code. Constants are deduplicated: if the same 256-bit value appears multiple times in the bytecode, it's stored once in the RoData section and referenced from multiple points in the generated code.
This matters for contracts that use the same constant in multiple contexts -- for example, a contract that uses a specific bitmask in several operations will load that mask from a single memory location rather than embedding it repeatedly in the instruction stream.
The interpreter
The interpreter is the fallback execution path and must be fast in its own right. It uses several key optimizations:
Computed-goto dispatch: Rather than a switch statement, the interpreter uses a dispatch table (instruction_table) of computed goto targets -- one label address per opcode. At each step, the interpreter loads the next opcode, indexes into the table, and jumps directly to the handler. This is faster than a switch because it avoids a branch-prediction-hostile indirect branch or range check.
Assembly trampoline: The interpreter's inner loop is implemented in assembly (or with assembly hints) to give precise control over register usage, ensure the instruction pointer and stack pointer stay in registers across iterations, and avoid compiler-generated prologue overhead.
Shared infrastructure: The interpreter and compiled code share the same memory pool, host interface, runtime function table, and state infrastructure. A contract that transitions from interpreted to compiled during its lifetime (as it accumulates gas) does so transparently -- callers see no difference.
The code cache
Compiled code is stored in a weight-based LRU cache keyed by code hash. The cache has a default capacity of 4 GB and a "warm" threshold at 75% capacity.
Each entry's weight is proportional to the compiled code size in kilobytes, plus a fixed 3 kB overhead per entry to account for metadata. When the cache exceeds capacity, the least recently used entries are evicted. The warm threshold is used to determine when the cache has enough entries to be effective -- until the cache is warm, the compiler may prioritize compilation more aggressively.
The cache is shared across all threads via the Varcode abstraction: the concurrent hash map allows multiple fibers to look up compiled code simultaneously without contention.
EVM revision support
Monad's execution engine supports multiple EVM revisions (Frontier, Homestead, Istanbul, Berlin, London, Shanghai, Cancun, Prague, and others). The compiler is parameterized by a traits type that encodes which opcodes are available, which gas schedules apply, and which behaviors differ between revisions.
Two trait implementation strategies are used:
Explicit traits: A template instantiated once per revision, generating separate native code paths per revision. This is the primary strategy for the JIT compiler, where the overhead of a runtime revision check would be paid on every compiled call.
Switch traits: A runtime dispatch strategy used in contexts where template instantiation overhead would be prohibitive.
When a new revision is added, both the compiler and interpreter are updated in lockstep, ensuring that the compiled and interpreted paths produce identical semantics.
When the interpreter is still used
The interpreter handles several cases where compilation is not applicable or not yet ready:
Cold contracts: Contracts that haven't accumulated enough gas usage to trigger compilation.
CREATE/CREATE2: Contract creation always runs through the interpreter, because the bytecode is new and potentially unique.
Compilation in progress: While the background thread is compiling a contract, calls to that contract use the interpreter.
Compilation errors: If compilation fails (e.g., the bytecode exceeds the code size bound or uses patterns the compiler doesn't handle), the contract falls back to the interpreter permanently.
Correctness
The compiler preserves exact EVM semantics. Gas accounting is identical; every opcode is charged the same gas whether interpreted or compiled. Stack behavior is identical: the same values are pushed and popped in the same order. State effects are identical: the same storage reads and writes occur.
The Category Labs team runs a fuzzer that generates random EVM bytecode sequences and verifies that the interpreter and JIT compiler produce the same results.
The only difference is speed. A compiled contract executes faster because per-instruction overhead is eliminated, constants are pre-computed, and values live in registers instead of a software stack. The optimization is invisible to smart contract developers, users, and any tool that interacts with the EVM.
Interested to learn more? Check out the code.
Originally published on X on April 7, 2026.