How a Stack Machine Runs Bytecode
The previous lesson introduced bytecode as a flat list of instructions for a machine defined in software. This one makes it concrete, because the idea stops being mysterious the instant you watch a machine actually run. We will use a stack machine, the most common design, with exactly four instructions, which is enough to evaluate any basic arithmetic and enough to see the whole trick.
If a machine with four instructions sounds too small to matter, remember the line cook from the bytecode lesson. A cook who can only chop, stir, heat, and plate can still produce any dish, given a long enough list of steps. Small instruction sets are not a limitation; they are the point. Simple pieces, arranged carefully, compute anything.
The whole instruction set
A stack machine keeps a single pile of values, the stack, and every instruction either pushes onto it or pops from it. Four operations:
- PUSH n: put the number n on top of the stack.
- ADD: pop the top two numbers, push their sum.
- SUB: pop the top two numbers, push their difference.
- MUL: pop the top two numbers, push their product.
That is the entire language. No variables, no names, no operators in the source sense, just opcodes and a stack. Yet this is enough to compute expressions of any depth, because the stack remembers intermediate results for you.
Compiling an expression to it
Take 2 + 3 * 4. A parser first builds a tree that respects precedence, so the multiplication sits deeper than the addition. Walking that tree produces this bytecode:
PUSH 2
PUSH 3
PUSH 4
MUL ; pop 4 and 3, push 12
ADD ; pop 12 and 2, push 14The precedence that lived in the tree is now baked into the order of the instructions. Nothing in this list contains a + or a *; it contains MUL and ADD as opcodes and the operands they consume from the stack. Change the expression to (2 + 3) * 4 and only the compiled order changes (PUSH 2, PUSH 3, ADD, PUSH 4, MUL), producing 20. Same four instructions, different arrangement, different answer.
Step through it yourself
The interactive machine below runs this bytecode for real, one instruction at a time, and shows the stack after every step. Press step to feed one instruction into the loop and watch a value get pushed or two get combined. This tiny loop, read an instruction, do what it says, repeat, is the beating heart of every bytecode VM, from this toy to a production language runtime.
A good exercise: before each press, say out loud what the instruction will do and what the stack will look like afterward, then press and check. If you can predict all five steps of 2 + 3 * 4, you understand stack machines. There is genuinely nothing more to the core idea.
Try hand-compiling before you step: write the bytecode for 5 * (2 + 3) yourself, then for 2 * 3 + 4 * 5, which needs an intermediate result held on the stack while a second product is computed. Notice that the stack depth an expression needs is decided entirely at compile time by its tree shape, deeper trees need deeper stacks, and real compilers compute exactly this number ahead of time to size each function's frame.
Now scale the intuition up
A real virtualizing obfuscator is this idea with the volume turned up. Instead of four opcodes it has many, covering variables, function calls, jumps, and comparisons. Instead of readable PUSH and ADD names, the opcodes are opaque numbers that differ per build. Instead of plain integer operands, constants come from an encoded table. And the interpreter loop that ties it together is itself run through the earlier techniques, renamed, flattened, sprinkled with opaque predicates, so that even the machine running your bytecode is hard to read.
When a decompiler opens virtualized output, this loop is what it finds, not your program. To recover your logic, an attacker must first understand this interpreter and its instruction format, then replay the bytecode through it in their head or their tools. That first step, absent from every other technique, is where most of virtualization's cost comes from.
And the ceiling still holds. Because the machine has to actually execute the bytecode to compute anything, an attacker who runs it can log every instruction and every stack value as they happen, and reconstruct behaviour from the trace. The toy above makes that visible: the stack panel is exactly the kind of state a run-and-dump attacker would capture. Virtualization makes that capture slow and per-build, which is the honest and considerable value it provides.
A stack machine is a loop plus a pile: fetch an opcode, pop what it needs, push what it produces, repeat. Precedence from the tree becomes instruction order, and four opcodes already compute arbitrary arithmetic. Next, the final lesson of this level descends one floor further, to the physical CPU running this same loop in silicon.
Frequently asked questions
What is a stack machine?
A virtual machine that holds values on a single stack and whose instructions push values onto it or pop values off it to compute results. It is the most common VM design because compiling expressions to it is simple: the stack automatically remembers intermediate results.
What is the difference between a stack machine and a register machine?
A stack machine keeps operands on a stack and instructions operate on the top of it; a register machine keeps operands in named slots (registers) and instructions reference them explicitly. Stack designs give smaller, simpler bytecode; register designs can be faster. Both are used in obfuscation VMs.
Why does watching the stack help against obfuscation?
It does not help you attack code; it helps you understand what virtualization does. The stack view shows that the machine must materialize real values to compute, which is precisely why a run-and-dump attacker can observe them, and why no VM can claim to be unbreakable.