A compiler performs one of the more remarkable translations in computing: it takes text written for humans and turns it into instructions a processor can execute, catching a whole class of your mistakes along the way. It can feel like a black box — source goes in, a program comes out, and occasionally a cryptic error stops the whole thing. But a compiler is not one monolithic step; it is a pipeline of stages, each with a clear job, each transforming the program into a form a little closer to the machine. Learn the stages and a lot of things click into place at once: why syntax errors read the way they do, why some code is faster than other code, and what language designers are actually trading off.
The shape of the pipeline
Most compilers are organised as a front end that understands the source language and a back end that produces code for a target machine, with an intermediate representation in between. That split is not academic: it is why a compiler infrastructure can support many languages on the front and many processors on the back, meeting in a common middle.
Let us walk each stage with a single tiny expression in mind, something like price = base * 2 + tax.
Stage 1: lexing (tokenizing)
The lexer (or scanner) reads the raw characters and groups them into tokens — the words of the language. Our line becomes a stream like: identifier price, equals, identifier base, star, number 2, plus, identifier tax. Whitespace and comments are usually discarded here. The lexer knows nothing about whether the arrangement makes sense; it only classifies the pieces. If you type a character the language does not allow, this is the stage that objects.
Stage 2: parsing (building the AST)
The parser takes that flat stream of tokens and discovers its grammatical structure, building an abstract syntax tree (AST) — a tree that captures how the pieces relate. This is where precedence lives: multiplication binds tighter than addition, so the tree nests the multiply inside the add.
Source: base * 2 + tax
AST:
(+)
/ \
(*) tax
/ \
base 2The AST is the single most important structure in a compiler; nearly every later stage walks or rewrites it. When you get a "syntax error", it means the parser reached a point where the tokens could not form a valid tree according to the language's grammar.
Stage 3: semantic analysis
A program can be grammatically valid and still meaningless — like a sentence that parses but says something nonsensical. Semantic analysis checks meaning: Is tax actually declared before use? Do the types line up — are you multiplying a number by a number, not a number by a string? In statically typed languages, this is where type checking happens, and where a large share of helpful compiler errors come from. The compiler also builds a symbol table here, tracking every name and what it refers to.
Stage 4: intermediate representation
With a validated AST, the compiler lowers the program into an intermediate representation — a simpler, more uniform form that is no longer tied to the source language's syntax but not yet tied to any specific processor. Think of it as a clean, machine-independent assembly. The IR is the meeting point of the front and back ends, and crucially it is the form on which optimization is easiest to perform.
Stage 5: optimization
Now the compiler improves the program without changing what it does. Optimizations range from the local — folding 2 + 3 into 5 at compile time (constant folding), deleting code whose result is never used (dead-code elimination) — to the global, like hoisting a repeated computation out of a loop. This stage is much of why the same algorithm can run faster when compiled with optimizations on: the machine is doing less work to get the same answer.
Why the front-end / back-end split pays off
Because optimization and code generation operate on the IR, a compiler infrastructure can add a new source language by writing only a front end that emits the IR, or support a new processor by writing only a back end that consumes it. The shared middle is reused. This is the architecture behind widely used toolchains.
Stage 6: code generation
Finally, the code generator translates the optimized IR into the target's actual instructions — machine code for a specific CPU architecture, or sometimes bytecode for a virtual machine. It handles the gritty realities the IR abstracted away: which values live in which registers, how the stack is laid out, how function calls are made. The output is something the machine (or a VM) can run directly.
Compilers, interpreters, and JIT
Not every language runs this whole pipeline ahead of time. An interpreter walks the AST or a bytecode form and executes it directly, without producing a standalone machine-code file — simpler and more flexible, typically slower. Many modern runtimes blend the two with just-in-time (JIT) compilation: they start by interpreting, watch which code runs hot, and compile those parts to fast machine code while the program runs.
Pros
- Ahead-of-time compilation: fast execution, errors caught before shipping, a standalone artifact.
- Interpretation: quick startup, portability, easier interactive development.
- JIT: adapts to real runtime behaviour, optimizing the paths that actually matter.
Cons
- Ahead-of-time: slower edit-run cycle; must target each platform.
- Interpretation: generally slower at steady state.
- JIT: warm-up cost and higher runtime memory and complexity.
Practical takeaway
The next time a compiler stops you, the stage names tell you what went wrong: a stray character is the lexer, a malformed structure is the parser, an undeclared variable or a type mismatch is semantic analysis. And the next time you wonder why optimized code is faster, it is the optimizer rewriting your program on the IR to do less work. You do not need to write a compiler to benefit from this map — but if you want to, building a small interpreter for a toy language is the single best way to make all six stages concrete, and the resources below are the standard on-ramps.
Sources & Further Reading
- 01Crafting Interpreters — Robert NystromA free, complete book that builds a language from lexer to runtime — the best hands-on introduction.
- 02LLVM Documentation — The LLVM ProjectThe docs for a widely used compiler infrastructure built around a shared intermediate representation.
- 03GCC Online Documentation — GNU ProjectReference material for the GNU Compiler Collection, including its optimization passes.
Editorial note — A conceptual explainer of the standard compiler pipeline. The example AST is illustrative; no product benchmarks or performance figures are quoted.


