Tour 42 · The dance of privilege · about 30 minutes · 21 steps
Every instruction a CPU executes runs with some value in sp, and that value is always a
claim: “this is my stack.” This tour follows the claim on hart 0 from the moment QEMU
powers it on, when sp is simply 0, to the moment init executes its first user
instruction on a stack that kexec built for it.
On the way, sp points into four different pieces of memory: nothing at all, hart 0’s slice
of the boot array stack0 (which quietly becomes the scheduler stack), the first
process’s empty kernel stack, and finally init’s user stack. Only four instructions
in the whole kernel ever move sp from one stack to another. This tour runs three of them
and points at the fourth.
Every address in this tour comes from this build: xv6/kernel/kernel.sym, and gdb attached to
QEMU, stopping at each point and reading sp. One surprise from the measurements shapes the
story: the first process doesn’t stay on hart 0. Partway through, the thread leaves the hart,
and the tour follows the stack, not the hart. That turns out to be the main lesson.
Best after: 2. Power-on to main, on every hart at once, 4. From the first process to the shell prompt, 41. Every transition: mode, stack and page table
Three harts power on together. All three run the same early code, each on its own stack; this tour watches hart 0, which builds the kernel.
| Hart | What it is doing |
|---|---|
| 0 | The tour’s hart: boots, runs main's setup, becomes a scheduler |
| 1 | Boots in parallel, then spins in main until hart 0 says started |
| 2 | The same as hart 1 |
Step 1 of 21
QEMU starts every hart in machine mode at 0x1000, a small boot ROM it provides. gdb’s
first look, before a single instruction ran:
0x1000: auipc t0, 0x0
0x1004: addi a2, t0, 40
0x1008: csrr a0, mhartid # a0 = this hart's number
0x100c: ld a1, 32(t0) # a1 = 0x87e00000, a device-tree address
0x1010: ld t0, 24(t0) # t0 = 0x80000000
0x1014: jr t0
The two words after the code hold 0x80000000 and 0x87e00000. So the ROM jumps every
hart to 0x80000000, where the linker put _entry (kernel/kernel.ld). xv6 ignores
the device tree in a1.
At this point sp is 0, on all three harts. There is no stack: nothing may call a
function or push anything. The ROM doesn’t need one, and the first five instructions of
_entry don’t either. Tour 2: Power-on to main, on every hart at once tells the boot story; this tour watches sp.
Step 2 of 21
Six instructions compute sp = stack0 + (hartid + 1) * 4096. In this build
stack0 is at 0x80007890 (kernel/kernel.sym), so:
| Hart | sp after line 17 (measured at start) |
|---|---|
| 0 | 0x80008890 |
| 1 | 0x80009890 |
| 2 | 0x8000a890 |
Why hartid + 1? A RISC-V stack grows down: a push first subtracts from sp. So
each hart’s sp starts at the top of its slice, which is the bottom of the next
hart’s slice.
This is the first of the four instructions in xv6 that move sp onto a different
stack (The stacks of xv6). Its stack is the boot stack.
Every hart computes its own answer from its own mhartid, so no lock and no
coordination are needed: hardware gives each hart a different number.
stack00x80008890Step 3 of 21
stack0 in this build, with what the linker put just below it (kernel/kernel.sym).
Each slice is one hart’s boot stack, later its scheduler stack; sp starts at the top
of a slice and moves down.stack0 is an ordinary global array: 4096 * NCPU bytes, and NCPU is 8, so 32 KiB, of
which three slices are used on our three-hart machine.
stack00x80008890 at entry, 0x80008880 after start’s 16-byte frameStep 4 of 21
call start (entry.S line 19) puts the return address in ra; RISC-V’s call pushes
nothing. start's compiled prologue then takes 16 bytes of stack for itself:
sp goes from 0x80008890 to 0x80008880.
start runs entirely in machine mode, with no page table (satp = 0: machine mode
doesn’t translate addresses at all). It is preparing a fake trap return: MPP = S
(“the trap came from supervisor mode”) and mepc = main (“…at this address”).
Tour 46: Boot: returning from a trap that never happened tells that story; here, notice what start does not prepare: a new stack.
The mret ahead won’t touch sp.
start also turns on two extensions in menvcfg: Svadu (line 41, the hardware updates
the accessed and dirty bits in page-table entries) and Sstc (in timerinit, a
supervisor timer). gdb read menvcfg = 0xa000000000000000 just before mret: bit 63
(STCE) and bit 61 (ADUE).
stack00x80008880Step 5 of 21
gdb on hart 0 at the mret (0x800000ca) and again at main’s first instruction:
at mret |
at main |
|
|---|---|---|
| mode | M | S |
pc |
0x800000ca |
0x80000e5e (mepc) |
mstatus.MPP |
1 (S) | 0 (U: mret resets it to the least-privileged mode) |
mstatus.MPIE |
0 | 1 |
sp |
0x80008880 |
0x80008880 |
tp |
0 | 0 |
sp is unchanged, so main starts on the same slice, just below start’s 16-byte frame.
That frame stays there forever: start never returns, and main never will either.
After this instruction, no hart re-enters machine mode for the rest of the run.
mtvec and mscratch are never written in this tree; gdb read both as 0 at main.
Every trap is delegated to supervisor mode, and the timer is a supervisor timer.
stack00x80008880 at entry; 0x80008870 after main’s prologueStep 6 of 21
main is in supervisor mode, and satp is still 0, so addresses are physical.
Hart 0 now builds the kernel: console, allocator, page table, process table and the
rest, all as ordinary C calls on this 4 KiB slice. Every call pushes a frame below
0x80008870 and pops it on return.
Interrupts are off (sstatus.SIE = 0; gdb read sstatus = 0x200000000), and stay off
until the scheduler. Nothing can interrupt the setup, so nothing needs a stack of its
own for interrupts yet.
The next three steps follow three things hart 0 builds that decide where sp will go
later.
stack0a few frames below 0x80008870Step 7 of 21
kvminit → kvmmake → proc_mapstacks allocates 64 pages, one per process
slot, and maps each at KSTACK(i) in the kernel page table it is building:
| Slot | Kernel stack (virtual) | Top (p->kstack + PGSIZE) |
|---|---|---|
| 0 | 0x3fffffd000 |
0x3fffffe000 |
| 1 | 0x3fffffb000 |
0x3fffffc000 |
| 2 | 0x3fffff9000 |
0x3fffffa000 |
KSTACK(i) is TRAMPOLINE - (i+1) * 2 * 4096 (kernel/memlayout.h:52): the stacks
are two pages apart, and the page below each one is never mapped. That unmapped page is
the guard page: an overflow faults instead of corrupting the next stack.
These stacks are allocated once, at boot, and never freed. They belong to the slot, not
to a process. allocproc doesn’t allocate a kernel stack; it finds one waiting.
(The stacks of xv6)
stack0below 0x80008870, same value before and afterStep 8 of 21
kvminithart writes satp = 0x8000000000087fff (measured): Sv39, root page table at
physical 0x87fff000. From the next instruction on, every address hart 0 uses,
including sp and the pc, goes through that table.
Why doesn’t the stack vanish? Because kvmmake mapped all of the kernel’s data and
free RAM at the same virtual address as its physical address
(kernel/vm.c:42), read-write. stack0 lives at 0x80007890, inside that range, so
sp = 0x8000886x means the same bytes before and after the switch. The kernel’s own
code is mapped the same way, which is why the pc survives too.
This csrw satp is the only satp write outside the trampoline (apart from
start’s w_satp(0)), and it runs once per hart. The trampoline will do the rest.
stack0main’s frames, below 0x80008870acquire in allocproc recorded intena 0init's p->lockStep 9 of 21
userinit calls allocproc, which takes slot 0 for pid 1, init. Lines 145–147 are
the most important lines for this tour:
p->context.ra = forkret (0x8000193a)
p->context.sp = p->kstack + 4096 (0x3fffffe000)
Nothing has run on this process yet, so there is no real context to save. allocproc
writes one by hand. When a scheduler later switches to it, swtch will “return”
to forkret with sp at the top of slot 0’s kernel stack: empty.
The process also gets a trapframe page from kalloc, filled with the junk byte 0x05.
gdb printed every field of init’s trapframe as 0x0505050505050505 at this point.
That will matter at the very end.
allocproc returns holding init’s p->lock; userinit marks it RUNNABLE and
releases it.
stack00x80008870 at scheduler’s entry; 0x80008810 after its 96-byte framewas: boot stackStep 10 of 21
main calls scheduler, which never returns. No instruction moved sp to a new
stack: gdb read 0x80008870 at scheduler’s first instruction, just below main’s
16-byte frame. Same memory, new role. From now on this slice is hart 0’s
scheduler stack.
hart 0's slice of stack0
0x80008890 ─ top
0x80008880 start's frame (16 bytes, dead)
0x80008870 main's frame (16 bytes, dead)
0x80008810 scheduler's frame (96 bytes, live forever)
0x80007890 ─ bottom
This is why the scheduler stack isn’t part of struct cpu and isn’t allocated anywhere:
it is simply the stack the hart was already standing on. The frame is 96 bytes because
the prologue saves ra and s0–s9, the callee-saved registers in which the loop keeps
all its variables. The functions it calls (acquire, release, swtch) push their own
frames briefly below it.
stack00x80008810intr_off() on line 442: intena 0init's p->lockStep 11 of 21
The loop opens the interrupt window (lines 441–442, the subject of Tour 44: One interrupt, three landing sites), then
scans the table. Slot 0 is RUNNABLE. Under init’s p->lock, the scheduler marks it
RUNNING, sets c->proc, and calls swtch(&c->context, &p->context).
In our run hart 0 got there first. That isn’t guaranteed: hart 0 set started and walked
straight into scheduler(), while harts 1 and 2 still had to print “hart N starting” and
set up paging, so hart 0 has a head start, but only a head start.
The lock is held across the switch. It will be released on the other side, by code running on a different stack (Tour 13: swtch and the lock handed across a context switch).
stack00x80008810init's p->lockStep 12 of 21
Fourteen stores into cpus[0].context. gdb printed the result later, from forkret:
| Field | Value | What it is |
|---|---|---|
ra |
0x80001df0 |
the instruction after jal swtch in scheduler |
sp |
0x80008810 |
the scheduler’s frame on hart 0’s slice |
s0 |
0x80008870 |
the frame pointer: the frame’s top |
s1 |
0x8000fdd0 |
p, the loop variable: &proc[0] |
This is all the scheduler needs to continue later: where to return, which stack, and
its callee-saved registers. struct context is a save area, not a stack
(The stacks of xv6).
swtch is plain code: no CSR, no mode change, no page-table change. It is S-mode on
the kernel page table before, during and after.
ld sp, 8(a1) in swtch (kernel/swtch.S:26)init’s thread along with the lock; gdb read noff 1, intena 0init's p->lockStep 13 of 21
ld sp, 8(a1) loads the forged p->context.sp: 0x3fffffe000. This is the third of
the four stack-switching instructions. The remaining loads fill s0–s11 with zeros
(allocproc cleared the context), and ret jumps to ra, which is forkret.
gdb at forkret’s first instruction:
| Register | Value |
|---|---|
sp |
0x3fffffe000 |
ra |
0x8000193a: forkret itself, not a return address |
sstatus |
SIE = 0, SPIE = 1 |
cpus[0].noff, intena |
1, 0 |
forkret has no caller. Its stack is empty, and its ra points at itself. It can never
return; it must leave some other way, and it will, through the trampoline.
release took noff to 0, and with intena 0 pop_off left SIE offStep 14 of 21
Line 520 releases the p->lock that hart 0’s scheduler took. release calls pop_off,
which turns interrupts on only if intena is 1. It is 0: when the scheduler did
acquire, interrupts were already off (line 442). The lock and its noff level crossed
swtch together, and the release brings noff from 1 to 0 without touching SIE
(Locks and interrupt state, Locks and interrupt state). So interrupts stay off, and they will
stay off for the whole of forkret on this first run (gdb later read SIE = 0 inside
kexec too).
Then the one-time work: fsinit reads the superblock and recovers the log, and
kexec loads /init. Both need the disk, and the disk means sleeping, which
needs a process with a kernel stack. That is why this happens here and not in main
(kernel/proc.c:525).
forkret, so sleep’s acquire recorded intena 0block 1's buffer (sleep-lock)init's p->lockStep 15 of 21
readsb → bread → virtio_disk_rw starts a read of block 1 and must wait. In
sleep, init marks itself SLEEPING and calls sched. gdb at sched’s entry
read sp = 0x3fffffdf00 on hart 0. Then swtch loads cpus[0].context, and hart 0’s
sp goes back to 0x80008810, its scheduler stack, exactly where step 12 left it.
Here the tour has to choose between the hart and the stack. Hart 0 now loops in its
scheduler. init’s kernel stack holds five frames (readsb is inlined into fsinit), six once
sched pushes its own, and waits. When the disk interrupt
wakes init, whichever hart’s scheduler finds it first loads init’s saved sp
and continues on this stack.
In our run init went to sleep 19 times before it reached user mode (disk reads in
fsinit and kexec), and resumed on harts 0, 1 and 2 in turn. Its final stretch ran on
hart 1. (In another run it was hart 2.) The stack is the thread; the hart is only
where it happens to run.
Step 16 of 21
kexec runs on init’s kernel stack (on hart 1, in our run) and builds a whole new
address space in a page table that isn’t installed anywhere. After loading the program
(which ends below 0x2000), it allocates two more pages and makes the lower one a guard
by clearing PTE_U (uvmclear). gdb at line 137:
| Variable | Value |
|---|---|
sz |
0x4000 |
| guard page | 0x2000–0x2fff (no PTE_U) |
stackbase |
0x3000 |
user sp |
0x3fe0 |
elf.entry |
0xbc, start in user/ulib.c |
kernel sp |
0x3fffffddb0 |
0x3fe0 holds the argv array: argv[0] = 0x3ff0 (the string "/init", 6 bytes
rounded to 16), argv[1] = 0. Lines 136–137 store 0xbc and 0x3fe0 into the
trapframe. Nothing runs on this stack yet: it is a promise, written into memory that
only init’s new page table can reach.
Step 17 of 21
Back in forkret, prepare_return fills in what the trap hardware and the
trampoline will need (Tour 43: A system call, CSR by CSR walks through it CSR by CSR). On hart 1 (the CSR rows
measured; the trapframe rows follow from lines 117 and 119 and the junk kalloc left):
| Register / field | before | after |
|---|---|---|
stvec |
0x800055b0 |
0x3ffffff000 |
sepc |
0x80001e10 |
0xbc |
trapframe->kernel_sp |
0x0505… |
0x3fffffe000 |
trapframe->kernel_hartid |
0x0505… |
1 |
sepc held 0x80001e10 before: an address inside hart 1’s scheduler, left by the
last interrupt hart 1 took there. CSRs belong to the hart, and they hold whatever that
hart did last.
kernel_sp is the top again, 0x3fffffe000. The next time init traps, its kernel
stack will be empty, whatever is on it now.
Step 18 of 21
A process returning from a trap gets to userret because usertrap returns there
(its ra was set by jalr in uservec). forkret has no such return address, so it
calls userret directly, through a function pointer:
| value | |
|---|---|
trampoline_userret |
0x3ffffff09c (TRAMPOLINE + userret’s offset 0x9c) |
argument satp (a0) |
0x8000000000087f52: Sv39, init’s root at 0x87f52000 |
The jump lands on the trampoline’s high address. The kernel page table maps it there
too (kernel/vm.c:47), which is what lets the very next instruction survive the
switch to init’s page table. We checked what happens if it doesn’t: with line 47
removed, the first thing to fail is this jump. The fetch at 0x3ffffff09c faults
(scause = 12), and the hart loops on faults at stvec forever (quiz, Tour 41: Every transition: mode, stack and page table).
Step 19 of 21
| Register | before line 111 | after |
|---|---|---|
satp |
0x8000000000087fff |
0x8000000000087f52 |
sp |
0x3fffffdfd0 |
0x3fffffdfd0 |
sp didn’t change, but its meaning did. init’s page table has the trampoline and
the trapframe at the top, and the program from 0 to 0x4000. Nothing at
0x3fffffdfd0. From here until sret hart 1 has no usable stack (for five instructions
this unmapped address, then, from line 118, the user’s sp), and userret uses none.
The frames still sitting on slot 0’s kernel stack (forkret’s, prepare_return’s) are
simply abandoned. Nobody pops them. The next trap reloads sp from kernel_sp, the top,
and writes over them.
sretld sp, 48(a0) in userret (kernel/trampoline.S:118) loads the user’s spStep 20 of 21
The fourth and last stack switch: ld sp, 48(a0) loads trapframe->sp, the value
kexec stored, 0x3fe0. gdb at the sret (0x3ffffff120):
| Register | Value | From |
|---|---|---|
sp |
0x3fe0 |
kernel/exec.c:137 |
a0 |
1 |
argc, returned by kexec into trapframe->a0 |
a1 |
0x3fe0 |
argv (kernel/exec.c:124) |
tp |
0x0505050505050505 |
nobody: kalloc’s junk, restored faithfully |
sepc |
0xbc |
|
sstatus |
SPP=0 SPIE=1 SIE=0 |
sp now points into init’s user stack, but the hart is still in supervisor mode,
which cannot use user pages, so until sret there is no usable stack and the code
pushes nothing. This is the mirror of uservec's line 76: the address is loaded first,
and something else (there csrw satp, here sret) makes it usable.
Every user register except sp, a0 and a1 is junk on init’s first entry. That’s
harmless: start and main set what they use. The tp junk is inherited by every
process init forks, which is why the user tp we see in Tour 43: A system call, CSR by CSR has the same
value.
sret (kernel/trampoline.S:153) returned to user mode, making the user stack usableStep 21 of 21
sp on the way from power-on to init’s first instruction, in our run. Only these
instructions move sp to another stack: kernel/entry.S:17,
kernel/swtch.S:26 (onto init’s stack on hart 0, back to hart 0’s scheduler stack
when init sleeps, and onto init’s stack again on hart 1), and
kernel/trampoline.S:118, whose user stack becomes usable only at the sret on
kernel/trampoline.S:153. The thread changed harts at sleep; the stack did not.sret sets the mode to U, SIE to 1, and the pc to 0xbc, and leaves sp alone;
the change of mode is what makes the user stack usable.
gdb, one instruction after sret: pc = 0xbc, priv = 0, sp = 0x3fe0, and the
instruction there is addi sp, sp, -16. The first thing init does in user mode is
push a frame onto the stack kexec built for it.
Almost. In a second run, without single-stepping, the first event at 0xbc was a
timer interrupt: sepc = 0xbc, scause = 0x8000000000000005. Interrupts had been
off on this thread for its entire trip through forkret, so the tick was already
pending. In user mode, supervisor interrupts are taken regardless of SIE, so it fired
before 0xbc executed even once. init went straight back through uservec onto its
empty kernel stack, then on to yield.
The figure above shows the whole journey, as sp saw it.
Tour 42 · wrap-up
| Lock | Taken in | Protects |
|---|---|---|
none (stack0) | _entry | Each hart computes its own slice from its own mhartid, so the stacks never overlap, and no lock is needed |
kmem.lock (spinlock) | kalloc, from proc_mapstacks and allocproc | The free-page list that the kernel stacks and trapframes come from |
init's p->lock (spinlock) | allocproc, the scheduler (released in forkret), sleep (released by the scheduler) | p->state, and with it the rule that only one hart at a time runs on init’s kernel stack |
block 1's buffer (sleep-lock) | bread in readsb | The buffer being filled from disk while init sleeps |
first (no lock) | forkret | One-time file-system setup; only the first process can be there |
When fork creates a child, what does the child’s kernel stack hold, and where does its sp point the first time a scheduler switches to it?
Nothing: allocproc gave the child the kernel stack of a free slot and a forged context with ra = forkret and sp = p->kstack + PGSIZE, the empty top. The child’s user stack is a copy of the parent’s (part of uvmcopy), and its trapframe is a copy too, except a0 = 0. So it leaves forkret through userret and returns from fork with 0, at the parent’s user sp.
Why is sp still valid right after kvminithart writes satp, even though every address now goes through a page table?
The boot stack is in stack0 at 0x80007890, inside the range kvmmake maps one-to-one, virtual = physical, read-write (kernel/vm.c:42). So the same sp value names the same bytes before and after. The kernel’s code is mapped the same way, which keeps the pc valid too.
Hart 1’s scheduler stack overflows by 200 bytes. What gets damaged, and when would you notice?
The writes land at the top of hart 0’s slice, just below 0x80008890: the dead start/main frames and scheduler’s frame, which is never read again because scheduler never returns and keeps its variables in registers (saved in cpus[0].context, not on the stack). So most of the damage is harmless. The danger is the bottom 72 bytes (0x800087c8–0x80008810). If hart 0 is inside acquire, release or a kernelvec frame there at that moment, it reloads a corrupted register within microseconds and crashes. Nothing detects the overflow, so you’d see a rare, timing-dependent crash.
You want a debugging function that prints which stack the current code is on. How can it decide from sp alone, and why can’t it also print the current privilege mode?
Compare sp with known ranges: inside stack0 (0x80007890–0x8000f890) means boot or scheduler stack, and the slice gives the hart. Inside KSTACK(i)'s page means slot i’s kernel stack. A low address under a user satp means a user stack. The mode is harder, because no CSR reports the current mode. sstatus.SPP is the mode before the last trap. Code that is running at all in the kernel is in S-mode, so the question only makes sense for code that runs in several modes, and there the answer has to come from context.
On init’s first run, interrupts stay off through all of forkret, fsinit and kexec, even across 19 sleeps. Why, and what is the visible consequence?
The scheduler acquired p->lock with interrupts already off, so intena was 0, and release in forkret left them off. Each sleep resumes through sched, which restores that saved intena (0). Disk interrupts are taken by harts idling in their schedulers, so nothing breaks. The consequence: a timer tick goes pending and stays pending, and in our run it fired the instant init reached user mode, before its first instruction ran.
Why did init’s first user instruction run on hart 1 in our run, not hart 0, which did all the setup?
init slept inside fsinit and kexec. Each time it was woken, whichever hart’s scheduler acquired its p->lock and found it RUNNABLE first switched to it, loading its saved sp and continuing on its kernel stack. The kernel stack, the trapframe and the context belong to the process; the hart is interchangeable. That is also why prepare_return records kernel_hartid fresh every time.
Keys: ← → step · Home start