Tour 47 · The dance of privilege · about 28 minutes · 17 steps
A process that is not running is not anywhere in the CPU. No hart holds its registers, no
pc points into its code. Yet some hart can pick it up at any moment and continue it as if
nothing had happened. So everything needed to continue it must be in memory. Where,
exactly?
The short answer is three places: the trapframe (the user registers), the kernel
stack (the kernel’s call chain), and p->context (14 registers saved by swtch).
The short answer is often followed by a tidy claim: the three never overlap, and together
they hold every register exactly once. This tour checks that claim against real memory,
dumped with gdb from this build while usertests preempt ran on three harts. It turns out
to be half right, and the half that is wrong teaches the most: some register values are
stored twice, sometimes as the same number and sometimes as different
ones, and the design is correct anyway.
Then we follow the three places through a process’s life: what fork gives a child, what
exec replaces, and what exit and wait free, and what they never free.
Best after: 7. The trampoline and the trapframe, 13. swtch and the lock handed across a context switch, 20. fork, 21. exit, wait and zombies, 22. exec, 45. One complete time slice on three harts
We stop the machine at three moments, from two runs. In the first two (usertests preempt),
hart 2’s scheduler has just come back from a swtch. In the third (usertests writebig),
it is hart 1’s. Either way, the process it left is fully suspended:
| Process | Slot | Why it stopped | Kernel stack used |
|---|---|---|---|
| pid 5, a spinning child | proc[4] |
timer, in user mode | 112 bytes |
| pid 4, the test | proc[3] |
asleep in read on a pipe |
336 bytes |
| pid 3, the shell’s child | proc[2] |
timer, in the kernel, during exec |
1920 bytes |
Meanwhile the other harts are running other processes, and the scheduler stacks are in use. Each hart runs at most one process; the suspended ones are pure data.
stack00x8000a810cpus[2] was 0 too)pid 5's p->lockStep 1 of 17
s4–s11 appear in both the trapframe and p->context. s3 matches only by
coincidence. s0–s3 also sit in prologue spills on the kernel stack.Three fields of struct proc point at the three places:
| Field | What it locates | Filled by |
|---|---|---|
trapframe |
a page holding the user registers | uservec on every trap from user mode |
kstack |
the kernel stack, KSTACK(slot) |
ordinary C function calls (and kernelvec) |
context |
14 registers of the kernel thread | swtch |
A fourth item is often forgotten: pagetable. The user’s memory, including its user
stack, is reached only through it.
We are on hart 2 at kernel/proc.c:456, just after pid 5’s swtch returned into the
scheduler. Pid 5’s p->lock is still held, so pid 5 cannot be resumed while we look:
this is exactly the moment the frozen process is complete. In this tour we dump all
three places for pid 5, and then compare.
stack0pid 5's p->lockStep 2 of 17
The trapframe is one page from kalloc, pointed to by p->trapframe and mapped at
TRAPFRAME (0x3fffffe000) in pid 5’s page table only. It is not on any stack.
gdb read pid 5’s:
| Field | Value | |
|---|---|---|
epc |
0x3a7a |
the j the tick interrupted |
ra, sp |
0x3a72, 0x11e80 |
user return address, user stack |
s0–s2 |
0x11ec0, 0, 0x8430 |
|
s3–s11 |
0, 1, 0, 2, 0x9010, 0x81e8, 0x9460, 1, 0x8220 |
|
a0–a7 |
0, 0x11def, 1, 0, 9, 0x81a1, 0x100000000000, 1 |
a7 = 1: its last system call was fork |
gp, tp, t0–t6 |
0x0505050505050505 |
never used by this program |
kernel_sp, kernel_hartid |
0x3fffff6000, 2 |
for the next trap |
0x0505050505050505 is kalloc's junk fill (kernel/kalloc.c:80). init’s first
trapframe page started as junk, userret loaded it into the registers, and since
these programs never write gp, tp or t0, t2–t6, and write t1 only in
printf’s printint (not yet called in this run), the junk has been copied through
every fork since. It is harmless, and it is a fine fingerprint.
That is the whole user state: 31 general registers (x0 is always 0) plus the pc.
Nothing else from user mode needs saving; the user’s memory stays put in its page table.
Step 3 of 17
Four writers, at four different moments:
| Fields | Written by | When |
|---|---|---|
| 31 registers | uservec, lines 40–73 |
every entry from user mode, before anything else runs |
epc |
usertrap, kernel/trap.c:52 |
right after, from sepc |
a0 (and epc += 4) |
syscall, usertrap |
system calls only: the return value, skip the ecall |
kernel_satp, kernel_sp, kernel_trap, kernel_hartid |
prepare_return |
every exit to user mode, for the next entry |
The trapframe also has readers besides the trampoline: argraw reads system-call
arguments from it, kfork copies it, kexec rewrites epc and sp.
There is no lock on it. Apart from its birth (the parent’s kfork fills a child’s
trapframe while the child is still USED and invisible) and its death (the parent’s
freeproc, after the child is a ZOMBIE), only the process’s own kernel thread and the
trampoline running on its behalf touch it, and a process runs on one hart at a time. While pid 5 is
suspended, nobody touches it at all.
yield’s acquire ran with SIE off (a timer trap from user mode): intena 0, which sched keeps in s3pid 5's p->lockStep 4 of 17
Pid 5’s KSTACK(4) page, from context.sp up to the top, word by word as gdb read it
(gcc keeps frame pointers, so each frame’s saved s0 links to the next):
address value frame meaning
0x3fffff5ff8 0x3ffffff09c usertrap ra: returns into userret (trampoline)
0x3fffff5ff0 0x11ec0 usertrap saved s0 = the USER's s0
0x3fffff5fe8 0 usertrap saved s1 = the USER's s1
0x3fffff5fe0 0x8430 usertrap saved s2 = the USER's s2
0x3fffff5fd8 0x800026d0 yield ra: back into usertrap (line 86)
0x3fffff5fd0 0x3fffff6000 yield saved s0 (usertrap's frame)
0x3fffff5fc8 0x80010370 yield saved s1 = &proc[4]
0x3fffff5fc0 2 yield (slot not written)
0x3fffff5fb8 0x80001f02 sched ra: back into yield
0x3fffff5fb0 0x3fffff5fe0 sched saved s0 (yield's frame)
0x3fffff5fa8 0x80010370 sched saved s1
0x3fffff5fa0 2 sched saved s2 = which_dev, from usertrap
0x3fffff5f98 0 sched saved s3 = the USER's s3
0x3fffff5f90 0x3fffff5fc0 sched (slot not written) ◄─ context.sp
No register file was dumped here wholesale. What is on a kernel stack is what any C
call chain leaves: return addresses, frame links, locals, and callee-saved registers
spilled by prologues. usertrap uses s0, s1 and s2, so its prologue saved the
values those registers held on entry, and on entry they held the user’s values,
because uservec never touches s registers. sched uses s3 (for intena, the
thread’s own copy of the flag, Locks and interrupt state), so
its prologue saved the user’s s3, which nothing between had touched. So four user
registers (s0–s2 spilled by usertrap, s3 by sched) are now stored twice: in the
trapframe and here.
s3 = 0 that swtch saves is this intenapid 5's p->lockStep 5 of 17
Back on hart 2 a moment earlier, inside swtch on pid 5’s stack: lines 10–23 store
ra, sp and s0–s11 into proc[4].context. gdb read the result:
p->context |
trapframe (user) | |
|---|---|---|
ra |
0x80001e9c (in sched) |
0x3a72 |
sp |
0x3fffff5f90 |
0x11e80 |
s0 |
0x3fffff5fc0 |
0x11ec0 |
s1 |
0x80010370 |
0 |
s2 |
0x8000f9a0 |
0x8430 |
s3 |
0: sched’s intena |
0: the same number, by coincidence |
s4–s11 |
1, 0, 2, 0x9010, 0x81e8, 0x9460, 1, 0x8220 |
the same eight values |
Why only these 14? Because swtch is called like an ordinary function. The calling
convention already lets a callee destroy a0–a7 and t0–t6, so sched cannot be
relying on them, and saving them would be wasted work. ra says where to continue,
sp where the stack is, and s0–s11 are exactly the registers a callee promises to
give back. gp is unused by the kernel; tp holds the hart number and deliberately
stays with the hart.
stack0pid 5's p->lockStep 6 of 17
Put the three dumps side by side for the callee-saved registers:
| Register | trapframe | kernel stack | p->context |
|---|---|---|---|
s0, s1, s2, s3 |
user values | user values again (spilled by usertrap’s and sched’s prologues) |
kernel values (s3 is sched’s intena, 0 here by coincidence) |
s4–s11 |
user values | absent | user values again |
ra, sp |
user values | absent | kernel values |
So “no overlap” is false: every user s register is stored at least twice. And “each
place holds different registers” is false too: the trapframe and p->context both have
s0–s11, ra and sp, holding the user’s values in one and the kernel
thread’s values in the other, which sometimes coincide. s4–s11 coincide because
nothing on the path usertrap → yield → sched used them, so they still held what the
user left in them when swtch saved them. s3 only looks the same: sched spilled the
user’s s3 onto the stack and reused the register for intena, which is also 0 here.
The design is still correct, because each copy has exactly one reader:
swtch restores the kernel thread from p->context;userret restores the user from the trapframe.A duplicate that only one reader ever uses can never disagree with anything. The
accurate claim is: the trapframe is the complete user state, p->context plus the
kernel stack are the complete kernel-thread state, and the two may share values.
sretld sp, 48(a0) (kernel/trampoline.S:118) loads the user’s spStep 7 of 17
Follow the user’s s registers through a trap. uservec leaves them untouched.
usertrap is an ordinary C function called with jalr, so the calling convention
obliges it to return with s0–s11 exactly as it found them, which were the user’s
values. Everything it calls obeys the same rule, swtch included. So when usertrap
returns into userret, the s registers already hold the user’s values, and the
twelve ld s0 … ld s11 among lines 124–142 write the same numbers back.
Then why load them at all? Because one path breaks the chain: every new process’s first
return, through forkret: each fork child, and init. forkret reaches userret through a function pointer, from a kernel thread that
never came in through uservec: its s registers hold kernel values (they started as
zeros in a forged context). For that thread, the trapframe copy, made from the parent’s
by kfork, is the only place the child’s user s registers exist.
This is a fact about this build’s paths, not a promise the code makes. A future kernel
that edits trapframe->s0, or reaches userret some other way, would need these loads,
and the trampoline does not try to be clever: it restores everything.
pi->lock was released on line 125 with intena 1, so SIE is on; sleep will take p->lock with SIE on, and sched’s s3 will be 1Step 8 of 17
Pid 4 called read(3, buf, 12288) on an empty pipe and slept. Its kernel stack, from
gdb’s frame-pointer walk:
KSTACK(3), top 0x3fffff8000
usertrap 0x3fffff7fe0 ra → userret; spills user s0, s1, s2 = 0x11ec0, 5, 0x8430
syscall 0x3fffff7fc0
sys_read 0x3fffff7f90
fileread 0x3fffff7f60
piperead 0x3fffff7f00 resumes at pipe.c:127, the line after sleep() (sleep's saved ra)
sleep 0x3fffff7ee0
sched 0x3fffff7eb0 ◄─ context.sp (336 bytes in all)
Its trapframe still holds the system call as the user made it: a7 = 5 (SYS_read),
a0 = 3 (the pipe’s read end), a1 = 0xccf8 (the user buffer), a2 = 0x3000, and
epc = 0x5772, already advanced past the ecall. a0 will be overwritten with the
result when read finishes.
Its context: s3 = 1 is sched’s copy of intena (interrupts were on when sleep
took p->lock; piperead’s own s3, the buffer address 0xccf8, was spilled into
sched’s frame). s4 = 0x80069218 (&pi->nread) and s5 = 0x3000 (n) are
piperead’s variables. s6–s11 again equal the user’s. The kernel stack
is the reason the process stopped, written as a call chain: a stack that ends in
piperead → sleep means “waiting for a pipe”.
pop_off had just dropped noff to 0 and, with intena 1, turned SIE on: the pending tick was taken at once. The inode’s sleep-lock is not countedroot directory inode (sleep-lock)Step 9 of 17
A third case, from a different moment (usertests writebig): the shell had forked
pid 3 to run usertests, and pid 3 was inside kexec, looking up the path, when hart
1’s timer fired in the kernel. The trap went through kernelvec, which pushed 256
bytes onto whatever stack was current, pid 3’s own, and kerneltrap called
yield. 1920 bytes of KSTACK(2) are in use:
usertrap → syscall → sys_exec → kexec → namei → namex → dirlookup
→ readi → brelse → releasesleep → release → pop_off
(interrupted here, just after intr_on)
→ [kernelvec frame, 256 bytes] → kerneltrap → yield → sched
Part of the 256-byte frame, as gdb read it:
| Offset | Slot | Value |
|---|---|---|
| 0 | ra |
0x80000c34 (inside pop_off) |
| 8 | sp (not saved) |
0x0505050505050505 |
| 24 | tp (not saved) |
0x0505050505050505 |
| 72–128 | a0–a7 |
0x8000fa50, 0x800199b0, 0x10, … |
| 56, 64, 136–208 | s0–s11 (not saved) |
leftovers |
Only gp and the 16 caller-saved registers (ra, t0–t6, a0–a7) are saved;
s registers are left to
kerneltrap’s own prologue, as for any C function. The unwritten slots still hold
junk: kernel stack pages come from kalloc at boot, filled with 0x05, and nobody
clears them.
yield will acquire with SIE off, so intena 0root directory inode (sleep-lock)Step 10 of 17
The kernelvec frame has no slot for sepc. Yet pid 3 must resume exactly where
pop_off was interrupted. Where is that address?
kerneltrap reads it into a local variable, and in this build gcc keeps the local in
s2 (sstatus in s1, scause in s3). Those are callee-saved registers, so when
kerneltrap calls yield, and yield calls sched, each prologue that wants those
registers spills them. gdb found them in sched’s frame:
| Address | Value | |
|---|---|---|
0x3fffff9890 |
0x80000c50 |
saved s2: sepc, pop_off just after intr_on() |
0x3fffff9888 |
0x8000000000000005 |
saved s3: scause, the timer |
0x3fffff98b8 |
0x200000120 |
saved s1 in yield’s frame: sstatus with SPP = 1, SPIE = 1 |
So the interrupted kernel pc is stored on the kernel stack by an ordinary function
prologue, in sched’s frame, three frames below the kernelvec frame (under
kerneltrap’s and yield’s). When pid 3 resumes (perhaps on another hart),
the epilogues put it back in s2, and kerneltrap writes it to that hart’s sepc
(kernel/trap.c:162) just before kernelvec’s sret. That is why kerneltrap keeps
these values in locals instead of re-reading the CSRs: by the time yield returns, the
hart’s sepc belongs to whatever ran in between (Tour 44: One interrupt, three landing sites).
This pc, 0x80000c50, also shows why the trap came here: the timer was already
pending, and fired the instant pop_off re-enabled interrupts. A tick can be taken
only where noff is 0, since SIE is on only then: pop_off turns it on only after
noff has dropped to 0, so the tick never lands inside a critical section
(Locks and interrupt state). Outside one it can land anywhere, even mid-work in
memset or uartwrite (Tour 44: One interrupt, three landing sites); here it was simply already waiting. The root
inode’s sleep-lock, still held, doesn’t count in noff.
allocproc took the child’s lock inside fork, with SIE on: intena 1pid 5's p->lockStep 11 of 17
A new process gets its three places from two functions. allocproc makes the
context: all zeros, except ra = forkret and sp = p->kstack + PGSIZE. When hart 2’s
scheduler first picked child pid 5 in our run, gdb showed exactly that:
proc[4].context: ra=0x8000193a (forkret) sp=0x3fffff6000 s0..s11 = 0
kernel stack used: 0 bytes
The kernel stack itself was not allocated here. KSTACK(4) was allocated and mapped at
boot by proc_mapstacks and belongs to the slot, so the child simply inherits an
empty page that previous occupants of proc[4] once used. Nothing of the parent’s
kernel stack is copied: the child will never return through sys_fork or syscall.
The trapframe page is new: kalloc at kernel/proc.c:129, full of 0x05 junk
until kfork fills it.
child's p->lock (pid 5)Step 12 of 17
kfork copies the parent’s whole trapframe, then sets the child’s a0 = 0. The
child’s user stack was copied by uvmcopy just before, page by page, so the copied
sp = 0x11e80 points at an identical stack in the child’s own memory. gdb, at the
child’s first scheduling:
| Child trapframe field | Value | |
|---|---|---|
epc |
0x5752 |
after fork’s ecall, same as the parent |
a0 |
0 | the only user register that differs |
sp, s0–s11 |
the parent’s | |
kernel_sp |
0x3fffff8000 |
the parent’s kernel stack top! |
kernel_hartid |
2 | where the parent last left user mode |
The kernel fields are stale, and that is fine: the child cannot trap before it has
returned to user mode once, and the only way there is forkret →
prepare_return, which rewrites all four (kernel/trap.c:116–119). One line, at
the right moment, makes a whole copied structure correct.
Line 279 runs while kfork still holds the child’s p->lock from allocproc
(released at line 294), and the child is still USED: no other hart can see it yet.
ld sp, 8(a1) in swtch loaded the forged spforkret’s release of the scheduler’s lock: intena was 0, so SIE stays offStep 13 of 17
swtch “returns” to forkret with sp at the very top of KSTACK(4). forkret
releases the lock the scheduler took (with intena 0, recorded by the scheduler, so
interrupts stay off: Locks and interrupt state), calls prepare_return, and jumps to userret
through a pointer. Its frame is never popped: once sret happens, the next trap starts
again from kernel_sp, the top, and writes over it.
Now count the child’s three places at the moment it reaches user mode:
a0 = 0, fresh kernel fields;p->context: the forged values, never used again; the next swtch overwrites them.The child’s user s registers come only from the trapframe, through userret’s loads.
That is the path that makes those loads necessary.
Step 14 of 17
kexec runs on the process’s kernel stack (the 1920-byte dump above came from inside
it) and builds a new page table with a new user stack, while the old ones stay
intact for a possible failure (Tour 22: exec). At the commit point:
| Place | What exec does |
|---|---|
| user memory and user stack | replaced: new page table installed, old one freed at line 138 |
| trapframe | same page; epc = the ELF entry, sp = the new stack; a1 = argv (line 124); the other 28 registers left as they were |
| kernel stack | the same page, still holding kexec’s own frames |
p->context |
untouched |
The return value, argc, reaches a0 through syscall like any result. The new
program starts with whatever s and t registers the old one had, a reason exec’d
programs must not assume registers start at zero.
kexit took wait_lock (noff 1, intena 0: it came from a timer trap), then its own p->lock (noff 2), and released wait_lock at line 360, leaving the 1 that sched demandspid 5's p->lockStep 15 of 17
Pid 5 was killed. At its next timer trap, usertrap saw killed and called
kexit directly, on pid 5’s kernel stack. kexit marked it ZOMBIE and called
sched, which never returns. gdb, afterwards:
proc[4] pid 5 ZOMBIE context: ra=0x80001e9c sp=0x3fffff5f80 s4=0xffffffffffffffff
frames: usertrap → kexit → sched (128 bytes)
s4 = -1 is kexit’s status argument, kept in a callee-saved register and saved by
swtch into a context that nobody will ever load. (Pid 4, which exited normally, left
usertrap → syscall → sys_exit → kexit → sched, 192 bytes.)
kexit holds two spinlocks for a moment, wait_lock and its own p->lock, and
releases wait_lock (line 360) before calling sched, which would panic with
“sched locks” at noff 2 (Locks and interrupt state).
The zombie cannot free its own kernel stack: it is standing on it until swtch
finishes, and it still holds p->lock. Its hart’s scheduler releases that lock only
after sp has left this page. Its trapframe and user memory remain too, until the
parent collects it.
wait_lock and the child’s p->lock, taken inside wait with SIE on; kfree on line 159 adds kmem.lock for a moment: noff 3, the deepest nesting measured (Locks and interrupt state)wait_lockpid 5's p->lockStep 16 of 17
The parent’s kwait finds the zombie under wait_lock and the child’s p->lock, and
calls freeproc:
| Place | Fate |
|---|---|
| trapframe page | kfree (line 159) |
| user page table, user memory, user stack | proc_freepagetable (line 162) |
| kernel stack | not freed: it belongs to slot 4, mapped since boot |
p->context |
not cleared: still says sp = 0x3fffff5f80, until the next allocproc overwrites it |
Then state = UNUSED, and slot 4, with its kernel stack, is free for the next
process. When the next fork lands in proc[4], its forged context will start at the
top of the same page, and the zombie’s old frames will be overwritten as the new
process’s call chains grow down over them.
This is why a design where each process owned a freshly allocated kernel stack would
need extra care: something other than the dying process would have to free it, after
the last swtch off it. xv6 avoids the question by tying stacks to slots.
stack0Step 17 of 17
Everything that makes a process resumable, with its size in our dumps and its owner:
| Place | Holds | Size here | Allocated | Freed |
|---|---|---|---|---|
| trapframe | user x1–x31, pc, 4 kernel fields |
288 bytes in a 4 KiB page | allocproc |
freeproc |
| kernel stack | the kernel’s call chain, spills, at most one kernelvec frame per trap |
112 / 336 / 1920 bytes of 4 KiB | at boot, per slot | never |
p->context |
ra, sp, s0–s11 of the kernel thread |
112 bytes | forged by allocproc, saved by swtch |
never cleared |
| page table | user memory, user stack | a few pages | fork, exec |
freeproc, exec |
And what is not in any of them, because it belongs to a hart: tp, the scheduler
stack, struct cpu, stvec, stimecmp. A process carries nothing hart-specific, which
is precisely what lets it resume on any hart.
The three places are not three disjoint thirds of one register file. They are the saved state of two threads of control, the user program and its kernel thread, plus the kernel thread’s memory. Overlaps between them are harmless because each copy has a single reader.
Tour 47 · wrap-up
| Lock | Taken in | Protects |
|---|---|---|
p->lock | yield / sleep → sched → released by the scheduler; kwait; allocproc | The moment of suspension: no hart may load p->context until swtch has saved it and left the stack |
wait_lock | kwait, kexit, kfork | p->parent; who may free a zombie’s trapframe and page table |
(no lock) the trapframe | uservec, usertrap, prepare_return, kfork | Nothing needed: apart from its birth (the parent’s kfork, while the child is USED) and its death (the parent’s freeproc, after ZOMBIE), only the process’s own thread touches it, and a process runs on one hart at a time |
(no lock) the kernel stack | every kernel function | Nothing needed: exactly one thread uses it, and p->lock covers the hand-over between harts |
A textbook table says: trapframe = user registers, kernel stack = caller-saved registers, context = callee-saved registers, “no overlap”. Using pid 5’s dump, give one register value stored in two of the places, and one register that has different values in two places.
User s7 = 0x9010 is in both the trapframe and p->context (also s4–s11; s3 matches only by coincidence, since sched reuses it for intena). User s2 = 0x8430 is in the trapframe and in usertrap’s frame on the kernel stack. s2 itself has different values in the trapframe (0x8430, the user’s) and the context (0x8000f9a0, the kernel thread’s).
Why is it safe that the same register has two different saved values?
Each copy has exactly one reader. swtch reads only p->context (the kernel thread’s value), epilogues read only their own spills, and userret reads only the trapframe (the user’s value). The two copies describe two different threads of control and are never compared.
Explain why userret’s loads of s0–s11 usually write back values the registers already hold, and name the case where they are essential.
usertrap is a C function reached from uservec, which never touches s registers, so by the calling convention it returns with the user’s s values intact. The loads matter on one path: every new process’s first return, through forkret (each fork child, and init). forkret reaches userret from a forged kernel thread whose s registers are not the user’s; only the copied trapframe has them.
Where is the pc at which a process was interrupted kept, (a) if it was in user mode, (b) if it was in the kernel?
(a) In trapframe->epc, copied from sepc by usertrap. (b) In kerneltrap’s local variable sepc, which in this build lives in s2 and is spilled onto the kernel stack by the prologue of a function it calls (sched in our dump); kerneltrap writes it back to the hart’s sepc before sret.
A child’s trapframe, copied from the parent, says kernel_sp = 0x3fffff8000, the parent’s stack. What stops the child from trapping onto its parent’s kernel stack?
The child can only enter user mode through forkret, which calls prepare_return first, and that rewrites kernel_sp (and the other kernel fields) for the child. A trap can only happen from user mode, so the stale value is never used.
After wait frees a zombie, its p->context still holds a saved sp. Why is that harmless, and what would make it dangerous?
No scheduler loads a context unless the process is RUNNABLE, and an UNUSED slot becomes runnable only after allocproc forges a fresh context. It would be dangerous if some code could set an UNUSED or USED slot RUNNABLE without going through allocproc.
Keys: ← → step · Home start