xv6, line by line
concepts

Concept 1

The stacks of xv6

Every C function needs a stack: memory for its local variables, the registers it must preserve, and the address to return to. The register sp points at the current top of that memory, and every call moves it further down. A hart (one CPU core) has exactly one sp, so, apart from a few instructions at power-on and two short stretches in the trampoline (described below), at any instant it is running on exactly one stack: whichever memory sp points into. We call that the hart’s active stack.

This page answers two questions for every moment of xv6’s life: which stack is active, and what changed it? xv6 has three kinds of stack, plus a few instructions at power-on and two short stretches in which sp points at nothing the running code may use:

Kind How many Size Virtual address Created by Guard page
user stack one per process 1 page just above the program in its user page table (0x3000–0x3fff for init) kexec yes: mapped, but PTE_U cleared
kernel stack one per proc slot: 64 1 page KSTACK(slot), high in the kernel page table (0x3fffffd000 for slot 0) proc_mapstacks, once, at boot yes: unmapped
boot stack, which becomes the scheduler stack one per hart 4096 bytes a slice of stack0 at 0x80007890 + 4096 × hart (virtual = physical) the build: stack0 is a global array in .bss none

The tour strip at the top of every tour step names the active stack with the same words, and lines in the source view that move sp to another stack, or change a stack’s role, carry a ⇄ mark in the line-number gutter. The companion page Mode, stack and page table: the master question adds the other two things you should always be able to name: the privilege mode and the page table.

One hart's sp moving between stacks Five stacks side by side: process A's user stack, A's kernel stack, the hart's scheduler stack, process B's kernel stack and B's user stack. The stack pointer moves from A's user stack to A's kernel stack (uservec), to the scheduler stack (swtch called by sched), to B's kernel stack (swtch called by scheduler), and to B's user stack (userret). high addresses at the top; stacks grow down A's user stack user page table U-mode main f g (frames illustrative) A's kernel stack KSTACK(A's slot) S-mode usertrap yield sched empty in user mode scheduler stack this hart's slice of stack0 S-mode start main scheduler the same three frames since boot B's kernel stack KSTACK(B's slot) S-mode usertrap yield sched B was preempted earlier; its frames waited here B's user stack B's user page table U-mode main h (frames illustrative) 1 2 3 4 1 · uservec (trap) ld sp, 8(a0) trampoline.S:76 2 · swtch from sched ld sp, 8(a1) swtch.S:26 3 · swtch from scheduler ld sp, 8(a1) swtch.S:26 4 · userret ld sp, 48(a0) trampoline.S:118 sp A runs in user mode. sp points into A's user stack. Timer interrupt. uservec saves A's sp in the trapframe and loads the top of A's empty kernel stack. usertrap, then yield, then sched: ordinary C calls push frames onto A's kernel stack. swtch loads sp from cpu->context: the hart is on its scheduler stack. A's frames stay behind. The scheduler picks B. swtch loads B's saved sp; sched, yield and usertrap return, popping B's frames. userret loads B's user sp from B's trapframe; sret returns to user mode, making it usable. One round uses three of the four stack-switching instructions (swtch's twice); the fourth runs once, at power-on. One hart. A and B are any two processes; both kernel stacks belong to their proc slots.
One hart’s sp on its way from one process to another. Process A is interrupted by a timer while in user mode, gives up the CPU, and the hart’s scheduler resumes process B, which had been interrupted the same way earlier. Only three instructions move sp between stacks in this loop; every other change to sp is a function call or return within one stack.

Four instructions move sp to another stack

Function calls move sp all the time, but only within one stack: a function’s first instruction is usually addi sp, sp, -N and its last ones undo it. Moving sp to a different stack is rare. In all of xv6 it happens at exactly four instructions:

# Instruction Where From → to When
1 add sp, sp, a0 kernel/entry.S:17 in _entry nothing → boot stack once per hart, at power-on
2 ld sp, 8(a0) kernel/trampoline.S:76 in uservec user stack → the process’s kernel stack (usable once csrw satp on line 92 installs the kernel page table) every trap from user mode
3 ld sp, 8(a1) kernel/swtch.S:26 in swtch kernel stack → scheduler stack (called from sched), or scheduler stack → kernel stack (called from scheduler) twice per context switch
4 ld sp, 48(a0) kernel/trampoline.S:118 in userret kernel stack → user stack (usable once sret, line 153, returns to user mode) every return to user mode

(Strictly, kernel/entry.S:12, la sp, stack0, writes sp too, as the first step of switch 1’s calculation: it puts stack0’s base address in sp, and line 17 adds the hart’s offset. Only after line 17 does sp point at the top of a usable slice.)

Instructions 2 to 4 load sp from memory, so some earlier code must have written the right value there. Knowing who wrote it is half of understanding each switch:

Some things look like stack switches and are not:

Where the stacks live

Where the stacks live Three columns. Left: init's user address space, with the user stack page at 0x3000 and a guard page at 0x2000. Middle: physical RAM, with stack0 inside the kernel's bss, the 64 kernel-stack pages side by side near the top of RAM, and init's user stack page. Right: the kernel page table, with the trampoline, kernel stacks at KSTACK(0) to KSTACK(63) separated by unmapped guard pages, and the direct map of all RAM. init's user page table virtual addresses (one process) physical RAM 128 MiB from 0x80000000 kernel page table virtual addresses (shared by all harts) 0x3ffffff000 trampoline (code) 0x3fffffe000 trapframe (save area) unmapped (heap would grow up into here) 0x3000 user stack, 1 page "/init" and argv[] at the top first sp = 0x3fe0 0x2000 guard page mapped, but PTE_U cleared 0x1000 data 0x0 text init's size: sz = 0x4000 No kernel stack appears here: user page tables never map one. 0x88000000 (PHYSTOP) page tables, other pages 64 kernel-stack pages, side by side no gaps here: guards exist only virtually top: 0x87f99000 = slot 0 bottom: 0x87f5a000 = slot 63 free and allocated pages 0x87f4a000 init's stack page free and allocated pages (at boot, kalloc hands out the top first) 0x80020bb0 (end) rest of .bss: cons, kmem, cpus, proc[] ... stack0: 0x80007890–0x8000f890 slices 3–7: unused with 3 harts 0x80009890–0x8000a890 hart 2 0x80008890–0x80009890 hart 1 0x80007890–0x80008890 hart 0 ticks, initproc, kernel_pagetable ... .rodata and .data from 0x80007000 0x80000000 kernel text 0x3ffffff000 trampoline 0x3fffffe000 unmapped 0x3fffffd000 KSTACK(0) 0x3fffffc000 guard (unmapped) 0x3fffffb000 KSTACK(1) 0x3fffffa000 guard (unmapped) ⋮ every slot: one stack page, one guard 0x3ffff7f000 KSTACK(63) 0x3ffff7e000 guard (unmapped) unmapped direct map: virtual = physical 0x80000000 – 0x88000000 every RAM page again, including the kernel-stack pages (no guards) stack0, at its own address boot and scheduler stacks no guard pages (R W) kernel text (R X) below it: UART, virtio, PLIC registers, also direct-mapped Not to scale. stack0 and the KSTACK addresses are fixed by this build (kernel.sym, memlayout.h). Physical pages handed out by kalloc are from one boot of this build under QEMU; they can differ.
Where the memory for each kind of stack comes from, and the virtual addresses sp holds while using it. The kernel-stack pages are ordinary free pages; only their virtual addresses have gaps between them. stack0 is part of the kernel’s .bss and has no gaps at all.

Three points in this map matter for the rest of the page:

The boot stack

When QEMU starts, every hart runs a few instructions in QEMU’s boot ROM at 0x1000 and then jumps to 0x80000000 (kernel/entry.S), in machine mode with paging off, and with nothing useful in sp. C code cannot run yet: the first function call would store its return address through a garbage pointer. stack0 provides the memory: one array of 4096 * NCPU bytes declared in kernel/start.c:11. NCPU is 8, so the array is 32 KiB, one 4096-byte slice per hart, and the build places it at 0x80007890 (kernel/kernel.sym). The Makefile starts QEMU with 3 harts (CPUS := 3), so only three slices are ever used:

stack0 in this build
  0x8000f890   end of stack0
               slices 3 to 7: never used with 3 harts
  0x8000a890   top of hart 2's slice  <- hart 2's first sp
  0x80009890   top of hart 1's slice  <- hart 1's first sp
  0x80008890   top of hart 0's slice  <- hart 0's first sp
  0x80007890   stack0 (start of hart 0's slice)

Switch number 1 (kernel/entry.S:17) computes sp = stack0 + (hartid + 1) * 4096, the top of the hart’s own slice (stacks grow down, so the top is where they start). All the harts run this line at about the same time, and each gets a different answer.

The boot stack then carries the hart through three stages without another switch:

  1. _entry and start run in machine mode with paging off. 0x80008890 is a physical address.
  2. mret (kernel/start.c:51) drops to supervisor mode and jumps to main. It does not touch sp.
  3. In main, kvminithart turns paging on (kernel/main.c:21 on hart 0, kernel/main.c:39 on the others). The next instruction still uses the same sp, and it still works, because the kernel page table maps 0x80008890 to physical 0x80008890.

Hart 0 does almost all the work of booting on this stack: every initialization call in main, from consoleinit to userinit, runs on it and returns. The other harts wait for hart 0’s signal and make a few short calls of their own.

Two frames are never popped. start never returns (it leaves with mret), so its 16-byte frame stays at the top of the slice; main never returns either, so its 16-byte frame stays below that. With gdb you can see them on every hart for as long as the machine runs.

The scheduler stack

kernel/main.c:44 is an ordinary call: scheduler();. No instruction moves sp to another stack. scheduler's frame goes right below main's, on the same slice, and because scheduler never returns, that slice is from now on this hart’s scheduler stack. It is the same memory as the boot stack with a new role. xv6 allocates nothing else for it, and it is not inside struct cpu.

In this build scheduler’s frame is 96 bytes, so inside the loop sp sits at a fixed spot:

hart 0's scheduler stack (its slice of stack0)
  0x80008890   top
               start       16 bytes   left over from boot
               main        16 bytes   left over from boot
               scheduler   96 bytes
  0x80008810   <- sp in the scheduler loop; saved in cpus[0].context.sp
               about 3.9 KiB unused
  0x80007890   bottom, with no guard page below

(On hart 1 the loop’s sp is 0x80009810, on hart 2 0x8000a810. All three were read with gdb from a running system.)

Leaving it. When the loop finds a RUNNABLE process, it calls swtch(&c->context, &p->context) (kernel/proc.c:453). swtch stores ra, sp (0x80008810) and s0–s11 into cpus[0].context, then switch number 3 loads the process’s saved sp. The scheduler’s frame stays where it is, untouched, for as long as the process runs.

Coming back. When the process gives up the hart, sched calls swtch(&p->context, &mycpu()->context) (kernel/proc.c:495). Switch number 3 loads 0x80008810 back, and swtch’s ret lands just after line 453, as if the call there had returned normally.

Three facts follow from this:

Interrupts on the scheduler stack. Each pass of the loop opens a short window with interrupts enabled (kernel/proc.c:441–442). An interrupt that arrives there, including one that woke the hart from wfi, lands on the scheduler stack: kernelvec pushes its frame right below scheduler’s. Caught with gdb on hart 0: sp was 0x80008710 (256 bytes below 0x80008810) when kerneltrap started, and sepc pointed at the instruction that turns interrupts back off. There is no current process at that moment, so kerneltrap does not call yield (kernel/trap.c:157): the frame is popped and the loop continues.

Kernel stacks

A process needs a stack for the kernel code that runs on its behalf: system calls, page faults, interrupts that arrive while it runs, and the work of being suspended and resumed. That is its kernel stack. The user stack cannot serve. User code controls sp and may have pointed it anywhere, and the program could read or change anything the kernel stored in its memory.

Where they are. KSTACK in kernel/memlayout.h places stack i at TRAMPOLINE - (i + 1) * 2 * PGSIZE. With TRAMPOLINE = 0x3ffffff000:

kernel page table, near the top of the address space
  0x3ffffff000   trampoline
  0x3fffffe000   (unmapped)
  0x3fffffd000   KSTACK(0)    slot 0's stack; its top is 0x3fffffe000
  0x3fffffc000   (unmapped)   guard page for slot 0
  0x3fffffb000   KSTACK(1)    top 0x3fffffc000
  0x3fffffa000   (unmapped)   guard page for slot 1
   ...
  0x3ffff7f000   KSTACK(63)   top 0x3ffff80000
  0x3ffff7e000   (unmapped)   guard page for slot 63

Who creates them. proc_mapstacks, called once by kvmmake on hart 0’s boot stack, allocates 64 pages with kalloc and maps each at KSTACK(i), readable and writable, not executable, not accessible from user mode. In one boot of this build the pages were 0x87f99000 (slot 0) down to 0x87f5a000 (slot 63): 64 physically adjacent pages. The guard gaps exist only among the virtual addresses. (The direct map also maps each of these pages at its physical address, with no guard. xv6 never uses those addresses as sp.)

Who owns them. Slots, not processes. procinit sets p->kstack = KSTACK(i) once (kernel/proc.c:57), and nothing changes it again: allocproc and freeproc never touch kstack, and the pages are never freed. When the process in slot 3 exits and a later process gets slot 3, the new process runs on the same page. Whatever the old one left there is stale, and harmless: the new process starts at the top and never reads below its own sp.

Who switches to them.

What lives on one. During a system call, the frames of the kernel’s C call chain, starting with usertrap. Here is the shell, read with gdb while it waited for a key:

sh (pid 2, slot 1), asleep inside read(): its kernel stack
  0x3fffffc000   top; trapframe->kernel_sp points here
                 usertrap       32 bytes   saved ra = 0x3ffffff09c (userret)
                 syscall        32 bytes
                 sys_read       48 bytes
                 fileread       48 bytes
                 consoleread    96 bytes
                 sleep          32 bytes
                 sched          48 bytes
  0x3fffffbeb0   <- p->context.sp, saved by swtch
                 3760 bytes unused
  0x3fffffb000   bottom; guard page below

Notice usertrap’s saved return address. uservec calls usertrap with jalr t0 (kernel/trampoline.S:98), which puts the address of the next instruction in ra. That next instruction is userret, at 0x3ffffff09c in the trampoline page, so when usertrap returns it falls straight into the return-to-user path.

Most chains are this shallow. The deepest ones go through exec: sys_exec keeps the path and an array of MAXARG argument pointers in its frame (480 bytes in this build), and kexec another 32-entry array, ustack (544 bytes). In the interrupt example below, the stack was nearly 1.9 KiB deep when the process gave up the hart.

User stacks

Who creates one. kexec, on every successful exec (kernel/exec.c:89–97). After loading the program’s segments it rounds the size up to a page boundary and allocates USERSTACK + 1 = 2 more pages (USERSTACK is 1). uvmclear clears PTE_U on the lower one, which becomes the guard page; the upper one is the stack. For init that gives:

init's user stack right after kexec("/init")
  0x4000   top (= the process size, sz)
  0x3ff0   "/init\0"                    each string starts on a multiple of 16
  0x3fe0   argv[0] = 0x3ff0, argv[1] = 0
           <- sp = 0x3fe0; also in a1, and argc = 1 in a0
  0x3000   bottom of the stack page
  0x2000   guard page (PTE_U cleared)
  0x1000   data
  0x0000   text

The physical page is whatever kalloc returned (0x87f4a000 for init in one boot).

Building it is not using it. While kexec fills in the new stack, the hart is running on the process’s kernel stack. The sp in kexec is a C variable, not the register, and the arguments reach the new page through copyout and the new page table. Even the array ustack (line 32) is a local variable on the kernel stack; it is copied to the user stack at line 117. Nothing changes for the hart until kexec stores the final value in trapframe->sp (kernel/exec.c:137), and even then the register gets it only on the way out.

Who switches to it. Only switch number 4, userret's ld sp, 48(a0) (kernel/trampoline.S:118), moves sp to the user stack’s address. From there until sret (kernel/trampoline.S:153) the hart is still in supervisor mode with the user page table, so it cannot use that stack yet: supervisor code cannot access user pages (sstatus.SUM is never set in this tree). The code only loads registers and pushes nothing, and tour strips say “no usable stack” there. sret returns to user mode, and from then on the program uses the stack as it likes. The kernel never relies on what the program does with sp: uservec saves whatever value it finds and never stores anything through it.

What protects it. The stack is one page and does not grow. A program that needs more runs into the guard page. The page is mapped, but without PTE_U a user-mode access faults; usertrap asks vmfault whether this is a lazily allocated heap page, vmfault refuses because the page is already mapped, and the process is killed. A single frame bigger than a page could jump over the guard into the program’s data; the guard only catches gradual growth.

fork and exit. uvmcopy copies every page below sz into the child, the stack and the guard page included, with their flags, so the child’s guard keeps PTE_U cleared. The child’s saved sp is the parent’s, because the trapframe is copied too: the same virtual address, in the child’s own copy. A user stack is freed with the rest of the user memory, by freeproc when the parent’s kwait collects the child.

No usable stack

Twice on every round trip through the kernel, for a short stretch of the trampoline page, sp holds a value that the code running may not use. Both stretches exist because the hardware switches privilege mode at a trap or sret, while sp and satp must be switched by separate instructions. Each stretch has two halves, split by the instruction that moves sp.

  1. uservec, from its first instruction (line 22) to csrw satp at line 92. A trap from user mode changes the mode and the program counter, not sp and not satp.

    • Line 22 to line 76: the hart is in supervisor mode with the user’s sp and the user’s page table, and supervisor code cannot use user pages. uservec pushes nothing: it frees a0 by moving it into sscratch (kernel/trampoline.S:32) and saves every user register into the trapframe through a0.
    • Line 76 to line 92: switch number 2 loads the kernel-stack address, but the user page table is still installed and does not map kernel stacks, so nothing may be pushed until csrw satp switches to the kernel page table at kernel/trampoline.S:92.

    Tour strips say “no usable stack” for this whole stretch and “kernel stack” from line 92 on.

  2. userret, from csrw satp (kernel/trampoline.S:111) to sret (kernel/trampoline.S:153).

    • Line 111 to line 118: the user page table is installed, and sp still holds a kernel-stack address it does not map. The code uses no stack.
    • Line 118 to line 153: switch number 4 has put the user’s stack pointer in sp, but the hart is still in supervisor mode, which cannot use user pages, the same situation as the first half of uservec. The code keeps loading registers from the trapframe and pushes nothing.

    sret changes the mode to user, and only then is the user stack usable.

The second stretch mirrors the first: on the way in, the trap makes the user stack unusable (line 22), line 76 loads the kernel-stack address, and line 92 makes it usable; on the way out, line 111 makes the kernel stack unusable, line 118 loads the user’s sp, and sret (line 153) makes it usable.

Notice the order. On the way in, sp changes before satp; on the way out, satp changes before sp. So the user’s sp value is only ever paired with the user page table, and C code only ever runs with a kernel sp and the kernel page table. The mismatched pair in between is covered by a few lines of assembly that touch only registers and the trapframe. Interrupts are off in both stretches: the trap cleared SIE on the way in, and prepare_return turned them off (kernel/trap.c:108) before the way out.

Interrupts borrow the current stack

A trap taken while the hart is already in supervisor mode goes to kernelvec: an interrupt during a system call, or in the scheduler’s interrupt window. (Traps from user mode go to uservec instead and land on the empty kernel stack.) kernelvec has no stack of its own. Its first instruction, addi sp, sp, -256 (kernel/kernelvec.S:14), makes room on the stack that is already active, and it saves 17 registers there: ra, gp, t0–t2, a0–a7 and t3–t6. That block is the kernelvec frame. s0–s11 are left out because kerneltrap, being C, preserves them; sp because the matching addi restores it; tp for a reason we come to in a moment.

Depending on what was interrupted, the frame lands on one of two stacks:

Yielding with a frame in the middle. If the interrupt was the timer and the hart has a current process, kerneltrap calls yield (kernel/trap.c:157), and the process is suspended with the kernelvec frame in the middle of its kernel stack, exactly as in the picture. Later, some hart’s scheduler resumes it: swtch returns into sched, yield returns, kerneltrap puts back the sepc and sstatus it saved on entry (kernel/trap.c:162), because other traps in the meantime will have overwritten those registers, and kernelvec pops its frame and srets into the interrupted code.

Why tp is not restored. The process may be resumed by a different hart. Its kernel stack, and the kernelvec frame on it, go wherever the process goes, but tp holds the ID of the hart, and swtch deliberately leaves it alone. When the process resumes on hart 2, tp already says 2. Restoring the value saved on hart 0 would make cpuid wrong from then on, which is the point of the comment at kernel/kernelvec.S:44. (For the same reason kernelvec does not save it.)

Never more than one. The trap clears SIE, so kernelvec and kerneltrap run with interrupts off. kerneltrap never turns them on, and a process that yields from it resumes with them still off (the lock yield took was acquired with interrupts off, so releasing it does not turn them on) until kernelvec’s sret restores the previous state. So in normal operation interrupt handlers do not nest, and a stack holds at most one kernelvec frame. Exceptions in kernel code are not handled at all: kerneltrap panics.

Save areas are not stacks

Four places hold registers while their owner is not running. None of them is a stack: no frames, no pushing or popping, just fixed slots that are overwritten each time. The glossary calls them save areas.

Save area Where it lives Written by Read by Holds
trapframe its own page per process, from allocproc; at TRAPFRAME (0x3fffffe000) in the user page table, and at its physical address in the kernel uservec (user registers), usertrap (epc), prepare_return (kernel fields), syscall, kfork, kexec userret, uservec (kernel fields), syscall (arguments) 31 user registers, the user pc, and four values uservec needs
p->context inside struct proc, in the proc array swtch called from sched; allocproc for a new process swtch called from scheduler ra, sp and s0–s11 of a suspended kernel thread
cpus[i].context inside struct cpu, in the cpus array swtch called from scheduler swtch called from sched the same 14 registers of hart i’s scheduler
sscratch a CSR (a special register), one per hart kernel/trampoline.S:32 kernel/trampoline.S:72 the user’s a0, for about 35 instructions

Two of these contain a value of sp: trapframe->sp points into the user stack, context.sp into a kernel or scheduler stack. Holding a pointer into a stack does not make something a stack.

One coincidence in the numbers: TRAPFRAME is 0x3fffffe000, which is also the top of KSTACK(0). They never meet. The trapframe is mapped there only in user page tables, the kernel stack only in the kernel page table, and the top of a stack is one past its last byte: the first push goes to 0x3fffffdff8.

Where a suspended process lives

Put the pieces together for a process that is not running, say the shell waiting for a key:

Where a sleeping shell lives The shell, pid 2 in proc slot 1, asleep inside read. Left: its struct proc, whose context.sp points at 0x3fffffbeb0 in its kernel stack. Middle: the kernel stack page KSTACK(1), holding the frames usertrap, syscall, sys_read, fileread, consoleread, sleep and sched, with the rest of the page unused and an unmapped guard page below. Right: the trapframe page, whose kernel_sp points at the top of the kernel stack and whose sp field points into the user stack at 0x4f10. proc[1]: struct proc in the proc table (kernel .bss) kernel stack KSTACK(1) one page in the kernel page table trapframe page TRAPFRAME in sh's page table state = SLEEPING pid = 2, name = "sh" chan = &cons.r kstack = 0x3fffffb000 pagetable = sh's user table trapframe = its page (PA) context: saved by swtch ra = sched+114 sp = 0x3fffffbeb0 s0 … s11 Not a stack: 14 registers. sp is a pointer into the stack. 0x3fffffc000 top (empty stack starts here) usertrap32 B syscall32 B sys_read48 B fileread48 B consoleread96 B sleep32 B sched48 B 0x3fffffbeb0 unused: 3760 of 4096 bytes (swtch has no frame of its own) 0x3fffffb000 bottom of the page guard page, unmapped 0x3fffffa000 kernel_satp = kernel table kernel_sp = 0x3fffffc000 kernel_trap = usertrap epc = after the ecall kernel_hartid = last hart sp = 0x4f10 a0, a1, a2 = 0, &c, 1 a7 = 5 (SYS_read) … all 31 user registers A save area, not a stack. 0x5000 sh's user stack argv for sh start, main getcmd, gets 0x4f10 0x4000 guard page, PTE_U cleared 0x3000 swtch usertrap's saved return address is 0x3ffffff09c: userret, in the trampoline. uservec called usertrap with jalr, so its ret lands there. No hart is running the shell. To resume it, any hart's scheduler loads context.sp (swtch.S:26); sched, sleep, consoleread … return up this stack, and userret later reloads sp = 0x4f10 from the trapframe. Values read with gdb from one boot of this build (3 harts), shell idle at the prompt. Physical addresses omitted; they vary.
The shell, asleep in read. Everything needed to resume it is in three places: p->context (where its kernel thread stopped), its kernel stack (the C call chain), and its trapframe (the user registers, including the user sp). No hart is using any of them.

Resuming it takes all three, in order. Some hart’s scheduler loads p->context, which puts sp back at 0x3fffffbeb0 (switch 3). The functions on the kernel stack return one by one, up to usertrap, which returns into userret. userret loads sp = 0x4f10 from the trapframe (switch 4) and returns to user mode. Tour 47: Where a suspended process lives follows this in detail.

fork, exec and exit, seen from the stacks

fork. The parent runs kfork on its own kernel stack (usertrap → syscall → sys_fork → kfork). The child gets:

The child’s first run: a scheduler’s swtch loads the top of its kernel stack (switch 3) and returns into forkret, which therefore runs on an empty stack (gdb shows sp = 0x3fffffe000 at its first instruction for pid 1, slot 0). forkret calls prepare_return and then calls userret through a function pointer (kernel/proc.c:542); switch 4 moves sp to the user stack’s address, and sret makes that stack usable. forkret’s frame is abandoned, never popped; the next trap starts again at the top and writes over it. For the very first process, forkret also runs fsinit and kexec("/init") on this stack before returning to user mode (kernel/proc.c:528–532).

exec. kexec runs on the calling process’s kernel stack, and that stack, the slot and the trapframe page are the same before and after. Only the user side is replaced. It builds a complete new page table with a new user stack, as described above, while the old user stack stays intact in the old page table, so a failure before the commit (kernel/exec.c:133) can still return -1 to the old program. At the commit it installs the new page table, epc and sp, then frees the old image, old user stack included (kernel/exec.c:138). The system call returns through usertrap and userret as usual, and switch 4 loads the new sp: the first instruction of the new program runs on the new stack.

exit. kexit runs on the process’s kernel stack (called from sys_exit, or from usertrap when the process has been killed). It closes files, hands children to init, marks itself ZOMBIE and calls sched, which never returns (kernel/proc.c:363). The frames usertrap → syscall → sys_exit → kexit → sched stay on the kernel stack, and p->context points into them, but nobody will ever resume them. A process could not free the stack it is standing on, and it still needs it while swtch runs; xv6 sidesteps the problem because kernel stacks belong to slots and are never freed. The parent, in kwait, calls freeproc (kernel/proc.c:398), which frees the trapframe and the user page table with all user memory, the user stack included, and marks the slot UNUSED. The kernel stack stays mapped, waiting for the slot’s next process.

The parent cannot reach the zombie too early. The child holds its p->lock from kexit through swtch, and its hart’s scheduler releases that lock only after swtch has moved sp off the child’s stack (kernel/proc.c:463). kwait must acquire the same lock before it can even see ZOMBIE.

Three harts at once

Each hart has its own sp, so with QEMU’s three harts up to three stacks are active at the same moment, one per hart. A typical snapshot (the addresses are real, the combination is illustrative):

Hart Running Mode Active stack Page table (satp)
0 its scheduler loop, nothing to run S hart 0’s scheduler stack, sp = 0x80008810 kernel
1 sh inside read, about to sleep S sh’s kernel stack, KSTACK(1) kernel
2 cat copying a file U cat’s user stack cat’s user page table
none init, asleep in wait its frames wait on KSTACK(0); no hart uses them

The rules that keep this safe:

At every instant, then, each hart has a privilege mode, an active stack and a page table, and only these combinations occur at this commit:

Mode Active stack Page table When
M none (sp = 0, then stack0’s base, not yet usable) none (paging off) QEMU’s boot ROM at 0x1000, _entry before line 17
M boot none (paging off) _entry from line 17, start
S boot none, then kernel main before it calls scheduler
S scheduler kernel the scheduler loop, and kerneltrap taken in its interrupt window
S kernel kernel system calls, faults, interrupts and yields of a process; forkret; kexit
S none (sp holds the user’s value) user uservec until line 76, and userret from line 118 until sret (line 153)
S none (sp holds a kernel-stack address) user uservec lines 76–92 and userret lines 111–118
U user user the program

Mode, stack and page table: the master question walks through how each of the eight kinds of transition moves a hart from one row to another.

Guard pages, and the stack that has none

A stack that grows past its end silently overwrites whatever lies below it. A guard page turns that into a fault.

Common misconceptions

Where to go next