Tour 5 · Traps and system calls · about 42 minutes · 26 steps
You type echo hi at the xv6 shell, and two letters appear on your screen. Between those
two events, a user program asks the kernel for help, the CPU changes privilege mode twice,
two page tables take turns, a process goes to sleep on one CPU and wakes up on another, and
a device interrupt is handled by a third.
This tour follows one system call, write(1, "hi", 2), from the moment echo makes it
to the moment echo continues with the next line, through every layer of xv6:
the user-space system call stub, the trampoline page, usertrap, the system
call table, the file layer, the console driver and the UART, and all the way
back.
It is the single most important path in the kernel. Almost every other tour is a detour from it.
Best after: 2. Power-on to main, on every hart at once, 4. From the first process to the shell prompt
The machine has three harts. When the tour starts:
| Hart | What it is doing |
|---|---|
| 0 | Its scheduler is looping, looking for something to run |
| 1 | Running echo (pid 3) in user mode: the process this tour follows |
| 2 | Running whatever else is runnable, or idle in its own scheduler |
The shell (pid 2) is asleep in kwait, waiting for echo to finish.
Step 1 of 26
The shell started echo with argv = {"echo", "hi", 0}, so the loop runs once, with
argv[1] = "hi". Line 11 calls write(1, "hi", 2).
echo cannot print anything by itself. It runs in user mode, where the hardware
forbids touching device registers, and even if it could, the kernel must decide who
gets the screen. So printing means asking the kernel, and asking the kernel means a
system call.
File descriptor 1 is standard output. echo never opened it.
It inherited it from the shell, which inherited it from init, which opened the
console device at boot (user/init.c). All three processes’ descriptor 1 point at
the same open file (struct file) in the kernel. Remember that: it will matter when we reach
the kernel’s file layer.
ecall will not change spStep 2 of 26
write is not a C function. It is a three-instruction system call stub generated by
user/usys.pl. By the RISC-V calling convention, the C call already left the
arguments in registers: a0 = 1, a1 = the address of "hi", a2 = 2.
li a7, SYS_write adds the one thing the kernel needs to know: which service is
requested. SYS_write is 16 (kernel/syscall.h). In the built program it
is the 2-byte instruction li a7,16 at address 0x352 (user/echo.asm).What ecall does, all at once, in hardware:
| Register | Becomes |
|---|---|
sepc |
the address of this ecall |
scause |
8, “environment call from U-mode” |
sstatus.SPP |
0: the trap came from user mode |
sstatus.SPIE / SIE |
the old interrupt-enable bit is saved, and interrupts are turned off |
| privilege mode | supervisor (because start delegated all exceptions to S-mode with medeleg, kernel/start.c:31; otherwise the trap would go to machine mode) |
pc |
the address in stvec |
One thing does not change: satp. The CPU is now in supervisor mode but still translating addresses with echo’s page table.
sp still holds echo’s user stack pointer; uservec never pusheswas: user stackStep 3 of 26
stvec pointed at uservec, so this is where hart 1 now executes. The problem it
must solve: the CPU is in the kernel’s privilege mode, but it is still using the user
page table, and every register still holds a user value that must be preserved.
This works only because the trampoline page page is mapped at the same virtual address
(TRAMPOLINE, the top page of the address space) in every page table, user and
kernel alike (kernel/vm.c:47, kernel/proc.c:189). Whatever satp says, the next
instruction is there.
The code needs one free register to work with, but all 31 hold user data. So it parks
the user’s a0 in sscratch, a spare CSR, and loads TRAPFRAME into
a0. TRAPFRAME is the page just below the trampoline, and in this process’s page
table it maps to this process’s trapframe.
Notice what the hart does not have: a stack. ecall did not touch sp, so it
still holds echo’s user stack pointer, a value the user program chose and that may
not even be a valid address. The kernel must never push onto it, and uservec is
written so that it never pushes at all. Until the kernel page table is installed at
line 92, the hart has no usable stack (The stacks of xv6): first it holds the user’s
sp, then, after line 76, a kernel-stack address that echo’s page table does not
map.
sp, saved into the trapframe by line 41Step 4 of 26
Thirty-one sd instructions copy every general-purpose register into the trapframe,
at the offsets of struct trapframe (kernel/proc.h:40):
ra at 40, sp at 48, and so on. The user’s a0, parked in sscratch, goes to
offset 112 on lines 72–73.
From now until the return, the trapframe is echo’s user state. The kernel will
read the system-call number from a7 (offset 168) and the arguments from a0–a2,
and it will write the return value into the saved a0. When echo resumes, it gets
these registers back, with a0 changed to the result.
Interrupts are still off (the ecall turned them off), so nothing can interrupt this
half-saved state.
sp holds the top of echo’s empty kernel stack, but the user page table does not map it; it becomes usable when satp switches at line 92Step 5 of 26
Now uservec loads what the kernel left in the trapframe last time echo returned
to user space:
sp ← kernel_sp: the top of echo’s kernel stack. Every process has its own,
so the C code about to run has a stack that no other hart is using.tp ← kernel_hartid: the kernel keeps the hart’s ID in tp (kernel/start.c:48),
but user code may have overwritten tp with anything.t0 ← kernel_trap, the address of usertrap.t1 ← kernel_satp, the kernel page table.Line 76, ld sp, 8(a0), is the stack switch of this trip into the kernel
(The stacks of xv6). Echo’s user sp is safe in the trapframe (line 41), and
the hart now points at the top of echo’s kernel stack. That stack is empty: every
return to user mode sets kernel_sp to p->kstack + PGSIZE (kernel/trap.c:117),
so a process in user mode has nothing on its kernel stack. If echo is the first
command after boot, it is in proc[2], and its stack is the page
0x3fffff9000–0x3fffffa000, so sp = 0x3fffffa000. For four more instructions that
address is not even mapped, because echo’s page table is still installed; nothing
pushes until the kernel page table is in place.
Why is kernel_hartid still correct? It was saved by prepare_return on whatever
hart last sent echo to user mode, and a process in user mode never changes hart:
only the kernel’s scheduler moves processes between harts. So it was saved on hart 1,
and it is still hart 1.
csrw satp (kernel/trampoline.S:92) makes the stack loaded on line 76 reachableStep 6 of 26
Then csrw satp, t1, bracketed by two sfence.vma instructions,
switches this hart to the kernel page table and discards cached user translations
from the TLB (translation lookaside buffer). The trampoline is mapped at the same address in the kernel table,
so the very next fetch still works. Finally jalr t0 calls usertrap.
The same switch makes sp usable: the kernel page table maps every process’s kernel
stack (all 64 were mapped once, at boot, by proc_mapstacks), so the first push in
usertrap’s prologue lands in echo’s stack page.
usertrap’s frame onlyStep 7 of 26
We are in C, on echo’s kernel stack, with the kernel page table. That stack was empty
a moment ago; usertrap’s prologue has just pushed the first frame, 32 bytes holding
ra (the trampoline address of userret) and some saved registers. usertrap:
SPP is 0). A kernel trap
arriving here would be a disaster.kernelvec. If an interrupt or exception happens
while the kernel is running, it must go to the kernel’s own handler, not back into
the trampoline.myproc: hart number from tp, then
cpus[hart].proc. That is why tp had to be right.sepc, the address of the ecall, into the trapframe, before anything can
overwrite it.Step 8 of 26
scause is 8, so this is a system call (trap cause (scause values)).
echo has been killed (perhaps by a kill from another hart a moment ago), it
exits now instead of doing the work.epc += 4: when echo resumes, it must continue after the ecall, not run it
again. ecall is always 4 bytes.intr_on: interrupts back on.That last line changes everything for the rest of the tour. From here on, a timer or
device interrupt can arrive on hart 1 in the middle of the system call. The kernel
waited until now because an interrupt overwrites sepc, scause and sstatus, and
those have now been read and saved.
Step 9 of 26
syscall reads the number from the saved a7: 16. It checks it is in range and
calls syscalls[16], which the table at kernel/syscall.c:126 maps to sys_write
using C designated initializers.
Look at line 146: whatever sys_write returns is stored into p->trapframe->a0, the
saved user a0. That is how a system call “returns a value”: when echo resumes, its
a0 register will hold the result, exactly as if write had been an ordinary
function.
A bad number (a buggy or hostile program could put anything in a7) gets -1 and a
message, not a crash.
Step 10 of 26
System-call handlers take no C arguments. They fetch them from the trapframe (system call arguments):
argaddr(1, &p): p is the user address of "hi". It is only a number; the
kernel cannot dereference it, because it means something only in echo’s page table.argint(2, &n): n = 2.argfd(0, 0, &f): turns descriptor 1 into a struct file *.Notice what is not checked: whether p points to readable memory. Checking now would
only duplicate work: copyin must translate every page of the buffer through echo’s
page table anyway when it copies the bytes, and it fails cleanly if an address is bad.
So the check happens exactly when the bytes are copied.
Step 11 of 26
argfd checks that the descriptor is in range and refers to an open file:
myproc()->ofile[1], the per-process table of open files (file descriptor).
There is no lock around that array read. Why is that safe on a three-hart machine?
Because ofile[] belongs to echo alone, and only echo itself ever changes it
(through open, close, dup, pipe). xv6 processes have one thread, so while echo
is here, nothing else can be modifying echo’s descriptor table.
Step 12 of 26
filewrite is the file layer: one entry point for pipes, disk files and devices.
After checking the file was opened for writing, it looks at f->type.
The console is FD_DEVICE with major number CONSOLE (1). Line 147 calls
devsw[1].write, a function pointer that consoleinit set to
consolewrite at boot (kernel/console.c:202). The first argument, 1, means “the
address is a user address”.
This table of function pointers is how a tiny kernel gets device independence: the file layer knows nothing about UARTs.
Step 13 of 26
consolewrite moves the data in batches through a 32-byte buffer on the kernel
stack. Ours is one batch of 2 bytes.
either_copyin with user_src = 1 calls copyin to bring "hi" from echo’s
memory into buf. If the address is bad, the copy fails and the loop stops, returning
how much it managed to write. A bad pointer from user space costs the program its
write, never the kernel its life.
Each batch then goes to uartwrite.
Step 14 of 26
The hart is using the kernel page table, so it cannot simply read address p: in
the kernel’s map, that number means something else entirely. copyin does in
software what the hardware would do with echo’s page table:
va0).walkaddr looks it up in echo’s page table and returns the physical address,
or 0 if it is not mapped (or not a user page).vmfault may supply a lazily-allocated page; otherwise the
copy fails.pa0 + offset. That works because the kernel maps all physical memory
at the same address (direct map)."hi" is two bytes on one page, so one iteration.
tx_lock.lk (spinlock)Step 15 of 26
uartwrite begins with acquiresleep(&tx_lock) (kernel/uart.c:82): only one
thread at a time may feed bytes to the UART, or two writers’ bytes would be shuffled
together.
Why a sleep lock and not a spinlock? Because the holder may have to wait for the UART, maybe for a long time on real hardware. A spinlock holder must never sleep, and a waiter spins with interrupts off, burning a whole CPU. With a sleep-lock, a waiting thread gives up its CPU instead.
Inside, a sleep-lock is a spinlock (lk->lk) protecting a locked flag. If it is
taken, the caller registers to wait, releases the spinlock and sleeps.
Look at the display above: throughout lines 24–33 (unless the loop at 25–29 has to
wait) hart 1 holds the sleep-lock’s inner spinlock tx_lock.lk, so interrupts are off. Only after line 31 sets
locked and line 33 releases the spinlock does echo hold tx_lock the sleep-lock,
with interrupts back on, which is the state you will see in the next step.
They come back on because the acquire on line 24 found them on and recorded
intena = 1; line 32’s myproc pushes noff to 2 and back for a moment without
changing that (Locks and interrupt state).
tx_lock (sleep-lock)Step 16 of 26
For each byte, uartwrite checks the UART’s line-status register: is the transmitter
ready (LSR_TX_IDLE)? If so, it writes the byte to the transmit-holding register
THR, a memory-mapped device register. The UART then shifts it out, one bit
at a time.
If the transmitter is busy, the thread must wait for the UART’s “I’m ready” interrupt. On QEMU the emulated UART is almost always ready immediately. On a real 38,400-baud line each byte takes about a quarter of a millisecond, an eternity for a CPU. Let’s follow the busy case.
The order of the three operations is the whole trick:
sleep_prepare(&tx_chan): “I am about to wait on tx_chan”.sleep().Registering before checking is what prevents a lost wakeup. See the box below.
usertrap … sleep; sched’s swtch will leave ittx_lock (sleep-lock)echo's p->lockStep 17 of 26
sleep takes echo’s p->lock, a spinlock, so interrupts are now off on hart 1. If
p->chan is still set (no wakeup yet), it marks the process SLEEPING and calls
sched, which switches to hart 1’s scheduler thread. sched insists that
noff is exactly 1 (p->lock only) and saves echo’s intena, 1, in a local
variable that this build keeps in the callee-saved register s3, which swtch
stores in echo’s p->context, because the next thread on hart 1 will set its own
(Locks and interrupt state).
Hart 1 is now free. Its scheduler will release echo’s p->lock and look for another
process to run. echo is frozen inside sleep, with its whole kernel call stack
preserved on its kernel stack, still holding tx_lock.
Inside sched, the ld sp, 8(a1) in swtch (kernel/swtch.S:26) moves hart 1’s
sp to its scheduler stack. Echo’s kernel stack, empty when uservec set sp to its
top, now holds everything this system call has done so far, and no hart’s sp points
into it:
echo's kernel stack (KSTACK(2), one page)
top ─► usertrap
syscall
sys_write
filewrite
consolewrite (buf[32] is in this frame)
uartwrite
sleep
sched ◄─ echo's p->context.sp, saved by swtch
Holding a sleep-lock while asleep is allowed, and it is the reason sleep-locks exist.
Holding a spinlock while asleep is forbidden, except for p->lock itself, which the
scheduler releases on the sleeper’s behalf (context switch).
stack0hart 2’s slice of stack0, with a 256-byte kernelvec frame on topStep 18 of 26
The UART finishes sending and raises its interrupt line. The
PLIC has the UART interrupt enabled for every hart (plicinithart), so
whichever hart claims it handles it. Here hart 2 wins. It arrives in devintr,
through kerneltrap if it was in the kernel or usertrap if it was in user mode.
In this telling hart 2 was idle, looping in its scheduler, and took the interrupt
in the moment at kernel/proc.c:441 when the scheduler turns interrupts on, with
noff = 0, as every interrupt must be: a hart holding any spinlock has its interrupts
off (Locks and interrupt state). There
is no separate interrupt stack in xv6: kernelvec pushed its 256-byte save area
onto the stack hart 2 was already using, its scheduler stack in stack0, and
kerneltrap and devintr run on top of it. (Had hart 2 been in a system call, the
same frame would have gone onto that process’s kernel stack; had it been in user mode,
uservec would have switched to that process’s kernel stack.)
scause says “supervisor external interrupt”. plic_claim asks the PLIC which
device it was: UART0_IRQ. After uartintr runs, plic_complete tells the PLIC
the UART may interrupt again.
This hart is handling an interrupt on behalf of a process (echo) that is not running
anywhere. Interrupt handlers serve the machine, not the current process.
stack0hart 2’s scheduler stack, under the kernelvec frameeach p->lock in turnStep 19 of 26
uartintr sees the transmitter idle and calls wakeup(&tx_chan)
(kernel/uart.c:145). wakeup walks the whole process table. For each process it
takes p->lock, and if that process is waiting on tx_chan, it clears p->chan and,
if it is fully asleep, sets it RUNNABLE. Each of those acquires runs inside an
interrupt handler with interrupts already off, so it records intena = 0 and its
release leaves them off (Locks and interrupt state).
echo is now RUNNABLE. It is not running yet; it is eligible. (The source comment
says “RUNNING”; the code correctly sets RUNNABLE.)
When the handler is done, kerneltrap does not yield: it yields only when
myproc() is a process, and hart 2 has none. kernelvec pops its frame and srets
back into the scheduler loop, where hart 2 could now find echo RUNNABLE too. In this
telling, hart 0’s scheduler gets there first.
ld sp, 8(a1) in swtch (kernel/swtch.S:26), called by hart 0’s schedulertx_lock (sleep-lock)Step 20 of 26
Hart 0’s scheduler was looking for work. It finds echo RUNNABLE, marks it
RUNNING, and switches to it. echo returns out of sleep exactly where it
stopped, still inside uartwrite, still holding tx_lock, now running on
hart 0.
The switch is the ld sp, 8(a1) in hart 0’s swtch: it loads echo’s
p->context.sp, saved on hart 1, so hart 0’s sp now points into echo’s kernel stack,
at the very frames hart 1 pushed. Hart 0’s own scheduler frames stay behind in its
slice of stack0.
Interrupts are on again because sched put back echo’s own intena, 1, saved on
hart 1, so sleep's release(&p->lock) turned them on, although hart 0’s scheduler
had taken that lock with interrupts off (Locks and interrupt state).
The loop goes around again: register, check, the transmitter is idle, write the byte
it was waiting to send. When both bytes are out, releasesleep gives up tx_lock and wakes anyone
waiting for it, such as the cat from step 15.
Nothing in uartwrite noticed the change of CPU, and nothing needed to. The kernel
stack, the local variables and the locks held all belong to the thread, not to the
hart.
Step 21 of 26
uartwrite returns to consolewrite, which returns 2 (bytes written).
filewrite returns 2, sys_write returns 2, and back in syscall line 146
stores it into p->trapframe->a0.
All of those returns pop frames off echo’s kernel stack. They are running on hart 0
now, but they are the same frames pushed on hart 1.
Step 22 of 26
usertrap continues after syscall():
kill from another hart only sets a
flag and wakes the victim if it is sleeping; it takes effect here, at the boundary.)which_dev is 0 for a system call, so no yield.prepare_return sets up the return to user space.MAKE_SATP builds the satp value for echo’s page table, and usertrap returns
it: in a0, by the calling convention, which is exactly where the trampoline
expects it.kernel_sp is set to its top for next timeStep 23 of 26
prepare_return first turns interrupts off. It is about to point stvec back
at uservec in the trampoline, and from that moment a kernel interrupt would land in
user-trap code. That would be a disaster.
Then it refills the trapframe’s kernel fields for the next trap: the kernel page
table, the top of the kernel stack (always the top: the next trap starts on an empty
stack), the address of usertrap, and, on line 119,
kernel_hartid = r_tp(): 0, this hart. The next time echo traps, it will be on
hart 0, because a process in user mode stays where it was sent. So hart 0 is the right
answer, and the hart 1 written on the last return is now stale.
Finally sstatus (SPP = 0: return to user mode; SPIE = 1: interrupts on once
there) and sepc (the saved user PC, already advanced past the ecall).
sp is the top of echo’s kernel stack, which echo’s page table does not mapwas: kernel stackStep 24 of 26
usertrap returns into the trampoline, right after jalr t0, which is userret.
In reverse order:
echo’s code was written into its pages with
ordinary data stores when exec loaded it, on whichever hart ran exec, and hart
0’s instruction cache could still hold stale bytes for those physical pages, for
example from a program that used them earlier. That matters the first time a process
runs on a hart; xv6 doesn’t track that, so it executes fence.i on every return to
user space.satp to echo’s page table, again between two sfence.vmas.From here until sret the hart has no usable stack. sp is back at the
top of echo’s kernel stack (usertrap popped its frame when it returned), but echo’s
page table does not map kernel stacks; from line 118 it holds echo’s user sp, which
supervisor code cannot use. The trampoline never pushes, so it does not matter.
sp holds echo’s user stack pointer, where write left it; supervisor code cannot use itld sp, 48(a0) in userret (kernel/trampoline.S:118) loads the user’s spStep 25 of 26
a0, which now holds 2.
The first of them that matters for stacks is line 118, ld sp, 48(a0): sp is
back on echo’s user stack, at exactly the value that uservec saved, still
holding main’s frame. From here to sret the hart is still in supervisor mode with
the user page table, and supervisor code cannot use user pages, so there is no usable
stack and nothing is pushed.sepc, with interrupts enabled. It
does not touch sp, but by changing the mode it makes the user stack usable again.In one instruction the hart drops from supervisor to user mode, and echo continues
as if write were an ordinary function call that took a while. Its kernel stack is
empty again, as it is whenever echo runs in user mode.
sret (kernel/trampoline.S:153) returned to user mode, making the user stack usableStep 26 of 26
The stub’s ret returns to main with a0 = 2. echo ignores the result, and on
line 15 writes the newline: the whole journey again, for one byte. Then exit(0),
another system call, which has a tour of its own: Tour 21: exit, wait and zombies.
Count what one write took in our story: 2 privilege-mode changes (ecall and
sret), 2 satp writes (one in uservec, one in userret, each between two
sfence.vmas), 31 registers saved and 31 restored, 1 sleep and the UART interrupt
that ended it (the UART interrupts again after the last byte, with nobody waiting),
and 3 harts. It also took 4 stack switches: uservec’s ld sp onto echo’s kernel
stack, the swtch in sched onto hart 1’s scheduler stack, hart 0’s swtch
back onto echo’s kernel stack, and userret’s ld sp back to the user stack (usable once sret returned to user mode). Hart 2’s
kernelvec frame was pushed onto the stack it was already on, which is not a switch. On QEMU, whose UART is almost always ready, there would usually be no
sleep at all. A system call is not a
function call. It is a carefully staged handover of the CPU, and the lock at every
shared point is what lets three CPUs do this at once without tripping over each other.
Tour 5 · wrap-up
| Lock | Taken in | Protects |
|---|---|---|
cpus[].proc (no lock: interrupts off) | myproc | Reading this hart’s current process; interrupts off keeps the hart from changing mid-read |
tx_lock.lk (spinlock) | acquiresleep, releasesleep | The sleep-lock’s own locked flag and holder pid |
tx_lock (sleep-lock) | uartwrite | The UART transmitter: one writer feeds it at a time |
p->lock (spinlock) | sleep_prepare, sleep, wakeup, killed (twice in usertrap), the schedulers | p->chan, p->state and p->killed: who is waiting, who may run, who must exit |
kmem.lock (spinlock) | kalloc, only if copyin meets a lazily allocated page | The shared free-page list |
When ecall executes, satp still holds the user page table. Why doesn’t the CPU crash fetching the next instruction?
The trampoline page is mapped at the same virtual address (TRAMPOLINE) in every user page table and in the kernel page table, and stvec points into it. Whatever page table is active, the trampoline code is there.
echo’s system call began on hart 1 and finished on hart 0. Which line makes sure the next trap from echo restores the right hart ID into tp, and why is the answer guaranteed to be correct?
p->trapframe->kernel_hartid = r_tp() in prepare_return (kernel/trap.c:119) records hart 0. It stays correct because a process in user mode never changes hart; only the kernel scheduler moves processes. So the next trap from echo is taken on hart 0.
Why does uartwrite use a sleep-lock (tx_lock) instead of a spinlock?
The holder may wait a long time for the UART and sleeps while doing so. A spinlock holder must never sleep (and has interrupts off), and its waiters would spin uselessly. A sleep-lock lets waiters give up their CPU.
The UART interrupt arrives on hart 2 between uartwrite’s check of LSR (busy) on hart 1 and its call to sleep(). Walk through why echo does not sleep forever.
echo called sleep_prepare(&tx_chan) before checking, so p->chan was already set. wakeup(&tx_chan) on hart 2 clears p->chan under p->lock. When echo then calls sleep(), it sees p->chan == 0 and returns immediately, loops, and finds the transmitter idle.
argfd reads myproc()->ofile[fd] without taking any lock. Why is that safe even with three harts?
Only the process itself changes its own ofile[] table, and xv6 processes have a single thread. While echo runs argfd, no other code can be modifying echo’s descriptor table. The shared struct file it points to is kept alive by its reference count.
Why does usertrap wait until after saving sepc and checking scause before calling intr_on()?
An interrupt overwrites sepc, scause and sstatus. Those registers had to be read and saved first, or the information about the system call would be lost.
Keys: ← → step · Home start