Tour 10 · Traps and system calls · about 39 minutes · 22 steps
A system call is a trap the program asks for. An exception is a trap the program did not ask for: it touched memory it has no page for, wrote to its own code, or executed an instruction it is not allowed to execute. The hardware stops the instruction halfway, records what went wrong in three CSRs, and jumps into the kernel. The kernel then has exactly three choices: fix the problem and retry the instruction, kill the process, or, if the faulting code was the kernel itself, stop the whole machine.
This tour shows all three, using real programs and real numbers from runs of this build (scenes 1–2) and of a scratch copy we modified (scenes 3–4):
usertests lazy_alloc asks for a gigabyte with sbrklazy and then stores into it. Each
first touch of a page is a page fault that vmfault quietly repairs.usertests nowrite stores into its own read-only code page. The fault cannot be
repaired, so the kernel prints usertrap(): unexpected scause 0xf and kills the child.csrr mhartid,
a machine-mode instruction, and is killed for an illegal instruction.panic: kerneltrap, and we watch what the panic does to the other two harts.The trap entry itself (trampoline, trapframe, usertrap's first lines) is covered
by Tour 5: Life of a system call and Tour 7: The trampoline and the trapframe. Here we go deeper into what happens after scause says
“this was not a system call”.
Best after: 5. Life of a system call, 7. The trampoline and the trapframe
The machine has three harts. The four faults come from separate runs; for the first one, in the run we traced with gdb:
| Hart | What it is doing |
|---|---|
| 0 | Idle in its scheduler (init and sh are asleep in wait), or running whatever is runnable |
| 1 | Running pid 4, the child that usertests forked to run lazy_alloc: the process we follow first |
| 2 | Idle in its own scheduler, or running whatever else is runnable |
usertests itself (pid 3) is asleep in kwait, waiting for pid 4. Later scenes
move to hart 2 (the nowrite children and the kernel fault) and back to hart 0 (a hart
that is caught printing during a panic, in a hypothetical scene).
Step 1 of 22
lazy_alloc asks for REGION_SZ, 1 GiB, with sbrklazy. That is
eight times the machine’s whole 128 MiB of RAM. Then it writes one pointer into every
64th page, 4096 pages in all (16 MiB), and reads them back.
sbrklazy(n) is a one-line wrapper (user/ulib.c:159): it calls the sys_sbrk
system call with a second argument, SBRK_LAZY (2), instead of SBRK_EAGER (1). The
difference is the whole point of this scene:
sbrk allocates and maps every page immediately (growproc →
uvmalloc). A gigabyte would fail: there is not that much memory.sbrk only promises the memory. Pages appear one at a time, when the
program first touches them, through page faults.In our run usertests had size 0x12000 (its code, data, a guard page and one stack
page), so sbrklazy returns prev_end = 0x12000, and the first store on line 2640
goes to prev_end + PGSIZE = 0x13000.
The allocation strategy itself is the subject of Tour 26: sbrk, eager and lazy, and page faults. This tour follows what the trap does with it.
ld sp, 8(a0) in uservec (kernel/trampoline.S:76)Step 2 of 22
sys_sbrk fetches n = 1073741824 and t = SBRK_LAZY from the trapframe
(argint, see Tour 6: System-call arguments and user pointers) and records addr = p->sz = 0x12000. Because the request
is lazy and positive, it takes the else branch:
addr + n < addr rejects a request that would wrap around the 64-bit address space.addr + n > TRAPFRAME rejects a heap that would run into the trapframe page at
0x3fffffe000.myproc()->sz += n: the size becomes 0x40012000.That is all. No page is allocated, no PTE (page-table entry) is written. The page table still ends
at 0x12000; above it, every PTE is invalid. p->sz is now a promise: “any address
below this is legitimately mine”.
Remember that number. In a moment the hardware will complain about address 0x13000,
and p->sz is the only thing that tells the kernel the complaint is about a page it
owes the process, not about a wild pointer.
ld sp, 48(a0) in userret (kernel/trampoline.S:118) and sret (kernel/trampoline.S:153), returning from sbrkStep 3 of 22
Back in user mode, line 2640 compiles to sd a5,0(a5) at address 0x495e
(user/usertests.asm), with a5 = 0x13000. The hart’s page-table walker looks up
0x13000 in pid 4’s page table and finds an invalid PTE. The store cannot complete,
so the hart raises a store/AMO page fault instead of executing it:
| Register | Becomes |
|---|---|
sepc |
0x495e, the address of the faulting sd itself |
scause |
15, “store/AMO page fault” (high bit 0: an exception, not an interrupt) |
stval |
0x13000, the address it tried to store to |
sstatus.SPP / SPIE / SIE |
0 (from user mode) / old SIE / interrupts off |
pc |
stvec, which points at uservec in the trampoline page |
The trap goes to supervisor mode only because start delegated all exceptions
there with medeleg at boot (kernel/start.c:31). Without that,
it would go to machine mode, where xv6 has no handler.
Compare with ecall in Tour 5: Life of a system call: there sepc pointed at an instruction that
finished its job, so the kernel added 4. Here the sd did not happen. If the
kernel fixes the cause, it must run this very instruction again.
usertrap’s frame is the only oneld sp, 8(a0) in uservec (kernel/trampoline.S:76)Step 4 of 22
uservec saves pid 4’s 31 registers into its trapframe and calls usertrap on
pid 4’s kernel stack, exactly as for a system call (Tour 7: The trampoline and the trapframe covers every
instruction). The exception, like any trap, left sp alone; uservec saved the user
sp in the trapframe and loaded the top of pid 4’s kernel stack, which is empty
whenever a process is in user mode (The stacks of xv6). The kernel never runs on
the user’s stack: the kernel page table does not map user memory at all, and the
user’s sp could hold any value, even an address in the guard page. usertrap saves sepc (0x495e) into p->trapframe->epc, and then
asks its three questions in order:
scause == 8? A system call. No.devintr? It tests scause for the two interrupt codes (high bit set). 15 has
the high bit clear, so devintr returns 0: not a device.scause is 15 (store) or 13 (load) and vmfault can supply a page? Lines
71–73 call vmfault(p->pagetable, p->sz, r_stval(), 0).Notice two things that are not here. There is no epc += 4: the instruction must be
retried. And there is no intr_on: on this path interrupts stay off for the whole
time in the kernel. Only the system-call branch turns them on. Repairing a page fault
is short work, and nothing in it needs to sleep.
Only 13 and 15 are candidates. An instruction page fault (12) never reaches
vmfault, so jumping into a lazily promised page kills the process; lazy pages are
mapped without execute permission anyway.
Step 5 of 22
vmfault has one job: tell a promised-but-missing page apart from everything else.
gdb, stopped here in our run: va = 0x13000, psz = 0x40012000, read = 0.
va >= psz? No: 0x13000 is below the size sys_sbrk recorded. The process owns it.PGROUNDDOWN(va): 0x13000 is already page-aligned. (A fault at 0x13008 would be
rounded to the same page.)ismapped? If the page is already mapped, the fault was not about a missing
page: it was a permission violation, such as a store to read-only memory. That is
not vmfault’s to fix, so it returns 0. We will hit this case in scene 2.The order of the checks matters. va >= psz comes first, and psz is never above
TRAPFRAME, so ismapped (which calls walk) only ever sees addresses below
MAXVA. walk would panic on an address at or above MAXVA; a user pointer like
0xffffffffffffffff must never get that far.
The read argument (1 for a load fault) is not used by this version of vmfault.
Loads and stores to a missing page are treated alike.
Step 6 of 22
ismapped calls walk(pagetable, va, 0). The final 0 means “do not allocate”:
if an intermediate page-table page is missing, walk returns 0 instead of creating
one, so looking can never change the page table.
For 0x13000 the two upper levels exist: top-level entry 0 (covering
0–0x3fffffff) and middle-level entry 0 (covering 0–0x1fffff) exist, and they
were created when exec mapped the
program’s text at address 0. So walk returns a pointer to the bottom-level PTE for
0x13000, and that PTE is 0: PTE_V clear. ismapped returns 0, “not mapped”, and
vmfault goes on to allocate.
This is the same walk the hardware did a few microseconds ago when it refused the store, now done in software by the kernel, on pid 4’s page table, while the hart itself runs on the kernel page table.
intr_on, so SIE was already off at the push_offkmem.lockStep 7 of 22
kalloc pops the first page off kmem.freelist. This is the first point on the path
that touches state shared by all harts, so it is the first lock: kmem.lock.
acquire calls push_off first. Interrupts were already off (we are in
usertrap’s fault path), so push_off records intena = 0 and the matching
pop_off will leave them off. The pop of the list is a handful of instructions inside
the critical section; the 4096-byte junk fill with 5 happens after the
release, so the lock is held only for a few instructions.
In our run the page was physical 0x80039000. Its old contents (junk, or whatever a
freed process left there) are about to be wiped. The allocator itself is
Tour 27: The physical page allocator.
Step 8 of 22
Back in vmfault:
kalloc returned 0? Memory is exhausted. vmfault returns 0, usertrap treats the
fault as unexplained, and the process is killed. A lazy promise can be broken; that
is the price of not reserving memory up front.memset(mem, 0, PGSIZE): the page must start as zeros. It is the kernel writing to
it, through the direct map at 0x80039000. Without this, pid 4 could read
whatever the page’s previous owner left behind: another user’s file contents, a
password typed into a shell.mappages installs it at 0x13000 with PTE_W | PTE_U | PTE_R: readable and
writable by user code, not executable.mappages fails (it may need to allocate a page-table page and find none), the
page goes back with kfree, which takes kmem.lock again.vmfault returns the physical address, non-zero, which usertrap reads as “handled”.
Step 9 of 22
mappages calls walk(pagetable, 0x13000, 1), this time with alloc = 1, so a
missing middle or bottom page-table page would be created with kalloc. Here none is
missing (see step 6), so walk returns the same PTE slot ismapped inspected.
The remap check on line 166 would panic if the slot were already valid: mapping over a live page would leak it, and in xv6 that can only mean a kernel bug.
Line 168 is the moment the promise becomes real: one 64-bit store,
PA2PTE(0x80039000) | PTE_W | PTE_U | PTE_R | PTE_V, into pid 4’s page table.
No sfence.vma is needed right here: hart 1 is running on the
kernel page table, and no other hart is using pid 4’s. One is needed before pid 4
retries the store: without the optional Svvptc extension, the spec does not promise
that a PTE changed from invalid to valid is seen by the page-table walker until an
sfence.vma, and a stale “invalid” view would fault again, and this time vmfault
would find the page mapped and kill pid 4. userret provides the fence: it writes
satp between two sfence.vma zero, zero instructions, ordering the PTE store before
later walks and discarding every cached translation on hart 1.
Step 10 of 22
usertrap continues:
killed: has someone killed pid 4 meanwhile? No.which_dev is 0, so no yield.prepare_return writes p->trapframe->epc, still 0x495e, into sepc.satp; userret restores all 31 registers and executes
sret.Pid 4 resumes at 0x495e, the same sd a5,0(a5). This time the walker finds a valid,
writable PTE and the store succeeds. From the program’s point of view, nothing
happened: one instruction took a few microseconds longer than usual.
This is what makes a fault restartable: the hardware guarantees that a faulting
instruction has no effect, and sepc points at it, so the kernel can fix the world
and simply let it try again.
ld sp, 48(a0) in userret (kernel/trampoline.S:118) and sret (kernel/trampoline.S:153), after each faultStep 11 of 22
The first loop faults once per touched page: 4096 round trips through the trampoline,
usertrap, vmfault, kalloc and mappages, and 16 MiB of zeroed pages. The second
loop only reads pages that now exist, so it runs without a single trap.
The cost per first touch: two privilege changes, two satp writes, four
sfence.vmas, 31 registers saved and restored, one fence.i, one kmem.lock acquisition (two on the
511 faults that also need a new page-table page), and a 4096-byte memset (plus
kalloc’s own 4096-byte junk fill). Every 8th fault enters a new 2 MiB region, so
walk must also allocate a bottom-level page-table page. The win: 1 GiB promised,
16 MiB of data pages plus 511 page-table pages (about 2 MiB) actually used.
When pid 4 exits, uvmfree frees only the pages that actually got mapped:
uvmunmap skips the invalid PTEs in the promised range.
Step 12 of 22
Scene 2, a fresh run of usertests nowrite: usertests is pid 3, the test runs in pid
4, and pid 4 forks one child per forbidden address. The first child, pid 5, executes
*addr = 10 with addr = 0: the store sw at 0x1ce8.
Address 0 is mapped: it is the first page of the program’s text. kexec mapped it
in pid 3 with the ELF segment’s flags, R E, which flags2perm turned into
PTE_X, plus PTE_R | PTE_U from uvmalloc, and each fork’s uvmcopy copied
those flags into pid 4 and pid 5. No PTE_W. The walker finds a valid
PTE without write permission and raises the same exception as before: scause = 15,
stval = 0.
All six addresses fault, each for its own reason:
| Address | Why the hardware refuses |
|---|---|
0 |
mapped, but read-only text |
0x80000000 |
not mapped in the user page table at all |
0x3fffffe000 |
TRAPFRAME: mapped, but without PTE_U |
0x3ffffff000 |
TRAMPOLINE: mapped, but without PTE_U |
0x4000000000 |
MAXVA: bits 63–39 are not copies of bit 38, not a valid Sv39 address |
0xffffffffffffffff |
a valid Sv39 address, but nothing is mapped there (the 4-byte sw is also misaligned there; QEMU reports the page fault) |
Step 13 of 22
usertrap takes the same branch as in scene 1, but vmfault returns 0:
0 < p->sz, so it proceeds to ismapped, which finds the text page valid. A
valid page that faulted is a permission violation, not a missing page. For the other
five addresses vmfault stops even earlier, at va >= psz.
So the else branch runs. These are the real lines from our run:
usertrap(): unexpected scause 0xf pid=5
sepc=0x1ce8 stval=0x0
usertrap(): unexpected scause 0xf pid=6
sepc=0x1ce8 stval=0x80000000
...
usertrap(): unexpected scause 0xf pid=10
sepc=0x1ce8 stval=0xffffffffffffffff
stval names the culprit address; sepc names the instruction (the same sw in all
six children, which are copies of one program). Running usertests stacktest gives a
load fault instead: scause 0xd (13), stval=0x10e90, an address in the
guard page below the stack. That page is valid but has PTE_U cleared
(uvmclear), so ismapped says yes and vmfault refuses too. stacktest only
imitates an overflow: its sp is still fine, and one load 4096 bytes below it hits
the guard page. In a real overflow sp itself would point into the guard page. The
kernel would still be unaffected, because uservec only saves the user’s sp and
runs usertrap on the process’s kernel stack.
Then setkilled, and at line 81 killed sees the flag and calls kexit(-1).
The kernel does not try to “repair” a bad store. It records why and removes the
process.
pr.lockStep 14 of 22
printk takes pr.lock (when no panic is in progress) and holds it while it emits
every character of one formatted line through consputc and uartputc_sync,
which busy-waits on the UART instead of sleeping. Spinning is acceptable here: these
messages are rare, and printk must work from places that cannot sleep, such as this
fault path with interrupts off.
The lock is per call. usertrap’s message is two calls, so another hart’s line can
land between usertrap(): unexpected scause... and sepc=.... The lock guarantees
lines are not shuffled character by character, not that a two-line message stays
together.
Console output in general, and why user writes go a different way, is Tour 38: Output to the console from three harts.
pid 5's p->lockStep 15 of 22
setkilled sets p->killed = 1 under p->lock. The process is not destroyed here;
it is only marked. Destruction happens at a safe point, which for a process in
usertrap is a few lines later, when killed reads the flag (again under p->lock)
and usertrap calls kexit(-1).
Why go through a flag at all, when usertrap could call kexit directly? Because the
flag is the same mechanism kkill uses from another process (Tour 23: kill). Using it
here keeps a single rule: a killed process dies only by noticing its own killed flag at a
point where it holds no locks and has nothing half-done.
wait_lockpid 5's p->lockStep 16 of 22
kexit closes pid 5’s files, releases its current directory, takes wait_lock, hands
any children to init, wakes the parent (pid 4, waiting in kwait), and then, holding
both wait_lock and its own p->lock, stores xstate = -1 and becomes a ZOMBIE.
sched switches away for the last time. Exit and reaping are Tour 21: exit, wait and zombies.
Pid 4’s wait returns 5 and stores -1 in xstatus. nowrite checks only that the
child did not exit with 0 (had the store succeeded, the child would have printed
write to ... did not fail! and called exit(0)), and moves on to the next address.
kexit runs to the end on pid 5’s own kernel stack, even after pid 5 is a ZOMBIE;
sched’s swtch moves hart 2 to its scheduler stack and never comes back. A process
cannot free the stack it is standing on, and xv6 never needs to: kernel stacks belong
to process-table slots, mapped once at boot. When pid 4 reaps pid 5, freeproc
frees the trapframe and the user page table (with the user stack), and the slot’s
kernel stack simply waits for its next occupant.
A detail worth noticing: since this trap was not a system call, interrupts were never
turned back on, and the outermost push_off of every acquire along the way records
intena = 0 (Locks and interrupt state). So kexit runs with interrupts off from start to finish, except while it sleeps
(for example in begin_op), when another process runs on this hart with its own
interrupt state. If kexit does sleep, pid 5 may wake on a different hart; in this
run nothing it closes needs the disk, so it stays on hart 2.
Step 17 of 22
Scene 3. No xv6 program does this, so we wrote a three-line test program, ill, and
added it to a scratch copy of the source; it is not part of xv6. Its only interesting
line is asm volatile("csrr %0, mhartid" : "=r"(x));. It assembles to
f14025f3 at address 8. That is the instruction r_mhartid runs in
start (kernel/start.c:47) in machine mode, where it is legal. In user
mode, reading a machine-mode CSR is an illegal-instruction exception,
scause = 2.
Nothing in usertrap special-cases 2. It is not 8, devintr returns 0, and it is
neither 13 nor 15, so it falls to the else branch. Our run printed:
usertrap(): unexpected scause 0x2 pid=3
sepc=0x8 stval=0xf14025f3
For an illegal instruction, the RISC-V spec lets stval hold either 0 or the bits of
the offending instruction; QEMU supplies the bits, so stval here is the instruction.
The process is killed exactly as in scene 2.
The same branch catches the rest of the exception list, such as ebreak (3) and
instruction page faults (12). xv6’s policy for
every one of them is “kill the process and keep the kernel running”.
usertrap → syscall → sys_uptimetickslockStep 18 of 22
Scene 4 needs a kernel bug, so we planted one in a scratch copy: line 109 became
xticks = *(volatile uint *)0;, and a two-line program up, also only in that scratch copy, called uptime(). The
planted load compiled to lw a5,0(zero) at 0x80002ae6 in that build.
The kernel page table maps nothing at virtual address 0 (its lowest mapping is the
PLIC at 0x0c000000), so the load raises a load page fault, scause = 13,
stval = 0, exactly as a user load would. Exceptions cannot be masked: interrupts are
off (the hart holds tickslock, a spinlock) but the fault happens anyway. SIE
controls only interrupts.
The trap came from supervisor mode, so stvec points at kernelvec, not the
trampoline (usertrap set it on entry). kernelvec pushes the registers on pid 3’s
kernel stack, right below sys_uptime’s frame, and calls kerneltrap. That entry
path is Tour 8: Traps taken inside the kernel.
Note what the hart is holding: tickslock. It will never release it.
kernelvec frame now sits below sys_uptime’stickslockStep 19 of 22
The fault did not change stacks. kernelvec's addi sp, sp, -256
(kernel/kernelvec.S:14) pushed its register frame onto the stack that was already
in use, pid 3’s kernel stack:
pid 3's kernel stack
top ─► usertrap
syscall
sys_uptime
kernelvec 256-byte register frame
kerneltrap ◄─ sp
kerneltrap reads sepc, sstatus and scause, checks that the trap really came
from supervisor mode (SPP = 1) and that interrupts are off (the trap turned SIE off),
and then asks devintr whether this was a device. scause = 13 is not an
interrupt, so devintr returns 0.
In user mode, that would lead to “kill the process”. In the kernel there is no such
option. The kernel has no vmfault path of its own and no other plan: a fault in
kernel code means a kernel bug, and the half-finished operation (here, a critical
section holding tickslock) cannot be safely abandoned or retried. So it prints the
evidence and panics. Our run:
scause=0xd sepc=0x80002ae6 stval=0x0
panic: kerneltrap
The first line is an ordinary printk: no panic is in progress yet, so it takes
pr.lock like any other line.
tickslockStep 20 of 22
panic does four things:
panicking = 1. From now on, printk on every hart skips pr.lock, and
uartputc_sync skips its push_off/pop_off (a printk already past that
check keeps the lock it took or is waiting for).panic: kerneltrap, in two printk calls, without the lock.panicked = 1: freeze every other hart that tries to print.Why must the panic message bypass pr.lock? Because the panicking hart may already
hold it. If the fault had happened inside printk itself, or if printk were the
caller that tripped acquire's own panic("acquire") check, a panic that called
acquire(&pr.lock) would find the lock held by its own hart: acquire would
panic("acquire") again, recursively, and never print a word. Another hart holding
the lock forever (frozen in the middle of a line) would block it the same way. The one
message that matters most must not depend on a lock.
Both flags are volatile int, so each check in printk and uartputc_sync really
reads memory, not a copy the compiler kept in a register.
panicking flipped in the middle of a uartputc_sync (its pop_off is then skipped); once the flag is seen at line 105, no push_offpr.lockStep 21 of 22
This scene is hypothetical: a second process, pid 7, faults on hart 0 just as hart 2
panics. panic does not stop the other harts directly. xv6 has no inter-processor interrupt to
send them. Instead, every path that prints to the UART synchronously checks the flags:
panicking set: no push_off/pop_off. The interrupt bookkeeping is skipped.panicked set: spin forever, right here, before writing the character.So a hart keeps running processes after another hart has panicked, until the moment it
tries to emit a character through uartputc_sync (kernel printk, or the echo of a
typed key in consoleintr), or until it blocks on a lock the dead hart held. The
console usually shows a panic: line as the last thing on the screen (possibly mixed
with one in-flight line, see below) instead of a mess.
This is a simplification of “stop the machine”, and an honest one: user writes go
through uartwrite, which does not check panicked, so in principle a user program
on a still-running hart could print after the panic message.
Step 22 of 22
Every exception in this tour reached one question: can the kernel make this instruction succeed?
| Fault | Who answers | Outcome | Cost |
|---|---|---|---|
store to a promised page (lazy_alloc) |
vmfault: below p->sz, not mapped |
page allocated, instruction retried | one trap, kmem.lock, a zeroed page |
store to read-only text (nowrite) |
vmfault: already mapped |
usertrap(): unexpected scause 0xf, kexit(-1) |
two printk lines, p->lock, wait_lock |
illegal instruction (csrr mhartid) |
nobody: not a page fault | killed, scause 0x2 |
same as above |
| kernel load from address 0 | kerneltrap: nothing to try |
panic: kerneltrap, machine stops |
everything |
The key ideas: sepc points at the faulting instruction, so a repaired fault is
invisible; p->sz is the kernel’s record of what the process is owed, and the PTE
bits say how it may be used; a user fault costs the user its process, never the
kernel its life; a kernel fault costs everything, and panic is written to need no
lock because it may already be holding them. And with three harts, even dying is a
concurrent operation: the other harts stop only when they next touch the console or
a lock the dead hart kept.
Tour 10 · wrap-up
| Lock | Taken in | Protects |
|---|---|---|
kmem.lock (spinlock) | kalloc from vmfault, and from walk when mappages needs a page-table page; kfree if mappages fails | The free-page list shared by all harts: no two faults get the same page |
p->sz and the user page table (no lock) | sys_sbrk, vmfault, mappages | Changed only on behalf of the process itself, which is single-threaded; other harts never touch them while it is alive |
pr.lock (spinlock) | printk, only while panicking == 0 | One formatted line at a time on the console; not a multi-line message |
p->lock (spinlock) | setkilled, killed, kkill, kexit | p->killed and p->state: a kill set on one hart is seen on another |
wait_lock (spinlock) | kexit | Parent links while the dying child reparents and wakes its parent |
ftable.lock (spinlock) | fileclose | Taken by kexit; see Tour 21: exit, wait and zombies |
log.lock (spinlock) | begin_op, end_op | Taken by kexit; see Tour 21: exit, wait and zombies |
itable.lock (spinlock) | iput | Taken by kexit; see Tour 21: exit, wait and zombies |
every other process's p->lock | wakeup from kexit | p->chan and p->state while waking the parent; see Tour 21: exit, wait and zombies |
tickslock (spinlock) | sys_uptime, clockintr | ticks; in scene 4 it is held forever by the panicked hart, so hart 0’s clock interrupt spins on it |
panicking, panicked (no lock: volatile flags) | panic, printk, uartputc_sync | Written once by the panicking hart and only ever set, never cleared; a lock would deadlock a hart that panics while holding pr.lock |
sepc, scause, stval (no lock: per-hart CSRs) | usertrap, kerneltrap | Each hart has its own trap CSRs, so simultaneous faults on three harts never see each other’s values |
After vmfault maps the page, why does usertrap not add 4 to p->trapframe->epc, as it does for a system call?
For ecall the instruction has completed and execution must continue after it. A faulting store has not happened; sepc points at it, and it must run again now that the page exists. Adding 4 would silently skip the store.
vmfault returns 0 if the page is already mapped. Which fault in this tour depends on that check, and what would happen without it?
The nowrite store to address 0: the text page is valid but read-only. Without the ismapped check, vmfault would call mappages on a valid PTE and hit panic("mappages: remap"), so a user program could crash the kernel by writing to its own code.
Why does vmfault test va >= psz before calling ismapped?
ismapped calls walk, which panics for va >= MAXVA. A user can make stval anything, for example 0xffffffffffffffff. Since psz never exceeds TRAPFRAME, the size check rejects every such address before it reaches walk.
Two harts take lazy page faults at the same moment, for two different processes. Which data do they share, and what keeps them from interfering?
They share only the free-page list, protected by kmem.lock in kalloc. Their trapframes, kernel stacks, page tables, p->sz and trap CSRs are all separate, so everything else runs without locks.
panic sets panicking = 1 before printing. Describe a situation in which taking pr.lock inside panic would hang the machine without printing anything.
If the panicking hart already holds pr.lock (it faulted inside printk), acquire sees it is already held by this hart and calls panic("acquire"), which would try the lock again, forever. If another hart holds it and is frozen, acquire would spin forever. Skipping the lock lets the message out in both cases.
In scene 4 the planted bug faults while holding tickslock on hart 2. Trace what happens to hart 0 over the next tenth of a second.
Hart 0’s next timer interrupt enters clockintr, which on hart 0 calls acquire(&tickslock). Hart 2 will never release it, so hart 0 spins in acquire forever with interrupts off. ticks stops advancing, and any process in pause() never wakes.
Keys: ← → step · Home start