Lab 15 · reveal · 18 steps · 10 commits
In this tree a process reaches a file’s bytes only through read and write: the kernel
copies them between the buffer cache and a buffer the program owns. In this lab you
add mmap, which puts part of a file into the address space: the program reads the file
by loading from memory and changes it by storing to memory. munmap takes the mapping
away again. Pages are read from the file only when they are first touched, and a shared
mapping’s changes go back to the file when it is unmapped.
The system call is two lines in a man page; the kernel work touches almost every part of
this tree. Where in the address space does a mapping go, when a heap already grows into
the same empty space? How does a page fault know which file and which bytes a
missing page stands for? Which pages were changed, and who noticed: the hardware, the
kernel, or neither? Writing a page back to a file is a file-system operation: what does
that require, and how much may one page cost? What do kfork, kexit and kexec,
which knew nothing about mappings, now get wrong? And the fault handler now reads a file:
what happens when the fault comes from inside a system call that already holds locks,
perhaps the lock of the very file being mapped? The think section asks these questions in
the order a designer meets them.
The reference solution is ten small commits. With it, touching 1 page of a 3-page mapping costs exactly 1 page of memory, and unmapping a 64-page mapping in which one page was changed writes one page back, not 64.
Each step shows one change on the branch ext/15-mmap, the code around it, and the state of the machine when that code runs.
kernel/proc.hStep 1 of 18 · commit 1: Add a per-process table of file mappings
The story of this tour was recorded with gdb on the finished branch, three harts:
mmaptest shared maps a file and changes it, fork hands a mapping to a child, locked
and self call read, write and wait with buffers inside a mapping, and stress
runs three processes against one file. Each step says which run its state comes from.
Commit 1 adds the record that question 1 designed. struct vma (Linux’s name, “virtual
memory area”) holds where the mapping starts and how long it is, prot and flags, the
file’s offset of start, and f, a counted reference to the open file. struct proc
gets a fixed array of NVMA (16, in param.h) of them; a slot with len 0 is free.
Each one is 40 bytes, so every struct proc grows by 640 bytes: gdb printed
sizeof(struct proc) = 1000 in this build, which is why the third process slot shows up
as proc+2000 in the backtraces below.
The array sits below the comment “these are private to the process, so p->lock need not
be held”, and that is where it belongs: only the process itself changes its mappings,
and kfork reads the parent’s while the parent is inside fork. That is also why the
fault path, later, reads the table with no lock at all.
sp = 0x3fffff9f60kernel/mmap.cStep 2 of 18 · commit 2: Add mmap: record a mapping of a file
sys_mmap fetches six arguments, the most any system call in this tree takes: a0
to a5 are all the argument registers argraw knows. Then it refuses, in order:
addr other than 0 (the kernel chooses), a zero or absurd len, an off that is
not a page boundary, and an off + len that wraps around;prot but read or read-write (W without R is reserved in a RISC-V PTE);flags but exactly one of shared and private;mmaptest private checks this one.This is mmaptest shared (pid 3), on hart 0, in a system call: interrupts on, nothing
held. gdb recorded the state at line 95, after all the checks; nothing on the way there
takes a lock.
Step 3 of 18 · commit 2: Add mmap: record a mapping of a file
A free slot, then a place: vmaplace (lines 34-51) starts with end at TRAPFRAME,
tries [end - len, end), and whenever an existing mapping overlaps it, moves end down
to that mapping’s start and tries again. Each retry lowers end, so the loop ends, at
the first gap that fits or at the heap, which it never crosses (line 41).
gdb at line 95: start = 0x3fffffb000, len = 0x3000, p->sz = 0x7000, file
inode 25 with f->ref 1. The mapping ends exactly at TRAPFRAME (0x3fffffe000), and
between it and the top of the heap lie about 256 GiB of empty address space.
Line 99 takes the reference: filedup makes f->ref 2, one for the descriptor and one
for the mapping, under ftable.lock for a moment (noff 1 inside filedup). The test
closes the descriptor right after mmap returns; the mapping keeps the file open. Nothing
else happens: no page is allocated, no PTE is written, no byte is read. mmap costs the
same for 1 page as for 1000.
kernel/proc.cStep 4 of 18 · commit 3: Keep the heap below the mappings
The heap had one fence, TRAPFRAME. Now it has vmabottom(p): the lowest address of
any mapping, or TRAPFRAME if there is none. growproc (eager sbrk) checks it here,
and sys_sbrk's lazy branch checks the same in sysproc.c. Without the fence, an eager
sbrk that starts just below a mapping (after lazy growth brought the heap there) would
run uvmalloc over it and mappages would panic on a loaded page (mappages: remap, reproduced on such a build); lazy growth would raise p->sz past mapped addresses, which
would confuse every function that walks [0, sz).
With no mappings vmabottom returns TRAPFRAME, so every sbrk behaves exactly as
before, which is why usertests cannot tell this commit apart from the original.
(No breakpoint was set here; the state is the ordinary one for a system call.)
sp = 0x3fffff9f40 in mmapfaultld sp, 8(a0) in uservec (kernel/trampoline.S:76), after the store page faultkernel/vm.cStep 5 of 18 · commit 4: Load a mapped page from its file on a fault
mmaptest shared closed the file and executed strcpy(m + 10, "hello"): a store to
0x3fffffb00a, in a page with no PTE. The hart raised a store page fault (scause 15),
usertrap called vmfault with read 0, and here is the one new test: the page
table is the process’s own, and vmalookup finds a mapping containing the address, so
the fault belongs to mmapfault.
The test must come before line 470. A mapped address is far above psz (p->sz,
0x7000 here), and the old first line would reject it. Below line 470 nothing changed:
the lazy heap works as before. The comparison with p->pagetable keeps kexec's
copies into a brand-new page table away from the old image’s mappings.
Interrupts are off, as on every page fault: usertrap turns them on only for system
calls (kernel/trap.c:66). No lock is held. gdb recorded exactly that on hart 0:
noff 0, SIE 0.
sp = 0x3fffff9f40kernel/mmap.cStep 6 of 18 · commit 4: Load a mapped page from its file on a fault
Two refusals first. A store (read 0) to a mapping without PROT_WRITE returns 0, and
usertrap kills the process: that is mmaptest private’s scause 0xf line. And a
fault on a page that is already mapped is about permissions, not about a missing page;
mapping it again would make mappages panic with remap.
Then a page from kalloc, zeroed at once. kalloc fills pages with 0x05
(kernel/kalloc.c:80), and any byte the file does not supply must read as 0: the end
of a last page that the file only partly fills, and whole pages past the end of the file.
Zeroing first and reading over it gives both for free.
gdb at line 92 for this fault: va = 0x3fffffb000, read 0, off 0, inode 25 with
size 0x2064 (8292 bytes: two pages and 100 bytes) and ref 1. The inode’s reference
is the struct file’s; the file’s own ref is 1 too, the mapping’s, because the test
has closed its descriptor.
the mapped file's inode lock (sleep-lock)Step 7 of 18 · commit 4: Load a mapped page from its file on a fault
The page’s file offset is v->off plus how far va is into the mapping. If that is
inside the file, readi copies min(4096, size - off) bytes into the front of the page,
with the inode locked: ilock is a sleep lock, ip->size is only stable under
it, and readi may wait for the disk in bread. Past the end of the file there is
nothing to read, and the zero page is the answer.
Then the PTE: PTE_R | PTE_U, and PTE_W for a writable mapping. When munmap later
read this PTE it was 0x21fc94d7: flags 0xd7 = V|R|W|U|A|D. A and D were set by
the hart when the retried strcpy used the page; the write-back steps rely on D.
gdb at line 93: the lock is held, lock.pid 3, interrupts still off, noff 0: the
sleep-lock is not a spinlock and is not counted. Sleeping here is legal: sched wants
exactly noff 1 (p->lock) and interrupts off. In this run the blocks were cached (the
test had just written the file); in the stress run, whose loads often missed the cache,
gdb caught loads asleep in virtio_disk_rw on all three harts, with noff 1
(disk.vdisk_lock) and intena 0, as a fault from user mode should show.
sp = 0x3fffff9f20 in vmaunmapkernel/mmap.cStep 8 of 18 · commit 5: Add munmap
sys_munmap (lines 188-207) accepts a range that is the whole mapping, or a part
of it starting at its start or ending at its end. A hole in the middle would leave two
pieces, which need two records; the reference refuses it (mmaptest partial checks the
refusal) rather than splitting.
vmaunmap frees the loaded pages with uvmunmap, which skips the pages that were
never touched (kernel/vm.c:205): an unmapped range with holes in it frees cleanly.
Then the record shrinks. Removing a part at the start moves start and off, so that
the remaining pages still find their own bytes in the file. When nothing is left,
fileclose drops the mapping’s reference; if it was the last, the file is closed, and
an unlinked file is freed inside fileclose's own transaction. Clinic 4 forgets this
line.
No TLB flush is needed here, though pages were just freed: this process will return to
user mode through userret, which runs sfence.vma after loading satp
(kernel/trampoline.S:112), and any other hart that still caches a translation of
this process flushes the same way before it runs user code again. Until then no user
code runs with this page table. (The sbrk that shrinks the heap relies on the same
argument.)
The state is from the finished branch’s munmap in mmaptest shared, recorded at the
first line of the write-back loop that commit 7 adds above line 173.
sp = 0x3fffff7f70kernel/proc.cStep 9 of 18 · commit 6: Unmap every mapping at exit and exec
Mapped pages lie above p->sz. The parent’s kwait frees this process’s page table
later with freeproc → uvmfree, which unmaps [0, sz) and then calls
freewalk; freewalk panics on any valid leaf it finds (kernel/vm.c:276). So
kexit removes every mapping before anything else, while the process can still sleep
(the write-back of commit 7 will need to) and holds nothing. A copy of the finished branch
without this call panicked with freewalk: leaf in mmaptest fork, inside the parent’s
wait on hart 2, with noff 2 (wait_lock and the child’s p->lock).
gdb stopped here in mmaptest fork’s child, pid 5, on hart 2: its p->vma[0] still had
len 0x2000 at 0x3fffffc000, the mapping it inherited (commit 8) and used.
Unmapping also drops the mappings’ file references, so an exiting process leaks no
struct file through its mappings, exactly as the loop below closes p->ofile.
sp = 0x3fffff7b30 in vmaunmapkernel/exec.cStep 10 of 18 · commit 6: Unmap every mapping at exit and exec
proc_freepagetable on line 139 frees the old page table with the same uvmfree
and freewalk, so kexec has the same problem as exit, and the same answer as
POSIX: a new program starts with no mappings.
The call sits after the commit (lines 133-137), not before. Up to line 132 exec can
still fail, and a failed exec must return to the old program with its mappings intact.
After line 134 the process runs on the new page table, so vmaunmapall is given the old
one explicitly: the write-back reads the old pages through their PTEs there.
gdb recorded this in mmaptest exit’s second child (pid 11), which stored into page 1
of a shared mapping and called exec: here the unmap found that page’s PTE
0x21fcb0d7, dirty, and wrote it back, starting on hart 2 and finishing (the
fileclose of the mapping) on hart 0. The new program never knew.
sp = 0x3fffff9f20kernel/mmap.cStep 11 of 18 · commit 7: Write dirty pages of shared mappings back to the file
Commit 7 makes unmapping write back. For a shared, writable mapping, the loop walks every
page of the range: a page that is loaded and has D in its PTE goes back to the file,
before uvmunmap frees it. PTE_D (bit 7) and PTE_A (bit 6) are new names in
riscv.h; the hardware has always had them.
gdb at line 204, mmaptest shared unmapping its 3 pages:
| page | PTE | flags | written back? |
|---|---|---|---|
0x3fffffb000 |
0x21fc94d7 |
V R W U A D |
yes (the strcpy) |
0x3fffffc000 |
0 |
never loaded | no |
0x3fffffd000 |
0x21fc98d7 |
V R W U A D |
yes (the X and Y) |
In mmaptest fork, the parent’s page 1, loaded after the child exited and only read, had
PTE 0x21fc8c57: flags 0x57 = V R W U A, accessed but not dirty, and it was not
written. Private mappings and read-only ones skip the loop entirely.
Why not write every loaded page? Clinic 6 tries it: 64 pages written per munmap instead
of 1, and in stress a process writes its stale copy of another process’s page over the
newer one.
the mapped file's inode lock (sleep-lock)Step 12 of 18 · commit 7: Write dirty pages of shared mappings back to the file
begin_op first, then ilock: the order filewrite uses, and the lock order of
the file system (a transaction is entered before any inode lock is taken). writei
copies from pa, the page’s physical address, which the kernel can use directly through
its direct map (user_src 0). Clinic 1 leaves out the transaction and gets panic: log_write outside of trans on the first write-back.
Only min(4096, size - off) bytes, and nothing for a page past the end: a mapping never
makes its file longer (clinic 2 writes the whole page and the file grows). Since nothing
past the end is written, nothing is allocated, and one page costs 4 data blocks plus the
inode block that writei's iupdate logs: 5, within MAXOPBLOCKS.
gdb at line 179: va 0x3fffffb000, pa 0x87f25000, off 0, inode 25 locked by pid
3, size 8292. In the frames, vmawrite is inlined into vmaunmap in this build (gdb
shows both at one address); the state’s stack lists it as the C code does.
pi->lockkernel/vm.cStep 13 of 18 · commit 7: Write dirty pages of shared mappings back to the file
copyout writes into a user page with memmove to pa0, the page’s physical address,
through the kernel’s direct map. The hart uses the kernel’s PTE for that store
and sets D there, if anywhere; the user PTE is untouched. A read() into a shared
mapping would put the file’s (or the pipe’s) bytes in the page and leave it “clean”, and
munmap would throw them away. So copyout sets PTE_D itself, right after it has
checked PTE_W.
mmaptest locked reads 8 bytes from a pipe into page 1 of a shared mapping: this line,
called from piperead with pi->lock held. At munmap gdb read that page’s PTE:
0x21fcc497, flags 0x97 = V R W U D, D from this line and no A at all, since no
user instruction ever used the page. A kernel without this line printed locked: FAIL (file contents) and self: FAIL (file contents).
(No breakpoint was set at this line in the pipe case; the state is reasoned: piperead
holds pi->lock, taken in a system call after interrupts were turned on, hence noff
1, intena 1. Hart 2 is where gdb saw this process’s prefault for the same read, in
commit 9’s code.)
np->lock (the child's, pid 5)kernel/proc.cStep 14 of 18 · commit 8: Give a forked child its parent's mappings
kfork copies the whole table and takes one more file reference for each mapping in
use, exactly as it does for p->ofile above. No page is copied: uvmcopy copied
[0, sz), and the child will load each mapped page from the file when it first touches
it. Clinic 3 removes these lines; the child’s first touch of the mapping falls through
to the heap code and kills it.
gdb at line 296 in mmaptest fork: parent pid 4 on hart 2, child pid 5, the mapping at
0x3fffffc000, f->ref 1 before this filedup (the parent had closed its
descriptor), 2 after. kfork still holds the child’s np->lock here (from
allocproc until line 302): noff 1, intena 1, interrupts off. Inside filedup,
ftable.lock nests under it for a few instructions: the p->lock → ftable edge that
kfork's p->ofile loop already creates (Locks and interrupt state).
What the child does not get: the parent stored P into page 0 before forking and has
not written it back. The child’s page 0 will come from the file, without the P. That is
the cost of this simple design; question 5 discusses what fixing it would take.
sp = 0x3fffff7ee0Step 15 of 18 · commit 8: Give a forked child its parent's mappings
The child (pid 5) read page 1 of the inherited mapping, which loaded it from the file on
hart 2, stored C into it, and called exit without munmap. kexit unmaps
everything (commit 6), and here the loop finds page 0 never loaded (PTE 0) and page 1
dirty (0x21fd0cd7), and writes page 1 back at file offset 0x1000 with inode 25 locked
by pid 5.
The write-back slept (log, inode lock, disk), and the child finished its unmap on hart 0:
gdb’s next stop, the fileclose of the mapping, was on hart 0, with the file’s ref
2 (the child’s mapping and the parent’s) about to become 1.
Then the parent (pid 4) returns from wait and loads page 1 through its mapping for
the first time: from the file, so it sees C. At its own munmap its page 0 (the P)
was dirty and was written; its page 1 was clean and was not. The file ends with both,
which is what mmaptest fork checks.
sp = 0x3fffff9f40kernel/mmap.cStep 16 of 18 · commit 9: Load mapped pages before read, write and wait take locks
vmaprefault(p, va, n) intersects [va, va+n) with every mapping and loads each page in
the intersection that is not mapped yet, through the same mmapfault. It holds nothing
while it does, so the load may sleep on the disk and take the inode lock freely. If a
page cannot be loaded, the system call fails with -1 before it has done anything.
This is mmaptest locked (pid 6) calling wait(&status) with status in page 2 of a
shared mapping, 0x3fffffd000. gdb stopped in this loop on hart 2 with interrupts on and
no lock; the load that followed locked inode 25 for pid 6 and returned. Moments later
kwait, holding wait_lock and the child’s p->lock, copied 7 into a page that was
there: no load, no wakeup, no panic: acquire (clinic 5).
Why may the page not vanish between this loop and the copy? Only this process can unmap
its pages (munmap, exit, exec, sbrk), and this process is busy in this system
call; xv6 processes have one thread. Lab 12 relies on the same argument for pages of the
program.
kernel/sysfile.cStep 17 of 18 · commit 9: Load mapped pages before read, write and wait take locks
sys_read and sys_write call the prefault after decoding their arguments and
before fileread or filewrite, the last point at which the call holds no lock of
any kind. sys_wait does the same for the status word in sysproc.c. What each copy
would otherwise hold:
| the copy | locks held during it | without the prefault |
|---|---|---|
piperead, pipewrite |
pi->lock |
panic: sched locks whenever the load sleeps |
consoleread |
cons.lock |
the same |
kwait |
wait_lock, child’s p->lock |
panic: acquire, every time |
fileread, filewrite |
inode lock, buffer lock (sleep-locks) | hang, if it is the mapped file |
This is mmaptest self (pid 8) on hart 1, reading its file into page 1 of its own
mapping. The prefault locked inode 25, loaded the page and unlocked it; then
fileread locked inode 25 again and readi copied into a page that was present.
Clinic 5 without these lines: the same read sleeps forever on its own inode lock.
user/mmaptest.cStep 18 of 18 · commit 10: Add mmaptest, a test program for mmap and munmap
The last check. Three children each run 30 rounds: open, map all 12 pages shared, check
that every page is uniform (a page is written back in one writei under the inode lock,
and loaded in one readi under it, so no process may ever see half of a write-back),
fill its own 4 pages with (c + 1) * 1000 + r, unmap. The parent then reads the file and
expects every child’s last round: 1029, 2029, 3029. It also counts free pages before and
after.
gdb on the stress run saw it spread over the machine: loads and write-backs on all three
harts, loads asleep on the disk, write-backs asleep in begin_op waiting for log
space, and processes moving between harts from one write-back to the next. The check passed 10 times out of 10 on the reference and failed 5 times out of 6 with
clinic 6’s clean write-backs.
What the branch cost: ten commits, one new file of about 280 lines, a dozen lines in
vm.c, proc.c, exec.c, sysproc.c and sysfile.c. What it bought: a file in the
address space, read lazily (one page of memory per page touched) and written back only
where it changed. What it exposed: three structures that walk [0, sz) and never saw
the new region, a PTE bit the kernel’s own writes do not set, and a fault handler that
sleeps on an inode lock, reachable from code that holds locks.
Lab 15 · wrap-up
On the branch (ext/15-mmap, 10 commits), built with the project toolchain and run on 3
harts (-smp 3 -m 128M), one boot:
$ mmaptest
mmaptest: read: touching 1 of 3 mapped pages took 1 free pages
usertrap(): unexpected scause 0xd pid=4
sepc=0x1e stval=0x3fffffb000
mmaptest: read: OK
mmaptest: shared: OK
usertrap(): unexpected scause 0xf pid=5
sepc=0xa stval=0x3fffffc000
mmaptest: private: OK
usertrap(): unexpected scause 0xd pid=6
sepc=0x1e stval=0x3fffffa000
usertrap(): unexpected scause 0xd pid=7
sepc=0x1e stval=0x3fffffd000
mmaptest: partial: OK
mmaptest: fork: OK
mmaptest: exit: OK
mmaptest: many: OK
mmaptest: locked: OK
mmaptest: self: OK
mmaptest: stress: OK
mmaptest: ALL OK
$ usertests -q
usertests starting
test copyin: OK
test copyout: OK
[...]
test nowrite: usertrap(): unexpected scause 0xf pid=6576
[...]
ALL TESTS PASSED
$ mmaptest
mmaptest: read: touching 1 of 3 mapped pages took 1 free pages
usertrap(): unexpected scause 0xd pid=6662
sepc=0x1e stval=0x3fffffb000
mmaptest: read: OK
[...]
mmaptest: stress: OK
mmaptest: ALL OK
The usertrap() lines in mmaptest are the expected kills: in read a child loads from the
mapping after munmap (stval is its first page); in private a child stores into a
read-only mapping (scause 0xf); in partial two children load from the first and the
last page after they were unmapped. A second boot printed the same, with the same pids.
usertests -q passing shows that what the original kernel did still works: the heap’s
sbrk tests (eager, lazy, shrinking, running out of memory) see the same limit as before
when no mapping exists, copyout into text still fails, and no page is lost. Every commit
builds with make kernel/kernel fs.img, and usertests -q printed ALL TESTS PASSED on 3
harts at each of the 10 commits.
mmaptest stress alone, 10 times in a row on one boot: 10 times OK.
Keys: ← → step · Home start