xv6, line by line
lab 15

Extension labs · lab 15 · Memory · ★★★★★

mmap and munmap of files

In this tree a process reaches a file’s bytes only through read and write: the kernel copies them between the buffer cache and a buffer the program owns. In this lab you add mmap, which puts part of a file into the address space: the program reads the file by loading from memory and changes it by storing to memory. munmap takes the mapping away again. Pages are read from the file only when they are first touched, and a shared mapping’s changes go back to the file when it is unmapped.

The system call is two lines in a man page; the kernel work touches almost every part of this tree. Where in the address space does a mapping go, when a heap already grows into the same empty space? How does a page fault know which file and which bytes a missing page stands for? Which pages were changed, and who noticed: the hardware, the kernel, or neither? Writing a page back to a file is a file-system operation: what does that require, and how much may one page cost? What do kfork, kexit and kexec, which knew nothing about mappings, now get wrong? And the fault handler now reads a file: what happens when the fault comes from inside a system call that already holds locks, perhaps the lock of the very file being mapped? The think section asks these questions in the order a designer meets them.

The reference solution is ten small commits. With it, touching 1 page of a 3-page mapping costs exactly 1 page of memory, and unmapping a 64-page mapping in which one page was changed writes one page back, not 64.

Read first: Tour 20: fork, Tour 21: exit, wait and zombies, Tour 22: exec, Tour 25: A user address space, Tour 26: sbrk, eager and lazy, and page faults, Tour 28: Crossing the user/kernel boundary in memory, Tour 31: The log: begin_op, commit and group commit, Tour 33: The life of an inode, Tour 36: Reading and writing a file, Tour 39: Pipes · The stacks of xv6, Locks and interrupt state

What this lab teaches

  • What a kernel must remember about a mapping so that a page fault, possibly much later, can build the missing page byte for byte, and why that record has to hold a counted reference to an open file.
  • How to fit a new region into an address space that already has a growing heap, and why the functions that copy and free address spaces in this tree do not see it.
  • How a RISC-V PTE records that a page has been written, who sets that bit, and which writes into user memory do not set it.
  • What it takes to write memory back into a file: which locks, which transaction, how many log blocks, and which bytes must never be written.
  • How fork, exit and exec each have to learn about a new kind of per-process state, and what each one does wrong when it does not.
  • Which kernel paths can reach a fault handler that sleeps and takes an inode’s sleep-lock, and three different ways it fails when the path already holds a lock.

The reference branch

ext/15-mmap in ShowMeTheStack/xv6-riscv-labs, branched from the frozen commit 06aad25; 10 commits.

git clone https://github.com/ShowMeTheStack/xv6-riscv-labs
cd xv6-riscv-labs
git checkout -b my-mmap 06aad25   # start your own
git diff 06aad25 origin/ext/15-mmap   # only when you want the answer

1. The spec

Behaviour. Two new system calls, in user/user.h:

void *mmap(void *addr, uint64 len, int prot, int flags, int fd, uint64 off);
int munmap(void *addr, uint64 len);

What must not change.

The test program, mmaptest, prints one line per check (mmaptest NAME runs one):

$ mmaptest
mmaptest: read: touching 1 of 3 mapped pages took 1 free pages
usertrap(): unexpected scause 0xd pid=4
            sepc=0x1e stval=0x3fffffb000
mmaptest: read: OK
mmaptest: shared: OK
usertrap(): unexpected scause 0xf pid=5
            sepc=0xa stval=0x3fffffc000
mmaptest: private: OK
usertrap(): unexpected scause 0xd pid=6
            sepc=0x1e stval=0x3fffffa000
usertrap(): unexpected scause 0xd pid=7
            sepc=0x1e stval=0x3fffffd000
mmaptest: partial: OK
mmaptest: fork: OK
mmaptest: exit: OK
mmaptest: many: OK
mmaptest: locked: OK
mmaptest: self: OK
mmaptest: stress: OK
mmaptest: ALL OK

The usertrap() lines are expected: each is a child the test starts on purpose to touch an address it must not touch, and the test checks that the child was killed.

2. Think first

Answer each question in your head (or on paper) before opening a hint. Hints get more specific; the reference answer comes last.

1What does mmap have to remember?

mmap maps no page. The first touch of a page might come seconds later, from user code or from inside a system call. At that moment the kernel must build the page exactly: which file, which bytes of it, which PTE bits. What does mmap have to record, and where? The program may close(fd) right after mmap, or even unlink the file. What keeps the file usable, and what kind of reference is the right one to keep?

Check yourself

1warm-upChoose all that apply

A page of a mapping is first touched long after mmap returned. Which of these must mmap have recorded so that the fault can build the page?

2solidTrue or false, and why

True or false: a program that calls close(fd) right after mmap(..., fd, 0) can no longer use the mapping.

Why?

2Where does a mapping go?

Look at a user address space in this tree (kernel/memlayout.h:54): program text and data at 0, the guard page and the stack, then the heap growing up with sbrk toward TRAPFRAME. Where will you put a mapping? What stops the heap from growing into it later, eagerly or lazily, and what stops a mapping from landing on the heap? And which functions that walk or copy an address space will see mapped pages, and which will not?

Check yourself

1solidType a number

mmaptest read maps 3 pages when the process has no other mapping. TRAPFRAME is 0x3fffffe000 in this tree. At what address does the reference place the mapping?

decimal, 0x hex or 0b binary
2solidChoose one

A kernel has the mapping table and mmap, but growproc and sys_sbrk still check against TRAPFRAME. A process maps a file and touches the mapping’s first page, grows its heap lazily (sbrklazy, many calls) to one page below the mapping, then calls the eager sbrk(8192). What happens?

3What happens on the first touch?

A program loads from a page of a mapping that has not been touched. What does the hart report, which code in this tree sees it first, and what must change there so that the page comes from the file and the lazy heap keeps working? What must the new page contain, byte by byte, when the mapping extends past the end of the file? Which PTE bits does it get? And what must happen to a store to a read-only mapping, or to a page that is already loaded?

Check yourself

1solidChoose one

mmaptest private maps a file PROT_READ, MAP_SHARED, and a child stores one byte into it. The page is not loaded yet. What happens, on the reference branch?

2solidType a number

A file is 2 pages and 1000 bytes long (9192 bytes) and is mapped with len 3 pages at offset 0. When page 2 of the mapping is loaded, how many bytes does mmapfault read from the file?

decimal, 0x hex or 0b binary

4Which pages go back to the file, and how?

When a shared writable mapping is unmapped, its changes must reach the file. Which pages must be written: all loaded ones, or only some? How does the kernel know which? Remember that the kernel itself also writes into user pages (a read into a mapped buffer). How do you write one page into a file safely: which locks, which transaction, how many log blocks? And which bytes of the page must never be written?

Check yourself

1solidDecode the bits

gdb read this PTE in mmaptest locked, when page 1 of a shared writable mapping was unmapped. Before that the program had called read from a pipe into the page, and had never touched it from user code. Decode the flags.

Value: 0x21fcc497

2deepType a number

vmawrite writes one page of a mapping back, inside a file whose size is a multiple of the page size and with blocks of 1024 bytes. How many blocks does that transaction log?

decimal, 0x hex or 0b binary

5What do fork, exit and exec now get wrong?

Three functions in this tree make, copy or destroy whole address spaces, and none of them knows about mappings: kfork, kexit (with freeproc in the parent’s wait) and kexec. For each one: what goes wrong if you leave it as it is, and what is the simplest correct change? For fork, decide what the child should see in a mapped page that the parent changed and has not yet written back.

Check yourself

1solidChoose one

A kernel is complete except that kexit does not unmap the process’s mappings. In mmaptest fork the child loads and changes page 1 of a mapping, then exits. What happens?

2deepTrue or false, and why

True or false: on the reference branch, if a parent stores into page 0 of a shared mapping and then forks, the child reading page 0 sees the parent’s store.

Why?

6Where can a load run, and what may it hold?

Loading a mapped page takes the file’s inode lock (a sleep-lock) and may wait for the disk. Which paths can reach the load? List them, and for each, which locks the thread already holds at that moment. What goes wrong on each path that holds a lock, and what will you do about it? Think in particular about a read from a file into a mapping of that same file.

Check yourself

1deepFill in the machine state

Without the prefault, mmaptest locked calls wait with the status in an unloaded mapped page. kwait's copyout loads the page; readi releases a buffer, and wakeup is about to acquire the zombie child’s p->lock. Fill in the hart’s state at that moment, as gdb recorded it at the resulting panic.

kernel/proc.c
578 struct proc *p;
580 for (p = proc; p < &proc[NPROC]; p++) {
582 if (p->chan == chan) {
583 // If the process is waiting for wakeups on this channel,
584 // signal that the wakeup happened by clearing p->chan.
2solidChoose one

Without the prefault, mmaptest self reads its file into an unloaded page of a mapping of the same file. What does the console show?

7How do you know it works, and that nothing leaks?

Many of this lab’s bugs leave the kernel running: a page goes to the wrong place, a file reference is never dropped, an update is lost when two processes write. What can a user program check, and how can it make such a bug visible? In particular, if munmap forgot to drop the mapping’s file reference, what would a user program notice, and after how long?

Check yourself

1solidChoose one

A kernel’s munmap frees the pages and clears the mapping but never calls fileclose. You run mmaptest many, then mmaptest, then ls. What do you see?

3. Build it

Start.

git checkout -b my-15 06aad25

Add the system calls first, with nothing behind them: numbers in kernel/syscall.h, entries in kernel/syscall.c and user/usys.pl, prototypes in user/user.h, the PROT_ and MAP_ constants in kernel/fcntl.h. A new file kernel/mmap.c (add $K/mmap.o to OBJS in the Makefile) keeps the mapping code in one place. Then write user/mmaptest.c from the spec’s list of checks and add $U/_mmaptest\ to UPROGS. Every check fails at first; that is the test working.

Milestones, in an order that keeps the system bootable after each one. usertests never calls mmap, so usertests -q must pass after every milestone.

  1. The table. struct vma and p->vma[NVMA] in kernel/proc.h. Test: it boots.
  2. mmap records. Check the arguments, find a free slot and a place, take a reference to the file, record. Test: mmap returns an address near the top (print it).
  3. The fence. growproc and sys_sbrk stop at the lowest mapping. Test: usertests -q (its sbrk tests).
  4. Load on fault. The test at the top of vmfault, and the loader. Test: the first check of mmaptest read. Until milestone 6, a process that exits with a loaded mapped page panics the kernel with freewalk: leaf (question 5), so expect that when the test exits.
  5. munmap, whole and at either end, without write-back. Test: mmaptest read, then partial.
  6. exit and exec unmap everything. Test: mmaptest read, private.
  7. Write-back. The D test, the transaction, the size limit, and copyout setting D. Test: shared, partial, exit.
  8. fork. Test: fork.
  9. Load before you lock. Test: locked, self, then the whole of mmaptest and usertests -q, and mmaptest once more.

Debugging advice. Start QEMU with make qemu-gdb and attach gdb on the port it prints; mapping code runs only when your test runs, so let the system boot, set your breakpoints, then run the test.

4. Debugging clinic

Each of these bugs was put into the reference solution on purpose and run on three harts. The symptom is exactly what happened. Try to explain it before revealing why.

1The write-back runs outside a transaction

vmawrite locks the inode and calls writei, but nobody opened a transaction: it looks like a plain write of kernel memory into a file.

-  begin_op();
   ilock(ip);
   if (off < ip->size) {
     [...]
     writei(ip, 0, pa, off, n);
   }
   iunlock(ip);
-  end_op();

What happened when we ran it

xv6 kernel is booting

hart 1 starting
hart 2 starting
init: starting sh
$ mmaptest shared
panic: log_write outside of trans

# the same boot, gdb breakpoint on panic:
#0  panic (s=s@entry=0x80008548 "log_write outside of trans") at kernel/printk.c:139
#1  0x0000000080003f5a in log_write (b=b@entry=0x800274a8 <bcache+27824>) at kernel/log.c:232
#2  0x00000000800037e8 in writei (ip=ip@entry=0x80029000 <itable+296>, user_src=user_src@entry=0, src=2280804352, off=off@entry=0, n=4096) at kernel/fs.c:566
#3  0x0000000080005028 in vmawrite (v=0x80011708 <proc+2344>, va=274877886464, pa=2280804352) at kernel/mmap.c:207
#4  vmaunmap (pagetable=0x87f24000, v=0x80011708 <proc+2344>, addr=addr@entry=274877886464, len=len@entry=12288) at kernel/mmap.c:228
#5  0x0000000080005184 in sys_munmap () at kernel/mmap.c:276
#6  0x0000000080002944 in syscall () at kernel/syscall.c:150
#7  0x00000000800026ca in usertrap () at kernel/trap.c:68
#8  0x0000003ffffff09c in ?? ()
  Id   Target Id                    Frame
* 1    Thread 1.1 (CPU#0 [running]) panic (s=s@entry=0x80008548 "log_write outside of trans") at kernel/printk.c:139
  2    Thread 1.2 (CPU#1 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
  3    Thread 1.3 (CPU#2 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
hart 0 noff 1 intena 1 pid 3 name mmaptest

2The whole page is written back

vmawrite keeps the “inside the file” test for the start of the page but always writes 4096 bytes:

   if (off < ip->size) {
     n = PGSIZE;
-    if (ip->size - off < PGSIZE)
-      n = ip->size - off;
     writei(ip, 0, pa, off, n);
   }

What happened when we ran it

$ mmaptest
mmaptest: read: touching 1 of 3 mapped pages took 1 free pages
usertrap(): unexpected scause 0xd pid=4
            sepc=0x1e stval=0x3fffffb000
mmaptest: read: OK
mmaptest: shared: FAIL (the file changed size)
mmaptest: shared: FAIL (page 2 not written back)
usertrap(): unexpected scause 0xf pid=5
            sepc=0xa stval=0x3fffffc000
mmaptest: private: OK
[...]
mmaptest: self: OK
mmaptest: stress: OK
mmaptest: SOME FAILED

3fork does not give the child the mappings

kfork copies the open files and the current directory, as before, and nothing else:

   np->cwd = idup(p->cwd);

-  memmove(np->vma, p->vma, sizeof(p->vma));
-  for (i = 0; i < NVMA; i++)
-    if (np->vma[i].len != 0)
-      filedup(np->vma[i].f);

What happened when we ran it

$ mmaptest
[...]
mmaptest: partial: OK
usertrap(): unexpected scause 0xd pid=8
            sepc=0x116 stval=0x3fffffd000
mmaptest: fork: FAIL (child)
mmaptest: fork: FAIL (the child's store did not reach the file)
mmaptest: fork: FAIL (file)
mmaptest: exit: OK
mmaptest: many: OK
mmaptest: locked: OK
mmaptest: self: OK
mmaptest: stress: OK
mmaptest: SOME FAILED

4munmap forgets the file reference

When the last page of a mapping goes, the record is cleared but the file is not closed:

   v->len -= len;
-  if (v->len == 0) {
-    fileclose(v->f);
+  if (v->len == 0)
     v->f = 0;
-  }

What happened when we ran it

xv6 kernel is booting

hart 1 starting
hart 2 starting
init: starting sh
$ mmaptest many
mmaptest: many: FAIL (open failed: files leak)
$ mmaptest
mmaptest: cannot create mm.read
$ ls
ls: cannot open .
$

5No prefault before read, write and wait

sys_read, sys_write and sys_wait go straight to the code that copies; a missing mapped page is loaded wherever the copy happens to find it:

   if (argfd(0, 0, &f) < 0)
     return -1;
-  // load the buffer's mapped pages now: fileread holds locks while
-  // it copies, and loading a mapped page needs to sleep.
-  if (n > 0 && vmaprefault(myproc(), p, n) < 0)
-    return -1;
   return fileread(f, p, n);

(and the same in sys_write, one line in sys_wait.)

What happened when we ran it

xv6 kernel is booting

hart 2 starting
hart 1 starting
init: starting sh
$ mmaptest locked
panic: acquire

# the same boot, gdb breakpoint on panic:
#0  panic (s=s@entry=0x80008048 "acquire") at kernel/printk.c:139
#1  0x0000000080000c28 in acquire (lk=lk@entry=0x80011998 <proc+3000>) at kernel/spinlock.c:26
#2  0x0000000080002008 in wakeup (chan=chan@entry=0x80022230 <bcache+6712>) at kernel/proc.c:592
#3  0x0000000080004092 in releasesleep (lk=lk@entry=0x80022230 <bcache+6712>) at kernel/sleeplock.c:42
#4  0x0000000080002d22 in brelse (b=b@entry=0x80022220 <bcache+6696>) at kernel/bio.c:122
#5  0x00000000800036da in readi (ip=ip@entry=0x80029000 <itable+296>, user_dst=user_dst@entry=0, dst=dst@entry=2280845312, off=off@entry=8192, n=n@entry=4096) at kernel/fs.c:531
#6  0x0000000080004cd8 in mmapfault (p=p@entry=0x800115b0 <proc+2000>, v=0x80011708 <proc+2344>, va=va@entry=274877894656, read=read@entry=0) at kernel/mmap.c:97
#7  0x000000008000151a in vmfault (pagetable=pagetable@entry=0x87f24000, psz=psz@entry=28672, va=va@entry=274877894656, read=read@entry=0) at kernel/vm.c:471
#8  0x00000000800015d0 in copyout (pagetable=0x87f24000, psz=28672, dstva=dstva@entry=274877894656, src=src@entry=0x800119c4 <proc+3044> "\a", len=len@entry=4) at kernel/vm.c:357
#9  0x0000000080002250 in kwait (addr=274877894656) at kernel/proc.c:402
#10 0x00000000800029d6 in sys_wait () at kernel/sysproc.c:37
#11 0x0000000080002944 in syscall () at kernel/syscall.c:150
#12 0x00000000800026ca in usertrap () at kernel/trap.c:68
#13 0x0000003ffffff09c in ?? ()
  Id   Target Id                    Frame
* 1    Thread 1.1 (CPU#0 [running]) panic (s=s@entry=0x80008048 "acquire") at kernel/printk.c:139
  2    Thread 1.2 (CPU#1 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
  3    Thread 1.3 (CPU#2 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
hart 0 noff 4 intena 1 pid 3 name mmaptest

# another boot of the same kernel: `mmaptest self`, which never returns
$ mmaptest self
# (nothing more; after 60 seconds gdb was attached to the hung kernel)
=== harts
  Id   Target Id                    Frame
* 1    Thread 1.1 (CPU#0 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
  2    Thread 1.2 (CPU#1 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
  3    Thread 1.3 (CPU#2 [halted ]) s_sstatus (x=2) at kernel/riscv.h:67
=== processes
proc[0] pid 1 SLEEPING init chan 0x80010de0 <proc>
proc[1] pid 2 SLEEPING sh chan 0x800111c8 <proc+1000>
proc[2] pid 3 SLEEPING mmaptest chan 0x80029010 <itable+312>
  vma[0] start 0x3fffffb000 len 0x3000 inode 25 lock.locked 1 lock.pid 3 &lock 0x80029010 <itable+312>
[...]
  kernel stack of pid 3 (from p->context):
#0  r_tp () at kernel/riscv.h:344
#1  cpuid () at kernel/proc.c:67
#2  mycpu () at kernel/proc.c:76
#3  sched () at kernel/proc.c:507
#4  0x0000000080001fba in sleep () at kernel/proc.c:580
#5  0x0000000080004044 in acquiresleep (lk=lk@entry=0x80029010 <itable+312>) at kernel/sleeplock.c:28
#6  0x00000000800032ac in ilock (ip=ip@entry=0x80029000 <itable+296>) at kernel/fs.c:303
#7  0x0000000080004ca4 in mmapfault (p=0x800109b0 <pid_lock>, p@entry=0x800115b0 <proc+2000>, v=0x80011708 <proc+2344>, va=1, va@entry=274877890560, read=read@entry=0) at kernel/mmap.c:92
#8  0x000000008000151a in vmfault (pagetable=pagetable@entry=0x87f24000, psz=28672, psz@entry=0, va=va@entry=274877890560, read=read@entry=0) at kernel/vm.c:471
#9  0x00000000800015d0 in copyout (pagetable=0x87f24000, psz=0, dstva=dstva@entry=274877890560, src=src@entry=0x80027500 <bcache+27912> "", len=len@entry=1024) at kernel/vm.c:357
#10 0x000000008000232a in either_copyout (user_dst=user_dst@entry=1, dst=dst@entry=274877890560, src=src@entry=0x80027500 <bcache+27912>, len=len@entry=1024) at kernel/proc.c:662
#11 0x00000000800036d0 in readi (ip=0x80029000 <itable+296>, user_dst=user_dst@entry=1, dst=dst@entry=274877890560, off=0, n=n@entry=4096) at kernel/fs.c:526
#12 0x0000000080004326 in fileread (f=0x8002ab30 <ftable+104>, addr=274877890560, n=4096) at kernel/file.c:122
#13 0x0000000080005422 in sys_read () at kernel/sysfile.c:81
#14 0x0000000080002944 in syscall () at kernel/syscall.c:150
#15 0x00000000800026ca in usertrap () at kernel/trap.c:68
#16 0x0000003ffffff09c in ?? ()

6Every loaded page is written back, clean or not

The D test is dropped, as if writing a clean page back were merely wasted work:

      pte = walk(pagetable, a, 0);
-     if (pte && (*pte & PTE_V) && (*pte & PTE_D))
+     if (pte && (*pte & PTE_V))
        vmawrite(v, a, PTE2PA(*pte)); // written since it was loaded

What happened when we ran it

$ mmaptest
[...]
mmaptest: locked: OK
mmaptest: self: OK
mmaptest: stress: FAIL (file contents)
mmaptest: SOME FAILED
$ mmaptest stress
mmaptest: stress: FAIL (file contents)
$ mmaptest stress
mmaptest: stress: FAIL (file contents)
$ mmaptest stress
mmaptest: stress: FAIL (file contents)
$ mmaptest stress
mmaptest: stress: FAIL (file contents)
$ mmaptest stress
mmaptest: stress: OK

# another boot: a copy of this kernel that prints how many pages each large munmap
# wrote back, running a scratch benchmark (not on the branch): 10 rounds of map 64
# pages, read every page, store into one, munmap
$ mmapbench
munmap: 64 pages written back
[...]
munmap: 64 pages written back
mmapbench: 10 rounds of: map 64 pages, read all, write 1, munmap: 60 ticks
$ mmapbench
[...]
mmapbench: 10 rounds of: map 64 pages, read all, write 1, munmap: 100 ticks
$ mmapbench
[...]
mmapbench: 10 rounds of: map 64 pages, read all, write 1, munmap: 53 ticks

# the same instrumentation on the reference branch, another boot:
$ mmapbench
munmap: 1 pages written back
[...]
mmapbench: 10 rounds of: map 64 pages, read all, write 1, munmap: 8 ticks
[...]
mmapbench: 10 rounds of: map 64 pages, read all, write 1, munmap: 8 ticks
[...]
mmapbench: 10 rounds of: map 64 pages, read all, write 1, munmap: 7 ticks

5. The reference solution

Take the guided tour through the reference solution, one commit at a time, with the machine state at every step:

Open the reveal tour →

Or read the commits

  1. 7f7a26e Add a per-process table of file mappings

    kernel/fcntl.h

    @@ -2,4 +2,10 @@
    22#define O_WRONLY 0x001
    33#define O_RDWR 0x002
    44#define O_CREATE 0x200
    55#define O_TRUNC 0x400
    6
    7// mmap
    8#define PROT_READ 0x1
    9#define PROT_WRITE 0x2
    10#define MAP_SHARED 0x01 // stores reach the file
    11#define MAP_PRIVATE 0x02 // stores stay in this process

    kernel/param.h

    @@ -5,8 +5,9 @@
    55#define NINODE 50 // maximum number of active i-nodes
    66#define NDEV 10 // maximum major device number
    77#define ROOTDEV 1 // device number of file system root disk
    88#define MAXARG 32 // max exec arguments
    9#define NVMA 16 // max file mappings per process
    910#define MAXOPBLOCKS 10 // max # of blocks any FS op writes
    1011#define LOGBLOCKS (MAXOPBLOCKS * 3) // max data blocks in on-disk log
    1112#define NBUF (MAXOPBLOCKS * 3) // size of disk block cache
    1213#define FSSIZE 2000 // size of file system in blocks

    kernel/proc.h

    @@ -77,8 +77,18 @@ struct trapframe {
    7777};
    7878
    8080
    81// A mapping of part of a file into the address space (mmap).
    82struct vma {
    83 uint64 start; // first virtual address, page-aligned
    84 uint64 len; // bytes, a multiple of PGSIZE; 0 = slot unused
    85 int prot; // PROT_READ, PROT_WRITE
    86 int flags; // MAP_SHARED or MAP_PRIVATE
    87 struct file *f; // the mapped file: a counted reference
    88 uint64 off; // offset in the file of start, page-aligned
    89};
    90
    8191// Per-process state
    8292struct proc {
    8393 struct spinlock lock;
    8494
    @@ -99,6 +109,7 @@ struct proc {
    99109 struct trapframe *trapframe; // data page for trampoline.S
    100110 struct context context; // swtch() here to run process
    101111 struct file *ofile[NOFILE]; // Open files
    102112 struct inode *cwd; // Current directory
    113 struct vma vma[NVMA]; // File mappings (mmap)
    103114 char name[16]; // Process name (debugging)
    104115};
  2. 4f8474c Add mmap: record a mapping of a file

    Makefile

    @@ -24,8 +24,9 @@ OBJS = \
    2424 $K/sleeplock.o \
    2525 $K/file.o \
    2626 $K/pipe.o \
    2727 $K/exec.o \
    28 $K/mmap.o \
    2829 $K/sysfile.o \
    2930 $K/kernelvec.o \
    3031 $K/plic.o \
    3132 $K/virtio_disk.o

    kernel/defs.h

    @@ -8,8 +8,9 @@ struct proc;
    88struct spinlock;
    99struct sleeplock;
    1010struct stat;
    1111struct superblock;
    12struct vma;
    1213
    1314// bio.c
    1415void binit(void);
    1516struct buf* bread(uint, uint);
    @@ -66,8 +67,11 @@ void initlog(int, struct superblock*);
    6667void log_write(struct buf*);
    6768void begin_op(void);
    6869void end_op(void);
    6970
    71// mmap.c
    72struct vma* vmalookup(struct proc*, uint64);
    73
    7074// pipe.c
    7175int pipealloc(struct file**, struct file**);
    7276void pipeclose(struct pipe*, int);
    7377int piperead(struct pipe*, uint64, int);

    kernel/mmap.c

    @@ -0,0 +1,102 @@
    1//
    2// mmap and munmap: mappings of files into a process's address
    3// space. mmap only records a mapping in p->vma[]; its pages are
    4// read from the file when they are first touched.
    5//
    6
    7#include "types.h"
    8#include "riscv.h"
    9#include "defs.h"
    10#include "param.h"
    11#include "memlayout.h"
    12#include "spinlock.h"
    13#include "proc.h"
    14#include "fs.h"
    15#include "sleeplock.h"
    16#include "file.h"
    17#include "fcntl.h"
    18
    19// The mapping of p that contains va, or 0.
    20struct vma *
    21vmalookup(struct proc *p, uint64 va)
    22{
    23 struct vma *v;
    24
    25 for (v = p->vma; v < &p->vma[NVMA]; v++)
    26 if (v->len != 0 && va >= v->start && va < v->start + v->len)
    27 return v;
    28 return 0;
    29}
    30
    31// Choose where a new mapping of len bytes goes: the highest free
    32// range below TRAPFRAME, and never below the heap. Returns its
    33// start, or 0 if there is no room.
    34static uint64
    35vmaplace(struct proc *p, uint64 len)
    36{
    37 uint64 end = TRAPFRAME, a;
    38 struct vma *v;
    39
    40 for (;;) {
    41 if (len > end || end - len < PGROUNDUP(p->sz))
    42 return 0;
    43 a = end - len;
    44 for (v = p->vma; v < &p->vma[NVMA]; v++)
    45 if (v->len != 0 && v->start < end && a < v->start + v->len)
    46 break;
    47 if (v == &p->vma[NVMA])
    48 return a; // [a, end) overlaps no mapping
    49 end = v->start; // try again just below the mapping in the way
    50 }
    51}
    52
    53// void *mmap(void *addr, uint64 len, int prot, int flags, int fd,
    54// uint64 off)
    55// addr must be 0: the kernel chooses the address.
    56uint64
    57sys_mmap(void)
    58{
    59 uint64 addr, len, off, start;
    60 int prot, flags, fd;
    61 struct proc *p = myproc();
    62 struct file *f;
    63 struct vma *v;
    64
    65 argaddr(0, &addr);
    66 argaddr(1, &len);
    67 argint(2, &prot);
    68 argint(3, &flags);
    69 argint(4, &fd);
    70 argaddr(5, &off);
    71
    72 if (addr != 0 || len == 0 || len > TRAPFRAME || off % PGSIZE != 0 ||
    73 off + len < off)
    74 return -1;
    75 if (prot != PROT_READ && prot != (PROT_READ | PROT_WRITE))
    76 return -1;
    77 if (flags != MAP_SHARED && flags != MAP_PRIVATE)
    78 return -1;
    79 if (fd < 0 || fd >= NOFILE || (f = p->ofile[fd]) == 0)
    80 return -1;
    81 if (f->type != FD_INODE || !f->readable)
    82 return -1; // pages are read from an inode: no pipes, no devices
    83 if (flags == MAP_SHARED && (prot & PROT_WRITE) && !f->writable)
    84 return -1; // stores would reach a file opened read-only
    85
    86 for (v = p->vma; v < &p->vma[NVMA]; v++)
    87 if (v->len == 0)
    88 break;
    89 if (v == &p->vma[NVMA])
    90 return -1; // no free slot
    92 if ((start = vmaplace(p, len)) == 0)
    93 return -1;
    94
    95 v->start = start;
    96 v->len = len;
    97 v->prot = prot;
    98 v->flags = flags;
    99 v->f = filedup(f); // the mapping outlives close(fd)
    100 v->off = off;
    101 return start;
    102}

    kernel/syscall.c

    @@ -102,8 +102,9 @@ extern uint64 sys_unlink(void);
    102102extern uint64 sys_link(void);
    103103extern uint64 sys_mkdir(void);
    104104extern uint64 sys_close(void);
    105105extern uint64 sys_sync(void);
    106extern uint64 sys_mmap(void);
    106107
    107108// An array mapping syscall numbers from syscall.h
    108109// to the function that handles the system call.
    109110static uint64 (*syscalls[])(void) = {
    @@ -129,8 +130,9 @@ static uint64 (*syscalls[])(void) = {
    129130 [SYS_link] = sys_link,
    130131 [SYS_mkdir] = sys_mkdir,
    131132 [SYS_close] = sys_close,
    132133 [SYS_sync] = sys_sync,
    134 [SYS_mmap] = sys_mmap,
    133135 // clang-format on
    134136};
    135137
    136138void

    kernel/syscall.h

    @@ -20,4 +20,5 @@
    2020#define SYS_link 19
    2121#define SYS_mkdir 20
    2222#define SYS_close 21
    2323#define SYS_sync 22
    24#define SYS_mmap 23

    user/user.h

    @@ -1,5 +1,6 @@
    11#define SBRK_ERROR ((char *)-1)
    2#define MAP_FAILED ((void *)-1)
    23
    34struct stat;
    45
    56// system calls
    @@ -24,8 +25,9 @@ int getpid(void);
    2425char *sys_sbrk(int, int);
    2526int pause(int);
    2627int uptime(void);
    2728int sync(void);
    29void *mmap(void *, uint64, int, int, int, uint64);
    2830
    2931// ulib.c
    3032int stat(const char *, struct stat *);
    3133char *strcpy(char *, const char *);

    user/usys.pl

    @@ -42,4 +42,5 @@ entry("getpid");
    4242entry("sbrk");
    4343entry("pause");
    4444entry("uptime");
    4545entry("sync");
    46entry("mmap");
  3. c143c84 Keep the heap below the mappings

    kernel/defs.h

    @@ -69,8 +69,9 @@ void begin_op(void);
    6969void end_op(void);
    7070
    7171// mmap.c
    7272struct vma* vmalookup(struct proc*, uint64);
    73uint64 vmabottom(struct proc*);
    7374
    7475// pipe.c
    7576int pipealloc(struct file**, struct file**);
    7677void pipeclose(struct pipe*, int);

    kernel/mmap.c

    @@ -27,8 +27,22 @@ vmalookup(struct proc *p, uint64 va)
    2727 return v;
    2828 return 0;
    2929}
    3030
    31// The lowest address of any mapping of p, or TRAPFRAME if it has
    32// none. The heap must stay below it.
    33uint64
    34vmabottom(struct proc *p)
    35{
    36 uint64 bottom = TRAPFRAME;
    37 struct vma *v;
    38
    39 for (v = p->vma; v < &p->vma[NVMA]; v++)
    40 if (v->len != 0 && v->start < bottom)
    41 bottom = v->start;
    42 return bottom;
    43}
    44
    3145// Choose where a new mapping of len bytes goes: the highest free
    3246// range below TRAPFRAME, and never below the heap. Returns its
    3347// start, or 0 if there is no room.
    3448static uint64

    kernel/proc.c

    @@ -239,10 +239,10 @@ growproc(int n)
    239239 struct proc *p = myproc();
    240240
    241241 sz = p->sz;
    242242 if (n > 0) {
    243 if (sz + n > TRAPFRAME) {
    244 return -1;
    243 if (sz + n > vmabottom(p)) {
    244 return -1; // would run into a mapping, or the trapframe
    245245 }
    246246 if ((sz = uvmalloc(p->pagetable, sz, sz + n, PTE_W)) == 0) {
    247247 return -1;
    248248 }

    kernel/sysproc.c

    @@ -56,10 +56,10 @@ sys_sbrk(void)
    5656 // size but don't allocate memory. If the processes uses the
    5757 // memory, vmfault() will allocate it.
    5858 if (addr + n < addr)
    5959 return -1;
    60 if (addr + n > TRAPFRAME)
    61 return -1;
    60 if (addr + n > vmabottom(myproc()))
    61 return -1; // would run into a mapping, or the trapframe
    6262 myproc()->sz += n;
    6363 }
    6464 return addr;
    6565}
  4. a31499a Load a mapped page from its file on a fault

    kernel/defs.h

    @@ -70,8 +70,9 @@ void end_op(void);
    7070
    7171// mmap.c
    7272struct vma* vmalookup(struct proc*, uint64);
    7373uint64 vmabottom(struct proc*);
    74uint64 mmapfault(struct proc*, struct vma*, uint64, int);
    7475
    7576// pipe.c
    7677int pipealloc(struct file**, struct file**);
    7778void pipeclose(struct pipe*, int);

    kernel/mmap.c

    @@ -63,8 +63,55 @@ vmaplace(struct proc *p, uint64 len)
    6363 end = v->start; // try again just below the mapping in the way
    6464 }
    6565}
    6666
    67// Load the page at va of mapping v from the file and map it.
    68// Called by vmfault. Locks the file's inode and may wait for the
    69// disk, so the caller must hold no spinlock, and must not hold
    70// that inode's lock. Returns the page's physical address, or 0 if
    71// the access is not allowed, the page is already mapped, or
    72// memory is short.
    73uint64
    74mmapfault(struct proc *p, struct vma *v, uint64 va, int read)
    75{
    76 struct inode *ip = v->f->ip;
    77 uint64 off;
    78 uint n;
    79 char *mem;
    80 int perm = PTE_R | PTE_U;
    81
    82 if (!read && (v->prot & PROT_WRITE) == 0)
    83 return 0; // a store to a read-only mapping
    84 va = PGROUNDDOWN(va);
    85 if (ismapped(p->pagetable, va))
    86 return 0; // mapped, so the fault was about permissions
    87
    88 if ((mem = kalloc()) == 0)
    89 return 0;
    90 memset(mem, 0, PGSIZE); // past the end of the file: zeros
    91 off = v->off + (va - v->start);
    92 ilock(ip);
    93 if (off < ip->size) {
    94 n = PGSIZE;
    95 if (ip->size - off < PGSIZE)
    96 n = ip->size - off;
    97 if (readi(ip, 0, (uint64)mem, off, n) != n) {
    98 iunlock(ip);
    99 kfree(mem);
    100 return 0;
    101 }
    102 }
    103 iunlock(ip);
    104
    105 if (v->prot & PROT_WRITE)
    106 perm |= PTE_W;
    107 if (mappages(p->pagetable, va, PGSIZE, (uint64)mem, perm) != 0) {
    108 kfree(mem);
    109 return 0;
    110 }
    111 return (uint64)mem;
    112}
    113
    67114// void *mmap(void *addr, uint64 len, int prot, int flags, int fd,
    68115// uint64 off)
    69116// addr must be 0: the kernel chooses the address.
    70117uint64

    kernel/vm.c

    @@ -451,15 +451,22 @@ copyinstr(pagetable_t pagetable, uint64 psz, char *dst, uint64 srcva,
    451451 }
    452452}
    453453
    454454// allocate and map user memory if process is referencing a page
    455// that was lazily allocated in sys_sbrk().
    455// that was lazily allocated in sys_sbrk(), or a page of a file
    456// mapping (mmap) that has not been read yet.
    456457// returns 0 if va is invalid or already mapped, or if
    457458// out of physical memory, and physical address if successful.
    458459uint64
    459460vmfault(pagetable_t pagetable, uint64 psz, uint64 va, int read)
    460461{
    461462 uint64 mem;
    463 struct proc *p = myproc();
    464 struct vma *v;
    465
    466 // a page of a mapped file: read it from the file.
    467 if (pagetable == p->pagetable && (v = vmalookup(p, va)) != 0)
    468 return mmapfault(p, v, va, read);
    462469
    463470 if (va >= psz)
    464471 return 0;
    465472 va = PGROUNDDOWN(va);
  5. 500fde3 Add munmap

    kernel/defs.h

    @@ -71,8 +71,9 @@ void end_op(void);
    7171// mmap.c
    7272struct vma* vmalookup(struct proc*, uint64);
    7373uint64 vmabottom(struct proc*);
    7474uint64 mmapfault(struct proc*, struct vma*, uint64, int);
    75void vmaunmap(pagetable_t, struct vma*, uint64, uint64);
    7576
    7677// pipe.c
    7778int pipealloc(struct file**, struct file**);
    7879void pipeclose(struct pipe*, int);

    kernel/mmap.c

    @@ -160,4 +160,48 @@ sys_mmap(void)
    160160 v->f = filedup(f); // the mapping outlives close(fd)
    161161 v->off = off;
    162162 return start;
    163163}
    164
    165// Remove [addr, addr+len) from mapping v, in the address space of
    166// pagetable: free its loaded pages and shrink the record. addr and
    167// len are page-aligned, and the range is the whole of v, or a part
    168// of it that starts at its start or ends at its end. The last page
    169// gone drops the mapping's reference to the file.
    170void
    171vmaunmap(pagetable_t pagetable, struct vma *v, uint64 addr, uint64 len)
    172{
    174 if (addr == v->start) {
    175 v->start += len;
    176 v->off += len;
    177 }
    178 v->len -= len;
    179 if (v->len == 0) {
    180 fileclose(v->f);
    181 v->f = 0;
    182 }
    183}
    184
    185// int munmap(void *addr, uint64 len)
    186// Unmaps a whole mapping, or a part at its start or at its end.
    187uint64
    188sys_munmap(void)
    189{
    190 uint64 addr, len;
    191 struct proc *p = myproc();
    192 struct vma *v;
    193
    194 argaddr(0, &addr);
    195 argaddr(1, &len);
    196 if (addr % PGSIZE != 0 || len == 0 || len > TRAPFRAME)
    197 return -1;
    198 len = PGROUNDUP(len);
    199 if ((v = vmalookup(p, addr)) == 0)
    200 return -1;
    201 if (addr + len > v->start + v->len)
    202 return -1; // runs past the end of the mapping
    203 if (addr != v->start && addr + len != v->start + v->len)
    204 return -1; // would leave a hole in the middle
    205 vmaunmap(p->pagetable, v, addr, len);
    206 return 0;
    207}

    kernel/syscall.c

    @@ -103,8 +103,9 @@ extern uint64 sys_link(void);
    103103extern uint64 sys_mkdir(void);
    104104extern uint64 sys_close(void);
    105105extern uint64 sys_sync(void);
    106106extern uint64 sys_mmap(void);
    107extern uint64 sys_munmap(void);
    107108
    108109// An array mapping syscall numbers from syscall.h
    109110// to the function that handles the system call.
    110111static uint64 (*syscalls[])(void) = {
    @@ -131,8 +132,9 @@ static uint64 (*syscalls[])(void) = {
    131132 [SYS_mkdir] = sys_mkdir,
    132133 [SYS_close] = sys_close,
    133134 [SYS_sync] = sys_sync,
    134135 [SYS_mmap] = sys_mmap,
    136 [SYS_munmap] = sys_munmap,
    135137 // clang-format on
    136138};
    137139
    138140void

    kernel/syscall.h

    @@ -21,4 +21,5 @@
    2121#define SYS_mkdir 20
    2222#define SYS_close 21
    2323#define SYS_sync 22
    2424#define SYS_mmap 23
    25#define SYS_munmap 24

    user/user.h

    @@ -26,8 +26,9 @@ char *sys_sbrk(int, int);
    2626int pause(int);
    2727int uptime(void);
    2828int sync(void);
    2929void *mmap(void *, uint64, int, int, int, uint64);
    30int munmap(void *, uint64);
    3031
    3132// ulib.c
    3233int stat(const char *, struct stat *);
    3334char *strcpy(char *, const char *);

    user/usys.pl

    @@ -43,4 +43,5 @@ entry("sbrk");
    4343entry("pause");
    4444entry("uptime");
    4545entry("sync");
    4646entry("mmap");
    47entry("munmap");
  6. 5992781 Unmap every mapping at exit and exec

    kernel/defs.h

    @@ -72,8 +72,9 @@ void end_op(void);
    7272struct vma* vmalookup(struct proc*, uint64);
    7373uint64 vmabottom(struct proc*);
    7474uint64 mmapfault(struct proc*, struct vma*, uint64, int);
    7575void vmaunmap(pagetable_t, struct vma*, uint64, uint64);
    76void vmaunmapall(struct proc*, pagetable_t);
    7677
    7778// pipe.c
    7879int pipealloc(struct file**, struct file**);
    7980void pipeclose(struct pipe*, int);

    kernel/exec.c

    @@ -134,8 +134,9 @@ kexec(char *path, char **argv)
    134134 p->pagetable = pagetable;
    135135 p->sz = sz;
    136136 p->trapframe->epc = elf.entry; // initial program counter = ulib.c:start()
    137137 p->trapframe->sp = sp; // initial stack pointer
    138 vmaunmapall(p, oldpagetable); // the new image starts with no mappings
    138139 proc_freepagetable(oldpagetable, oldsz);
    139140
    140141 return argc; // this ends up in a0, the first argument to main(argc, argv)
    141142

    kernel/mmap.c

    @@ -181,8 +181,21 @@ vmaunmap(pagetable_t pagetable, struct vma *v, uint64 addr, uint64 len)
    181181 v->f = 0;
    182182 }
    183183}
    184184
    185// Remove all of p's mappings from pagetable: at exit, and in exec
    186// for the old image. Their pages lie above p->sz, where uvmfree
    187// does not look, and freewalk panics on any page still mapped.
    188void
    189vmaunmapall(struct proc *p, pagetable_t pagetable)
    190{
    191 struct vma *v;
    192
    193 for (v = p->vma; v < &p->vma[NVMA]; v++)
    194 if (v->len != 0)
    195 vmaunmap(pagetable, v, v->start, v->len);
    196}
    197
    185198// int munmap(void *addr, uint64 len)
    186199// Unmaps a whole mapping, or a part at its start or at its end.
    187200uint64
    188201sys_munmap(void)

    kernel/proc.c

    @@ -329,8 +329,11 @@ kexit(int status)
    329329
    330330 if (p == initproc)
    331331 panic("init exiting");
    332332
    333 // Unmap all mapped files.
    334 vmaunmapall(p, p->pagetable);
    335
    333336 // Close all open files.
    334337 for (int fd = 0; fd < NOFILE; fd++) {
    335338 if (p->ofile[fd]) {
    336339 struct file *f = p->ofile[fd];
  7. 3b5bcad Write dirty pages of shared mappings back to the file

    kernel/mmap.c

    @@ -161,16 +161,51 @@ sys_mmap(void)
    161161 v->off = off;
    162162 return start;
    163163}
    164164
    165// Write the page at va of shared mapping v, whose contents are at
    166// physical address pa, back to the file. Only the part inside the
    167// file is written: a mapping never makes its file longer.
    168static void
    169vmawrite(struct vma *v, uint64 va, uint64 pa)
    170{
    171 struct inode *ip = v->f->ip;
    172 uint64 off = v->off + (va - v->start);
    173 uint n;
    174
    175 // one page is 4 blocks, all already in the file, plus the
    176 // inode: 5 log blocks, within the MAXOPBLOCKS of one transaction.
    177 begin_op();
    178 ilock(ip);
    179 if (off < ip->size) {
    180 n = PGSIZE;
    181 if (ip->size - off < PGSIZE)
    182 n = ip->size - off;
    183 writei(ip, 0, pa, off, n);
    184 }
    185 iunlock(ip);
    186 end_op();
    187}
    188
    165189// Remove [addr, addr+len) from mapping v, in the address space of
    166// pagetable: free its loaded pages and shrink the record. addr and
    167// len are page-aligned, and the range is the whole of v, or a part
    168// of it that starts at its start or ends at its end. The last page
    169// gone drops the mapping's reference to the file.
    190// pagetable: write back the dirty pages of a shared mapping, free
    191// the loaded pages and shrink the record. addr and len are
    192// page-aligned, and the range is the whole of v, or a part of it
    193// that starts at its start or ends at its end. The last page gone
    194// drops the mapping's reference to the file.
    170195void
    171196vmaunmap(pagetable_t pagetable, struct vma *v, uint64 addr, uint64 len)
    172197{
    198 uint64 a;
    199 pte_t *pte;
    200
    201 if (v->flags == MAP_SHARED && (v->prot & PROT_WRITE)) {
    202 for (a = addr; a < addr + len; a += PGSIZE) {
    203 pte = walk(pagetable, a, 0);
    204 if (pte && (*pte & PTE_V) && (*pte & PTE_D))
    205 vmawrite(v, a, PTE2PA(*pte)); // written since it was loaded
    206 }
    207 }
    173208 uvmunmap(pagetable, addr, len / PGSIZE, 1);
    174209 if (addr == v->start) {
    175210 v->start += len;
    176211 v->off += len;

    kernel/riscv.h

    @@ -396,8 +396,10 @@ typedef uint64 *pagetable_t; // 512 PTEs
    396396#define PTE_R (1L << 1)
    397397#define PTE_W (1L << 2)
    398398#define PTE_X (1L << 3)
    399399#define PTE_U (1L << 4) // user can access
    400#define PTE_A (1L << 6) // accessed: set by the hardware
    401#define PTE_D (1L << 7) // dirty: set by the hardware on a store
    400402
    401403// shift a physical address to the right place for a PTE.
    402404#define PA2PTE(pa) ((((uint64)pa) >> 12) << 10)
    403405

    kernel/vm.c

    @@ -362,8 +362,11 @@ copyout(pagetable_t pagetable, uint64 psz, uint64 dstva, char *src, uint64 len)
    362362 pte = walk(pagetable, va0, 0);
    363363 // forbid copyout over read-only user text pages.
    364364 if ((*pte & PTE_W) == 0)
    365365 return -1;
    366 // the copy goes through the kernel's own mapping of the page,
    367 // so the hardware won't set D in the user PTE: set it here.
    368 *pte |= PTE_D;
    366369
    367370 n = PGSIZE - (dstva - va0);
    368371 if (n > len)
    369372 n = len;
  8. 8473afe Give a forked child its parent's mappings

    kernel/proc.c

    @@ -286,8 +286,16 @@ kfork(void)
    286286 if (p->ofile[i])
    287287 np->ofile[i] = filedup(p->ofile[i]);
    288288 np->cwd = idup(p->cwd);
    289289
    290 // the child has the same mappings, with its own file references.
    291 // uvmcopy copied no mapped page (they lie above sz): the child
    292 // reads each page from the file when it first touches it.
    293 memmove(np->vma, p->vma, sizeof(p->vma));
    294 for (i = 0; i < NVMA; i++)
    295 if (np->vma[i].len != 0)
    296 filedup(np->vma[i].f);
    297
    290298 safestrcpy(np->name, p->name, sizeof(p->name));
    291299
    292300 pid = np->pid;
    293301
  9. c058c08 Load mapped pages before read, write and wait take locks

    kernel/defs.h

    @@ -73,8 +73,9 @@ struct vma* vmalookup(struct proc*, uint64);
    7373uint64 vmabottom(struct proc*);
    7474uint64 mmapfault(struct proc*, struct vma*, uint64, int);
    7575void vmaunmap(pagetable_t, struct vma*, uint64, uint64);
    7676void vmaunmapall(struct proc*, pagetable_t);
    77int vmaprefault(struct proc*, uint64, uint64);
    7778
    7879// pipe.c
    7980int pipealloc(struct file**, struct file**);
    8081void pipeclose(struct pipe*, int);

    kernel/mmap.c

    @@ -110,8 +110,33 @@ mmapfault(struct proc *p, struct vma *v, uint64 va, int read)
    110110 }
    111111 return (uint64)mem;
    112112}
    113113
    114// Load the pages of p's mappings in [va, va+n) that are not loaded
    115// yet. A system call that copies to or from user memory while it
    116// holds locks calls this first, holding none: loading a mapped page
    117// sleeps, and locks the mapped file's inode. Returns -1 if a page
    118// can't be loaded.
    119int
    120vmaprefault(struct proc *p, uint64 va, uint64 n)
    121{
    122 uint64 a, lo, hi;
    123 struct vma *v;
    124
    125 if (va + n < va)
    126 return -1;
    127 for (v = p->vma; v < &p->vma[NVMA]; v++) {
    128 if (v->len == 0)
    129 continue;
    130 lo = PGROUNDDOWN(va) > v->start ? PGROUNDDOWN(va) : v->start;
    131 hi = va + n < v->start + v->len ? va + n : v->start + v->len;
    132 for (a = lo; a < hi; a += PGSIZE)
    133 if (!ismapped(p->pagetable, a) && mmapfault(p, v, a, 1) == 0)
    134 return -1;
    135 }
    136 return 0;
    137}
    138
    114139// void *mmap(void *addr, uint64 len, int prot, int flags, int fd,
    115140// uint64 off)
    116141// addr must be 0: the kernel chooses the address.
    117142uint64

    kernel/sysfile.c

    @@ -75,8 +75,12 @@ sys_read(void)
    7575 argaddr(1, &p);
    7676 argint(2, &n);
    7777 if (argfd(0, 0, &f) < 0)
    7878 return -1;
    79 // load the buffer's mapped pages now: fileread holds locks while
    80 // it copies, and loading a mapped page needs to sleep.
    81 if (n > 0 && vmaprefault(myproc(), p, n) < 0)
    82 return -1;
    7983 return fileread(f, p, n);
    8084}
    8185
    8286uint64
    @@ -89,8 +93,11 @@ sys_write(void)
    8993 argaddr(1, &p);
    9094 argint(2, &n);
    9195 if (argfd(0, 0, &f) < 0)
    9296 return -1;
    97 // as in sys_read: filewrite copies in while holding locks.
    98 if (n > 0 && vmaprefault(myproc(), p, n) < 0)
    99 return -1;
    93100
    94101 return filewrite(f, p, n);
    95102}
    96103

    kernel/sysproc.c

    @@ -32,8 +32,11 @@ uint64
    3232sys_wait(void)
    3333{
    3434 uint64 p;
    3535 argaddr(0, &p);
    36 // kwait copies the status out holding spinlocks.
    37 if (p != 0 && vmaprefault(myproc(), p, sizeof(int)) < 0)
    38 return -1;
    3639 return kwait(p);
    3740}
    3841
    3942uint64
  10. a980c9a Add mmaptest, a test program for mmap and munmap

    Makefile

    @@ -150,8 +150,9 @@ UPROGS=\
    150150 $U/_logstress\
    151151 $U/_forphan\
    152152 $U/_dorphan\
    153153 $U/_sync\
    154 $U/_mmaptest\
    154155
    155156fs.img: mkfs/mkfs README $(UPROGS)
    156157 mkfs/mkfs fs.img README $(UPROGS)
    157158

    user/mmaptest.c

    @@ -0,0 +1,660 @@
    1// Tests for mmap and munmap. Prints one line per check;
    2// "mmaptest NAME" runs only that check.
    3
    4#include "kernel/param.h"
    5#include "kernel/types.h"
    6#include "kernel/stat.h"
    7#include "kernel/riscv.h"
    8#include "kernel/fcntl.h"
    9#include "user/user.h"
    10
    11int failed;
    12char buf[PGSIZE]; // the user stack is one page: keep big buffers off it
    13
    14void
    15fail(char *test, char *what)
    16{
    17 printf("mmaptest: %s: FAIL (%s)\n", test, what);
    18 failed = 1;
    19}
    20
    21// The byte at offset i of every test file.
    22char
    23fbyte(uint i)
    24{
    25 return (i % 251) ^ (i / PGSIZE);
    26}
    27
    28// Create name with size bytes of fbyte().
    29void
    30makefile(char *name, uint size)
    31{
    32 char buf[512];
    33 uint i, j, n;
    34 int fd;
    35
    36 if ((fd = open(name, O_CREATE | O_TRUNC | O_RDWR)) < 0) {
    37 printf("mmaptest: cannot create %s\n", name);
    38 exit(1);
    39 }
    40 for (i = 0; i < size; i += n) {
    41 n = size - i < sizeof(buf) ? size - i : sizeof(buf);
    42 for (j = 0; j < n; j++)
    43 buf[j] = fbyte(i + j);
    44 if (write(fd, buf, n) != n) {
    45 printf("mmaptest: cannot write %s\n", name);
    46 exit(1);
    47 }
    48 }
    49 close(fd);
    50}
    51
    52// Read n bytes at offset off of name into buf.
    53int
    54readat(char *name, uint off, char *buf, int n)
    55{
    56 static char skip[512];
    57 int fd, k;
    58
    59 if ((fd = open(name, O_RDONLY)) < 0)
    60 return -1;
    61 while (off > 0) {
    62 k = off < sizeof(skip) ? off : sizeof(skip);
    63 if (read(fd, skip, k) != k) {
    64 close(fd);
    65 return -1;
    66 }
    67 off -= k;
    68 }
    69 k = read(fd, buf, n);
    70 close(fd);
    71 return k;
    72}
    73
    74int
    75filesize(char *name)
    76{
    77 struct stat st;
    78
    79 if (stat(name, &st) < 0)
    80 return -1;
    81 return st.size;
    82}
    83
    84// Free pages, counted the way usertests counts them.
    85int
    87{
    88 int n = 0;
    89 uint64 sz0 = (uint64)sbrk(0);
    90
    91 while (sbrk(PGSIZE) != SBRK_ERROR)
    92 n++;
    93 sbrk(-((uint64)sbrk(0) - sz0));
    94 return n;
    95}
    96
    97// Run f in a child; return the child's exit status.
    98int
    99inchild(void (*f)(char *), char *arg)
    100{
    101 int pid, xs;
    102
    103 if ((pid = fork()) < 0) {
    104 printf("mmaptest: fork failed\n");
    105 exit(1);
    106 }
    107 if (pid == 0) {
    108 f(arg);
    109 exit(0);
    110 }
    111 if (wait(&xs) != pid)
    112 return -100;
    113 return xs;
    114}
    115
    116void
    117store(char *a)
    118{
    119 *(volatile char *)a = 1;
    120}
    121
    122void
    123load(char *a)
    124{
    125 printf("mmaptest: loaded %d: should have been killed\n",
    126 *(volatile char *)a);
    127}
    128
    129// Map a file read-only, close it, read it through memory.
    130void
    131readtest(void)
    132{
    133 char *t = "read";
    134 uint size = 2 * PGSIZE + 1000, i;
    135 int fd, f0, f1, bad = 0;
    136 char *m;
    137
    138 makefile("mm.read", size);
    139 if ((fd = open("mm.read", O_RDONLY)) < 0) {
    140 fail(t, "open");
    141 return;
    142 }
    143 m = mmap(0, 3 * PGSIZE, PROT_READ, MAP_PRIVATE, fd, 0);
    144 close(fd); // the mapping keeps its own reference
    145 if (m == MAP_FAILED) {
    146 fail(t, "mmap");
    147 return;
    148 }
    149 if ((uint64)m % PGSIZE != 0 || (uint64)m < (uint64)sbrk(0))
    150 fail(t, "address");
    151
    152 countfree(); // warm-up: the heap's page-table pages
    153 f0 = countfree();
    154 if (m[PGSIZE + 5] != fbyte(PGSIZE + 5))
    155 fail(t, "page 1");
    156 f1 = countfree();
    157 printf("mmaptest: read: touching 1 of 3 mapped pages took %d free pages\n",
    158 f0 - f1);
    159 if (f0 - f1 != 1)
    160 fail(t, "lazy");
    161
    162 for (i = 0; i < 3 * PGSIZE; i++)
    163 if (m[i] != (i < size ? fbyte(i) : 0))
    164 bad++;
    165 if (bad)
    166 fail(t, "contents");
    167 if (munmap(m, 3 * PGSIZE) != 0)
    168 fail(t, "munmap");
    169 if (inchild(load, m) != -1)
    170 fail(t, "still mapped after munmap");
    171 unlink("mm.read");
    172}
    173
    174// Stores through a shared mapping reach the file at munmap; the
    175// file never grows.
    176void
    177sharedtest(void)
    178{
    179 char *t = "shared";
    180 uint size = 2 * PGSIZE + 100, i;
    181 int fd, bad = 0;
    182 char *m;
    183
    184 makefile("mm.shared", size);
    185 if ((fd = open("mm.shared", O_RDWR)) < 0) {
    186 fail(t, "open");
    187 return;
    188 }
    189 m = mmap(0, 3 * PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    190 close(fd);
    191 if (m == MAP_FAILED) {
    192 fail(t, "mmap");
    193 return;
    194 }
    195 strcpy(m + 10, "hello"); // page 0
    196 m[2 * PGSIZE + 50] = 'X'; // page 2, inside the file
    197 m[2 * PGSIZE + 3000] = 'Y'; // page 2, past the end of the file
    198 if (munmap(m, 3 * PGSIZE) != 0) {
    199 fail(t, "munmap");
    200 return;
    201 }
    202
    203 if (filesize("mm.shared") != size)
    204 fail(t, "the file changed size");
    205 if (readat("mm.shared", 0, buf, PGSIZE) != PGSIZE ||
    206 strcmp(buf + 10, "hello") != 0)
    207 fail(t, "page 0 not written back");
    208 for (i = 0; i < PGSIZE; i++)
    209 if (i < 10 || i > 15)
    210 bad += buf[i] != fbyte(i);
    211 if (readat("mm.shared", PGSIZE, buf, PGSIZE) != PGSIZE)
    212 fail(t, "read page 1");
    213 for (i = 0; i < PGSIZE; i++)
    214 bad += buf[i] != fbyte(PGSIZE + i);
    215 if (readat("mm.shared", 2 * PGSIZE, buf, PGSIZE) != 100 || buf[50] != 'X')
    216 fail(t, "page 2 not written back");
    217 for (i = 0; i < 100; i++)
    218 if (i != 50)
    219 bad += buf[i] != fbyte(2 * PGSIZE + i);
    220 if (bad)
    221 fail(t, "other bytes changed");
    222 unlink("mm.shared");
    223}
    224
    225// Private stores never reach the file; prot is enforced.
    226void
    227privatetest(void)
    228{
    229 char *t = "private";
    230 int fd, pfd[2], i, bad = 0;
    231 char *m, *r;
    232
    233 makefile("mm.priv", PGSIZE);
    234 if ((fd = open("mm.priv", O_RDONLY)) < 0) {
    235 fail(t, "open");
    236 return;
    237 }
    238 if (mmap(0, PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0) !=
    239 MAP_FAILED)
    240 fail(t, "shared writable map of a read-only file");
    241 m = mmap(0, PGSIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);
    242 r = mmap(0, PGSIZE, PROT_READ, MAP_SHARED, fd, 0);
    243 close(fd);
    244 if (m == MAP_FAILED || r == MAP_FAILED) {
    245 fail(t, "mmap");
    246 return;
    247 }
    248 for (i = 0; i < PGSIZE; i++)
    249 m[i] = 'P';
    250 if (m[100] != 'P' || r[100] != fbyte(100))
    251 fail(t, "contents");
    252 if (inchild(store, r) != -1)
    253 fail(t, "store to a read-only mapping");
    254 if ((fd = open("mm.priv", O_RDONLY)) < 0 || read(fd, r, 10) != -1)
    255 fail(t, "read() into a read-only mapping");
    256 close(fd);
    257 if (munmap(m, PGSIZE) != 0 || munmap(r, PGSIZE) != 0)
    258 fail(t, "munmap");
    259 if (readat("mm.priv", 0, buf, PGSIZE) != PGSIZE)
    260 fail(t, "read");
    261 for (i = 0; i < PGSIZE; i++)
    262 bad += buf[i] != fbyte(i);
    263 if (bad)
    264 fail(t, "a private store reached the file");
    265
    266 pipe(pfd);
    267 if (mmap(0, PGSIZE, PROT_READ, MAP_PRIVATE, pfd[0], 0) != MAP_FAILED)
    268 fail(t, "mapped a pipe");
    269 close(pfd[0]);
    270 close(pfd[1]);
    271 unlink("mm.priv");
    272}
    273
    274// munmap of a part at the start or at the end of a mapping.
    275void
    276partialtest(void)
    277{
    278 char *t = "partial";
    279 int fd, i;
    280 char *m;
    281
    282 makefile("mm.part", 4 * PGSIZE);
    283 if ((fd = open("mm.part", O_RDWR)) < 0) {
    284 fail(t, "open");
    285 return;
    286 }
    287 m = mmap(0, 4 * PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    288 close(fd);
    289 if (m == MAP_FAILED) {
    290 fail(t, "mmap");
    291 return;
    292 }
    293 for (i = 0; i < 4; i++)
    294 m[i * PGSIZE] = 'A' + i;
    295 if (munmap(m + PGSIZE, PGSIZE) != -1)
    296 fail(t, "unmapped a hole in the middle");
    297 if (munmap(m, PGSIZE) != 0) // the first page
    298 fail(t, "munmap start");
    299 if (munmap(m + 3 * PGSIZE, PGSIZE) != 0) // the last page
    300 fail(t, "munmap end");
    301 if (inchild(load, m) != -1 || inchild(load, m + 3 * PGSIZE) != -1)
    302 fail(t, "still mapped after munmap");
    303 if (m[PGSIZE] != 'B' || m[2 * PGSIZE] != 'C')
    304 fail(t, "the rest lost its contents");
    305 m[2 * PGSIZE + 1] = 'c';
    306 if (munmap(m + PGSIZE, 2 * PGSIZE) != 0)
    307 fail(t, "munmap rest");
    308 for (i = 0; i < 4; i++)
    309 if (readat("mm.part", i * PGSIZE, buf, 2) != 2 || buf[0] != 'A' + i ||
    310 buf[1] != (i == 2 ? 'c' : fbyte(i * PGSIZE + 1)))
    311 fail(t, "not written back");
    312 unlink("mm.part");
    313}
    314
    315char *forkmap;
    316
    317void
    318forkchild(char *arg)
    319{
    320 int i;
    321
    322 // page 1: the parent never touched it.
    323 for (i = 0; i < PGSIZE; i++)
    324 if (forkmap[PGSIZE + i] != fbyte(PGSIZE + i))
    325 exit(1);
    326 forkmap[PGSIZE] = 'C';
    327 exit(0); // no munmap: exit writes page 1 back
    328}
    329
    330// A child has its parent's mappings.
    331void
    332forktest(void)
    333{
    334 char *t = "fork";
    335
    336 makefile("mm.fork", 2 * PGSIZE);
    337 int fd = open("mm.fork", O_RDWR);
    338 forkmap = mmap(0, 2 * PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    339 close(fd);
    340 if (forkmap == MAP_FAILED) {
    341 fail(t, "mmap");
    342 return;
    343 }
    344 forkmap[0] = 'P';
    345 if (inchild(forkchild, 0) != 0)
    346 fail(t, "child");
    347 if (forkmap[PGSIZE] != 'C') // loaded now, after the child's exit
    348 fail(t, "the child's store did not reach the file");
    349 if (munmap(forkmap, 2 * PGSIZE) != 0)
    350 fail(t, "munmap");
    351 if (readat("mm.fork", 0, buf, 1) != 1 || buf[0] != 'P' ||
    352 readat("mm.fork", PGSIZE, buf, 1) != 1 || buf[0] != 'C')
    353 fail(t, "file");
    354 unlink("mm.fork");
    355}
    356
    357void
    358exitchild(char *how)
    359{
    360 int fd = open("mm.exit", O_RDWR);
    361 char *m = mmap(0, 2 * PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    362 char *argv[] = {"mmaptest", "nop", 0};
    363
    364 close(fd);
    365 if (m == MAP_FAILED)
    366 exit(1);
    367 if (how[0] == 'x') {
    368 m[0] = 'x';
    369 exit(0);
    370 }
    371 m[PGSIZE] = 'e';
    372 exec("mmaptest", argv);
    373 exit(1);
    374}
    375
    376// exit and exec write dirty shared pages back.
    377void
    378exittest(void)
    379{
    380 char *t = "exit";
    381
    382 makefile("mm.exit", 2 * PGSIZE);
    383 if (inchild(exitchild, "x") != 0 || inchild(exitchild, "e") != 0)
    384 fail(t, "child");
    385 if (readat("mm.exit", 0, buf, 1) != 1 || buf[0] != 'x')
    386 fail(t, "exit did not write back");
    387 if (readat("mm.exit", PGSIZE, buf, 1) != 1 || buf[0] != 'e')
    388 fail(t, "exec did not write back");
    389 unlink("mm.exit");
    390}
    391
    392// The table fills up; nothing leaks.
    393void
    394manytest(void)
    395{
    396 char *t = "many";
    397 char *m[NVMA + 1];
    398 int fd, i, f0, f1;
    399
    400 makefile("mm.many", PGSIZE);
    401 countfree();
    402 f0 = countfree();
    403 if ((fd = open("mm.many", O_RDONLY)) < 0) {
    404 fail(t, "open");
    405 return;
    406 }
    407 for (i = 0; i < NVMA; i++) {
    408 m[i] = mmap(0, PGSIZE, PROT_READ, MAP_PRIVATE, fd, 0);
    409 if (m[i] == MAP_FAILED || m[i][7] != fbyte(7)) {
    410 fail(t, "mmap");
    411 return;
    412 }
    413 }
    414 if (mmap(0, PGSIZE, PROT_READ, MAP_PRIVATE, fd, 0) != MAP_FAILED)
    415 fail(t, "more than NVMA mappings");
    416 close(fd);
    417 for (i = 0; i < NVMA; i++)
    418 if (munmap(m[i], PGSIZE) != 0)
    419 fail(t, "munmap");
    420
    421 // each round takes a file and a page; if munmap failed to give
    422 // them back, open() would run out of files after NFILE rounds.
    423 for (i = 0; i < 3 * NFILE; i++) {
    424 if ((fd = open("mm.many", O_RDONLY)) < 0) {
    425 fail(t, "open failed: files leak");
    426 return;
    427 }
    428 m[0] = mmap(0, PGSIZE, PROT_READ, MAP_PRIVATE, fd, 0);
    429 close(fd);
    430 if (m[0] == MAP_FAILED || m[0][9] != fbyte(9) ||
    431 munmap(m[0], PGSIZE) != 0) {
    432 fail(t, "round");
    433 return;
    434 }
    435 }
    436 f1 = countfree();
    437 if (f1 != f0) {
    438 printf("mmaptest: many: %d free pages before, %d after\n", f0, f1);
    439 fail(t, "pages leak");
    440 }
    441
    442 // a mapping keeps an unlinked file alive.
    443 fd = open("mm.many", O_RDONLY);
    444 m[0] = mmap(0, PGSIZE, PROT_READ, MAP_PRIVATE, fd, 0);
    445 close(fd);
    446 unlink("mm.many");
    447 if (m[0] == MAP_FAILED || m[0][11] != fbyte(11))
    448 fail(t, "unlinked file");
    449 munmap(m[0], PGSIZE);
    450}
    451
    452// System calls that copy into or out of unloaded mapped pages
    453// while they hold spinlocks.
    454void
    455lockedtest(void)
    456{
    457 char *t = "locked";
    458 int fd, pfd[2], i, bad = 0;
    459 char *m;
    460
    461 makefile("mm.lock", 3 * PGSIZE);
    462 fd = open("mm.lock", O_RDWR);
    463 m = mmap(0, 3 * PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    464 close(fd);
    465 if (m == MAP_FAILED) {
    466 fail(t, "mmap");
    467 return;
    468 }
    469 // pipewrite copies page 0 in, piperead copies into page 1, both
    470 // holding the pipe's spinlock.
    471 pipe(pfd);
    472 if (write(pfd[1], m, 8) != 8 || read(pfd[0], m + PGSIZE, 8) != 8)
    473 fail(t, "pipe");
    474 close(pfd[0]);
    475 close(pfd[1]);
    476 // kwait copies the status into page 2 holding two spinlocks.
    477 if (fork() == 0)
    478 exit(7);
    479 wait((int *)(m + 2 * PGSIZE));
    480 if (*(int *)(m + 2 * PGSIZE) != 7)
    481 fail(t, "wait");
    482 if (munmap(m, 3 * PGSIZE) != 0) {
    483 fail(t, "munmap");
    484 return;
    485 }
    486 // what the kernel wrote into pages 1 and 2 went back to the file:
    487 // copyout marked them dirty.
    488 readat("mm.lock", PGSIZE, buf, 8);
    489 for (i = 0; i < 8; i++)
    490 bad += buf[i] != fbyte(i);
    491 readat("mm.lock", 2 * PGSIZE, buf, 4);
    492 bad += *(int *)buf != 7;
    493 if (bad)
    494 fail(t, "file contents");
    495 unlink("mm.lock");
    496}
    497
    498// read() and write() of a file, with the buffer in an unloaded page
    499// of a mapping of that same file.
    500void
    501selftest(void)
    502{
    503 char *t = "self";
    504 int fd, fd2, i, bad = 0;
    505 char *m;
    506
    507 makefile("mm.self", 3 * PGSIZE);
    508 fd = open("mm.self", O_RDWR);
    509 m = mmap(0, 3 * PGSIZE, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    510 if (m == MAP_FAILED) {
    511 fail(t, "mmap");
    512 return;
    513 }
    514 // fileread holds the file's inode lock while it copies into page 1.
    515 fd2 = open("mm.self", O_RDONLY);
    516 if (read(fd2, m + PGSIZE, PGSIZE) != PGSIZE)
    517 fail(t, "read into the file's own mapping");
    518 close(fd2);
    519 // filewrite holds the inode lock and a transaction while it
    520 // copies from page 2; the bytes go to the start of the file.
    521 if (write(fd, m + 2 * PGSIZE, 16) != 16)
    522 fail(t, "write from the file's own mapping");
    523 close(fd);
    524 if (munmap(m, 3 * PGSIZE) != 0) {
    525 fail(t, "munmap");
    526 return;
    527 }
    528 // page 1 (written by read) went back; page 2 was only read, and
    529 // page 0 was never loaded: neither was written back.
    530 readat("mm.self", 0, buf, PGSIZE);
    531 for (i = 0; i < PGSIZE; i++)
    532 bad += buf[i] != (i < 16 ? fbyte(2 * PGSIZE + i) : fbyte(i));
    533 readat("mm.self", PGSIZE, buf, PGSIZE);
    534 for (i = 0; i < PGSIZE; i++)
    535 bad += buf[i] != fbyte(i);
    536 readat("mm.self", 2 * PGSIZE, buf, PGSIZE);
    537 for (i = 0; i < PGSIZE; i++)
    538 bad += buf[i] != fbyte(2 * PGSIZE + i);
    539 if (bad)
    540 fail(t, "file contents");
    541 unlink("mm.self");
    542}
    543
    544#define NCHILD 3
    545#define PERCHILD 4
    546#define ROUNDS 30
    547
    548// Each child, ROUNDS times: map the whole file shared, check every
    549// page, fill its own pages, unmap.
    550void
    551stresschild(char *arg)
    552{
    553 int c = arg[0] - '0', r, i, k, fd;
    554 int *m;
    555
    556 for (r = 0; r < ROUNDS; r++) {
    557 fd = open("mm.stress", O_RDWR);
    558 m = mmap(0, NCHILD * PERCHILD * PGSIZE, PROT_READ | PROT_WRITE,
    559 MAP_SHARED, fd, 0);
    560 close(fd);
    561 if (m == MAP_FAILED)
    562 exit(1);
    563 for (k = 0; k < NCHILD * PERCHILD; k++) {
    564 int *pg = m + k * PGSIZE / sizeof(int);
    565 for (i = 1; i < PGSIZE / sizeof(int); i++)
    566 if (pg[i] != pg[0])
    567 exit(2); // a page must never be half written
    568 }
    569 for (k = c * PERCHILD; k < (c + 1) * PERCHILD; k++)
    570 for (i = 0; i < PGSIZE / sizeof(int); i++)
    571 m[k * PGSIZE / sizeof(int) + i] = (c + 1) * 1000 + r;
    572 if (munmap(m, NCHILD * PERCHILD * PGSIZE) != 0)
    573 exit(3);
    574 }
    575 exit(0);
    576}
    577
    578void
    579stresstest(void)
    580{
    581 char *t = "stress";
    582 int fd, c, k, xs, f0, f1, bad = 0;
    583 int *ibuf = (int *)buf;
    584 char arg[2];
    585
    586 fd = open("mm.stress", O_CREATE | O_TRUNC | O_RDWR);
    587 memset(buf, 0, sizeof(buf));
    588 for (k = 0; k < NCHILD * PERCHILD; k++)
    589 write(fd, buf, PGSIZE);
    590 close(fd);
    591 countfree();
    592 f0 = countfree();
    593 for (c = 0; c < NCHILD; c++) {
    594 if (fork() == 0) {
    595 arg[0] = '0' + c;
    596 arg[1] = 0;
    597 stresschild(arg);
    598 }
    599 }
    600 for (c = 0; c < NCHILD; c++) {
    601 wait(&xs);
    602 if (xs != 0) {
    603 printf("mmaptest: stress: a child exited with %d\n", xs);
    604 bad++;
    605 }
    606 }
    607 fd = open("mm.stress", O_RDONLY);
    608 for (k = 0; k < NCHILD * PERCHILD; k++) {
    609 read(fd, buf, PGSIZE);
    610 for (c = 0; c < PGSIZE / sizeof(int); c++)
    611 bad += ibuf[c] != (k / PERCHILD + 1) * 1000 + ROUNDS - 1;
    612 }
    613 close(fd);
    614 if (bad)
    615 fail(t, "file contents");
    616 f1 = countfree();
    617 if (f1 != f0) {
    618 printf("mmaptest: stress: %d free pages before, %d after\n", f0, f1);
    619 fail(t, "pages leak");
    620 }
    621 unlink("mm.stress");
    622}
    623
    624struct test {
    625 void (*f)(void);
    626 char *name;
    627} tests[] = {
    628 {readtest, "read"}, {sharedtest, "shared"}, {privatetest, "private"},
    629 {partialtest, "partial"}, {forktest, "fork"}, {exittest, "exit"},
    630 {manytest, "many"}, {lockedtest, "locked"}, {selftest, "self"},
    631 {stresstest, "stress"},
    632};
    633
    634int
    635main(int argc, char *argv[])
    636{
    637 struct test *t;
    638 int ran = 0;
    639
    640 if (argc > 1 && strcmp(argv[1], "nop") == 0)
    641 exit(0); // exittest's exec
    642 for (t = tests; t < &tests[sizeof(tests) / sizeof(tests[0])]; t++) {
    643 if (argc > 1 && strcmp(argv[1], t->name) != 0)
    644 continue;
    645 int before = failed;
    646 failed = 0;
    647 t->f();
    648 if (!failed)
    649 printf("mmaptest: %s: OK\n", t->name);
    650 failed |= before;
    651 ran++;
    652 }
    653 if (ran == 0) {
    654 printf("mmaptest: no test %s\n", argv[1]);
    655 exit(1);
    656 }
    657 if (argc == 1)
    658 printf(failed ? "mmaptest: SOME FAILED\n" : "mmaptest: ALL OK\n");
    659 exit(failed);
    660}

6. Verify and measure

On the branch (ext/15-mmap, 10 commits), built with the project toolchain and run on 3 harts (-smp 3 -m 128M), one boot:

$ mmaptest
mmaptest: read: touching 1 of 3 mapped pages took 1 free pages
usertrap(): unexpected scause 0xd pid=4
            sepc=0x1e stval=0x3fffffb000
mmaptest: read: OK
mmaptest: shared: OK
usertrap(): unexpected scause 0xf pid=5
            sepc=0xa stval=0x3fffffc000
mmaptest: private: OK
usertrap(): unexpected scause 0xd pid=6
            sepc=0x1e stval=0x3fffffa000
usertrap(): unexpected scause 0xd pid=7
            sepc=0x1e stval=0x3fffffd000
mmaptest: partial: OK
mmaptest: fork: OK
mmaptest: exit: OK
mmaptest: many: OK
mmaptest: locked: OK
mmaptest: self: OK
mmaptest: stress: OK
mmaptest: ALL OK
$ usertests -q
usertests starting
test copyin: OK
test copyout: OK
[...]
test nowrite: usertrap(): unexpected scause 0xf pid=6576
[...]
ALL TESTS PASSED
$ mmaptest
mmaptest: read: touching 1 of 3 mapped pages took 1 free pages
usertrap(): unexpected scause 0xd pid=6662
            sepc=0x1e stval=0x3fffffb000
mmaptest: read: OK
[...]
mmaptest: stress: OK
mmaptest: ALL OK

The usertrap() lines in mmaptest are the expected kills: in read a child loads from the mapping after munmap (stval is its first page); in private a child stores into a read-only mapping (scause 0xf); in partial two children load from the first and the last page after they were unmapped. A second boot printed the same, with the same pids.

usertests -q passing shows that what the original kernel did still works: the heap’s sbrk tests (eager, lazy, shrinking, running out of memory) see the same limit as before when no mapping exists, copyout into text still fails, and no page is lost. Every commit builds with make kernel/kernel fs.img, and usertests -q printed ALL TESTS PASSED on 3 harts at each of the 10 commits.

mmaptest stress alone, 10 times in a row on one boot: 10 times OK.

Pages cost only when touched. mmaptest read counts free pages around one touch of a 3-page mapping: touching 1 of 3 mapped pages took 1 free pages, on every run. mmap itself allocates nothing, and no page-table page is needed either: the mapping lies in the same 2 MiB as the trapframe, under a level-0 page table that exists from the start.

Write back what changed. A scratch benchmark (not on the branch) maps a 64-page file shared, reads one byte from every page, stores one byte into page 0, and unmaps, 10 times; a copy of each kernel printed how many pages each such munmap wrote back. Three runs each, one boot each:

kernel pages written back per munmap 10 rounds, ticks
reference (D bit) 1 8, 8, 7
clinic 6 (every loaded page) 64 60, 100, 53

A tick is 1,000,000 timer cycles (kernel/trap.c:179), 0.1 s on QEMU’s 10 MHz clock: about 0.8 s against 5 to 10 s for the same ten rounds. Each extra page is a transaction, five blocks (4 data, 1 inode) copied into the log and back, and a disk write per block; the loads (64 per round) are the same in both. (The second row varies a lot from run to run; that was not investigated.)

Correctness, measured. mmaptest stress: 10 of 10 OK on the reference; 1 of 6 OK with clean pages written back (5 FAIL (file contents)).

Where loads and write-backs happen. In the gdb runs of the finished branch, loads came from user page faults (interrupts off, no locks) and from the three prefaults (interrupts on, no locks), never from a copy made under a lock. Write-backs came from munmap, from kexit and from kexec. In the stress run both happened on all three harts, loads slept on the disk, and write-backs slept in begin_op for log space: with MAXOPBLOCKS 10 reserved per transaction and 30 log blocks, three write-backs fill the log’s reservation although each logs only 5 blocks.

7. Go further