xv6, line by line
lab 17

Extension labs · lab 17 · Memory · ★★★☆☆

Direct user access with sstatus.SUM

Every time a system call reads or writes user memory, this tree’s copyin and copyout translate the user’s address in software: walkaddr walks the process’s page table, finds the physical page, and the kernel copies through its own direct map of RAM. The user pointer itself is never dereferenced. Real kernels (Linux on RISC-V among them) do the opposite: they load and store through the user’s virtual address directly, and let the hardware translate it. In this lab you make xv6 do that.

RISC-V has a bit for exactly this purpose, SUM in the sstatus register. Setting it turns out to be the easy part. Which page table is in satp while a system call runs, and what does it map at a user address? If you put the user’s pages into it, what is already there? What becomes of a user pointer that points at the kernel, at a device, or at a page that does not exist yet, now that the hardware, not your code, does the lookup? A page fault inside the kernel has always meant panic in this tree: what must it mean now? And is the result actually faster? The think section asks these in the order a designer meets them, and the measure section answers the last one with numbers that may surprise you.

The reference solution is seven commits. usertests -q passes on 3 harts at every one of them, and a new test, sumtest, checks direct copies with good and bad pointers.

Read first: Tour 6: System-call arguments and user pointers, Tour 8: Traps taken inside the kernel, Tour 12: One scheduler per hart, Tour 20: fork, Tour 22: exec, Tour 24: The kernel page table and turning paging on, Tour 25: A user address space, Tour 26: sbrk, eager and lazy, and page faults, Tour 28: Crossing the user/kernel boundary in memory · Mode, stack and page table: the master question, The stacks of xv6, Locks and interrupt state

What this lab teaches

  • Which page table the hardware uses while the kernel runs a system call, and what a user address means in it.
  • What sstatus.SUM permits and forbids, why it is off by default, and what its neighbour MXR does.
  • How the kernel’s address space and a process’s address space fit (or do not fit) into one page table.
  • How to keep two page tables that describe the same user memory in step, which changes need a TLB flush, and when the hart must stop using a page table that is about to be freed.
  • What a user pointer can reach once the hardware translates it, and what the kernel must still check in software.
  • How a kernel can survive a page fault taken in its own code, and what problem Linux’s exception tables solve.
  • How to measure a change honestly on QEMU, and how to take a slowdown apart into its causes.

The reference branch

ext/17-sum in ShowMeTheStack/xv6-riscv-labs, branched from the frozen commit 06aad25; 7 commits.

git clone https://github.com/ShowMeTheStack/xv6-riscv-labs
cd xv6-riscv-labs
git checkout -b my-sum 06aad25   # start your own
git diff 06aad25 origin/ext/17-sum   # only when you want the answer

1. The spec

Behaviour. For the current process, copyin, copyout and copyinstr copy with ordinary loads and stores to the user’s virtual address, with sstatus.SUM set for the duration of the copy and clear at all other times. They no longer walk the process’s page table in software. Everything a system call returns stays the same:

usertests. If your design cannot give a process everything a test asks for, you may change that test’s sizes or targets, but each test must keep its purpose, and you must say exactly what you changed and what the test now exercises.

What must not change.

The test program, sumtest, prints one line per check:

$ sumtest
sumtest: copy: OK
sumtest: lazy: OK
sumtest: bad pointers: OK
sumtest: shrink: OK
sumtest: limit: OK
sumtest: fork: OK
sumtest: ALL OK

sumtest time counts fstat calls, 512-byte pipe round trips and 8 KB file reads in 20 ticks, for the measure section.

2. Think first

Answer each question in your head (or on paper) before opening a hint. Hints get more specific; the reference answer comes last.

1Why can't the kernel just use the pointer?

read(fd, buf, n) hands the kernel a user virtual address, buf. Today copyout never dereferences it. Suppose a system call simply did *(int *)buf = 1 in the kernel. Which page table translates that store, and what does that page table have at the address of a typical user buffer, say 0x2000? At 0x80001000? Commit to an answer, then work out what copyout does instead, and what that costs per page.

Check yourself

1warm-upChoose one

A buggy system call in the unmodified kernel executes *(int *)addr, where addr is 0x1000, the address of a global variable in the calling program. What happens?

2solidType a number

On the unmodified kernel, read asks copyout to write 4096 bytes to the user address 0x2010 (not page-aligned; both pages are mapped and writable). How many times does copyout call walk, directly or through walkaddr?

kernel/vm.c
344int
350 while (len > 0) {
352 if (va0 >= MAXVA)
353 return -1;
356 if (pa0 == 0) {
357 if ((pa0 = vmfault(pagetable, psz, va0, 0)) == 0) {
358 return -1;
359 }
360 }
363 // forbid copyout over read-only user text pages.
364 if ((*pte & PTE_W) == 0)
365 return -1;
367 n = PGSIZE - (dstva - va0);
368 if (n > len)
369 n = len;
370 memmove((void *)(pa0 + (dstva - va0)), src, n);
372 len -= n;
373 src += n;
375 }
376 return 0;
decimal, 0x hex or 0b binary

2What SUM does, and why it is off

Suppose satp did map the user’s pages, with the same PTEs as the user page table, PTE_U included. Would a supervisor-mode load from such a page succeed? Find the rule in the RISC-V privileged specification. Then decide why the architects made that the default, and what a kernel should do about it.

Check yourself

1solidDecode the bits

On the reference branch, gdb stopped inside the copy loop, during pipe() copying the two file descriptors out to sumtest, and printed sstatus. Decode it.

Value: 0x200040022

2warm-upTrue or false, and why

True or false: once the page table in satp maps a process’s pages, setting SUM lets supervisor mode load from, store to and execute those pages.

Why?

3Which page table?

The previous answers ask for a page table that maps both the kernel and the current process’s pages, at their own virtual addresses. Neither table in this tree is that: the user page table has no kernel text, data or stacks, the kernel page table no user pages. Make one that has both. Before choosing how, lay the two address spaces over each other. Where do they collide, and what will you give up to remove the collision?

Check yourself

1warm-upType a number

With user memory capped at the PLIC’s address, how many MB can a process have at most (code, data, stack and heap together)?

kernel/memlayout.h
30#define CLINT(hart) (CLINT_BASE + (hart) * 4)
32// qemu puts platform-level interrupt controller (PLIC) here.
33#define PLIC 0x0c000000L
34#define PLIC_PRIORITY (PLIC + 0x0)
35#define PLIC_PENDING (PLIC + 0x1000)
36#define PLIC_SENABLE(hart) (PLIC + 0x2080 + (hart) * 0x100)
37#define PLIC_SPRIORITY(hart) (PLIC + 0x201000 + (hart) * 0x2000)
38#define PLIC_SCLAIM(hart) (PLIC + 0x201004 + (hart) * 0x2000)
40// the kernel expects there to be RAM
41// for use by the kernel and user pages
42// from physical address 0x80000000 to PHYSTOP.
43#define KERNBASE 0x80000000L
44#define PHYSTOP (KERNBASE + 128 * 1024 * 1024)
decimal, 0x hex or 0b binary
2solidChoose all that apply

Which kernel mappings would a process collide with if its memory could still grow up to TRAPFRAME?

4Building a process's kernel page table

A process needs a page table with every kernel mapping and its own user pages. Copying the whole kernel page table for every process would cost dozens of pages. What can be shared with kernel_pagetable, and what must be private? Then: which of the process’s PTEs go into it? All of them, or only some? Think about the pages a user PTE can describe.

Check yourself

1solidDecode the bits

gdb printed this PTE from the shell’s user page table, for virtual address 0x3000. Decode it, and decide whether it belongs in the shell’s kernel page table.

Value: 0x21fccc07

2deepChoose one

Why can a process’s kernel page table share the kernel’s level-1 page for RAM (root entry 2) but not the one for the first gigabyte (root entry 0)?

5Two page tables, one truth

Every user mapping now exists twice. List every place in this tree where a user mapping is created, changed or removed, and say what the process’s kernel page table must do at each. Some of those places change the page table the hart is running on: which of them need an sfence.vma, and which can do without one? And when, exactly, must a hart stop using a process’s kernel page table?

Check yourself

1solidChoose all that apply

Which of these must flush this hart’s TLB (with sfence.vma) after updating the process’s kernel page table?

2deepChoose one

In the reference scheduler, why must kvminithart() (back to kernel_pagetable) come right after swtch returns and before release(&p->lock)?

6SUM is not a bounds check

With SUM set and p->kpagetable in satp, a copy loop can use any address the process passes. Which addresses that a process might pass will the hardware still accept, although they do not belong to the process? Decide what check the copy functions must make before the loop, and what they should do with a range that starts inside the process’s memory and ends outside it.

Check yourself

1solidChoose one

On a branch where the copy functions skip the range check, which usertests check fails first?

7A page fault inside the kernel

Even with the range check, a direct copy can fault: on a page reserved with sbrklazy and never touched, on the stack guard page, on a store into program text. The fault is taken in supervisor mode and lands in kerneltrap, which today panics. What should it do in each case, and how does it know the fault came from a copy and not from a kernel bug? Why can the copy loop not simply be memmove? Finally: a timer interrupt can arrive in the middle of a copy, with SUM set. What must kerneltrap do about SUM, given that it may call yield?

Check yourself

1solidPut in order

sumtest calls read from a pipe into an untouched sbrklazy page. Put the events in order, as gdb recorded them on the reference branch.

  1. piperead calls copyout, which calls ucopy; ucopy sets SUM
  2. the sb in ucopy raises a store page fault (scause 15) in supervisor mode
  3. kernelvec pushes a frame and calls kerneltrap, which clears SUM
  4. sret returns to the same sb, which now succeeds
  5. kerneltrap flushes the TLB and writes back the saved sepc and sstatus (SUM set again)
  6. uvmfault maps a zeroed page in the user page table and in the kernel page table
2deepTrue or false, and why

True or false: kerneltrap clearing SUM on entry is only tidiness; since the interrupted copy’s SUM is restored at the end, leaving it set during kerneltrap could not affect anything else.

Why?

8Will it be faster?

Before you measure, predict. The old path per copy call: two software walks per page, then a byte-copy memmove. The new path: a check of the range, setting and clearing SUM (two CSR writes), the copy loop. What else did the design add to every call, and to every context switch? On QEMU, which do you expect to win for a 1-byte copy, and for a 4096-byte copy? How would you find out why?

Check yourself

1solidChoose one

In the in-kernel timings on QEMU, the branch’s 1-byte copyout took 200 to 260 ns longer than the original’s. Which single ingredient accounts for most of that?

3. Build it

Start.

git checkout -b my-sum 06aad25

Write user/sumtest.c first, with the checks listed in the spec, and add $U/_sumtest\ to UPROGS in the Makefile. On the unmodified kernel every check passes except limit (sbrk beyond USERTOP succeeded): the old kernel already refuses all the bad pointers. sumtest guards against breaking that. Define USERTOP in milestone 1 before you build it.

Milestones, in an order that keeps usertests -q passing after each one.

  1. The cap. USERTOP in kernel/memlayout.h; use it in growproc, sys_sbrk and kexec (each segment’s end, and the stack). Lower REGION_SZ in usertests.c and point lazy_sbrk at USERTOP. Test: usertests -q (without the lazy_sbrk change it loops forever: its loop ignores sbrklazy failing).
  2. A kernel page table per process. p->kpagetable; create it in allocproc, free it in freeproc, install it in scheduler around swtch. No user pages in it yet. Test: usertests -q. Every system call now runs on it (it is what prepare_return stores as kernel_satp), so a mistake in the shared kernel part shows up at once.
  3. The mirror. A function that makes a range of the kernel page table agree with the user page table; call it in kfork, growproc, the user-mode lazy fault path and kexec. Test: usertests -q. Nothing uses the user part yet, so check it with gdb: after ls, compare the shell’s user and kernel PTEs for 0x0-0x4000.
  4. The copy loops. SSTATUS_SUM in riscv.h; ucopy and ucopystr in a new kernel/uaccess.S (add it to OBJS), with labels marking where the loops end and where a faulting copy resumes. Nothing calls them yet. Test: it builds.
  5. Fault recovery. In kerneltrap: clear SUM on entry; for page faults inside the loops, the lazy handler or the resume label. Test: usertests -q.
  6. Switch the copy functions. Range check, then the loops, for the current process; keep the walk for kexec's new page table. Test: sumtest, usertests -q, sumtest.

Debugging advice. Start QEMU with make qemu-gdb, which waits for gdb on the port it prints: set boot-time breakpoints before the first continue; for anything after the shell prompt, let it boot and interrupt gdb (Ctrl-C) when you need it.

4. Debugging clinic

Each of these bugs was put into the reference solution on purpose and run on three harts. The symptom is exactly what happened. Try to explain it before revealing why.

1No fault recovery in kerneltrap

Commit 5’s new branch in kerneltrap is missing (the line that clears SUM is kept):

-  if ((scause == 13 || scause == 15) && sepc >= (uint64)ucopy &&
-      sepc < (uint64)ucopyend) {
-    // a page fault in a user copy: map a lazily-allocated
-    // page and retry, or make the copy return -1.
-    if (uvmfault(myproc(), r_stval(), scause == 13) != 0)
-      sfence_vma(); // so that the retry sees the new PTE
-    else
-      sepc = (uint64)ucopyfail;
-  } else if ((which_dev = devintr()) == 0) {
+  if ((which_dev = devintr()) == 0) {

What happened when we ran it

$ usertests -q
usertests starting
test copyin: OK
test copyout: scause=0xf sepc=0x800058ec stval=0x0
panic: kerneltrap

(another boot)
$ sumtest
sumtest: copy: OK
scause=0xf sepc=0x800058ec stval=0x800a
panic: kerneltrap

2growproc does not update the kernel page table when memory grows

growproc mirrors a shrink but forgets the grow case:

     if ((sz = uvmalloc(p->pagetable, sz, sz + n, PTE_W)) == 0) {
       return -1;
     }
-    if (kvmsync(p->pagetable, p->kpagetable, p->sz, sz) < 0) {
-      uvmdealloc(p->pagetable, sz, p->sz);
-      kvmsync(p->pagetable, p->kpagetable, p->sz, sz);
-      return -1;
-    }
   } else if (n < 0) {

What happened when we ran it

init: starting sh
$ echo hi
exec echo failed
$ ls
exec ls failed
$ cat README
exec cat failed
$ cd /
$

# gdb, the same kernel, typing echo hi:
=== growproc: pid 3 sh n=65536 sz=0x5000
=== kerneltrap ucopy fault: hart 2 pid 3 sh
$1 = 0xd
$2 = 0x14f58
$3 = 0x15000
[...]
=== ucopyfail: pid 3
#0  ucopyfail () at kernel/uaccess.S:65
#1  0x000000008000166e in copyin (pagetable=<optimized out>, psz=86016, dst=dst@entry=0x3fffff9de0 "\005\005\005\005\005\005\005\005XO\001", srcva=srcva@entry=85848, len=len@entry=8) at kernel/vm.c:486
#2  0x0000000080002a4a in fetchaddr (addr=85848, ip=ip@entry=0x3fffff9de0) at kernel/syscall.c:18
#3  0x0000000080005714 in sys_exec () at kernel/sysfile.c:474
[...]

3User memory may still grow up to TRAPFRAME

Commit 1 is left out: growproc, sys_sbrk and kexec still allow memory up to TRAPFRAME, and usertests.c is the original one. Everything else is the reference.

What happened when we ran it

$ usertests lazy_alloc
usertests starting
test lazy_alloc: OK
FAILED -- lost some free pages 32003 (out of 32386)
$ usertests lazy_unmap
usertests starting
test lazy_unmap:
[... no more output; after several minutes, gdb:]
  Id   Target Id                    Frame
* 1    Thread 1.1 (CPU#0 [halted ]) s_sstatus (x=2) at kernel/riscv.h:69
  2    Thread 1.2 (CPU#1 [halted ]) s_sstatus (x=2) at kernel/riscv.h:69
  3    Thread 1.3 (CPU#2 [running]) kernelvec () at kernel/kernelvec.S:14
[...]
# thread 3: p/x $scause, $stval, $sepc, $sp
$1 = 0xf
$2 = 0x1ec2130ef0
$3 = 0x800058a2
$4 = 0x1ec2130ef0
[...]
# the kernel's level-0 page-table pages for the UART/virtio page and the
# PLIC's first pages:
0x87ffd000:	0x0000000000000000	0x0000000000000000
0x87ffc000:	0x0000000000000000	0x0000000000000000
0x87ffb000:	0x0000000000000000	0x0000000000000000

4The copy functions trust the user’s range

The range check is left out; the loops get the whole range:

-    n = ulen(dstva, len, psz);
-    if (ucopy((char *)dstva, src, n) < 0 || n < len)
-      return -1;
-    return 0;
+    return ucopy((char *)dstva, src, len);

and the same in copyin and copyinstr (which passes max unchanged).

What happened when we ran it

$ usertests -q
usertests starting
test copyin: write(fd, 0x0000000080000000, 8192) returned 8192, not -1
FAILED
SOME TESTS FAILED
$

(another boot)
$ sumtest
sumtest: copy: OK
sumtest: lazy: OK
sumtest: write(pipe, 0x0000000080000000) from kernel text (KERNBASE) returned 16
sumtest: read(fd, 0x0000000087FFF000) into kernel RAM (PHYSTOP - PGSIZE) returned 16
scause=0panic:

# gdb, attached afterwards (another boot of the same test, which stopped
# printing in the middle of the same line):
0x87fff000:	0x6120736920367678	0x6c706d692d657220
0x87fff000:	"xv6 is a re-impl\001p\377!"

5ucopystr leaves SUM set when it finds the ‘\0’

An early return that forgets to clean up, the classic mistake:

 2:
-        csrc sstatus, t0
         li a0, 0
         ret

To see the effect, both this kernel and the reference got a deliberately buggy system call, kpeek(int *addr), that returns *(int *)addr, a user pointer used without copyin. The test program calls open("README", ...) (a copyinstr that finds the '\0'), then kpeek(&secret) with secret = 42, then prints two lines with printf, then calls kpeek again.

What happened when we ran it

# the reference branch plus kpeek:
$ peektest
peektest: secret is at 0x0000000000001000
scause=0xd sepc=0x80002d9a stval=0x1000
panic: kerneltrap

# with the bug:
$ peektest
peektest: secret is at 0x0000000000001000
peektest: right after open(), kpeek returned 42
peektest: and after these printf()s?
scause=0xd sepc=0x80002d9a stval=0x1000
panic: kerneltrap

6The scheduler stays on the process’s kernel page table

The switch back after swtch is missing:

         kvmswitch(p->kpagetable);
         swtch(&c->context, &p->context);
-
-        // back to the kernel's own page table: p's may be
-        // freed once p->lock is released.
-        kvminithart();

What happened when we ran it

init: starting sh
$ usertests -q
[... no more output for 900 seconds]

(another boot)
init: starting sh
$
[... the typed command is never echoed; gdb:]
  Id   Target Id                    Frame
* 1    Thread 1.1 (CPU#0 [running]) kernelvec () at kernel/kernelvec.S:14
  2    Thread 1.2 (CPU#1 [halted ]) s_sstatus (x=2) at kernel/riscv.h:69
  3    Thread 1.3 (CPU#2 [running]) kernelvec () at kernel/kernelvec.S:14
[...]
Thread 3 (Thread 1.3 (CPU#2 [running])):
$1 = 0x8000000000087f52
Thread 2 (Thread 1.2 (CPU#1 [halted ])):
$2 = 0x8000000000087f31
Thread 1 (Thread 1.1 (CPU#0 [running])):
$3 = 0x8000000000087f52
[...]
0x87f52000:	0x8000000000087f31	0x0000003fffffc000
0x87f52010:	0x00000000800027ee	0x0000000000000ca4

5. The reference solution

Take the guided tour through the reference solution, one commit at a time, with the machine state at every step:

Open the reveal tour →

Or read the commits

  1. 32ba84a Limit user memory to below the PLIC

    kernel/exec.c

    @@ -67,8 +67,10 @@ kexec(char *path, char **argv)
    6767 if (ph.vaddr + ph.memsz < ph.vaddr)
    6868 goto bad;
    6969 if (ph.vaddr % PGSIZE != 0)
    7070 goto bad;
    71 if (ph.vaddr + ph.memsz > USERTOP)
    72 goto bad;
    7173 uint64 sz1;
    7274 if ((sz1 = uvmalloc(pagetable, sz, ph.vaddr + ph.memsz,
    7375 flags2perm(ph.flags))) == 0)
    7476 goto bad;
    @@ -86,8 +88,10 @@ kexec(char *path, char **argv)
    8688 // Allocate some pages at the next page boundary.
    8789 // Make the first inaccessible as a stack guard.
    8890 // Use the rest as the user stack.
    8991 sz = PGROUNDUP(sz);
    92 if (sz + (USERSTACK + 1) * PGSIZE > USERTOP)
    93 goto bad;
    9094 uint64 sz1;
    9195 if ((sz1 = uvmalloc(pagetable, sz, sz + (USERSTACK + 1) * PGSIZE, PTE_W)) ==
    9296 0)
    9397 goto bad;

    kernel/memlayout.h

    @@ -56,8 +56,13 @@
    5656// text
    5757// original data and bss
    5858// fixed-size stack
    5959// expandable heap
    60// ...
    60// ... up to USERTOP
    6161// TRAPFRAME (p->trapframe, used by the trampoline)
    6262// TRAMPOLINE (the same page as in the kernel)
    6363#define TRAPFRAME (TRAMPOLINE - PGSIZE)
    64
    65// user memory must end below the lowest device that the kernel
    66// maps (the PLIC), so that a kernel page table can also map a
    67// process's user pages at their user addresses.
    68#define USERTOP PLIC

    kernel/proc.c

    @@ -239,9 +239,9 @@ growproc(int n)
    239239 struct proc *p = myproc();
    240240
    241241 sz = p->sz;
    242242 if (n > 0) {
    243 if (sz + n > TRAPFRAME) {
    243 if (sz + n > USERTOP) {
    244244 return -1;
    245245 }
    246246 if ((sz = uvmalloc(p->pagetable, sz, sz + n, PTE_W)) == 0) {
    247247 return -1;

    kernel/sysproc.c

    @@ -56,9 +56,9 @@ sys_sbrk(void)
    5656 // size but don't allocate memory. If the processes uses the
    5757 // memory, vmfault() will allocate it.
    5858 if (addr + n < addr)
    5959 return -1;
    60 if (addr + n > TRAPFRAME)
    60 if (addr + n > USERTOP)
    6161 return -1;
    6262 myproc()->sz += n;
    6363 }
    6464 return addr;

    user/usertests.c

    @@ -2619,9 +2619,11 @@ badarg(char *s)
    26192619
    26202620 exit(0);
    26212621}
    26222622
    2623#define REGION_SZ (1024 * 1024 * 1024)
    2623// user memory ends at USERTOP (192 MB, the PLIC's address);
    2624// 176 MB still reserves more than all of RAM.
    2625#define REGION_SZ (176 * 1024 * 1024)
    26242626
    26252627// Touch a page every 64 pages, which with lazy allocation
    26262628// causes one page to be allocated.
    26272629void
    @@ -2777,31 +2779,22 @@ lazy_copyinstr(char *s)
    27772779
    27782780void
    27792781lazy_sbrk(char *s)
    27802782{
    2781 // sbrk() takes just int, so take 2^30-sized steps towards MAXVA
    2783 // grow to one page below USERTOP; less than 2^31, so one
    2784 // sbrk() call's int argument is enough.
    27822785 char *p = sbrk(0);
    2783 while ((uint64)p < MAXVA - (1 << 30)) {
    2784 p = sbrklazy(1 << 30);
    2785 if (p < 0) {
    2786 printf("sbrklazy(%d) returned %p\n", 1 << 30, p);
    2787 exit(1);
    2788 }
    2789
    2790 p = sbrklazy(0);
    2791 }
    2792
    2793 int n = TRAPFRAME - PGSIZE - (uint64)p;
    2786 int n = USERTOP - PGSIZE - (uint64)p;
    27942787
    27952788 char *p1 = sbrklazy(n);
    27962789 if (p1 < 0 || p1 != p) {
    27972790 printf("sbrklazy(%d) returned %p, not expected %p\n", n, p1, p);
    27982791 exit(1);
    27992792 }
    28002793
    28012794 p = sbrk(PGSIZE);
    2802 if (p < 0 || (uint64)p != TRAPFRAME - PGSIZE) {
    2803 printf("sbrk(%d) returned %p, not expected TRAPFRAME-PGSIZE\n", PGSIZE, p);
    2795 if (p < 0 || (uint64)p != USERTOP - PGSIZE) {
    2796 printf("sbrk(%d) returned %p, not expected USERTOP-PGSIZE\n", PGSIZE, p);
    28042797 exit(1);
    28052798 }
    28062799
    28072800 p[0] = 1;
  2. 139d9e4 Give each process its own kernel page table

    kernel/defs.h

    @@ -153,8 +153,11 @@ void uartputc_sync(int);
    153153
    154154// vm.c
    155155void kvminit(void);
    156156void kvminithart(void);
    157pagetable_t kvmcreate(void);
    158void kvmfree(pagetable_t);
    159void kvmswitch(pagetable_t);
    157160void kvmmap(pagetable_t, uint64, uint64, uint64, int);
    158161int mappages(pagetable_t, uint64, uint64, uint64, int);
    159162pagetable_t uvmcreate(void);
    160163uint64 uvmalloc(pagetable_t, uint64, uint64, int);

    kernel/proc.c

    @@ -139,8 +139,16 @@ found:
    139139 release(&p->lock);
    140140 return 0;
    141141 }
    142142
    143 // A kernel page table, with no user pages yet.
    144 p->kpagetable = kvmcreate();
    145 if (p->kpagetable == 0) {
    146 freeproc(p);
    147 release(&p->lock);
    148 return 0;
    149 }
    150
    143151 // Set up new context to start executing at forkret,
    144152 // which returns to user space.
    145153 memset(&p->context, 0, sizeof(p->context));
    146154 p->context.ra = (uint64)forkret;
    @@ -160,8 +168,11 @@ freeproc(struct proc *p)
    160168 p->trapframe = 0;
    161169 if (p->pagetable)
    162170 proc_freepagetable(p->pagetable, p->sz);
    163171 p->pagetable = 0;
    172 if (p->kpagetable)
    173 kvmfree(p->kpagetable);
    174 p->kpagetable = 0;
    164175 p->sz = 0;
    165176 p->pid = 0;
    166177 p->name[0] = 0;
    167178 p->chan = 0;
    @@ -449,10 +460,15 @@ scheduler(void)
    449460 // to release its lock and then reacquire it
    450461 // before jumping back to us.
    451462 p->state = RUNNING;
    452463 c->proc = p;
    464 kvmswitch(p->kpagetable);
    453465 swtch(&c->context, &p->context);
    454466
    467 // back to the kernel's own page table: p's may be
    468 // freed once p->lock is released.
    470
    455471 // Don't re-enable interrupts on release.
    456472 mycpu()->intena = 0;
    457473
    458474 // Process is done running for now.

    kernel/proc.h

    @@ -95,8 +95,9 @@ struct proc {
    9595 // these are private to the process, so p->lock need not be held.
    9696 uint64 kstack; // Virtual address of kernel stack
    9797 uint64 sz; // Size of process memory (bytes)
    9898 pagetable_t pagetable; // User page table
    99 pagetable_t kpagetable; // Kernel page table, used while p runs
    99100 struct trapframe *trapframe; // data page for trampoline.S
    100101 struct context context; // swtch() here to run process
    101102 struct file *ofile[NOFILE]; // Open files
    102103 struct inode *cwd; // Current directory

    kernel/vm.c

    @@ -82,8 +82,56 @@ kvminithart()
    8282 // flush stale entries from the TLB.
    8383 sfence_vma();
    8484}
    8585
    86// Make a kernel page table for one process. It shares the
    87// kernel's page-table pages for everything above the first
    88// gigabyte. The first gigabyte holds the devices and, later,
    89// the process's user pages below USERTOP, so it gets a private
    90// level-1 page that starts as a copy of the kernel's.
    91// returns 0 if out of memory.
    92pagetable_t
    93kvmcreate(void)
    94{
    95 pagetable_t kpt, l1;
    96
    97 if ((kpt = (pagetable_t)kalloc()) == 0)
    98 return 0;
    99 if ((l1 = (pagetable_t)kalloc()) == 0) {
    100 kfree(kpt);
    101 return 0;
    102 }
    104 memmove(l1, (void *)PTE2PA(kernel_pagetable[0]), PGSIZE);
    105 kpt[0] = PA2PTE(l1) | PTE_V;
    106 return kpt;
    107}
    108
    109// Free a process's kernel page table: the private level-1 page
    110// and the level-0 pages under it below USERTOP. The user pages
    111// belong to the user page table, and the rest to the kernel.
    112void
    113kvmfree(pagetable_t kpt)
    114{
    115 pagetable_t l1 = (pagetable_t)PTE2PA(kpt[0]);
    116
    117 for (int i = 0; i < PX(1, USERTOP); i++) {
    118 if (l1[i] & PTE_V)
    119 kfree((void *)PTE2PA(l1[i]));
    120 }
    121 kfree(l1);
    122 kfree(kpt);
    123}
    124
    125// Switch this hart to page table kpt and flush its TLB.
    126void
    127kvmswitch(pagetable_t kpt)
    128{
    129 sfence_vma();
    130 w_satp(MAKE_SATP(kpt));
    131 sfence_vma();
    132}
    133
    86134// Return the address of the PTE in page table pagetable
    87135// that corresponds to virtual address va. If alloc!=0,
    88136// create any required page-table pages.
    89137//
  3. fc0bfd9 Map user pages in the process's kernel page table too

    kernel/defs.h

    @@ -156,8 +156,9 @@ void kvminit(void);
    156156void kvminithart(void);
    157157pagetable_t kvmcreate(void);
    158158void kvmfree(pagetable_t);
    159159void kvmswitch(pagetable_t);
    160int kvmsync(pagetable_t, pagetable_t, uint64, uint64);
    160161void kvmmap(pagetable_t, uint64, uint64, uint64, int);
    161162int mappages(pagetable_t, uint64, uint64, uint64, int);
    162163pagetable_t uvmcreate(void);
    163164uint64 uvmalloc(pagetable_t, uint64, uint64, int);
    @@ -172,8 +173,9 @@ int copyout(pagetable_t, uint64, uint64, char *, uint64);
    172173int copyin(pagetable_t, uint64, char *, uint64, uint64);
    173174int copyinstr(pagetable_t, uint64, char *, uint64, uint64);
    174175int ismapped(pagetable_t, uint64);
    175176uint64 vmfault(pagetable_t, uint64, uint64, int);
    177uint64 uvmfault(struct proc *, uint64, int);
    176178
    177179// plic.c
    178180void plicinit(void);
    179181void plicinithart(void);

    kernel/exec.c

    @@ -33,8 +33,9 @@ kexec(char *path, char **argv)
    3333 struct elfhdr elf;
    3434 struct inode *ip;
    3535 struct proghdr ph;
    3636 pagetable_t pagetable = 0, oldpagetable;
    37 pagetable_t kpagetable = 0, oldkpagetable;
    3738 struct proc *p = myproc();
    3839
    3940 begin_op();
    4041
    @@ -132,21 +133,33 @@ kexec(char *path, char **argv)
    132133 if (*s == '/')
    133134 last = s + 1;
    134135 safestrcpy(p->name, last, sizeof(p->name));
    135136
    137 // A kernel page table that maps the new image.
    138 if ((kpagetable = kvmcreate()) == 0)
    139 goto bad;
    140 if (kvmsync(pagetable, kpagetable, 0, sz) < 0)
    141 goto bad;
    142
    136143 // Commit to the user image.
    137144 oldpagetable = p->pagetable;
    145 oldkpagetable = p->kpagetable;
    138146 p->pagetable = pagetable;
    147 p->kpagetable = kpagetable;
    139148 p->sz = sz;
    140149 p->trapframe->epc = elf.entry; // initial program counter = ulib.c:start()
    141150 p->trapframe->sp = sp; // initial stack pointer
    151 kvmswitch(kpagetable); // stop using the old one
    142152 proc_freepagetable(oldpagetable, oldsz);
    153 kvmfree(oldkpagetable);
    143154
    144155 return argc; // this ends up in a0, the first argument to main(argc, argv)
    145156
    146157bad:
    147158 if (pagetable)
    160 if (kpagetable)
    161 kvmfree(kpagetable);
    149162 if (ip) {
    150163 iunlockput(ip);
    151164 end_op();
    152165 }

    kernel/proc.c

    @@ -256,10 +256,17 @@ growproc(int n)
    256256 }
    257257 if ((sz = uvmalloc(p->pagetable, sz, sz + n, PTE_W)) == 0) {
    258258 return -1;
    259259 }
    260 if (kvmsync(p->pagetable, p->kpagetable, p->sz, sz) < 0) {
    261 uvmdealloc(p->pagetable, sz, p->sz);
    262 kvmsync(p->pagetable, p->kpagetable, p->sz, sz);
    263 return -1;
    264 }
    260265 } else if (n < 0) {
    261266 sz = uvmdealloc(p->pagetable, sz, sz + n);
    267 kvmsync(p->pagetable, p->kpagetable, sz, p->sz);
    268 sfence_vma(); // forget the freed pages' translations
    262269 }
    263270 p->sz = sz;
    264271 return 0;
    265272}
    @@ -285,8 +292,15 @@ kfork(void)
    285292 return -1;
    286293 }
    287294 np->sz = p->sz;
    288295
    296 // The child's kernel page table maps its user pages too.
    297 if (kvmsync(np->pagetable, np->kpagetable, 0, np->sz) < 0) {
    298 freeproc(np);
    299 release(&np->lock);
    300 return -1;
    301 }
    302
    289303 // copy saved user registers.
    290304 *(np->trapframe) = *(p->trapframe);
    291305
    292306 // Cause fork to return 0 in the child.

    kernel/trap.c

    @@ -68,10 +68,9 @@ usertrap(void)
    6868 syscall();
    6969 } else if ((which_dev = devintr()) != 0) {
    7070 // ok
    7171 } else if ((r_scause() == 15 || r_scause() == 13) &&
    72 vmfault(p->pagetable, p->sz, r_stval(),
    73 (r_scause() == 13) ? 1 : 0) != 0) {
    72 uvmfault(p, r_stval(), (r_scause() == 13) ? 1 : 0) != 0) {
    7473 // page fault on lazily-allocated page
    7574 } else {
    7675 printk("usertrap(): unexpected scause 0x%lx pid=%d\n", r_scause(), p->pid);
    7776 printk(" sepc=0x%lx stval=0x%lx\n", r_sepc(), r_stval());

    kernel/vm.c

    @@ -130,8 +130,32 @@ kvmswitch(pagetable_t kpt)
    130130 w_satp(MAKE_SATP(kpt));
    131131 sfence_vma();
    132132}
    133133
    134// Make the user part of kernel page table kpt agree with user
    135// page table pagetable for addresses [start, end): copy every
    136// user (PTE_U) leaf PTE, and clear the rest. If kpt is in use,
    137// the caller must flush the TLB after removing mappings.
    138// returns 0, or -1 if out of memory for a page-table page.
    139int
    140kvmsync(pagetable_t pagetable, pagetable_t kpt, uint64 start, uint64 end)
    141{
    142 uint64 va;
    143 pte_t *pte, *kpte;
    144
    145 for (va = PGROUNDDOWN(start); va < end; va += PGSIZE) {
    146 pte = walk(pagetable, va, 0);
    147 if (pte != 0 && (*pte & PTE_V) && (*pte & PTE_U)) {
    148 if ((kpte = walk(kpt, va, 1)) == 0)
    149 return -1;
    150 *kpte = *pte;
    151 } else if ((kpte = walk(kpt, va, 0)) != 0) {
    152 *kpte = 0;
    153 }
    154 }
    155 return 0;
    156}
    157
    134158// Return the address of the PTE in page table pagetable
    135159// that corresponds to virtual address va. If alloc!=0,
    136160// create any required page-table pages.
    137161//
    @@ -524,8 +548,26 @@ vmfault(pagetable_t pagetable, uint64 psz, uint64 va, int read)
    524548 }
    525549 return mem;
    526550}
    527551
    552// handle a page fault at user address va of process p: if va
    553// is a lazily-allocated page, map a new page in both of p's
    554// page tables. returns its physical address, or 0.
    555uint64
    556uvmfault(struct proc *p, uint64 va, int read)
    557{
    558 uint64 pa;
    559
    560 if ((pa = vmfault(p->pagetable, p->sz, va, read)) == 0)
    561 return 0;
    562 va = PGROUNDDOWN(va);
    563 if (kvmsync(p->pagetable, p->kpagetable, va, va + PGSIZE) < 0) {
    564 uvmunmap(p->pagetable, va, 1, 1);
    565 return 0;
    566 }
    567 return pa;
    568}
    569
    528570int
    529571ismapped(pagetable_t pagetable, uint64 va)
    530572{
    531573 pte_t *pte = walk(pagetable, va, 0);
  4. 92f0594 Add copy loops that use user addresses with SUM set

    Makefile

    @@ -26,8 +26,9 @@ OBJS = \
    2626 $K/pipe.o \
    2727 $K/exec.o \
    2828 $K/sysfile.o \
    2929 $K/kernelvec.o \
    30 $K/uaccess.o \
    3031 $K/plic.o \
    3132 $K/virtio_disk.o
    3233
    3334# riscv64-unknown-elf- or riscv64-linux-gnu-

    kernel/defs.h

    @@ -175,8 +175,12 @@ int copyinstr(pagetable_t, uint64, char *, uint64, uint64);
    175175int ismapped(pagetable_t, uint64);
    176176uint64 vmfault(pagetable_t, uint64, uint64, int);
    177177uint64 uvmfault(struct proc *, uint64, int);
    178178
    179// uaccess.S
    180int ucopy(char *, char *, uint64);
    181int ucopystr(char *, char *, uint64);
    182
    179183// plic.c
    180184void plicinit(void);
    181185void plicinithart(void);
    182186int plic_claim(void);

    kernel/riscv.h

    @@ -40,8 +40,10 @@ w_mepc(uint64 x)
    4040}
    4141
    4242// Supervisor Status Register, sstatus
    4343
    44#define SSTATUS_MXR (1L << 19) // Make eXecutable pages Readable
    45#define SSTATUS_SUM (1L << 18) // permit Supervisor User Memory access
    4446#define SSTATUS_SPP (1L << 8) // Previous mode, 1=Supervisor, 0=User
    4547#define SSTATUS_SPIE (1L << 5) // Supervisor Previous Interrupt Enable
    4648#define SSTATUS_UPIE (1L << 4) // User Previous Interrupt Enable
    4749#define SSTATUS_SIE (1L << 1) // Supervisor Interrupt Enable

    kernel/uaccess.S

    @@ -0,0 +1,68 @@
    1 #
    2 # copies between kernel memory and user memory that use
    3 # the user's virtual addresses directly. they set
    4 # sstatus.SUM, which lets supervisor-mode loads and stores
    5 # use PTE_U pages, for the copy loop only. the page table
    6 # in satp must map the user's pages: p->kpagetable.
    7 #
    8 # kerneltrap() sends a page fault taken on any instruction
    9 # from ucopy up to ucopyend to ucopyfail.
    10 #
    11
    12.equ SSTATUS_SUM, 1 << 18
    13
    14.globl ucopy
    15.globl ucopystr
    16.globl ucopyend
    17.globl ucopyfail
    18
    19 # int ucopy(char *dst, char *src, uint64 n)
    20 # copy n bytes from src to dst; return 0.
    21ucopy:
    22 li t0, SSTATUS_SUM
    23 csrs sstatus, t0
    241:
    25 beqz a2, 2f
    26 lbu t1, 0(a1)
    27 sb t1, 0(a0)
    28 addi a0, a0, 1
    29 addi a1, a1, 1
    30 addi a2, a2, -1
    31 j 1b
    322:
    33 csrc sstatus, t0
    34 li a0, 0
    35 ret
    36
    37 # int ucopystr(char *dst, char *src, uint64 max)
    38 # copy bytes up to and including a '\0', at most max.
    39 # return 0 if a '\0' was copied, -1 if not.
    40ucopystr:
    41 li t0, SSTATUS_SUM
    42 csrs sstatus, t0
    431:
    44 beqz a2, 3f
    45 lbu t1, 0(a1)
    46 sb t1, 0(a0)
    47 beqz t1, 2f
    48 addi a0, a0, 1
    49 addi a1, a1, 1
    50 addi a2, a2, -1
    51 j 1b
    522:
    53 csrc sstatus, t0
    54 li a0, 0
    55 ret
    563:
    57 csrc sstatus, t0
    58 li a0, -1
    59 ret
    60ucopyend:
    61
    62 # a copy that faulted on a bad user address resumes
    63 # here: turn SUM off and return -1 to the caller.
    64ucopyfail:
    65 li t0, SSTATUS_SUM
    66 csrc sstatus, t0
    67 li a0, -1
    68 ret
  5. 2f2222b Recover from page faults in user copies

    kernel/trap.c

    @@ -10,8 +10,12 @@ struct spinlock tickslock;
    1010uint ticks;
    1111
    1212extern char trampoline[], uservec[];
    1313
    14// in uaccess.S: the user copy loops, and where a faulting copy
    15// resumes.
    16extern char ucopyend[], ucopyfail[];
    17
    1418// in kernelvec.S, calls kerneltrap().
    1519void kernelvec();
    1620
    1721extern int devintr();
    @@ -144,9 +148,21 @@ kerneltrap()
    144148 panic("kerneltrap: not from supervisor mode");
    145149 if (intr_get() != 0)
    146150 panic("kerneltrap: interrupts enabled");
    147151
    148 if ((which_dev = devintr()) == 0) {
    152 // the trap may have interrupted a user copy: keep SUM off
    153 // until the copy resumes (w_sstatus below).
    154 c_sstatus(SSTATUS_SUM);
    155
    156 if ((scause == 13 || scause == 15) && sepc >= (uint64)ucopy &&
    157 sepc < (uint64)ucopyend) {
    158 // a page fault in a user copy: map a lazily-allocated
    159 // page and retry, or make the copy return -1.
    160 if (uvmfault(myproc(), r_stval(), scause == 13) != 0)
    161 sfence_vma(); // so that the retry sees the new PTE
    162 else
    163 sepc = (uint64)ucopyfail;
    164 } else if ((which_dev = devintr()) == 0) {
    149165 // interrupt or trap from an unknown source
    150166 printk("scause=0x%lx sepc=0x%lx stval=0x%lx\n", scause, r_sepc(),
    151167 r_stval());
    152168 panic("kerneltrap");
  6. 202fb4f Copy user memory directly in copyin, copyout, copyinstr

    kernel/vm.c

    @@ -409,8 +409,22 @@ uvmclear(pagetable_t pagetable, uint64 va)
    409409 panic("uvmclear");
    410410 *pte &= ~PTE_U;
    411411}
    412412
    413// how many of the len bytes at user address va lie below sz?
    414// the kernel must not copy beyond them: the page table in satp
    415// also maps kernel memory and devices, without PTE_U, and SUM
    416// does not keep supervisor-mode accesses away from those.
    417static uint64
    418ulen(uint64 va, uint64 len, uint64 sz)
    419{
    420 if (va >= sz)
    421 return 0;
    422 if (len > sz - va)
    423 return sz - va;
    424 return len;
    425}
    426
    413427// Copy from kernel to user.
    414428// Copy len bytes from src to virtual address dstva in a given page table.
    415429// Return 0 on success, -1 on error.
    416430int
    @@ -418,8 +432,18 @@ copyout(pagetable_t pagetable, uint64 psz, uint64 dstva, char *src, uint64 len)
    418432{
    419433 uint64 n, va0, pa0;
    420434 pte_t *pte;
    421435
    436 if (pagetable == myproc()->pagetable) {
    437 // the current process's pages are mapped in satp's page
    438 // table (p->kpagetable): store to dstva directly.
    439 n = ulen(dstva, len, psz);
    440 if (ucopy((char *)dstva, src, n) < 0 || n < len)
    441 return -1;
    442 return 0;
    443 }
    444
    445 // another page table (kexec's new image): walk it.
    422446 while (len > 0) {
    423447 va0 = PGROUNDDOWN(dstva);
    424448 if (va0 >= MAXVA)
    425449 return -1;
    @@ -453,27 +477,15 @@ copyout(pagetable_t pagetable, uint64 psz, uint64 dstva, char *src, uint64 len)
    453477// Return 0 on success, -1 on error.
    454478int
    455479copyin(pagetable_t pagetable, uint64 psz, char *dst, uint64 srcva, uint64 len)
    456480{
    457 uint64 n, va0, pa0;
    458
    459 while (len > 0) {
    460 va0 = PGROUNDDOWN(srcva);
    461 pa0 = walkaddr(pagetable, va0);
    462 if (pa0 == 0) {
    463 if ((pa0 = vmfault(pagetable, psz, va0, 1)) == 0) {
    464 return -1;
    465 }
    466 }
    467 n = PGSIZE - (srcva - va0);
    468 if (n > len)
    469 n = len;
    470 memmove(dst, (void *)(pa0 + (srcva - va0)), n);
    481 uint64 n;
    471482
    472 len -= n;
    473 dst += n;
    474 srcva = va0 + PGSIZE;
    475 }
    483 if (pagetable != myproc()->pagetable)
    484 panic("copyin: not the current page table");
    485 n = ulen(srcva, len, psz);
    486 if (ucopy(dst, (char *)srcva, n) < 0 || n < len)
    487 return -1;
    476488 return 0;
    477489}
    478490
    479491// Copy a null-terminated string from user to kernel.
    @@ -483,45 +495,11 @@ copyin(pagetable_t pagetable, uint64 psz, char *dst, uint64 srcva, uint64 len)
    483495int
    484496copyinstr(pagetable_t pagetable, uint64 psz, char *dst, uint64 srcva,
    485497 uint64 max)
    486498{
    487 uint64 n, va0, pa0;
    488 int got_null = 0;
    489
    490 while (got_null == 0 && max > 0) {
    491 va0 = PGROUNDDOWN(srcva);
    492 pa0 = walkaddr(pagetable, va0);
    493 if (pa0 == 0) {
    494 if ((pa0 = vmfault(pagetable, psz, va0, 1)) == 0) {
    495 return -1;
    496 }
    497 }
    498 n = PGSIZE - (srcva - va0);
    499 if (n > max)
    500 n = max;
    501
    502 char *p = (char *)(pa0 + (srcva - va0));
    503 while (n > 0) {
    504 if (*p == '\0') {
    505 *dst = '\0';
    506 got_null = 1;
    507 break;
    508 } else {
    509 *dst = *p;
    510 }
    511 --n;
    512 --max;
    513 p++;
    514 dst++;
    515 }
    516
    517 srcva = va0 + PGSIZE;
    518 }
    519 if (got_null) {
    520 return 0;
    521 } else {
    522 return -1;
    523 }
    499 if (pagetable != myproc()->pagetable)
    500 panic("copyinstr: not the current page table");
    501 return ucopystr(dst, (char *)srcva, ulen(srcva, max, psz));
    524502}
    525503
    526504// allocate and map user memory if process is referencing a page
    527505// that was lazily allocated in sys_sbrk().
  7. 3bee162 Add sumtest

    Makefile

    @@ -150,8 +150,9 @@ UPROGS=\
    150150 $U/_logstress\
    151151 $U/_forphan\
    152152 $U/_dorphan\
    153153 $U/_sync\
    154 $U/_sumtest\
    154155
    155156fs.img: mkfs/mkfs README $(UPROGS)
    156157 mkfs/mkfs fs.img README $(UPROGS)
    157158

    user/sumtest.c

    @@ -0,0 +1,341 @@
    1// sumtest: checks the kernel's direct copies to and from user
    2// memory (copyin, copyout and copyinstr with sstatus.SUM).
    3//
    4// sumtest run the checks
    5// sumtest time count copies per 20 ticks
    6
    7#include "kernel/types.h"
    8#include "kernel/stat.h"
    9#include "kernel/fcntl.h"
    10#include "kernel/memlayout.h"
    11#include "kernel/riscv.h"
    12#include "user/user.h"
    13
    14char data[100] = "initialized data, copied by the kernel";
    15
    16static int
    17fail(char *what)
    18{
    19 printf("sumtest: %s\n", what);
    20 return 1;
    21}
    22
    23// pass n bytes from src through a pipe into dst.
    24// returns read()'s result.
    25static int
    26viapipe(char *dst, char *src, int n)
    27{
    28 int fds[2], r;
    29
    30 if (pipe(fds) < 0)
    31 return -1;
    32 if (write(fds[1], src, n) != n) {
    33 close(fds[0]);
    34 close(fds[1]);
    35 return -1;
    36 }
    37 r = read(fds[0], dst, n);
    38 close(fds[0]);
    39 close(fds[1]);
    40 return r;
    41}
    42
    43// the kernel copies from and to the stack, data and heap.
    44int
    45copytest(void)
    46{
    47 char stack[100];
    48 struct stat st;
    49 char *heap = sbrk(2 * PGSIZE);
    50 char *edge = heap + PGSIZE - 3; // 6 bytes across a page boundary
    51 int fd;
    52
    53 if (heap == SBRK_ERROR)
    54 return fail("sbrk failed");
    55 if (viapipe(stack, data, 100) != 100 || memcmp(stack, data, 100) != 0)
    56 return fail("data -> pipe -> stack");
    57 if (viapipe(heap, stack, 100) != 100 || memcmp(heap, data, 100) != 0)
    58 return fail("stack -> pipe -> heap");
    59 if (viapipe(edge, "abcdef", 6) != 6 || memcmp(edge, "abcdef", 6) != 0)
    60 return fail("pipe -> across a page boundary");
    61
    62 fd = open("sumfile", O_CREATE | O_RDWR | O_TRUNC);
    63 if (fd < 0 || write(fd, edge, 6) != 6)
    64 return fail("write from across a page boundary");
    65 close(fd);
    66 fd = open("sumfile", O_RDONLY);
    67 if (fd < 0 || fstat(fd, &st) < 0 || st.size != 6)
    68 return fail("fstat");
    69 if (read(fd, stack, 6) != 6 || memcmp(stack, "abcdef", 6) != 0)
    70 return fail("file -> stack");
    71 close(fd);
    72 unlink("sumfile");
    73 sbrk(-2 * PGSIZE);
    74 return 0;
    75}
    76
    77// pages reserved with sbrklazy: the kernel's own access takes
    78// the page fault.
    79int
    80lazytest(void)
    81{
    82 char *p = sbrklazy(3 * PGSIZE);
    83 char buf[8];
    84 struct stat st;
    85 int fd;
    86
    87 if (p == SBRK_ERROR)
    88 return fail("sbrklazy failed");
    89 // copyout into a page nobody has touched
    90 if (viapipe(p + PGSIZE + 10, "hello", 6) != 6 ||
    91 strcmp(p + PGSIZE + 10, "hello") != 0)
    92 return fail("read() into an untouched lazy page");
    93 // copyin from a page nobody has touched: zeros
    94 memset(buf, 'x', sizeof(buf));
    95 if (viapipe(buf, p + 2 * PGSIZE, 8) != 8 ||
    96 memcmp(buf, "\0\0\0\0\0\0\0\0", 8) != 0)
    97 return fail("write() from an untouched lazy page");
    98 sbrklazy(-3 * PGSIZE);
    99 // copyinstr from an untouched page: the name "", which
    100 // open() takes to mean the current directory
    101 p = sbrklazy(PGSIZE);
    102 fd = open(p + 100, O_RDONLY);
    103 if (fd < 0 || fstat(fd, &st) < 0 || st.type != T_DIR)
    104 return fail("open() of a name in an untouched lazy page");
    105 close(fd);
    107 return 0;
    108}
    109
    110// every one of these addresses must be refused, in both
    111// directions, and the kernel must survive.
    112int
    113badtest(void)
    114{
    115 char c;
    116 uint64 stackpage = PGROUNDDOWN((uint64)&c);
    117 uint64 end = (uint64)sbrk(0);
    118 struct {
    119 uint64 addr;
    120 char *what;
    121 } bad[] = {
    122 {0x80000000L, "kernel text (KERNBASE)"},
    123 {PHYSTOP - PGSIZE, "kernel RAM (PHYSTOP - PGSIZE)"},
    124 {0x0c000000L, "the PLIC"},
    125 {0x10000000L, "the UART"},
    126 {stackpage - PGSIZE, "the stack guard page"},
    127 {end, "the first byte beyond sbrk(0)"},
    128 {TRAPFRAME, "the trapframe"},
    129 {0xffffffffffffffffL, "the last address"},
    130 };
    131 int fd, fds[2], n, bads = 0;
    132 char buf[16];
    133
    134 for (int i = 0; i < sizeof(bad) / sizeof(bad[0]); i++) {
    135 char *a = (char *)bad[i].addr;
    136
    137 // copyout
    138 fd = open("README", O_RDONLY);
    139 if ((n = read(fd, a, 16)) > 0) {
    140 printf("sumtest: read(fd, %p) into %s returned %d\n", a, bad[i].what, n);
    141 bads++;
    142 }
    143 close(fd);
    144 // copyin
    145 if (pipe(fds) < 0)
    146 return fail("pipe failed");
    147 if ((n = write(fds[1], a, 16)) > 0) {
    148 read(fds[0], buf, sizeof(buf));
    149 printf("sumtest: write(pipe, %p) from %s returned %d\n", a, bad[i].what,
    150 n);
    151 bads++;
    152 }
    153 close(fds[0]);
    154 close(fds[1]);
    155 // copyinstr
    156 if ((fd = open(a, O_RDONLY)) >= 0) {
    157 printf("sumtest: open(%p) of %s returned %d\n", a, bad[i].what, fd);
    158 close(fd);
    159 bads++;
    160 }
    161 }
    162
    163 // a range that starts below sbrk(0) and ends beyond it: the
    164 // kernel copies the 1 valid byte, then stops
    165 if (pipe(fds) < 0)
    166 return fail("pipe failed");
    167 if ((n = write(fds[1], (char *)end - 1, 16)) > 1) {
    168 printf("sumtest: write(pipe, sbrk(0) - 1, 16) returned %d\n", n);
    169 bads++;
    170 }
    171 close(fds[0]);
    172 close(fds[1]);
    173
    174 // text is readable but not writable
    175 fd = open("README", O_RDONLY);
    176 if ((n = read(fd, (char *)fail, 16)) > 0) {
    177 printf("sumtest: read() into the code of fail() returned %d\n", n);
    178 bads++;
    179 }
    180 close(fd);
    181 if (viapipe(buf, (char *)fail, 16) != 16 || memcmp(buf, fail, 16) != 0) {
    182 printf("sumtest: write() from the code of fail() failed\n");
    183 bads++;
    184 }
    185 return bads != 0;
    186}
    187
    188// after sbrk(-n), the kernel must not reach the old pages.
    189int
    190shrinktest(void)
    191{
    192 char *a = sbrk(2 * PGSIZE);
    193
    194 if (a == SBRK_ERROR)
    195 return fail("sbrk failed");
    196 if (viapipe(a + PGSIZE, "1234567", 8) != 8 || strcmp(a + PGSIZE, "1234567"))
    197 return fail("read() into new sbrk memory");
    198 sbrk(-2 * PGSIZE);
    199 if (viapipe(a + PGSIZE, "1234567", 8) > 0)
    200 return fail("read() into memory given back with sbrk(-n) succeeded");
    201 return 0;
    202}
    203
    204// user memory ends at USERTOP, and the kernel can copy right
    205// up to it.
    206int
    207limittest(void)
    208{
    209 uint64 end = (uint64)sbrk(0);
    210 int n = USERTOP - end;
    211 char *last = (char *)USERTOP - 8;
    212
    213 int r = 0;
    214
    215 if (sbrklazy(n) == SBRK_ERROR)
    216 return fail("sbrklazy up to USERTOP failed");
    217 if (sbrklazy(1) != SBRK_ERROR || sbrk(1) != SBRK_ERROR)
    218 r = fail("sbrk beyond USERTOP succeeded");
    219 last[-1] = 'x'; // a lazy page fault from user mode
    220 if (viapipe(last, "1234567", 8) != 8 || strcmp(last, "1234567") != 0)
    221 r = fail("read() into the last 8 bytes below USERTOP");
    222 // only the 4 bytes below USERTOP can be copied
    223 if (viapipe(last + 4, "1234567", 8) != 4)
    224 r = fail("read() across USERTOP did not stop at USERTOP");
    225 sbrk(-n);
    226 return r;
    227}
    228
    229// a forked child's kernel page table maps what it inherited.
    230int
    231forktest(void)
    232{
    233 char *heap = sbrk(PGSIZE);
    234 char stack[16];
    235 int pid, xstatus;
    236
    237 strcpy(heap, "parent");
    238 strcpy(stack, "parent");
    239 if ((pid = fork()) < 0)
    240 return fail("fork failed");
    241 if (pid == 0) {
    242 if (viapipe(heap, "child", 6) != 6 || strcmp(heap, "child") != 0)
    243 exit(1);
    244 if (viapipe(stack, "child", 6) != 6 || strcmp(stack, "child") != 0)
    245 exit(1);
    246 exit(0);
    247 }
    248 if (wait(&xstatus) != pid || xstatus != 0)
    249 return fail("the child could not read() into inherited memory");
    250 if (strcmp(heap, "parent") != 0 || strcmp(stack, "parent") != 0)
    251 return fail("the child's copies changed the parent's memory");
    252 sbrk(-PGSIZE);
    253 return 0;
    254}
    255
    256// count operations in 20 ticks (about 2 seconds).
    257static int
    258count(int (*op)(void))
    259{
    260 int n = 0, t0;
    261
    262 t0 = uptime();
    263 while (uptime() == t0)
    264 ;
    265 t0 = uptime();
    266 while (uptime() - t0 < 20) {
    267 for (int i = 0; i < 10; i++)
    268 op();
    269 n += 10;
    270 }
    271 return n;
    272}
    273
    274int timefd, timefds[2];
    275char timebuf[8192];
    276
    277static int
    278opfstat(void)
    279{
    280 struct stat st;
    281 return fstat(timefd, &st);
    282}
    283
    284static int
    285oppipe(void)
    286{
    287 write(timefds[1], timebuf, 512);
    288 return read(timefds[0], timebuf, 512);
    289}
    290
    291static int
    292opread(void)
    293{
    294 int fd = open("sumtime", O_RDONLY);
    295 int n = read(fd, timebuf, sizeof(timebuf));
    296 close(fd);
    297 return n;
    298}
    299
    300void
    301timing(void)
    302{
    303 int fd = open("sumtime", O_CREATE | O_RDWR | O_TRUNC);
    304
    305 write(fd, timebuf, sizeof(timebuf));
    306 close(fd);
    307 timefd = open("sumtime", O_RDONLY);
    308 pipe(timefds);
    309 printf("sumtest: fstat (copyout 24 bytes): %d in 20 ticks\n",
    310 count(opfstat));
    311 printf("sumtest: pipe write+read of 512 bytes: %d in 20 ticks\n",
    312 count(oppipe));
    313 printf("sumtest: open+read 8192 bytes+close: %d in 20 ticks\n",
    314 count(opread));
    315 unlink("sumtime");
    316}
    317
    318int
    319main(int argc, char *argv[])
    320{
    321 struct {
    322 int (*f)(void);
    323 char *name;
    324 } tests[] = {
    325 {copytest, "copy"}, {lazytest, "lazy"}, {badtest, "bad pointers"},
    326 {shrinktest, "shrink"}, {limittest, "limit"}, {forktest, "fork"},
    327 };
    328 int failed = 0;
    329
    330 if (argc > 1 && strcmp(argv[1], "time") == 0) {
    331 timing();
    332 exit(0);
    333 }
    334 for (int i = 0; i < sizeof(tests) / sizeof(tests[0]); i++) {
    335 int r = tests[i].f();
    336 printf("sumtest: %s: %s\n", tests[i].name, r ? "FAIL" : "OK");
    337 failed |= r;
    338 }
    339 printf("sumtest: %s\n", failed ? "SOME TESTS FAILED" : "ALL OK");
    340 exit(failed);
    341}

6. Verify and measure

On the branch (ext/17-sum, 7 commits), built with the project toolchain and run on 3 harts (-smp 3 -m 128M), in one boot:

$ sumtest
sumtest: copy: OK
sumtest: lazy: OK
sumtest: bad pointers: OK
sumtest: shrink: OK
sumtest: limit: OK
sumtest: fork: OK
sumtest: ALL OK
$ usertests -q
usertests starting
test copyin: OK
test copyout: OK
[...]
test MAXVAplus: usertrap(): unexpected scause 0xf pid=6519
[...]
test lazy_alloc: OK
[...]
test lazy_sbrk: OK
test partial_write: OK
test unlinkcwd: OK
ALL TESTS PASSED
$ sumtest
sumtest: copy: OK
sumtest: lazy: OK
sumtest: bad pointers: OK
sumtest: shrink: OK
sumtest: limit: OK
sumtest: fork: OK
sumtest: ALL OK
$ ls | wc
29 116 722

The usertrap() lines inside usertests are expected kills (MAXVAplus, nowrite and others store where they may not, on purpose). usertests -q passing shows that the bad-pointer tests (copyin, copyout, copyinstr1, lazy_copy), lazy allocation through system calls (lazy_copy, lazy_copyinstr), partial copies (partial_write), exec with arguments and the free-page count all behave as before.

usertests -q also printed ALL TESTS PASSED on 3 harts at each of commits 1 to 6, each built on its own. Another boot of the head also ran sumtest a third time after ls | wc: ALL OK.

The same sumtest on the original kernel (USERTOP defined just for the build) passes every check except limit, which prints sbrk beyond USERTOP succeeded and read() across USERTOP did not stop at USERTOP: the old kernel already refused all the bad pointers, by walking the page table. The point of sumtest is that the new kernel, which no longer walks, refuses them too.

All numbers are from QEMU 10.2.1, 3 harts, measured while the computer was busy with other work (your times will differ), so the two kernels were alternated run by run and the tables report ranges. The timing instrumentation is not on the branch: a scratch system call ran copyout or copyin 100,000 times (2,000 for 512 and 4096 bytes) on a user buffer and returned the elapsed time CSR ticks (10 MHz, 100 ns each).

Per call, in the kernel (two quietest pairs of runs, ns per call):

original direct (SUM)
copyout, 1 byte 179-220 395-479
copyout, 4096 bytes 5,496-6,289 8,022-8,659
copyin, 1 byte 144-168 400-435
copyin, 4096 bytes 5,356-6,085 8,379-8,870

Over all 11 pairs of runs (some noisier than others) the direct path was slower for 1, 8 and 64 bytes in every pair, about 2x. For 512 and 4096 bytes it was slower in 39 of 44 comparisons, about 1.4x; the 5 exceptions came from runs where the original kernel’s own times jumped by half.

Where the time goes (the same scratch system call timing one ingredient at a time; runs 9 and 10, both kernels):

ingredient ns
myproc (in the new copyout, copyin, copyinstr) 279-315
csrs + csrc of SUM 82-94
csrs + csrc of SPIE, for comparison 83-96
walkaddr (twice per page in the old path), original kernel 86-98
walkaddr, branch kernel, for comparison 47
memmove of 4096 bytes, kernel to kernel 5,051-6,131
ucopy of 4096 bytes to the user buffer 7,647-8,341

So the fixed cost of a direct copy on QEMU is mostly myproc, a push_off and pop_off whose CSR accesses are expensive under emulation, plus about 90 ns to set and clear SUM, which costs no more than toggling any other sstatus bit. The per-byte cost is the loop: 7 instructions per byte in ucopy against 5 in the compiled memmove (kernel/kernel.asm), close to the measured ratio of about 1.5. The two walks it saves cost less than the myproc it adds.

From user mode (sumtest time, operations in 20 ticks, 6 runs each):

original direct (SUM)
fstat (one 24-byte copyout) 36,170-45,920 30,740-46,140
write + read of 512 bytes through a pipe 2,990-4,110 2,340-3,040
open + 8 KB read + close 1,320-1,720 1,240-1,670

fstat and file reads are dominated by the system call and the file system: no difference beyond the noise. Pipes copy one byte per copyin or copyout call, 1,024 calls per round trip, and in every pair of runs the direct kernel completed 21 to 27% fewer round trips (27 to 37% more time per round trip).

Memory. At the first shell prompt, with init and sh running, gdb counted the pages on kmem.freelist: 32,545 on the original kernel and 32,539 on the branch, out of 32,735 pages between the kernel’s end and PHYSTOP in both builds (two separate counts agreed). That is 190 pages in use against 196 on the branch. That is 3 pages for each of init and sh: a root, a private level-1 page, and one level-0 page for their first 2 MB. A process with a 100 MB heap would need another 50 level-0 pages (reasoned, not measured).

Context switches now write satp (with two sfence.vma) before swtch and again after it. We did not measure that cost separately.

The honest conclusion: on QEMU, direct user access makes copies slower, because the ingredients the design adds (a myproc call per copy, two CSR writes) cost more under emulation than the page walks they replace. We have no hardware numbers.

7. Go further