xv6, line by line
lab 10

Extension labs · lab 10 · Memory · ★★☆☆☆

A shared read-only page: system calls without a trap

getpid() asks the kernel for one integer it already knows. To get it, the process executes ecall, the hart traps, uservec saves 31 registers, the kernel switches page tables, usertrap and syscall run, and the whole trip is undone on the way back. In this tree that is 1,103 instructions after the ecall, two writes to satp and four sfence.vma (counted with gdb). In this lab you map one extra page into every process, at a fixed address just below the trapframe, with the process’s pid in it, readable but not writable from user mode. ugetpid() then becomes one load: 12 instructions, no trap.

A second page, the same physical page in every process, carries a copy of the kernel’s ticks, written by clockintr and read with no lock by uuptime(). That raises the question every shared-memory design must answer: can a reader see a half-written value?

The code is small (about 70 lines of kernel, a third of them comments), but placing a page at the top of the user address space touches more than you might expect. What else in the kernel already decides what lives near the top of a user address space? Who creates a page table, who throws one away, and how often does that happen in the life of one process? Does the kernel itself obey a read-only PTE when a system call touches user memory? Answering those is the lab. On QEMU the result is about 500 times faster per call.

Read first: Tour 5: Life of a system call, Tour 7: The trampoline and the trapframe, Tour 20: fork, Tour 22: exec, Tour 25: A user address space, Tour 26: sbrk, eager and lazy, and page faults, Tour 28: Crossing the user/kernel boundary in memory · The stacks of xv6, Locks and interrupt state

What this lab teaches

  • The top of a user address space: TRAMPOLINE (no PTE_U, used only in S-mode), the trapframe page below it, and now two pages that user code may read; how a fixed virtual address becomes a contract between the kernel and the user library.
  • What the PTE permission bits do to each kind of access, and what happens in this tree, where store faults already mean lazy allocation, when a user program stores to a page it may only read.
  • How long a page lives compared with the page tables that map it, and what exec, which builds a whole new page table, does to a page that must outlive it.
  • Whether the kernel’s own copies to and from user memory obey the same PTE bits as the hardware does, and what that means for a page user code may read but should not hand to a system call.
  • When a value shared between harts can be read with no lock (one writer, one naturally aligned 32-bit word, a volatile read) and when it cannot (two values that must agree).
  • What a trap costs, measured two ways: instructions counted with gdb, and calls per tick timed from user mode, and why user code here cannot read the time CSR.

The reference branch

ext/10-vdso in ShowMeTheStack/xv6-riscv-labs, branched from the frozen commit 06aad25; 8 commits.

git clone https://github.com/ShowMeTheStack/xv6-riscv-labs
cd xv6-riscv-labs
git checkout -b my-vdso 06aad25   # start your own
git diff 06aad25 origin/ext/10-vdso   # only when you want the answer

1. The spec

Behaviour. Every process has two extra pages mapped at fixed addresses just below the trapframe, both readable and not writable from user mode:

address name contents physical page
0x3fffffd000 USYSCALL struct usyscall { int pid; } one per process
0x3fffffc000 USHARED struct ushared { uint ticks; } one for the whole system

The user library gains two functions that read them without a system call:

int ugetpid(void);   // == getpid()
int uuptime(void);   // == uptime(), give or take a tick in flight

What must not change. System calls behave as before for every address below p->sz. A system call handed a pointer into the new pages fails, as it did when nothing was mapped there (usertests lazy_copy checks read and write on exactly these addresses). No page leaks; the shared page is never freed. usertests -q must print ALL TESTS PASSED on 3 harts. (One test, lazy_sbrk, checks the old heap ceiling by its exact value; the reference branch updates that one expected value, see the reveal.)

The test program, vdsotest, prints one line per check:

$ vdsotest
vdsotest: pid: OK
vdsotest: uptime agrees: OK
vdsotest: uptime advances: OK
usertrap(): unexpected scause 0xf pid=4
            sepc=0x14 stval=0x3fffffd000
vdsotest: usyscall read-only: OK
usertrap(): unexpected scause 0xf pid=5
            sepc=0x2e stval=0x3fffffc000
vdsotest: ushared read-only: OK
vdsotest: read() into usyscall fails: OK
vdsotest: write() from ushared fails: OK
vdsotest: pid unchanged: OK
usertrap(): unexpected scause 0xf pid=6
            sepc=0xd2 stval=0x3fffffd000
vdsotest: heap ceiling: OK
vdsotest: fork: OK
vdsotest: exec: OK
vdsotest: 20 children: OK
vdsotest: ALL OK

vdsotest time measures: it reports whether user mode may execute rdtime, then how many getpid, ugetpid, uptime and uuptime calls fit into 20 ticks.

2. Think first

Answer each question in your head (or on paper) before opening a hint. Hints get more specific; the reference answer comes last.

1What does getpid cost, and which part of it is necessary?

Trace getpid() from the user stub to the return. List what the hardware and the kernel do on the way in and on the way out. Then ask: which of those steps is needed to produce the answer, and what would user code need in order to find the pid by itself, with no trap? Commit to an answer before reading the hints: could the kernel give the process its pid once, and keep it correct afterwards?

Check yourself

1warm-upType a number

In one getpid() round trip on this kernel, from ecall to the instruction after it, how many times is satp written?

kernel/trampoline.S
87 # wait for any previous memory operations to complete, so that
88 # they use the user page table.
89 sfence.vma zero, zero
91 # install the kernel page table.
92 csrw satp, t1
94 # flush now-stale user entries from the TLB.
95 sfence.vma zero, zero
97 # call usertrap()
98 jalr t0
100.globl userret
102 # usertrap() returns here, with user satp in a0.
103 # return from kernel to user.
105 # flush icache, in case this is the first time
106 # we're running this proc on this hart.
107 fence.i
109 # switch to the user page table.
110 sfence.vma zero, zero
111 csrw satp, a0
112 sfence.vma zero, zero
decimal, 0x hex or 0b binary
2solidChoose one

gdb counted 1,103 instructions from the first instruction of uservec to sret for one getpid. Which part of that work does ugetpid() keep?

2Where in the address space?

User code must find the page without asking the kernel, so its address must be fixed and known when the user library is compiled. Look at the user address space (Tour 25: A user address space): which addresses are already taken, and by what? Pick an address for the per-process page, and check it against everything else that can place memory in a user address space. What else in the kernel has to learn about your choice?

Check yourself

1warm-upType a number

MAXVA is 1L << 38, TRAMPOLINE is MAXVA - PGSIZE and TRAPFRAME is TRAMPOLINE - PGSIZE. What is USYSCALL = TRAPFRAME - PGSIZE, in hex?

kernel/memlayout.h
46// map the trampoline page to the highest address,
47// in both user and kernel space.
50// map kernel stacks beneath the trampoline,
51// each surrounded by invalid guard pages.
52#define KSTACK(p) (TRAMPOLINE - ((p) + 1) * 2 * PGSIZE)
54// User memory layout.
55// Address zero first:
56// text
57// original data and bss
58// fixed-size stack
59// expandable heap
60// ...
61// TRAPFRAME (p->trapframe, used by the trampoline)
62// TRAMPOLINE (the same page as in the kernel)
decimal, 0x hex or 0b binary
2solidChoose one

You map the two pages but leave the heap ceiling at TRAPFRAME in growproc and sys_sbrk. usertests lazy_sbrk grows the heap lazily to 0x3fffffb000, adds one page eagerly with sbrk(4096) so that p->sz is 0x3fffffc000, then calls sbrk(1). What happens?

kernel/proc.c
239 struct proc *p = myproc();
241 sz = p->sz;
242 if (n > 0) {
243 if (sz + n > TRAPFRAME) {
244 return -1;
245 }
246 if ((sz = uvmalloc(p->pagetable, sz, sz + n, PTE_W)) == 0) {
247 return -1;
248 }
249 } else if (n < 0) {
251 }
252 p->sz = sz;
253 return 0;

3Which bits in the PTE, and what does a store do?

Choose the permission bits for the two new PTEs. Then predict, step by step, what happens when a user program executes *(int *)USYSCALL = 1234;. Which scause? Which kernel function is asked to deal with it first, and what must it decide? Could it end up giving the process a fresh writable page at that address?

Check yourself

1solidDecode the bits

gdb read this PTE for USYSCALL in vdsotest’s page table, after the program had called ugetpid(). Decode the flag bits.

Value: 0x21fd4453

2solidChoose one

vdsotest’s child executes *(volatile int *)USYSCALL = 1234; with p->sz = 0x5000. Which statement describes what the kernel does?

kernel/trap.c
69 } else if ((which_dev = devintr()) != 0) {
70 // ok
71 } else if ((r_scause() == 15 || r_scause() == 13) &&
73 (r_scause() == 13) ? 1 : 0) != 0) {
74 // page fault on lazily-allocated page
75 } else {
76 printk("usertrap(): unexpected scause 0x%lx pid=%d\n", r_scause(), p->pid);
77 printk(" sepc=0x%lx stval=0x%lx\n", r_sepc(), r_stval());
79 }

4Who owns the page, and when does it live?

The per-process page needs a physical page and a mapping. Decide when the physical page is allocated, when the pid is written into it, when it is mapped, unmapped and freed. Walk the life of a process that the shell forks and that then execs a program: does your design give the exec’d program the page? What happens to the page table exec throws away?

Check yourself

1solidPut in order

The shell forks pid 3, which execs vdsotest, which exits and is reaped by the shell. Put these events for pid 3’s usyscall page in order.

  1. proc_freepagetable unmaps it from the old page table, without freeing it
  2. allocproc allocates the page and writes 3 into it
  3. proc_pagetable maps it into the child’s first page table
  4. freeproc, in the shell’s wait, frees the page and removes the last mapping
  5. kexec’s call to proc_pagetable maps the same page into the new image’s page table
2deepChoose one

You add uvmunmap(pagetable, USYSCALL, 1, 0) to proc_pagetable's error paths but forget it in proc_freepagetable. When does the kernel first notice?

kernel/vm.c
262// Recursively free page-table pages.
263// All leaf mappings must already have been removed.
264void
267 // there are 2^9 = 512 PTEs in a page table.
268 for (int i = 0; i < 512; i++) {
270 if ((pte & PTE_V) && (pte & (PTE_R | PTE_W | PTE_X)) == 0) {
271 // this PTE points to a lower-level page table.
275 } else if (pte & PTE_V) {
276 panic("freewalk: leaf");
277 }
278 }
279 kfree((void *)pagetable);
282// Free user memory pages,
283// then free page-table pages.
284void
287 if (sz > 0)

5What do fork and exec give the new image?

After your answer to the previous question, check two cases by reading code, not by assumption. A forked child must see its pid, not its parent’s: is there any way the parent’s page, or its contents, could end up in the child? And after exec, could the new image see a stale value?

Check yourself

1solidTrue or false, and why

True or false: in the reference design, uvmcopy copies the parent’s usyscall page into the child, and kfork then overwrites the pid in the copy.

Why?

6The kernel reads and writes user memory too

Your pages are mapped and read-only. Now a program passes their address to a system call: read(fd, (char *)USYSCALL, 8) asks the kernel to write the page, and write(fd, (char *)USHARED, 8) asks it to read it. The kernel does not use the user’s PTE the way the hardware does. What happens to each call? Is either result a problem? Run usertests -q at this point if you have built that far.

Check yourself

1solidMatch the pairs

Before the copyin fix, match each access to the new pages with what happens.

7One page for everyone, read with no lock

uptime() returns ticks, which clockintr increments on hart 0 under tickslock, and sys_uptime reads under the same lock. For uuptime() the kernel copies ticks into one page shared by all processes, and user code reads it with no lock at all (it could not take a kernel spinlock anyway). Where does the page come from, and who may free it? Then the hard part: on three harts, can a reader see a value that was never written? Can the compiler or the hardware make the read return something stale forever?

Check yourself

1solidChoose all that apply

uuptime() reads the shared page with no lock while clockintr on hart 0 writes it. Which statements are true for this tree on RISC-V?

2warm-upChoose one

proc_freepagetable unmaps the shared page. Which do_free argument is right, and why?

8How do you measure the difference?

You want numbers: how long one getpid() takes versus one ugetpid(). The best clock is the time CSR, which clockintr reads with rdtime. Can a user program execute rdtime in this tree? If not, what clock can you use from user mode, and how many calls must you make for the answer to be meaningful?

Check yourself

1solidChoose one

A user program on this kernel executes rdtime a5 (as vdsotest time’s probe does). What happens?

kernel/start.c
54// ask each hart to generate timer interrupts.
55void
58 // enable the sstc extension (i.e. stimecmp).
61 // allow supervisor to use stimecmp and time.
64 // ask for the very first timer interrupt.
65 w_stimecmp(r_time() + 1000000);

3. Build it

Start.

git checkout -b my-vdso 06aad25

Write the test first: copy the spec’s checks into user/vdsotest.c and add $U/_vdsotest\ to UPROGS in the Makefile. Until the library functions exist it will not link; that is fine, it tells you what to build.

Milestones, in an order that keeps the system bootable after each one.

  1. The layout. USYSCALL, USHARED, the two structs and the heap’s new ceiling in kernel/memlayout.h; use the ceiling in growproc and sys_sbrk; update the one expected value in usertests lazy_sbrk. Guard the structs with #ifndef __ASSEMBLER__: trampoline.S includes memlayout.h. Test: usertests -q.
  2. The physical page. A struct usyscall *usyscall field in struct proc; allocate, zero and fill it in allocproc, free it in freeproc. Test: usertests -q (the page exists but nothing maps it; a leak would show in the free-page checks).
  3. The mapping. Map it read-only in the function that builds user page tables, unmap it in the one that frees them, and undo it in the error path. Test: boot. If the kernel panics before the shell prompt, read clinic 3. From here until milestone 5, usertests -q fails at lazy_copy (write succeeded); that is expected.
  4. The shared page. Allocate it in trapinit, update it in clockintr, map and unmap it like the per-process page (but never free it).
  5. The kernel’s copies. Run usertests -q: lazy_copy fails with write succeeded. Make copyin and copyinstr refuse addresses at or above psz. Test: usertests -q.
  6. The library and the test. ugetpid and uuptime in user/ulib.c, prototypes in user/user.h. Test: vdsotest, then usertests -q, then vdsotest again.

Debugging advice. To catch anything during boot (clinic 3 panics before the shell starts), start QEMU halted with make qemu-gdb (it adds -S and a gdb port of its own, which it writes into .gdbinit), run ${TOOLPREFIX}gdb kernel/kernel in another terminal, and set breakpoints before the first continue. TOOLPREFIX is your RISC-V toolchain’s prefix, the same one xv6’s Makefile detects (riscv64-unknown-elf-, riscv64-linux-gnu- or riscv64-elf-); set it with export TOOLPREFIX=riscv64-unknown-elf- or whichever you have. On Debian/Ubuntu/WSL, gdb-multiarch also works as the debugger.

4. Debugging clinic

Each of these bugs was put into the reference solution on purpose and run on three harts. The symptom is exactly what happened. Try to explain it before revealing why.

1The PTE has no PTE_U

The usyscall mapping copies the trapframe’s mapping and drops PTE_W, leaving PTE_R only:

   if (mappages(pagetable, USYSCALL, PGSIZE, (uint64)(p->usyscall),
-               PTE_R | PTE_U) < 0) {
+               PTE_R) < 0) {

What happened when we ran it

$ vdsotest
usertrap(): unexpected scause 0xd pid=3
            sepc=0x980 stval=0x3fffffd000
$ usertests -q
usertests starting
[...]
ALL TESTS PASSED

# gdb, breakpoint on setkilled (rerun):
setkilled: {'hart': 0, 'noff': 0, 'intena': 0, 'sie': 0, 'pid': 3, 'name': 'vdsotest'} scause 0xd stval 0x3fffffd000 sepc 0x980
  frames ['setkilled', 'usertrap', '0x3ffffff09c']
  p->usyscall 0x87f51000 pid in page 3
  pagetable 0x87f21000
  PTE for 0x3ffffff000: 0x2000184b
  PTE for 0x3fffffe000: 0x21fcfcc7
  PTE for 0x3fffffd000: 0x21fd4403
  PTE for 0x3fffffc000: 0x21fd6413

2PTE_W is left on

Both pages are mapped like the trapframe, with PTE_W, plus PTE_U:

   if (mappages(pagetable, USYSCALL, PGSIZE, (uint64)(p->usyscall),
-               PTE_R | PTE_U) < 0) {
+               PTE_R | PTE_W | PTE_U) < 0) {
[...]
   if (mappages(pagetable, USHARED, PGSIZE, (uint64)ushared,
-               PTE_R | PTE_U) < 0) {
+               PTE_R | PTE_W | PTE_U) < 0) {

To see the effect we added a small program, forge (not part of the branch), that stores 1 at USYSCALL and 1000000 at USHARED.

What happened when we ran it

$ forge
forge: getpid() 3, ugetpid() 3
forge: after the store, ugetpid() 1
forge: uptime() 1, uuptime() 1000000
$ vdsotest
vdsotest: pid: OK
vdsotest: uptime agrees: OK
vdsotest: uptime advances: OK
vdsotest: usyscall read-only: FAIL
vdsotest: ushared read-only: FAIL
vdsotest: read() into usyscall fails: FAIL
vdsotest: write() from ushared fails: OK
vdsotest: pid unchanged: FAIL
vdsotest: heap ceiling: FAIL
vdsotest: fork: FAIL
vdsotest: exec: OK
vdsotest: 20 children: OK
vdsotest: SOME TESTS FAILED
$ usertests -q
usertests starting
[...]
test lazy_copy: read succeeded
FAILED
SOME TESTS FAILED

3The page is not unmapped in proc_freepagetable

The page is mapped in proc_pagetable, but its mirror forgets it:

 proc_freepagetable(pagetable_t pagetable, uint64 sz)
 {
   uvmunmap(pagetable, TRAMPOLINE, 1, 0);
   uvmunmap(pagetable, TRAPFRAME, 1, 0);
-  uvmunmap(pagetable, USYSCALL, 1, 0);
   uvmunmap(pagetable, USHARED, 1, 0); // shared: never free it
   uvmfree(pagetable, sz);
 }

What happened when we ran it

xv6 kernel is booting

hart 2 starting
hart 1 starting
panic: freewalk: leaf

# gdb, breakpoint on panic (rerun):
#0  panic (s=s@entry=0x80007138 "freewalk: leaf") at kernel/printk.c:139
#1  0x000000008000137c in freewalk (pagetable=0x87f51000) at kernel/vm.c:276
#2  0x000000008000139a in freewalk (pagetable=0x87f52000) at kernel/vm.c:273
#3  0x000000008000139a in freewalk (pagetable=pagetable@entry=0x87f53000) at kernel/vm.c:273
#4  0x00000000800013c8 in uvmfree (pagetable=pagetable@entry=0x87f53000, sz=sz@entry=0) at kernel/vm.c:289
#5  0x0000000080001b98 in proc_freepagetable (pagetable=0x87f53000, sz=sz@entry=0) at kernel/proc.c:248
#6  0x0000000080004c1a in kexec (path=path@entry=0x80007180 "/init", argv=argv@entry=0x3fffffdfd0) at kernel/exec.c:138
#7  0x0000000080001990 in forkret () at kernel/proc.c:566
[...]
  frame freewalk {'i': '0x1fd', 'pagetable': '0x87f51000', '$s0': '0x3fffffdd10', '$s1': '0x87f51fe8', '$s2': '0x87f52000', '$s3': '0x87f51000', '$a0': '0x80007138'}
  frame freewalk {'i': '0x1ff', 'pagetable': '0x87f52000', '$s0': '0x3fffffdd40', '$s1': '0x87f52ff8', '$s2': '0x87f53000', '$s3': '0x87f52000', '$a0': '0x80007138'}
  frame freewalk {'i': '0xff', 'pagetable': '0x87f53000', '$s0': '0x3fffffdd70', '$s1': '0x87f537f8', '$s2': '0x87f54000', '$s3': '0x87f53000', '$a0': '0x80007138'}

4The page is freed twice

proc_freepagetable unmaps the page with do_free = 1, while freeproc still frees p->usyscall itself:

   uvmunmap(pagetable, TRAPFRAME, 1, 0);
-  uvmunmap(pagetable, USYSCALL, 1, 0);
+  uvmunmap(pagetable, USYSCALL, 1, 1);

What happened when we ran it

$ vdsotest
vdsotest: pid: FAIL
vdsotest: uptime agrees: OK
vdsotest: uptime advances: OK
usertrap(): unexpected scause 0xf pid=4
            sepc=0x14 stval=0x3fffffd000
vdsotest: usyscall read-only: OK
usertrap(): unexpected scause 0xf pid=5
            sepc=0x2e stval=0x3fffffc000
vdsotest: ushared read-only: OK
vdsotest: read() into usyscall fails: OK
vdsotest: write() from ushared fails: OK
vdsotest: pid unchanged: FAIL
scause=0xd sepc=0x80000b2e stval=0x800e022e4061141
panic: kerneltrap

# gdb (rerun, same console output), breakpoints on the unmap in proc_freepagetable,
# on freeproc's kfree, on sys_getpid and on panic:
[...]
allocproc: {'hart': 0, 'noff': 1, 'intena': 0, 'sie': 0} pid 1 p->usyscall 0x87f54000
unmap-free USYSCALL: {'hart': 0, 'noff': 0, 'intena': 0, 'sie': 0, 'pid': 1, 'name': 'init'} pagetable 0x87f53000 PTE 0x21fd5013 frames ['proc_freepagetable', 'kexec', 'forkret', 'myproc']
allocproc: {'hart': 0, 'noff': 1, 'intena': 1, 'sie': 0, 'pid': 1, 'name': 'init'} pid 2 p->usyscall 0x87f52000
unmap-free USYSCALL: {'hart': 0, 'noff': 0, 'intena': 1, 'sie': 1, 'pid': 2, 'name': 'sh'} pagetable 0x87f51000 PTE 0x21fd4813 frames ['proc_freepagetable', 'kexec', 'sys_exec', 'syscall', 'usertrap', '0x3ffffff09c']
allocproc: {'hart': 1, 'noff': 1, 'intena': 1, 'sie': 0, 'pid': 2, 'name': 'sh'} pid 3 p->usyscall 0x87f51000
unmap-free USYSCALL: {'hart': 1, 'noff': 0, 'intena': 1, 'sie': 1, 'pid': 3, 'name': 'vdsotest'} pagetable 0x87f54000 PTE 0x21fd4413 frames ['proc_freepagetable', 'kexec', 'sys_exec', 'syscall', 'usertrap', '0x3ffffff09c']
sys_getpid: {'hart': 1, 'noff': 0, 'intena': 1, 'sie': 1, 'pid': 3, 'name': 'vdsotest'} p->usyscall 0x87f51000 first 8 bytes 0x87f19000 next 8 0x101010101010101
allocproc: {'hart': 0, 'noff': 1, 'intena': 1, 'sie': 0, 'pid': 3, 'name': 'vdsotest'} pid 4 p->usyscall 0x87f54000
freeproc kfree: {'hart': 0, 'noff': 2, 'intena': 1, 'sie': 0, 'pid': 3, 'name': 'vdsotest'} p->pid 4 p->usyscall 0x87f54000 frames ['freeproc', 'kwait', 'sys_wait', 'syscall', 'usertrap', '0x3ffffff09c']
unmap-free USYSCALL: {'hart': 0, 'noff': 2, 'intena': 1, 'sie': 0, 'pid': 3, 'name': 'vdsotest'} pagetable 0x87f47000 PTE 0x21fd5013 frames ['proc_freepagetable', 'freeproc', 'kwait', 'sys_wait', 'syscall', 'usertrap', '0x3ffffff09c']
[...]
panic: {'hart': 1, 'noff': 2, 'intena': 1, 'sie': 0, 'pid': 3, 'name': 'vdsotest'} frames ['panic', 'kerneltrap', 'kernelvec']
#0  panic (s=s@entry=0x80007390 "kerneltrap") at kernel/printk.c:139
#1  0x000000008000289c in kerneltrap () at kernel/trap.c:159
#2  0x0000000080005728 in kernelvec () at kernel/kernelvec.S:38
Backtrace stopped: frame did not save the PC
interrupted: sepc 0x80000b2e kalloc + 32 in section .text s1 0x800e022e4061141 fp 0x3fffff9e90
  saved ra 0x80000b24 kalloc + 22 in section .text
  ra 0x80000fb8 walk + 122 in section .text
  ra 0x8000105a mappages + 72 in section .text
  ra 0x8000144a uvmcopy + 100 in section .text
  ra 0x80001da6 kfork + 44 in section .text
  ra 0x80002a9e sys_fork + 12 in section .text
  ra 0x80002a2e syscall + 58 in section .text
  ra 0x800027b4 usertrap + 166 in section .text
  ra 0x3ffffff09c No symbol matches 274877903004.
kmem.freelist 0x800e022e4061141

5The mapping is made outside the page-table builder, so exec loses it

The page is mapped in allocproc, right after proc_pagetable returns, instead of inside it (the unmap in proc_freepagetable stays):

   p->pagetable = proc_pagetable(p);
   if (p->pagetable == 0) {
     freeproc(p);
     release(&p->lock);
     return 0;
   }
+
+  // map the usyscall page, read-only for user code.
+  if (mappages(p->pagetable, USYSCALL, PGSIZE, (uint64)(p->usyscall),
+               PTE_R | PTE_U) < 0) {
+    freeproc(p);
+    release(&p->lock);
+    return 0;
+  }

What happened when we ran it

$ vdsotest
usertrap(): unexpected scause 0xd pid=3
            sepc=0x980 stval=0x3fffffd000
$ usertests -q
usertests starting
[...]
ALL TESTS PASSED

# gdb, breakpoint on setkilled (rerun):
setkilled: {'hart': 2, 'noff': 0, 'intena': 0, 'sie': 0, 'pid': 3, 'name': 'vdsotest'} scause 0xd stval 0x3fffffd000 sepc 0x980
  frames ['setkilled', 'usertrap', '0x3ffffff09c']
  p->usyscall 0x87f51000 pid in page 3
  pagetable 0x87f21000
  PTE for 0x3ffffff000: 0x2000184b
  PTE for 0x3fffffe000: 0x21fcfcc7
  PTE for 0x3fffffd000: no PTE (level-0 entry 509 is 0x0)
  PTE for 0x3fffffc000: 0x21fd6413

6The heap’s ceiling stays at TRAPFRAME

Everything else as in the reference, but the old limit is kept in both places:

-    if (sz + n > MAXHEAP) {
+    if (sz + n > TRAPFRAME) {
[...]
-    if (addr + n > MAXHEAP)
+    if (addr + n > TRAPFRAME)

What happened when we ran it

# one boot:
$ vdsotest
[...]
vdsotest: pid unchanged: OK
vdsotest: heap ceiling: FAIL
vdsotest: fork: OK
vdsotest: exec: OK
vdsotest: 20 children: OK
vdsotest: SOME TESTS FAILED
$ usertests lazy_sbrk
usertests starting
test lazy_sbrk: panic: mappages: remap

# a second boot, gdb attached, breakpoint on panic, `usertests lazy_sbrk`:
#0  panic (s=s@entry=0x80007108 "mappages: remap") at kernel/printk.c:139
#1  0x00000000800010ac in mappages (pagetable=pagetable@entry=0x80023000, va=va@entry=274877890560, size=size@entry=4096, pa=<optimized out>, pa@entry=2147725312, perm=perm@entry=22) at kernel/vm.c:167
#2  0x0000000080001302 in uvmalloc (pagetable=0x80023000, oldsz=274877890560, newsz=274877890561, xperm=xperm@entry=4) at kernel/vm.c:234
#3  0x0000000080001d4a in growproc (n=1) at kernel/proc.c:281
#4  0x0000000080002b28 in sys_sbrk () at kernel/sysproc.c:51
#5  0x0000000080002a2e in syscall () at kernel/syscall.c:146
#6  0x00000000800027b4 in usertrap () at kernel/trap.c:74
[...]

# two more boots with a probe program (ceilprobe, not on the branch) whose children
# grow p->sz lazily to 0x3fffffe000, over both new pages:
$ ceilprobe store
ceilprobe: sz 0x0000003FFFFFE000, storing to USYSCALL
usertrap(): unexpected scause 0xf pid=4
            sepc=0x9e stval=0x3fffffd000
ceilprobe: store child status -1
ceilprobe: sz 0x0000003FFFFFE000, write(USHARED) returns 8
$
[...]
$ ceilprobe shrink
ceilprobe: sz 0x0000003FFFFFE000, sbrk(-8192)
usertrap(): unexpected scause 0xd pid=4
            sepc=0x490 stval=0x3fffffc000
ceilprobe: shrink child status -1, uuptime 3
$ ceilprobe store
usertrap(): unexpected scause 0xf pid=5
            sepc=0xa70 stval=0x0
$ vdsotest
usertrap(): unexpected scause 0xf pid=6
            sepc=0xa70 stval=0x0

5. The reference solution

Take the guided tour through the reference solution, one commit at a time, with the machine state at every step:

Open the reveal tour →

Or read the commits

  1. 086102f Reserve two user pages below the trapframe

    kernel/memlayout.h

    @@ -57,7 +57,26 @@
    5757// original data and bss
    5858// fixed-size stack
    5959// expandable heap
    6060// ...
    61// USHARED (one page shared by all processes, read-only)
    62// USYSCALL (this process's struct usyscall, read-only)
    6163// TRAPFRAME (p->trapframe, used by the trampoline)
    6264// TRAMPOLINE (the same page as in the kernel)
    6365#define TRAPFRAME (TRAMPOLINE - PGSIZE)
    66#define USYSCALL (TRAPFRAME - PGSIZE)
    67#define USHARED (USYSCALL - PGSIZE)
    68
    69// the heap may grow up to, but not into, these pages.
    70#define MAXHEAP USHARED
    71
    72#ifndef __ASSEMBLER__
    73// what user code can read at USYSCALL without a system call.
    74struct usyscall {
    75 int pid; // this process's pid
    76};
    77
    78// what every process can read at USHARED.
    79struct ushared {
    80 uint ticks; // a copy of ticks, updated by clockintr()
    81};
    82#endif

    kernel/proc.c

    @@ -239,9 +239,9 @@ growproc(int n)
    239239 struct proc *p = myproc();
    240240
    241241 sz = p->sz;
    242242 if (n > 0) {
    243 if (sz + n > TRAPFRAME) {
    243 if (sz + n > MAXHEAP) {
    244244 return -1;
    245245 }
    246246 if ((sz = uvmalloc(p->pagetable, sz, sz + n, PTE_W)) == 0) {
    247247 return -1;

    kernel/sysproc.c

    @@ -56,9 +56,9 @@ sys_sbrk(void)
    5656 // size but don't allocate memory. If the processes uses the
    5757 // memory, vmfault() will allocate it.
    5858 if (addr + n < addr)
    5959 return -1;
    60 if (addr + n > TRAPFRAME)
    60 if (addr + n > MAXHEAP)
    6161 return -1;
    6262 myproc()->sz += n;
    6363 }
    6464 return addr;

    user/usertests.c

    @@ -2789,19 +2789,19 @@ lazy_sbrk(char *s)
    27892789
    27902790 p = sbrklazy(0);
    27912791 }
    27922792
    2793 int n = TRAPFRAME - PGSIZE - (uint64)p;
    2793 int n = MAXHEAP - PGSIZE - (uint64)p;
    27942794
    27952795 char *p1 = sbrklazy(n);
    27962796 if (p1 < 0 || p1 != p) {
    27972797 printf("sbrklazy(%d) returned %p, not expected %p\n", n, p1, p);
    27982798 exit(1);
    27992799 }
    28002800
    28012801 p = sbrk(PGSIZE);
    2802 if (p < 0 || (uint64)p != TRAPFRAME - PGSIZE) {
    2803 printf("sbrk(%d) returned %p, not expected TRAPFRAME-PGSIZE\n", PGSIZE, p);
    2802 if (p < 0 || (uint64)p != MAXHEAP - PGSIZE) {
    2803 printf("sbrk(%d) returned %p, not expected MAXHEAP-PGSIZE\n", PGSIZE, p);
    28042804 exit(1);
    28052805 }
    28062806
    28072807 p[0] = 1;
  2. e635a6e Give each process a usyscall page

    kernel/proc.c

    @@ -131,8 +131,17 @@ found:
    131131 release(&p->lock);
    132132 return 0;
    133133 }
    134134
    135 // Allocate the page that user code reads at USYSCALL.
    136 if ((p->usyscall = (struct usyscall *)kalloc()) == 0) {
    137 freeproc(p);
    138 release(&p->lock);
    139 return 0;
    140 }
    141 memset(p->usyscall, 0, PGSIZE);
    142 p->usyscall->pid = p->pid;
    143
    135144 // An empty user page table.
    136145 p->pagetable = proc_pagetable(p);
    137146 if (p->pagetable == 0) {
    138147 freeproc(p);
    @@ -157,8 +166,11 @@ freeproc(struct proc *p)
    157166{
    158167 if (p->trapframe)
    159168 kfree((void *)p->trapframe);
    160169 p->trapframe = 0;
    170 if (p->usyscall)
    171 kfree((void *)p->usyscall);
    172 p->usyscall = 0;
    161173 if (p->pagetable)
    162174 proc_freepagetable(p->pagetable, p->sz);
    163175 p->pagetable = 0;
    164176 p->sz = 0;

    kernel/proc.h

    @@ -96,8 +96,9 @@ struct proc {
    9696 uint64 kstack; // Virtual address of kernel stack
    9797 uint64 sz; // Size of process memory (bytes)
    9898 pagetable_t pagetable; // User page table
    9999 struct trapframe *trapframe; // data page for trampoline.S
    100 struct usyscall *usyscall; // read-only page for user code
    100101 struct context context; // swtch() here to run process
    101102 struct file *ofile[NOFILE]; // Open files
    102103 struct inode *cwd; // Current directory
    103104 char name[16]; // Process name (debugging)
  3. 403eca4 Map the usyscall page read-only into user space

    kernel/proc.c

    @@ -212,8 +212,18 @@ proc_pagetable(struct proc *p)
    212212 uvmfree(pagetable, 0);
    213213 return 0;
    214214 }
    215215
    216 // map this process's usyscall page below the trapframe.
    217 // user code may read it (PTE_U) but not write it (no PTE_W).
    218 if (mappages(pagetable, USYSCALL, PGSIZE, (uint64)(p->usyscall),
    219 PTE_R | PTE_U) < 0) {
    223 return 0;
    224 }
    225
    216226 return pagetable;
    217227}
    218228
    219229// Free a process's page table, and free the
    @@ -222,8 +232,9 @@ void
    222232proc_freepagetable(pagetable_t pagetable, uint64 sz)
    223233{
    224234 uvmunmap(pagetable, TRAMPOLINE, 1, 0);
    225235 uvmunmap(pagetable, TRAPFRAME, 1, 0);
    236 uvmunmap(pagetable, USYSCALL, 1, 0);
    226237 uvmfree(pagetable, sz);
    227238}
    228239
    229240// Set up first user process.
  4. 26cb112 Share one ticks page with every process

    kernel/defs.h

    @@ -8,8 +8,9 @@ struct proc;
    88struct spinlock;
    99struct sleeplock;
    1010struct stat;
    1111struct superblock;
    12struct ushared;
    1213
    1314// bio.c
    1415void binit(void);
    1516struct buf* bread(uint, uint);
    @@ -142,8 +143,9 @@ void syscall();
    142143extern uint ticks;
    143144void trapinit(void);
    144145void trapinithart(void);
    145146extern struct spinlock tickslock;
    147extern struct ushared *ushared;
    146148void prepare_return(void);
    147149
    148150// uart.c
    149151void uartinit(void);

    kernel/proc.c

    @@ -222,8 +222,19 @@ proc_pagetable(struct proc *p)
    222222 uvmfree(pagetable, 0);
    223223 return 0;
    224224 }
    225225
    226 // map the page that all processes share below it, also
    227 // read-only. it is the same physical page in every page table.
    228 if (mappages(pagetable, USHARED, PGSIZE, (uint64)ushared,
    229 PTE_R | PTE_U) < 0) {
    230 uvmunmap(pagetable, USYSCALL, 1, 0);
    234 return 0;
    235 }
    236
    226237 return pagetable;
    227238}
    228239
    229240// Free a process's page table, and free the
    @@ -233,8 +244,9 @@ proc_freepagetable(pagetable_t pagetable, uint64 sz)
    233244{
    234245 uvmunmap(pagetable, TRAMPOLINE, 1, 0);
    235246 uvmunmap(pagetable, TRAPFRAME, 1, 0);
    236247 uvmunmap(pagetable, USYSCALL, 1, 0);
    248 uvmunmap(pagetable, USHARED, 1, 0); // shared: never free it
    237249 uvmfree(pagetable, sz);
    238250}
    239251
    240252// Set up first user process.

    kernel/trap.c

    @@ -7,8 +7,9 @@
    77#include "defs.h"
    88
    1010uint ticks;
    11struct ushared *ushared; // mapped at USHARED in every process
    1112
    1213extern char trampoline[], uservec[];
    1314
    1415// in kernelvec.S, calls kerneltrap().
    @@ -19,8 +20,13 @@ extern int devintr();
    1920void
    2021trapinit(void)
    2122{
    2223 initlock(&tickslock, "time");
    24
    25 // one page, shared read-only by every process, never freed.
    26 if ((ushared = (struct ushared *)kalloc()) == 0)
    27 panic("trapinit: ushared");
    28 memset(ushared, 0, PGSIZE);
    2329}
    2430
    2531// set up to take exceptions and traps while in the kernel.
    2632void
    @@ -168,8 +174,9 @@ clockintr()
    168174{
    169175 if (cpuid() == 0) {
    170176 acquire(&tickslock);
    171177 ticks++;
    178 ushared->ticks = ticks; // user code reads it with no lock
    172179 wakeup(&ticks);
    173180 release(&tickslock);
    174181 }
    175182
  5. 358b5d6 Keep system calls out of the read-only pages

    kernel/vm.c

    @@ -385,8 +385,12 @@ copyin(pagetable_t pagetable, uint64 psz, char *dst, uint64 srcva, uint64 len)
    385385 uint64 n, va0, pa0;
    386386
    387387 while (len > 0) {
    388388 va0 = PGROUNDDOWN(srcva);
    389 // system calls read only the process's memory, below
    390 // psz, not the read-only pages above it.
    391 if (va0 >= psz)
    392 return -1;
    389393 pa0 = walkaddr(pagetable, va0);
    390394 if (pa0 == 0) {
    391395 if ((pa0 = vmfault(pagetable, psz, va0, 1)) == 0) {
    392396 return -1;
    @@ -416,8 +420,10 @@ copyinstr(pagetable_t pagetable, uint64 psz, char *dst, uint64 srcva,
    416420 int got_null = 0;
    417421
    418422 while (got_null == 0 && max > 0) {
    419423 va0 = PGROUNDDOWN(srcva);
    424 if (va0 >= psz) // as in copyin()
    425 return -1;
    420426 pa0 = walkaddr(pagetable, va0);
    421427 if (pa0 == 0) {
    422428 if ((pa0 = vmfault(pagetable, psz, va0, 1)) == 0) {
    423429 return -1;
  6. c07e83a Add ugetpid and uuptime to the user library

    user/ulib.c

    @@ -1,8 +1,9 @@
    11#include "kernel/types.h"
    22#include "kernel/stat.h"
    33#include "kernel/fcntl.h"
    44#include "kernel/riscv.h"
    5#include "kernel/memlayout.h"
    56#include "kernel/vm.h"
    67#include "user/user.h"
    78
    89//
    @@ -159,4 +160,22 @@ char *
    159160sbrklazy(int n)
    160161{
    161162 return sys_sbrk(n, SBRK_LAZY);
    162163}
    164
    165// getpid() without a system call: read the pid that
    166// allocproc() wrote into this process's usyscall page.
    167int
    168ugetpid(void)
    169{
    170 struct usyscall *u = (struct usyscall *)USYSCALL;
    171 return u->pid;
    172}
    173
    174// uptime() without a system call. volatile: the kernel
    175// changes the value behind the compiler's back.
    176int
    177uuptime(void)
    178{
    179 volatile struct ushared *u = (struct ushared *)USHARED;
    180 return u->ticks;
    181}

    user/user.h

    @@ -39,8 +39,10 @@ int atoi(const char *);
    3939int memcmp(const void *, const void *, uint);
    4040void *memcpy(void *, const void *, uint);
    4141char *sbrk(int);
    4242char *sbrklazy(int);
    43int ugetpid(void);
    44int uuptime(void);
    4345
    4446// printf.c
    4547void fprintf(int, const char *, ...) __attribute__((format(printf, 2, 3)));
    4648void printf(const char *, ...) __attribute__((format(printf, 1, 2)));
  7. 1cad61d Add vdsotest

    Makefile

    @@ -149,8 +149,9 @@ UPROGS=\
    149149 $U/_logstress\
    150150 $U/_forphan\
    151151 $U/_dorphan\
    152152 $U/_sync\
    153 $U/_vdsotest\
    153154
    154155fs.img: mkfs/mkfs README $(UPROGS)
    155156 mkfs/mkfs fs.img README $(UPROGS)
    156157

    user/vdsotest.c

    @@ -0,0 +1,217 @@
    1// test the read-only pages at USYSCALL and USHARED:
    2// ugetpid() and uuptime() must agree with the system calls,
    3// the pages must be read-only, and fork and exec must give
    4// each process the right pid.
    5
    6#include "kernel/types.h"
    7#include "kernel/riscv.h"
    8#include "kernel/memlayout.h"
    9#include "kernel/fcntl.h"
    10#include "user/user.h"
    11
    12int failed = 0;
    13
    14void
    15check(char *name, int ok)
    16{
    17 printf("vdsotest: %s: %s\n", name, ok ? "OK" : "FAIL");
    18 if (!ok)
    19 failed = 1;
    20}
    21
    22// run f in a child; return the child's exit status.
    23int
    24inchild(void (*f)(void))
    25{
    26 int pid, xstatus;
    27
    28 pid = fork();
    29 if (pid < 0) {
    30 printf("vdsotest: fork failed\n");
    31 exit(1);
    32 }
    33 if (pid == 0) {
    34 f();
    35 exit(0); // f survived
    36 }
    37 if (wait(&xstatus) != pid)
    38 return 99;
    39 return xstatus;
    40}
    41
    42void
    43pidtest(void)
    44{
    45 check("pid", ugetpid() == getpid());
    46}
    47
    48void
    49uptimetest(void)
    50{
    51 int a, u, b, t0;
    52
    53 a = uptime();
    54 u = uuptime();
    55 b = uptime();
    56 check("uptime agrees", a <= u && u <= b);
    57
    58 // the copy must move on its own: wait for 2 ticks.
    59 t0 = uuptime();
    60 pause(2);
    61 check("uptime advances", uuptime() >= t0 + 2);
    62}
    63
    64void
    65store_usyscall(void)
    66{
    67 *(volatile int *)USYSCALL = 1234;
    68}
    69
    70void
    71store_ushared(void)
    72{
    73 *(volatile uint *)USHARED = 0;
    74}
    75
    76void
    77readonlytest(void)
    78{
    79 // a store from user mode must kill the child (status -1),
    80 // not succeed (status 0).
    81 check("usyscall read-only", inchild(store_usyscall) == -1);
    82 check("ushared read-only", inchild(store_ushared) == -1);
    83}
    84
    85void
    86kerneltest(void)
    87{
    88 int fd;
    89
    90 // system calls may not use the pages either: read() would
    91 // write them (copyout() needs PTE_W), and write() would read
    92 // them (copyin() stops at p->sz).
    93 fd = open("README", 0);
    94 if (fd < 0) {
    95 printf("vdsotest: cannot open README\n");
    96 exit(1);
    97 }
    98 check("read() into usyscall fails", read(fd, (char *)USYSCALL, 8) < 0);
    99 close(fd);
    100 fd = open("vdsojunk", O_CREATE | O_WRONLY);
    101 if (fd < 0) {
    102 printf("vdsotest: cannot create vdsojunk\n");
    103 exit(1);
    104 }
    105 check("write() from ushared fails", write(fd, (char *)USHARED, 8) < 0);
    106 close(fd);
    107 unlink("vdsojunk");
    108 check("pid unchanged", ugetpid() == getpid());
    109}
    110
    111void
    112ceilingtest(void)
    113{
    114 char *p;
    115
    116 // grow the heap lazily as far as it may go: up to MAXHEAP.
    117 p = sbrklazy(0);
    118 while ((uint64)p + (1 << 30) <= MAXHEAP)
    119 p = sbrklazy(1 << 30) + (1 << 30);
    120 if (sbrklazy(MAXHEAP - (uint64)p) == (char *)-1)
    121 exit(1);
    122 if ((uint64)sbrklazy(0) != MAXHEAP)
    123 exit(2);
    124 // one more byte would reach USHARED.
    125 if (sbrklazy(1) != (char *)-1)
    126 exit(3);
    127 if (ugetpid() != getpid())
    128 exit(4);
    129 // the pages are above p->sz: vmfault() must not "fix"
    130 // this store by allocating a page.
    131 store_usyscall();
    132}
    133
    134void
    135forktest(void)
    136{
    137 int parent = getpid(), pid, xstatus;
    138
    139 pid = fork();
    140 if (pid < 0) {
    141 printf("vdsotest: fork failed\n");
    142 exit(1);
    143 }
    144 if (pid == 0) {
    145 exit(ugetpid() == getpid() && ugetpid() != parent ? 0 : 1);
    146 }
    147 wait(&xstatus);
    148 check("fork", xstatus == 0 && ugetpid() == parent);
    149}
    150
    151void
    152exectest(void)
    153{
    154 int pid, xstatus;
    155 char *argv[] = {"vdsotest", "exec", 0};
    156
    157 pid = fork();
    158 if (pid < 0) {
    159 printf("vdsotest: fork failed\n");
    160 exit(1);
    161 }
    162 if (pid == 0) {
    163 exec("vdsotest", argv);
    164 exit(2);
    165 }
    166 wait(&xstatus);
    167 check("exec", xstatus == 0);
    168}
    169
    170void
    171manytest(void)
    172{
    173 int i, j, pid, xstatus, ok = 1;
    174
    175 // 20 children at once, on all harts, each reading its own page.
    176 for (i = 0; i < 20; i++) {
    177 pid = fork();
    178 if (pid < 0) {
    179 printf("vdsotest: fork failed\n");
    180 exit(1);
    181 }
    182 if (pid == 0) {
    183 pid = getpid();
    184 for (j = 0; j < 100000; j++)
    185 if (ugetpid() != pid)
    186 exit(1);
    187 exit(0);
    188 }
    189 }
    190 for (i = 0; i < 20; i++) {
    191 wait(&xstatus);
    192 if (xstatus != 0)
    193 ok = 0;
    194 }
    195 check("20 children", ok);
    196}
    197
    198int
    199main(int argc, char *argv[])
    200{
    201 if (argc == 2 && strcmp(argv[1], "exec") == 0) {
    202 // after exec: the new page table maps the same pid.
    203 exit(ugetpid() == getpid() ? 0 : 1);
    204 }
    205
    206 pidtest();
    207 uptimetest();
    208 readonlytest();
    209 kerneltest();
    210 check("heap ceiling", inchild(ceilingtest) == -1);
    211 forktest();
    212 exectest();
    213 manytest();
    214
    215 printf("vdsotest: %s\n", failed ? "SOME TESTS FAILED" : "ALL OK");
    216 exit(failed);
    217}
  8. 3ae5e66 Measure in vdsotest what the trap costs

    user/vdsotest.c

    @@ -194,15 +194,63 @@ manytest(void)
    194194 }
    195195 check("20 children", ok);
    196196}
    197197
    198// try to read the time CSR from user mode.
    199void
    200userrdtime(void)
    201{
    202 uint64 x;
    203 asm volatile("rdtime %0" : "=r"(x));
    204}
    205
    206// call f for 20 ticks; print the calls made and the cost of one.
    207// the clock is read with uuptime(), which costs no trap.
    208void
    209timeit(char *name, int (*f)(void))
    210{
    211 int i, t0;
    212 uint64 n = 0;
    213
    214 t0 = uuptime();
    215 while (uuptime() == t0) // start on a tick boundary
    216 ;
    217 t0 = uuptime();
    218 while (uuptime() - t0 < 20) {
    219 for (i = 0; i < 1000; i++)
    220 f();
    221 n += 1000;
    222 }
    223 // a tick is about 0.1 s: 20 ticks are about 2,000,000,000 ns.
    224 printf("vdsotest: %s: %lu calls in 20 ticks, about %lu ns each\n", name, n,
    225 2000000000UL / n);
    226}
    227
    228void
    229measure(void)
    230{
    231 // the timer CSR would be the best clock, if user mode may read it.
    232 if (inchild(userrdtime) == -1)
    233 printf("vdsotest: rdtime in user mode: killed, so time is in ticks\n");
    234 else
    235 printf("vdsotest: rdtime in user mode: allowed\n");
    236 timeit("getpid()", getpid);
    237 timeit("ugetpid()", ugetpid);
    238 timeit("uptime()", uptime);
    239 timeit("uuptime()", uuptime);
    240}
    241
    198242int
    199243main(int argc, char *argv[])
    200244{
    201245 if (argc == 2 && strcmp(argv[1], "exec") == 0) {
    202246 // after exec: the new page table maps the same pid.
    203247 exit(ugetpid() == getpid() ? 0 : 1);
    204248 }
    249 if (argc == 2 && strcmp(argv[1], "time") == 0) {
    250 measure();
    251 exit(0);
    252 }
    205253
    206254 pidtest();
    207255 uptimetest();
    208256 readonlytest();

6. Verify and measure

On the branch (ext/10-vdso, 8 commits, every commit builds), built with the project toolchain and run on 3 harts (-smp 3 -m 128M):

$ vdsotest
vdsotest: pid: OK
vdsotest: uptime agrees: OK
vdsotest: uptime advances: OK
usertrap(): unexpected scause 0xf pid=4
            sepc=0x14 stval=0x3fffffd000
vdsotest: usyscall read-only: OK
usertrap(): unexpected scause 0xf pid=5
            sepc=0x2e stval=0x3fffffc000
vdsotest: ushared read-only: OK
vdsotest: read() into usyscall fails: OK
vdsotest: write() from ushared fails: OK
vdsotest: pid unchanged: OK
usertrap(): unexpected scause 0xf pid=6
            sepc=0xd2 stval=0x3fffffd000
vdsotest: heap ceiling: OK
vdsotest: fork: OK
vdsotest: exec: OK
vdsotest: 20 children: OK
vdsotest: ALL OK
$ usertests -q
usertests starting
test copyin: OK
test copyout: OK
[...]
test nowrite: usertrap(): unexpected scause 0xf pid=6590
[...]
test lazy_copy: OK
test lazy_copyinstr: OK
test lazy_sbrk: OK
test partial_write: OK
test unlinkcwd: OK
ALL TESTS PASSED
$ vdsotest
[...]
vdsotest: ALL OK

The three usertrap() lines in vdsotest are the expected kills of the read-only and ceiling checks; the test verifies each from the child’s exit status. usertests -q passing shows that the rest of the system is unchanged: lazy_copy (system calls refuse the addresses of the new pages), lazy_sbrk (the heap reaches its new ceiling and no further), nowrite (stores to TRAPFRAME and above still kill), and the free-page counts that usertests compares. vdsotest passes again after usertests.

For comparison, at commit 4 (both pages mapped, copyin not yet changed) usertests -q stops at test lazy_copy: write succeeded, FAILED.

Time per call, from vdsotest time on 3 harts (QEMU 10.2.1; a tick is about 0.1 s, 20 ticks per measurement; times on QEMU depend on the computer and on what else it is doing, so yours will differ):

calls in 20 ticks about ratio
getpid() 175,000 11,428 ns
ugetpid() 88,831,000 22 ns 508×
uptime() 169,000 11,834 ns
uuptime() 89,029,000 22 ns 527×

Two earlier runs gave 174,000 / 168,000 getpid() and 89,731,000 / 87,894,000 ugetpid() calls: the numbers are stable to a few percent.

Instructions per call, single-stepped with gdb (QEMU’s single-step with interrupts masked, one hart locked):

user mode kernel and trampoline csrw satp sfence.vma
ugetpid() 12 0 0 0
getpid() 3 (li a7,11, ecall, ret) 1,103 (first instruction of uservec to sret) 2 4

Where the 1,103 go: 83 in the trampoline (uservec 44, userret 39), 392 in mycpu, 144 in push_off, 118 in pop_off, 84 in myproc, 86 in acquire and holding, 34 in release, 38 in killed, 43 in usertrap, 42 in prepare_return, 29 in syscall and 10 in sys_getpid. Most of the cost is bookkeeping around the trap: finding the current process four times and checking killed twice, each under p->lock. The answer itself is one load in both cases.

Instructions against time. The timing loop adds 3 instructions per call (jalr s2, addiw, bnez at 0x53c-0x540 in user/vdsotest.asm), so one iteration is 1,109 instructions with getpid (3 + 3 + 1,103) and 15 with ugetpid (3 + 12): a factor of about 74. The measured factor is about 508, so each instruction on the getpid path costs about 7 times as much time. A run with the four sfence.vma removed from trampoline.S (not valid on hardware; on QEMU the satp writes still flush): getpid went from 11,695 and 12,269 ns to 8,230 and 7,874 ns, ugetpid stayed at 21-24 ns. Translation-cache flushes are therefore a large share of the extra time per instruction; the rest was not measured.

Memory. One page per process (the usyscall page, allocated next to the trapframe) and one page in total for USHARED. Two pages of user virtual address space are taken from the top of the heap’s range.

7. Go further