xv6, line by line
tour 25
Tours25 A user address space

Tour 25 · Memory · about 27 minutes · 17 steps

A user address space

The shell sh (pid 2) believes it owns a private memory that starts at address 0 and runs up to 256 GiB. Its code is at 0x0, its command buffer at 0x2020, its stack near 0x5000. Every other process believes the same thing about its own memory, with the same addresses meaning entirely different bytes. This tour takes that illusion apart.

First you will see the layout of sh’s address space, drawn with the real addresses from readelf and from a gdb dump of its page table taken in this build. Then you will follow one load, buf[0] at virtual address 0x2020, through the three levels of an Sv39 page table by hand: bits, indices, page-table entries and the physical byte at the end. Finally you will see why the kernel, which can reach every byte of physical memory, still cannot simply dereference 0x2020.

Physical addresses below are from our runs of this build. They depend on the order of kalloc calls, so another build, a different -smp, or a different history of commands can change them; virtual addresses, sizes and flags do not.

Best after: 5. Life of a system call, 22. exec, 24. The kernel page table and turning paging on

Who is running where

The machine has three harts. When the tour starts:

Hart What it is doing
0 Running sh (pid 2) in user mode, at its prompt: getcmd has just returned from reading a line
1 Idle in its scheduler
2 Idle in its scheduler. In a moment it will pick up sh’s child (pid 3), which will parse the command and become ls

Each hart has its own satp register and its own TLB (translation lookaside buffer). Hart 0’s satp names sh’s page table; hart 2’s names ls’s.

Three harts are running. This tour follows one path through the code, but the machine has three CPUs executing at the same time. Watch the locks held display at the top of each step, and read the Meanwhile, on other harts boxes: they show what the other CPUs could be doing at that very moment.
The route
  1. 1The linker decides the addresses user/user.ld
  2. 2The whole map kernel/memlayout.h
  3. 3Why 256 GiB and not 512 kernel/riscv.h
  4. 4Two pages every process has kernel/proc.c
  5. 5The register that names the page table kernel/riscv.h
  6. 6One load from buf user/sh.c
  7. 7Cutting the address into fields kernel/vm.c
  8. 8PX, the field extractor kernel/riscv.h
  9. 9Level 2 and level 1 kernel/vm.c
  10. 10The leaf, and the physical byte kernel/vm.c
  11. 11Every leaf in sh's table kernel/riscv.h
  12. 12The guard page is mapped, but not for you kernel/vm.c
  13. 13Walking to the top of the address space kernel/memlayout.h
  14. 14Why the kernel cannot dereference 0x2020 kernel/vm.c
  15. 15So the kernel walks the user table itself kernel/vm.c
  16. 16Where the heap would go kernel/proc.c
  17. 17What an address space costs kernel/memlayout.h

Keys: ← → step · Home start