kernel/riscv.h
Included by 26 files
kernel/bio.c, kernel/console.c, kernel/exec.c, kernel/file.c, kernel/fs.c, kernel/kalloc.c, kernel/log.c, kernel/main.c, kernel/pipe.c, kernel/plic.c, kernel/printk.c, kernel/proc.c, kernel/sleeplock.c, kernel/spinlock.c, kernel/start.c, kernel/syscall.c, kernel/sysfile.c, kernel/sysproc.c, kernel/trampoline.S, kernel/trap.c, kernel/uart.c, kernel/virtio_disk.c, kernel/vm.c, user/grind.c, user/ulib.c, user/usertests.cAbout this file
The kernel’s vocabulary for talking to the RISC-V hardware. C has no way to name a
CSR (control and status register) or a special instruction, so this header wraps each one xv6 needs in a tiny
function containing one line of inline assembly: r_sstatus reads
sstatus, w_satp writes satp, sfence_vma flushes the TLB (translation lookaside buffer), and so on.
Each is static inline, so a call compiles to the single instruction
it wraps.
The second half defines the constants and macros of Sv39 paging, which
kernel/vm.c is built on: the page size, the bits of a PTE (page-table entry), and the
arithmetic that converts between physical addresses, PTEs and page-table indices.
The file is long but repetitive. Most of it is pairs of accessors, r_X() to read
register X and w_X(v) to write it; the notes explain the pattern once (lines 3–10)
and then say what each register is for and which bits xv6 uses.
Read before: kernel/start.c uses the machine-mode half. Read next:
kernel/vm.c.
Everything up to line 387 is C, which an assembler cannot read. GCC defines
__ASSEMBLER__ when it preprocesses an assembly (.S) file, so such a file that
includes this header sees only the macros after line 387.
kernel/trampoline.S does exactly that: through kernel/memlayout.h it needs
MAXVA and PGSIZE to compute TRAPFRAME.
The accessor pattern: read mhartid
Every CSR reader in this file looks like this one. Reading it closely once explains them all:
static inline: see static inline. Inkernel/kernel.asmyou find thecsrrinstruction inside each caller, not a call.asm volatile("csrr %0, mhartid" : "=r"(x));is GCC’s extended inline assembly. The string is the instruction (csrr), with%0standing for operand number 0. After the first colon come the outputs:"=r"(x)says “operand 0 is written by the instruction (=), is a general-purpose register (r) of the compiler’s choice, and afterwards holds the value ofx”. So the compiler picks a register, puts it in place of%0, and treats that register asxfrom then on.volatilestops the compiler from deleting or merging the instruction, which it might otherwise do, since it cannot see that reading a CSR twice can give different answers.
Writers (w_mstatus and the rest) have an empty output list and an input after
the second colon: "r"(x) means “put x in some register and use it as %0”.
This reader returns the hart's ID from mhartid, a
machine-mode CSR, so only start can call it.
mstatus: machine status
mstatus holds the machine-mode control fields. xv6 uses only MPP,
bits 12–11: the privilege mode mret switches to. The encoding is
the standard one for privilege levels: 3 = machine, 1 = supervisor, 0 = user (2 is
reserved). start clears the field with MSTATUS_MPP_MASK and sets it to
MSTATUS_MPP_S so that mret enters supervisor mode.
The L suffix on these constants makes them long (64 bits), so shifting further
left, as some later constants do (1L << 63), does not overflow a 32-bit int.
The MPP field, bits 12–11 (binary 11 shifted to bit 11), used to clear it.
MPP = 3: machine mode. Unused.
MPP = 1: supervisor mode. start sets this.
MPP = 0: user mode. Unused.
mepc: where mret jumps
mepc is the address mret jumps to. start sets it
to main. Only a writer exists because nothing reads it.
sstatus: the bits xv6 uses
sstatus is the supervisor’s view of the status register. The bits:
| Bit | Name | Meaning |
|---|---|---|
| 8 | SPP | mode a trap came from: 1 = supervisor, 0 = user; sret returns to it |
| 5 | SPIE | value SIE had before the trap; sret copies it back into SIE |
| 1 | SIE | supervisor interrupts enabled (the global switch) |
usertrap and kerneltrap check SPP to make sure the trap came from where
they expect, and prepare_return clears SPP and sets SPIE so that sret
enters user mode with interrupts enabled.
SSTATUS_UPIE (bit 4) and SSTATUS_UIE (bit 0) belong to user-mode interrupt
handling (the “N” extension), which never became part of the ratified RISC-V
specification. Older versions of the privileged specification listed them (as
present only with N); since version 1.12 they are reserved. xv6 never uses
them.
SPP, bit 8: 1 if the trap came from supervisor mode, 0 from user mode.
SPIE, bit 5: the interrupt-enable state before the trap, restored by sret.
Bit 4: from the abandoned user-interrupt extension; unused.
SIE, bit 1: supervisor interrupts on or off.
Bit 0: from the abandoned user-interrupt extension; unused.
Read and write sstatus
The plain accessors, used by the trap code (kernel/trap.c) to inspect the
status and to prepare sret. Changing one bit with them takes three steps (read,
modify, write); the next three functions do it in one instruction.
Set and clear sstatus bits in one instruction
s_sstatus and c_sstatus wrap csrs and csrc, which set or
clear chosen bits of a CSR in one instruction. rc_sstatus wraps
csrrc, which also returns the old value. Doing it in one instruction
matters for the interrupt-enable bit: an interrupt cannot arrive between reading the
old value and changing it. push_off relies on this to turn interrupts off and
learn whether they were on, atomically.
Two details of the operands differ from the accessors above:
- The constraint
"rK"allows either a register or a 5-bit unsigned constant (K). Every caller passesSSTATUS_SIE= 2, a constant, so GCC emits the immediate forms:csrsi sstatus,2,csrci sstatus,2andcsrrci a5,sstatus,2inkernel/kernel.asm. - The
"memory"clobber tells the compiler that the instruction may read or write any memory, so it must not move memory accesses from one side of it to the other or keep memory values cached in registers across it. Without it, the compiler could move a store that must happen with interrupts off to after interrupts are back on.
These three use the spelling __asm__ __volatile__, which means the same as
asm volatile.
Set the bits that are 1 in x; with a constant x this becomes csrsi.
Clear the bits that are 1 in x; with a constant x this becomes csrci.
Return the old value of sstatus and clear the bits of x, in one instruction.
sip: pending interrupts (unused)
sip shows which supervisor interrupts are waiting to be taken. Nothing in this version of xv6 calls these two functions.
sie: which interrupts are enabled
sie has one enable bit per kind of supervisor interrupt. start sets
the two that xv6 uses: external interrupts from devices through the
PLIC (SIE_SEIE, bit 9) and timer interrupts (SIE_STIE, bit 5).
These are per-source switches; the SIE bit of sstatus is the global one.
Bit 9: external (device) interrupts, which arrive through the PLIC.
Bit 5: timer interrupts.
mie: machine interrupt enable (unused)
mie is the machine-mode enable register; bit 5 enables supervisor timer
interrupts at the machine level. Nothing in this version uses MIE_STIE, r_mie or
w_mie; they appear to be left over from earlier versions of xv6, whose
timerinit wrote mie.
Writing mie is unnecessary for the supervisor timer: for interrupts delegated to
supervisor mode, the bits of sie are the same storage as the
corresponding bits of mie. So start's write of SIE_STIE into sie already
sets mie.STIE.
sepc: where sret returns
On a trap into supervisor mode the hardware saves the address of the interrupted
instruction in sepc; sret jumps back to whatever sepc holds.
usertrap saves it in the trapframe (another trap could overwrite the CSR before
the return), prepare_return writes it back before returning to user mode, and
kerneltrap saves and restores it around a possible yield.
medeleg / mideleg: send traps to supervisor mode
The delegation registers (medeleg and mideleg) decide which traps
go to supervisor mode instead of machine mode. start writes both once
(trap delegation); the read accessors are unused.
stvec: the trap handler's address
stvec holds the address the hart jumps to on a trap into supervisor
mode. The kernel switches it back and forth: kernelvec while running kernel code
(trapinithart, usertrap), the trampoline’s uservec just before returning
to user mode (prepare_return).
“Low two bits are mode”: bits 1–0 are the MODE field, 0 = Direct (every trap jumps
to the address) and 1 = Vectored (interrupts jump to address + 4 × cause). xv6’s
handlers are aligned to 16 bytes (.align 4), so their low bits are 0 and the mode
is Direct. r_stvec is unused.
stimecmp: the timer deadline
stimecmp (Sstc extension) sets when the next timer
interrupt fires: when time reaches it. timerinit and
clockintr write it; the reader is unused.
The instructions name the CSR by its number, 0x14d, instead of by name. The
commented-out lines show the named form: assemblers older than the Sstc extension
do not know the name, and the number works with all of them.
Read CSR number 0x14d, which is stimecmp.
Write CSR number 0x14d, stimecmp.
menvcfg: optional features
menvcfg turns on features for the lower privilege modes.
start and timerinit set two bits: MENVCFG_STCE (bit 63) enables
stimecmp, and MENVCFG_ADUE (bit 61) makes the hardware maintain the A and D
bits of PTEs (Svadu extension (A and D bits)). As with stimecmp, the CSR is named by number,
0x30a, for the sake of older assemblers.
Bit 63, STCE: enable stimecmp.
Bit 61, ADUE: hardware updates of the PTE A and D bits.
Read CSR number 0x30a, which is menvcfg.
Write CSR number 0x30a, menvcfg.
pmpcfg0 / pmpaddr0: physical memory protection
Writers for the first PMP entry (pmpcfg0 and pmpaddr0),
which start sets up to give supervisor and user mode access to all physical
memory.
Build a satp value
satp turns paging on and says where the root page table is. Its
layout: MODE in bits 63–60, an address-space ID in bits 59–44 (xv6 leaves it 0),
and the root page’s physical page number in bits 43–0.
MAKE_SATP builds that value from a page-table pointer: mode 8 (SATP_SV39)
in the top four bits, and the page number, which is the address shifted right by 12
(÷ 4096). The page table must start on a page boundary, which kalloc guarantees,
or the bits shifted out would be lost. Used by kvminithart for the kernel page
table, and by usertrap and forkret for the process’s page table, which the
trampoline loads.
MODE = 8 in bits 63–60 selects Sv39.
The satp value for a page table: Sv39 mode and the root page’s physical page number.
ASID (bits 59–44) stays 0.
Read and write satp
w_satp switches page tables; kvminithart uses it to turn paging on.
prepare_return uses r_satp to record the kernel’s satp in the trapframe,
for the trampoline to switch back to on the next trap. (The trampoline itself, being
assembly, uses csrw satp directly.)
Writing satp alone does not make the TLB (translation lookaside buffer) forget old translations. Whenever
xv6 switches to a page table (kvminithart, and kernel/trampoline.S:89–95 and
110–112) it surrounds the write with sfence.vma; the w_satp(0) in start needs
no fence because Bare mode translates nothing.
scause and stval: why the trap happened
After a trap, scause says what caused it: bit 63 is 1 for an
interrupt and 0 for an exception, and the low bits give the cause number (for
example 8 = ecall from user mode, 13 = load page fault, 15 = store page fault).
stval holds extra information; for a page fault it is the
faulting virtual address. usertrap uses both to decide what to do; see
trap cause (scause values).
mcounteren: let supervisor mode read the clock
timerinit sets bit 1 of mcounteren so that supervisor mode
may read time and use stimecmp.
time: the real-time clock
Reads time, a counter that ticks at a constant rate (10 MHz on QEMU’s
virt machine). clockintr and timerinit add an interval to it to schedule
the next timer interrupt.
The comment “machine-mode cycle counter” is wrong on both counts. time is not the
cycle counter (that is the separate cycle CSR, which counts clock cycles of the
hart), and it is not machine-mode only: this function runs in supervisor mode, which
is allowed because timerinit set mcounteren.TM.
Interrupts on, off, and are they on?
These three manipulate the global supervisor interrupt switch, the SIE bit of
sstatus. While it is 0, the hart takes no interrupts in supervisor
mode; pending ones wait until it is set again. (Exceptions, such as page faults, are
not affected.)
The comments say “device interrupts”, but the switch covers all supervisor
interrupts, the timer included. Code that holds a spinlock must keep it off
(see interrupts and spinlocks (push_off / pop_off)); that is why most of the kernel uses push_off
and pop_off rather than calling these directly.
Set SIE: the hart may now take supervisor interrupts.
Clear SIE: no supervisor interrupts until it is set again.
Read sstatus and test SIE: 1 if interrupts are on, 0 if off.
Read the sp, tp and ra registers
These read or write ordinary registers, not CSRs, using mv.
r_tpandw_tp: the tp register holds this hart’s ID, stored there bystartand read bycpuid. As the comment says, it is the index intocpus.r_spreads the stack pointer sp. Only a user program,user/usertests.c, uses it (this header is also included by user code).r_rareads the return-address register ra; nothing calls it.
Copy the stack pointer into x.
Copy tp, the hart ID, into x.
Store x in tp.
Flush the TLB
sfence_vma executes sfence.vma, needed after changing page
tables or satp. It orders the hart’s earlier writes to page tables before later
address translations and discards cached translations from the TLB (translation lookaside buffer). The two
zero operands mean “all virtual addresses” and “all address spaces”, so it flushes
everything. The "memory" clobber keeps the compiler from moving memory accesses
across it, for example a page-table write from before the flush to after it.
Flush all TLB entries, for all address spaces.
Order device accesses
io_fence executes fence iorw, iorw (fence): every device
input/output and memory read/write before it is ordered before every one after it.
kernel/virtio_disk.c uses it so that the disk device sees the request
descriptors in memory before it sees the index that announces them. See
memory barrier (fence).
Order all earlier device and memory accesses before all later ones.
Synchronize the instruction cache (unused)
icache_fence wraps fence.i, which makes later instruction
fetches on this hart see code that was stored to memory earlier. No C code calls it;
the trampoline executes fence.i itself before returning to user mode
(kernel/trampoline.S).
Make instruction fetches see earlier stores.
Types for page tables
pte_t is one page-table entry, 64 bits. pagetable_t is a pointer to
a page-table page, viewed as an array of 512 PTEs: pagetable[i] is entry i. Both
are plain integer types, so C does not stop you from mixing a PTE with an address;
the macros below do the conversions.
End of the C-only part that started at line 1.
Pages
A page is PGSIZE = 4096 = 2^12 bytes, so the low PGSHIFT = 12 bits of
an address are the offset within the page.
PGROUNDUP and PGROUNDDOWN round to a page boundary. Both clear the low 12
bits with the mask ~(PGSIZE - 1). PGSIZE - 1 is 0xfff; ~ makes it -4096,
an int, which C converts to 0xfffffffffffff000 when it is combined with a 64-bit
unsigned value, so the mask keeps every bit above bit 11. PGROUNDUP adds 4095 first,
so that any value not already on a boundary is pushed past the next one:
PGROUNDUP(1) = PGROUNDUP(4096) = 4096, PGROUNDUP(4097) = 8192.
Bytes per page: 4096.
log2 of PGSIZE: an address shifted right by 12 is its page number.
Round up to a multiple of 4096 (a size up to whole pages).
Round down to a multiple of 4096 (an address down to the start of its page).
The PTE permission bits
The low bits of a PTE (page-table entry) (layout in the glossary entry):
| Macro | Bit | Meaning |
|---|---|---|
PTE_V |
0 | valid: the entry is in use. If 0, every other bit is ignored and any access faults. |
PTE_R |
1 | readable |
PTE_W |
2 | writable |
PTE_X |
3 | executable: instructions may be fetched from the page |
PTE_U |
4 | user mode may access the page. Supervisor mode then may not (unless sstatus.SUM is set, which xv6 never does). |
A valid PTE with R, W and X all 0 is not a mapping but a pointer to the next level of
the page table; walk and freewalk depend on that rule. W without R is
reserved by the specification; xv6 never creates it.
Bits 5–7 (G, A, D) are not defined here because xv6 never sets or tests them; with Svadu extension (A and D bits) enabled the hardware maintains A and D itself.
Converting between addresses and PTEs
A PTE stores a physical address in compressed form: the page number (address ÷ 4096) in bits 53–10, leaving bits 9–0 for flags.
PA2PTE: shift right 12 to drop the offset, giving the page number; shift left 10 to move it into place above the flags. The result has all flag bits 0; the caller ORs the flags in.PTE2PA: shift right 10 to drop the flags; shift left 12 to turn the page number back into an address.PTE_FLAGS: keep only bits 9–0 (0x3FF), the flags.
Example: the page at 0x87fff000 has page number 0x87fff; PA2PTE gives
0x21fffc00, and with R, W, U and V (0x17) the PTE is 0x21fffc17. PTE2PA of
that gives 0x87fff000 back.
Physical address → the page-number part of a PTE (bits 53–10).
PTE → the physical address of the page it maps or points to.
The 10 flag bits of a PTE.
Page-table indices from a virtual address
PX(level, va) extracts the 9-bit index into the page-table page at level
(2 = root, 1, 0 = last): shift right past the 12 offset bits and 9 bits for each
lower level (PXSHIFT gives 12, 21, 30), then keep 9 bits (PXMASK = 0x1FF
= 511).
va |
PX(2) | PX(1) | PX(0) |
|---|---|---|---|
0x1234 |
0 | 0 | 1 |
0x10000000 (UART) |
0 | 128 | 0 |
0x80000000 (KERNBASE) |
2 | 0 | 0 |
0x3ffffff000 (TRAMPOLINE) |
255 | 511 | 511 |
9 bits: an index from 0 to 511.
Bit position where the index for level starts: 12, 21 or 30.
The index into the page-table page at level for address va.
The top of the address space
MAXVA = 2^38 = 0x4000000000 (256 GiB), one past the highest virtual address xv6
uses.
Sv39 addresses have 39 bits, but bits 63–39 of the 64-bit address must be copies of
bit 38, so addresses with bit 38 set are “negative”: 0xffffffc000000000 and up.
By staying below 2^38, xv6 only ever uses addresses whose upper bits are all 0, and
never has to sign-extend. The cost is half the address space: root-table entries
256–511 are never used. walk panics, and walkaddr and copyout fail, for
addresses at or above MAXVA.
1L << 38 = 0x4000000000. Written as 9 + 9 + 9 + 12 − 1 to show where it comes
from: three 9-bit indices, a 12-bit offset, minus the sign bit.