kernel/proc.h
Included by 15 files
kernel/console.c, kernel/exec.c, kernel/file.c, kernel/fs.c, kernel/pipe.c, kernel/printk.c, kernel/proc.c, kernel/sleeplock.c, kernel/spinlock.c, kernel/syscall.c, kernel/sysfile.c, kernel/sysproc.c, kernel/trap.c, kernel/uart.c, kernel/vm.cAbout this file
The data structures behind processes and CPUs. Four structs and one enum:
struct context: the registersswtchsaves when a kernel thread is switched out.struct cpu: per-CPU state, one entry per hart incpus.struct trapframe: the page where a process’s user registers are saved when it enters the kernel (trapframe).enum procstate: the states a process moves through.struct proc: everything the kernel knows about one process, one entry per slot in theproctable.
The most important thing to learn here is which lock protects which field. The source
states it in comments on struct proc; the notes below check those comments against how
kernel/proc.c and the rest of the kernel actually use each field.
Read before: kernel/spinlock.h. Read next: kernel/proc.c and
kernel/swtch.S.
The registers saved by a context switch
A kernel thread that is not running is described completely by its stack and these
14 values. swtch stores them in the old context and loads them from the new
one, at offsets 0, 8, …, 104, which must match the sd/ld offsets in
kernel/swtch.S.
ra is where the thread resumes and sp is its stack. s0–s11 are the
callee-saved registers: under the calling convention a called function must
preserve them, so the code that called swtch may still need their values. The other
registers are not here because the caller of swtch cannot rely on them surviving
any function call anyway. See kernel/swtch.S for the full argument.
Each process has one (p->context) and each CPU has one for its scheduler thread
(c->context).
Return address: where the thread resumes when swtch executes ret. For a new
process, allocproc sets it to forkret.
Stack pointer. Points into the process’s kernel stack, or into the CPU’s boot stack for a scheduler context.
struct cpu: per-CPU state
One per hart, in cpus, indexed by the hart ID that cpuid reads from
tp. mycpu returns the current one.
None of these fields has a lock, and none needs one: each CPU only ever touches its own entry, and always with interrupts off, so nothing else can run on that CPU in the middle of an update.
proc: the process running on this CPU, or 0 while thescheduleritself runs. Set and cleared byscheduler; read throughmyproc.context: the scheduler thread’s saved registers.schedswitches to it.noff,intena: the nesting depth ofpush_offand the interrupt state before the outermost one (interrupts and spinlocks (push_off / pop_off)).
In the build, sizeof(struct cpu) is 128 bytes: mycpu computes
cpus + (id << 7).
The current process on this CPU, or null when the scheduler is running. myproc
returns it.
The scheduler thread’s saved registers. sched does swtch(&p->context, &mycpu()->context) to return to scheduler.
How many push_off calls are not yet matched by pop_off; normally the number of
spinlocks this CPU holds.
1 if interrupts were enabled before the outermost push_off. pop_off uses it to
decide whether to turn them back on; sched saves and restores it.
The CPU array
Declares the array defined in kernel/proc.c:9, so that other files can see it.
NCPU is 8, the most harts xv6 supports.
What the trapframe is for
When a user program traps into the kernel (a system call, an interrupt, a fault),
uservec in kernel/trampoline.S must save all 31 user registers before it can
run any C code, and it needs a few kernel values to get into the kernel at all. The
trapframe holds both. Each process has its own trapframe page, allocated in
allocproc; it is mapped at TRAPFRAME, just below the
trampoline page page, in every user page table (proc_pagetable).
The comment is accurate for this version: uservec saves the user registers and
loads kernel_sp, kernel_hartid and kernel_satp, then jumps to the address in
kernel_trap (usertrap). On the way out, prepare_return fills in the
kernel_* fields and userret restores the user registers. The kernel reaches the
same page through p->trapframe, a physical address, which works because the kernel
maps RAM at equal virtual and physical addresses.
Values the trap entry code needs
The first five slots are not user registers; they are filled in by the kernel for
uservec's benefit:
kernel_satp: the kernel page table, to load into satp;kernel_sp: the top of this process’s kernel stack;kernel_trap: the address ofusertrap;epc: the user program counter.usertrapcopies sepc here, andprepare_returncopies it back beforesret. It is kept here and not only insepcbecausesepcwould be overwritten by the next trap taken in the kernel, or by another process if this one is switched out.kernel_hartid: the hart ID, reloaded intotpbecause user code may have changedtp.
The /* 0 */, /* 8 */ comments are the byte offsets that kernel/trampoline.S
hard-codes; changing the order here would break it.
The kernel’s satp value. Written by prepare_return before every
return to user space.
Top of this process’s kernel stack, p->kstack + PGSIZE. uservec loads it into
sp.
Saved user program counter. On a system call usertrap adds 4 so the program
continues after the ecall.
The hart ID; uservec restores it into tp so cpuid works in the kernel.
The saved user registers
All 31 general-purpose registers except zero, in register-number order (ra is
x1, sp x2, …, t6 x31), at offsets 40 to 280. The whole struct is 288 bytes,
far less than its 4096-byte page.
The kernel reads and writes these slots to talk to the user program: a system call’s
arguments come from a0–a5 and its number from a7 (system call arguments); its
return value is written to a0. kfork copies the whole trapframe to the child and
then sets the child’s a0 to 0; kexec sets sp, epc and a1 for the new
program.
The process states
The process state machine, in the order the enum numbers them (0 to 5):
UNUSED: a free slot. Static storage starts zeroed, andUNUSEDis 0, butprocinitsets it explicitly anyway.USED: claimed byallocprocand being filled in; not yet allowed to run.SLEEPING: waiting insleepfor awakeup.RUNNABLE: ready; any CPU’sschedulermay pick it.RUNNING: currently executing on some CPU.ZOMBIE: has exited; waits for its parent’skwait(see zombie).
The six process states. UNUSED is 0, so a zeroed struct proc is a free slot.
struct proc, and its lock
One slot of the process table proc (NPROC = 64 slots). The build has
sizeof(struct proc) = 360 bytes.
lock is a spinlock that protects the fields on lines 86–90, and it does a
second job: it is held across every context switch of this process. The
scheduler acquires it before switching to the process, and the process holds it when
it switches back (see scheduler, sched and yield). That keeps another CPU
from picking the process while its registers are still being saved.
Fields protected by p->lock
These fields are read or written by other processes and other CPUs: the scheduler
looks at state, wakeup at chan, kkill writes killed, the parent reads
xstate and pid. So every access normally holds p->lock.
The comment is true for every write. A few reads skip the lock, safely:
- A process reads its own
pidwithout the lock (sys_getpid,acquiresleep).pidchanges only inallocproc(before the process first runs) andfreeproc(after it has exited for good), so it cannot change under a running process. procdumpreadsstateandpidwith no lock at all, on purpose (it is a debugging aid that must work even when locks are stuck).
Current process state. The scheduler reads it on every pass.
The wait channel: non-zero after sleep_prepare until a
wakeup on the same address clears it. Can be left non-zero after the process stops
waiting (see sleep); that is harmless.
Set by kkill or setkilled. Read with killed; the process exits next time
it passes through usertrap.
The exit status passed to kexit, kept until the parent’s kwait copies it out.
The PID (process ID); 0 for a free slot.
The parent pointer, protected by wait_lock
parent is protected by the global wait_lock, not by p->lock. Every write
holds it: kfork sets the child’s parent, reparent hands orphans to init, and
kwait clears it. Every read does too: kwait scans for its children and
kexit reads p->parent to wake it.
Why a separate lock? A parent’s kwait has to examine many children’s parent
fields and sleep without missing the moment one of them exits; one lock covering all
parent links makes that check-then-sleep atomic. kernel/proc.c:23 explains the
lock order this implies.
The process that created this one with fork (or init, after reparent). Protected
by wait_lock.
Fields private to the process
These are used by the process itself, so it needs no lock to touch them. The comment is right, with the understanding that “private” also allows two other parties at times when the process cannot be running:
kforkfills in a new child’s fields while it is stillUSED(not yet runnable).- The parent’s
kwaitcallsfreeprocon aZOMBIEchild, which freestrapframeandpagetableand resetsszandname. The child will never run again. procdumpreadsnamewith no lock (debug output only).
context is special: it is written by swtch when the process is switched out,
and that always happens with p->lock held.
Virtual address of the lowest byte of this process’s kernel stack, KSTACK(i) for slot
i. Set once by procinit and never changed.
Size of the user address space in bytes: user memory is [0, sz). With lazy sbrk,
not every page below sz is necessarily allocated yet.
The user page table, built by proc_pagetable and replaced by kexec.
Kernel pointer to this process’s trapframe page (a physical address, usable directly because of the kernel’s direct map).
Saved kernel registers while the process is not running; see struct context above.
Open files, indexed by file descriptor number; 0 means the descriptor is not
in use. NOFILE is 16.
The current working directory, an in-memory inode reference.
Program name for procdump: the last path element given to kexec, at most 15
characters plus the terminating 0.