Lab 5 · reveal · 19 steps · 8 commits
You add two system calls. sigalarm(n, handler) asks the kernel to call handler in user
space after every n timer ticks of CPU time the process uses; sigreturn(), called at
the end of the handler, puts the process back exactly where the timer interrupted it, with
every register as it was. sigalarm(0, 0) turns the alarm off. It is a small version of
what Unix calls a signal: the kernel makes a running program jump to a function it never
called, then makes it carry on as if nothing had happened.
About sixty lines of kernel code (not counting comments), and almost every one of them
touches something central. You will decide what “a tick of this process” means when three
harts each take their own timer interrupt but only hart 0 advances
ticks. You will find out where the hart decides which user instruction comes next, and
whether you can change it from usertrap. You will decide what to save so that the
interrupted code notices nothing, and where to keep it. Along the way: is user address 0 a
valid handler? What happens to a0 when sigreturn, itself a system call, returns a
value? And what if the handler runs longer than its interval?
Each step shows one change on the branch ext/05-sigalarm, the code around it, and the state of the machine when that code runs.
kernel/syscall.hStep 1 of 19 · commit 1: Add the sigalarm and sigreturn system calls
The story for this tour is a recorded run of alarmtest on three harts with gdb
attached to the reference kernel (the branch head). The machine state in every step
comes from that run; at this commit there is no test program yet. alarmtest is pid 3; for its first test it runs on
hart 2. sh (pid 2) and init (pid 1) are asleep in wait.
The first commit adds the plumbing, the same five places as any new system call (lab 1
walks through each one): numbers 23 and 24 here, two entry(...) lines in
user/usys.pl, two prototypes in user/user.h, and two externs and table entries
in kernel/syscall.c. In this build the stubs are at 0xd6c (sigalarm: li a7,23; ecall; ret) and 0xd74 (sigreturn) in alarmtest.
The prototype in user.h is int sigalarm(int ticks, void (*handler)());: a
function pointer. When test_called calls sigalarm(2, periodic), a0 holds 2
and a1 holds the address of periodic, which in this build is 0.
kernel/sysproc.cStep 2 of 19 · commit 1: Add the sigalarm and sigreturn system calls
gdb stopped here (at the line that stores the interval) on hart 2: sstatus.SIE = 1,
noff = 0, and pid 3’s trapframe held a7 = 0x17 (23) and epc = 0xd72, the
stub’s ecall at 0xd6e plus the 4 that usertrap added.
argint reads the interval from the saved a0; argaddr reads the handler from
the saved a1, as a plain 64-bit number. Neither checks anything, and nothing needs
checking: the kernel never dereferences the handler. It will only ever copy it into the
saved program counter, so if it is garbage, the process jumps to garbage in user mode,
and the hardware handles that like any other bad jump.
sigalarm(0, 0) stores interval 0, which is what “off” means; a negative interval is
refused. Resetting alarmticks makes every sigalarm start a fresh interval.
sys_sigreturn is a placeholder at this commit. It returns -1, and syscall will
put that in a0.
kernel/proc.hStep 3 of 19 · commit 1: Add the sigalarm and sigreturn system calls
The three fields go after name, in the section that is “private to the process, so
p->lock need not be held” (kernel/proc.h:95). The comment on alarmhandler says what
the first think question was about: 0 is a valid handler. The interval is the switch.
Who touches these fields? sys_sigalarm (this step), the timer path in usertrap
(next), and later sys_sigreturn, kexec and freeproc. All but freeproc run as
process 3 itself; freeproc runs when it is dead. That is the whole argument for no
lock.
ld sp, 8(a0) in uservec (kernel/trampoline.S:76)kernel/trap.cStep 4 of 19 · commit 2: Count a process's timer ticks in usertrap
A timer interrupt arrived on hart 2 while pid 3 was spinning in waitfor: gdb saw sepc
= 0x14c, the lw a5,-68(s0) of the volatile loop. The hart went through
uservec into usertrap, devintr() returned 2, and now, before yield(),
alarmtick(p) counts the tick.
This is the one place where all three facts are known at once: a timer interrupt
happened, it interrupted user code, and p is the process that was running. It works
the same on every hart. The compiler inlined alarmtick into usertrap (there is no
separate symbol in kernel.asm); gdb still shows it as a frame.
Interrupts are off: usertrap turns them on only for system calls
(kernel/trap.c:66), so nothing can interrupt the counting on this hart. gdb read
SIE = 0 and noff = 0 here. The saved a7 was 0xe (14), left over from waitfor’s
last uptime() call: the trapframe holds whatever the registers held.
Ticks taken in the kernel (in kerneltrap) do not reach this code. The alarm counts
user time only.
Step 5 of 19 · commit 2: Count a process's timer ticks in usertrap
A moment earlier on the same path, devintr called clockintr. On hart 2,
cpuid() == 0 is false, so the whole block is skipped: ticks is not touched. The
only thing every hart does is re-arm its own timer, stimecmp = time + 1000000, a tenth
of a second ahead at QEMU’s 10 MHz timebase (Sstc extension).
So ticks is hart 0’s wall clock, protected by tickslock because sys_uptime and
sys_pause read it from any hart. It cannot tell you how much CPU time pid 3 used: it
advances while pid 3 sleeps, and it never advances because of anything pid 3’s hart
does. Clinic 5 put the count inside this if: a lone process got 28 handler calls in 100
ticks instead of about 100, and with three processes the one that sat on hart 0 got
nearly all of them.
kernel/trap.cStep 6 of 19 · commit 3: Call the alarm handler when the interval runs out
With sigalarm(2, periodic), the second tick reaches the interval. gdb stopped on line
116 at ticks = 2: the count had just been reset to 0, the trapframe said epc =
0x14c, sp = 0x4f30, ra = 0x136, a0 = 1.
Line 116 copies all 288 bytes of the trapframe into p->alarmframe: the 31 registers
uservec saved, the user program counter, and the four kernel_* fields. (In
kernel.asm the structure assignment is an inline loop, 9 rounds of 32 bytes.) Only
then does line 117 change epc. The order matters: copy first, or the copy would
remember the handler instead of the interrupted instruction.
Line 117 is the redirect. Nothing is jumped to here. usertrap will yield(), then
prepare_return will copy epc into sepc, and the sret at the end of
userret will land on the handler.
kernel/proc.hStep 7 of 19 · commit 3: Call the alarm handler when the interval runs out
alarmframe is a whole struct trapframe embedded in struct proc, at offset 376 at
this commit (addi a4,s1,376 in the inlined copy; 384 at the branch head, after
alarmactive is added). The real trapframe is a separate
page, because the trampoline must reach it at the fixed address TRAPFRAME in every
user page table. The copy is never touched by the trampoline, so ordinary kernel
memory will do. Embedding it means no allocation and no failure path; each of the 64
slots grows by 304 bytes at this commit (360 to 664), and by 312 at the head (360 to
672).
Why not keep it on the user stack, as Unix does? That needs a copyout that can
fail on a full stack, and sigreturn would read back a frame the program could have
changed. Why not only epc? Clinic 1 answers that: the handler’s sp, s0 and ra
leak into the interrupted code.
Step 8 of 19 · commit 3: Call the alarm handler when the interval runs out
prepare_return is unchanged by this lab, and it does exactly what we need. It turns
interrupts off, points stvec at the trampoline, refreshes the four kernel_* fields
for the next trap, sets sstatus for a return to user mode with interrupts enabled, and
last, copies the saved program counter into sepc.
gdb, stopped just after prepare_return returned to usertrap, read sepc = 0x0:
the handler’s address, from the trapframe’s epc. The trapframe still holds all of the
interrupted code’s registers, sp = 0x4f30 included, so the handler will start with
them.
Notice line 142, kernel_hartid = r_tp(). Those four fields are rewritten before every
return to user space. That is why a saved copy can be restored wholesale later, even on
a different hart, without breaking the next trap.
sp holds the user’s 0x4f30 again (line 118), which S-mode cannot use until sretwas: kernel stackStep 9 of 19 · commit 3: Call the alarm handler when the interval runs out
usertrap returned the user satp to uservec's jalr, which falls into
userret. It installed pid 3’s page table (line 111), reloaded every register from
the trapframe, a0 last (line 149), and executes sret: user mode, interrupts on, and
the program counter from sepc, which is 0.
So the first user instruction is addi sp,sp,-16 at address 0, the start of periodic.
Every register holds what the interrupted waitfor had. periodic was never called,
so its ra is waitfor’s ra (0x136), and if it ever executed ret it would
“return” into the middle of waitfor. That is why a handler must end with
sigreturn().
The handler’s frame goes on the same user stack, just below waitfor’s sp: at
0x4f20. The RISC-V calling convention says procedures “must not rely upon the
persistence of stack-allocated data whose addresses lie below the stack pointer” (it
defines no red zone), so periodic can be dropped onto the stack at any instruction
without harming waitfor.
kernel/sysproc.cStep 10 of 19 · commit 4: Restore the interrupted registers in sigreturn
periodic did count++ and called sigreturn(). Before line 141, gdb read the
trapframe as the handler left it: epc = 0xd7a (the stub’s ecall at 0xd76, plus
4), sp = 0x4f20, s0 = 0x4f30, ra = 0x1e, a7 = 0x18. After line 141 it was
epc = 0x14c, sp = 0x4f30, s0 = 0x4f80, ra = 0x136, a7 = 0xe, a0 =
1: exactly the values saved at the interrupt.
One structure assignment undoes everything: the handler’s registers, and also the 4
that usertrap added to epc for this ecall. The process will not continue after
the ecall; it will continue at 0x14c, the instruction the timer interrupted, which
has not run yet.
Interrupts are on (this is a system call), and that is fine: a timer interrupt here goes
to kerneltrap, which never touches the alarm fields.
Step 11 of 19 · commit 4: Restore the interrupted registers in sigreturn
Back in syscall, line 150 stores sys_sigreturn’s return value in
p->trapframe->a0, after the restore. That is the line that makes every system call’s
result appear in a0, and here it would destroy the restored a0 unless the two are
the same value. sys_sigreturn returns p->trapframe->a0, so line 150 writes the
restored value onto itself (in kernel.asm: ld a0,112(a5), offset 112 being a0 in
the trapframe).
Clinic 4 returned 0 instead: the a0 check failed on every one of its ten interrupted
runs, while every other check passed.
kernel/trap.cStep 12 of 19 · commit 5: Do not re-enter a running alarm handler
The re-entry test: sigalarm(1, slow), and slow spins for 3 ticks. By now
alarmtest runs on hart 1 (it moved during the register test; more in a later step).
gdb caught several timer interrupts that landed inside slow, at sepc = 0x9e
with sp = 0x4ef0: alarmactive was 1 and the count stayed 0. Line 106 returned at
once.
alarmactive is set on line 116 at delivery and cleared by sigreturn. While it is
set, nothing is counted, so the next interval starts after the handler returns. Without
the check (clinic 2), slow nested 56 times and the process died in its stack’s guard
page.
kernel/sysproc.cStep 13 of 19 · commit 5: Do not re-enter a running alarm handler
sigreturn now refuses to run unless a handler is running, and clears the flag when it
does. Without this, a stray sigreturn() would copy whatever alarmframe held, which is
the frame of some earlier interrupt, or all zeros, and the process would jump back
in time with an old stack pointer. alarmtest’s first test calls sigreturn() from
main and expects -1.
Clearing the flag before the copy or after makes no difference here: interrupts are on,
but an interrupt in this system call goes to kerneltrap, never to the timer path
that reads the flag.
kernel/exec.cStep 14 of 19 · commit 6: Turn the alarm off in exec and in freeproc
In test_forkexec, a child turns its alarm on with sigalarm(1, periodic) and execs
alarmtest exec. kexec has built the new image, and from line 133 on it is
committed: the new page table, size, epc and sp are in place and the old memory is
freed. Only then are the alarm fields cleared. If exec fails earlier (at bad:), the
old program goes on running with its alarm intact.
The new program is the same alarmtest, so periodic is at the same address 0 in the
new image. That makes the check sharp: a surviving alarm would really call periodic
in the new program and count. The recorded run printed after exec: 0.
alarmhandler and alarmticks are left as they are: with the interval at 0 nothing
reads them, and the next sigalarm overwrites all three.
wait_lockthe child's p->lockkernel/proc.cStep 15 of 19 · commit 6: Turn the alarm off in exec and in freeproc
alarmtest reaps a child in kwait, which holds wait_lock and the child’s
p->lock while it calls freeproc: two spinlocks, noff 2, interrupts off, and
intena 1 because the first acquire happened during a system call with interrupts
on.
The four alarm fields are cleared with the rest. allocproc does not clear them, so
this is what makes a forked child start with the alarm off: kfork does not copy the
fields, and the slot it gets was last cleared here (or is zero from boot). The first
test_forkexec child spins for 10 ticks after fork and saw 0 handler calls.
sret in userret (kernel/trampoline.S:153)user/alarmtest.cStep 16 of 19 · commit 7: Add alarmtest, a test program for sigalarm
This is the code that ran after the sret three steps back. periodic is the first
function in the file, and alarmtest.asm shows 0000000000000000 <periodic>:; the
test prints the address, so a reader can see that 0 is real code.
The stack list above is an honest oddity: waitfor never called periodic. The
kernel dropped it on top of waitfor’s frame, so its saved ra is waitfor’s ra
(0x136), and only sigreturn() gets the process back. count is volatile because
main’s loop reads it while the handler changes it behind the compiler’s back.
slow (from line 21) is the long handler for the re-entry test: it records its nesting
depth and prints a line if it is ever entered twice. In clinic 2 that line appeared 56
times.
Step 17 of 19 · commit 7: Add alarmtest, a test program for sigalarm
fillspin is inline assembly (asm volatile): it loads 26 registers with recognizable patterns
(x11 gets 0x1111111111111111), spins 20 million times with the counter in a0, then
stores all 26 into regs[]. The clobber list tells the compiler that every one of them
changes, so it saves s1-s11 and ra in the prologue. sp, s0, gp, tp and a0
are not compared; a0 gets its own test.
gdb caught the deliveries: every one interrupted the loop at 0x52e (addi a0,a0,-1),
with ra = 0x0101010101010101, a7 = 0x1717171717171717, sp = 0x4ed0, and a0
a counter such as 0xd049b3. At the matching sigreturn, the trapframe held the
handler’s ra = 0x1e and sp = 0x4ec0; after the copy, the patterns were back.
During this test the process moved: timer interrupts 18 and 19 of the gdb run landed on
hart 2 and then on hart 1, and the trapframe’s kernel_hartid changed from 2 to 1.
alarmframe and the trapframe moved with the process, and the test passed.
Step 18 of 19 · commit 7: Add alarmtest, a test program for sigalarm
a0spin puts a pattern in a0 and spins with the counter in t0, so a0 is live at
every instruction of the loop. gdb saw the timer interrupts land at 0x7b8
(addi t0,t0,-1) on hart 1, with the trapframe’s a0 = 0x0a0a0a0a0a0a0a0a, the value
alarmframe then preserved.
This is the check that catches the sigreturn-returns-0 bug. Elsewhere a0 is often
dead at the interrupted instruction, so a clobbered a0 changes nothing. A good test
makes the value you are worried about matter.
user/alarmrate.cStep 19 of 19 · commit 8: Add alarmrate, to measure how often handlers run
alarmtest checks behaviour with generous time limits; it cannot tell a correct rate
from a wrong one (clinic 5 passed all six checks). alarmrate INTERVAL SECONDS NPROC
forks NPROC spinners, each with sigalarm(INTERVAL, tick), and lets them burn user
time for SECONDS of wall-clock time measured with uptime(). Each child reports its
handler count through its exit status, so the parent prints whole lines that
cannot interleave.
With the reference kernel on three harts: one spinner, interval 1, 10 seconds: 97 calls; three spinners: 91, 96, 94; six spinners: 45 to 49 each. The alarm follows the CPU time each process actually gets, not the clock on the wall. The full numbers are under Measure.
Lab 5 · wrap-up
On the branch head, on three harts (QEMU -smp 3), alarmtest passes, usertests -q
passes, and alarmtest passes again afterwards:
$ alarmtest
alarmtest: periodic is at address 0x0000000000000000
alarmtest: 5 calls; 0 more in 10 ticks after sigalarm(0, 0)
alarmtest: handler called, then stopped: OK
alarmtest: sigreturn outside a handler fails: OK
alarmtest: the handler ran during 10 fillspin runs (10 calls)
alarmtest: registers preserved: OK
alarmtest: slow ran 3 times, nested at most 1 deep
alarmtest: no re-entry: OK
alarmtest: the handler ran during 10 a0spin runs (10 calls)
alarmtest: sigreturn restores a0: OK
alarmtest: handler calls in the forked child: 0, after exec: 0
alarmtest: fork and exec start with the alarm off: OK
alarmtest: 6 of 6 checks OK
$ usertests -q
usertests starting
[...]
test lazy_sbrk: OK
test partial_write: OK
test unlinkcwd: OK
ALL TESTS PASSED
$ alarmtest
alarmtest: periodic is at address 0x0000000000000000
[...]
alarmtest: 6 of 6 checks OK
What this shows: the handler is called and stopped; all 26 compared registers and a0
survive ten interrupted runs each; a long handler is never nested; a forked child and an
exec’d program start with the alarm off. No program in usertests calls sigalarm, so
for every other process the only cost is one load and compare per user-mode timer
interrupt (alarminterval == 0), and 312 more bytes in each of the 64 struct proc
slots (struct proc grows from 360 to 672 bytes in this build).
Every commit on the branch was built from clean (make clean; make kernel/kernel fs.img).
Keys: ← → step · Home start