Lab 6 · reveal · 17 steps · 9 commits
In this tree, Ctrl-C is just a byte. Type cat, press Ctrl-C, and the UART
delivers 0x03 to consoleintr, which stores it in the input buffer like any letter.
cat keeps waiting. A program that computes forever can only be stopped by rebooting. In
this lab you make Ctrl-C kill the shell’s current foreground job, every process of a
pipeline, while the shell itself survives and prints a new prompt.
The keystroke arrives in an interrupt handler, on whichever hart the PLIC picked,
usually while some unrelated process or nobody at all is running there. That raises the
questions this lab is about. The console is shared by init, the shell and every job: who
is “in front”, and who tells the console? What does “kill” mean in a kernel where a victim
can only die at certain checkpoints, and how long can it take for a process asleep in the
kernel, or one that never makes a system call? What may an interrupt handler do with a
spinlock already held, and which locks may it never take? A process asleep in
consoleread is waiting for exactly the device that is now interrupting it: does it wake?
And what happens at the edges: a ^C at an empty prompt, a job that exits just as the key
is pressed, a job that forks at that very moment?
The reference solution is nine small commits: three small system calls, one new case
in consoleintr, six lines in the shell, a pause command, an in-kernel test program
and a QEMU driver that types Ctrl-C from outside the machine.
Each step shows one change on the branch ext/06-ctrl-c, the code around it, and the state of the machine when that code runs.
kernel/proc.hStep 1 of 17 · commit 1: Add a process group id to struct proc
The story of this tour was recorded on three harts with gdb attached: the machine boots,
you type echo hi, then cat and Ctrl-C, pause 1000 and Ctrl-C, cat | grep x and
Ctrl-C, and finally a background job and one more Ctrl-C. Pids, harts and lock depths
below come from that run.
Commit 1 adds one int to struct proc. Where it goes matters: in the block headed “p->lock
must be held when using these”, next to killed and state. The kill loop will read
pgid, killed and state of another process together, and one lock must cover all
three.
0 means “in no group”. Nobody sets the field for init: the process table is a
zero-initialised global, so init starts with 0, and every process forked from it
inherits 0 until something puts it into a group. That includes the shell, which is
what keeps it safe later.
sp = 0x3fffffbf70np->lock (pid 3's)kernel/proc.cStep 2 of 17 · commit 1: Add a process group id to struct proc
You typed echo hi; the shell (pid 2) forks pid 3 on hart 1. The child must end up in
its parent’s group, and this is the last moment to do it: on line 308 it becomes
RUNNABLE and may start running on another hart at once.
The parent’s pgid may be written by someone else (the parent’s own parent can call
setpgid on it, next commit), so it is read under p->lock (lines 302-304) and copied
into a local. Then the child’s lock is taken and the group stored, in the same critical
section as the RUNNABLE. The two locks are never held together: holding a second
p->lock while holding one is a pattern this kernel avoids, and here it is not needed.
gdb at line 307: hart 1, pid 2, noff 1 (the child’s lock), intena 1 (a system call,
after usertrap's intr_on); pgid was 0, because the shell is in no group.
Note what is not copied: killed. Until line 307 runs, the child is in the table
with group 0. The last think question shows what a Ctrl-C at that moment can do.
wait_lockthe zombie's p->lockStep 3 of 17 · commit 1: Add a process group id to struct proc
freeproc runs when a zombie is reaped and resets the slot so that the next process
starts clean. Line 166 adds the group. Without it, a free slot would still “belong” to
a group, a Ctrl-C for that group would mark the free slot killed, and the next process
born there would die at its first system call, because allocproc never resets
killed (clinic 5 is exactly that: ls | wc printing 0 0 0).
The state shown is the usual one for freeproc (reasoned, not recorded for this step):
called from kwait with wait_lock and the zombie’s p->lock held.
sp = 0x3fffffbf60wait_lockpid 4's p->lockkernel/proc.cStep 4 of 17 · commit 2: Add the setpgid system call
The new system call, setpgid(pid, pgid) (number 23, wired up in syscall.h,
syscall.c, sysproc.c, user.h and usys.pl as in lab 1). pid 0 means the caller;
pgid 0 means “a new group named after pid”.
The target must be the caller or one of its children. “Child” means p->parent == me,
and p->parent is protected by wait_lock, so the scan holds wait_lock and takes
each candidate’s p->lock inside it: the kernel’s existing order (wait_lock, then
p->lock, as in kwait). The pid and state checks and the store need p->lock.
Recorded: you typed cat; the shell forked pid 4 and, on hart 2, set its group at line
652: pid 4 was still RUNNABLE (it had not run yet) with pgid 0, now 4. noff 2:
wait_lock and pid 4’s lock. A little later pid 4 itself made the same call, on the
same hart, and found its group already 4.
stack0hart 1’s slice of stack0, sp = 0x80009640, with a kernelvec frame belowcons.lockpid 6's p->lockkernel/proc.cStep 5 of 17 · commit 3: Add killpgrp and the killpg system call
The heart of the lab, and almost a copy of kkill: for every slot, under its
p->lock, if the group matches, set killed and turn SLEEPING into RUNNABLE.
Two differences: it does not stop at the first match, and it counts (the console will
need to know whether it found anyone). Group 0 is refused at line 640: that is what keeps
init and the shell out (clinic 4).
It takes nothing but p->lock, one slot at a time, so it may run in an interrupt
handler that holds cons.lock. The state shown is from commit 5’s point of view, in
the recorded run: you pressed Ctrl-C during cat | grep x, hart 1 was idle in
scheduler and took the UART interrupt, and at line 645 gdb saw three victims in a
row, all SLEEPING and all in group 6:
| slot | pid | name | asleep in |
|---|---|---|---|
| 2 | 6 | sh (the pipe’s parent) |
kwait |
| 3 | 7 | cat |
consoleread |
| 4 | 8 | grep |
piperead |
noff 2 (cons.lock plus the victim’s lock), intena 0, SIE 0. Nobody dies here. Each
victim is only made runnable; it notices killed when it runs.
kernel/sysproc.cStep 6 of 17 · commit 3: Add killpgrp and the killpg system call
killpg(pgid) (number 24) calls the same loop from a system call. The return
convention follows kill: -1 if nothing was found (including group 0), 0 otherwise.
It exists for testing. A Ctrl-C can only come from outside the machine, but everything
after the keystroke (which processes are marked, whether each one wakes, how it dies) is
the same loop, and pgtest can drive it with four sleepers in four different places.
From a system call the loop runs with interrupts on between its critical sections and
intena 1 inside them; from the console handler, with interrupts off throughout.
kernel/console.cStep 7 of 17 · commit 4: Give the console a foreground group, set by setfg
The console cannot work out “the foreground job” by itself, so it is told. The number
lives in cons, next to the input buffer, under cons.lock. That is the lock the
interrupt handler takes first anyway, so reading cons.fg there costs nothing, and a
setfg can never change it halfway through a Ctrl-C.
sp = 0x3fffffbf80cons.lockStep 8 of 17 · commit 4: Give the console a foreground group, set by setfg
setfg(pgid) (number 25; the handler in sysproc.c refuses negative numbers) ends
here. Recorded: right after the setpgid of the previous steps, the shell made group 4
the foreground on hart 2: noff 1, intena 1.
cons.lock is taken both here, in a system call with interrupts on before the
acquire, and in consoleintr, in an interrupt handler. That is safe for the reason
Locks and interrupt state gives: while this hart holds cons.lock, its interrupts are
off, so a UART interrupt cannot arrive on this hart and spin forever on a lock its own
hart holds. One on another hart waits a few instructions.
stack0hart 2’s slice of stack0, sp = 0x8000a690, above a kernelvec framecons.lockkernel/console.cStep 9 of 17 · commit 5: Kill the foreground group when Ctrl-C is typed
The new case. Recorded for the first Ctrl-C (cat waiting): the interrupt was taken
by hart 2, which had nothing to run. It was in scheduler's intr_on(); intr_off();
window (an idle hart sleeps in wfi with interrupts off and takes the pending interrupt
the next time round that window), so kernelvec pushed its frame onto hart 2’s
scheduler stack and kerneltrap → devintr → uartintr → consoleintr run
on top of it. All five Ctrl-Cs in the recording landed on an idle hart, in turn on
harts 2, 2, 1, 1 and 2. cons.lock is held from line 152: noff 1, intena 0.
Line 159 drops the line being edited (the characters from cons.w to cons.e, the
same region ^U erases), so a half-typed command never reaches the next reader. Lines
160-162 echo ^C and a newline through consputc, which writes the UART directly
with interrupts off: safe here, unlike the sleeping uartwrite. Line 163 kills the
foreground group (previous steps): at this moment cons.fg was 4, cat’s group, and
cat (pid 4, slot 2) was SLEEPING.
hart 2's scheduler stack (stack0 slice), during the Ctrl-C
top ─► start, main (boot frames, never popped)
scheduler
kernelvec (256-byte frame: the interrupted registers)
kerneltrap
devintr
uartintr
consoleintr (sp = 0x8000a690)
stack0cons.lockStep 10 of 17 · commit 5: Kill the foreground group when Ctrl-C is typed
If the kill loop found nobody (at the prompt, or a group whose processes are gone), the
reader gets an empty line: one \n stored and published by moving cons.w, exactly as
a typed Enter would, minus the echo (the handler already echoed a newline). The shell’s
read returns it, the shell skips the blank command and prints a fresh $ . The test
on line 163 also checks for room, as the default case does: a buffer full of
unread lines gets nothing more.
Line 171 is the subtle one, and it runs in both cases. A reader that was killed while
fully asleep is already RUNNABLE. But a reader that had released cons.lock and not
yet reached sleep is still RUNNING, with p->chan = &cons.r; the kill loop
only set its killed. wakeup clears its chan, so sleep returns at once
instead of sleeping until the next Enter. When nobody was killed, the same call wakes
the shell for its empty line. (It helps only readers of the console; a victim in the
same window in piperead or kwait is not covered, as the think section explains.)
Recorded: the wakeup at line 171 ran on hart 2 with noff 1 still (each p->lock inside
wakeup makes it 2 for a moment).
sp = 0x3fffff9f00Step 11 of 17 · commit 5: Kill the foreground group when Ctrl-C is typed
Back to the first Ctrl-C. cat (pid 4) had been asleep inside sleep (line 109),
called from this loop. Made RUNNABLE by the kill loop on hart 2, it was picked up by
hart 0’s scheduler within the same tick, returned from sleep, took cons.lock
again, found the buffer still empty (cons.r = cons.w = 12: the Ctrl-C added
nothing), and this time the test on line 103 said yes. gdb recorded it at line 105,
after the release: hart 0, noff 0, SIE on.
Note the migration: cat went to sleep on whatever hart it was running on and wakes
on hart 0, chosen by whichever scheduler found it first. consoleread returns -1;
read in cat never sees it.
sp = 0x3fffff9fe0: only usertrap’s frame is leftStep 12 of 17 · commit 5: Kill the foreground group when Ctrl-C is typed
The system call has returned into usertrap, and before going back to user space it
checks killed again (kernel/trap.c:81, unchanged by the branch). cat calls
kexit(-1) here, on hart 0. Every killed process in the recording died at this line:
cat 4, pause 5, the pipeline’s sh 6, cat 7, grep 8 and the last cat, 12. pause and grep
returned from sys_pause and piperead with -1 first; sh 6 from kwait.
kexit closes the files, wakes the parent and becomes a zombie; the shell (pid 2),
asleep in wait, reaps it. Its exit status is -1, the same value a process killed for
a bad memory access gets, which is why the shell does not try to print “interrupted”.
The last of the kernel commits is done. Up to here the shell never sets a group, so every Ctrl-C finds nobody and only gives an empty line. The next commit lets it say which job is in front.
user/sh.cStep 13 of 17 · commit 6: Run each shell job in its own foreground group
Six lines in the shell, and the only place that knows where a job begins and ends.
runcmd can fork anything. Everything the job forks from here on inherits it.RUNNABLE.fork1() on, the
foreground group is still 0: a Ctrl-C there only produces an empty line, which the
new job may read.)wait, nobody is in front. A Ctrl-C at the prompt finds group 0,
is refused, and gets an empty line.The shell itself never calls setpgid for itself: it stays in group 0 with init,
the group no Ctrl-C can reach.
The recorded sequence for cat: fork → pid 4; shell setpgid(4, 4); shell
setfg(4); pid 4 setpgid(0, 0); pid 4 execs cat; Ctrl-C; wait returns 4;
setfg(0).
Step 14 of 17 · commit 6: Run each shell job in its own foreground group
For (pause 30; echo survived) & the shell’s child (pid 9, group 9, in front) forks
pid 10 for the command and exits at once; the shell’s wait returns and it calls
setfg(0). Pid 10 inherited group 9, so before it does anything else it moves to a
group of its own (line 128). In the recording it did so on hart 0 (gdb saw its pgid
go from 9 to 10), and the pause it forked a moment later was born into group 10.
Group 10 is never made the foreground one, so no Ctrl-C reaches it: in the test, cat
and a Ctrl-C ran while the job slept, and survived appeared on time.
The shell’s child does not need to set pid 10’s group as well: nothing here races with
a setfg. The cost of that simplicity is a window of a few microseconds after the
fork in which a Ctrl-C for job 9 would still kill pid 10.
user/pause.cStep 15 of 17 · commit 7: Add the pause user program
A two-line program around the pause system call: sys_pause sleeps on &ticks and
checks killed every tick (Tour 14: pause(n) and the tick counter). It is the second kind of foreground job the
lab needs: asleep in the kernel but not waiting for the console, so it shows that the
kill reaches jobs that the console has nothing to do with. If pause is killed, the
system call returns -1 and usertrap ends the process before main sees the value,
so the program does not check it.
user/pgtest.cStep 16 of 17 · commit 8: Add pgtest, a test of process groups and killpg
The in-kernel test. A leader in a new group forks three members; together they are
asleep in four different places: kwait (the leader), piperead on a pipe whose
write end the member holds itself (so it never sees end-of-file), sys_pause, and
consoleread. Each writes its pid to ready just before blocking; pgtest waits
half a second more so that all four are really asleep, then calls killpg once.
How does a test know they all exited, without a way to ask? Every member inherited
done[1] and pgtest closed its own copy (line 87), so the read on line 109 returns
0 (end-of-file) only when all four have exited and closed it. If the kill failed to
wake someone, that read would block forever, so a watchdog (lines 95-105) writes one
byte to the pipe wd, then kills the members one by one with kill() after 3 seconds.
pgtest reports FAIL if that byte is there once the watchdog has exited: clinic 3
shows that line failing. (The byte, rather than the watchdog’s exit status, is the
evidence: a watchdog preempted right after its last kill() would be killed by pgtest
itself and exit with -1 just like one that never acted.)
A process in another group (forked first, line 59) must still answer on its pipe after
the killpg. The last line of the output says what this program cannot test: the
keystroke.
ctrlc-test.pyStep 17 of 17 · commit 9: Add ctrlc-test.py, which types Ctrl-C at a booted xv6
The only way to press Ctrl-C is from outside the machine. QEMU’s -nographic connects
the UART to its standard input and output, so the script starts QEMU with both on
pipes, writes cat\n, waits a second, writes the byte \x03, and waits for
^C\n$ to come back. Each scenario of the spec is one check: cat, pause 1000
(it also prints how long the prompt took), cat | grep x, a background job surviving,
Ctrl-C at the prompt, a half-typed line dropped, and finally that init: starting sh
appeared only once, so the shell was never replaced.
A pitfall: a driver that waits for the text survived and then immediately sends the
next \x03 breaks. echo writes its argument and the newline with two separate
write calls, so the kernel’s ^C\n can land between them and the next check sees
^C\n\n$ . The test waits for survived\n. Output from three harts and a test
driver interleaves at byte granularity; tests must wait for whole lines.
A second pitfall: a half-typed-line check that only looks for the text junecho is
too weak. A kernel without cons.e = cons.w passes it, because with no job the Ctrl-C
completes the line with its \n and the shell simply runs echo jun. So the check
demands echo jun^C followed directly by a prompt (line 122). With that line removed
from the kernel, the check fails (9 of 10); the console shows echo jun^C, then jun.
Commands after -- are typed at the end: the verification below ran pgtest,
usertests -q and pgtest again in the same boot.
Lab 6 · wrap-up
On the branch (ext/06-ctrl-c, 9 commits), run on 3
harts (-smp 3 -m 128M) by ctrlc-test.py kernel fs.img -- pgtest "usertests -q" pgtest,
all in one boot:
$ cat
^C
$
### ctrlc-test: cat waiting for input, then ^C: OK
pause 1000
^C
$
### ctrlc-test: pause 1000, then ^C (prompt after 0.001 s): OK
cat | grep x
^C
$
### ctrlc-test: cat | grep x, then ^C: OK
[...]
### ctrlc-test: the shell was never restarted: OK
pgtest
pgtest: setpgid on a process that is not our child fails: OK
[...]
pgtest: ALL OK (the Ctrl-C keystroke itself is tested by ctrlc-test.py)
$ usertests -q
usertests starting
test copyin: OK
test copyout: OK
[...]
ALL TESTS PASSED
$ pgtest
[...]
pgtest: ALL OK (the Ctrl-C keystroke itself is tested by ctrlc-test.py)
$
ctrlc-test: 10 of 10 checks OK
ctrlc-test: ALL OK
(The ### ctrlc-test: lines are the driver’s own, printed between the console output it
passes through.) Two more boots with the same commands gave the same result: 10 of 10,
pgtest ALL OK twice, ALL TESTS PASSED.
What it shows: Ctrl-C reaches a reader of the console, a sleeper in pause, and all
three processes of a pipeline; a background job lives through it; at the prompt it only
produces a new prompt; and the shell is never restarted. pgtest shows that killpg wakes
sleepers in kwait, piperead, sys_pause and consoleread without help.
usertests -q passing shows that nothing else changed, including kill itself
(killstatus, kill-based tests) and every program that forks and waits, now through a
shell that calls three new system calls per command.
The same ctrlc-test.py against the original kernel: ctrlc-test: 1 of 10 checks OK
(only “the shell was never restarted”). ^C is stored as the byte 0x03 and the first
cat is never interrupted, so it reads every line typed after it: the driver’s
echo alive came back twice, once as the console’s echo and once printed by cat.
Every commit builds on its own (make kernel/kernel fs.img, only the usual RWX link
warning).
Keys: ← → step · Home start