xv6, line by line
lab 6

Extension labs · lab 6 · Traps and control flow · ★★★☆☆

Ctrl-C: interrupting the foreground job

In this tree, Ctrl-C is just a byte. Type cat, press Ctrl-C, and the UART delivers 0x03 to consoleintr, which stores it in the input buffer like any letter. cat keeps waiting. A program that computes forever can only be stopped by rebooting. In this lab you make Ctrl-C kill the shell’s current foreground job, every process of a pipeline, while the shell itself survives and prints a new prompt.

The keystroke arrives in an interrupt handler, on whichever hart the PLIC picked, usually while some unrelated process or nobody at all is running there. That raises the questions this lab is about. The console is shared by init, the shell and every job: who is “in front”, and who tells the console? What does “kill” mean in a kernel where a victim can only die at certain checkpoints, and how long can it take for a process asleep in the kernel, or one that never makes a system call? What may an interrupt handler do with a spinlock already held, and which locks may it never take? A process asleep in consoleread is waiting for exactly the device that is now interrupting it: does it wake? And what happens at the edges: a ^C at an empty prompt, a job that exits just as the key is pressed, a job that forks at that very moment?

The reference solution is nine small commits: three small system calls, one new case in consoleintr, six lines in the shell, a pause command, an in-kernel test program and a QEMU driver that types Ctrl-C from outside the machine.

Read first: Tour 9: Device interrupts and the PLIC, Tour 16: sleep and wakeup, and the lost-wakeup problem, Tour 21: exit, wait and zombies, Tour 23: kill, Tour 37: A keystroke's journey, Tour 40: Capstone: the shell running ls | wc, Tour 44: One interrupt, three landing sites, Tour 51: The lock-order graph, measured · Locks and interrupt state, The stacks of xv6

What this lab teaches

  • Can the console tell by itself which processes are “in front”? If not, who knows, and what is the smallest piece of state that describes a whole pipeline?
  • What does kill actually do to a process in this kernel, and where does each kind of victim (asleep in the kernel, running in user mode) finally notice it?
  • What may code running in an interrupt handler do: which locks may it take, in which order, and what happens on three harts if it gets the order wrong?
  • Can a process be killed in the instant between deciding to sleep and actually sleeping, and what does it take for the kill to reach it anyway?
  • How do init and the shell stay out of harm’s way, and what can a field left behind in a freed process slot do to the next process born there?
  • How do you test a feature that starts with a keystroke no program can type?

The reference branch

ext/06-ctrl-c in ShowMeTheStack/xv6-riscv-labs, branched from the frozen commit 06aad25; 9 commits.

git clone https://github.com/ShowMeTheStack/xv6-riscv-labs
cd xv6-riscv-labs
git checkout -b my-ctrl-c 06aad25   # start your own
git diff 06aad25 origin/ext/06-ctrl-c   # only when you want the answer

1. The spec

Behaviour. Typing Ctrl-C at the console kills every process of the shell’s current foreground job: a single command, every process of a pipeline, a list. The shell itself is not killed; as soon as the job’s processes have exited it prints a new prompt.

$ cat
^C
$ pause 1000
^C
$ cat | grep x
^C
$ (pause 30; echo survived) &
$ cat
^C
$ survived
^C
$ echo jun^C
$ echo clean
clean

(pause N is a new command that sleeps for N ticks of 1/10 s, so pause 1000 would take 100 seconds. In the line echo jun^C, ^C is the kernel’s echo; the half-typed echo jun never reaches the shell.)

Three new system calls, which the shell and the test use:

What must not change. kill(pid) and every existing program behave as before; init is never killed by Ctrl-C (the kernel panics if init exits). usertests -q must print ALL TESTS PASSED on 3 harts.

The tests. A Ctrl-C is a byte arriving at the UART, which no program inside xv6 can produce, so there are two:

2. Think first

Answer each question in your head (or on paper) before opening a hint. Hints get more specific; the reference answer comes last.

1Who should receive the Ctrl-C?

Ctrl-C arrives as one byte, 0x03, in consoleintr (kernel/console.c:147), called from an interrupt handler. You want it to kill “the program the user is running”. The spec has the shell call setfg to tell the console which processes those are. Why must the kernel be told at all? While cat | grep x runs, which processes have the console open, and which of them should die? Try every rule the kernel could apply with what it already knows (open files, parent pointers, who is reading the console) and find the case each one gets wrong. Then: what must the thing the shell passes to setfg name, so that one call covers a whole pipeline?

Check yourself

1warm-upType a number

In the unmodified tree, you type cat | grep x and both programs are running. How many processes have the console open on at least one file descriptor?

decimal, 0x hex or 0b binary
2solidChoose one

Suppose Ctrl-C killed every process whose file table contains the console. You type cat and press Ctrl-C. What happens?

2How does a job get its group, and who decides?

You have decided on process groups and a foreground group. Now the mechanics. When does a process get its group number? The shell forks one child per command line, and that child may fork again (a pipeline) before or after the shell gets to run again: on three harts the order is not fixed. Who should call your “set the group” system call, the shell, the child, or both, so that every process of the job is in the group before the shell makes that group the foreground one? What about a background job, cmd &? And what should init and the shell’s own group be?

Check yourself

1solidPut in order

Put the shell’s own steps (the parent side) for one command line in the order the reference uses.

  1. setpgid(pid, pid) puts the leader in a new group
  2. setfg(pid) makes that group the foreground one
  3. setfg(0): no foreground job at the prompt
  4. fork1() creates the job’s leader
  5. wait(0) until the leader exits
2deepTrue or false, and why

True or false: it would be enough for the shell’s child to call setpgid(0, 0) as its first action; the shell’s own setpgid(pid, pid) is redundant.

Why?

3What does "kill" mean here, and when does the victim notice?

You will mark the processes of the foreground group as killed from the interrupt handler. Look at how this tree kills one process with kill(pid). What exactly does it change in the victim, and when does the victim actually stop running? Go through the victims of a Ctrl-C one by one: cat asleep reading the console, pause 1000 asleep for 100 seconds, grep asleep reading an empty pipe, the pipeline’s parent asleep in wait, and a program that loops in user mode without ever making a system call. For each one, where does it notice?

Check yourself

1solidMatch the pairs

Match each victim of a Ctrl-C with the place where it first notices that it was killed.

2deepChoose one

A Ctrl-C arrives on hart 0 while spin, a program that loops forever in user mode and never makes a system call, is running on hart 1. The handler sets spin’s killed flag. What happens next?

4The kill runs in an interrupt handler. What may it do there?

Your kill loop will be called from consoleintr. Trace how consoleintr is reached and what is true at that moment: which locks are held, whether interrupts are on, which stack, and which process (if any) is “current” on that hart. Then decide what your loop may lock. It must look at every process’s group number. Which lock protects that field? Could it also take wait_lock (say, to walk parent pointers instead of using groups)? Could it sleep, or wait for the victims to die?

Check yourself

1solidFill in the machine state

Hart 1 is idle in scheduler when you press Ctrl-C. It takes the UART interrupt, and the reference’s kill loop has just acquired the p->lock of the first victim. Fill in the hart’s state.

kernel/console.c
151 switch (c) {
152 case C('P'): // Print process list.
154 break;
2deepChoose one

Instead of groups, a learner kills “the foreground leader and all its descendants”: for each process it acquires p->lock, then wait_lock to follow the parent pointers up, releases wait_lock, marks the process if the leader was found, and releases p->lock. Run on 3 harts while the job forks and waits a lot, what is the risk?

5Does the console reader wake up?

Now look at the victim that is waiting for this very device: cat in consoleread. Your kill loop does what kkill does: under p->lock, set killed, and make a SLEEPING process RUNNABLE. Is that enough for a reader of the console? Look closely at the steps between the reader’s killed test and the moment it is really asleep. Which locks does it hold at each step, and can your handler run in between? What does the reader’s state look like then, and what will it do next?

Check yourself

1deepChoose one

The handler’s kill loop sets killed on cat while cat is between release(&cons.lock) and sleep() in consoleread. Without the extra wakeup(&cons.r), what happens to cat?

kernel/console.c
96 while (n > 0) {
97 // wait until interrupt handler has put some
98 // input into cons.buffer.
99 while (cons.r == cons.w) {
100 if (killed(myproc())) {
102 return -1;
103 }
108 }
2solidTrue or false, and why

True or false: if the kill loop only set killed (no RUNNABLE, no wakeup(&cons.r)), a Ctrl-C would leave pause 1000 running for the full 100 seconds.

Why?

6Keeping the shell and init alive, and Ctrl-C with no job

The shell must never die from a Ctrl-C, and init must not die at all. With the design so far, which group is the foreground group while the shell sits at its prompt? What should a Ctrl-C do then, and what should happen to a half-typed command line? Who prints the newline and the next prompt after a job is killed: the kernel or the shell? Can the shell tell, from wait, that its job died from a Ctrl-C?

Check yourself

1solidChoose one

A learner’s kill loop forgets to refuse group 0. The shell resets the foreground group to 0 after every job, as in the reference. You type echo hi, then press Ctrl-C at the empty prompt. What happens?

2warm-upChoose all that apply

In the reference, what does a Ctrl-C typed at an empty prompt (no foreground job) do?

7What happens at the edges of a job's life?

Two races remain. First: the job exits just as you press Ctrl-C. The shell has reaped it but not yet called setfg(0), so the foreground group names a group whose processes are gone. Can that number now belong to someone else? What about the process-table slots those processes used, and the fields left in them? Second: the job’s leader is in the middle of fork when the Ctrl-C arrives. The kill loop visits the slots one at a time. Can the new child escape?

Check yourself

1solidTrue or false, and why

True or false: in this tree, after the processes of group 7 have all exited and been reaped, a later job may also be given group number 7.

Why?

2deepChoose one

The job’s leader is inside kfork, between allocproc and the line that copies its group into the child, when the Ctrl-C handler’s loop runs over the table. What can happen in the reference?

3. Build it

Start.

git checkout -b my-ctrl-c 06aad25

Begin with the pause program (user/pause.c: pause(atoi(argv[1])), add $U/_pause\ to UPROGS) so that you have a job that sleeps without reading the console. On the original kernel, type cat, then Ctrl-C, then Enter, then Ctrl-D: cat prints the line holding the 0x03 byte, and only Ctrl-D ends it. That is the behaviour you are replacing.

Milestones, each one bootable.

  1. The group field. int pgid in struct proc (under p->lock); kfork copies it from the parent into the child before making it RUNNABLE; freeproc clears it. Test: boot and run usertests -q (nothing uses groups yet).
  2. setpgid(pid, pgid). The usual five places for a new system call (syscall.h, syscall.c, sysproc.c, user.h, usys.pl; see lab 1). The target must be the caller or a child: hold wait_lock to read p->parent, then the target’s p->lock.
  3. The kill loop and killpg. A group version of kkill, refusing group 0. Write pgtest now (four sleepers in four places, a watchdog, a survivor in another group) and make it pass: everything below the keystroke is tested here.
  4. The foreground group. A field in cons under cons.lock, and setfg(pgid).
  5. The keystroke. A case C('C') in consoleintr: drop the edited line, echo, kill, wakeup(&cons.r), and the empty line when nobody was found. Test by hand: Ctrl-C at the prompt must give a new prompt (the shell does not set a group yet, so this is the only case you can try).
  6. The shell. setpgid on both sides of the fork in main, setfg around wait, and setpgid(0, 0) in the background case. Now cat, pause 1000 and cat | grep x can be interrupted.
  7. The driver. A script that boots QEMU with its standard input on a pipe and writes \x03 there (the reference’s ctrlc-test.py). Run it, then pgtest, usertests -q, pgtest again.

Typing Ctrl-C into QEMU. With -nographic and a terminal, Ctrl-C reaches xv6 (QEMU’s own escape key is Ctrl-A). From a script, write the byte \x03 to QEMU’s standard input.

Debugging advice. For breakpoints at boot, start QEMU halted with -S -gdb tcp::GDBPORT (GDBPORT is any free port, e.g. 25501), or use make qemu-gdb, which picks its own port; for a Ctrl-C you can also attach later.

4. Debugging clinic

Each of these bugs was put into the reference solution on purpose and run on three harts. The symptom is exactly what happened. Try to explain it before revealing why.

1Killing everything except init

No groups at all: Ctrl-C kills every process but init. In the reference kill loop:

-  if (pgid <= 0)
-    return 0; // group 0 is "no group": init and the shell
   for (p = proc; p < &proc[NPROC]; p++) {
     acquire(&p->lock);
-    if (p->pgid == pgid) {
+    if (p->pid > 1) { // everyone but init

What happened when we ran it

init: starting sh
$ cat
^C
init: starting sh
$ (pause 30; echo survived) &
$ cat
^C
init: starting sh
$ echo alive
alive
$

2Taking wait_lock inside the kill loop

A learner kills “the foreground leader and its descendants” and, knowing that p->parent needs wait_lock, takes it inside the loop, while holding p->lock:

   for (p = proc; p < &proc[NPROC]; p++) {
     acquire(&p->lock);
-    if (p->pgid == pgid) {
+    acquire(&wait_lock); // p->parent needs wait_lock
+    mine = 0;
+    for (q = p; q != 0; q = q->parent)
+      if (q->pid == pgid)
+        mine = 1;
+    release(&wait_lock);
+    if (mine) {

The workload, forkwait (not part of the branch): three workers, each forking a child that exits at once and reaping it, forever. Type forkwait, wait 1 second, Ctrl-C.

What happened when we ran it

$ forkwait
^C

# gdb attached 20 s after the Ctrl-C, same boot:
[...]
  Id   Target Id                    Frame 
* 1    Thread 1.1 (CPU#0 [running]) acquire (lk=lk@entry=0x8000f9c8 <wait_lock>) at kernel/spinlock.c:37
  2    Thread 1.2 (CPU#1 [running]) acquire (lk=lk@entry=0x8000f9c8 <wait_lock>) at kernel/spinlock.c:37
  3    Thread 1.3 (CPU#2 [running]) acquire (lk=lk@entry=0x80010380 <proc+1440>) at kernel/spinlock.c:37

Thread 3 (Thread 1.3 (CPU#2 [running])):
#0  acquire (lk=lk@entry=0x80010380 <proc+1440>) at kernel/spinlock.c:37
#1  0x000000008000248a in kwait (addr=0) at kernel/proc.c:391
#2  0x0000000080002bae in sys_wait () at kernel/sysproc.c:36
#3  0x0000000080002b1c in syscall () at kernel/syscall.c:152
#4  0x00000000800028a2 in usertrap () at kernel/trap.c:68
#5  0x0000003ffffff09c in ?? ()

Thread 2 (Thread 1.2 (CPU#1 [running])):
#0  acquire (lk=lk@entry=0x8000f9c8 <wait_lock>) at kernel/spinlock.c:37
#1  0x0000000080001de8 in kfork () at kernel/proc.c:297
#2  0x0000000080002b8c in sys_fork () at kernel/sysproc.c:28
#3  0x0000000080002b1c in syscall () at kernel/syscall.c:152
#4  0x00000000800028a2 in usertrap () at kernel/trap.c:68
#5  0x0000003ffffff09c in ?? ()

Thread 1 (Thread 1.1 (CPU#0 [running])):
#0  acquire (lk=lk@entry=0x8000f9c8 <wait_lock>) at kernel/spinlock.c:37
#1  0x000000008000226c in killpgrp (pgid=3) at kernel/proc.c:648
#2  0x00000000800003fc in consoleintr (c=<optimized out>) at kernel/console.c:163
#3  0x0000000080000ab2 in uartintr () at kernel/uart.c:153
#4  0x00000000800027d2 in devintr () at kernel/trap.c:199
#5  0x000000008000283a in usertrap () at kernel/trap.c:69
#6  0x0000003ffffff09c in ?? ()
$1 = 3
$2 = 1
$3 = 2
[...]
$6 = {locked = 1, name = 0x80007168 "wait_lock", cpu = 0x8000fae0 <cpus+256>}
[...]
$8 = 11
[...]
slot 4 lock held by cpus+0
slot 4 pid 5 forkwait RUNNABLE killed 0 pgid 3 parent 3 chan 0x0 lock 1

3The kill only sets the flag

The learner reasons that “kill means killed = 1” and writes the loop like setkilled: no SLEEPING → RUNNABLE, and no wakeup(&cons.r) in consoleintr:

       p->killed = 1;
-      if (p->state == SLEEPING) {
-        // Wake process from sleep().
-        p->state = RUNNABLE;
-      }
       n++;
[...]
-    // a killed reader may be between sleep_prepare() and
-    // sleep() in consoleread(); this clears its chan.
-    wakeup(&cons.r);
     break;

What happened when we ran it

$ cat
^C

$ pause 1000
^C
$ cat | grep x
^C

$ pgtest
pgtest: setpgid on a process that is not our child fails: OK
pgtest: killpg(0), the group of init and the shell, is refused: OK
pgtest: killpg finds the group (wait, pipe read, pause, console read): OK
pgtest: every member exited: OK
pgtest: ...within 3 seconds, without the watchdog's kill(): FAIL
pgtest: the leader exited with status -1, like any killed process: OK
pgtest: a process in another group survives: OK
pgtest: SOME TESTS FAILED
$

4Group 0 is not refused

The kill loop forgets that 0 means “no group”:

-  if (pgid <= 0)
-    return 0; // group 0 is "no group": init and the shell
   for (p = proc; p < &proc[NPROC]; p++) {

What happened when we ran it

init: starting sh
$ echo hi
hi
$ ^C
panic: init exiting

5A freed slot keeps its group number

freeproc resets every field it knows about, but the learner forgot the new one:

   p->pid = 0;
-  p->pgid = 0;
   p->name[0] = 0;

What happened when we ran it

init: starting sh
$ echo hi; cat
hi
^C
$ ls | wc
0 0 0
$ echo after
after
$

# a second boot under gdb, same steps; one record from a breakpoint at
# usertrap's kexit(-1) for a killed system call (trap.c:58):
{"tag": "syscall-kexit", "n": 1, "hart": 2, "cur": {"pid": 6, "name": "sh", "state": "RUNNING", "pgid": 5, "slot": 3, "killed": 1, "kstack": "0x3fffff7000"}, "noff": 0, "intena": 0, "sie": 0, "sp": "0x3fffff7fe0", "satp": "0x8000000000087fff", "ticks": 14, "fg": 5, "bt": ["usertrap", "0x3ffffff09c"], "a7": 21}

5. The reference solution

Take the guided tour through the reference solution, one commit at a time, with the machine state at every step:

Open the reveal tour →

Or read the commits

  1. 9708c9a Add a process group id to struct proc

    kernel/proc.c

    @@ -162,8 +162,9 @@ freeproc(struct proc *p)
    162162 proc_freepagetable(p->pagetable, p->sz);
    163163 p->pagetable = 0;
    164164 p->sz = 0;
    165165 p->pid = 0;
    166 p->pgid = 0;
    166167 p->name[0] = 0;
    167168 p->chan = 0;
    168169 p->killed = 0;
    169170 p->xstate = 0;
    @@ -257,9 +258,9 @@ growproc(int n)
    257258// Sets up child kernel stack to return as if from fork() system call.
    258259int
    259260kfork(void)
    260261{
    261 int i, pid;
    262 int i, pid, pgid;
    262263 struct proc *np;
    263264 struct proc *p = myproc();
    264265
    265266 // Allocate process.
    @@ -296,9 +297,15 @@ kfork(void)
    296297 acquire(&wait_lock);
    297298 np->parent = p;
    298299 release(&wait_lock);
    299300
    301 // the child joins the parent's process group.
    302 acquire(&p->lock);
    303 pgid = p->pgid;
    304 release(&p->lock);
    305
    300306 acquire(&np->lock);
    307 np->pgid = pgid;
    301308 np->state = RUNNABLE;
    302309 release(&np->lock);
    303310
    304311 return pid;

    kernel/proc.h

    @@ -87,8 +87,9 @@ struct proc {
    8787 void *chan; // If non-zero, sleeping on chan
    8888 int killed; // If non-zero, have been killed
    8989 int xstate; // Exit status to be returned to parent's wait
    9090 int pid; // Process ID
    91 int pgid; // Process group (0: none); ^C kills a group
    9192
    9293 // wait_lock must be held when using this:
    9394 struct proc *parent; // Parent process
    9495
  2. 94bc06e Add the setpgid system call

    kernel/defs.h

    @@ -86,8 +86,9 @@ int growproc(int);
    8686void proc_mapstacks(pagetable_t);
    8787pagetable_t proc_pagetable(struct proc *);
    8888void proc_freepagetable(pagetable_t, uint64);
    8989int kkill(int);
    90int ksetpgid(int, int);
    9091int killed(struct proc*);
    9192void setkilled(struct proc*);
    9293struct cpu* mycpu(void);
    9394struct proc* myproc();

    kernel/proc.c

    @@ -627,8 +627,40 @@ kkill(int pid)
    627627 }
    628628 return -1;
    629629}
    630630
    631// Put process pid, which must be the caller or one of its
    632// children, into process group pgid.
    633int
    634ksetpgid(int pid, int pgid)
    635{
    636 struct proc *p;
    637 struct proc *me = myproc();
    638
    639 if (pid == 0)
    640 pid = me->pid;
    641 if (pgid == 0)
    642 pgid = pid;
    643 if (pid < 0 || pgid < 0)
    644 return -1;
    645
    646 acquire(&wait_lock); // protects p->parent
    647 for (p = proc; p < &proc[NPROC]; p++) {
    648 if (p != me && p->parent != me)
    649 continue;
    650 acquire(&p->lock);
    651 if (p->pid == pid && p->state != ZOMBIE) {
    652 p->pgid = pgid;
    653 release(&p->lock);
    655 return 0;
    656 }
    657 release(&p->lock);
    658 }
    660 return -1;
    661}
    662
    631663void
    632664setkilled(struct proc *p)
    633665{
    634666 acquire(&p->lock);

    kernel/syscall.c

    @@ -102,8 +102,9 @@ extern uint64 sys_unlink(void);
    102102extern uint64 sys_link(void);
    103103extern uint64 sys_mkdir(void);
    104104extern uint64 sys_close(void);
    105105extern uint64 sys_sync(void);
    106extern uint64 sys_setpgid(void);
    106107
    107108// An array mapping syscall numbers from syscall.h
    108109// to the function that handles the system call.
    109110static uint64 (*syscalls[])(void) = {
    @@ -129,8 +130,9 @@ static uint64 (*syscalls[])(void) = {
    129130 [SYS_link] = sys_link,
    130131 [SYS_mkdir] = sys_mkdir,
    131132 [SYS_close] = sys_close,
    132133 [SYS_sync] = sys_sync,
    134 [SYS_setpgid] = sys_setpgid,
    133135 // clang-format on
    134136};
    135137
    136138void

    kernel/syscall.h

    @@ -20,4 +20,5 @@
    2020#define SYS_link 19
    2121#define SYS_mkdir 20
    2222#define SYS_close 21
    2323#define SYS_sync 22
    24#define SYS_setpgid 23

    kernel/sysproc.c

    @@ -97,8 +97,20 @@ sys_kill(void)
    9797 argint(0, &pid);
    9898 return kkill(pid);
    9999}
    100100
    101// setpgid(pid, pgid): move process pid (0: the caller)
    102// into group pgid (0: a new group numbered pid).
    103uint64
    104sys_setpgid(void)
    105{
    106 int pid, pgid;
    107
    108 argint(0, &pid);
    109 argint(1, &pgid);
    110 return ksetpgid(pid, pgid);
    111}
    112
    101113// return how many clock tick interrupts have occurred
    102114// since start.
    103115uint64
    104116sys_uptime(void)

    user/user.h

    @@ -24,8 +24,9 @@ int getpid(void);
    2424char *sys_sbrk(int, int);
    2525int pause(int);
    2626int uptime(void);
    2727int sync(void);
    28int setpgid(int, int);
    2829
    2930// ulib.c
    3031int stat(const char *, struct stat *);
    3132char *strcpy(char *, const char *);

    user/usys.pl

    @@ -42,4 +42,5 @@ entry("getpid");
    4242entry("sbrk");
    4343entry("pause");
    4444entry("uptime");
    4545entry("sync");
    46entry("setpgid");
  3. b7f9f6d Add killpgrp and the killpg system call

    kernel/defs.h

    @@ -87,8 +87,9 @@ void proc_mapstacks(pagetable_t);
    8787pagetable_t proc_pagetable(struct proc *);
    8888void proc_freepagetable(pagetable_t, uint64);
    8989int kkill(int);
    9090int ksetpgid(int, int);
    91int killpgrp(int);
    9192int killed(struct proc*);
    9293void setkilled(struct proc*);
    9394struct cpu* mycpu(void);
    9495struct proc* myproc();

    kernel/proc.c

    @@ -627,8 +627,34 @@ kkill(int pid)
    627627 }
    628628 return -1;
    629629}
    630630
    631// Kill every process in group pgid, the way kkill() kills
    632// one, and return how many there were. Takes only p->lock,
    633// one at a time, so an interrupt handler may call it.
    634int
    635killpgrp(int pgid)
    636{
    637 struct proc *p;
    638 int n = 0;
    639
    640 if (pgid <= 0)
    641 return 0; // group 0 is "no group": init and the shell
    642 for (p = proc; p < &proc[NPROC]; p++) {
    643 acquire(&p->lock);
    644 if (p->pgid == pgid) {
    645 p->killed = 1;
    646 if (p->state == SLEEPING) {
    647 // Wake process from sleep().
    648 p->state = RUNNABLE;
    649 }
    650 n++;
    651 }
    652 release(&p->lock);
    653 }
    654 return n;
    655}
    656
    631657// Put process pid, which must be the caller or one of its
    632658// children, into process group pgid.
    633659int
    634660ksetpgid(int pid, int pgid)

    kernel/syscall.c

    @@ -103,8 +103,9 @@ extern uint64 sys_link(void);
    103103extern uint64 sys_mkdir(void);
    104104extern uint64 sys_close(void);
    105105extern uint64 sys_sync(void);
    106106extern uint64 sys_setpgid(void);
    107extern uint64 sys_killpg(void);
    107108
    108109// An array mapping syscall numbers from syscall.h
    109110// to the function that handles the system call.
    110111static uint64 (*syscalls[])(void) = {
    @@ -131,8 +132,9 @@ static uint64 (*syscalls[])(void) = {
    131132 [SYS_mkdir] = sys_mkdir,
    132133 [SYS_close] = sys_close,
    133134 [SYS_sync] = sys_sync,
    134135 [SYS_setpgid] = sys_setpgid,
    136 [SYS_killpg] = sys_killpg,
    135137 // clang-format on
    136138};
    137139
    138140void

    kernel/syscall.h

    @@ -21,4 +21,5 @@
    2121#define SYS_mkdir 20
    2222#define SYS_close 21
    2323#define SYS_sync 22
    2424#define SYS_setpgid 23
    25#define SYS_killpg 24

    kernel/sysproc.c

    @@ -109,8 +109,20 @@ sys_setpgid(void)
    109109 argint(1, &pgid);
    110110 return ksetpgid(pid, pgid);
    111111}
    112112
    113// killpg(pgid): kill every process in group pgid.
    114uint64
    115sys_killpg(void)
    116{
    117 int pgid;
    118
    119 argint(0, &pgid);
    120 if (killpgrp(pgid) == 0)
    121 return -1;
    122 return 0;
    123}
    124
    113125// return how many clock tick interrupts have occurred
    114126// since start.
    115127uint64
    116128sys_uptime(void)

    user/user.h

    @@ -25,8 +25,9 @@ char *sys_sbrk(int, int);
    2525int pause(int);
    2626int uptime(void);
    2727int sync(void);
    2828int setpgid(int, int);
    29int killpg(int);
    2930
    3031// ulib.c
    3132int stat(const char *, struct stat *);
    3233char *strcpy(char *, const char *);

    user/usys.pl

    @@ -43,4 +43,5 @@ entry("sbrk");
    4343entry("pause");
    4444entry("uptime");
    4545entry("sync");
    4646entry("setpgid");
    47entry("killpg");
  4. a1c3ff2 Give the console a foreground group, set by setfg

    kernel/console.c

    @@ -52,8 +52,10 @@ struct {
    5252 char buf[INPUT_BUF_SIZE];
    5353 uint r; // Read index
    5454 uint w; // Write index
    5555 uint e; // Edit index
    56
    57 int fg; // foreground process group, for ^C (0: none)
    5658} cons;
    5759
    5860//
    5961// user write() system calls to the console go here.
    @@ -188,8 +190,17 @@ consoleintr(int c)
    188190
    189191 release(&cons.lock);
    190192}
    191193
    194// make group pgid the foreground job (0: none).
    195void
    196consolesetfg(int pgid)
    197{
    199 cons.fg = pgid;
    201}
    202
    192203void
    193204consoleinit(void)
    194205{
    195206 initlock(&cons.lock, "cons");

    kernel/defs.h

    @@ -21,8 +21,9 @@ void bunpin(struct buf*);
    2121// console.c
    2222void consoleinit(void);
    2323void consoleintr(int);
    2424void consputc(int);
    25void consolesetfg(int);
    2526
    2627// exec.c
    2728int kexec(char*, char**);
    2829

    kernel/syscall.c

    @@ -104,8 +104,9 @@ extern uint64 sys_mkdir(void);
    104104extern uint64 sys_close(void);
    105105extern uint64 sys_sync(void);
    106106extern uint64 sys_setpgid(void);
    107107extern uint64 sys_killpg(void);
    108extern uint64 sys_setfg(void);
    108109
    109110// An array mapping syscall numbers from syscall.h
    110111// to the function that handles the system call.
    111112static uint64 (*syscalls[])(void) = {
    @@ -133,8 +134,9 @@ static uint64 (*syscalls[])(void) = {
    133134 [SYS_close] = sys_close,
    134135 [SYS_sync] = sys_sync,
    135136 [SYS_setpgid] = sys_setpgid,
    136137 [SYS_killpg] = sys_killpg,
    138 [SYS_setfg] = sys_setfg,
    137139 // clang-format on
    138140};
    139141
    140142void

    kernel/syscall.h

    @@ -22,4 +22,5 @@
    2222#define SYS_close 21
    2323#define SYS_sync 22
    2424#define SYS_setpgid 23
    2525#define SYS_killpg 24
    26#define SYS_setfg 25

    kernel/sysproc.c

    @@ -121,8 +121,22 @@ sys_killpg(void)
    121121 return -1;
    122122 return 0;
    123123}
    124124
    125// setfg(pgid): make group pgid the console's foreground
    126// job, the one that Ctrl-C interrupts (0: none).
    127uint64
    128sys_setfg(void)
    129{
    130 int pgid;
    131
    132 argint(0, &pgid);
    133 if (pgid < 0)
    134 return -1;
    135 consolesetfg(pgid);
    136 return 0;
    137}
    138
    125139// return how many clock tick interrupts have occurred
    126140// since start.
    127141uint64
    128142sys_uptime(void)

    user/user.h

    @@ -26,8 +26,9 @@ int pause(int);
    2626int uptime(void);
    2727int sync(void);
    2828int setpgid(int, int);
    2929int killpg(int);
    30int setfg(int);
    3031
    3132// ulib.c
    3233int stat(const char *, struct stat *);
    3334char *strcpy(char *, const char *);

    user/usys.pl

    @@ -44,4 +44,5 @@ entry("pause");
    4444entry("uptime");
    4545entry("sync");
    4646entry("setpgid");
    4747entry("killpg");
    48entry("setfg");
  5. 1d9acb9 Kill the foreground group when Ctrl-C is typed

    kernel/console.c

    @@ -6,8 +6,9 @@
    66// control-h -- backspace
    77// control-u -- kill line
    88// control-d -- end of file
    99// control-p -- print process list
    10// control-c -- kill the foreground job
    1011//
    1112
    1213#include <stdarg.h>
    1314
    @@ -153,8 +154,23 @@ consoleintr(int c)
    153154 switch (c) {
    154155 case C('P'): // Print process list.
    155156 procdump();
    156157 break;
    158 case C('C'): // Interrupt the foreground job.
    159 cons.e = cons.w; // drop the line being typed
    160 consputc('^');
    161 consputc('C');
    162 consputc('\n');
    163 if (killpgrp(cons.fg) == 0 && cons.e - cons.r < INPUT_BUF_SIZE) {
    164 // no job to interrupt: hand the reader (the shell)
    165 // an empty line, so that it prompts again.
    166 cons.buf[cons.e++ % INPUT_BUF_SIZE] = '\n';
    167 cons.w = cons.e;
    168 }
    169 // a killed reader may be between sleep_prepare() and
    170 // sleep() in consoleread(); this clears its chan.
    171 wakeup(&cons.r);
    172 break;
    157173 case C('U'): // Kill line.
    158174 while (cons.e != cons.w &&
    159175 cons.buf[(cons.e - 1) % INPUT_BUF_SIZE] != '\n') {
    160176 cons.e--;
  6. 42af66a Run each shell job in its own foreground group

    user/sh.c

    @@ -123,10 +123,12 @@ runcmd(struct cmd *cmd)
    123123 break;
    124124
    125125 case BACK:
    126126 bcmd = (struct backcmd *)cmd;
    127 if (fork1() == 0)
    127 if (fork1() == 0) {
    128 setpgid(0, 0); // leave the foreground job: ^C spares it
    128129 runcmd(bcmd->cmd);
    130 }
    129131 break;
    130132 }
    131133 exit(0);
    132134}
    @@ -145,9 +147,9 @@ getcmd(char *buf, int nbuf)
    145147int
    146148main(void)
    147149{
    148150 static char buf[100];
    149 int fd;
    151 int fd, pid;
    150152
    151153 // Ensure that three file descriptors are open.
    152154 while ((fd = open("console", O_RDWR)) >= 0) {
    153155 if (fd >= 3) {
    @@ -168,11 +170,17 @@ main(void)
    168170 cmd[strlen(cmd) - 1] = 0; // chop \n
    169171 if (chdir(cmd + 3) < 0)
    170172 fprintf(2, "cannot cd %s\n", cmd + 3);
    171173 } else {
    172 if (fork1() == 0)
    174 pid = fork1();
    175 if (pid == 0) {
    176 setpgid(0, 0); // the job gets a group of its own
    173177 runcmd(parsecmd(cmd));
    178 }
    179 setpgid(pid, pid); // here too: whichever runs first
    180 setfg(pid); // ^C now interrupts this job
    174181 wait(0);
    182 setfg(0); // back at the prompt: no foreground job
    175183 }
    176184 }
    177185 exit(0);
    178186}
  7. 11cd0ac Add the pause user program

    Makefile

    @@ -149,8 +149,9 @@ UPROGS=\
    149149 $U/_logstress\
    150150 $U/_forphan\
    151151 $U/_dorphan\
    152152 $U/_sync\
    153 $U/_pause\
    153154
    154155fs.img: mkfs/mkfs README $(UPROGS)
    155156 mkfs/mkfs fs.img README $(UPROGS)
    156157

    user/pause.c

    @@ -0,0 +1,15 @@
    1#include "kernel/types.h"
    2#include "user/user.h"
    3
    4// pause ticks: sleep in the kernel for that many clock ticks
    5// (a tick is 1/10 s). Handy for trying Ctrl-C and &.
    6int
    7main(int argc, char **argv)
    8{
    9 if (argc != 2) {
    10 fprintf(2, "usage: pause ticks\n");
    11 exit(1);
    12 }
    13 pause(atoi(argv[1]));
    14 exit(0);
    15}
  8. 1970c39 Add pgtest, a test of process groups and killpg

    Makefile

    @@ -150,8 +150,9 @@ UPROGS=\
    150150 $U/_forphan\
    151151 $U/_dorphan\
    152152 $U/_sync\
    153153 $U/_pause\
    154 $U/_pgtest\
    154155
    155156fs.img: mkfs/mkfs README $(UPROGS)
    156157 mkfs/mkfs fs.img README $(UPROGS)
    157158

    user/pgtest.c

    @@ -0,0 +1,132 @@
    1#include "kernel/types.h"
    2#include "user/user.h"
    3
    4// Tests process groups and killpg, the kernel half of Ctrl-C.
    5// The keystroke itself (consoleintr, the shell's setfg) can only
    6// be typed from outside: ctrlc-test.py does that.
    7
    8#define NMEMBER 4 // the leader and its three children
    9#define PATIENCE 30 // ticks the members get to die (3 s)
    10
    11int fails;
    12
    13void
    14check(int ok, char *what)
    15{
    16 printf("pgtest: %s: %s\n", what, ok ? "OK" : "FAIL");
    17 if (!ok)
    18 fails++;
    19}
    20
    21// a child of the leader: report our pid, then block in the
    22// kernel in one of three ways until killed.
    23void
    24member(int how, int ready)
    25{
    26 int pid = getpid();
    27 int block[2];
    28 char c;
    29
    30 if (how == 0) {
    31 // keep both ends, so that the read can never see EOF.
    32 pipe(block);
    33 write(ready, &pid, sizeof(pid));
    34 read(block[0], &c, 1); // piperead
    35 } else if (how == 1) {
    36 write(ready, &pid, sizeof(pid));
    37 pause(10000); // sys_pause
    38 } else {
    39 write(ready, &pid, sizeof(pid));
    40 read(0, &c, 1); // consoleread; nobody types
    41 }
    42 exit(0);
    43}
    44
    45int
    46main(void)
    47{
    48 int ready[2], done[2], ping[2], pong[2], wd[2];
    49 int pids[NMEMBER];
    50 int leader, bg, watchdog, pid, i, n, st;
    51 char c;
    52
    53 check(setpgid(1, 1) < 0, "setpgid on a process that is not our child fails");
    54 check(killpg(0) < 0, "killpg(0), the group of init and the shell, is refused");
    55
    56 // a process in another group, which must survive.
    57 pipe(ping);
    58 pipe(pong);
    59 bg = fork();
    60 if (bg == 0) {
    61 setpgid(0, 0);
    62 read(ping[0], &c, 1);
    63 write(pong[1], "!", 1);
    64 exit(0);
    65 }
    66 setpgid(bg, bg);
    67
    68 // the group: a leader that waits, and its three children.
    69 // every member inherits done[1]; done[0] sees EOF when all
    70 // of them have exited.
    71 pipe(ready);
    72 pipe(done);
    73 leader = fork();
    74 if (leader == 0) {
    75 setpgid(0, 0);
    76 close(done[0]);
    77 for (i = 0; i < NMEMBER - 1; i++)
    78 if (fork() == 0)
    79 member(i, ready[1]);
    80 pid = getpid();
    81 write(ready[1], &pid, sizeof(pid));
    82 for (;;)
    83 if (wait(0) < 0)
    84 pause(10000); // only if a child died early
    85 }
    86 setpgid(leader, leader);
    87 close(done[1]);
    88 for (i = 0; i < NMEMBER; i++)
    89 if (read(ready[0], &pids[i], sizeof(int)) != sizeof(int))
    90 pids[i] = -1;
    91
    92 // a watchdog that kills the members one by one with kill()
    93 // if killpg has not got rid of them in time. It writes a
    94 // byte to wd first, so that pgtest can tell that it acted.
    95 pipe(wd);
    96 watchdog = fork();
    97 if (watchdog == 0) {
    98 close(done[0]);
    99 pause(PATIENCE);
    100 write(wd[1], "!", 1);
    101 for (i = 0; i < NMEMBER; i++)
    102 kill(pids[i]);
    103 exit(1);
    104 }
    105 close(wd[1]);
    106
    107 pause(5); // let all four fall asleep
    108 check(killpg(leader) == 0, "killpg finds the group (wait, pipe read, pause, console read)");
    109 n = read(done[0], &c, 1); // EOF once every member has exited
    110 check(n == 0, "every member exited");
    111
    112 kill(watchdog);
    113 st = 0;
    114 for (i = 0; i < 2; i++) {
    115 pid = wait(&n);
    116 if (pid == leader)
    117 st = n;
    118 }
    119 // the watchdog has exited: wd is at EOF unless it acted.
    120 check(read(wd[0], &c, 1) == 0, "...within 3 seconds, without the watchdog's kill()");
    121 check(st == -1, "the leader exited with status -1, like any killed process");
    122
    123 write(ping[1], "?", 1);
    124 check(read(pong[0], &c, 1) == 1, "a process in another group survives");
    125 wait(0);
    126
    127 if (fails == 0)
    128 printf("pgtest: ALL OK (the Ctrl-C keystroke itself is tested by ctrlc-test.py)\n");
    129 else
    130 printf("pgtest: SOME TESTS FAILED\n");
    131 exit(fails != 0);
    132}
  9. 66bb387 Add ctrlc-test.py, which types Ctrl-C at a booted xv6

    ctrlc-test.py

    @@ -0,0 +1,135 @@
    1#!/usr/bin/env python3
    2
    3#
    4# Types Ctrl-C at a booted xv6 and checks what happens.
    5#
    6# ./ctrlc-test.py [kernel/kernel [fs.img]] [-- extra shell commands]
    7#
    8# Boots QEMU directly (3 harts), types commands and \x03 bytes on
    9# its standard input (the UART), and prints one OK/FAIL line per
    10# check. Extra commands after -- are run at the end (for example
    11# pgtest or "usertests -q"), and their output is only shown.
    12#
    13
    14import os, select, subprocess, sys, time
    15
    16args = sys.argv[1:]
    17extra = []
    18if "--" in args:
    19 i = args.index("--")
    20 args, extra = args[:i], args[i + 1:]
    21kernel = args[0] if len(args) > 0 else "kernel/kernel"
    22fsimg = args[1] if len(args) > 1 else "fs.img"
    23
    24q = subprocess.Popen(
    25 ["qemu-system-riscv64", "-machine", "virt", "-bios", "none",
    26 "-kernel", kernel, "-m", "128M", "-smp", "3", "-nographic",
    27 "-global", "virtio-mmio.force-legacy=false",
    28 "-drive", f"file={fsimg},if=none,format=raw,id=x0",
    29 "-device", "virtio-blk-device,drive=x0,bus=virtio-mmio-bus.0"],
    30 stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
    31
    32out = b"" # everything the console printed
    33seen = 0 # how much of out the checks have consumed
    34
    35
    36def pump(t):
    37 global out
    38 r, _, _ = select.select([q.stdout], [], [], t)
    39 if r:
    40 d = os.read(q.stdout.fileno(), 65536)
    41 if d:
    42 out += d
    43 sys.stdout.buffer.write(d)
    44 sys.stdout.flush()
    45
    46
    47def expect(pat, tmo):
    48 """Wait for pat after the last match; True if it came in time."""
    49 global seen
    50 pat = pat.encode()
    51 end = time.time() + tmo
    52 while time.time() < end:
    53 i = out.find(pat, seen)
    54 if i >= 0:
    55 seen = i + len(pat)
    56 return True
    57 pump(0.1)
    58 return False
    59
    60
    61def quiet(t):
    62 end = time.time() + t
    63 while time.time() < end:
    64 pump(0.1)
    65
    66
    67def type_(s):
    68 q.stdin.write(s.encode())
    69 q.stdin.flush()
    70
    71
    72results = []
    73
    74
    75def check(ok, what):
    76 results.append((ok, what))
    77 print(f"\n### ctrlc-test: {what}: {'OK' if ok else 'FAIL'}", flush=True)
    78
    79
    80assert expect("$ ", 60), "no shell prompt"
    81start = time.time()
    82
    83# 1. a job reading the console
    84type_("cat\n")
    85quiet(1)
    86type_("\x03")
    87check(expect("^C\n$ ", 5), "cat waiting for input, then ^C")
    88
    89# 2. a job asleep in pause (100 s)
    90type_("pause 1000\n")
    91quiet(1)
    92t = time.time()
    93type_("\x03")
    94ok = expect("^C\n$ ", 5)
    95check(ok, f"pause 1000, then ^C (prompt after {time.time() - t:.3f} s)")
    96
    97# 3. a pipeline: cat in consoleread, grep in piperead, their parent in wait
    98type_("cat | grep x\n")
    99quiet(1)
    100type_("\x03")
    101check(expect("^C\n$ ", 5), "cat | grep x, then ^C")
    102
    103# 4. a background job survives a ^C for the foreground job
    104type_("(pause 30; echo survived) &\n")
    105check(expect("$ ", 5), "background job started")
    106type_("cat\n")
    107quiet(1)
    108type_("\x03")
    109check(expect("^C\n$ ", 5), "cat, then ^C, with a background job running")
    110check(expect("survived\n", 10), "the background job finished its work")
    111
    112# 5. ^C at the prompt: no job, the shell must not die
    113type_("\x03")
    114check(expect("^C\n$ ", 5), "^C at an empty prompt gives a new prompt")
    115type_("echo jun")
    116quiet(0.5)
    117type_("\x03")
    118quiet(0.5)
    119type_("echo clean\n")
    120# the reference prompts straight after the ^C line; had the line
    121# survived, the shell would have run "echo jun" and printed jun there.
    122check(expect("echo jun^C\n$ ", 5) and expect("clean\n$ ", 5), "^C drops a half-typed line")
    123type_("echo alive\n")
    124check(expect("alive\n$ ", 5), "the shell still runs commands")
    125check(out.count(b"init: starting sh") == 1, "the shell was never restarted")
    126
    127for c in extra:
    128 type_(c + "\n")
    129 expect("$ ", 3600)
    130
    131q.kill()
    132bad = [w for ok, w in results if not ok]
    133print(f"\nctrlc-test: {len(results) - len(bad)} of {len(results)} checks OK", flush=True)
    134print("ctrlc-test: " + ("ALL OK" if not bad else "SOME CHECKS FAILED"))
    135sys.exit(1 if bad else 0)

6. Verify and measure

On the branch (ext/06-ctrl-c, 9 commits), run on 3 harts (-smp 3 -m 128M) by ctrlc-test.py kernel fs.img -- pgtest "usertests -q" pgtest, all in one boot:

$ cat
^C
$
### ctrlc-test: cat waiting for input, then ^C: OK
pause 1000
^C
$
### ctrlc-test: pause 1000, then ^C (prompt after 0.001 s): OK
cat | grep x
^C
$
### ctrlc-test: cat | grep x, then ^C: OK
[...]
### ctrlc-test: the shell was never restarted: OK
pgtest
pgtest: setpgid on a process that is not our child fails: OK
[...]
pgtest: ALL OK (the Ctrl-C keystroke itself is tested by ctrlc-test.py)
$ usertests -q
usertests starting
test copyin: OK
test copyout: OK
[...]
ALL TESTS PASSED
$ pgtest
[...]
pgtest: ALL OK (the Ctrl-C keystroke itself is tested by ctrlc-test.py)
$
ctrlc-test: 10 of 10 checks OK
ctrlc-test: ALL OK

(The ### ctrlc-test: lines are the driver’s own, printed between the console output it passes through.) Two more boots with the same commands gave the same result: 10 of 10, pgtest ALL OK twice, ALL TESTS PASSED.

What it shows: Ctrl-C reaches a reader of the console, a sleeper in pause, and all three processes of a pipeline; a background job lives through it; at the prompt it only produces a new prompt; and the shell is never restarted. pgtest shows that killpg wakes sleepers in kwait, piperead, sys_pause and consoleread without help. usertests -q passing shows that nothing else changed, including kill itself (killstatus, kill-based tests) and every program that forks and waits, now through a shell that calls three new system calls per command.

The same ctrlc-test.py against the original kernel: ctrlc-test: 1 of 10 checks OK (only “the shell was never restarted”). ^C is stored as the byte 0x03 and the first cat is never interrupted, so it reads every line typed after it: the driver’s echo alive came back twice, once as the console’s echo and once printed by cat.

Every commit builds on its own (make kernel/kernel fs.img, only the usual RWX link warning).

How fast does a Ctrl-C take effect? Time from writing \x03 to QEMU’s standard input until the next $ came back, measured by the driver on the host (no gdb attached; QEMU 10.2.1, 3 harts):

foreground job where the victims were prompt after
cat asleep in consoleread 0.001 s
pause 1000 asleep in sys_pause (would have taken 100 s) 0.001 s
cat | grep x kwait, consoleread, piperead 0.006 s
spin (a loop in user mode, no system calls; not on the branch) running on another hart 0.001 s, three times

Everything is bounded by the time the victims need to be scheduled and exit, and the shell to be woken: a few milliseconds of emulated time. The surprise is spin. It is running when the kill arrives and only dies at its next trap; one would expect that to be its hart’s next timer interrupt, up to 1/10 s later. gdb (four Ctrl-Cs) showed something else: the Ctrl-C was always taken by an idle hart, and spin died on another hart a moment later at an external-interrupt trap (scause 0x8000000000000009), in the same tick. An instrumented devintr showed that in 5 of 5 kills, spin’s hart took that trap but its plic_claim() returned 0. The PLIC signals the interrupt to every hart, another hart claimed and handled the UART, and the empty trap was still enough to bring spin into usertrap's killed test. Do not rely on it: the timer is the guarantee.

The original kernel. The same keystroke on 06aad25 (control bytes shown as cat -v shows them; ^C here is the single byte 0x03, echoed back by the UART, not the two characters the branch prints):

$ cat
^C
^C
^D$ echo after
after

cat waits for Enter, then prints the line holding the 0x03 byte, and only Ctrl-D ends it. And with spin (added for this measurement), Ctrl-C changes nothing; five seconds later Ctrl-P still lists 3 run spin. There is no way back to the prompt but a reboot.

The windows, made wide on purpose (scratch copies with an added delay, not the branch):

7. Go further