Tour 50 · Locks and interrupt state · about 37 minutes · 22 steps
Every hart keeps three small facts about interrupts. SIE is one bit in sstatus: may an
interrupt be taken right now? noff (mycpu()->noff) counts how many push_off calls
are still open. intena (mycpu()->intena) remembers whether SIE was on just before the
outermost push_off. Together they decide, at every release, whether interrupts come
back on (Locks and interrupt state, Locks and interrupt state).
This tour watches those three values on hart 0, instruction by instruction, through
one real stretch of time. The shell goes to sleep in read. The scheduler takes over and
resets intena. A timer tick lands in the scheduler. A process that was preempted on
another hart resumes here, inside a trap handler, and leaves with interrupts still off,
as it must. A brand-new process starts in forkret. Finally a tick lands inside that
process’s kernel code and shows how the preempted process got into its state in the first
place.
Nothing here is invented. We booted this build on three harts, typed grind & (a test
program that runs random system calls in several processes forever), and attached gdb to
QEMU. Breakpoints on hart 0 recorded sstatus, cpus[0].noff (offset 120 of struct cpu),
cpus[0].intena (offset 124), sp and the relevant registers at each instruction we
quote. Every number below comes from that run, or from the build’s kernel/kernel.asm.
Tour 13: swtch and the lock handed across a context switch and Tour 15: Spinlocks from the hardware up explained the code; this tour is about the values.
Best after: 8. Traps taken inside the kernel, 11. From a timer tick to a context switch, 13. swtch and the lock handed across a context switch, 15. Spinlocks from the hardware up, 16. sleep and wakeup, and the lost-wakeup problem
The run starts at tick 12, about 1.2 s after boot. grind & has just been typed.
| Hart | What it is doing |
|---|---|
| 0 | Running sh (pid 2, slot proc[1]): it printed $ and is about to call read(0, &c, 1) |
| 1 | In its scheduler loop, idle |
| 2 | Running pid 4 (slot proc[3]), the shell’s background grandchild, inside exec("grind"): still named sh (pid 3, the forked copy of sh that started it, has already exited) |
sh uses KSTACK(1) (top 0x3fffffc000), pid 4 uses KSTACK(3) (top 0x3fffff8000).
Hart 0’s scheduler stack is its slice of stack0, 0x80007890–0x80008890.
In the state strip, intena is shown only while noff > 0. When noff is 0 the field
still holds a stale value; we mention it where it matters.
Step 1 of 22
sh executed ecall for read. The trap cleared SIE, so usertrap starts with
interrupts off. gdb at the csrsi sstatus,2 of line 66 (0x80002668) shows:
sstatus |
SIE | noff | intena | |
|---|---|---|---|---|
| before line 66 | 0x200000020 |
0 | 0 | 0 (stale) |
| after line 66 | 0x200000022 |
1 | 0 | 0 (stale) |
Bit 1 (0x2) is SIE. Bit 5 (0x20) is SPIE, the copy of SIE the trap saved; the high
0x200000000 is the UXL field and never changes.
Look at the stale intena. It is 0 because myproc on line 49, and then the
acquire inside killed on line 57, each ran an outermost push_off while SIE was
still off, and a push_off at noff == 0 always overwrites intena. That is the first lesson of this tour: intena is rewritten
every time noff goes from 0 to 1. It is not a long-lived setting; it is a note about
the latest outermost push_off.
Why turn interrupts on at all? A system call may run for a long time, and while it does,
the hart should still take timer and device interrupts. Line 66 comes after line 52
saved sepc, because an interrupt would overwrite it (Tour 8: Traps taken inside the kernel).
Step 2 of 22
syscall calls myproc to find the trapframe. myproc brackets its read of
c->proc with push_off/pop_off, so that a timer interrupt cannot move this thread
to another hart between “which hart am I?” and “read that hart’s proc”.
This is the first push_off with SIE on. gdb at the csrrci a5,sstatus,2
(0x80000bb8):
sstatus |
a5 |
noff | intena | |
|---|---|---|---|---|
| before | 0x200000022 |
0 | 0 | |
after csrrci |
0x200000020 |
0x200000022 |
0 | 0 |
end of push_off |
0x200000020 |
1 | 1 |
csrrci read the old sstatus into a5 and cleared SIE in the same instruction. Since
noff was 0, intena took the old SIE bit: 1. Then pop_off took noff back to 0,
saw intena == 1, and executed csrsi sstatus,2 (0x80000c4c): SIE on again.
Before consoleread was even reached, the same pair ran four more times (three in
argraw, via argaddr, argint and argfd’s argint, and one in argfd
itself), each time turning SIE off and back on. A
push_off level is not always a lock. myproc and uartputc_sync use it too, which
is why noff can be larger than the number of locks held.
cons.lockStep 3 of 22
consoleread was entered with SIE on (sstatus = 0x200000022, a2 = 1: sh reads
one byte at a time). Its first act is acquire(&cons.lock), and acquire's first act
is push_off:
0x80000bb8 csrrci a5,sstatus,2 a5 = 0x200000022 sstatus -> 0x200000020
0x80000bc2 lw a5,120(a0) noff was 0, so ...
0x80000be4 sw a5,124(a0) intena = 1 (cpus[0] + 124 = 0x8000fa4c)
0x80000bce sw a5,120(a0) noff = 1 (cpus[0] + 120 = 0x8000fa48)
Only then does the amoswap.w.aq spin take cons.lock (at 0x8000f890). The order
matters: interrupts are off before the lock is held, so there is no instant where
this hart holds cons.lock with SIE on. If there were, a keystroke interrupt could
arrive, consoleintr would try to take cons.lock on the same hart, and spin
forever, because the holder is the code it interrupted
(Locks and interrupt state).
From here on, intena = 1 records one fact: “when sh’s thread started holding things
on this hart, interrupts were on”. Every release below consults it.
cons.lockStep 4 of 22
The buffer is empty (cons.r == cons.w), so line 100 calls killed(myproc()).
myproc’s push_off now runs inside the critical section:
a5 (old sstatus) |
noff | intena | |
|---|---|---|---|
csrrci |
0x200000020 (SIE already 0) |
1 | 1 |
end of push_off |
2 | 1 (untouched) | |
end of pop_off |
1 | 1 |
This is why push_off only writes intena when noff == 0. The old SIE here is 0,
because acquire already cleared it. If the nested call recorded that 0, the final
release would leave interrupts off for the rest of the system call. And this pop_off
at noff 2 → 1 must not turn interrupts on, or cons.lock would be held with SIE on.
The counter handles both.
Then killed acquires sh’s p->lock: noff 2 again, this time with two real
locks, and back to 1 at its release.
cons.locksh's p->lockStep 5 of 22
sleep_prepare registers sh on the channel &cons.r. First a myproc (noff
1 → 2 → 1, as before). Then acquire(&p->lock) for sh’s lock at 0x8000ff38:
csrrci returns a5 = 0x200000020: SIE was already off.noff becomes 2. intena stays 1.Two spinlocks are held, cons.lock and sh’s p->lock, in that order, which matches
the measured edge cons.lock → p->lock (Locks and interrupt state). p->chan is set,
and the release at line 555 drops noff to 1. Its pop_off sees noff == 1 and
leaves SIE alone.
Why is it safe to register before letting go of cons.lock? Because a
wakeup that runs any time after this point finds chan == &cons.r and clears it,
so the coming sleep will return at once instead of sleeping through it
(Locks and interrupt state).
Step 6 of 22
Line 105 of consoleread releases cons.lock. release stores 0 into the lock word
(after fence rw,w) and then calls pop_off:
0x80000c42 addiw a5,a5,-1 noff 1 -> 0
0x80000c48 lw a5,124(a0) intena = 1
0x80000c4c csrsi sstatus,2 sstatus 0x200000020 -> 0x200000022
This is the only way a release turns interrupts on: the counter reaches 0 and
intena says they were on before. From here until the next acquire, sh can be
interrupted. Notice what the window is for: if the Enter key (completing a line) arrives
now, the UART interrupt can run on this very hart, take cons.lock, and call
wakeup(&cons.r), which
clears sh’s chan. The coming sleep() would then return at once.
Note the address 0x80000c50, the instruction right after this csrsi. It will come
back three times in this tour.
sh's p->lockStep 7 of 22
sleep calls myproc (another 0 → 1 → 0, SIE off and on), then acquire(&p->lock).
gdb at that push_off:
sstatus = 0x200000022, noff = 0csrrci: a5 = 0x200000022, so the old SIE is 1noff = 1, intena = 1intena was already 1, but it was written again: this is a new outermost
push_off. The value that matters for the coming switch is this one, recorded with
sh’s p->lock as the only lock.
p->chan is still &cons.r (no key yet), so sh becomes SLEEPING and calls
sched. It holds exactly one spinlock, its own p->lock, which the scheduler will
release for it (Locks and interrupt state).
sh's p->lockStep 8 of 22
sched's own myproc() takes noff to 2 and back to 1 (the trace shows both). Then
the checks, as compiled:
lw a4,168(a5) at 0x80001e54 loads cpus[0].noff (168 is 120 plus the
48 bytes from pid_lock, the base the compiler uses, to cpus). It is 1. Had
consoleread called sleep while still holding cons.lock, it would be 2 and the
kernel would panic("sched locks").csrr a5,sstatus at 0x80001e66 reads 0x200000020: SIE is 0, so no
panic("sched interruptible").The noff == 1 rule is the code version of “never sleep holding a spinlock”. A sleeping
thread cannot release anything, so a second lock would stay held, with interrupts off on
whatever hart wants it, for as long as sh sleeps. Sleep-locks are not counted in
noff, which is why they may be held across a sleep (Locks and interrupt state).
sh's p->lockStep 9 of 22
Line 494 compiles to one load:
0x80001e7e lw s3,172(a5) a5 = 0x8000f9a0, so this reads 0x8000fa4c = cpus[0].intena
s3 = 1
0x80001e98 jal swtch swtch(&p->context, &cpus[0].context)
s3 is a callee-saved register, so swtch stores it into sh’s p->context
along with ra and sp. From this moment, sh’s “interrupts were on” no longer lives
in cpus[0]. It lives in sh’s saved context, and it will come back wherever and
whenever sh runs again, possibly on hart 1 or 2.
noff is not saved. It is 1 on both sides of every swtch, because the one lock held
is the one being handed over. The comment above sched (kernel/proc.c:472) says
why intena is not simply a field of struct proc: some code holds locks with no
process at all (the scheduler, boot).
stack0sp = 0x80008810, hart 0’s slice of stack0ld sp, 8(a1) in swtch (kernel/swtch.S:26)sh's p->lockStep 10 of 22
swtch returns into the scheduler at 0x80001df0. gdb, before and after line 456
(sw zero,172(a5) at 0x80001df8):
| noff | intena | SIE | |
|---|---|---|---|
after swtch |
1 | 1 (left by sh) |
0 |
| after line 456 | 1 | 0 | 0 |
cpus[0] is a per-hart structure, so the scheduler inherits whatever the last thread
left there. Here that is sh’s 1. The scheduler’s own truth is different: its
interrupts are off on purpose (line 442). Without line 456, the release at line 463
would see noff 0 and intena 1 and execute csrsi, turning interrupts on in the
middle of the scan. The later acquires in that same pass would then record intena = 1,
and a new process picked later in that pass would carry the wrong 1 into forkret
(compare step 18). Lines 441–442 reset SIE before the next pass, so the damage lasts
one pass at a time.
So there are two safeguards, one for each direction:
intena;sched's save and restore protects each process from the scheduler’s intena.stack0sp = 0x80008810Step 11 of 22
c->proc = 0, then release(&p->lock) on sh’s lock. sh is now SLEEPING (state
2 in our trace). The release’s pop_off takes noff from 1 to 0, finds intena == 0,
and skips the csrsi. SIE stays 0.
One lock was acquired by sh’s thread in sleep and released by the scheduler’s
thread. holding() accepted it because the lock records the hart, cpus[0], not the
thread (Locks and interrupt state).
The scan goes on through the remaining 62 slots (2 to 63). Each acquire in the scan is
an outermost push_off made with SIE off, so each one records intena = 0. Slot 3,
pid 4, is RUNNING (on hart 2). Nothing else is RUNNABLE. found is already 1 (line
461, because sh ran in this pass), so this pass skips wfi; the next pass finds
nothing and reaches it.
stack0sp = 0x80008810Step 12 of 22
The loop goes round. In this build the instructions are laid out as:
0x80001e08 wfi only if found == 0
0x80001e0c csrsi sstatus,2 line 441: sstatus 0x200000020 -> 0x200000022
0x80001e10 csrci sstatus,2 line 442: back to 0x200000020
gdb saw the next pass: SIE 0, then 1 for exactly one instruction, then 0, then the scan,
then wfi. wfi waits for an interrupt to become pending; it does not need SIE on,
and with SIE off nothing is taken while it waits. When it returns, the csrsi opens the
window and a pending interrupt is taken right there.
Why off around wfi? If interrupts were on, a device interrupt could be handled
between the scan and the wfi and make a process RUNNABLE; the hart would then sleep
until the next interrupt with work waiting. Why on at all? So that, with every process
asleep, the interrupt that wakes one of them can be taken. And note noff here: 0.
This window is where the invariant is easiest to see: SIE is
only ever 1 when noff is 0.
stack0a 256-byte kernelvec frame on top: sp = 0x800086e0 in kerneltraptickslockStep 13 of 22
The wfi returned, the csrsi at 0x80001e0c ran, and the timer interrupt was taken
before the next instruction: kerneltrap reads sepc = 0x80001e10, scause = 0x8000000000000005 (supervisor timer), sstatus = 0x200000120. That sstatus now has
SPP set (0x100, came from S-mode) and SPIE set (0x20, SIE was 1 when the trap hit).
kernelvec's addi sp,sp,-256 pushed its frame onto the scheduler stack: there
is no separate interrupt stack.
At clockintr’s acquire(&tickslock) gdb shows noff = 0, intena = 0 (stale), SIE
0. The push_off reads an sstatus whose SIE is 0, so it records intena 0 and
noff becomes 1. Hart 0 is the only hart that counts time: ticks goes from 12 to 13,
and wakeup scans all 64 p->locks at noff 2.
Two facts about interrupt context, both visible here:
noff == 0. It has to: SIE can only be 1 when noff is 0. In a
separate run of grind (about a minute of wall-clock time, slowed by gdb) with a
breakpoint on kerneltrap, gdb counted 43,104 kernel-mode traps on the three
harts, and none of those whose noff it could read had noff ≠ 0 (one read
failed). Most hit idle schedulers; 2,103 interrupted a process and 929 were timer
ticks. In the tour’s own run, all 310 kernel traps checked were at noff 0 too.intena = 0, so its releases never turn
interrupts on inside the handler.myproc() is 0 here, so there is no yield. kernelvec’s sret returns to
0x80001e10 with SIE restored from SPIE, and the csrci turns it off again.
stack0sp = 0x80008810pid 4's p->lockStep 14 of 22
The next scan finds slot 3 RUNNABLE. Pid 4 had just been preempted on hart 2: at
hart 2’s own tick, a timer interrupt hit pid 4 in kernel code, and kerneltrap called
yield. Our breakpoint on sched recorded that suspension: path kerneltrap → yield → sched, saved s3 = 0.
Hart 0’s acquire(&p->lock) for pid 4: push_off at noff 0 with SIE 0, so
cpus[0].intena = 0, noff = 1. gdb confirms noff 1, intena 0. Line 451 sets
RUNNING, line 452 c->proc, line 453 calls swtch.
Remember this 0. Whenever the scheduler switches to any process, cpus[i].intena is
0, because the scheduler’s own acquire just wrote it with interrupts off. That is the
value every resumed thread would see if sched did not restore its own.
ld sp, 8(a1) in swtch (kernel/swtch.S:26)pid 4's p->lockStep 15 of 22
swtch loads pid 4’s context and returns to 0x80001e9c, after the swtch call in
sched. gdb: s3 = 0, cpus[0].intena = 0 before line 496, and 0 after
sw s3,172(s2) (0x80001ea4).
Here are the values side by side:
cpus[0].intena on arrival |
saved in s3 |
correct value | |
|---|---|---|---|
sh (slept in a system call) |
— | 1 | 1 |
pid 4 (preempted in kerneltrap) |
0 (scheduler’s acquire) |
0 | 0 |
For pid 4 the restore writes 0 over 0. Pid 4 does not depend on line 496: without it,
pid 4 would get the scheduler’s 0, which is right, thanks to line 456 (and the
intr_off at line 442). A thread that yielded from a trap is protected by either
safeguard. The thread that
line 496 really rescues is sh. When sh resumes, it will find the scheduler’s 0 in
cpus[i].intena. Without the restore, sleep’s release would leave SIE off, and the
rest of sh’s read (and every later acquire, which would re-record 0) would run with
interrupts off until the return to user mode.
For pid 4 to inherit a wrong 1, both safeguards would have to be missing: no line
456 (so sh’s 1 survives and the rest of that scan runs with SIE on, its acquires
recording 1) and no line 496, and pid 4 would have to be picked in that same scan
pass. In this run it was not: pid 4 was picked two passes later, after lines 441–442
had turned SIE off again. Had it been, the next step would turn interrupts on inside a
trap handler.
Step 16 of 22
yield ends with release(&p->lock). pop_off: noff 1 → 0, intena 0, no
csrsi. gdb at the end of that pop_off: sstatus = 0x200000020, SIE 0.
Back in kerneltrap, lines 162–163 put back the trap registers it saved on entry
on hart 2. They survived the switch on pid 4’s kernel stack, in the saved-register slots
of yield’s and sched’s frames, and were reloaded into s2 and s1 on the way back:
0x80002708 csrw sepc,s2 sepc = 0x80000c50
0x8000270c csrw sstatus,s1 sstatus = 0x200000120 (SPP=1, SPIE=1, SIE=0)
Why must SIE still be 0 here? (Reasoned from the code, not observed.) An interrupt taken
before line 162 would be harmless: lines 162–163 rewrite both registers afterwards. After
csrw sstatus,s1, SIE is 0 again from the saved value. The one exposed point is the
boundary between the two csrws: an interrupt there would set sepc to 0x8000270c
itself, and kernelvec’s sret would later jump back into the middle of kerneltrap
with its frame already popped. The comment at lines 160–161 shows the authors knew
yield may let traps happen; intena = 0 keeps that last boundary closed.
Step 17 of 22
kernelvec restores the caller-saved registers and executes sret (0x800055fc), with
sepc = 0x80000c50. sret sets the mode from SPP (S), sets SIE from SPIE (1), and
jumps to sepc. gdb after the sret: pc = 0x80000c50, sstatus = 0x200000022,
SIE 1, noff 0.
0x80000c50 is the instruction right after pop_off’s csrsi. gdb’s backtrace from
there: pop_off ← release ← releasesleep ← brelse ← readi ← dirlookup ← namex ← namei ← kexec ← sys_exec. So on hart 2, pid 4 was releasing a buffer’s sleep-lock while
looking up grind; the release of the sleep-lock’s inner spinlock re-enabled
interrupts; the pending tick was taken on the very next instruction; and pid 4 now
resumes that instruction on hart 0, with SIE on, exactly as if nothing had happened.
Interrupts came back on only here, by sret, at noff 0, in the code that had them on
before the trap. That is the whole contract: intena 0 kept the handler closed, and
SPIE carried the “they were on” fact across the trap instead.
ld sp, 8(a1) in swtch from hart 0’s schedulerpid 7's p->lockStep 18 of 22
Tick 14. grind has forked; hart 0’s scheduler picks the child, pid 7 in slot 5, and
its first swtch lands in forkret on an empty kernel stack (Tour 13: swtch and the lock handed across a context switch). gdb at
entry: noff = 1, intena = 0, sstatus = 0x200000020.
There is no sched frame to restore anything: pid 7 has never run. The intena it
sees is simply the one the scheduler’s acquire recorded, 0. myproc() takes noff
to 2 and back. Then line 520’s release: noff 1 → 0, intena 0, no csrsi. gdb
after it: SIE 0, noff 0.
That is the right answer. forkret is about to call prepare_return, which turns
interrupts off anyway (line 108 of trap.c) before changing stvec; sret in
userret will turn them on in user mode through SPIE. If line 456 were missing,
this is the release that would act on a leftover 1: interrupts would come on for the
few instructions before prepare_return turns them off. Mostly harmless here; the real
damage would have been done earlier, in the scheduler’s scan.
Step 19 of 22
Three ticks later pid 7 is in a system call. The timer fires on hart 0 and
kerneltrap reads sepc = 0x80000c50 (once again the instruction after
pop_off’s csrsi), sstatus = 0x200000120, scause = 0x8000000000000005. gdb at
entry: noff = 0, SIE 0, and a stale intena = 1, written by the last outermost
push_off of pid 7’s system call, when SIE was on.
kerneltrap checks SPP (it came from S-mode) and checks that SIE is off (the hardware
cleared it). The stale 1 does no harm, because nothing reads intena until a
pop_off reaches noff 0, and the next push_off at noff 0 will overwrite it.
We did not stop on the instructions before the trap, so we cannot say which release
pid 7 was in. What we can say is what the invariant guarantees: the trap arrived at
noff 0, so pid 7 held no spinlock when it was interrupted.
tickslockinit's p->lock (the first of 64)Step 20 of 22
acquire(&tickslock) (the lock named time, at 0x800157d0). gdb at its push_off:
a5 from csrrci |
noff | intena | |
|---|---|---|---|
| before | 0 | 1 (stale) | |
| after | 0x200000120 (SIE 0) |
1 | 0 |
The stale 1 is gone: an outermost push_off in a handler records the handler’s truth.
ticks goes from 17 to 18 and wakeup walks all 64 slots, each acquire taking
noff to 2 and each release back to 1, intena 0 throughout. gdb saw init
(pid 1), sh (pid 2) and grind pid 5 SLEEPING among them.
This is the pair the whole mechanism exists for. sys_pause holds tickslock in
process context. If a timer interrupt could arrive on the same hart while it does,
clockintr would spin on tickslock forever, since its holder cannot run until the
handler returns. push_off in sys_pause’s acquire makes that impossible. The evidence: none of the
2,103 measured traps that interrupted process code, nor the 310 checked in this run,
arrived at noff above 0.
The release(&tickslock): noff 1 → 0, intena 0, SIE stays off.
pid 7's p->lockStep 21 of 22
myproc() is pid 7, so kerneltrap calls yield. Its acquire(&p->lock) (lock at
0x800104d8) is another outermost push_off with SIE 0: noff 1, intena 0. In
sched, noff is exactly 1, SIE is 0, and lw s3,172(a5) loads 0. gdb at the
swtch call: s3 = 0.
This is precisely the state pid 4 was in on hart 2 before step 15: suspended in
sched under kerneltrap, with a saved intena of 0 and saved sepc/sstatus in
kerneltrap’s s2/s1. When pid 7 resumes, on whatever hart, steps 15–17 will replay
for it: restore 0, release with SIE off, csrw sepc, csrw sstatus, sret, SIE on.
Compare the two suspensions on hart 0 in this tour:
how it got to sched |
s3 |
after it resumes, release does |
|
|---|---|---|---|
sh |
system call, sleep |
1 | csrsi: SIE on, system call continues |
| pid 7 | timer in kernel, yield |
0 | nothing: SIE stays off until sret |
stack0sp = 0x80008810pid 7's p->lockStep 22 of 22
Hart 0’s values, from gdb, step by step:
| Step | Where | SIE | noff | intena |
|---|---|---|---|---|
| 1 | usertrap line 66 |
0 → 1 | 0 | (0, stale) |
| 2 | myproc in syscall |
0 | 1 | 1 |
| 3 | acquire(&cons.lock) |
0 | 1 | 1 |
| 4 | myproc inside it |
0 | 2 | 1 |
| 5 | sleep_prepare |
0 | 2 | 1 |
| 6 | release(&cons.lock) |
1 | 0 | (1) |
| 7 | sleep’s acquire |
0 | 1 | 1 (rewritten) |
| 8–9 | sched, s3 = 1 |
0 | 1 | 1 |
| 10 | scheduler line 456 | 0 | 1 | 1 → 0 |
| 11 | release sh’s lock |
0 | 0 | (0) |
| 12 | lines 441–442 | 1 for one instruction | 0 | (0) |
| 13 | clockintr, tickslock |
0 | 1 | 0 |
| 14 | acquire pid 4’s lock |
0 | 1 | 0 |
| 15 | sched restores s3 = 0 |
0 | 1 | 0 |
| 16 | yield’s release |
0 | 0 | (0) |
| 17 | sret |
1 | 0 | (0) |
| 18 | forkret release |
0 | 1 → 0 | 0 |
| 19 | tick in pid 7 | 0 | 0 | (1, stale) |
| 20 | clockintr, wakeup |
0 | 1, 2 | 0 |
| 21 | yield, sched, s3 = 0 |
0 | 1 | 0 |
Read the SIE column: it is 1 only on rows where noff is 0, and it was only ever turned
on by four instructions: usertrap’s csrsi, pop_off’s csrsi at noff 0 with
intena 1, the scheduler’s csrsi, and sret. Read the intena column: it changes
only at an outermost push_off, at line 456, and at line 496. And read the two s3
values: 1 for a thread that went to sleep with interrupts on, 0 for one that was
preempted in a handler. The hart’s field is a scratch copy; the thread’s truth travels in
its saved registers (Locks and interrupt state).
Tour 50 · wrap-up
| Lock | Taken in | Protects |
|---|---|---|
cons.lock | consoleread (step 3), released before sleep (step 6) | The console input buffer cons.buf, cons.r, cons.w; also taken by consoleintr, so it must never be held with SIE on |
sh's p->lock | killed, sleep_prepare, sleep (acquired by sh); released by hart 0’s scheduler (step 11) | sh’s chan and state; held across swtch as the hand-off |
pid 4's p->lock | hart 0’s scheduler (step 14); released by yield (step 16) | Pid 4’s state, so only one hart can start it after it moved from hart 2 |
pid 7's p->lock | hart 0’s scheduler → released in forkret (step 18); yield in step 21 → released by the scheduler | Pid 7’s state across both of its switches on hart 0 |
tickslock | clockintr on hart 0 only, in interrupt context (steps 13 and 20) | ticks; shared with sys_pause and sys_uptime in process context, which is why it may never be held with SIE on |
every p->lock, in turn | wakeup inside clockintr, at noff 2 | Each process’s chan and state |
(no lock) cpus[i].noff and cpus[i].intena | push_off, pop_off, sched, scheduler line 456 | Per-hart fields: only the owning hart touches them, always with SIE off, so no lock is needed |
(not a lock) myproc's push_off | myproc, everywhere (steps 2, 4, 7, 8, 18) | Nothing shared: it keeps the thread on one hart while it reads c->proc, and raises noff by one |
In step 4 the nested push_off saw an old SIE of 0. Why must it not write that 0 into intena, and what would the user notice if it did?
intena must describe the state before the outermost push_off, which was 1. If the nested call stored 0, the final release(&cons.lock) in step 6 would skip the csrsi, and sh would continue its system call, and every later outermost push_off would re-record 0, with interrupts off until it returned to user mode. Timer preemption and device interrupts on that hart would be delayed for the rest of the system call.
Steps 9 and 21 both end in a call to swtch from sched, with noff = 1. Why does s3 differ, and where does each value come from?
s3 is loaded from cpus[0].intena, which the last outermost push_off wrote. In step 9 the value came from sleep’s acquire (step 7), run in a system call with SIE on, so 1. In step 21 it was yield’s acquire, run inside kerneltrap with SIE off, so 0.
Remove only line 496 (sched’s restore). Which of this tour’s three threads (sh, pid 4, pid 7) behave differently, and how?
Of the suspensions shown, only sh’s. Every thread resumes with the 0 the scheduler’s acquire recorded. That is correct for pid 4 and pid 7 in the suspensions shown here, which saved 0 anyway; but any thread that sleeps in a system call is affected like sh (pid 4 and pid 7 also do, later in this run, with s3 = 1). sh saved 1; it would resume with 0, sleep’s release would leave SIE off, and the rest of sh’s read would run with interrupts off.
Remove both line 456 and line 496. Trace what happens in steps 10–16 with the values you saw.
In step 10 cpus[0].intena stays 1 (sh’s). The release in step 11 then executes csrsi, and the rest of that scan runs with SIE on, each acquire recording 1; a process picked later in that same pass would get 1. Pid 4 is not: lines 441–442 turn SIE off before the next pass, so its acquire in step 14 still records 0. Had pid 4 been RUNNABLE further along that first pass, then with no restore its release in yield would turn SIE on inside kerneltrap, and an interrupt between csrw sepc and csrw sstatus would leave sepc pointing into kerneltrap, and kernelvec’s sret would jump there with the frame already popped.
gdb counted 43,104 kernel-mode traps and found none at noff > 0, and most of them at 0x80001e10 or 0x80000c50. Explain both facts.
An interrupt can be taken in S-mode only while SIE is 1 (a kernel exception would panic), and SIE is 1 only when noff is 0 (every push_off clears it first; only pop_off at noff 0, usertrap’s intr_on, the scheduler’s intr_on and sret set it). An interrupt that becomes pending while SIE is off waits, so it is delivered on the first instruction after SIE comes back on: right after the scheduler’s csrsi (0x80001e10) or pop_off’s csrsi (0x80000c50).
Why does forkret need no saved intena, and why is 0 the right value for it?
A new process has no sched frame to restore from; it inherits whatever the scheduler’s acquire recorded, which is always 0 because the scheduler acquires with SIE off. 0 is right because forkret goes straight to prepare_return, which turns interrupts off before changing stvec, and sret turns them on in user mode through SPIE.
Keys: ← → step · Home start