Lab 25 · reveal · 19 steps · 7 commits
xv6 already has one device driver that works by DMA: the virtio disk, where the kernel writes descriptors into memory, pokes a register, and an interrupt says the device is done (Tour 29: A disk read, end to end). In this lab you write a second one, for QEMU’s emulation of the Intel e1000 network card, and put a small network stack and a socket interface on top of it, so that a user program can exchange UDP packets with a program on the host.
The disk is a polite device: it only speaks when spoken to, and exactly one request waits for each answer. A network card is not. Packets arrive whenever the outside world sends them, on whichever hart the interrupt controller picks, while processes on the other two harts are sending. That raises the questions this lab is about. Where is the card, and how does the kernel even reach its registers? How do the driver and the card hand descriptors back and forth without a lock they could share? What may code do that runs inside an interrupt handler, and what must it never do? How does a process sleep until a packet arrives without missing the wakeup that an interrupt on another hart delivers? And what does “the device sees memory in the right order” mean on RISC-V?
The reference solution is seven commits: find the card on the PCI bus, give it descriptor rings, transmit, receive from the interrupt, answer ARP and parse IPv4/UDP, add sockets, and add the test with its host-side partner. On three harts a UDP round trip through QEMU’s user-mode network took a median of 149 to 243 microseconds, measured from the host on three boots (measured while the computer was busy with other work; your times will differ). Some driver bugs are invisible on QEMU: Clinic 2 is one that no test here can show, and Clinic 6 one that shows only when the card is made to fall behind.
Each step shows one change on the branch ext/25-e1000, the code around it, and the state of the machine when that code runs.
MakefileStep 1 of 19 · commit 1: Find the e1000 on the PCI bus
The story of this tour: one boot of the finished branch on 3 harts, with gdb attached
from the first instruction, running nettest against the host’s echo server. Its
first check, a child process of nettest (pid 4), sends one byte to the host and
waits for the echo. The states below come from that run unless a step says they are
reasoned.
Commit 1 starts outside the kernel. -netdev user is QEMU’s user-mode network,
slirp: a small TCP/IP stack inside QEMU that plays router, DHCP and DNS server for
a private network 10.0.2.0/24 and turns the guest’s packets into ordinary socket
calls on the host. hostfwd=udp::FWDPORT-:2000 makes slirp forward UDP packets that
arrive at the host’s port FWDPORT to the guest’s port 2000. -device e1000 adds an
Intel 82540EM network card, and bus=pcie.0 puts it on the virt machine’s PCIe bus,
the first bus QEMU creates.
FWDPORT is derived from your user id, as GDBPORT already is, so that two people on
one machine do not collide.
stack0hart 0’s slice of stack0, sp = 0x80009900kernel/vm.cStep 2 of 19 · commit 1: Find the e1000 on the PCI bus
gdb stopped here in the recorded boot: hart 0, building the kernel page table
(kpgtbl = 0x87fff000, the first page kalloc handed out, from the top of RAM).
Paging is still off: kvminithart turns it on only after kvmmake returns.
Two mappings join the UART, the disk and the PLIC, all identity-mapped and readable and writable, never executable:
0x30000000, 1 MiB: the PCIe configuration space of bus 0 (ECAM), one 4 KiB
header for each of 32 devices × 8 functions. QEMU’s virt machine reserves 256 MiB
there for 256 buses; xv6 only needs bus 0.0x40000000, 128 KiB: where the next step will tell the e1000 to put its
registers. It is inside the window QEMU reserves for PCIe memory (0x40000000 up to
the start of RAM at 0x80000000), so it collides with nothing.Both lie below KERNBASE, like every device in xv6, and both are mapped once at boot
and never changed: kvmmap panics if a mapping fails.
stack0sp = 0x80009920kernel/pci.cStep 3 of 19 · commit 1: Find the e1000 on the PCI bus
Recorded at line 39 (at commit 7; the lines are the same here): dev = 1, h =
0x30008000, h[0] = 269385862 = 0x100e8086, the e1000’s device and vendor IDs.
Device 0, the host bridge, was skipped. Then at line 46: size = 131072.
Three steps make the card usable:
~(0xfffe0000 & ~0xf) + 1 = 0x20000.0x40000000 to BAR 0 is a request: “answer the 128 KiB from
here”. The card does not choose; the kernel does.Word 15 of the header holds the interrupt pin: gdb read h[15] = 256 = 0x100, so
the pin byte (bits 8-15) is 1, INTA, and the line byte is 0 (nobody routed it; there
is no firmware). The boot line follows:
pci: e1000 at 0:1.0, 131072 bytes of registers at 0x40000000, INTA
pci_init runs only on hart 0, from main, before the other harts are released,
so it needs no lock. Paging is on by now (satp: kernel): the two mappings from the
last step are what make these accesses legal.
kernel/e1000.cStep 4 of 19 · commit 2: Initialize the e1000 and its descriptor rings
The driver’s state. Each ring is an array of 16-byte descriptors (layouts in
kernel/e1000_dev.h: an address, a length, a command byte for transmit, a status
byte with the DD bit for both). 16 slots × 16 bytes = 256 bytes, so one kalloc
page holds a ring with room to spare. The card requires the ring length to be a
multiple of 128 bytes and the base address to be 16-byte aligned; a page satisfies
both.
The comment at lines 22-24 is the reason DMA is easy in xv6: the kernel maps RAM one-to-one (direct map), so the kernel address of a descriptor is the physical address the card needs. Real kernels with a different layout must translate.
tx_bufs and rx_bufs remember which page each slot holds, so that the driver can
free a transmitted page later and hand a received one up. lock is for harts, not for
the card: think question 2.
stack0sp = 0x800098f0kernel/e1000.cStep 5 of 19 · commit 2: Initialize the e1000 and its descriptor rings
Recorded: e1000.tx = 0x87f53000 with tx[0].status = 1 (DD), and e1000.rx =
0x87f52000, whose slot 0 points at the buffer 0x87f51000 (still full of the
junk byte 5 that kalloc writes into every page).
The register writes are volatile stores to the mapped BAR (the R() macro, line
16), the same idiom as kernel/virtio_disk.c. Lines 86-97 (in view) program the
MAC address 52:54:00:12:34:56 into the receive filter (QEMU’s default for a card)
and enable transmit and receive.
sp = 0x3fffff7ef0e1000.lockkernel/e1000.cStep 6 of 19 · commit 3: Transmit a packet
The callers sys_send and net_tx_udp arrive in commits 6 and 5; this state was
recorded at commit 7, where the function is the same, 4 lines further down. pid 4
sends its one byte: at the “full?” test (line 116 here), i = TDT = 0, TDH = 0,
tx[0].status = 1.
The test has two halves, and both matter:
e1000.lock is the only lock here, taken in a system call after usertrap turned
interrupts on: noff 1, intena 1, SIE 0 (gdb read sstatus = 0x200000020). It
serializes the harts that transmit; the card itself never sees it.
In six measured runs of nettest (2,003 packets each) the test never refused: QEMU’s
card sends during the TDT write itself, so by the next call TDH has caught up (next
step). The -1 path is for real hardware, where the card runs in parallel.
Line 122, io_fence() after the test, keeps the reads and writes of the old slot
from moving above the read of DD; the old page is freed only now (125-126).
sp = 0x3fffff7ef0e1000.lockStep 7 of 19 · commit 3: Transmit a packet
Recorded at the publish fence (line 136 here, 140 in the final file), in a second
boot where pid 4 happened to run on hart 1: tx[0] = {addr = 0x87f33000, length
= 43, cmd = 9, status = 0}, while TDT and TDH were still 0. 43 bytes is 14
(Ethernet) + 20 (IP) + 8 (UDP) + 1 byte of data; the card pads it to the 60-byte
minimum (TCTL.PSP). cmd 9 is EOP (this descriptor ends the packet) plus RS (set
DD when done).
The fence at line 136 is the “publish” fence of think question 3. The store at 139 is the moment of hand-over: from then on slot 0 belongs to the card.
The next transmit in that boot (from inside an interrupt, the ARP reply) found TDT = 1 and TDH = 1, and slot 0’s status back at 1: QEMU’s model had already sent the packet. A real card works in parallel with the CPU, and TDH would lag behind.
stack0sp = 0x80009920kernel/plic.cStep 8 of 19 · commit 4: Receive packets from the e1000 interrupt
E1000_IRQ is 33 (kernel/memlayout.h, with the computation in its comment): QEMU
maps PCI pin INTA of device 1 to PLIC source 32 + (0 + 1) mod 4. Two things change
in the PLIC:
plicinit (line 17) gives source 33 priority 1. Priority 0 means “never”.plicinithart (line 29) sets bit 1 of the second 32-bit enable word: each word
enables 32 sources, and 33 is the second bit of sources 32-63.gdb recorded line 29 on hart 0, 1 and 2 in turn, each enabling the source for its own S-mode context. So the PLIC may offer IRQ 33 to any of the three harts, and whoever claims first handles it; the next steps show that this is what happened.
sp = 0x3fffff7dd0no switch: kernelvec pushed its 256-byte frame onto the stack hart 2 was already on, pid 4’s kernel stack (kernel/kernelvec.S:14)kernel/trap.cStep 9 of 19 · commit 4: Receive packets from the e1000 interrupt
pid 4 sent its byte on hart 2, entered recv, found the queue empty and registered
on the socket’s channel. Before it got to sleep, slirp’s ARP request arrived and gdb
stopped at line 203 on hart 2: irq = 33. That is the check of the IRQ number:
the card’s interrupt arrives as 33.
The interrupt landed on pid 4’s own hart, on pid 4’s kernel stack, in the short stretch
after recv released its socket lock (noff 0, interrupts on) and before sleep
took p->lock. (gdb cannot unwind past kernelvec; the sys_recv frames are
reasoned from sp, which is inside pid 4’s kernel stack, and from the order of the
breakpoints.) The irony is instructive: the frame is not even for pid 4. It is an ARP
request, and pid 4 is merely the process that happened to be running. Interrupt
handlers serve the device.
devintr calls e1000_intr and, afterwards, plic_complete. Between claim and
complete the PLIC will not offer IRQ 33 to any hart, so e1000_intr never runs on
two harts at once. The driver still takes e1000.lock, because the transmit path,
which any hart may run, shares the same structure.
The next interrupt of this run, the echo itself, landed on hart 1, idle in its
scheduler, on hart 1’s slice of stack0. In the measurement runs each hart
handled 424 to 869 of the roughly 1,800 interrupts per run.
sp = 0x3fffff7cc0, below the kernelvec frameacquire pushed the first level, so intena is 0 (Locks and interrupt state)e1000.lockkernel/e1000.cStep 10 of 19 · commit 4: Receive packets from the e1000 interrupt
Recorded at line 163: RDH = 1, RDT = 15. The card has filled slot 0 (the slot after
RDT) and moved its head to 1. Recorded at line 194, just before the RDT store of the
first iteration: i = 0, n = 1, the slot’s length 60 (slirp’s ARP request,
padded), status 0 again, and the buffer in the slot is a fresh page. The next
interrupt of the run found RDT = 0: slot 0 was the card’s again.
Read in order:
kalloc never sleeps, so it is
allowed here. If memory runs out, the old buffer stays and the packet is lost.In the measurement runs, 87% of interrupts found exactly one packet, 9% two, 1% three, and about 2% none: a packet that arrived after ICR was read was taken by the running loop, and its own interrupt then found nothing.
Step 11 of 19 · commit 4: Receive packets from the e1000 interrupt
Recorded at line 202: noff 0, SIE still 0 (this is still the interrupt handler), and
lens[0] = 60. The packets in bufs no longer belong to the ring, so they need
no lock, and net_rx is free to transmit. We ran a copy that calls it before the
release: panic: acquire on the first ARP request (think question 5).
At this commit net_rx (in kernel/net.c) frees the page: the driver works, but
nothing listens yet.
kernel/net.hStep 12 of 19 · commit 5: Answer ARP, parse IPv4 and UDP, send UDP
Every multi-byte field on the wire is big-endian; RISC-V is little-endian.
htons/htonl reverse the bytes; ntohs and ntohl are the same operations under
the names that say which way the value travels.
The structs are packed: an IP header starts 14 bytes into the frame, so its
32-bit addresses sit at offsets 26 and 30, not multiples of 4. Packed tells the
compiler to assume nothing about alignment and to access them a byte at a time.
gdb shows what byte order means in practice. In the recorded net_tx_udp, print *udp gave sport = 53511, dport = 44904: 53511 is 0xd107, the bytes 07 d1 of
port 2001 read back as a little-endian number; 44904 is the echo port (ECHOPORT) the
same way.
ip_len = 7424 is 0x1d00, 29 bytes. The fields are right; it is the debugger that
reads them in the host’s order.
sp = 0x3fffff7c90kernel/net.cStep 13 of 19 · commit 5: Answer ARP, parse IPv4 and UDP, send UDP
Slirp knows that the guest is 10.0.2.15, but not its MAC address, so before it can
deliver the echo it broadcasts “who has 10.0.2.15? tell 10.0.2.2”. In the recorded run
this was the very first frame xv6 received, after its first transmit: slirp asks
only when it has something to deliver.
arp_rx first checks that this really is an Ethernet/IPv4 request for us, with 6-
and 4-byte addresses (lines 84-86), then turns the request into the reply in the same
page: op 2, target = the old sender, sender = us. Recorded at line 100: sha =
52:54:00:12:34:56, op = 512 (2 in network order), tip = 33685514 =
0x0202000a, the bytes 0a 00 02 02: 10.0.2.2.
The transmit happens from inside the interrupt: gdb recorded e1000_transmit on hart
2 with e1000.lock held, noff 1, intena 0, into slot 1. It is legal because
e1000_intr released the lock before calling net_rx.
stack0kernel/net.cStep 14 of 19 · commit 5: Answer ARP, parse IPv4 and UDP, send UDP
(State reasoned from the code: the call chain of the recorded sock_rx step, one
frame up, on hart 1.) Everything that arrives is checked before anyone trusts a
length: version 4 with a 20-byte header (0x45), protocol UDP, addressed to us, not a
fragment, a correct header checksum, and lengths that fit inside what the card
actually received. A UDP length larger than the frame would otherwise make recv
copy bytes past the end of the packet.
At this commit udp_rx (lines 104-110) drops the packet; commit 6 replaces it with
the socket layer. Recorded at commit 5: gdb stopped in udp_rx with the data hello
at buf + 42, sent from the host.
sp = 0x3fffff7f30kernel/net.cStep 15 of 19 · commit 5: Answer ARP, parse IPv4 and UDP, send UDP
Recorded at line 69 (commit 7): no lock held, interrupts on (sstatus =
0x200000022, SIE set), the normal state of a system call between critical sections.
The caller left 42 bytes free at the front of the page (NET_HDRS), so the headers are
written in front of the data without copying it (lines 51-67). Every frame goes to
52:55:0a:00:02:02, the MAC address with which slirp answers ARP for the host. That
works for the DNS server too, although slirp answers ARP for 10.0.2.3 with
52:55:0a:00:02:03: slirp accepts IP frames whatever their destination MAC. So xv6
needs no ARP requests of its own (a real network would; see “Further”). The UDP
checksum 0 means “none”, which IPv4 allows; the IP header checksum is required, and
slirp checks it.
kernel/socket.cStep 16 of 19 · commit 6: Add UDP sockets: bind, unbind, send and recv
A socket is a port and a bounded queue of arrived packets. Two kinds of lock:
socktab.lock protects which socket has which port: bind checks for a duplicate
and claims a free socket in one critical section.s->lock protects that socket’s queue (head, n, q[]); the port field
changes only with both locks held, so reading it under either is safe.Port 0 means “free”, which is why sock_rx drops packets for port 0 before it looks
(next step): otherwise they would be queued on a free socket, and the bind that
later takes it would start with n = 0 and lose those pages. We sent 40 raw frames
to port 0 into a copy without that test (through QEMU’s -netdev dgram, since
slirp’s forwarding cannot produce port 0): free pages went from 32447 to 32431, a
bind/unbind of all 16 ports left 32431, and 40 more frames took it to 32415. With the
test, 32447 throughout.
sockfind returns the socket locked, while the caller still holds socktab.lock:
the order is socktab.lock → s->lock → p->lock (in wakeup,
sleep_prepare, killed). No path takes them in another order, and none of them
is taken while e1000.lock is held, so the network adds no cycle to the lock-order
graph of Tour 51: The lock-order graph, measured.
stack0sp = 0x8000a600, hart 1’s slice of stack0socket 2001's s->lockkernel/socket.cStep 17 of 19 · commit 6: Add UDP sockets: bind, unbind, send and recv
Recorded at line 94, the second e1000 interrupt of the run, on hart 1, which was
idle: the echo of pid 4’s byte (sport ECHOPORT, len 1) has been queued on port
2001’s socket (n = 1), and wakeup is next. noff 1: socktab.lock was released
at line 76.
wakeup runs with s->lock held and takes each process’s p->lock in turn
(noff 2 inside). It finds pid 4 with chan = s and makes it RUNNABLE. Because the
append and the wakeup happen under the lock that recv holds while it checks the
queue and registers on the channel (next step), the wakeup cannot fall between the
two. Clinic 3 removes this lock and loses wakeups.
A full queue drops the packet: the handler must not wait. No packet was dropped in the measured runs; every socket there had one packet in flight at a time.
sp = 0x3fffff7f50socket 2001's s->lockkernel/socket.cStep 18 of 19 · commit 6: Add UDP sockets: bind, unbind, send and recv
Recorded at line 221, just before the interrupt of the “devintr” step: pid 4 on hart
2, port 2001, n = 0, holding only its socket’s lock (noff 1, intena 1: a system
call).
The loop is virtio_disk_rw's, with s as the channel:
kkill makes a sleeper runnable but does not clear its
channel, so the loop must look, also when packets are queued;sleep_prepare(s) while holding s->lock (221);sleep, re-acquire, around again (222-224).Any sock_rx for this socket either finished before the check (the queue is not
empty) or waits for s->lock until step 3 is done, so its wakeup finds p->chan
= s and clears it. sleep then returns at once, or wakes from SLEEPING. In this
run the ARP interrupt hit pid 4’s hart between line 222 and the
acquire inside sleep; the echo came on hart 1.
One window stays open, the same as in piperead: a kill after the killed(p)
check and before sleep finds the process not yet SLEEPING, so it only sets
killed; the process then sleeps until the next wakeup.
After the loop the packet comes off the queue under the lock, and the copy to user
space happens with no lock held (236-246): copyout may fault in a lazily
allocated heap page, and there is no reason to make the interrupt handler wait for
it.
user/nettest.cStep 19 of 19 · commit 7: Add nettest and its host-side partner, nettest.py
The three-process check: three children, one per hart if the scheduler places them so, each with its own port (2010-2012), each doing 300 round trips with lengths from 1 to 1472 bytes. Every byte of every echo is compared with what was sent, and the pattern depends on the child’s number, so a packet delivered to the wrong socket fails the comparison.
run (lines 69-90) forks the check and a watchdog that kills it after a time limit;
a lost packet becomes FAIL (timed out) instead of a hang, and the next check still
runs. This mattered: in Clinics 1, 3 and 5 the watchdog turned every stall into a
line of output.
Every check first calls unbind on its ports. A process left over from an earlier,
failed run may still sleep in recv there; unbind wakes it and frees the port.
Without that, one failure would make every later run fail at bind, which is a false
claim of its own.
The host side is nettest.py at the top of the tree: echo PORT sends every
packet back; ping FWDPORT N times round trips to nettest server N.
Lab 25 · wrap-up
On the branch (ext/25-e1000, 7 commits), built with the project toolchain and run on 3
harts (-smp 3 -m 128M), with python3 nettest.py echo ECHOPORT 600 on the host:
$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: OK
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes took 3 ticks
nettest: ALL OK
$ usertests -q
usertests starting
ALL TESTS PASSED
$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: OK
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes took 1 ticks
nettest: ALL OK
And the host-side ping against nettest server 1000:
$ nettest server 1000
nettest server: echoing 1000 packets on port 2000
nettest server: done
# on the host:
nettest.py: 1000 round trips: OK
Every commit was built on its own, and usertests -q printed ALL TESTS PASSED on 3
harts at each of commits 1 to 6 as well. nettest printed ALL OK on all three boots
of the head that ran it (the run above and two gdb traces). A kernel without three of
the branch’s fixes (the transmit “full” test, port 0, the port re-check in recv)
passed the same sequence, and nettest on all 8 boots of its measurement copies.
The fixes themselves were checked with scratch copies: Clinic 6’s run for the transmit ring, and the port-0 frames described in the reveal step on sockets. Clinics 1 to 5 were run on a kernel without those fixes; none of the fixes touches the code those clinics change.
DNS: slirp’s DNS server at 10.0.2.3 forwards queries to the DNS resolver the host
itself uses, and every recorded run got an answer. On a host without a working
resolver this check fails and the test says so (ALL OK except dns).
What this demonstrates: packets go both ways through both rings, many times around;
three processes on three harts send and receive at once without a packet reaching the
wrong socket or a byte changing; kill ends a recv; and the rest of the kernel is
unaffected. It does not demonstrate the absence of races: Clinic 3’s lost wakeup passed
9 of 9 runs without help, and Clinic 2’s missing fences cannot fail on QEMU at all.
Keys: ← → step · Home start