xv6, line by line
lab 25

Extension labs · lab 25 · Devices · ★★★★★

An e1000 network driver and UDP sockets

xv6 already has one device driver that works by DMA: the virtio disk, where the kernel writes descriptors into memory, pokes a register, and an interrupt says the device is done (Tour 29: A disk read, end to end). In this lab you write a second one, for QEMU’s emulation of the Intel e1000 network card, and put a small network stack and a socket interface on top of it, so that a user program can exchange UDP packets with a program on the host.

The disk is a polite device: it only speaks when spoken to, and exactly one request waits for each answer. A network card is not. Packets arrive whenever the outside world sends them, on whichever hart the interrupt controller picks, while processes on the other two harts are sending. That raises the questions this lab is about. Where is the card, and how does the kernel even reach its registers? How do the driver and the card hand descriptors back and forth without a lock they could share? What may code do that runs inside an interrupt handler, and what must it never do? How does a process sleep until a packet arrives without missing the wakeup that an interrupt on another hart delivers? And what does “the device sees memory in the right order” mean on RISC-V?

The reference solution is seven commits: find the card on the PCI bus, give it descriptor rings, transmit, receive from the interrupt, answer ARP and parse IPv4/UDP, add sockets, and add the test with its host-side partner. On three harts a UDP round trip through QEMU’s user-mode network took a median of 149 to 243 microseconds, measured from the host on three boots (measured while the computer was busy with other work; your times will differ). Some driver bugs are invisible on QEMU: Clinic 2 is one that no test here can show, and Clinic 6 one that shows only when the card is made to fall behind.

Read first: Tour 9: Device interrupts and the PLIC, Tour 15: Spinlocks from the hardware up, Tour 16: sleep and wakeup, and the lost-wakeup problem, Tour 19: Memory ordering across harts, Tour 24: The kernel page table and turning paging on, Tour 29: A disk read, end to end, Tour 44: One interrupt, three landing sites, Tour 50: noff and intena through a sleep, a yield and an interrupt · Locks and interrupt state, The stacks of xv6

What this lab teaches

  • How a kernel finds a PCI device on a machine with no firmware help: memory-mapped configuration space, sizing and placing a BAR, and mapping both into the kernel page table.
  • How a driver and a DMA device share a ring of descriptors: who owns which slot, what the head and tail registers mean, and why one slot of each ring must stay empty.
  • Why the order in which the CPU writes memory and device registers matters, what fence does about it, and why QEMU cannot show you the bug.
  • How a PCI interrupt pin becomes a PLIC interrupt number on QEMU’s virt machine, and what the handler may and may not do on whatever hart and stack it lands on.
  • How to deliver data from an interrupt handler to a sleeping process with this tree’s sleep_prepare/sleep/wakeup protocol, and which lock makes the hand-off safe.
  • How to measure the result: round-trip latency from the host, interrupts and packets per hart, and where three harts wait for each other’s locks.

The reference branch

ext/25-e1000 in ShowMeTheStack/xv6-riscv-labs, branched from the frozen commit 06aad25; 7 commits.

git clone https://github.com/ShowMeTheStack/xv6-riscv-labs
cd xv6-riscv-labs
git checkout -b my-e1000 06aad25   # start your own
git diff 06aad25 origin/ext/25-e1000   # only when you want the answer

1. The spec

Behaviour. QEMU gets an e1000 card on its PCIe bus, connected to QEMU’s user-mode network (“slirp”). On that network xv6 is 10.0.2.15, the host is 10.0.2.2 and slirp’s DNS server is 10.0.2.3; UDP packets sent to the host’s port FWDPORT arrive at xv6’s port 2000. The QEMU flags (the Makefile chooses FWDPORT from your user id, as it chooses GDBPORT: echo $(( $(id -u) % 5000 + 30000 )) prints it, on Linux, macOS and WSL alike):

-netdev user,id=net0,hostfwd=udp::FWDPORT-:2000
-device e1000,netdev=net0,bus=pcie.0

Four new system calls:

int bind(int port);       // start queuing packets that arrive for port
int unbind(int port);     // stop; drop what is queued; wake a waiting recv
int send(uint32 dst, uint16 dport, uint16 sport, char *buf, int len);
int recv(uint16 port, uint32 *src, uint16 *sport, char *buf, int len);

What must not change. Everything else. usertests -q must print ALL TESTS PASSED on 3 harts, and the kernel must stay usable while packets arrive.

The test. nettest ECHOPORT talks to python3 nettest.py echo ECHOPORT on the host, which sends every packet back to its sender. ECHOPORT is any free UDP port on the host other than FWDPORT (QEMU already holds that one), e.g. 25501; port numbers in the recorded runs are shown as ECHOPORT and FWDPORT. Each check runs in a child process with a watchdog that kills it after a time limit, so a lost packet becomes a FAIL line rather than a hang:

$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: OK
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes took 3 ticks
nettest: ALL OK

nettest server N echoes N packets that arrive at port 2000; python3 nettest.py ping FWDPORT N on the host then times each round trip.

2. Think first

Answer each question in your head (or on paper) before opening a hint. Hints get more specific; the reference answer comes last.

1Where is the card, and how does the kernel reach its registers?

The UART and the virtio disk sit at fixed physical addresses that kernel/memlayout.h simply lists, and kvmmake maps them. The e1000 is different: QEMU puts it on a PCIe bus, and nothing tells xv6 where its registers are, because there is no firmware (-bios none) to set the bus up. Before writing any driver code, decide: how does the kernel find the card, how does it choose where the card’s registers appear, and what must change in the kernel page table so that the kernel can touch either? Which of these happen before paging is turned on, and does that matter?

Check yourself

1warm-upChoose one

Configuration space on this machine starts at 0x30000000, and a device’s header is at 0x30000000 + (bus << 20) + (device << 15) + (function << 12). At which address does xv6 read the e1000’s ID word (bus 0, device 1, function 0)?

2solidType a number

pci_init writes 0xffffffff to the e1000’s BAR 0 and reads back 0xfffe0000. The low 4 bits of a memory BAR are flags (zero here). How many bytes of registers does the card need?

decimal, 0x hex or 0b binary

2How do the driver and the card share a ring without sharing a lock?

The card reads packets to send from memory and writes packets it receives into memory, by DMA, at any moment. It cannot take a spinlock. The virtio disk solved this with its descriptor table and two rings (Tour 29: A disk read, end to end); the e1000 has a simpler scheme: for each direction, one ring of 16 descriptors in RAM and two registers, head and tail. Design the rules. Which slots may the driver write, and which belong to the card? How does the driver learn that the card has finished with a slot? When may the page a transmitted packet lives in be freed? And with the receive tail starting at 15 and the head at 0, how many packets can arrive before the driver has given anything back?

Check yourself

1solidType a number

At initialization RDH = 0 and RDT = 15 in a 16-slot receive ring. If the driver never gives a slot back, how many packets can the card receive?

decimal, 0x hex or 0b binary
2solidPut in order

Put the steps of e1000_transmit for one packet in order.

  1. write TDT = the next slot
  2. acquire e1000.lock
  3. fence
  4. release e1000.lock
  5. write the new packet’s address, length and command into the slot, and clear DD
  6. read TDT, and check that the slot after it is not TDH and that the slot’s DD bit is set
  7. free the page of the packet this slot sent last time

3In what order does the card see your stores?

e1000_transmit writes four fields of a descriptor in RAM and then the TDT register. The card reacts to the TDT write by reading the descriptor by DMA. What guarantees that it reads the new descriptor, not the old one? Consider both the compiler and the processor. kernel/virtio_disk.c has an answer for its own rings; is the same needed here, and in which other places of the driver? Then predict: on QEMU, what happens if you leave it out?

Check yourself

1solidChoose all that apply

Which statements about io_fence() before *R(E1000_TDT) = ... are true?

4Which interrupt, and who runs the handler?

The card raises an interrupt when packets have arrived. It is wired to the PCI interrupt pin INTA, not to a PLIC source number. Which PLIC interrupt number does xv6 have to enable and dispatch, and how would you check your answer on the running machine? Then think about where the handler runs: on which hart, with what interrupt state, on which stack, and what it must do to the card so that the interrupt does not fire again at once.

Check yourself

1solidType a number

On QEMU’s virt machine a PCI device’s pin becomes PLIC source 32 + (pin + device) mod 4, with INTA = 0. A second card on bus 0 as device 2, also using INTA: which PLIC source does it raise?

decimal, 0x hex or 0b binary
2solidChoose one

gdb stopped in e1000_intr on hart 0 while no process was running there. Where is sp?

5What may the receive path do inside the interrupt handler?

Everything from “a slot has DD set” to “the packet is queued for a process” runs inside e1000_intr, with SIE off, on an arbitrary hart, possibly with no process at all. Some packets need an answer (ARP), so the receive path sometimes transmits. List what the receive path must never do there, and decide how e1000_intr should use e1000.lock given that an ARP reply calls e1000_transmit, which takes the same lock.

Check yourself

1solidChoose one

A learner calls net_rx for each packet before release(&e1000.lock) in e1000_intr. The first packet to arrive is slirp’s ARP request. What happens?

6How does recv sleep until a packet arrives without missing it?

recv runs in a process on one hart; the packet it waits for is queued by the interrupt handler on another hart, at any moment. Design the socket’s queue and its locking: which lock protects the queue, who takes it, and in what order relative to the process locks that wakeup takes? Then write the wait loop with this tree’s sleep_prepare/sleep/wakeup (not the book’s sleep(chan, lock)), and explain why no packet can arrive “between the check and the sleep” unnoticed. What should a full queue do, given that the handler may not wait?

Check yourself

1solidTrue or false, and why

True or false: sock_rx runs with interrupts off, so it may update the socket’s queue without the socket’s lock.

Why?

2deepPut in order

Put the steps of recv on an empty queue in order (reference code).

  1. acquire s->lock and check the queue again
  2. take the oldest packet and release s->lock
  3. release s->lock and sleep()
  4. copyout the data, then kfree the page
  5. sleep_prepare(s)
  6. release socktab.lock
  7. acquire socktab.lock and find the socket, which acquires s->lock
  8. see that the port is still bound, the process is not killed, and the queue is empty

7Which bytes are big-endian?

RISC-V stores multi-byte integers little-endian. Ethernet, ARP, IP and UDP headers are big-endian on the wire. The system calls take ports and addresses as ordinary numbers. Decide where conversions happen, which fields need them, and predict what goes wrong, and where you would see it, if net_tx_udp stores the destination port without converting it.

Check yourself

1solidType a number

net_tx_udp stores udp->dport = dport; with dport = 25501 (0x639d) and no htons. Which destination port does the host see?

decimal, 0x hex or 0b binary

3. Build it

Start.

git checkout -b my-e1000 06aad25

Add the two QEMU flags from the spec to QEMUOPTS first and boot: nothing changes yet, but info pci in QEMU’s monitor (Ctrl-A C) lists the card, which tells you its bus, device number and IDs before you write a line of code.

Milestones, in an order where each one can be checked on its own.

  1. Find the card. Map the configuration window and the register window in kvmmake, scan bus 0, size and place BAR 0, enable memory and bus mastering, and print what you found. Check: one boot line, with 131,072 bytes at 0x40000000, and usertests -q still passes (the two new mappings must not overlap anything).
  2. Rings. Reset the card, allocate the two rings and the receive buffers with kalloc, program the ring registers, the MAC address filter, and transmit and receive control. Check with gdb: x/4wx 0x40002810 shows RDH, a reserved word, RDT (0x40002818) as 0 … 15.
  3. Transmit. Write e1000_transmit. You cannot see a packet leave yet without a protocol, so check with gdb that TDH follows TDT after each call (on QEMU the card sends during the TDT write itself). Because of that, QEMU never fills your ring: test the “full” case on purpose (Clinic 6 shows one way).
  4. Receive. Enable IRQ 33 in the PLIC, dispatch it in devintr, write the handler, and for now free every packet. Check: send one UDP packet to FWDPORT from the host (a three-line Python script) and break on your handler; the first frame is 60 bytes, a broadcast ARP request from slirp (52:55:0a:00:02:02) asking for 10.0.2.15.
  5. ARP, IPv4, UDP. Answer the ARP request; parse IPv4/UDP; build headers for sending. Check with gdb: after the ARP reply, slirp delivers the UDP packet, and its data hello sits at offset 42 of the buffer.
  6. Sockets and the system calls. Then the test program and its host partner. Run nettest, usertests -q and nettest again on 3 harts.

Debugging advice. Run QEMU with -smp 3 and attach gdb (make qemu-gdb picks its own gdb port).

4. Debugging clinic

Each of these bugs was put into the reference solution on purpose and run on three harts. The symptom is exactly what happened. Try to explain it before revealing why.

1The receive slot is never given back

The last statement of the harvest loop, which moves RDT onto the slot just emptied, is forgotten:

     // give the slot back to the device.
-    *R(E1000_RDT) = i;
   }

Many drivers keep their own index of the next slot to look at instead of computing it from RDT. We built that variant too (a field rx_next, advanced at the end of the loop), with the same mistake, and as a control the same variant with the RDT write.

What happened when we ran it

# nettest server 30 in xv6; nettest.py ping FWDPORT 30 on the host.
# as written above (the next slot is computed from RDT):
nettest.py: ping 0: no reply in 5 s: FAIL
# gdb, attached afterwards:
RDH 2 RDT 15
slot  0 status 0 length 60
slot  1 status 7 length 60
slot  2 status 0 length 0
[...]

# the variant with its own index rx_next (two boots, the same result):
nettest.py: ping 14: no reply in 5 s: FAIL
# gdb, attached afterwards:
RDH 15 RDT 15
$1 = 15
slot  0 status 0 length 60
slot  1 status 0 length 60
[...]
slot 14 status 0 length 60
slot 15 status 0 length 0

# the control (rx_next, with the RDT write):
nettest.py: 30 round trips: OK

2No fence before the tail writes (not visible on QEMU)

Both “publish” fences are removed: the one before the TDT write in e1000_transmit and the one before the RDT write in e1000_intr.

   e1000.tx[i].status = 0;
-
-  // the descriptor must be in memory before the device can see
-  // the new tail and read it.
-  io_fence();

   // give the slot to the device.
   *R(E1000_TDT) = (i + 1) % TX_RING_SIZE;

(and the same three lines before *R(E1000_RDT) = i;).

What happened when we ran it

$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: OK
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes took 7 ticks
nettest: ALL OK
[... two more runs on the same boot: ALL OK, ALL OK ...]

# e1000_transmit in this build (objdump -d), the end of the descriptor stores
# and the TDT store, with nothing in between but the index arithmetic:
    80005ef2:	01373023          	sd	s3,0(a4)
    80005ef6:	639c                	ld	a5,0(a5)
    80005ef8:	97ca                	add	a5,a5,s2
    80005efa:	01479423          	sh	s4,8(a5)
    80005efe:	4725                	li	a4,9
    80005f00:	00e785a3          	sb	a4,11(a5)
    80005f04:	00078623          	sb	zero,12(a5)
    80005f08:	2485                	addiw	s1,s1,1
[...]
    80005f18:	400047b7          	lui	a5,0x40004
    80005f1c:	8097ac23          	sw	s1,-2024(a5) # 40003818 <_entry-0x3fffc7e8>

3The handler updates the socket queue without the socket’s lock

“Interrupts are off in the handler, so nothing can interfere”: sock_rx lets go of the socket’s lock as soon as it has found the socket.

   if (s == 0) {
     kfree(buf);
     return;
   }
+  // interrupts are off here, so nothing can interfere.
+  release(&s->lock);

   if (s->n == NQ) {
     // nobody is reading fast enough: drop the packet.
-    release(&s->lock);
     kfree(buf);
     return;
   }
   [...]
   s->n++;
   wakeup(s);
-  release(&s->lock);
 }

We ran it as written, and a second copy with a delay loop (200,000 iterations of a volatile counter) in recv between the “queue is empty” check and sleep_prepare(s), to widen the window. As a control, the reference with the same delay.

What happened when we ran it

# as written: 3 boots, 3 runs of nettest each: 9 of 9 ALL OK.

# with the delay (one of six boots; five of the six failed at least one check):
$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: FAIL (timed out)
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes: FAIL (timed out)
nettest: SOME TESTS FAILED

# gdb snapshot of the same boot, 20 s after the first FAIL, during the last check:
pid 1 init SLEEPING chan 0x80010e90 <proc>
pid 2 sh SLEEPING chan 0x80010ff8 <proc+360>
pid 3 nettest SLEEPING chan 0x80011160 <proc+720>
pid 18 nettest SLEEPING chan 0x80021f58 <socktab+448>
pid 19 nettest SLEEPING chan 0x80008940 <ticks>
sock 0 at 0x80021db0 <socktab+24>: port 2002 n 0 head 10
sock 1 at 0x80021f58 <socktab+448>: port 2003 n 1 head 5

# the control (reference with the same delay): 1 boot, 3 runs: ALL OK.

4The socket is protected by a sleep-lock

“recv sleeps anyway, so why not a sleep-lock?” Every struct spinlock lock of a socket becomes a struct sleeplock, every acquire(&s->lock) an acquiresleep(&s->lock), every release(&s->lock) a releasesleep(&s->lock):

 struct sock {
-  struct spinlock lock;
+  struct sleeplock lock;
[...]
     if (s->port == port) {
-      acquire(&s->lock);
+      acquiresleep(&s->lock);
       return s;

What happened when we ran it

$ nettest ECHOPORT
scause=0xd sepc=0x80004028 stval=0x30
panic: kerneltrap

# another boot, gdb attached:
$ nettest ECHOPORT
nettest: one round trip: OK
scause=0xd sepc=0x80004028 stval=0x30
panic: kerneltrap
# at the breakpoint on panic: hart 0, no process (cpus[0].proc = 0), noff 2,
# sp = 0x80009490 <stack0+2880>

5Forgot htons on the destination port

   udp->sport = htons(sport);
-  udp->dport = htons(dport);
+  udp->dport = dport;

QEMU recorded every frame (-object filter-dump,...,file=r.pcap); a short Python script decoded it.

What happened when we ran it

$ nettest ECHOPORT
nettest: one round trip: FAIL (timed out)
nettest: 100 round trips: FAIL (timed out)
nettest: 3 processes x 300 round trips: FAIL (timed out)
nettest: kill wakes recv: OK
nettest: dns query to 10.0.2.3: FAIL (timed out)
[...]
# the host's echo server, the whole run:
nettest.py: echoing on port ECHOPORT
# the capture:
  1 UDP 10.0.2.15:2001 -> 10.0.2.2:43880 udp-length 9 (dst port bytes ab 68)
  2 ARP request 10.0.2.2 -> 10.0.2.15
  3 ARP reply 10.0.2.15 -> 10.0.2.2
  4 ICMP 127.0.0.1 -> 10.0.2.15 type 3 code 3
[...]
 13 UDP 10.0.2.15:2030 -> 10.0.2.3:13568 udp-length 44 (dst port bytes 35 00)
 14 ICMP 10.0.2.2 -> 10.0.2.15 type 3 code 0

6The transmit ring is allowed to fill all 16 slots

The natural first version of the “ring full” test checks only the DD bit of the slot at TDT:

   int i = *R(E1000_TDT);
-  if ((i + 1) % TX_RING_SIZE == *R(E1000_TDH) ||
-      (e1000.tx[i].status & E1000_TXD_STAT_DD) == 0) {
+  if ((e1000.tx[i].status & E1000_TXD_STAT_DD) == 0) {
     release(&e1000.lock);
     return -1;
   }

QEMU’s card sends during the TDT write, so on QEMU the ring never holds more than one packet and this can never matter. To make the card fall behind, a scratch copy (not on the branch) adds a test switch: a send from source port 65535 sets a flag, and while it is set e1000_transmit clears the transmit-enable bit in TCTL, so the card stops sending. A small test program then sends 20 packets with the card stopped, releases it, and sends 3 more (printing each send’s result), and then 3 more in a second command. The host’s echo server counts what arrives. We ran the same test on the broken check and on the reference.

What happened when we ran it

# broken check (DD at TDT only):
$ nabuse txwedge ECHOPORT
held: 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 -1 -1 -1 -1
released: -1 -1 -1
$ nabuse tx ECHOPORT 3
-1 -1 -1
tx: done
# on the host:
nettest.py: echoed 0 packets, then 8 s without one
# gdb, another boot of the same kernel, after "released":
TDH 0 TDT 0
slot  0 status 0 length 52
slot  1 status 0 length 52
[...]
slot 15 status 0 length 52

# the reference (one slot always empty):
$ nabuse txwedge ECHOPORT
held: 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 -1 -1 -1 -1 -1
released: 0 0 0
$ nabuse tx ECHOPORT 3
0 0 0
tx: done
# on the host:
nettest.py: echoed 21 packets, then 8 s without one

5. The reference solution

Take the guided tour through the reference solution, one commit at a time, with the machine state at every step:

Open the reveal tour →

Or read the commits

  1. 537db01 Find the e1000 on the PCI bus

    Makefile

    @@ -27,9 +27,10 @@ OBJS = \
    2727 $K/exec.o \
    2828 $K/sysfile.o \
    2929 $K/kernelvec.o \
    3030 $K/plic.o \
    31 $K/virtio_disk.o
    31 $K/virtio_disk.o \
    32 $K/pci.o
    3233
    3334# riscv64-unknown-elf- or riscv64-linux-gnu-
    3435# perhaps in /opt/riscv/bin
    3536#TOOLPREFIX =
    @@ -178,8 +179,15 @@ QEMUOPTS = -machine virt -bios none -kernel $K/kernel -m 128M -smp $(CPUS) -nogr
    178179QEMUOPTS += -global virtio-mmio.force-legacy=false
    179180QEMUOPTS += -drive file=fs.img,if=none,format=raw,id=x0
    180181QEMUOPTS += -device virtio-blk-device,drive=x0,bus=virtio-mmio-bus.0
    181182
    183# an e1000 network card behind QEMU's user-mode network ("slirp").
    184# UDP packets sent to the host's port $(FWDPORT) arrive at the
    185# guest's port 2000; the guest reaches the host as 10.0.2.2.
    186FWDPORT = $(shell expr `id -u` % 5000 + 30000)
    187QEMUOPTS += -netdev user,id=net0,hostfwd=udp::$(FWDPORT)-:2000
    188QEMUOPTS += -device e1000,netdev=net0,bus=pcie.0
    189
    182190qemu: check-qemu-version $K/kernel fs.img
    183191 $(QEMU) $(QEMUOPTS)
    184192
    185193.gdbinit: .gdbinit.tmpl-riscv

    kernel/defs.h

    @@ -176,8 +176,11 @@ void plicinit(void);
    176176void plicinithart(void);
    177177int plic_claim(void);
    178178void plic_complete(int);
    179179
    180// pci.c
    181void pci_init(void);
    182
    180183// virtio_disk.c
    181184void virtio_disk_init(void);
    182185void virtio_disk_rw(struct buf *, int);
    183186void virtio_disk_intr(void);

    kernel/main.c

    @@ -27,8 +27,9 @@ main()
    2727 binit(); // buffer cache
    2828 iinit(); // inode table
    2929 fileinit(); // file table
    3030 virtio_disk_init(); // emulated hard disk
    31 pci_init(); // find the e1000 network card
    3132 userinit(); // first user process
    3233
    3334 __atomic_store_n(&started, 1, __ATOMIC_RELEASE);
    3435 } else {

    kernel/memlayout.h

    @@ -7,8 +7,10 @@
    77// 02000000 -- CLINT
    88// 0C000000 -- PLIC
    99// 10000000 -- uart0
    1010// 10001000 -- virtio disk
    11// 30000000 -- PCIe configuration space (ECAM)
    12// 40000000 -- PCIe memory window; we put the e1000's registers here
    1113// 80000000 -- qemu's boot ROM loads the kernel here,
    1214// then jumps here.
    1315// unused RAM after 80000000.
    1416
    @@ -24,8 +26,18 @@
    2426// virtio mmio interface
    2527#define VIRTIO0 0x10001000
    2628#define VIRTIO0_IRQ 1
    2729
    30// qemu's PCIe host bridge: configuration space for every
    31// bus/device/function, memory-mapped ("ECAM"). bus 0 is enough.
    32#define PCIE_ECAM 0x30000000L
    33#define PCIE_ECAM_SIZE 0x100000L
    34
    35// where pci_init() asks the e1000 to decode its registers
    36// (BAR 0), inside the PCIe memory window at 0x40000000.
    37#define E1000_REGS 0x40000000L
    38#define E1000_REGS_SIZE 0x20000L
    39
    2840// core-local interrupt controller (CLINT)
    2941#define CLINT_BASE 0x02000000L
    3042#define CLINT(hart) (CLINT_BASE + (hart) * 4)
    3143

    kernel/pci.c

    @@ -0,0 +1,55 @@
    1//
    2// just enough PCI to find qemu's e1000 network card.
    3//
    4// qemu -machine virt has a PCIe host bridge whose configuration
    5// space is memory-mapped ("ECAM"): every bus/device/function has a
    6// 4096-byte block of registers at
    7// PCIE_ECAM + (bus << 20) + (device << 15) + (function << 12).
    8// xv6 looks only at bus 0, function 0 of each device.
    9//
    10
    11#include "types.h"
    12#include "param.h"
    13#include "memlayout.h"
    14#include "riscv.h"
    15#include "defs.h"
    16
    17// 32-bit words of a configuration header.
    18#define PCI_ID 0 // vendor ID (low 16 bits), device ID (high 16)
    19#define PCI_COMMAND 1 // command (low 16 bits), status (high 16)
    20#define PCI_BAR0 4 // base address register 0
    21#define PCI_INTR 15 // interrupt line (bits 0-7), pin (bits 8-15)
    22
    23#define PCI_COMMAND_MEMORY 0x2 // respond to accesses inside the BARs
    24#define PCI_COMMAND_MASTER 0x4 // allow the device to do DMA
    25
    26// device 0x100e (82540EM) from vendor 0x8086 (Intel).
    27#define E1000_ID 0x100e8086
    28
    29void
    30pci_init(void)
    31{
    32 for (int dev = 0; dev < 32; dev++) {
    33 volatile uint32 *h = (volatile uint32 *)(PCIE_ECAM + (dev << 15));
    34 if (h[PCI_ID] != E1000_ID)
    35 continue;
    36
    37 // how many bytes of registers does BAR 0 decode? write all
    38 // ones: the device keeps the address bits below its size zero.
    39 h[PCI_BAR0] = 0xffffffff;
    40 uint32 size = ~(h[PCI_BAR0] & ~0xf) + 1;
    41 if (size > E1000_REGS_SIZE)
    42 panic("pci: e1000 registers too big");
    43
    44 // place the registers at E1000_REGS, then let the device
    45 // decode them and do DMA.
    46 h[PCI_BAR0] = E1000_REGS;
    47 h[PCI_COMMAND] = PCI_COMMAND_MEMORY | PCI_COMMAND_MASTER;
    48
    49 int pin = (h[PCI_INTR] >> 8) & 0xff; // 1 = INTA ... 4 = INTD
    50 printk("pci: e1000 at 0:%d.0, %d bytes of registers at 0x%lx, INT%c\n",
    51 dev, size, E1000_REGS, 'A' + pin - 1);
    52 return;
    53 }
    54 panic("pci: no e1000");
    55}

    kernel/vm.c

    @@ -31,8 +31,14 @@ kvmmake(void)
    3131
    3232 // virtio mmio disk interface
    3333 kvmmap(kpgtbl, VIRTIO0, VIRTIO0, PGSIZE, PTE_R | PTE_W);
    3434
    35 // PCIe configuration space for bus 0
    36 kvmmap(kpgtbl, PCIE_ECAM, PCIE_ECAM, PCIE_ECAM_SIZE, PTE_R | PTE_W);
    37
    38 // e1000 registers
    39 kvmmap(kpgtbl, E1000_REGS, E1000_REGS, E1000_REGS_SIZE, PTE_R | PTE_W);
    40
    3541 // PLIC
    3642 kvmmap(kpgtbl, PLIC, PLIC, 0x4000000, PTE_R | PTE_W);
    3743
    3844 // map kernel text executable and read-only.
  2. cf55c0c Initialize the e1000 and its descriptor rings

    Makefile

    @@ -28,9 +28,10 @@ OBJS = \
    2828 $K/sysfile.o \
    2929 $K/kernelvec.o \
    3030 $K/plic.o \
    3131 $K/virtio_disk.o \
    32 $K/pci.o
    32 $K/pci.o \
    33 $K/e1000.o
    3334
    3435# riscv64-unknown-elf- or riscv64-linux-gnu-
    3536# perhaps in /opt/riscv/bin
    3637#TOOLPREFIX =

    kernel/defs.h

    @@ -176,8 +176,11 @@ void plicinit(void);
    176176void plicinithart(void);
    177177int plic_claim(void);
    178178void plic_complete(int);
    179179
    180// e1000.c
    181void e1000_init(void);
    182
    180183// pci.c
    181184void pci_init(void);
    182185
    183186// virtio_disk.c

    kernel/e1000.c

    @@ -0,0 +1,98 @@
    1//
    2// driver for qemu's e1000 network card (Intel 82540EM).
    3//
    4// qemu ... -netdev user,id=net0 -device e1000,netdev=net0,bus=pcie.0
    5//
    6
    7#include "types.h"
    8#include "param.h"
    9#include "memlayout.h"
    10#include "riscv.h"
    11#include "spinlock.h"
    12#include "defs.h"
    13#include "e1000_dev.h"
    14
    15// the address of e1000 register r (a byte offset).
    16#define R(r) ((volatile uint32 *)(E1000_REGS + (r)))
    17
    18#define TX_RING_SIZE 16
    19#define RX_RING_SIZE 16
    20
    21static struct e1000 {
    22 // the rings of descriptors the device reads by DMA. each lives
    23 // in a page from kalloc(); the kernel maps RAM one-to-one, so
    24 // the address of a descriptor is also its physical address.
    25 struct tx_desc *tx;
    26 struct rx_desc *rx;
    27
    28 // the packet buffer (a kalloc() page) in each slot.
    29 char *tx_bufs[TX_RING_SIZE];
    30 char *rx_bufs[RX_RING_SIZE];
    31
    32 struct spinlock lock;
    33} e1000;
    34
    35// QEMU's default MAC address for a network card.
    36uint8 local_mac[6] = {0x52, 0x54, 0x00, 0x12, 0x34, 0x56};
    37
    38// called by pci_init() once the registers are mapped at E1000_REGS.
    39void
    40e1000_init(void)
    41{
    42 int i;
    43
    44 initlock(&e1000.lock, "e1000");
    45
    46 // reset the device, with its interrupts masked.
    47 *R(E1000_IMC) = 0xffffffff;
    48 *R(E1000_CTL) |= E1000_CTL_RST;
    49 *R(E1000_IMC) = 0xffffffff;
    50 io_fence();
    51
    52 // transmit ring: every slot starts out "done", with no packet.
    53 e1000.tx = (struct tx_desc *)kalloc();
    54 if (e1000.tx == 0)
    55 panic("e1000_init: kalloc");
    56 memset(e1000.tx, 0, PGSIZE);
    57 for (i = 0; i < TX_RING_SIZE; i++) {
    58 e1000.tx[i].status = E1000_TXD_STAT_DD;
    59 e1000.tx_bufs[i] = 0;
    60 }
    61 *R(E1000_TDBAL) = (uint64)e1000.tx;
    62 *R(E1000_TDBAH) = (uint64)e1000.tx >> 32;
    63 *R(E1000_TDLEN) = TX_RING_SIZE * sizeof(struct tx_desc);
    64 *R(E1000_TDH) = 0;
    65 *R(E1000_TDT) = 0;
    66
    67 // receive ring: every slot gets an empty buffer.
    68 e1000.rx = (struct rx_desc *)kalloc();
    69 if (e1000.rx == 0)
    70 panic("e1000_init: kalloc");
    71 memset(e1000.rx, 0, PGSIZE);
    72 for (i = 0; i < RX_RING_SIZE; i++) {
    73 e1000.rx_bufs[i] = kalloc();
    74 if (e1000.rx_bufs[i] == 0)
    75 panic("e1000_init: kalloc");
    76 e1000.rx[i].addr = (uint64)e1000.rx_bufs[i];
    77 }
    78 *R(E1000_RDBAL) = (uint64)e1000.rx;
    79 *R(E1000_RDBAH) = (uint64)e1000.rx >> 32;
    80 *R(E1000_RDLEN) = RX_RING_SIZE * sizeof(struct rx_desc);
    81 // the device may fill the slots from RDH up to, but not
    82 // including, RDT.
    83 *R(E1000_RDH) = 0;
    84 *R(E1000_RDT) = RX_RING_SIZE - 1;
    85
    86 // accept packets for our MAC address (and broadcasts, below).
    87 *R(E1000_RA) = local_mac[0] | (local_mac[1] << 8) | (local_mac[2] << 16) |
    88 (local_mac[3] << 24);
    89 *R(E1000_RA + 4) = local_mac[4] | (local_mac[5] << 8) | E1000_RA_AV;
    90 for (i = 0; i < 128; i++)
    91 *R(E1000_MTA + 4 * i) = 0;
    92
    93 // turn on transmit and receive.
    94 *R(E1000_TCTL) = E1000_TCTL_EN | E1000_TCTL_PSP | E1000_TCTL_CT |
    95 E1000_TCTL_COLD;
    96 *R(E1000_TIPG) = 10 | (8 << 10) | (6 << 20);
    97 *R(E1000_RCTL) = E1000_RCTL_EN | E1000_RCTL_BAM | E1000_RCTL_SECRC;
    98}

    kernel/e1000_dev.h

    @@ -0,0 +1,71 @@
    1//
    2// the parts of the Intel 8254x ("e1000") interface that xv6 uses.
    3// see Intel's "PCI/PCI-X Family of Gigabit Ethernet Controllers
    4// Software Developer's Manual" (document 317453006EN).
    5//
    6
    7// registers: byte offsets from the start of BAR 0.
    8#define E1000_CTL 0x00000 // device control
    9#define E1000_ICR 0x000C0 // interrupt cause read (reading clears it)
    10#define E1000_IMS 0x000D0 // interrupt mask set
    11#define E1000_IMC 0x000D8 // interrupt mask clear
    12#define E1000_RCTL 0x00100 // receive control
    13#define E1000_TCTL 0x00400 // transmit control
    14#define E1000_TIPG 0x00410 // transmit inter-packet gap
    15#define E1000_RDBAL 0x02800 // receive ring address, low 32 bits
    16#define E1000_RDBAH 0x02804 // receive ring address, high 32 bits
    17#define E1000_RDLEN 0x02808 // receive ring length in bytes
    18#define E1000_RDH 0x02810 // receive head (the device's next slot)
    19#define E1000_RDT 0x02818 // receive tail (the driver's next slot)
    20#define E1000_RDTR 0x02820 // receive interrupt delay
    21#define E1000_TDBAL 0x03800 // transmit ring address, low 32 bits
    22#define E1000_TDBAH 0x03804 // transmit ring address, high 32 bits
    23#define E1000_TDLEN 0x03808 // transmit ring length in bytes
    24#define E1000_TDH 0x03810 // transmit head (the device's next slot)
    25#define E1000_TDT 0x03818 // transmit tail (the driver's next slot)
    26#define E1000_MTA 0x05200 // multicast table, 128 words
    27#define E1000_RA 0x05400 // receive address (our MAC), 2 words
    28
    29#define E1000_CTL_RST 0x04000000 // reset the device
    30
    31#define E1000_TCTL_EN 0x00000002 // enable transmit
    32#define E1000_TCTL_PSP 0x00000008 // pad short packets
    33#define E1000_TCTL_CT 0x00000100 // collision threshold 16
    34#define E1000_TCTL_COLD 0x00040000 // collision distance 64 (full duplex)
    35
    36#define E1000_RCTL_EN 0x00000002 // enable receive
    37#define E1000_RCTL_BAM 0x00008000 // accept broadcast packets
    38#define E1000_RCTL_SECRC 0x04000000 // strip the Ethernet CRC
    39// buffer size bits 0 in RCTL: 2048-byte receive buffers.
    40
    41#define E1000_RA_AV 0x80000000 // receive address valid
    42
    43#define E1000_ICR_RXT0 0x00000080 // receiver timer interrupt
    44
    45// transmit descriptor (legacy format, section 3.3.3).
    46struct tx_desc {
    47 uint64 addr; // physical address of the packet
    48 uint16 length;
    49 uint8 cso;
    50 uint8 cmd;
    51 uint8 status;
    52 uint8 css;
    53 uint16 special;
    54};
    55
    56#define E1000_TXD_CMD_EOP 0x01 // end of packet
    57#define E1000_TXD_CMD_RS 0x08 // report status: set DD when done
    58#define E1000_TXD_STAT_DD 0x01 // descriptor done
    59
    60// receive descriptor (section 3.2.3).
    61struct rx_desc {
    62 uint64 addr; // physical address of a 2048-byte buffer
    63 uint16 length;
    64 uint16 csum;
    65 uint8 status;
    66 uint8 errors;
    67 uint16 special;
    68};
    69
    70#define E1000_RXD_STAT_DD 0x01 // descriptor done
    71#define E1000_RXD_STAT_EOP 0x02 // end of packet

    kernel/pci.c

    @@ -48,8 +48,9 @@ pci_init(void)
    4848
    4949 int pin = (h[PCI_INTR] >> 8) & 0xff; // 1 = INTA ... 4 = INTD
    5050 printk("pci: e1000 at 0:%d.0, %d bytes of registers at 0x%lx, INT%c\n",
    5151 dev, size, E1000_REGS, 'A' + pin - 1);
    52 e1000_init();
    5253 return;
    5354 }
    5455 panic("pci: no e1000");
    5556}
  3. dddfa9a Transmit a packet

    kernel/defs.h

    @@ -178,8 +178,9 @@ int plic_claim(void);
    178178void plic_complete(int);
    179179
    180180// e1000.c
    181181void e1000_init(void);
    182int e1000_transmit(char *, int);
    182183
    183184// pci.c
    184185void pci_init(void);
    185186

    kernel/e1000.c

    @@ -95,4 +95,49 @@ e1000_init(void)
    9595 E1000_TCTL_COLD;
    9696 *R(E1000_TIPG) = 10 | (8 << 10) | (6 << 20);
    9797 *R(E1000_RCTL) = E1000_RCTL_EN | E1000_RCTL_BAM | E1000_RCTL_SECRC;
    9898}
    99
    100// hand the packet in buf (a kalloc() page) to the device. on
    101// success the driver owns buf and frees it once the device has
    102// sent it. returns -1 if the ring is full; the caller then still
    103// owns buf.
    104// may be called on any hart, from system calls and interrupts.
    105int
    106e1000_transmit(char *buf, int len)
    107{
    108 acquire(&e1000.lock);
    109
    110 // TDT is the next slot the driver may fill. the device sends
    111 // the slots from TDH up to, but not including, TDT, so TDT
    112 // equal to TDH means "nothing to send": TDT must never move
    113 // onto TDH, and one slot always stays empty. the slot must
    114 // also be done (DD): the device may still be reading it.
    115 int i = *R(E1000_TDT);
    116 if ((i + 1) % TX_RING_SIZE == *R(E1000_TDH) ||
    117 (e1000.tx[i].status & E1000_TXD_STAT_DD) == 0) {
    118 release(&e1000.lock);
    119 return -1;
    120 }
    121 // order our reads of the slot after the device's write of DD.
    122 io_fence();
    123
    124 // the device has sent the slot's previous packet: free it.
    125 if (e1000.tx_bufs[i])
    126 kfree(e1000.tx_bufs[i]);
    127 e1000.tx_bufs[i] = buf;
    128
    129 e1000.tx[i].addr = (uint64)buf;
    130 e1000.tx[i].length = len;
    131 e1000.tx[i].cmd = E1000_TXD_CMD_EOP | E1000_TXD_CMD_RS;
    132 e1000.tx[i].status = 0;
    133
    134 // the descriptor must be in memory before the device can see
    135 // the new tail and read it.
    136 io_fence();
    137
    138 // give the slot to the device.
    139 *R(E1000_TDT) = (i + 1) % TX_RING_SIZE;
    140
    141 release(&e1000.lock);
    142 return 0;
    143}
  4. 9d059c8 Receive packets from the e1000 interrupt

    Makefile

    @@ -29,9 +29,10 @@ OBJS = \
    2929 $K/kernelvec.o \
    3030 $K/plic.o \
    3131 $K/virtio_disk.o \
    3232 $K/pci.o \
    33 $K/e1000.o
    33 $K/e1000.o \
    34 $K/net.o
    3435
    3536# riscv64-unknown-elf- or riscv64-linux-gnu-
    3637# perhaps in /opt/riscv/bin
    3738#TOOLPREFIX =

    kernel/defs.h

    @@ -179,8 +179,12 @@ void plic_complete(int);
    179179
    180180// e1000.c
    181181void e1000_init(void);
    182182int e1000_transmit(char *, int);
    183void e1000_intr(void);
    184
    185// net.c
    186void net_rx(char *, int);
    183187
    184188// pci.c
    185189void pci_init(void);
    186190

    kernel/e1000.c

    @@ -94,8 +94,12 @@ e1000_init(void)
    9494 *R(E1000_TCTL) = E1000_TCTL_EN | E1000_TCTL_PSP | E1000_TCTL_CT |
    9595 E1000_TCTL_COLD;
    9696 *R(E1000_TIPG) = 10 | (8 << 10) | (6 << 20);
    9797 *R(E1000_RCTL) = E1000_RCTL_EN | E1000_RCTL_BAM | E1000_RCTL_SECRC;
    98
    99 // interrupt as soon as a packet has arrived (no delay).
    100 *R(E1000_RDTR) = 0;
    101 *R(E1000_IMS) = E1000_ICR_RXT0;
    98102}
    99103
    100104// hand the packet in buf (a kalloc() page) to the device. on
    101105// success the driver owns buf and frees it once the device has
    @@ -140,4 +144,60 @@ e1000_transmit(char *buf, int len)
    140144
    141145 release(&e1000.lock);
    142146 return 0;
    143147}
    148
    149// the e1000 raised an interrupt: some packets have arrived.
    150// called from devintr() with interrupts off, on whichever hart
    151// claimed the interrupt.
    152void
    153e1000_intr(void)
    154{
    155 char *bufs[RX_RING_SIZE];
    156 int lens[RX_RING_SIZE];
    157 int n = 0;
    158
    159 acquire(&e1000.lock);
    160
    161 // reading the interrupt cause clears it. a packet that arrives
    162 // after this read raises a new interrupt.
    163 *R(E1000_ICR);
    164
    165 // the slot after the tail is the oldest one the device may have
    166 // filled. take filled slots until there are none, but go at most
    167 // once around the ring.
    168 for (int k = 0; k < RX_RING_SIZE; k++) {
    169 int i = (*R(E1000_RDT) + 1) % RX_RING_SIZE;
    170 if ((e1000.rx[i].status & E1000_RXD_STAT_DD) == 0)
    171 break;
    172
    173 // order our reads of the descriptor and the packet after
    174 // the device's write of DD.
    175 io_fence();
    176
    177 // pass the full buffer up and put a fresh one in the slot.
    178 // if there is no memory, drop the packet and reuse the buffer.
    179 char *fresh = kalloc();
    180 if (fresh) {
    181 bufs[n] = e1000.rx_bufs[i];
    182 lens[n] = e1000.rx[i].length;
    183 n++;
    184 e1000.rx_bufs[i] = fresh;
    185 e1000.rx[i].addr = (uint64)fresh;
    186 }
    187 e1000.rx[i].status = 0;
    188
    189 // the descriptor must be in memory before the device can see
    190 // the new tail and fill the slot again.
    191 io_fence();
    192
    193 // give the slot back to the device.
    194 *R(E1000_RDT) = i;
    195 }
    196
    197 release(&e1000.lock);
    198
    199 // hand the packets to the network stack without holding
    200 // e1000.lock: answering a packet calls e1000_transmit().
    201 for (int j = 0; j < n; j++)
    202 net_rx(bufs[j], lens[j]);
    203}

    kernel/memlayout.h

    @@ -36,8 +36,13 @@
    3636// (BAR 0), inside the PCIe memory window at 0x40000000.
    3737#define E1000_REGS 0x40000000L
    3838#define E1000_REGS_SIZE 0x20000L
    3939
    40// qemu wires the PCIe interrupt pins INTA-INTD to PLIC IRQs
    41// 32-35, rotated by device number ("swizzled"): pin INTA of
    42// device 1, the e1000, is IRQ 32 + (0 + 1) % 4 = 33.
    43#define E1000_IRQ 33
    44
    4045// core-local interrupt controller (CLINT)
    4146#define CLINT_BASE 0x02000000L
    4247#define CLINT(hart) (CLINT_BASE + (hart) * 4)
    4348

    kernel/net.c

    @@ -0,0 +1,19 @@
    1//
    2// the network stack.
    3//
    4
    5#include "types.h"
    6#include "param.h"
    7#include "memlayout.h"
    8#include "riscv.h"
    9#include "spinlock.h"
    10#include "defs.h"
    11
    12// a packet of len bytes has arrived in buf (a kalloc() page).
    13// called from e1000_intr() with interrupts off; must not sleep.
    14// for now, drop it.
    15void
    16net_rx(char *buf, int len)
    17{
    18 kfree(buf);
    19}

    kernel/plic.c

    @@ -13,8 +13,9 @@ plicinit(void)
    1313{
    1414 // set desired IRQ priorities non-zero (otherwise disabled).
    1515 *(uint32 *)(PLIC + UART0_IRQ * 4) = 1;
    1616 *(uint32 *)(PLIC + VIRTIO0_IRQ * 4) = 1;
    17 *(uint32 *)(PLIC + E1000_IRQ * 4) = 1;
    1718}
    1819
    1920void
    2021plicinithart(void)
    @@ -23,8 +24,10 @@ plicinithart(void)
    2324
    2425 // set enable bits for this hart's S-mode
    2526 // for the uart and virtio disk.
    2627 *(uint32 *)PLIC_SENABLE(hart) = (1 << UART0_IRQ) | (1 << VIRTIO0_IRQ);
    28 // the enable bits for IRQs 32-63 are in the next word.
    29 *(uint32 *)(PLIC_SENABLE(hart) + 4) = (1 << (E1000_IRQ - 32));
    2730
    2831 // set this hart's S-mode priority threshold to 0.
    2932 *(uint32 *)PLIC_SPRIORITY(hart) = 0;
    3033}

    kernel/trap.c

    @@ -198,8 +198,10 @@ devintr()
    198198 if (irq == UART0_IRQ) {
    199199 uartintr();
    200200 } else if (irq == VIRTIO0_IRQ) {
    202 } else if (irq == E1000_IRQ) {
    203 e1000_intr();
    202204 } else if (irq) {
    203205 printk("unexpected interrupt irq=%d\n", irq);
    204206 }
    205207
  5. 36fada5 Answer ARP, parse IPv4 and UDP, send UDP

    kernel/defs.h

    @@ -180,11 +180,13 @@ void plic_complete(int);
    180180// e1000.c
    181181void e1000_init(void);
    182182int e1000_transmit(char *, int);
    183183void e1000_intr(void);
    184extern uint8 local_mac[6];
    184185
    185186// net.c
    186187void net_rx(char *, int);
    188int net_tx_udp(char *, int, uint32, uint16, uint16);
    187189
    188190// pci.c
    189191void pci_init(void);
    190192

    kernel/net.c

    @@ -1,19 +1,159 @@
    11//
    2// the network stack.
    2// the network stack: just enough Ethernet, ARP, IPv4 and UDP to
    3// talk to QEMU's user-mode network ("slirp").
    34//
    45
    56#include "types.h"
    67#include "param.h"
    78#include "memlayout.h"
    89#include "riscv.h"
    910#include "spinlock.h"
    1011#include "defs.h"
    12#include "net.h"
    13
    14// xv6's address on slirp's network.
    15static uint32 local_ip = MAKE_IP_ADDR(10, 0, 2, 15);
    16
    17// the MAC address with which slirp answers ARP for the host
    18// (10.0.2.2). slirp takes IP packets for the host, its DNS
    19// server (10.0.2.3) and the rest of the world without checking
    20// their destination MAC, so xv6 sends every packet here and
    21// needs no ARP requests of its own.
    22static uint8 host_mac[6] = {0x52, 0x55, 0x0a, 0x00, 0x02, 0x02};
    23
    24// the Internet checksum (RFC 1071) of len bytes at p,
    25// in network byte order.
    26static uint16
    27in_cksum(uint8 *p, int len)
    28{
    29 uint32 sum = 0;
    30
    31 for (; len > 1; len -= 2, p += 2)
    32 sum += (p[0] << 8) | p[1];
    33 if (len == 1)
    34 sum += p[0] << 8;
    35 while (sum >> 16)
    36 sum = (sum & 0xffff) + (sum >> 16);
    37 return htons(~sum);
    38}
    39
    40// send a UDP packet to dip:dport from local port sport. buf is a
    41// kalloc() page holding len bytes of data at offset NET_HDRS;
    42// this fills in the headers in front of the data. buf belongs to
    43// the driver afterwards, or is freed on failure.
    44int
    45net_tx_udp(char *buf, int len, uint32 dip, uint16 dport, uint16 sport)
    46{
    47 struct eth *eth = (struct eth *)buf;
    48 struct ip *ip = (struct ip *)(eth + 1);
    49 struct udp *udp = (struct udp *)(ip + 1);
    50
    51 memmove(eth->dhost, host_mac, 6);
    52 memmove(eth->shost, local_mac, 6);
    53 eth->type = htons(ETHTYPE_IP);
    54
    55 memset(ip, 0, sizeof(*ip));
    56 ip->ip_vhl = 0x45;
    57 ip->ip_len = htons(sizeof(*ip) + sizeof(*udp) + len);
    58 ip->ip_ttl = 64;
    59 ip->ip_p = IPPROTO_UDP;
    60 ip->ip_src = htonl(local_ip);
    61 ip->ip_dst = htonl(dip);
    62 ip->ip_sum = in_cksum((uint8 *)ip, sizeof(*ip));
    63
    64 udp->sport = htons(sport);
    65 udp->dport = htons(dport);
    66 udp->ulen = htons(sizeof(*udp) + len);
    67 udp->sum = 0;
    68
    69 if (e1000_transmit(buf, NET_HDRS + len) < 0) {
    70 kfree(buf);
    71 return -1;
    72 }
    73 return 0;
    74}
    75
    76// answer an ARP request for local_ip, turning the request's
    77// buffer into the reply.
    78static void
    79arp_rx(char *buf, int len)
    80{
    81 struct eth *eth = (struct eth *)buf;
    82 struct arp *arp = (struct arp *)(eth + 1);
    83
    84 if (len < sizeof(*eth) + sizeof(*arp) || ntohs(arp->hrd) != ARP_HRD_ETHER ||
    85 ntohs(arp->pro) != ETHTYPE_IP || arp->hln != 6 || arp->pln != 4 ||
    86 ntohs(arp->op) != ARP_OP_REQUEST || ntohl(arp->tip) != local_ip) {
    87 kfree(buf);
    88 return;
    89 }
    90
    91 arp->op = htons(ARP_OP_REPLY);
    92 memmove(arp->tha, arp->sha, 6);
    93 arp->tip = arp->sip;
    94 memmove(arp->sha, local_mac, 6);
    95 arp->sip = htonl(local_ip);
    96
    97 memmove(eth->dhost, eth->shost, 6);
    98 memmove(eth->shost, local_mac, 6);
    99
    100 if (e1000_transmit(buf, sizeof(*eth) + sizeof(*arp)) < 0)
    101 kfree(buf);
    102}
    103
    104// a UDP packet with len bytes of data at buf + NET_HDRS has
    105// arrived for local port dport. for now, drop it.
    106static void
    107udp_rx(char *buf, int len, uint32 sip, uint16 sport, uint16 dport)
    108{
    109 kfree(buf);
    110}
    111
    112// check an IPv4 packet's headers and pass UDP packets on.
    113static void
    114ip_rx(char *buf, int len)
    115{
    116 struct ip *ip = (struct ip *)(buf + sizeof(struct eth));
    117 struct udp *udp = (struct udp *)(ip + 1);
    118
    119 // xv6 understands only plain UDP packets for itself: no IP
    120 // options, no fragments.
    121 if (len < NET_HDRS || ip->ip_vhl != 0x45 || ip->ip_p != IPPROTO_UDP ||
    122 ntohl(ip->ip_dst) != local_ip || (ntohs(ip->ip_off) & 0x3fff) != 0 ||
    123 in_cksum((uint8 *)ip, sizeof(*ip)) != 0) {
    124 kfree(buf);
    125 return;
    126 }
    127
    128 // the lengths in the headers must fit in what arrived.
    129 int iplen = ntohs(ip->ip_len);
    130 int ulen = ntohs(udp->ulen);
    131 if (iplen > len - (int)sizeof(struct eth) || ulen < (int)sizeof(*udp) ||
    132 ulen > iplen - (int)sizeof(*ip)) {
    133 kfree(buf);
    134 return;
    135 }
    136
    137 udp_rx(buf, ulen - sizeof(*udp), ntohl(ip->ip_src), ntohs(udp->sport),
    138 ntohs(udp->dport));
    139}
    11140
    12141// a packet of len bytes has arrived in buf (a kalloc() page).
    13// called from e1000_intr() with interrupts off; must not sleep.
    14// for now, drop it.
    142// called from e1000_intr() with interrupts off: must not sleep.
    143// every path ends by freeing buf or passing it on.
    15144void
    16145net_rx(char *buf, int len)
    17146{
    18 kfree(buf);
    147 struct eth *eth = (struct eth *)buf;
    148
    149 if (len < sizeof(*eth)) {
    150 kfree(buf);
    151 return;
    152 }
    153 if (ntohs(eth->type) == ETHTYPE_ARP)
    154 arp_rx(buf, len);
    155 else if (ntohs(eth->type) == ETHTYPE_IP)
    156 ip_rx(buf, len);
    157 else
    158 kfree(buf);
    19159}

    kernel/net.h

    @@ -0,0 +1,81 @@
    1//
    2// packet formats for Ethernet, ARP, IPv4 and UDP.
    3// all multi-byte fields are big-endian ("network byte order").
    4//
    5
    6// RISC-V is little-endian: swap the bytes of a 16- or 32-bit
    7// value between host and network byte order.
    8static inline uint16
    9htons(uint16 x)
    10{
    11 return (x << 8) | (x >> 8);
    12}
    13
    14static inline uint32
    15htonl(uint32 x)
    16{
    17 return ((x & 0xff) << 24) | ((x & 0xff00) << 8) | ((x >> 8) & 0xff00) |
    18 (x >> 24);
    19}
    20
    21#define ntohs htons
    22#define ntohl htonl
    23
    24#define MAKE_IP_ADDR(a, b, c, d) \
    25 (((uint32)(a) << 24) | ((uint32)(b) << 16) | ((uint32)(c) << 8) | (uint32)(d))
    26
    27// the headers are packed: inside a packet, a 32-bit field may
    28// sit at an address that is not a multiple of 4.
    29
    30struct eth {
    31 uint8 dhost[6];
    32 uint8 shost[6];
    33 uint16 type;
    34} __attribute__((packed));
    35
    36#define ETHTYPE_IP 0x0800
    37#define ETHTYPE_ARP 0x0806
    38
    39struct arp {
    40 uint16 hrd; // hardware type: 1 = Ethernet
    41 uint16 pro; // protocol type: ETHTYPE_IP
    42 uint8 hln; // hardware address length: 6
    43 uint8 pln; // protocol address length: 4
    44 uint16 op;
    45 uint8 sha[6]; // sender's MAC address
    46 uint32 sip; // sender's IP address
    47 uint8 tha[6]; // target's MAC address
    48 uint32 tip; // target's IP address
    49} __attribute__((packed));
    50
    51#define ARP_HRD_ETHER 1
    52#define ARP_OP_REQUEST 1
    53#define ARP_OP_REPLY 2
    54
    55struct ip {
    56 uint8 ip_vhl; // version (4) and header length in words (5)
    57 uint8 ip_tos;
    58 uint16 ip_len; // header and data, in bytes
    59 uint16 ip_id;
    60 uint16 ip_off; // fragment flags and offset
    61 uint8 ip_ttl;
    62 uint8 ip_p; // protocol: IPPROTO_UDP
    63 uint16 ip_sum;
    64 uint32 ip_src;
    65 uint32 ip_dst;
    66} __attribute__((packed));
    67
    68#define IPPROTO_UDP 17
    69
    70struct udp {
    71 uint16 sport;
    72 uint16 dport;
    73 uint16 ulen; // header and data, in bytes
    74 uint16 sum; // 0: no checksum
    75} __attribute__((packed));
    76
    77// a UDP packet's data starts this many bytes into the frame.
    78#define NET_HDRS (sizeof(struct eth) + sizeof(struct ip) + sizeof(struct udp))
    79
    80// the most data one UDP packet may carry (1500-byte IP packets).
    81#define NET_MAXDATA (1500 - sizeof(struct ip) - sizeof(struct udp))
  6. 9149741 Add UDP sockets: bind, unbind, send and recv

    Makefile

    @@ -30,9 +30,10 @@ OBJS = \
    3030 $K/plic.o \
    3131 $K/virtio_disk.o \
    3232 $K/pci.o \
    3333 $K/e1000.o \
    34 $K/net.o
    34 $K/net.o \
    35 $K/socket.o
    3536
    3637# riscv64-unknown-elf- or riscv64-linux-gnu-
    3738# perhaps in /opt/riscv/bin
    3839#TOOLPREFIX =

    kernel/defs.h

    @@ -186,8 +186,12 @@ extern uint8 local_mac[6];
    186186// net.c
    187187void net_rx(char *, int);
    188188int net_tx_udp(char *, int, uint32, uint16, uint16);
    189189
    190// socket.c
    191void sockinit(void);
    192void sock_rx(char *, int, uint32, uint16, uint16);
    193
    190194// pci.c
    191195void pci_init(void);
    192196
    193197// virtio_disk.c

    kernel/main.c

    @@ -27,8 +27,9 @@ main()
    2727 binit(); // buffer cache
    2828 iinit(); // inode table
    2929 fileinit(); // file table
    3030 virtio_disk_init(); // emulated hard disk
    31 sockinit(); // UDP sockets
    3132 pci_init(); // find the e1000 network card
    3233 userinit(); // first user process
    3334
    3435 __atomic_store_n(&started, 1, __ATOMIC_RELEASE);

    kernel/net.c

    @@ -100,16 +100,8 @@ arp_rx(char *buf, int len)
    100100 if (e1000_transmit(buf, sizeof(*eth) + sizeof(*arp)) < 0)
    101101 kfree(buf);
    102102}
    103103
    104// a UDP packet with len bytes of data at buf + NET_HDRS has
    105// arrived for local port dport. for now, drop it.
    106static void
    107udp_rx(char *buf, int len, uint32 sip, uint16 sport, uint16 dport)
    108{
    109 kfree(buf);
    110}
    111
    112104// check an IPv4 packet's headers and pass UDP packets on.
    113105static void
    114106ip_rx(char *buf, int len)
    115107{
    @@ -133,10 +125,10 @@ ip_rx(char *buf, int len)
    133125 kfree(buf);
    134126 return;
    135127 }
    136128
    137 udp_rx(buf, ulen - sizeof(*udp), ntohl(ip->ip_src), ntohs(udp->sport),
    138 ntohs(udp->dport));
    129 sock_rx(buf, ulen - sizeof(*udp), ntohl(ip->ip_src), ntohs(udp->sport),
    130 ntohs(udp->dport));
    139131}
    140132
    141133// a packet of len bytes has arrived in buf (a kalloc() page).
    142134// called from e1000_intr() with interrupts off: must not sleep.

    kernel/socket.c

    @@ -0,0 +1,247 @@
    1//
    2// UDP sockets: bind a local port, then send and receive
    3// datagrams through it.
    4//
    5
    6#include "types.h"
    7#include "param.h"
    8#include "memlayout.h"
    9#include "riscv.h"
    10#include "spinlock.h"
    11#include "proc.h"
    12#include "defs.h"
    13#include "net.h"
    14
    15#define NSOCK 16 // sockets in the whole system
    16#define NQ 16 // packets queued on one socket, at most
    17
    18struct sock {
    19 struct spinlock lock;
    20 int port; // the bound local port, or 0 if free. changes only
    21 // with both socktab.lock and this lock held.
    22
    23 // packets that arrived and wait for recv(): q[head] is the
    24 // oldest, n of them in all.
    25 int head;
    26 int n;
    27 struct {
    28 char *buf; // the packet's kalloc() page; data at NET_HDRS
    29 int len;
    30 uint32 sip;
    31 uint16 sport;
    32 } q[NQ];
    33};
    34
    35struct {
    36 struct spinlock lock; // protects the choice of each sock's port
    37 struct sock sock[NSOCK];
    38} socktab;
    39
    40void
    41sockinit(void)
    42{
    43 initlock(&socktab.lock, "socktab");
    44 for (int i = 0; i < NSOCK; i++)
    45 initlock(&socktab.sock[i].lock, "sock");
    46}
    47
    48// find the socket bound to port and return it locked, or 0.
    49// the caller must hold socktab.lock.
    50static struct sock *
    51sockfind(int port)
    52{
    53 for (int i = 0; i < NSOCK; i++) {
    54 struct sock *s = &socktab.sock[i];
    55 if (s->port == port) {
    56 acquire(&s->lock);
    57 return s;
    58 }
    59 }
    60 return 0;
    61}
    62
    63// called from ip_rx(), inside the e1000 interrupt: queue the
    64// packet on the socket bound to dport, or drop it. must not sleep.
    65void
    66sock_rx(char *buf, int len, uint32 sip, uint16 sport, uint16 dport)
    67{
    68 // port 0 marks a free socket; no packet may go there.
    69 if (dport == 0) {
    70 kfree(buf);
    71 return;
    72 }
    73
    74 acquire(&socktab.lock);
    75 struct sock *s = sockfind(dport);
    76 release(&socktab.lock);
    77 if (s == 0) {
    78 kfree(buf);
    79 return;
    80 }
    81
    82 if (s->n == NQ) {
    83 // nobody is reading fast enough: drop the packet.
    84 release(&s->lock);
    85 kfree(buf);
    86 return;
    87 }
    88 int i = (s->head + s->n) % NQ;
    89 s->q[i].buf = buf;
    90 s->q[i].len = len;
    91 s->q[i].sip = sip;
    92 s->q[i].sport = sport;
    93 s->n++;
    94 wakeup(s);
    95 release(&s->lock);
    96}
    97
    98// int bind(int port)
    99uint64
    100sys_bind(void)
    101{
    102 int port;
    103
    104 argint(0, &port);
    105 if (port <= 0 || port > 0xffff)
    106 return -1;
    107
    108 acquire(&socktab.lock);
    109 struct sock *s = sockfind(port);
    110 if (s) {
    111 // already bound.
    112 release(&s->lock);
    113 release(&socktab.lock);
    114 return -1;
    115 }
    116 s = sockfind(0);
    117 if (s == 0) {
    118 release(&socktab.lock);
    119 return -1;
    120 }
    121 s->port = port;
    122 s->head = 0;
    123 s->n = 0;
    124 release(&s->lock);
    125 release(&socktab.lock);
    126 return 0;
    127}
    128
    129// int unbind(int port)
    130uint64
    131sys_unbind(void)
    132{
    133 int port;
    134
    135 argint(0, &port);
    136 if (port <= 0 || port > 0xffff)
    137 return -1;
    138
    139 acquire(&socktab.lock);
    140 struct sock *s = sockfind(port);
    141 if (s == 0) {
    142 release(&socktab.lock);
    143 return -1;
    144 }
    145 s->port = 0;
    146 release(&socktab.lock);
    147
    148 for (; s->n > 0; s->n--) {
    149 kfree(s->q[s->head].buf);
    150 s->head = (s->head + 1) % NQ;
    151 }
    152 // a process sleeping in recv() on this port must give up.
    153 wakeup(s);
    154 release(&s->lock);
    155 return 0;
    156}
    157
    158// int send(uint32 dst, int dport, int sport, char *buf, int len)
    159uint64
    160sys_send(void)
    161{
    162 struct proc *p = myproc();
    163 int dst, dport, sport, len;
    164 uint64 addr;
    165
    166 argint(0, &dst);
    167 argint(1, &dport);
    168 argint(2, &sport);
    169 argaddr(3, &addr);
    170 argint(4, &len);
    171 if (len < 0 || len > NET_MAXDATA || dport <= 0 || dport > 0xffff ||
    172 sport <= 0 || sport > 0xffff)
    173 return -1;
    174
    175 char *buf = kalloc();
    176 if (buf == 0)
    177 return -1;
    178 if (copyin(p->pagetable, p->sz, buf + NET_HDRS, addr, len) < 0) {
    179 kfree(buf);
    180 return -1;
    181 }
    182 return net_tx_udp(buf, len, dst, dport, sport);
    183}
    184
    185// int recv(int port, uint32 *src, uint16 *sport, char *buf, int len)
    186// wait for a packet on a bound port. returns the number of data
    187// bytes copied to buf (the rest of a longer packet is lost).
    188uint64
    189sys_recv(void)
    190{
    191 struct proc *p = myproc();
    192 int port, maxlen;
    193 uint64 srcaddr, sportaddr, addr;
    194
    195 argint(0, &port);
    196 argaddr(1, &srcaddr);
    197 argaddr(2, &sportaddr);
    198 argaddr(3, &addr);
    199 argint(4, &maxlen);
    200 if (port <= 0 || port > 0xffff || maxlen < 0)
    201 return -1;
    202
    203 acquire(&socktab.lock);
    204 struct sock *s = sockfind(port);
    205 release(&socktab.lock);
    206 if (s == 0)
    207 return -1;
    208
    209 // sleep until a packet arrives. sock_rx() queues it and calls
    210 // wakeup() while holding s->lock, so registering on the channel
    211 // before releasing s->lock means the wakeup cannot be missed.
    212 // give up if the port was unbound (and maybe bound again to
    213 // another port) or the process was killed meanwhile.
    214 while (1) {
    215 if (s->port != port || killed(p)) {
    216 release(&s->lock);
    217 return -1;
    218 }
    219 if (s->n > 0)
    220 break;
    222 release(&s->lock);
    223 sleep();
    224 acquire(&s->lock);
    225 }
    226
    227 // take the oldest packet off the queue.
    228 char *buf = s->q[s->head].buf;
    229 int len = s->q[s->head].len;
    230 uint32 sip = s->q[s->head].sip;
    231 uint16 sport = s->q[s->head].sport;
    232 s->head = (s->head + 1) % NQ;
    233 s->n--;
    234 release(&s->lock);
    235
    236 // copy out without holding a spinlock.
    237 if (len > maxlen)
    238 len = maxlen;
    239 int r = len;
    240 if (copyout(p->pagetable, p->sz, srcaddr, (char *)&sip, sizeof(sip)) < 0 ||
    241 copyout(p->pagetable, p->sz, sportaddr, (char *)&sport,
    242 sizeof(sport)) < 0 ||
    243 copyout(p->pagetable, p->sz, addr, buf + NET_HDRS, len) < 0)
    244 r = -1;
    245 kfree(buf);
    246 return r;
    247}

    kernel/syscall.c

    @@ -102,8 +102,12 @@ extern uint64 sys_unlink(void);
    102102extern uint64 sys_link(void);
    103103extern uint64 sys_mkdir(void);
    104104extern uint64 sys_close(void);
    105105extern uint64 sys_sync(void);
    106extern uint64 sys_bind(void);
    107extern uint64 sys_unbind(void);
    108extern uint64 sys_send(void);
    109extern uint64 sys_recv(void);
    106110
    107111// An array mapping syscall numbers from syscall.h
    108112// to the function that handles the system call.
    109113static uint64 (*syscalls[])(void) = {
    @@ -129,8 +133,12 @@ static uint64 (*syscalls[])(void) = {
    129133 [SYS_link] = sys_link,
    130134 [SYS_mkdir] = sys_mkdir,
    131135 [SYS_close] = sys_close,
    132136 [SYS_sync] = sys_sync,
    137 [SYS_bind] = sys_bind,
    138 [SYS_unbind] = sys_unbind,
    139 [SYS_send] = sys_send,
    140 [SYS_recv] = sys_recv,
    133141 // clang-format on
    134142};
    135143
    136144void

    kernel/syscall.h

    @@ -20,4 +20,8 @@
    2020#define SYS_link 19
    2121#define SYS_mkdir 20
    2222#define SYS_close 21
    2323#define SYS_sync 22
    24#define SYS_bind 23
    25#define SYS_unbind 24
    26#define SYS_send 25
    27#define SYS_recv 26

    user/user.h

    @@ -24,8 +24,12 @@ int getpid(void);
    2424char *sys_sbrk(int, int);
    2525int pause(int);
    2626int uptime(void);
    2727int sync(void);
    28int bind(int);
    29int unbind(int);
    30int send(uint32, uint16, uint16, char *, int);
    31int recv(uint16, uint32 *, uint16 *, char *, int);
    2832
    2933// ulib.c
    3034int stat(const char *, struct stat *);
    3135char *strcpy(char *, const char *);

    user/usys.pl

    @@ -42,4 +42,8 @@ entry("getpid");
    4242entry("sbrk");
    4343entry("pause");
    4444entry("uptime");
    4545entry("sync");
    46entry("bind");
    47entry("unbind");
    48entry("send");
    49entry("recv");
  7. e3b8f88 Add nettest and its host-side partner, nettest.py

    Makefile

    @@ -153,8 +153,9 @@ UPROGS=\
    153153 $U/_logstress\
    154154 $U/_forphan\
    155155 $U/_dorphan\
    156156 $U/_sync\
    157 $U/_nettest\
    157158
    158159fs.img: mkfs/mkfs README $(UPROGS)
    159160 mkfs/mkfs fs.img README $(UPROGS)
    160161

    nettest.py

    @@ -0,0 +1,68 @@
    1#!/usr/bin/env python3
    2#
    3# the host side of xv6's nettest. QEMU's user-mode network (slirp)
    4# shows the host to xv6 as 10.0.2.2, and forwards the host's UDP
    5# port FWDPORT (see the Makefile) to xv6's port 2000.
    6#
    7# python3 nettest.py echo PORT echo every UDP packet that
    8# arrives at PORT back to its sender;
    9# then run "nettest PORT" in xv6.
    10# python3 nettest.py ping FWDPORT N first run "nettest server N" in
    11# xv6; then send N packets to it,
    12# one at a time, and time each
    13# round trip.
    14#
    15
    16import socket
    17import sys
    18import time
    19
    20
    21def echo(port, idle=60):
    22 s = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
    23 s.bind(("127.0.0.1", port))
    24 s.settimeout(idle)
    25 n = 0
    26 print(f"nettest.py: echoing on port {port}", flush=True)
    27 try:
    28 while True:
    29 data, addr = s.recvfrom(4096)
    30 s.sendto(data, addr)
    31 n += 1
    32 except socket.timeout:
    33 pass
    34 print(f"nettest.py: echoed {n} packets, then {idle} s without one", flush=True)
    35
    36
    37def ping(port, n):
    38 s = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
    39 s.settimeout(5)
    40 rtts = []
    41 for i in range(n):
    42 msg = b"ping %d" % i
    43 t0 = time.perf_counter()
    44 s.sendto(msg, ("127.0.0.1", port))
    45 try:
    46 data, _ = s.recvfrom(4096)
    47 except socket.timeout:
    48 print(f"nettest.py: ping {i}: no reply in 5 s: FAIL")
    49 sys.exit(1)
    50 rtts.append(time.perf_counter() - t0)
    51 if data != msg:
    52 print(f"nettest.py: ping {i}: wrong reply {data!r}: FAIL")
    53 sys.exit(1)
    54 rtts.sort()
    55 us = lambda x: round(x * 1e6)
    56 print(f"nettest.py: {n} round trips: OK")
    57 print(f"nettest.py: round trip in microseconds: min {us(rtts[0])}, "
    58 f"median {us(rtts[n // 2])}, 99th percentile {us(rtts[n * 99 // 100])}, "
    59 f"max {us(rtts[-1])}")
    60
    61
    62if len(sys.argv) >= 3 and sys.argv[1] == "echo":
    63 echo(int(sys.argv[2]), int(sys.argv[3]) if len(sys.argv) > 3 else 60)
    64elif len(sys.argv) == 4 and sys.argv[1] == "ping":
    65 ping(int(sys.argv[2]), int(sys.argv[3]))
    66else:
    67 print("usage: nettest.py echo PORT [IDLE] | nettest.py ping FWDPORT N")
    68 sys.exit(1)

    user/nettest.c

    @@ -0,0 +1,295 @@
    1//
    2// test the e1000 driver, the network stack and UDP sockets
    3// against nettest.py on the host. see the comment there.
    4//
    5// host: python3 nettest.py echo PORT then xv6: nettest PORT
    6// xv6: nettest server N then host: python3 nettest.py ping FWDPORT N
    7//
    8
    9#include "kernel/types.h"
    10#include "kernel/stat.h"
    11#include "user/user.h"
    12
    13#define HOST MAKE_IP(10, 0, 2, 2) // the host, as slirp shows it to xv6
    14#define DNS MAKE_IP(10, 0, 2, 3) // slirp's DNS server
    15#define MAKE_IP(a, b, c, d) (((a) << 24) | ((b) << 16) | ((c) << 8) | (d))
    16
    17int hostport;
    18
    19// the bytes of round seq of process id.
    20void
    21fill(char *buf, int id, int seq, int len)
    22{
    23 for (int i = 0; i < len; i++)
    24 buf[i] = 'a' + (id * 7 + seq * 3 + i) % 26;
    25}
    26
    27// one round trip through the host's echo server from local
    28// port lport. returns 0 if the echo is exactly what was sent.
    29int
    30roundtrip(int lport, int id, int seq, int len)
    31{
    32 static char out[1500], in[1500];
    33 uint32 src;
    34 uint16 sport;
    35
    36 fill(out, id, seq, len);
    37 if (send(HOST, hostport, lport, out, len) < 0) {
    38 printf("nettest: send failed\n");
    39 return -1;
    40 }
    41 int n = recv(lport, &src, &sport, in, sizeof(in));
    42 if (n != len || src != HOST || sport != hostport || memcmp(in, out, len)) {
    43 printf("nettest: process %d round %d: bad echo (%d bytes, from 0x%x:%d)\n",
    44 id, seq, n, src, sport);
    45 return -1;
    46 }
    47 return 0;
    48}
    49
    50// n round trips from lport with lengths from 1 to 1472 bytes.
    51int
    52roundtrips(int lport, int id, int n)
    53{
    54 unbind(lport); // in case a failed earlier run left it bound
    55 if (bind(lport) < 0) {
    56 printf("nettest: bind %d failed\n", lport);
    57 return -1;
    58 }
    59 for (int seq = 0; seq < n; seq++) {
    60 if (roundtrip(lport, id, seq, 1 + (seq * 97 + id * 31) % 1472) < 0)
    61 return -1;
    62 }
    63 unbind(lport);
    64 return 0;
    65}
    66
    67// run f(arg) in a child; kill it after timeout ticks.
    68// returns the child's exit status, or -1 if it was killed.
    69int
    70run(int (*f)(int), int arg, int timeout)
    71{
    72 int pid = fork();
    73 if (pid == 0)
    74 exit(f(arg) == 0 ? 0 : 1);
    75 int dog = fork();
    76 if (dog == 0) {
    77 pause(timeout);
    78 kill(pid);
    79 exit(0);
    80 }
    81 // wait for both; the order in which they exit varies.
    82 int xs = -1, st, w;
    83 while ((w = wait(&st)) >= 0) {
    84 if (w == pid) {
    85 xs = st;
    86 kill(dog);
    87 }
    88 }
    89 return xs;
    90}
    91
    92int
    93one(int arg)
    94{
    95 return roundtrips(2001, 0, 1);
    96}
    97
    98int
    99many(int arg)
    100{
    101 return roundtrips(2002, 0, 100);
    102}
    103
    104int
    105pingpong(int n)
    106{
    107 unbind(2003);
    108 if (bind(2003) < 0)
    109 return -1;
    110 for (int seq = 0; seq < n; seq++) {
    111 if (roundtrip(2003, 0, seq, 64) < 0)
    112 return -1;
    113 }
    114 unbind(2003);
    115 return 0;
    116}
    117
    118// three processes, one per hart, each with its own port.
    119int
    120parallel(int npkt)
    121{
    122 for (int i = 0; i < 3; i++) {
    123 if (fork() == 0)
    124 exit(roundtrips(2010 + i, i + 1, npkt) == 0 ? 0 : 1);
    125 }
    126 int ok = 1;
    127 for (int i = 0; i < 3; i++) {
    128 int xs;
    129 wait(&xs);
    130 if (xs != 0)
    131 ok = 0;
    132 }
    133 return ok ? 0 : -1;
    134}
    135
    136// recv() on a port nobody sends to sleeps; kill() must end it.
    137int
    138killrecv(int arg)
    139{
    140 char buf[16];
    141 uint32 src;
    142 uint16 sport;
    143
    144 unbind(2020);
    145 if (bind(2020) < 0)
    146 return -1;
    147 int pid = fork();
    148 if (pid == 0) {
    149 recv(2020, &src, &sport, buf, sizeof(buf));
    150 exit(0); // not reached if recv() returns because of the kill
    151 }
    152 pause(5);
    153 kill(pid);
    154 int xs;
    155 wait(&xs);
    156 unbind(2020);
    157 return xs == -1 ? 0 : -1;
    158}
    159
    160// ask slirp's DNS server for pdos.csail.mit.edu's address.
    161int
    162dns(int arg)
    163{
    164 static char q[64], buf[512];
    165 uchar *r = (uchar *)buf;
    166 char *name = "\4pdos\5csail\3mit\3edu";
    167 uint32 src;
    168 uint16 sport;
    169
    170 // header: id 0x1234, recursion desired, one question.
    171 memset(q, 0, sizeof(q));
    172 q[0] = 0x12;
    173 q[1] = 0x34;
    174 q[2] = 0x01;
    175 q[5] = 1;
    176 int n = 12;
    177 strcpy(q + n, name);
    178 n += strlen(name) + 1;
    179 q[n + 1] = 1; // type A
    180 q[n + 3] = 1; // class IN
    181 n += 4;
    182
    183 unbind(2030);
    184 if (bind(2030) < 0)
    185 return -1;
    186 if (send(DNS, 53, 2030, q, n) < 0)
    187 return -1;
    188 int len = recv(2030, &src, &sport, buf, sizeof(buf));
    189 unbind(2030);
    190 if (len < 12 || src != DNS || sport != 53 || r[0] != 0x12 || r[1] != 0x34 ||
    191 (r[3] & 0xf) != 0) {
    192 printf("nettest: dns: bad reply (%d bytes)\n", len);
    193 return -1;
    194 }
    195
    196 // skip the question (the same name, type and class), then look
    197 // for an A record among the answers.
    198 int nans = (r[6] << 8) | r[7];
    199 int i = n;
    200 for (int a = 0; a < nans && i + 12 <= len; a++) {
    201 if ((r[i] & 0xc0) == 0xc0) {
    202 i += 2; // the name, compressed
    203 } else {
    204 while (i < len && r[i] != 0)
    205 i += r[i] + 1;
    206 i += 1;
    207 }
    208 int type = (r[i] << 8) | r[i + 1];
    209 int rdlen = (r[i + 8] << 8) | r[i + 9];
    210 i += 10;
    211 if (type == 1 && rdlen == 4 && i + 4 <= len) {
    212 printf("nettest: dns: pdos.csail.mit.edu is %d.%d.%d.%d\n", r[i],
    213 r[i + 1], r[i + 2], r[i + 3]);
    214 return 0;
    215 }
    216 i += rdlen;
    217 }
    218 printf("nettest: dns: no A record in the reply\n");
    219 return -1;
    220}
    221
    222void
    223check(char *what, int xs, int *ok)
    224{
    225 if (xs == 0) {
    226 printf("nettest: %s: OK\n", what);
    227 } else {
    228 printf("nettest: %s: FAIL%s\n", what, xs == -1 ? " (timed out)" : "");
    229 *ok = 0;
    230 }
    231}
    232
    233// echo n packets that arrive at port 2000 back to their sender.
    234void
    235server(int n)
    236{
    237 static char buf[1500];
    238 uint32 src;
    239 uint16 sport;
    240
    241 unbind(2000);
    242 if (bind(2000) < 0) {
    243 printf("nettest server: bind 2000 failed\n");
    244 exit(1);
    245 }
    246 printf("nettest server: echoing %d packets on port 2000\n", n);
    247 for (int i = 0; i < n; i++) {
    248 int len = recv(2000, &src, &sport, buf, sizeof(buf));
    249 if (len < 0 || send(src, sport, 2000, buf, len) < 0) {
    250 printf("nettest server: failed\n");
    251 exit(1);
    252 }
    253 }
    254 unbind(2000);
    255 printf("nettest server: done\n");
    256 exit(0);
    257}
    258
    259int
    260main(int argc, char *argv[])
    261{
    262 int ok = 1;
    263
    264 if (argc == 3 && strcmp(argv[1], "server") == 0)
    265 server(atoi(argv[2]));
    266 if (argc != 2) {
    267 printf("usage: nettest hostport | nettest server n\n");
    268 exit(1);
    269 }
    270 hostport = atoi(argv[1]);
    271
    272 check("one round trip", run(one, 0, 50), &ok);
    273 check("100 round trips", run(many, 0, 100), &ok);
    274 check("3 processes x 300 round trips", run(parallel, 300, 600), &ok);
    275 check("kill wakes recv", run(killrecv, 0, 50), &ok);
    276 int dnsok = 1;
    277 check("dns query to 10.0.2.3", run(dns, 0, 50), &dnsok);
    278
    279 // latency, in clock ticks (about 0.1 s each).
    280 int t0 = uptime();
    281 int xs = run(pingpong, 1000, 600);
    282 int t1 = uptime();
    283 if (xs == 0)
    284 printf("nettest: 1000 round trips of 64 bytes took %d ticks\n", t1 - t0);
    285 else
    286 check("1000 round trips of 64 bytes", xs, &ok);
    287
    288 if (ok && dnsok)
    289 printf("nettest: ALL OK\n");
    290 else if (ok)
    291 printf("nettest: ALL OK except dns\n");
    292 else
    293 printf("nettest: SOME TESTS FAILED\n");
    294 exit(ok ? 0 : 1);
    295}

6. Verify and measure

On the branch (ext/25-e1000, 7 commits), built with the project toolchain and run on 3 harts (-smp 3 -m 128M), with python3 nettest.py echo ECHOPORT 600 on the host:

$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: OK
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes took 3 ticks
nettest: ALL OK
$ usertests -q
usertests starting
ALL TESTS PASSED
$ nettest ECHOPORT
nettest: one round trip: OK
nettest: 100 round trips: OK
nettest: 3 processes x 300 round trips: OK
nettest: kill wakes recv: OK
nettest: dns: pdos.csail.mit.edu is 128.52.129.126
nettest: dns query to 10.0.2.3: OK
nettest: 1000 round trips of 64 bytes took 1 ticks
nettest: ALL OK

And the host-side ping against nettest server 1000:

$ nettest server 1000
nettest server: echoing 1000 packets on port 2000
nettest server: done
# on the host:
nettest.py: 1000 round trips: OK

Every commit was built on its own, and usertests -q printed ALL TESTS PASSED on 3 harts at each of commits 1 to 6 as well. nettest printed ALL OK on all three boots of the head that ran it (the run above and two gdb traces). A kernel without three of the branch’s fixes (the transmit “full” test, port 0, the port re-check in recv) passed the same sequence, and nettest on all 8 boots of its measurement copies.

The fixes themselves were checked with scratch copies: Clinic 6’s run for the transmit ring, and the port-0 frames described in the reveal step on sockets. Clinics 1 to 5 were run on a kernel without those fixes; none of the fixes touches the code those clinics change.

DNS: slirp’s DNS server at 10.0.2.3 forwards queries to the DNS resolver the host itself uses, and every recorded run got an answer. On a host without a working resolver this check fails and the test says so (ALL OK except dns).

What this demonstrates: packets go both ways through both rings, many times around; three processes on three harts send and receive at once without a packet reaching the wrong socket or a byte changing; kill ends a recv; and the rest of the kernel is unaffected. It does not demonstrate the absence of races: Clinic 3’s lost wakeup passed 9 of 9 runs without help, and Clinic 2’s missing fences cannot fail on QEMU at all.

Round-trip latency, from the host. nettest server 1000 in xv6, nettest.py ping FWDPORT 1000 on the host, which times each round trip with time.perf_counter. The head and a kernel without the three fixes listed in Verify, booted alternately, three times each (measured while the computer was busy with other work; timings on QEMU depend on what else the computer is doing, so yours will differ):

kernel boot min median 99th percentile max
before the fixes 1 166 µs 253 µs 925 µs 96,946 µs
head 1 131 µs 243 µs 345 µs 19,155 µs
before the fixes 2 132 µs 217 µs 815 µs 19,546 µs
head 2 137 µs 219 µs 848 µs 38,970 µs
before the fixes 3 157 µs 229 µs 375 µs 11,400 µs
head 3 122 µs 149 µs 275 µs 40,566 µs

The extra register read of TDH in every transmit makes no difference that this measurement can see; the spread between boots of the same kernel is larger. On a quieter computer, three boots of the kernel without the fixes gave medians of 142, 163 and 202 µs.

A round trip crosses the host’s kernel twice, slirp twice, the e1000 model twice, one xv6 interrupt, a wakeup, a context switch into the server, and two system calls. The rare maxima of 10-100 ms are host scheduling delays (QEMU’s threads share the computer with other work). Inside xv6, nettest’s 1000 round trips of 64 bytes took 1 to 4 ticks of about 0.1 s in the runs without gdb, which agrees within the clock’s resolution.

Where the work and the waiting go. A measurement copy (not on the branch, built without the three fixes, which change nothing these counters see) counts, per hart, e1000 interrupts, packets per interrupt, transmits, full transmit rings, dropped packets, and contended spinlock acquisitions (an acquire whose first amoswap failed); Ctrl-P prints and resets the counters. One run of nettest per boot, six boots:

per run values
packets received / sent 2,003 / 2,003 every run (2,001 echoes, 1 ARP, 1 DNS)
e1000 interrupts 1,803 to 1,838
interrupts per hart between 424 and 869; every hart took a share in every run
packets per interrupt: 0 / 1 / 2 / 3 about 30 / 1,600 / 165 / 20
full transmit ring 0
packets dropped by a full socket queue 0
recv calls that had to sleep 641 to 923 of about 2,000

Contended acquisitions per run (the last three boots also recorded whether the waiter was the interrupt handler):

lock contended of which the handler waited
e1000.lock 176 to 196 26 to 44
socktab.lock 175 to 226 119 to 154
the socket locks 1,450 to 1,561 257 to 303
for scale: kmem.lock 18 to 29

7. Go further