user/ulib.c
About this file
xv6 user programs have no standard C library. This file, with user/printf.c,
user/umalloc.c and the system call stubs in user/usys.S, is everything they get
instead: the four object files that the Makefile calls ULIB are linked into
every program (Makefile:104 defines the list, Makefile:107 is the link
command).
It contains:
start, the start-up code: where every program begins, and what callsmain;- the string and memory functions programs need (
strcpy,strcmp,strlen,memset,strchr,memmove,memcmp,memcpy), small versions of the standard ones with a few differences noted below; - two helpers built on system calls,
gets(read a line) andstat(file information by name), andatoi(text to number); sbrkandsbrklazy, two ways of growing a process’s memory through the one system callsys_sbrk.
The kernel has its own copies of the memory functions in kernel/string.c: kernel and
user programs are linked separately and share no code.
Read before: kernel/exec.c (how a program is loaded and started).
Read next: user/user.h, user/usys.S.
Headers shared with the kernel
User programs include some kernel headers directly, so that both sides agree on the
same definitions: kernel/types.h for uint and uchar, kernel/stat.h for
struct stat (filled in by fstat; strictly, this file only passes a pointer to it,
so the declaration in user/user.h would do), kernel/fcntl.h for the open flags such as
O_RDONLY, and kernel/vm.h for SBRK_EAGER and SBRK_LAZY, the two modes
of sys_sbrk. Nothing in this file uses kernel/riscv.h either; including it is
harmless.
user/user.h must come after types.h, because it uses uint without including
anything itself.
start(): where every program begins
When kexec loads a program, it sets the trapframe so that the return
to user mode lands at the program’s entry point with:
a0=argc, the number of arguments: the return value ofexec, whichsyscallstores in the saveda0(kernel/exec.c:140);a1= the user address of theargvarray that kexec copied onto the new user stack (kernel/exec.c:124);sp= the top of that stack, and the program counter = the entry point (kernel/exec.c:136).
The entry point is start. user/user.ld names no entry symbol, and in that case
GNU ld uses the address of a symbol called start (with ${TOOLPREFIX}readelf -h user/_cat the entry is 0xf6, start’s address in user/cat.sym; TOOLPREFIX is your RISC-V toolchain’s prefix; see Tour 1: From make qemu to a disk image and a kernel).
By the RISC-V calling convention a function’s first two arguments arrive in a0
and a1, so the values kexec left there are start’s parameters. Nobody actually
called start, though: there is no valid return address, so start must never
return. It calls main and then hands main’s result to exit, which does not
return. That is the meaning of the comment on line 9: a program may end by returning
from main instead of calling exit itself.
start receives argc in a0 and argv in a1, put there by the kernel, not by a
caller. The name matters: it is how the linker finds the entry point, since
user/user.ld has no ENTRY line. (With no symbol named start either, ld would
use the start of .text, address 0, which is whatever function comes first in the
program’s own file.)
Declares main inside the function. ulib.o is linked into every program, and each
program defines its own main (main, main, …); this
declaration lets the reference compile here and be resolved at link time to whichever
program ulib.o is linked with.
Run the program. When main returns, r holds its return value.
End the process with main’s return value as its exit status. The call goes
through the exit stub to sys_exit and kexit, which never
return: the process becomes a zombie until its parent’s wait collects r.
Because user/user.h declares exit with __attribute__((noreturn)), the compiler
emits no return code after this call: in user/cat.asm, start ends with
jal exit and has no ret. If exit could return, start would “return” to whatever
happened to be in ra.
strcpy(): copy a string
Copies the string t, including its terminating zero byte, to s, and returns the
original s (saved in os, because the loop moves s). As in standard C, nothing
checks that s has room; the caller must make sure. user/ls.c:57 uses it to copy
a directory’s path into a buffer, after checking on line 53 that the path fits.
Copy one byte and advance both pointers. The value of the assignment is the byte just
copied, so the loop ends right after copying the terminating zero. The empty statement
; on line 27 is the loop body: all the work is in the condition.
strcmp(): compare two strings
Returns a negative number, zero or a positive number when p sorts before, equal
to, or after q, comparing byte by byte as unsigned values. The loop advances while
both strings have the same character and p has not ended. When it stops, either
the characters differ or both strings ended together (then both are 0 and the
result is 0).
Advance while p has not ended and the two current characters are equal. If q ends
first, *q is 0, which differs from the non-zero *p, so the loop stops there too.
Subtract the two characters as uchar (unsigned) values, so that bytes above 0x7f
sort after ordinary ASCII instead of being negative on machines where char is signed.
strlen(): length of a string
Counts the bytes before the terminating zero. The result type is uint where
standard C uses size_t; xv6 strings are short, so the int counter is enough.
Step n forward until s[n] is the terminating zero; n is then the length.
memset(): fill memory with one byte value
Stores the low byte of c into each of the n bytes starting at dst, and returns
dst. Programs mostly use it to zero a buffer or a struct (user/sh.c:138 clears
the command buffer before reading a line into it).
The compiler may also generate calls to memset on its own, for example to zero a
large local array, so it must exist even in a program that never names it. GCC
documents that a freestanding (vs. hosted) C program must provide memcpy, memmove, memset
and memcmp; this file defines all four.
The int c is converted to char on assignment, so only its low 8 bits are stored.
strchr(): find a character in a string
Returns a pointer to the first c in s, or 0 if there is none. One difference
from standard C: the standard strchr(s, '\0') returns a pointer to the terminating
zero, but here the loop stops at the zero before comparing it, so the answer is 0.
user/sh.c uses strchr to ask “is this character one of these symbols?”, for
example strchr(whitespace, *s).
Walk the string until its zero byte, returning the first position holding c. The
cast removes const: the result points into the caller’s string, which the caller may
be allowed to modify even though strchr itself promised not to.
gets(): read one line from standard input
Reads characters from file descriptor 0 into buf until a newline or carriage
return, end of input, or until the buffer is full, then terminates the string. The
newline, if one was read, stays in the buffer.
It reads one byte per system call. That is slow but correct: when standard
input is the console, the kernel’s consoleread makes the first
read wait until the user has typed a whole line and pressed Enter
(line editing (cooked input)); the following one-byte reads then return the buffered
characters immediately. Reading exactly one byte at a time also means gets never
consumes input beyond the end of the line, which matters when the descriptor is
shared with another process.
The only caller is the shell’s getcmd (user/sh.c:139), which treats an empty
result as end of input (Ctrl-D at the console).
i + 1 < max keeps one byte free for the terminating zero, so at most max - 1
characters are stored and the buffer cannot overflow. (A caller passing max of 0
would still get a zero written to buf[0]; no caller does.)
Read one byte from file descriptor 0, standard input. The call goes through the
read stub to sys_read and, for the console, consoleread.
read returns 0 at end of input and -1 on error; either way there is nothing more to
read, so stop with what has been collected so far.
A line ends at '\n' or '\r'. The console turns the Enter key’s carriage return into
'\n' (consoleintr), but checking both is harmless.
Terminate the string. i is at most max - 1 here, so this stays inside the buffer.
stat(): information about a file, by name
The kernel offers only fstat, which takes an open file descriptor. stat is
the by-name version, built in user space: open the file, fstat it, close it.
sys_fstat copies a struct stat (kernel/stat.h) into the caller’s
memory: device, inode number, type (file, directory or device), number of links
and size. user/ls.c:65 uses it on every name in a directory.
Because it goes through open, stat needs a free descriptor (each process has
NOFILE = 16) and briefly creates an open file (struct file). It opens read-only
because sys_open refuses to open a directory any other way, and stat
must work on directories.
Open the file to get a descriptor (sys_open). This follows the
path name and fails if it does not exist.
If the file cannot be opened, report failure with -1, as the system calls do.
Ask the kernel to fill *st (sys_fstat → filestat, which copies
the inode's information out with copyout).
Close the descriptor again so stat does not leak it. r keeps fstat’s result,
which is returned.
atoi(): decimal text to a number
Converts the leading decimal digits of s to an int and stops at the first
character that is not a digit. Unlike standard atoi it does not skip leading
spaces, does not accept a - or + sign (atoi("-1") is 0), and does not detect
overflow. main uses it to turn its arguments into process IDs.
For each digit, shift the number so far one decimal place left and add the digit.
*s++ - '0' turns a character such as '7' into the number 7, because the digit
characters are consecutive in ASCII.
memmove(): copy bytes, even if the areas overlap
Copies n bytes from vsrc to vdst and returns vdst. The areas may overlap, and
the copy direction is what makes that safe:
- If the source is above the destination, copying from the front is safe: each source byte is read before the copy could overwrite it.
- Otherwise (source below or equal to destination) the function copies from the back, starting with the last byte, for the same reason.
For example, moving "abcd" one byte to the right with a front-to-back copy would
produce "aaaaa"; back-to-front produces "aabcd" as intended. user/grep.c uses
it to shift the unconsumed part of its buffer to the start.
n is an int here (the declaration in user/user.h agrees), while memcpy
passes it a uint; a count above 2^31 − 1 would become negative and nothing would be
copied. The kernel’s version is memmove.
Choose the copy direction by comparing addresses. (Comparing pointers into different objects is not defined by the C standard, but on RISC-V it compares the addresses, as intended.)
Forward copy, front to back. n-- > 0 tests the old value of n, so the body runs
exactly n times, and a negative n copies nothing.
Backward copy: move both pointers one past the end, then pre-decrement before each byte, so the last byte is copied first and the first byte last.
memcmp(): compare two memory areas
Compares n bytes and returns the difference of the first pair that differs, or 0.
The bytes are read as plain char. On RISC-V, plain char is unsigned (the RISC-V
ABI says so; (char)-1 < 0 compiles to 0 with GCC for riscv64), so the result has
the same sign as an unsigned comparison, which is what the C standard requires. On a
machine where char is signed this code would get bytes above 0x7f backwards.
No xv6 user program calls it; it exists because the compiler may emit calls to it
(see the note on memset above).
Return the difference of the first unequal bytes. Its sign says which area is larger.
memcpy(): copy non-overlapping memory
Standard C says memcpy’s areas must not overlap, which lets a library use a faster
copy. xv6 does not bother and forwards to memmove, which is correct for any
areas. No xv6 user program calls memcpy by name (it appears in no user/*.asm call
instruction); it is here for the compiler, which may generate a call to memcpy to
copy a large struct.
Forward to memmove, which also handles overlapping areas.
sbrk() and sbrklazy(): grow or shrink the process's memory
The kernel has a single system call for changing the size of a process’s memory,
sys_sbrk, with two arguments: how many bytes to add (negative to remove),
and how. These two wrappers fix the second argument. The stub they call is named
sys_sbrk in user space too (user/usys.pl:12 gives it that name), so that the
name sbrk is free for the friendlier wrapper.
Both return the old end of memory, which is the start of the newly added bytes, or
SBRK_ERROR ((char *)-1) on failure. The kernel stores -1 in a0 and the
char * return type turns it into that pointer value.
sbrk(n)passesSBRK_EAGER: the kernel allocates and maps zeroed pages right away throughgrowproc.mallocuses it (user/umalloc.c:54).sbrklazy(n)passesSBRK_LAZY: the kernel only raisesp->sz. A page is allocated the first time the program touches it, when the resulting page fault reachesvmfault. Reserving a large area this way is fast and costs no memory for parts that are never used.user/usertests.ctests it.
For a negative n the kernel always shrinks immediately, whichever mode was asked
for.
Call the sys_sbrk system call stub with SBRK_EAGER (1) as the second argument,
in a1: allocate the memory now. user/cat.asm shows this as li a1,1 followed by a
call to sys_sbrk; n is already in a0.
Same call with SBRK_LAZY (2): only reserve the address range, and let page faults
allocate the pages when they are first used (kernel/sysproc.c:54).