kernel/string.c
About this file
The kernel’s own versions of the C library’s memory and string functions. A
freestanding (vs. hosted) C kernel has no C library, so anything it needs from <string.h> it must
write itself. These eight functions are short loops, but they are used everywhere:
memset clears pages and structures, memmove copies disk blocks and data between
user and kernel memory, and the string functions handle file names and process names.
They follow the standard C functions of the same names closely, with a few differences worth knowing:
- sizes are
uint(32-bit) orint, notsize_t(64-bit on RV64); memcpyismemmove, so it is safe even when the two areas overlap;safestrcpyhas no standard counterpart: it isstrncpythat always terminates the result;strlenreturnsint.
The Makefile passes -fno-builtin-memset and similar flags
(-fno-builtin-NAME) so that the compiler does not substitute its own built-in
versions for these names. User programs have a separate set in user/ulib.c.
Read next: any caller, for example kernel/fs.c (names) or kernel/vm.c
(copyout and copyin use memmove).
Header
Only kernel/types.h, for uint and uchar. These functions use nothing else in
the kernel.
memset(): fill memory with one byte value
Sets n bytes starting at dst to the value c, and returns dst, as the standard
function does. The kernel uses it to clear freshly allocated memory and structures
(for example a new page-table page in kernel/vm.c), and kfree uses it to fill
freed pages with the byte 1, so that code still using a freed page reads garbage
and fails quickly instead of seeming to work.
Line 6 converts the void * to a char * because C does not allow indexing or
dereferencing a void *: it has no element size. Assigning the int c to a
char keeps only its low 8 bits, which matches the standard (“converted to unsigned
char”).
Difference from standard C: n is a uint (at most 4 GiB) instead of size_t.
View the destination as an array of bytes.
Store the low 8 bits of c in the i-th byte.
memcmp(): compare two memory areas
Compares n bytes and returns 0 if they are all equal; otherwise the difference
between the first pair of bytes that differ, which is negative if v1’s byte is
smaller. The bytes are compared as unsigned (uchar), as the C standard requires,
so a byte 0xff counts as larger than 0x01.
Nothing in this version of the kernel calls it (kernel/kernel.asm contains its code
but no call to it). It is still worth having: GCC documents that even freestanding code
must provide memcmp, memcpy, memmove and memset, because the compiler may
generate calls to them on its own.
Treat both areas as arrays of unsigned bytes.
Loop n times; n-- tests the old value, so the body runs exactly n times.
The first difference decides: negative, or positive, by the unsigned byte values.
memmove(): copy memory, even if the areas overlap
Copies n bytes from src to dst and returns dst. It is the kernel’s general
copying function: copyout and copyin use it to move data between kernel and
user memory, the log (kernel/log.c) copies whole disk blocks with it, and
kernel/fs.c copies inode and directory data.
The standard promises that memmove works correctly when the source and destination
overlap, and this one does. It chooses the copying direction for that.
Choosing the direction
A plain front-to-back copy fails in one case: when the destination starts inside
the source, after its start. Example: moving "abcd" one byte to the right. Copying
forward writes a over b before b has been read, and the result is aaaaa.
Line 41 detects exactly that case: s < d (destination after source) and
s + n > d (the source still runs into the destination). Then the loop starts at
the end of both areas and copies backwards (*--d = *--s steps back first, then
copies), so every source byte is read before it is overwritten.
In every other case, including overlap with the destination before the source, a forward copy is safe, and lines 47–48 do that.
The n == 0 test on line 36 returns early; the loops below would also copy nothing,
so it is only a shortcut.
Does the destination start inside the source? Only then is a backward copy needed.
Point one past the end of each area.
Step back one byte in both, then copy: from the last byte down to the first.
The ordinary forward copy.
memcpy(): only so the compiler finds one
Standard memcpy is allowed to fail when the areas overlap; this one forwards to
memmove, so it never does. xv6’s own code always calls memmove, as the
comment says.
The function exists because GCC may emit calls to memcpy by itself (for example
to copy a large structure), and the link would fail if no such function existed. In
this build that never happens: kernel/kernel.asm contains memcpy’s code but no call
to it.
Forward to memmove, which handles any overlap.
strncmp(): compare two strings, at most n characters
Advances through both strings while characters remain to compare, the end of p has
not been reached, and the characters match. If all n characters matched, the
strings are equal for this purpose and it returns 0. Otherwise it returns the
difference of the first differing characters, compared as unsigned bytes. If q
ends first, its 0 differs from p’s character, so the shorter string compares as
smaller, as in the standard.
The kernel uses it in namecmp (kernel/fs.c:586) to compare file names, which
are at most DIRSIZ (14) characters and, in a directory entry, need not end in a
zero byte. Comparing at most 14 characters handles both cases.
Skip the common prefix, but never past n characters or the end of p.
All n characters matched (or n was 0): equal.
Compare the differing characters as unsigned bytes.
strncpy(): copy a string into a fixed-size field
This has the standard (and surprising) strncpy behavior:
- it copies characters from
tuntil it has copied the terminating zero or writtennbytes; - if
twas shorter thann, it fills the rest of thenbytes with zeros; - if
tisncharacters or longer, the result is not zero-terminated.
That suits exactly one job: filling a fixed-width name field. dirlink uses it to
write a name into a directory entry’s 14-byte name (kernel/fs.c:640), where a
14-character name with no terminator is allowed and unused bytes should be zero, so
that the entry on disk does not hold stale bytes.
Difference from standard C: n is an int; a zero or negative n copies nothing.
Copy one character at a time, stopping after the zero has been copied or after n
bytes. The assignment *s++ = *t++ copies one character and advances both pointers.
Fill whatever remains of the n bytes with zeros.
safestrcpy(): copy a string, always terminated
The copy to use when the destination must end up a valid C string. It copies at most
n - 1 characters and always writes a terminating zero, truncating t if it does not
fit. Unlike strncpy it does not fill the rest with zeros. There is no such function
in standard C. The kernel uses it for process names: kexec stores the program’s
name in p->name (kernel/exec.c:130) and kfork copies the parent’s
(kernel/proc.c:290).
The loop condition --n > 0 decrements first, so at most n - 1 characters are copied
and room for the zero always remains. The copy stops early once the zero itself has
been copied (the assignment’s value is the character copied). Line 94 then writes a
zero at s. That is the terminator when the source was truncated; when the source’s
own zero was already copied, it writes one more zero right after it, which is still
inside the n bytes. A non-positive n writes nothing at all (line 90), not even a
terminator, since there is no room.
No room even for a terminator: do nothing.
Copy at most n - 1 characters, stopping after the source’s zero.
Always terminate.
strlen(): the length of a string
Counts characters until the terminating zero. The loop body is empty (;); all the
work is in the for header. It returns an int where standard C returns size_t.
kexec uses it to size each argument string when building the new program’s
stack (kernel/exec.c:102).
Count up to the terminating zero.