user/ls.c
About this file
ls lists files. ls alone lists the current directory; ls d lists directory d;
ls f for an ordinary file prints one line about that file. Each line has four columns:
the name (padded to 14 characters), the type (1 directory, 2 file, 3 device), the
inode number and the size in bytes. In a freshly booted xv6, ls prints:
. 1 1 1024
.. 1 1 1024
README 2 2 2441
cat 2 3 38128
...
console 3 23 0
(Sizes depend on the build.) Notice that . and .. are both inode 1 in the root
directory, and that console is a device with size 0.
The program shows two things you cannot see from cat. First, how to ask the kernel
about a file without reading it: fstat fills in a struct stat (kernel/stat.h).
Second, that in xv6 a directory is an ordinary file whose content is an array of
16-byte struct dirent records (kernel/fs.h); ls reads them with plain read.
Read before: user/cat.c, kernel/stat.h, the dirent definition in
kernel/fs.h. Read next: kernel/fs.c (dirlookup reads the same
entries inside the kernel).
Headers
kernel/stat.h defines struct stat and the type codes T_DIR, T_FILE,
T_DEVICE; kernel/fs.h defines struct dirent and DIRSIZ (14, the longest
file name); kernel/fcntl.h defines O_RDONLY. ls thus includes kernel headers
directly, so user programs and the kernel agree on these layouts by sharing the
definitions.
fmtname(): the last path component, padded to 14 characters
fmtname turns a path such as ./README into the column-aligned string
README (14 characters), so that the numbers after it line up.
Lines 14–16 find the last component. The loop starts at the string’s terminating NUL
and walks backwards until it reaches a / or falls off the front; p++ then steps
forward onto the first character of the name. For a path without any /, the loop
ends with p one position before path, and p++ brings it back. (Strictly, C
does not allow computing a pointer before the start of an array; it works with
xv6’s compiler, but it is undefined behavior.)
Lines 19–24 do the padding in a static buffer: copy the name, fill with spaces up
to DIRSIZ, terminate. A name of 14 characters or more is returned as it is, with
no padding.
Because buf is static, every call returns the same buffer and overwrites the
previous result. That is safe here only because each result is printed before the
next call.
DIRSIZ + 1 = 15 bytes: 14 for the padded name and one for the NUL. static makes
it survive after fmtname returns, so returning a pointer to it is valid.
Start at the NUL at the end of path and step backwards until p is on a / or has
passed the beginning.
An empty loop body: all the work is in the for header.
Step forward from the / (or from just before the start) to the first character of
the last component.
A name of exactly DIRSIZ characters needs no padding, and a longer one would not fit
in buf, so both are returned as they are. A longer name can only come from the
command line, and it opens at all only because the kernel’s skipelem cuts
each path element to DIRSIZ characters during lookup: ls abcdefghijklmnopqrstu
opens abcdefghijklmn.
Return a pointer into the caller’s path, not into buf.
Copy the name’s characters, without its NUL.
Fill the rest of the 14 characters with spaces.
Terminate the string at buf[14].
ls(): open the path and find out what it is
ls first opens the path and asks for its metadata with fstat. In the kernel,
sys_fstat calls filestat, which locks the inode, copies
its device, inode number, type, link count (nlink) and size into a struct stat with
stati, and copies that out to the user buffer st.
Opening a directory works because ls asks for O_RDONLY; sys_open
refuses to open a directory for writing (kernel/sysfile.c:355).
Errors go to standard error and ls moves on to the next
argument: unlike user/cat.c, a bad name does not stop the run. The error path
after a failed fstat closes the descriptor, so it does not leak.
Open the path read-only (sys_open); this works for files, directories and
devices alike.
Report the failure on standard error and skip this argument.
Fill st with the inode’s metadata through the open descriptor.
A file or a device prints one line
For a T_FILE or T_DEVICE, ls prints the single line for the path itself. The
size of a device is 0: a device inode has no data blocks; its read and write go
to a driver chosen by its major number (devsw).
st.size is a uint64, but %d expects an int. The (int) cast makes the
argument match the format; without it, the format(printf, ...) attribute on
printf in user/user.h makes the compiler warn, and xv6 builds with -Werror.
Sizes over 2 GiB would print wrongly, but no xv6 file can be that big.
Choose by the inode’s type.
Name, type, inode number, size. fmtname(path) strips any leading directories:
ls /README prints README.
A directory needs a path buffer
To show each entry’s type and size, ls must stat each entry, which needs its full
path: the directory path, a /, the entry’s name. buf (512 bytes, on the
stack) holds that path. Lines 57–59 copy the directory path into it and append
the /; p then points where each entry’s name will be written.
Line 53 is the guard against overflowing buf: the directory path, the /, up to
DIRSIZ name characters and the terminating NUL must fit in 512 bytes. Without it, a
long path would make strcpy write past the end of buf and over whatever
lies beyond it on the stack, such as saved registers and the return address.
In practice the check cannot fire in xv6: the kernel’s argstr rejects any
path of MAXPATH (128) characters or more, so open on line 35 would
already have failed for a path long enough to trip it. The guard is still the right
habit: ls should not depend on a kernel limit to keep its own buffer safe.
Would the longest possible entry path overflow buf? See the block note.
This message goes to standard output.
Start the entry path with the directory path.
p points at the NUL at the end of the copied path.
Append / and leave p just after it, where each entry’s name goes.
Read the directory one entry at a time
Each read of sizeof(de) (16) bytes returns the next dirent: a 2-byte inode
number and a 14-byte name. In the kernel, fileread calls
readi on the directory’s inode at the current offset, the same code that
reads ordinary files. The loop ends when read returns 0 at the end of the
directory (or anything other than 16).
An entry with inode number 0 is a free slot, left when a file is removed
(sys_unlink overwrites the entry with zeros) or never used, so it is
skipped. The name is copied into the path buffer and NUL-terminated explicitly,
because a name of exactly 14 characters fills the field and has no NUL of its own.
stat (in user/ulib.c) opens the full path, calls fstat and closes it
again, so listing a directory costs three system calls per entry. It can fail if the
entry was removed after read returned it, or if the process or system is out of
descriptors or open-file slots. Such failures are reported on standard output
(line 66) and the listing continues.
Reusing st for each entry is fine: the directory’s own struct stat is no longer
needed once the switch has chosen this case.
Read the next 16-byte directory entry (sys_read); stop at the end.
Inode number 0 marks an empty slot.
Skip it.
Copy all 14 name bytes after the /.
Terminate the name, which may fill all 14 bytes.
Look up the entry’s metadata by path (stat: open + fstat + close).
Report the failure (on standard output) and go on with the next entry.
Print the entry. fmtname(buf) strips the directory part again, leaving the name.
Done with this path
All three cases end here, and the descriptor opened on line 35 is closed so that a long argument list does not run out of descriptors.
Close the file or directory opened on line 35.
main(): list . or each argument
With no arguments, ls lists ., the current directory (each process has one,
p->cwd, inherited from its parent and changed by cd via sys_chdir).
Otherwise each argument is listed in turn. Errors are reported per argument, and
the exit status is always 0.
No arguments.
List the current directory.
List each argument in turn.
One call per argument; errors in one do not stop the others.