pete > courses > CS 315 Fall 26 > lecture 03: file mode, extern, code spacing, first syscalls
Lecture 03: file mode, extern, code spacing, first syscalls
Goals
- use bit masks
- explain the meaning of the extern keyword
- explain tradeoffs in the use of tabs and spaces in source code
- implement copy-file using system calls
- measure its efficiency relative to the original and explain the difference
- implement a wrapper to make the syscall implementation more efficient
- use strace to record a process’ use of syscalls
in the syscall version of the copy-file program, we encoutered open(2) and its weird way of passing flags: it bitwise-or’d the various options and sent the result as its second parameter
you may have wondered at the time what happens on the other end: how does open(2) interpret this parameter such that it knows exactly which flags are set?
I do not present this question purely for curiosity reasons, as you will be faced with a similar question in writing myls
specifically, recall the mode (permission) bits, which look like so:
-rwxr-xr-x
they are also recorded as an integer, where individual bits carry separate meanings (in fact, for a directory with the above permissions, the value of that integer will be decimal 16877)
the value of the least-significant bit indicates whether "other" users have execute permission: if it’s a 1, they do; if it’s a 0, they don’t
the next least-significant bit is write perms for "other", then read perms for "other", then execute perms for "group", etc
therefore, to figure out whether to print an 'x' in that final spot, you need to figure out whether the least-significant bit of the mode field is 1
this is relatively straightforward, but in the spirit of "no magic numbers", I invite you to look at the bottom of the stat(2) manpage, which in turn invites you to read (among others) the inode(7) manpage
the number in parentheses is the section, so…
$ man 7 inode
near the end, it talks about "The file type and mode", in which it tells you that the macro S_ISDIR(m) will tell you whether it’s a directory (where m is the mode field of the stat struct)
just past that, it describes a bunch of values for interpreting the file mode component of the st_mode field
eg:
S_IRUSR 00400 owner has read permission S_IWUSR 00200 owner has write permission S_IXUSR 00100 owner has execute permission
this means that S_IRUSR is defined to have the value 00400 (note that it is octal because of the leading zero)
interestingly, if we write the value 00400 as binary, we get 000100000000
and that 1 is in the exact spot where the "owner has read permission" bit is in the st_mode field
how can we take advantage of this?
bitwise-AND!
recall that a bitwise operation is one that works on individual bits
so bitwise-AND will take two inputs and compare the first bit of each to produce the first bit of output, the second bit of each to produce the second bit of output, and so on
so if we write 16877 as binary and then bitwise-AND it with S_IRUSR…
16877 = 100000111101101 S_IRUSR = 000000100000000 ------------------------- result = 000000100000000
the result being non-zero tells us that the single bit identified by S_IRUSR is indeed set in the other number, which tells us that the "owner has read permission" is set
in C, the bitwise operator is a single ampersand:
has_owner_read_perm = m & S_IRUSR;
values like S_IRUSR are called bit masks because they act like (human) masks: they hide part of the input and only reveal the stuff we want to see
when writing the code to output the file mode characters for myls, you might find it annoyingly wordy
to make it more concise (though you don’t need to), check out the ternary operator in C
there’s a curious thing in the getopt manpage
as you recall, the SYNOPSIS part of the manpage lists the function prototypes and the header file you need to include
but the getopt manpage also includes these two lines:
extern char *optarg; extern int optind, opterr, optopt;
these don’t look like functions, they look like variables
and indeed they are
when a variable declaration is preceded with the extern keyword, it means: "this variable is declared in another file, and you can use it without having to allocate any memory for it"
(the meaning of these particular variables is described in that manpage)
the larger implications and mechanics of this will become clear in the coming weeks, I just wanted to recognize that it’s something new and you can use those variables without declaring them
when you’re editing source code and you hit the Tab key, the cursor magically moves to the right a certain number of columns
the number of columns it moves is usually a configuration setting of your editor (if it’s not, find a new editor. seriously.)
many text editors default to 4 columns in this case, but some people prefer 8 or 2, and I’ve even seen some degens use 3, where no self-respecting computer scientist would use anything but a power of 2
and if we look at our text editor window after hitting Tab, it sure looks and feels like our editor just inserted the right number of space characters to get the cursor to the correct column
(recall that these are text files, often using the ASCII character encoding, which means 1 byte per character)
what may be interesting and surprising to you is that a) there is an ASCII character that represents a horizontal tab and b) that horizontal tab character isn’t necessarily being inserted into your file when you press the Tab key
there are (at least) two cases:
case 1: when you press the Tab key, your editor puts the ASCII character for a horizontal tab in the file you’re editing—this is called using "hard tabs"
case 2: when you press the Tab key, your editor instead inserts the correct number of ASCII space characters—this is called using "soft tabs"
if a file contains hard tabs, the editor also needs to know how many columns a hard tab should represent so they can show well-formatted code to you
and this is where the badness happens, because as I said before, different people have different opinions about "correct" indentation
consider this source file: chaos.c
(the following snippets show what it looks like in my editor: I’m not actually changing the contents of the file, I’m changing the editor settings; and I had to use spaces in the notes to represent all the badness, even hough the file linked above uses hard tabs)
/*
* chaos.c
*/
void chaos(void)
{
printf("this");
printf("is");
printf("a");
printf("disaster");
}
it looks okay, but when I change my editor to use 8 columns for tabs, it looks like this instead:
/*
* chaos.c
*/
void chaos(void)
{
printf("this");
printf("is");
printf("a");
printf("disaster");
}
this is because the first and last lines use 4 spaces for indentation and the middle two lines use hard tabs
if I send this file to someone who has their editor configured for 8-column tabs, all my carefully manicured code will look like a mess to them
and I saw in complete seriousness: the aesthetic appeal of code has direct benefits to its understandability
regularly of form is super important! (and I will require it in your submissions)
there are a few different approaches to solving this problem
the first is to configure your editor to insert soft tabs, as described above
here is a file that I created using soft tabs: spaces-only.c
anybody who opens it in their editor sees the exact same thing because all that indentation is done using spaces and a space is a space is a space
/*
* tabs-only.c
*/
void spaces_only(void)
{
printf("this function also demonstrates multi-line string literals in C, "
"but it also demonstrates why using spaces to indent is"
"splendiforous\n");
for(;;) {
while(1) {
i_am_a_function(parameter_number_one, parameter_number_two,
parameter_number_three, parameter_number_four);
}
}
}
another solution is to use always use hard tabs, as demonstrated in this file: tabs-only.c
in my editor, it looks pretty much the same as the previous file
this is because I have my editor set to use 4 columns for hard tabs
/*
* tabs-only.c
*/
void tabs_only(void)
{
printf("this function also demonstrates multi-line string literals in C, "
"but it also demonstrates why using spaces to indent is"
"splendiforous\n");
for(;;) {
while(1) {
i_am_a_function(parameter_number_one, parameter_number_two,
parameter_number_three, parameter_number_four);
}
}
}
and so when I reconfigure my editor to use, for example, 8 columns for hard tabs, I see this:
/*
* tabs-only.c
*/
void tabs_only(void)
{
printf("in adddition to demonstrating how to make multi-line string "
"literals work in C, this function demonstrates why using tabs to "
"indent is a bad thing, mkay?\n");
for(;;) {
while(1) {
i_am_a_function(parameter_number_one, parameter_number_two,
parameter_number_three, parameter_number_four);
}
}
}
it’s mostly correct, in that the various blocks of code are indented consistently
where it falls apart, though, is where tabs are being used for both indentation and alignment
note the second two lines of the call to printf(3) and the second line of the parameter list to i_am_a_function
when a single line of code becomes too long, we often want to break it into multiple lines for readability
and we may want to start it on a column further right than the current indentation level would otherwise dictate
in the case of this file, pushing the continuation of the printf(3) and i_am_a_function parameters further right makes it very clear that they are, in fact, continuations of the previous and not new, separate lines of code—this is what I’m referring to as alignment
and due to the facts that a) the number of columns to insert depends on the number of characters on the previous line and b) different people may have different preferences regarding hard tabs, we cannot use hard tabs for alignment
and so we are left with the idea of using hard tabs for indentation and spaces for alignment, as demonstrated by this file: tabs-with-spaces-for-alignment.c
in the rendition below, I’ve replaced columns inserted by tabs with the ">" symbol, and have left spaces intact
/*
* tabs-with-spaces-for-alignment.c
*/
void tabs_with_spaces_for_alignment(void)
{
>>>>printf("this function also demonstrates multi-line string literals in C, "
>>>> "but it also demonstrates why using spaces to indent is"
>>>> "splendiforous\n");
>>>>for(;;) {
>>>>>>>>while(1) {
>>>>>>>>>>>>i_am_a_function(parameter_number_one, parameter_number_two,
>>>>>>>>>>>> parameter_number_three, parameter_number_four);
>>>>>>>>}
>>>>}
}
as indicated by the name, we use hard tabs for indentation, but spaces for alignment
I have seen plenty of codebases using soft tabs and plenty using hard tabs, so you are welcome to use either approach
no matter what you choose, you MUST be consistent: you must not mix tabs and spaces for indentation
when you start your own projects, you can do whatever you like, but I really truly honestly believe that, above all, consistency results in better code, so I do strongly recommend that
in the real world, any project you join will have their own conventions and you will have to conform to them
any editor worth its salt will allow you to configure it to use soft tabs
if it doesn’t, find a new editor (again, seriously)
hidden in the discussion of tabs above is a criticism of super-long lines
we are told by Science that humans are best at reading lines of modest width
that is, code is more readable if it doesn’t have ridiculously long lines
concrete requirement: each line of code should be no longer than 80 characters
(there is some wiggle room here, but not much—going over 90 will not fly, not least because I’ve configured my computing situation to work really, really well with lines no more than 90 characters wide—and you might want to do the same for yourself)
and honestly: super-long lines just are not quick and easy to read
it’s worth taking the time to write easy-to-read code; Future You will be grateful
with that stuff out of the way, let’s return to the copy-file program from last time: copy-file.c
recalling our discussion about system calls, we concluded that operations like opening files, reading from files, and writing to files all ought to be system calls because they are both permission-sensitive and prime targets for optimizing
but you may have noticed that the manpages of the functions we used (fopen, fread, fwrite, and fclose) were from section 3, which is not the section of the manpage that covers syscalls
and indeed those functions are not syscalls!
so what are they? and how do we even operate on files without syscalls?
now recall our discussion of printf from last week
the conclusion we came to is that the functionality of printf falls into two categories: formatting output (that is, constructing the string to show) and actually showing that string
the former does not need to involve system calls because it’s "only" computation: it doesn’t require special permissions
but the latter most definitely does require a system call because it accesses a shared hardware resource: the screen (this glosses over details, but the principle holds: it needs to be a syscall)
thus, printf is a wrapper around some input/output (I/O) syscall(s)
meaning that it takes those more primitive operation(s) and builds (wraps) logic around it
the fopen, fclose, fread, and fwrite functions are also wrappers around I/O syscalls
this raises two questions:
- why would one want to wrap the I/O syscalls?
- what do these particular wrappers actually do?
to begin answering this, let’s look at a program that achieves the same effect as copy-file, but is written using system calls: copy-file-syscalls.c
its structure is almost identical, but there are some notable differences worth pointing out
one of the most important similarities (for our purposes today) is that both programs read/write 1024 bytes each loop iteration—this will become important later
instead of fopen, fclose, fread, and fwrite, we have open, close, read, and write
they take some parameters of different types and sometimes have return values of different types than their non-syscall counterparts
to start with, recall that fopen returns a value of type FILE *, and that this pointer is used to refer to the opened file in subsequent functions
in this program, open returns an integer, called a file descriptor, which serves the same purpose
brief history digression!
when Thompson and Ritchie were inventing UNIX (of which Linux is a spiritual descendant) many ages ago, they decided that the central organizing abstraction would be the file
meaning that, from the operating system’s perspective, nearly every resource you’d want to interact with is represented as a file: "normal" files, directories, pieces of hardware, network connections, the screen, etc
therefore, all the functions that perform these operations themselves operate on file descriptors
so the idea of the file descriptor is important, which you will see throughout this course and UNIX system-level programming in general, which is why I’m wasting so much time talking about such a seemingly-minor point
in contrast to the non-syscall version, another significant difference is the parameters to open
to demystify these, your first reaction should be "check the manpage!" and you would be correct
but this is not a normal function; it’s a syscall
therefore, its documentation lies in section 2 of the manual:
$ man 2 open
and we see that the various parameters and their values are explained
in the second call to open (line 37), the second parameter is interesting and possibly unusual:
O_CREAT | O_WRONLY
under the hood, they’re both integers
specifically, O_CREAT has value 64 and O_WRONLY has value 1 (these values are defined in one of the header files you are instructed to include by the manpage)
the | operation is bitwise-OR
here’s how that works in raw binary:
O_CREAT = 64 = 01000000 O_WRONLY = 1 = 00000001 O_CREAT | O_WRONLY = 65 = 01000001
note that the values for both constants consist of all zeroes and a single 1
and that the bitwise-OR result essentially gathers together all the ones that were set in the inputs
this is a common way to specify a set of boolean flags that all apply to the same operation
take a bit vector and assign one bit position to each flag
if a given bit is set to 1, that flag is on, otherwise it is off
then we define a bunch of integers that specify each individual flag and then bitwise-OR them together to specify the set of flags we want
before we compile and run the two programs, check out the Makefile for today
first thing to note is that this one has way more recipes
nothing specifically surprising there, but it’s good to see that this is possible and even a good idea if your situation calls for it
next, look at the first rule:
PHONY: all all: copy-file copy-file-syscalls
there are two new things about it: it has multiple prerequisites and it has no commands to produce those prerequisites!
to build the all target, make will determine that it needs both copy-file and copy-file-syscalls and will thus build both of them (if they need to be built)
no commands are necessary because the Makefile includes separate rules to build both of them
so if I run make all, it should build both
it’s actually even simpler: recall that, if I don’t specify a target on the command-line, make will automatically build the first target
in this case, that’s the all target
so I only need to type…
$ make gcc -Wall -pedantic -o copy-file copy-file.c gcc -Wall -pedantic -o copy-file-syscalls copy-file-syscalls.c
super convenient
now let’s create a file as copy-fodder
the dd program is useful here:
dd if=/dev/urandom of=in-file bs=1024 count=102400
this is a suuuuuper old-school command (like, from the early 70s), which is why its parameters look so weird
the if parameter specifies the input file
the out parameter specifies the output file
the bs parameter specifies the block size (ie, the size of each chunk to be copied)
and the count parameter specifies how many chunks to copy
thus, this command will copy 1024x102400 bytes from /dev/urandom into in-file:
$ dd if=/dev/urandom of=in-file bs=1024 count=102400 102400+0 records in 102400+0 records out 104857600 bytes (105 MB, 100 MiB) copied, 0.725024 s, 145 MB/s $ ls -lh in-file -rw-r--r-- 1 pete pete 100M Sep 19 21:52 in-file
which is, conveniently, 100 megabytes
(there’s also a rule in the Makefile to do this)
so now let’s run the two programs, and let’s time how long they take to run
the time program will do this for us:
$ time ./copy-file in-file out-file real 0m0.092s user 0m0.003s sys 0m0.089s
this output tells us that the program took .092s to run from start to finish
that it spent almost all of that time (.089s) executing system calls
and that it spent only a wee bit of time (.003s) executing "user" code (ie, not system calls)
now the syscall version:
$ time ./copy-file-syscalls in-file out-file real 0m0.152s user 0m0.020s sys 0m0.132s
a bit slower, but not egregiously so
I’m now going to change the number of bytes read/written per loop iteration to 10
(in copy-file.c, change the BUFFER_SIZE constant to value 10; and in copy-file-syscalls.c, change the BYTES_PER_ITERATION constant to 10 as well)
once I’ve saved those changes, I just run make again, it notices that the corresponding source files have changed, and rebuilds the two programs
$ make gcc -Wall -pedantic -o copy-file copy-file.c gcc -Wall -pedantic -o copy-file-syscalls copy-file-syscalls.c
and let’s run our experiments again
$ time ./copy-file in-file out-file real 0m0.412s user 0m0.274s sys 0m0.101s
worse than before, but only by a factor of 4
$ time ./copy-file-syscalls in-file out-file real 0m12.619s user 0m2.075s sys 0m10.510s
MUCH worse than before: by a factor of 100
what’s going on here?
first consider in isolation the behavior of the syscall version
I intend this example to demonstrate that syscalls are expensive operations
indeed, they need to be because they do a lot more than a "normal" function: they actually switch the processor to a different, privileged state in which it can execute the sensitive operating-system code
which does its thing (checks permissions, performs the requested operation)
and then reverts the processor to normal mode before returning
(more on this in Operating Systems; for now, this level of detail is sufficient)
therefore, because these operations are expensive, it behooves a program-writer to perform them as seldom as possible: it is far more efficient to perform a single 1000-byte read than to perform 100 10-byte reads
(note that, in our experiment, we increased the number of syscalls by a factor of 100 and we saw a 100x performance degradation—not a coincidence)
how can we use this information to our advantage?
how can we insulate a program from this massive performance hit on small I/O operations?
we could write a wrapper around, eg, read that performs like so: - the first time our wrapper is called, instead of reading the number of bytes requested, it actually reads a lot more (using the syscall) and saves them; it then fulfills the requested read out of that big bunch - the next time our wrapper is called, the read can be satisfied out of the saved chunk rather than having to execute the read syscall again - and so on and so forth - when the saved chunk doesn’t have enough data in it to satisfy a request, only then will our wrapper call the read syscall again
what about write?
well, instead of writing immediately when we are asked, our wrapper could instead save the data requested to be written in some local memory, wait for it to fill up with successive writes, and only when it’s full will it actually call the write syscall to send it to disk
in effect, our wrapper around read would front-load the corresponding syscall, and our wrapper around write would back-load it
I claim that this is what fread and fwrite do
why should you believe me?
no proof by intimidation this time, I’m going to show you
using a program called strace that makes a record of every system call invoked by a process
this is a very powerful diagnostic tool that you will use a lot in this course (and afterwards, if your life takes you in the direction of this kind of work)
it works like the time program in that it measures whatever program follows it on the command-line
but first, I’m going to make a smaller input file so our record is more tractable:
$ dd if=/dev/urandom of=in-file bs=1k count=10 10+0 records in 10+0 records out 10240 bytes (10 kB, 10 KiB) copied, 0.00016966 s, 60.4 MB/s
and now let’s measure:
$ strace -o strace-no-syscalls.out ./copy-file in-file out-file $ strace -o strace-syscalls.out ./copy-file-syscalls in-file out-file
the -o option tells strace to save its output in the named file (rather than printing to the screen, which is a pain to work with)
the strace output file contains one line per syscall invoked
we can perform a superficial initial test just by counting the lines in the two files:
$ wc -l *.out
47 strace-no-syscalls.out
2085 strace-syscalls.out
2132 total
the wc program ("wordcount") with the -l parameter counts the number of lines in files
*.out means "all files that end in .out"
at first approximation, the non-syscall version is waaaaaay more efficient
(the following examples are cherry-picked from strace-syscalls.out)
since program are (for now) linear, the order they appear in the file is the order in which they were called and returned
openat(AT_FDCWD, "/usr/lib/libc.so.6", O_RDONLY|O_CLOEXEC) = 5
this line tells us that the openat syscall was invoked
that its three parameters were AT_FDCWD, "/usr/lib/libc.so.6", and O_RDONLY|O_CLOEXEC
and that it returned the value 5
read(5, "\177ELF\2\1\1\3\0\0\0\0\0\0\0\0\3\0>\0\1\0\0\0000C\2\0\0\0\0\0"..., 832) = 832
this follows the same pattern, except that the ellipsis in the second parameter indicates that strace is not showing us the full value (there’s a command-line parameter to change this behavior)
don’t worry about what all the inscrutable syscalls near the top do—we’ll get to most of them throughout this course
if we go down about 30 lines in strace-syscalls.out, though, we see records for all the calls to read and write that our loop made
(recall this was the version that reads/writes 10 bytes per loop iteration, and the file is 10240 bytes long, so there should be 1024 reads and 1024 writes down there)
now, open up strace-no-syscalls.out
what do we see?
three calls to read and three calls to write
therefore, even though we called fread and fwrite 1024 times each, they only decided to call the underlying syscall three times
hypothesis proven.
QED.
a couple questions arise, though
if we look a bit closer, we see that both read and write work in chunks of 4096, and only retreat (in this case, to 2048) when the file data is exhausted
why 4096?
because that’s the typical amount in which memory is allocated to a process
it’s a unit called a page, and we’ll talk much more about that—again, in Operating Systems
next question: if fwrite saves up the data being written and only calls write when it has "enough"… what happens when the program exits before it gets "enough"?
the easy answer is that every program calls fclose, which contains logic to flush the buffer (ie, send its data to disk even though it hasn’t been filled)
but what if a program doesn’t call fclose?
it turns out that the stdio library hooks into the program-exit logic to make sure that the buffer-flushing song and dance happens no matter what, so you’re safe
but this is why, if a process writes data and then segfaults, some of the data may not actually appear at its intended destination
when a program segfaults, it misses that buffer-flushing song and dance, and thus the data disappears