pete > courses > CS 315 Fall 26 > lecture 04: basic linking, make variables, gitignore
Lecture 04: basic linking, make variables, gitignore
Goals
- use function prototypes
- compile multiple source files into a single program
- use gcc to separately compile and link
- streamline your Makefiles with user-defined variables and automatic variables
- use .gitignore to keep your git repos clean
consider that your implementations of myread, mywrite, myopen, and myclose are quite useful
so useful, in fact, that you might want to use them in multiple different programs
you could cut and paste the function implementations into each new project
but it would be more convenient to put them in their own file and just copy the whole file between projects
so you’d have one file containing the "main" project code, which calls functions that are defined in this other file
which begs the question: how do we compile together multiple source files into a single program?
it turns out to be simple: you just feed both source files to gcc and magic happens
here are some source files: main.c, file1.c, file2.c
note that code in main.c calls functions defined in the other two files
here is a Makefile that will compile all three of those files into a single program:
main: main.c file1.c file2.c
gcc -Wall -pedantic -o main main.c file1.c file2.c
.PHONY: clean
clean:
rm -f main
running make gives a warning, though:
$ make
gcc -Wall -pedantic -o main main.c file1.c file2.c
main.c: In function ‘main’:
main.c:8:5: warning: implicit declaration of function ‘one’ [-Wimplicit-function-declaration]
one();
^~~
main.c:9:5: warning: implicit declaration of function ‘two’ [-Wimplicit-function-declaration]
two();
^~~
why?
check out the code in main.c
note that, when the compiler reaches the line that calls one(); it has never heard that a function called one even exists, let alone what kind of arguments it takes or what type of value it returns
now, we the programmers know that those functions exist, we know their arguments, we know their return value
we can share this information with the compiler by adding a pair of function prototypes at the top of main.c (just after the header comment):
void one(void); void two(void);
these lines tell the compiler about the functions, and therefore it won’t be worried when they’re called a few lines later
they are explicit declarations (in contrast to before, when they were, as the compiler complained, only implicitly defined)
hearkening back to Tuesday’s lecture, these are essentially extern declarations for functions: introducing the name so the compiler doesn’t complain, but not actually defining them because they’re defined elsewhere and will be compiled in at the appropriate time
now it builds and runs without complaint:
$ make gcc -Wall -pedantic -o main main.c file1.c file2.c $ ./main A long time ago in a galaxy far, far away...
thinking back to what we know from Architecture, what’s actually happening here?
the compiler is taking the source code in the three .c files, compiling it all to assembly, compiling all of that into machine code, and cramming all that machine code into a file named main, which we can then run
cool
look back at the Makefile above; why might compiling this way be inconvenient?
what if only file1.c changes?
to rebuild, the compiler will take the source code in the three .c files, compile it all to assembly, compile all of that into machine code, and cram all that machine code into a file named main, which we can then run
this is wasteful, because the code in main.c and file2.c didn’t change
here’s a better solution:
- compile each source file to machine code and save the results in separate files
- combine those machine-code files to produce the program
- then, when a single source file changes, recompile only that source file to produce a new machine-code file
- combine this new machine-code file with the others that didn’t need to be recompiled to produce a new program
the "combine machine-code files" step is called linking and we will see it in great detail in a couple weeks
for now, it suffices to know how to make it happen: gcc will actually do it for us with no fuss
but first, we need to know how to ask gcc to only compile to machine code and not try to produce a program
this is accomplished with the -c flag to gcc
by convention, the machine code is saved in a file with the .o extension, reflecting another name for machine code: object code
eg:
gcc -Wall -pedantic -c -o main.o main.c
once we’ve got the three .o files, we can ask gcc to link them together to form the program
we do so by just feeding it the .o files as input, instead of the .c files from before (gcc will notice the different contents and do the right thing)
gcc -o main main.o file1.o file2.o
here’s the full Makefile:
main: main.o file1.o file2.o
gcc -o main main.o file1.o file2.o
main.o: main.c
gcc -Wall -pedantic -c -o main.o main.c
file1.o: file1.c
gcc -Wall -pedantic -c -o file1.o file1.c
file2.o: file2.c
gcc -Wall -pedantic -c -o file2.o file2.c
.PHONY: clean
clean:
rm -f main main.o file1.o file2.o
let’s test it:
$ make gcc -Wall -pedantic -c -o main.o main.c gcc -Wall -pedantic -c -o file1.o file1.c gcc -Wall -pedantic -c -o file2.o file2.c gcc -o main-in-parts main.o file1.o file2.o
and the contents of the directory corroborate this:
$ ls file1.c file1.o file2.c file2.o main.c main main.o Makefile
now, I’m going to use the touch command to change the modification-time of file1.c, which will cause make to think that file1.o is now out of date and needs to be re-compiled
$ touch file1.c $ make gcc -Wall -pedantic -c -o file1.o file1.c gcc -o main main.o file1.o file2.o
we can see that make only felt the need to recompile the one file but then linked its result with the two .o files left over from before
this may seem like a pretty pathetic win, but if we’re working on a program that comprises hundreds of source files (which is not at all unreasonable to imagine) it’s potentially huge
our Makefile from last time still has some annoying aspects
notably, it has a ton of repetition, especially of things we probably want to keep the same throughout:
most glaringly, we probably always want all three .c files to be compiled with the same flags
and thus repeating them three times is morally reprehensible
fortunately, we can define and use variables in our Makefile
by convention, flags (options) passed to the C compiler are defined in a variable called CFLAGS, like so:
CFLAGS=-Wall -pedantic
main: main.o file1.o file2.o
gcc -o main main.o file1.o file2.o
main.o: main.c
gcc $(CFLAGS) -c -o main.o main.c
file1.o: file1.c
gcc $(CFLAGS) -c -o file1.o file1.c
file2.o: file2.c
gcc $(CFLAGS) -c -o file2.o file2.c
.PHONY: clean
clean:
rm -f main main.o file1.o file2.o
note that when defining the variable, you don’t use the $ or ( ), but when you use it, you do
you may notice that I’ve specifically left two parts of the gcc command out of the CFLAGS variable: -c and -o
the reason I left out -c is because it’s possible that sometimes you want to use all the standard compiler flags but actually produce a program immediately
putting -c in the CFLAGS would make that difficult
secondly, I left out -o because it’s a two-part option, and putting the -o in CFLAGS would mean that, in the recipe, I always need to follow $(CFLAGS) immediately by the name of the output file
this requirement is not at all clear, and is thus easily forgotten, which makes it more likely to result in a failure requiring debugging, for which most people have better things to do
lesson: implicit requirements are bad and often waste time in the long run
but the Makefile is still a bit repetitive
specifically, the rules for building main.o, file1.o, and file2.o are all very similar: the only things that differ are the names of the input and output files
it would be nice if we could tell make something like "to build any .o file, take the corresponding .c file and…"
good news: we can
here’s the rule:
%.o: %.c
gcc $(CFLAGS) -o $@ $^
the target here is %.o, which matches any requested target that ends in .o
thus, if we run make foo.o, this rule will match
additionally, make will remember the part of the filename that came before the .o part—this is called the stem
the prerequisite is %.c, where % is replaced with the stem, so when we run make foo.o, make infers that the prereq is actually foo.c
now, since the target and prereqs are both variable, we need some way to refer to them in the commands below
make provides the automatic variables $@ and @^ for precisely this purpose
$@ is automatically expanded to the target
$^ is automatically expanded to the list of prereqs
(the full list of automatic variables is available here
we can do the same thing with the primary rule, and now our Makefile looks like this:
CFLAGS=-Wall -pedantic
.PHONY: all
all: main
main: main.o file1.o file2.o
gcc -o $@ $^
%.o: %.c
gcc $(CFLAGS) -c -o $@ $^
.PHONY: clean
clean:
rm -f main main.o file1.o file2.o
one final thing to note: for the "main" rule, we do not use CFLAGS
for the very specific reason that this rule is not performing compilation: it is only linking together files that have already been compiled
(it is reasonable to provide linker flags at this stage, but we don’t have a need to do so…. yet)
minor-ish note on git
as you develop, you will produce lots of files that probably don’t make sense to store in git
currently, the most obvious example is compiled programs
there isn’t much need to commit them to your repo, as anyone with the source code can reproduce them
another is all the .o files we saw this morning: these files can be easily reproduced by the compiler, so it is wasteful to put them in your git repo
it’s also annoying to have to manually filter the results of git status to consciously ignore these files every time you perform a commit
enter: .gitignore
this is a text file in the top-most directory of your repo that lists files that git should ignore
thus, if you put the name of your program in this file, git won’t try to commit it, it won’t list it in the output of git status, and it won’t complain when you don’t git add it
it is a good idea to commit the .gitignore file itself to your repo, however, so that people using your repo get the benefit of your ignores
from the copy-file example from this morning, here is what a .gitignore file would look like:
$ cat .gitignore copy-file copy-file-syscalls *.o
this file must be in the top-level directory of your repository
and yes: you should commit it into your repo, so anybody working with your repo gets its benefit
(the cat program just dumps the contents of a file to the terminal)
the *.o syntax says "every file ending in .o"