pete > courses > CS 315 Fall 26 > lecture 04: basic linking, make variables, gitignore


Lecture 04: basic linking, make variables, gitignore

Goals


consider that your implementations of myread, mywrite, myopen, and myclose are quite useful

so useful, in fact, that you might want to use them in multiple different programs

you could cut and paste the function implementations into each new project

but it would be more convenient to put them in their own file and just copy the whole file between projects

so you’d have one file containing the "main" project code, which calls functions that are defined in this other file


which begs the question: how do we compile together multiple source files into a single program?

it turns out to be simple: you just feed both source files to gcc and magic happens

here are some source files: main.c, file1.c, file2.c

note that code in main.c calls functions defined in the other two files

here is a Makefile that will compile all three of those files into a single program:

main: main.c file1.c file2.c
    gcc -Wall -pedantic -o main main.c file1.c file2.c

.PHONY: clean
clean:
    rm -f main

running make gives a warning, though:

$ make
gcc -Wall -pedantic -o main main.c file1.c file2.c
main.c: In function ‘main’:
main.c:8:5: warning: implicit declaration of function ‘one’ [-Wimplicit-function-declaration]
     one();
     ^~~
main.c:9:5: warning: implicit declaration of function ‘two’ [-Wimplicit-function-declaration]
     two();
     ^~~

why?

check out the code in main.c

note that, when the compiler reaches the line that calls one(); it has never heard that a function called one even exists, let alone what kind of arguments it takes or what type of value it returns

now, we the programmers know that those functions exist, we know their arguments, we know their return value

we can share this information with the compiler by adding a pair of function prototypes at the top of main.c (just after the header comment):

void one(void);
void two(void);

these lines tell the compiler about the functions, and therefore it won’t be worried when they’re called a few lines later

they are explicit declarations (in contrast to before, when they were, as the compiler complained, only implicitly defined)

hearkening back to Tuesday’s lecture, these are essentially extern declarations for functions: introducing the name so the compiler doesn’t complain, but not actually defining them because they’re defined elsewhere and will be compiled in at the appropriate time


now it builds and runs without complaint:

$ make
gcc -Wall -pedantic -o main main.c file1.c file2.c
$ ./main
A long time ago
in a galaxy far, far away...

thinking back to what we know from Architecture, what’s actually happening here?

the compiler is taking the source code in the three .c files, compiling it all to assembly, compiling all of that into machine code, and cramming all that machine code into a file named main, which we can then run

cool


look back at the Makefile above; why might compiling this way be inconvenient?

what if only file1.c changes?

to rebuild, the compiler will take the source code in the three .c files, compile it all to assembly, compile all of that into machine code, and cram all that machine code into a file named main, which we can then run

this is wasteful, because the code in main.c and file2.c didn’t change


here’s a better solution:

the "combine machine-code files" step is called linking and we will see it in great detail in a couple weeks

for now, it suffices to know how to make it happen: gcc will actually do it for us with no fuss


but first, we need to know how to ask gcc to only compile to machine code and not try to produce a program

this is accomplished with the -c flag to gcc

by convention, the machine code is saved in a file with the .o extension, reflecting another name for machine code: object code

eg:

gcc -Wall -pedantic -c -o main.o main.c

once we’ve got the three .o files, we can ask gcc to link them together to form the program

we do so by just feeding it the .o files as input, instead of the .c files from before (gcc will notice the different contents and do the right thing)

gcc -o main main.o file1.o file2.o

here’s the full Makefile:

main: main.o file1.o file2.o
    gcc -o main main.o file1.o file2.o

main.o: main.c
    gcc -Wall -pedantic -c -o main.o main.c

file1.o: file1.c
    gcc -Wall -pedantic -c -o file1.o file1.c

file2.o: file2.c
    gcc -Wall -pedantic -c -o file2.o file2.c

.PHONY: clean
clean:
    rm -f main main.o file1.o file2.o

let’s test it:

$ make
gcc -Wall -pedantic -c -o main.o main.c
gcc -Wall -pedantic -c -o file1.o file1.c
gcc -Wall -pedantic -c -o file2.o file2.c
gcc -o main-in-parts main.o file1.o file2.o

and the contents of the directory corroborate this:

$ ls
file1.c  file1.o  file2.c  file2.o  main.c  main  main.o  Makefile

now, I’m going to use the touch command to change the modification-time of file1.c, which will cause make to think that file1.o is now out of date and needs to be re-compiled

$ touch file1.c
$ make
gcc -Wall -pedantic -c -o file1.o file1.c
gcc -o main main.o file1.o file2.o

we can see that make only felt the need to recompile the one file but then linked its result with the two .o files left over from before

this may seem like a pretty pathetic win, but if we’re working on a program that comprises hundreds of source files (which is not at all unreasonable to imagine) it’s potentially huge


our Makefile from last time still has some annoying aspects

notably, it has a ton of repetition, especially of things we probably want to keep the same throughout:

most glaringly, we probably always want all three .c files to be compiled with the same flags

and thus repeating them three times is morally reprehensible

fortunately, we can define and use variables in our Makefile

by convention, flags (options) passed to the C compiler are defined in a variable called CFLAGS, like so:

CFLAGS=-Wall -pedantic

main: main.o file1.o file2.o
    gcc -o main main.o file1.o file2.o

main.o: main.c
    gcc $(CFLAGS) -c -o main.o main.c

file1.o: file1.c
    gcc $(CFLAGS) -c -o file1.o file1.c

file2.o: file2.c
    gcc $(CFLAGS) -c -o file2.o file2.c

.PHONY: clean
clean:
    rm -f main main.o file1.o file2.o

note that when defining the variable, you don’t use the $ or ( ), but when you use it, you do


you may notice that I’ve specifically left two parts of the gcc command out of the CFLAGS variable: -c and -o

the reason I left out -c is because it’s possible that sometimes you want to use all the standard compiler flags but actually produce a program immediately

putting -c in the CFLAGS would make that difficult

secondly, I left out -o because it’s a two-part option, and putting the -o in CFLAGS would mean that, in the recipe, I always need to follow $(CFLAGS) immediately by the name of the output file

this requirement is not at all clear, and is thus easily forgotten, which makes it more likely to result in a failure requiring debugging, for which most people have better things to do

lesson: implicit requirements are bad and often waste time in the long run


but the Makefile is still a bit repetitive

specifically, the rules for building main.o, file1.o, and file2.o are all very similar: the only things that differ are the names of the input and output files

it would be nice if we could tell make something like "to build any .o file, take the corresponding .c file and…"

good news: we can

here’s the rule:

%.o: %.c
    gcc $(CFLAGS) -o $@ $^

the target here is %.o, which matches any requested target that ends in .o

thus, if we run make foo.o, this rule will match

additionally, make will remember the part of the filename that came before the .o part—this is called the stem

the prerequisite is %.c, where % is replaced with the stem, so when we run make foo.o, make infers that the prereq is actually foo.c

now, since the target and prereqs are both variable, we need some way to refer to them in the commands below

make provides the automatic variables $@ and @^ for precisely this purpose

$@ is automatically expanded to the target

$^ is automatically expanded to the list of prereqs

(the full list of automatic variables is available here


we can do the same thing with the primary rule, and now our Makefile looks like this:

CFLAGS=-Wall -pedantic

.PHONY: all
all: main

main: main.o file1.o file2.o
    gcc -o $@ $^

%.o: %.c
    gcc $(CFLAGS) -c -o $@ $^

.PHONY: clean
clean:
    rm -f main main.o file1.o file2.o

one final thing to note: for the "main" rule, we do not use CFLAGS

for the very specific reason that this rule is not performing compilation: it is only linking together files that have already been compiled

(it is reasonable to provide linker flags at this stage, but we don’t have a need to do so…. yet)


minor-ish note on git

as you develop, you will produce lots of files that probably don’t make sense to store in git

currently, the most obvious example is compiled programs

there isn’t much need to commit them to your repo, as anyone with the source code can reproduce them

another is all the .o files we saw this morning: these files can be easily reproduced by the compiler, so it is wasteful to put them in your git repo

it’s also annoying to have to manually filter the results of git status to consciously ignore these files every time you perform a commit


enter: .gitignore

this is a text file in the top-most directory of your repo that lists files that git should ignore

thus, if you put the name of your program in this file, git won’t try to commit it, it won’t list it in the output of git status, and it won’t complain when you don’t git add it

it is a good idea to commit the .gitignore file itself to your repo, however, so that people using your repo get the benefit of your ignores


from the copy-file example from this morning, here is what a .gitignore file would look like:

$ cat .gitignore
copy-file
copy-file-syscalls
*.o

this file must be in the top-level directory of your repository

and yes: you should commit it into your repo, so anybody working with your repo gets its benefit

(the cat program just dumps the contents of a file to the terminal)

the *.o syntax says "every file ending in .o"

Last modified: