C · 6 · Advanced / Job-Ready34 / 35 · 97%
Profiling & Performance
Measure first. The bottleneck is almost never where you guessed.
Examples: perf, gprof, cache misses, -O2
shortcuts: ← prev · → next · M mark
1
Tools
Sampling profilers show where the time actually goes.
Syntax
syntax
perf stat ./app # cycles, IPC, cache misses
perf record -g ./app && perf report
gcc -pg app.c -o app && ./app && gprof ./app2
Cache Locality Wins
Row-major traversal is often several times faster than column-major.
Example
example
// fast: sequential memory
for (int i = 0; i < N; i++)
for (int j = 0; j < N; j++) sum += m[i][j];
// slow: strides the cache on every step
for (int j = 0; j < N; j++)
for (int i = 0; i < N; i++) sum += m[i][j];TIP
Same complexity, same instruction count — the only difference is cache behaviour.
3
Compiler Flags
The optimizer is your cheapest speedup.
Complexity
| -O0 | no opt | debug builds only |
| -O2 | default release | |
| -O3 | more inlining/vectorizing | measure, can be slower |
| -march=native | uses this CPU's ISA | not portable |
| -flto | cross-file inlining |
4
Method
A repeatable loop beats intuition.
Interview question
How do you approach optimizing a slow C program?