Lecture 9 - Complexity: how to compare two programs
Announcements
Learning objectives
By the end of today, you should be able to:
- Time two programs that do the same thing, and explain why debug and release give different numbers
- Count the steps in each part of a program, and say how that count grows with the input
- Use Big O notation to describe time and space complexity
- Recognize common complexity classes: O(1), O(log n), O(n), O(n^2), O(2^n)
- Apply key rules: drop constants, keep dominant terms
Part 1: Which one is faster?
Monday's question
Two ways to find the nth prime (the 4th prime is 7):
A. Check each number. Count up from 2. For each number, try dividing it by every smaller number. Stop once you've found n primes.
B. Cross out multiples. Write down every number up to 100,000. Cross out every multiple of 2, then every multiple of 3, then of the next number that isn't crossed out, and so on. Then count through what's left.
Which would you bet is faster? Does it depend on n?
The two programs in Rust
/// Count up from 2, checking each number, until we've found n primes.
fn nth_prime_a(n: u32) -> u32 {
let mut count = 0;
let mut candidate = 1;
while count < n {
candidate += 1;
let mut is_prime = true;
for d in 2..candidate {
if candidate % d == 0 {
is_prime = false;
break;
}
}
if is_prime {
count += 1;
}
}
candidate
}
/// Cross out every multiple of every number up to a limit,
/// then count through whatever is left.
fn nth_prime_b(n: u32) -> u32 {
let mut maybe_prime = [true; 100_000];
maybe_prime[0] = false;
maybe_prime[1] = false;
for i in 2..maybe_prime.len() {
if maybe_prime[i] {
let mut multiple = i * 2;
while multiple < maybe_prime.len() {
maybe_prime[multiple] = false;
multiple += i;
}
}
}
let mut count = 0;
for i in 0..maybe_prime.len() {
if maybe_prime[i] {
count += 1;
if count == n {
return i as u32;
}
}
}
0 // n was too big for this limit
}
Let's time them
cargo run
On my machine, in debug mode:
| n | nth prime | A | B |
|---|---|---|---|
| 4 | 7 | 709 ns | 2.4 ms |
| 100 | 541 | 0.2 ms | 2.8 ms |
| 1,000 | 7,919 | 25 ms | 2.3 ms |
| 2,000 | 17,389 | 87 ms | 1.4 ms |
| 4,000 | 37,813 | 343 ms | 1.7 ms |
| 8,000 | 81,799 | 1.5 s | 1.7 ms |
Now in release mode
cargo run --release
| n | A | B |
|---|---|---|
| 4 | 42 ns | 0.33 ms |
| 100 | 0.03 ms | 0.34 ms |
| 1,000 | 4.2 ms | 0.31 ms |
| 2,000 | 16 ms | 0.26 ms |
| 4,000 | 50 ms | 0.21 ms |
| 8,000 | 225 ms | 0.27 ms |
Time things in release mode. Debug is for building and fixing: it compiles fast and runs slow. Release takes longer to compile and runs about 6-7x faster here.
What just happened?
- A wins for small
n. B wins, by a lot, for bign - B takes about the same time no matter which prime you ask for
- Each time
ndoubles, A gets about 4x slower, not 2x
Our intuition says a job twice as big should take twice as long. It's often not that simple. It depends on the algorithm.
T/P/S - Why? Look at each loop in the code. What decides how many times it runs?
It's about more than getting the right answer
- A and B are both correct. Which one is better depends on how it gets used
- Ask an AI for "the nth prime" and you might get A
- Better asks:
- "Show me a few different approaches"
- "How does this scale as
ngrows?" - "I'll need lots of primes. Can we find them once and look them up after?"
- You can only ask those if you know they're the right questions
Counting the steps in each part
A has a loop inside a loop, and both grow with n:
- The
whileruns once per number checked, all the way up to thenth prime (81,799 forn= 8,000) - For each prime, the
fortries every smaller number
Double n and you check about twice as many numbers, each with about twice as many divisors to try: about 4x the work.
B never looks at n until the very end:
- Crossing out always covers all 100,000 numbers
- Counting through what's left is at most 100,000 more steps
So B does about the same work every time.
That's why it loses for the 4th prime and wins for the 8,000th.
Part 2: Big O notation - The math of "about how fast?"
Think-pair-share: Counting operations
Part 1: Given this code:
#![allow(unused)] fn main() { fn sum_to(n: u64) -> u64 { let mut total = 0; for i in 1..=n { total += i; } total } }
Question: How many addition operations happen?
Part 2: Now consider this code:
#![allow(unused)] fn main() { fn count_pairs(n: usize) -> usize { let mut count = 0; for i in 1..n { for _j in i..n { count += 1; } } count } }
Question: If we call count_pairs(n), how many times does the inner loop execute in total?
Can't figure out a formula? Try tracing it by hand with n = 3 or 4.
Part 1: n additions, one per time through the loop.
Part 2: For n = 4 the inner loop runs 3 + 2 + 1 = 6 times. In general (n-1) + (n-2) + ... + 1 = n(n-1)/2, which grows like n^2.
What is Big O?
Big O notation describes how runtime/memory grows as input size grows.
Key idea: We ignore:
- Exact number of operations
- Constants and performance on small inputs
- Hardware / OS dependent values
We focus on: The growth rate as n goes to infinity
So in Big O terms, B is the "constant" one, even though it lost for small n. Big O is about what happens as n gets big.
Example: Linear growth
#![allow(unused)] fn main() { fn print_up_to(n: u32) { for i in 0..n { // n iterations println!("{}", i); } } }
- n = 10: ~10 operations
- n = 100: ~100 operations
- Any n: ~n operations
This is O(n) - "linear time"
Example: Quadratic growth
#![allow(unused)] fn main() { fn print_all_pairs(n: u32) { for i in 0..n { // n iterations for j in 0..n { // n iterations for EACH i println!("{}, {}", i, j); } } } }
- n = 10: ~100 operations (10 × 10)
- n = 100: ~10,000 operations (100 × 100)
- Any n: ~n^2 operations
This is O(n^2) - "quadratic time"
Example: Exponential growth
A door opens for exactly one on/off setting of n light switches. Try them all:
#![allow(unused)] fn main() { fn try_every_setting(n: u32) -> u64 { let mut tried = 0; for _setting in 0..2u64.pow(n) { // the _ means we never use the variable tried += 1; } tried } }
- 10 switches: 1,024 settings (2^10)
- 20 switches: 1,048,576 settings (2^20)
- 40 switches: 1,099,511,627,776 settings (2^40)
Each extra switch doubles the work, like a population of bunnies doubling every generation.
This is O(2^n) - "exponential time" (explodes quickly!)
Example: Logarithmic growth
Now run the bunnies backwards. Start with 2, double every generation. How many generations until there are at least n?
#![allow(unused)] fn main() { fn generations_until(n: u64) -> u32 { let mut bunnies = 2; let mut generations = 0; while bunnies < n { bunnies *= 2; generations += 1; } generations } }
- 1,000 bunnies: 9 generations
- 1,000,000 bunnies: 19 generations
- 1,000,000,000 bunnies: 29 generations
Ask for 1,000x more bunnies and it only takes 10 more generations.
This is O(log n) - "logarithmic time" (very fast!)
Example: Constant time
#![allow(unused)] fn main() { fn last_digit(n: u64) -> u64 { n % 10 } }
- n = 10: 1 operation
- n = 1,000,000,000: 1 operation
- Any n: still 1 operation!
This is O(1) - "constant time" (doesn't depend on n)
Think about: What's the complexity?
#![allow(unused)] fn main() { fn evens_times_odds(n: u64) -> u64 { let mut evens = 0; for i in 0..n { if i % 2 == 0 { evens += 1; } } let mut odds = 0; for i in 0..n { if i % 2 == 1 { odds += 1; } } evens * odds } }
T/P/S - How does the work grow with n?
O(n). Two loops of n one after the other is 2n steps, and Big O drops the 2. A loop after a loop adds; a loop inside a loop multiplies.
Common complexity classes (from best to worst)
| Notation | Name | Example |
|---|---|---|
| O(1) | Constant | Array access by index |
| O(log n) | Logarithmic | Doubling until you reach n (the bunnies) |
| O(n) | Linear | Loop from 0 to n once |
| O(n log n) | Linearithmic | Good sorting algorithms |
| O(n^2) | Quadratic | Nested loops |
| O(2^n) | Exponential | Trying every on/off setting |
| O(n!) | Factorial | Trying all permutations |
Each step down this list is MUCH slower!
Rules for analyzing code
-
Loops: Multiply complexity by number of iterations
- Loop n times doing O(1) work = O(n)
- Loop n times doing O(n) work = O(n^2)
- Outer loop n times, inner loop m times = O(n m)
-
Drop constants and lower-order terms:
- O(3n) -> O(n)
- O(n^2 + n) -> O(n^2)
- O(5) -> O(1)
Let's do this one together
#![allow(unused)] fn main() { fn mystery_function(n: u64) -> u64 { let mut count = 0; for i in 0..n { count += i; } for _ in 0..10 { count += 1; } for i in 0..n { for j in 0..n { if i == j { count += 1; } } } count } }
n + 10 + n^2 steps. Drop the constant and the smaller term: O(n^2).
Space complexity exists too!
Big O also applies to memory usage. Back to the two prime finders:
- A keeps a few variables (
count,candidate,d) no matter how bigngets: O(1) space - B makes an array of 100,000
true/falsevalues before it does anything else. The memory it needs grows with its limit: O(limit) space
B buys its speed with memory. You make that trade all the time:
- Your browser caches websites. It keeps copies on disk, so a page you've visited loads without downloading it all again
- Looking for one thing in a box? Dump the whole box out on the floor. It takes up the whole floor, but now you can see everything at once
Best case vs. worst case vs. average case
Example: checking whether a single number is prime
#![allow(unused)] fn main() { fn is_prime(n: u64) -> bool { if n < 2 { return false; } for d in 2..n { if n % d == 0 { return false; } } true } }
- Best case: O(1) -
nis even, so it stops atd = 2 - Worst case: O(n) -
nis prime, so it tries everydfrom 2 to n - 1 - Average case: somewhere in between, and harder to work out
Usually we care most about worst case!
Sometimes the algorithm isn't even the whole story
Two ways to add up every number in a 10,000 by 10,000 grid:
// Across each row, then down to the next row
for r in 0..rows {
for c in 0..cols {
sum = sum + matrix[r][c];
}
}
// Down each column, then over to the next column
for c in 0..cols {
for r in 0..rows {
sum = sum + matrix[r][c];
}
}
Same 100 million additions. Same Big O. On my machine, going down the columns is about 6x slower (in release mode).
Why? We'll come back to that when we get to memory.
Activity Time
Complexity cheat sheet
Fast to Slow:
- O(1) - Instant, no matter the size
- O(log n) - Doubling the input adds one step
- O(n) - Proportional to size
- O(n log n) - The best we can do for sorting
- O(n^2) - Nested loops, gets bad quickly
- O(2^n) - Explodes! Avoid if possible