warming up your workspace

Why floating point lies, and the arithmetic you have to unlearn

Type 0.1 + 0.2 into almost any programming language and you get 0.30000000000000004. This is the single most reported non-bug in software history, and the reactions split people into two groups: those who shrug it off as a rounding quirk, and those who understand that it is the visible tip of a set of rules that quietly govern every scientific calculation, every financial total, every physics simulation. Floating-point arithmetic is not the real-number arithmetic you learned in school. It breaks laws you assume are unbreakable, and the only defense is to know exactly which ones and why.

The one idea

A 64-bit float has room for about 15 to 17 significant decimal digits, and it stores numbers in binary. Most decimal fractions, including 0.1, have no exact binary representation, the same way 1/3 has no exact decimal one. So the moment you write 0.1, the computer stores the nearest value it can, which is slightly off. Every operation then rounds its result back to the nearest representable number. Individually these errors are tiny, around one part in ten quadrillion. The danger is that they interact: certain operations magnify the tiny errors into large ones, and certain sequences of operations lose digits that never come back. The rules of real arithmetic that fail are the ones that assume infinite precision.

The four things that break

Equality is not exact. The stored values of 0.1 and 0.2 are each a hair off, and their sum rounds to a value a hair off from the stored 0.3.

println(0.1 + 0.2)            # 0.30000000000000004
println(0.1 + 0.2 == 0.3)     # false

Never compare floats with ==. Compare with a tolerance: is the difference smaller than some tiny epsilon.

Addition is not associative. Grouping changes the answer, because a small number added to a huge one can vanish entirely, and whether it vanishes depends on the order.

a, b, c = 1e16, -1e16, 1.0
println((a + b) + c)          # 1.0  -- the 1.0 survives
println(a + (b + c))          # 0.0  -- the 1.0 is swallowed by 1e16, then cancelled

In a + (b + c), the 1.0 is added to -1e16 first, but 1e16 has no bits left to represent a difference of one, so b + c rounds straight back to -1e16 and the 1.0 is gone. This is why parallel sums, which regroup additions across threads, can give different results run to run.

Catastrophic cancellation. Subtracting two nearly equal numbers cancels their shared leading digits, leaving only the noisy trailing ones. The classic trap is the "computational" variance formula, mean of squares minus square of mean:

xs = [1e8 + i for i in 0:4]                    # true variance is 2.0
naive  = (sum(x^2 for x in xs) - sum(xs)^2/length(xs)) / length(xs)
stable = sum((x - mean(xs))^2 for x in xs) / length(xs)
naive variance:  3.2       (wrong)
stable variance: 2.0       (right)

Both formulas are algebraically identical. But the naive one subtracts two numbers near 10^16, and the answer, a mere 2, lives in digits that got rounded away before the subtraction. The stable form subtracts first, while the numbers are small, and keeps every digit.

Small numbers get lost in long sums. Add ten million copies of 0.1 and the running total grows large enough that each new 0.1 falls below its precision and is under-counted.

v = fill(0.1, 10_000_000)
println(sum(v))              # 999999.9999999986, not 1000000.0
println(kahan(v))            # 1000000.0  -- Kahan summation recovers it

Kahan summation keeps a second variable tracking the low-order bits that each addition throws away, and feeds them back in. It is a few extra lines that turn a drifting sum into an exact one.

Three details that matter:

  • The errors are not random, they are deterministic and analyzable. Numerical analysis is the field that predicts, for a given formula, how much precision you will lose, so you can pick a formula that keeps it.
  • The fix is almost always to rearrange the math, not to add more precision. The stable variance and Kahan summation cost no extra memory; they just order the operations so cancellation never happens.
  • Higher precision delays the problem, it does not remove it. 128-bit floats push the errors smaller but the same laws break at the same places. Correct algorithms beat bigger numbers.

Where this shows up

This is why a bank never stores money in floats, it uses integers of cents or a decimal type, so 0.10 + 0.20 is exactly 0.30. It is why two runs of the same parallel simulation can diverge, why a naive statistics library can report a negative variance, and why numerical code is full of formulas that look needlessly complicated, they are the rearranged, cancellation-proof versions. Every serious numerical library, LAPACK, the routines under NumPy and Julia, is built by people who treat these rules as the ground truth they are.

If you want to build the stable algorithms, condition numbers, error analysis, and the linear algebra that respects them, that is what the scientific-computing track on IWTLP teaches in Julia, where numerical correctness is the whole game.