Root-causing a crash: from stack trace to primitive

A crash is not a vulnerability. Turning one into the other is most of the work, and it’s the part that’s least written about.

The sequence I follow:

  1. Reproduce deterministically. Non-deterministic crashes waste enormous amounts of time. Disable ASLR, fix the heap layout, record with rr if you can.
  2. Classify the crash. Read vs. write, near-null vs. wild, controlled offset vs. not. !exploitable and its clones are a rough first pass, not an answer.
  3. Find the originating allocation. Which object, what size, who allocated it. ASAN or a heap tracer earns its overhead here.
  4. Identify the primitive. What do you actually get? Relative write of N bytes? Arbitrary free? Type confusion between two structs? Say it precisely — “heap overflow” is not precise.
  5. Assess reachability. Can the primitive be driven from attacker-controlled input, and how much of it do you control?

Step 4 is where people rush. If you can’t state the primitive in one sentence, you don’t understand the bug yet, and every exploitation attempt after that is guesswork.

What’s your process? Particularly interested in how people handle step 3 on targets where ASAN isn’t an option.