← Charlie Yan

boostkit

2026 — Python · numpy · LightGBM

Repository ↗

A histogram GBDT written out in numpy: quantile binning, leaf-wise growth with histogram subtraction, LightGBM’s gain and leaf formulas, L2 and binary objectives, early stopping. Plus exact TreeSHAP.

What’s tested

The histogram and gain identities. LightGBM parity on identical bins: same first split, prediction correlation above 0.9999. TreeSHAP against brute-force Shapley enumeration to 1e-9. Additivity on every prediction, phi.sum(1) + base == pred, to 1e-9.

The one rule

Every number in the README is produced by boostkit bench on the machine it names. CI runs boostkit bench --check, which regenerates the benchmark and compares: first splits must match exactly, metrics within 1%.

LightGBM already does all of this, faster. The point is that every step is written out and checked against it.