● for teams shipping XGBoost / LightGBM / scikit-learn models as ONNX

Your model was converted to ONNX.
Is it still the same model?

Testing on a sample can't tell you. Conversion bugs live in slivers of the input space one float step wide — around a threshold, at exactly zero, or only for missing values. Leafparity doesn't sample: it reads both models and proves, for every possible input, whether they agree — or shows you exactly where they don't.

leafparity check model.txt model.onnx
$ leafparity check model.txt model.onnx --background sample.csv

VERDICT: NOT EQUIVALENT
Not equivalent: 4 distinct problem(s) at 420 place(s) in the trees. For some
    inputs the raw outputs differ by 290.952 (proven to be the largest possible
    difference).

#1  zero / near-zero values are handled differently  [feature 'age']
    occurs at 30 split node(s) in 29 tree(s); largest effect of a single tree: 23.02
    example - tree 3:
      original  node 0: x <= -17.066402760988257 (double), missing_type=Zero, default -> left
      converted node 0: x <= -17.066402435302734 (float32), missing -> true branch
      inputs routed differently here: [-1.0000000180025095e-35, 1.0000000180025095e-35]
    witness: {age=0.0, income=114.41658720372287, ...}
      original predicts -196.355; converted predicts 57.9331  (verified by running both real runtimes)

How it works

Exact, not sampled.

Leafparity reproduces each runtime's exact arithmetic — every cast, preprocessing step, comparison operator and missing-value rule — then walks both models together over the whole input space at floating-point precision. There are only two outcomes.

✓

EQUIVALENT

A proof that for every possible input, the two models reach corresponding leaves. Outputs can differ only by a stated, proven floating-point rounding allowance — not an estimate, a bound.

✕

NOT EQUIVALENT

Every place in every tree where the models disagree, the exact set of inputs affected, and a concrete witness input for each problem — verified by running both real runtimes so you can check it yourself.

Real, reproduced bugs

What it catches

Reproduced in the shipped examples/ folder — not hypothetical.

Case What leafparity reports
The sklearn-onnx docs' own "float switch" example (StandardScaler + DecisionTreeRegressor) Their own test set shows a largest error of ~190. Leafparity proves the largest possible error is 556 (721 with missing values), shows exactly where each discrepancy sits, and certifies their CastTransformer fix as EQUIVALENT for float32 — but proves the fix does not hold for float64 inputs.
LightGBM with zero_as_missing=True, converted via onnxmltools The converter silently ignores LightGBM's missing_type=Zero rule, so an input of exactly 0.0 takes a different path in every tree. Raw scores differ by up to ~290.
scikit-learn trees (≥1.3) receiving NaN, converted via skl2onnx scikit-learn routes NaN using missing_go_to_left; the converted model sends NaN the other way at the affected nodes — a silent class flip.
LightGBM (double thresholds) served through float32 ONNX Narrow bands of float64 input next to thresholds route differently at almost every node. Leafparity lists them, bounds their effect, and tells you whether your actual float32 data can even reach them.

Why trust the verdict

It checks its own work, three times.

The report shows the result of every layer below — nothing is asserted without evidence.

1

Self-check, before any analysis

Inputs are constructed to reach every split node and sit exactly on — and one float step either side of — its decision boundary, including zeros and NaN. They run through the real libraries and through leafparity's exact model. If a single tree decision disagrees, leafparity refuses to certify rather than guess.

2

Witness verification

Every reported problem, and the worst case, comes with an input that has actually been run through both real runtimes. The report states whether the observed difference matches the predicted one.

3

Independent cross-check

Thousands of adversarial and boundary inputs run through both real runtimes without using the analysis engine at all. Every observed difference must fall within the proven bound, or the certificate is refused.

Coverage

Supported today

Not yet supported constructs are refused outright, never guessed at.

Original model

  • XGBoost — gbtree, numerical splits; regression, binary, multiclass
  • LightGBM — gbdt / rf, all missing-value modes (None / Zero / NaN)
  • scikit-learn — DecisionTree, ExtraTree, RandomForest, ExtraTrees, GradientBoosting
  • Optional Pipeline preprocessing: StandardScaler, MinMaxScaler, MaxAbsScaler, RobustScaler, CastTransformer

Converted model

  • ONNX ai.onnx.ml TreeEnsembleRegressor / TreeEnsembleClassifier (opset ≤ 3)
  • Optional Cast / Scaler / Add / Sub / Mul / Div preprocessing by constants
  • Float or double input, as produced by skl2onnx and onnxmltools
not yet: categorical splits XGBoost dart LightGBM linear trees opset-5 TreeEnsemble PMML

Pricing

Simple, honest, self-serve.

Founding-customer pricing — no "% off" games. If you sign up now, you keep this rate when it rises later.

Per-model certificate
One model, one conversion, one answer — for a governance file.
$1,200 one-off
Delivered as a report you can hand to auditors or risk sign-off.
  • Full EQUIVALENT / NOT EQUIVALENT proof
  • Every discrepancy, with a verified witness
  • JSON + human-readable report
Request a certificate
CI plan founding rate
Run it on every model you ship, automatically.
$120/month
Priced to not need procurement approval.
  • GitHub Action, drop-in CI step
  • Exit codes for CI gating (--fail-above)
  • Unlimited checks on your models
  • Rate locked in for life as an early customer
Start the CI plan
Enterprise
Multiple teams, models and pipelines.
$9,000/year
For organizations standardizing equivalence checks across teams.
  • Everything in the CI plan, org-wide
  • Priority support
  • Direct line for new format requests
Talk to us

Don't find out in production.

Send us your model and its ONNX conversion. We'll tell you, with proof, whether they agree.