{
  "id": 11753301,
  "title": "I converted 126 tree models to ONNX by following the docs. Here is what changed.",
  "url": "https://urgent.news/2026/10/03/i-converted-126-tree-models-to-onnx-by-following-the-docs-here-is",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-03T19:27:58.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/milivoje_simonovi_ddc92b/i-converted-126-tree-models-to-onnx-by-following-the-docs-here-is-what-changed-3a3j"
  },
  "original_language": "en",
  "account": "A reporter examined how accurately the documented method of converting tree models to ONNX format preserves model equivalence. The tester constructed 126 model pairs using the standard procedures outlined in the skl2onnx and onnxmltools documentation, employing default parameters and 32-bit floating point inputs. The model types included XGBoost, scikit-learn GradientBoosting, scikit-learn RandomForest, scikit-learn ExtraTrees, scikit-learn DecisionTree, and LightGBM. Each pair underwent dual testing, once considering only finite inputs and once incorporating NaN values as well. The ONNX files were then evaluated against their original counterparts using an equivalence checker. Out of the 126 pairs, only one LightGBM model exhibited discrepancies due to differences in how the two platforms handle float32 versus float64 threshold values. This discrepancy could be significant when dealing with decimals as features, as round decimals were precisely the values users might input. In contrast, all XGBoost and scikit-learn GradientBoosting models maintained equivalence, even when encountering NaN values during predictions. The author, the creator of the equivalence checker, invites further discussion on any discovered issues or confirmations of the findings.",
  "summary": "I wanted to know how often the documented way of converting a tree model to ONNX produces a file that is not the same model. So I built the pairs the way the skl2onnx and onnxmltools documentation shows, with default settings and float32 input, and checked every one with an equivalence checker. The setup: 7 datasets (iris, wine, breast cancer, digits, diabetes, a synthetic regression, and a…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}