Gradient boosting
Boosted trees for classification or regression, with holdout fit and feature importance.
When to use it
Use gradient boosting when you want an ensemble of trees that can run as classification or as regression from one dialog. Mode starts at classification. Random forest is the bagged counterpart: classification or regression.
Assumptions
Classification: categorical outcome, numeric features, classes with fewer than two rows dropped. Regression: numeric outcome. Holdout fraction 0.25, 100 trees. HTTP and chat dispatch both honour test_fraction. Learning rate and maximum depth are not exposed.
Running it in Tensr
Options
Prop
Type
Reading the output
The same strong pass/fail sample, classification mode, 100 stages, 25% holdout.
| Actual \ Predicted | 0 | 1 |
|---|---|---|
| 0 | 27 | 6 |
| 1 | 11 | 16 |
| Variable | Importance |
|---|---|
| hours | 0.792 |
| anxiety | 0.208 |
Train accuracy = 0.994, test accuracy = 0.717, n = 240. Hours was built to move the pass rate strongly, so holdout accuracy should sit clearly above chance.
Reporting (APA 7)
Holdout classification of passed from hours and anxiety, n = 240. Test accuracy = 0.717, train accuracy = 0.994.
Coming from SPSS
SPSS has no gradient-boosting dialog. Tensr puts classification and regression on the ML menu as two labels of this one procedure. The path string stored for this item says ML → Gradient Boosting.
Related
Random forest classification is bagging. Neural network is the other mode-switched model.