Analyses

Gradient boosting

Boosted trees for classification or regression, with holdout fit and feature importance.

Edit on GitHub

When to use it

Use gradient boosting when you want an ensemble of trees that can run as classification or as regression from one dialog. Mode starts at classification. Random forest is the bagged counterpart: classification or regression.

Assumptions

Classification: categorical outcome, numeric features, classes with fewer than two rows dropped. Regression: numeric outcome. Holdout fraction 0.25, 100 trees. HTTP and chat dispatch both honour test_fraction. Learning rate and maximum depth are not exposed.

Running it in Tensr

ML → Classification → Gradient Boosting (Classification), or ML → Regression → Gradient Boosting (Regression). In chat: “Gradient boosting of passed on hours and anxiety.”
Set mode to classification or regression. Target matches that mode. Features need at least one numeric column.

Options

Prop

Type

Reading the output

The same strong pass/fail sample, classification mode, 100 stages, 25% holdout.

Actual \ Predicted01
0276
11116
VariableImportance
hours0.792
anxiety0.208

Train accuracy = 0.994, test accuracy = 0.717, n = 240. Hours was built to move the pass rate strongly, so holdout accuracy should sit clearly above chance.

Reporting (APA 7)

Holdout classification of passed from hours and anxiety, n = 240. Test accuracy = 0.717, train accuracy = 0.994.

Coming from SPSS

SPSS has no gradient-boosting dialog. Tensr puts classification and regression on the ML menu as two labels of this one procedure. The path string stored for this item says ML → Gradient Boosting.

Random forest classification is bagging. Neural network is the other mode-switched model.