Analyses

Random forest classification

Bagged trees for a categorical outcome, with holdout accuracy and feature importance.

Edit on GitHub

When to use it

Use random forest classification when the outcome is a category and you want holdout accuracy plus which numeric features split the classes, without reading one tree. A single tree is decision tree. The numeric-outcome version is random forest regression.

Assumptions

The outcome is categorical. Features are numeric. Classes with fewer than two rows are dropped. Holdout fraction 0.25, 100 trees, min_samples_leaf 5, no depth cap, random_state=0. HTTP and chat dispatch both honour test_fraction and min_samples_leaf. Kernel and class weights are not exposed.

Running it in Tensr

ML → Classification → Random Forest Classification. In chat: “Random forest of passed on hours and anxiety.”
Target is the categorical column. Features need at least one numeric column.

Options

Prop

Type

Reading the output

A separate 240-row sample where hours strongly shifts the pass rate. 100 trees, min_samples_leaf 5, no depth cap, 25% holdout. Test accuracy should sit clearly above chance.

Actual \ Predicted01
0267
11017
VariableImportance
hours0.798
anxiety0.202

Train accuracy = 0.9, test accuracy = 0.717, n = 240, trees = 100, min_samples_leaf = 5. Hours was built to move the pass rate strongly, so holdout accuracy should sit clearly above chance.

Reporting (APA 7)

Holdout classification of passed from hours and anxiety, n = 240. Test accuracy = 0.717, train accuracy = 0.9.

Coming from SPSS

SPSS has no random-forest dialog. Tensr’s item is on the ML menu, Classification → Random Forest Classification. The path string stored for this item matches that menu.

Decision tree is one tree. Gradient boosting is the boosting counterpart. SVM classification is another holdout classifier.