Random forest classification
Bagged trees for a categorical outcome, with holdout accuracy and feature importance.
When to use it
Use random forest classification when the outcome is a category and you want holdout accuracy plus which numeric features split the classes, without reading one tree. A single tree is decision tree. The numeric-outcome version is random forest regression.
Assumptions
The outcome is categorical. Features are numeric. Classes with fewer than two rows are dropped. Holdout fraction 0.25, 100 trees, min_samples_leaf 5, no depth cap, random_state=0. HTTP and chat dispatch both honour test_fraction and min_samples_leaf. Kernel and class weights are not exposed.
Running it in Tensr
Options
Prop
Type
Reading the output
A separate 240-row sample where hours strongly shifts the pass rate. 100 trees, min_samples_leaf 5, no depth cap, 25% holdout. Test accuracy should sit clearly above chance.
| Actual \ Predicted | 0 | 1 |
|---|---|---|
| 0 | 26 | 7 |
| 1 | 10 | 17 |
| Variable | Importance |
|---|---|
| hours | 0.798 |
| anxiety | 0.202 |
Train accuracy = 0.9, test accuracy = 0.717, n = 240, trees = 100, min_samples_leaf = 5. Hours was built to move the pass rate strongly, so holdout accuracy should sit clearly above chance.
Reporting (APA 7)
Holdout classification of passed from hours and anxiety, n = 240. Test accuracy = 0.717, train accuracy = 0.9.
Coming from SPSS
SPSS has no random-forest dialog. Tensr’s item is on the ML menu, Classification → Random Forest Classification. The path string stored for this item matches that menu.
Related
Decision tree is one tree. Gradient boosting is the boosting counterpart. SVM classification is another holdout classifier.