Cluster analysis
K-means or hierarchical partitions of rows on numeric columns.
When to use it
Use cluster analysis when you want to split people into a chosen number of groups based on numeric similarity, not a known label. “Two style groups from two well-separated columns” is the shape of this example. Density clusters that can leave points as noise are DBSCAN.
Assumptions
Columns are numeric. K-means is the dialog default, with three clusters, and variables standardized. Hierarchical uses Ward linkage on the same standardized columns. HTTP and chat dispatch both honour standardize. Number of clusters is its own field, not shared with PCA.
Running it in Tensr
Options
Prop
Type
Reading the output
Two numeric columns built as two well-separated blobs, 240 rows, k-means with k = 2, standardised. Silhouette should be high.
| Cluster | x | y |
|---|---|---|
| 0 | 2.215 | 1.972 |
| 1 | -2.121 | -2.044 |
| Cluster | Within-cluster SS | Silhouette |
|---|---|---|
| 0 | 9.494 | 0.872 |
| 1 | 10.465 | 0.865 |
kmeans, k = 2, silhouette = 0.869. The two columns were built as separate blobs, so a high silhouette is the expected reading.
Reporting (APA 7)
K-means clustering of x and y into 2 groups, n = 240, silhouette = 0.869.
Coming from SPSS
SPSS is Analyze → Classify → K-Means Cluster or Hierarchical Cluster. Tensr’s Analyze menu is Classification → Cluster Analysis, with k-means as the default. The path string stored for this item says Analyze → Classify → Hierarchical Cluster even when the method is k-means.
Related
DBSCAN finds dense regions and can label noise. A known group label is discriminant analysis.