Analyses

Cluster analysis

K-means or hierarchical partitions of rows on numeric columns.

Edit on GitHub

When to use it

Use cluster analysis when you want to split people into a chosen number of groups based on numeric similarity, not a known label. “Two style groups from two well-separated columns” is the shape of this example. Density clusters that can leave points as noise are DBSCAN.

Assumptions

Columns are numeric. K-means is the dialog default, with three clusters, and variables standardized. Hierarchical uses Ward linkage on the same standardized columns. HTTP and chat dispatch both honour standardize. Number of clusters is its own field, not shared with PCA.

Running it in Tensr

Analyze → Classification → Cluster Analysis. In chat: “K-means on hours and anxiety, three clusters.”
Put the numeric columns in the clustering list. Method starts at k-means. Number of clusters starts at 3. Standardize starts on.

Options

Prop

Type

Reading the output

Two numeric columns built as two well-separated blobs, 240 rows, k-means with k = 2, standardised. Silhouette should be high.

Clusterxy
02.2151.972
1-2.121-2.044
ClusterWithin-cluster SSSilhouette
09.4940.872
110.4650.865

kmeans, k = 2, silhouette = 0.869. The two columns were built as separate blobs, so a high silhouette is the expected reading.

Reporting (APA 7)

K-means clustering of x and y into 2 groups, n = 240, silhouette = 0.869.

Coming from SPSS

SPSS is Analyze → Classify → K-Means Cluster or Hierarchical Cluster. Tensr’s Analyze menu is Classification → Cluster Analysis, with k-means as the default. The path string stored for this item says Analyze → Classify → Hierarchical Cluster even when the method is k-means.

DBSCAN finds dense regions and can label noise. A known group label is discriminant analysis.