Analyses

DBSCAN

Density clustering that can label points as noise.

Edit on GitHub

When to use it

Use DBSCAN when you want clusters of dense points and you are willing to leave sparse points as noise, rather than forcing every row into k groups. A fixed k is cluster analysis.

Assumptions

Columns are numeric. Epsilon starts at 0.5, minimum samples at 5, standardize on. HTTP and chat dispatch both honour standardize. If more than 30% of rows are noise, the report warns. There is no k to choose.

Running it in Tensr

ML → Clustering → DBSCAN. In chat: “DBSCAN on hours and anxiety.”
Put the numeric columns in the feature list. Epsilon, minimum samples, and standardize start at 0.5, 5, and on.

Options

Prop

Type

Reading the output

The two blobs, ε = 0.8, min_samples = 5, standardised. Dense cores should appear, with a modest noise share.

ClusterSize
0120
1120

2 dense cluster(s), 0 noise points (0%). ε = 0.8 on two well-separated blobs after standardising should recover the cores and leave a modest noise share.

Reporting (APA 7)

DBSCAN on x and y, ε = 0.8, min_samples = 5, 2 cluster(s), 0 noise points.

Coming from SPSS

SPSS has no DBSCAN dialog. Tensr’s item is on the ML menu, Clustering → DBSCAN. The path string stored for this item matches that menu.

A chosen number of groups is cluster analysis.