Skip to contents

Clusters wide-format sequences using pairwise string dissimilarity and either PAM (Partitioning Around Medoids) or hierarchical clustering. Supports 9 distance metrics including temporal weighting for Hamming distance. When the stringdist package is available, uses C-level distance computation for 100-1000x speedup on edit distances.

Usage

build_clusters(
  data,
  k,
  dissimilarity = "hamming",
  method = "pam",
  na_syms = c("*", "%"),
  weighted = FALSE,
  lambda = 1,
  seed = NULL,
  q = 2L,
  p = 0.1,
  covariates = NULL,
  estimator = c("auto", "firth", "multinom", "chisq"),
  ...
)

# S3 method for class 'net_clustering'
print(x, digits = 3L, ...)

# S3 method for class 'net_clustering'
summary(object, ...)

# S3 method for class 'net_clustering'
plot(
  x,
  type = c("silhouette", "mds", "heatmap", "predictors"),
  combined = TRUE,
  ...
)

# S3 method for class 'tidy_covariates'
print(x, ...)

Arguments

data

Input data. Accepts multiple formats:

data.frame / matrix

Wide-format sequences (rows = sequences, columns = time points, values = state names).

netobject

A network object from build_network. Extracts the stored sequence data. Only valid for sequence-based methods (relative, frequency, co_occurrence, attention).

tna

A tna model from the tna package. Decodes the integer-encoded sequence data using stored labels.

cograph_network

A cograph network object. Extracts the stored sequence data.

k

Integer. Number of clusters (must be between 2 and nrow(data) - 1).

dissimilarity

Character. Distance metric. One of "hamming", "osa" (optimal string alignment), "lv" (Levenshtein), "dl" (Damerau-Levenshtein), "lcs" (longest common subsequence), "qgram", "cosine", "jaccard", "jw" (Jaro-Winkler). Default: "hamming".

method

Character. Clustering method. "pam" for Partitioning Around Medoids, or a hierarchical method: "ward.D2", "ward.D", "complete", "average", "single", "mcquitty", "median", "centroid". Default: "pam".

na_syms

Character vector. Symbols treated as missing values. Default: c("*", "%").

Missing-value distance rule: after symbols are converted to NA, missing values are encoded as a single comparable sentinel state – not pairwise-deleted. Two missing values in the same position match (distance contribution 0); a missing value paired with any observed state mismatches (distance contribution 1 for Hamming, etc.). This is the conventional behaviour for aligned sequence matrices because pairwise deletion would change the effective length of every pair and break the metric. If you want pairwise deletion or a different missing-value semantic, drop or recode the missing cells before passing the data in.

weighted

Logical. Apply exponential decay weighting to Hamming distance positions? Only valid when dissimilarity = "hamming". Default: FALSE.

lambda

Numeric. Non-negative decay rate for weighted Hamming. Higher values weight earlier positions more strongly. Default: 1.

seed

Integer or NULL. Random seed for reproducibility. Default: NULL.

q

Integer. Size of q-grams for "qgram", "cosine", and "jaccard" distances. Default: 2L.

p

Numeric. Winkler prefix penalty for Jaro-Winkler distance. Must be between 0 and 0.25. Default: 0.1.

covariates

Optional. Post-hoc covariate analysis of cluster membership. Accepts:

string

Single column name, e.g. "Age". Resolved against x$metadata (and x$data) for netobject or cograph_network input.

character vector

c("Age", "Gender"), same lookup.

formula

~ Age + Gender, same lookup; supports "Age + Gender" string form too.

data.frame

All columns used as covariates verbatim; must have one row per sequence.

NULL

No covariate analysis (default).

For netobject or cograph_network input, names are resolved against $metadata first and then non-state columns of $data, so a typical call looks like build_clusters(net, k = 3, covariates = "session_label") without pre-extracting a data.frame. tna input requires the data.frame form. Results are stored in $covariates.

estimator

Multinomial logit fitter for the covariate analysis. "auto" (default) inspects the cluster x covariate cross-tab and falls back to "firth" only when any cell has fewer than 5 observations (quasi-complete separation risk); otherwise uses the much faster "multinom". "firth" forces Firth's penalised likelihood via brglm2::brmultinom – bias-reduced and finite under separation, but ~200x slower than multinom on well-conditioned data. "multinom" forces classical ML via nnet::multinom; warns because rare-cell separation produces astronomical ORs with degenerate CIs (silent failure). "chisq" runs WeightedCluster-style descriptive tests (chi-square + Cramer's V + standardized adjusted residuals for factors; Kruskal-Wallis + eta-squared for numerics).

...

Unsupported. Supplying unused arguments raises an error. In plot.net_clustering(), print.net_clustering() and summary.net_clustering(): Unsupported. Supplying unused arguments raises an error. In print.tidy_covariates(): Ignored.

x

For the print() and plot() methods: an object of class net_clustering or tidy_covariates.

digits

Integer. Decimal places used for floating-point statistics in the printout. Default 3. Non-breaking: existing print(x) calls keep their previous formatting.

object

For the summary() method: an object of class net_clustering.

type

Character. Plot type: "silhouette" (per-observation silhouette bars), "mds" (2D MDS projection), "heatmap" (distance matrix heatmap ordered by cluster), or "predictors" (odds-ratio forest plot of the post-hoc covariate analysis; requires covariates and an estimator that produces coefficients). Default: "silhouette".

combined

Logical. For type = "predictors" only: when TRUE (default), covariate forest panels are combined into a single faceted plot; when FALSE, a list of separate ggplots is returned.

Value

An object of class "net_clustering" containing:

data

The original input data.

k

Number of clusters.

assignments

Named integer vector of cluster assignments.

silhouette

Overall average silhouette width.

sizes

Named integer vector of cluster sizes.

method

Clustering method used.

dissimilarity

Distance metric used.

distance

The computed dissimilarity matrix (dist object).

medoids

Integer vector of medoid row indices (PAM only; NULL for hierarchical methods).

seed

Seed used (or NULL).

weighted

Logical, whether weighted Hamming was used.

lambda

Lambda value used (0 if not weighted).

covariates

The post-hoc covariate analysis (a list; see the estimator argument), or NULL when covariates = NULL.

network_method, build_args

For netobject input, the source network's method and stored build arguments, so per-cluster networks can be rebuilt the same way. NULL otherwise.

metadata

For netobject input, its per-sequence metadata, one row per clustered sequence, so session_ids can name each sequence. NULL otherwise.

htna_partition

For HTNA input, the preserved node-to-actor partition used to restore HTNA children when networks are built.

In print.net_clustering(): The input object, invisibly.

In summary.net_clustering(): A data frame of per-cluster statistics, one row per cluster, with columns cluster, size and mean_within_dist, returned visibly. When the clustering was fitted with covariates, a tidy_covariates/data.frame (the tidied covariate table, with cluster sizes, fit statistics and profiles attached as attributes) is returned invisibly instead. In both cases the printed summary is a side effect.

In plot.net_clustering(): A ggplot object (invisibly); for type = "predictors" with combined = FALSE, a list of ggplot objects named by cluster (invisibly).

In print.tidy_covariates(): The input invisibly.

Methods

  • print.net_clustering(): Compact, fixed-width summary of a sequence-clustering result. The header carries the clustering method and dissimilarity; the per-cluster table carries cluster size (count and percentage) and mean within-cluster distance when available. Optional medoid and covariate lines surface only when those fields are populated.

  • print.tidy_covariates(): Prints a one-line header naming the estimator, then the data.frame. The full human-readable view (per-cluster stats, profiles, OR/test tables) was already printed by summary() when this object was produced, so this method intentionally stays minimal to avoid duplication. Auto-prints when the user types the variable at the REPL.

Examples

seqs <- data.frame(V1 = c("A","B","C","A","B"), V2 = c("B","C","A","B","A"),
                   V3 = c("C","A","B","C","B"))
cl <- build_clusters(seqs, k = 2)
cl
#> Sequence Clustering [pam]
#>   Sequences: 5  |  Clusters: 2
#>   Dissimilarity: hamming
#>   Quality: silhouette = 0.600
#> 
#>   Cluster  N          Mean within-dist  Medoid
#>   1        2 (40.0%)  0.000             4
#>   2        3 (60.0%)  2.000             5
# \donttest{
seqs <- data.frame(
  V1 = sample(LETTERS[1:3], 20, TRUE), V2 = sample(LETTERS[1:3], 20, TRUE),
  V3 = sample(LETTERS[1:3], 20, TRUE), V4 = sample(LETTERS[1:3], 20, TRUE)
)
cl <- build_clusters(seqs, k = 2)
print(cl)
#> Sequence Clustering [pam]
#>   Sequences: 20  |  Clusters: 2
#>   Dissimilarity: hamming
#>   Quality: silhouette = 0.222
#> 
#>   Cluster  N           Mean within-dist  Medoid
#>   1        9 (45.0%)   2.444             13
#>   2        11 (55.0%)  2.236             14
summary(cl)
#> Sequence Clustering Summary
#>   Method:        pam 
#>   Dissimilarity: hamming 
#>   Silhouette:    0.2222 
#> 
#> Per-cluster statistics:
#>  cluster size mean_within_dist
#>        1    9         2.444444
#>        2   11         2.236364
#>   cluster size mean_within_dist
#> 1       1    9         2.444444
#> 2       2   11         2.236364
# }