Clusters wide-format sequences using pairwise string dissimilarity and either PAM (Partitioning Around Medoids) or hierarchical clustering. Supports 9 distance metrics including temporal weighting for Hamming distance. When the stringdist package is available, uses C-level distance computation for 100-1000x speedup on edit distances.
Usage
build_clusters(
data,
k,
dissimilarity = "hamming",
method = "pam",
na_syms = c("*", "%"),
weighted = FALSE,
lambda = 1,
seed = NULL,
q = 2L,
p = 0.1,
covariates = NULL,
estimator = c("auto", "firth", "multinom", "chisq"),
...
)
# S3 method for class 'net_clustering'
print(x, digits = 3L, ...)
# S3 method for class 'net_clustering'
summary(object, ...)
# S3 method for class 'net_clustering'
plot(
x,
type = c("silhouette", "mds", "heatmap", "predictors"),
combined = TRUE,
...
)
# S3 method for class 'tidy_covariates'
print(x, ...)Arguments
- data
Input data. Accepts multiple formats:
- data.frame / matrix
Wide-format sequences (rows = sequences, columns = time points, values = state names).
- netobject
A network object from
build_network. Extracts the stored sequence data. Only valid for sequence-based methods (relative, frequency, co_occurrence, attention).- tna
A tna model from the tna package. Decodes the integer-encoded sequence data using stored labels.
- cograph_network
A cograph network object. Extracts the stored sequence data.
- k
Integer. Number of clusters (must be between 2 and
nrow(data) - 1).- dissimilarity
Character. Distance metric. One of
"hamming","osa"(optimal string alignment),"lv"(Levenshtein),"dl"(Damerau-Levenshtein),"lcs"(longest common subsequence),"qgram","cosine","jaccard","jw"(Jaro-Winkler). Default:"hamming".- method
Character. Clustering method.
"pam"for Partitioning Around Medoids, or a hierarchical method:"ward.D2","ward.D","complete","average","single","mcquitty","median","centroid". Default:"pam".- na_syms
Character vector. Symbols treated as missing values. Default:
c("*", "%").Missing-value distance rule: after symbols are converted to
NA, missing values are encoded as a single comparable sentinel state – not pairwise-deleted. Two missing values in the same position match (distance contribution 0); a missing value paired with any observed state mismatches (distance contribution 1 for Hamming, etc.). This is the conventional behaviour for aligned sequence matrices because pairwise deletion would change the effective length of every pair and break the metric. If you want pairwise deletion or a different missing-value semantic, drop or recode the missing cells before passing the data in.- weighted
Logical. Apply exponential decay weighting to Hamming distance positions? Only valid when
dissimilarity = "hamming". Default:FALSE.- lambda
Numeric. Non-negative decay rate for weighted Hamming. Higher values weight earlier positions more strongly. Default: 1.
- seed
Integer or NULL. Random seed for reproducibility. Default:
NULL.- q
Integer. Size of q-grams for
"qgram","cosine", and"jaccard"distances. Default:2L.- p
Numeric. Winkler prefix penalty for Jaro-Winkler distance. Must be between 0 and 0.25. Default:
0.1.- covariates
Optional. Post-hoc covariate analysis of cluster membership. Accepts:
- string
Single column name, e.g.
"Age". Resolved againstx$metadata(andx$data) fornetobjectorcograph_networkinput.- character vector
c("Age", "Gender"), same lookup.- formula
~ Age + Gender, same lookup; supports"Age + Gender"string form too.- data.frame
All columns used as covariates verbatim; must have one row per sequence.
- NULL
No covariate analysis (default).
For
netobjectorcograph_networkinput, names are resolved against$metadatafirst and then non-state columns of$data, so a typical call looks likebuild_clusters(net, k = 3, covariates = "session_label")without pre-extracting a data.frame.tnainput requires the data.frame form. Results are stored in$covariates.- estimator
Multinomial logit fitter for the covariate analysis.
"auto"(default) inspects the cluster x covariate cross-tab and falls back to"firth"only when any cell has fewer than 5 observations (quasi-complete separation risk); otherwise uses the much faster"multinom"."firth"forces Firth's penalised likelihood viabrglm2::brmultinom– bias-reduced and finite under separation, but ~200x slower than multinom on well-conditioned data."multinom"forces classical ML viannet::multinom; warns because rare-cell separation produces astronomical ORs with degenerate CIs (silent failure)."chisq"runs WeightedCluster-style descriptive tests (chi-square + Cramer's V + standardized adjusted residuals for factors; Kruskal-Wallis + eta-squared for numerics).- ...
Unsupported. Supplying unused arguments raises an error. In
plot.net_clustering(),print.net_clustering()andsummary.net_clustering(): Unsupported. Supplying unused arguments raises an error. Inprint.tidy_covariates(): Ignored.- x
For the
print()andplot()methods: an object of classnet_clusteringortidy_covariates.- digits
Integer. Decimal places used for floating-point statistics in the printout. Default
3. Non-breaking: existingprint(x)calls keep their previous formatting.- object
For the
summary()method: an object of classnet_clustering.- type
Character. Plot type:
"silhouette"(per-observation silhouette bars),"mds"(2D MDS projection),"heatmap"(distance matrix heatmap ordered by cluster), or"predictors"(odds-ratio forest plot of the post-hoc covariate analysis; requirescovariatesand an estimator that produces coefficients). Default:"silhouette".- combined
Logical. For
type = "predictors"only: whenTRUE(default), covariate forest panels are combined into a single faceted plot; whenFALSE, a list of separate ggplots is returned.
Value
An object of class "net_clustering" containing:
- data
The original input data.
- k
Number of clusters.
- assignments
Named integer vector of cluster assignments.
- silhouette
Overall average silhouette width.
- sizes
Named integer vector of cluster sizes.
- method
Clustering method used.
- dissimilarity
Distance metric used.
- distance
The computed dissimilarity matrix (
distobject).- medoids
Integer vector of medoid row indices (PAM only; NULL for hierarchical methods).
- seed
Seed used (or NULL).
- weighted
Logical, whether weighted Hamming was used.
- lambda
Lambda value used (0 if not weighted).
- covariates
The post-hoc covariate analysis (a list; see the
estimatorargument), or NULL whencovariates = NULL.- network_method, build_args
For
netobjectinput, the source network's method and stored build arguments, so per-cluster networks can be rebuilt the same way. NULL otherwise.- metadata
For
netobjectinput, its per-sequence metadata, one row per clustered sequence, sosession_idscan name each sequence. NULL otherwise.- htna_partition
For HTNA input, the preserved node-to-actor partition used to restore HTNA children when networks are built.
In print.net_clustering(): The input object, invisibly.
In summary.net_clustering(): A data frame of per-cluster statistics, one row per cluster, with columns cluster, size and mean_within_dist, returned visibly. When the clustering was fitted with covariates, a tidy_covariates/data.frame (the tidied covariate table, with cluster sizes, fit statistics and profiles attached as attributes) is returned invisibly instead. In both cases the printed summary is a side effect.
In plot.net_clustering(): A ggplot object (invisibly); for type = "predictors" with combined = FALSE, a list of ggplot objects named by cluster (invisibly).
In print.tidy_covariates(): The input invisibly.
Methods
print.net_clustering(): Compact, fixed-width summary of a sequence-clustering result. The header carries the clustering method and dissimilarity; the per-cluster table carries cluster size (count and percentage) and mean within-cluster distance when available. Optional medoid and covariate lines surface only when those fields are populated.print.tidy_covariates(): Prints a one-line header naming the estimator, then the data.frame. The full human-readable view (per-cluster stats, profiles, OR/test tables) was already printed bysummary()when this object was produced, so this method intentionally stays minimal to avoid duplication. Auto-prints when the user types the variable at the REPL.
Examples
seqs <- data.frame(V1 = c("A","B","C","A","B"), V2 = c("B","C","A","B","A"),
V3 = c("C","A","B","C","B"))
cl <- build_clusters(seqs, k = 2)
cl
#> Sequence Clustering [pam]
#> Sequences: 5 | Clusters: 2
#> Dissimilarity: hamming
#> Quality: silhouette = 0.600
#>
#> Cluster N Mean within-dist Medoid
#> 1 2 (40.0%) 0.000 4
#> 2 3 (60.0%) 2.000 5
# \donttest{
seqs <- data.frame(
V1 = sample(LETTERS[1:3], 20, TRUE), V2 = sample(LETTERS[1:3], 20, TRUE),
V3 = sample(LETTERS[1:3], 20, TRUE), V4 = sample(LETTERS[1:3], 20, TRUE)
)
cl <- build_clusters(seqs, k = 2)
print(cl)
#> Sequence Clustering [pam]
#> Sequences: 20 | Clusters: 2
#> Dissimilarity: hamming
#> Quality: silhouette = 0.222
#>
#> Cluster N Mean within-dist Medoid
#> 1 9 (45.0%) 2.444 13
#> 2 11 (55.0%) 2.236 14
summary(cl)
#> Sequence Clustering Summary
#> Method: pam
#> Dissimilarity: hamming
#> Silhouette: 0.2222
#>
#> Per-cluster statistics:
#> cluster size mean_within_dist
#> 1 9 2.444444
#> 2 11 2.236364
#> cluster size mean_within_dist
#> 1 1 9 2.444444
#> 2 2 11 2.236364
# }