Skip to contents

Discovers latent subgroups with different transition dynamics using Expectation-Maximization. Each mixture component has its own transition matrix. Sequences are probabilistically assigned to components.

Usage

build_mmm(
  data,
  k = 2L,
  n_starts = 50L,
  max_iter = 200L,
  tol = 1e-06,
  smooth = 0.01,
  seed = NULL,
  covariates = NULL,
  covariate_effect = c("em", "posthoc"),
  estimator = c("auto", "firth", "multinom", "chisq")
)

# S3 method for class 'net_mmm'
print(x, digits = 3L, ...)

# S3 method for class 'net_mmm'
summary(object, ...)

# S3 method for class 'net_mmm'
plot(x, type = c("posterior", "covariates"), combined = TRUE, ...)

# S3 method for class 'net_mmm_clustering'
print(x, digits = 3L, ...)

# S3 method for class 'net_mmm_clustering'
plot(
  x,
  type = c("posterior", "covariates", "predictors"),
  combined = TRUE,
  ...
)

Arguments

data

A data.frame (wide format), netobject, cograph_network, or tna model. For tna and cograph_network objects the stored (integer-encoded) data is extracted and decoded to state labels.

k

Integer. Whole finite number of mixture components, >= 2. Default: 2.

n_starts

Integer. Positive whole finite number of random restarts. Default: 50.

max_iter

Integer. Positive whole finite maximum EM iterations per start. Default: 200.

tol

Numeric. Finite positive convergence tolerance. Default: 1e-6.

smooth

Numeric. Finite non-negative Laplace smoothing constant. Default: 0.01.

seed

Integer or NULL. Random seed.

covariates

Optional. Covariates integrated into the EM algorithm to model covariate-dependent mixing proportions. Accepts a string, character vector, formula, or data.frame (same forms as build_clusters). For netobject or cograph_network input, names are resolved against $metadata first, so a typical call is build_mmm(net, k = 3, covariates = "session_label"). Unlike the post-hoc analysis in build_clusters(), these covariates directly influence cluster membership during EM estimation (see covariate_effect).

covariate_effect

How covariates enter the model. "em" (default) folds them into the EM as covariate-dependent mixing proportions, so they shape the cluster fit itself (and rows with missing covariates are dropped before fitting). "posthoc" fits a plain mixture on every sequence and uses the covariates only for the after-fit multinomial logit, so covariate values — and their missingness — never change which clusters are found. Ignored when covariates is NULL.

estimator

Multinomial fitter for the post-hoc covariate analysis (does not affect EM): "auto" (default) inspects the cluster x covariate cross-tab and falls back to "firth" only when any cell has fewer than 5 observations (separation risk), otherwise the much faster "multinom"; "firth" forces Firth's penalised likelihood via brglm2::brmultinom (finite under separation); "multinom" forces nnet::multinom (warns about separation risk); "chisq" runs descriptive tests (no logit). See build_clusters for full details.

x

For the print() and plot() methods: an object of class net_mmm or net_mmm_clustering.

digits

In print.net_mmm(): Integer. Decimal places for floating-point statistics. Default 3. Non-breaking: print(x) keeps the same alignment as before. In print.net_mmm_clustering(): Integer. Decimal places for floating-point statistics. Default 3.

...

In plot.net_mmm(), plot.net_mmm_clustering(), print.net_mmm(), print.net_mmm_clustering() and summary.net_mmm(): Unsupported. Supplying unused arguments raises an error.

object

For the summary() method: an object of class net_mmm.

type

In plot.net_mmm(): Character. Plot type: "posterior" (default) or "covariates". In plot.net_mmm_clustering(): Character. One of "posterior" (default; histogram of max posterior probability per sequence, coloured by cluster), "covariates" or its alias "predictors" (covariate forest plot when cluster_mmm() was run with covariates).

combined

In plot.net_mmm(): Logical. For type = "covariates" only: when TRUE (default), covariate forest panels are combined into a single faceted plot; when FALSE, a list of separate ggplots is returned. In plot.net_mmm_clustering(): Logical. For type in "covariates" or "predictors" only: when TRUE (default), forest panels are combined into a single faceted plot; when FALSE, a list of separate ggplots is returned.

Value

An object of class net_mmm with components:

data

The full N-row sequence frame used for estimation.

models

List of netobjects, one per component. Each component carries the rows assigned to that component in its $data slot, while its transition matrix is the EM-estimated component transition matrix.

k

Number of components.

mixing

Numeric vector of mixing proportions.

posterior

N x k matrix of posterior probabilities.

assignments

Integer vector of hard assignments (1..k).

quality

List: avepp (per-class), avepp_overall, entropy, relative_entropy, classification_error, class_entropy.

log_likelihood, BIC, AIC, ICL

Model fit statistics.

n_params

Number of free parameters behind BIC/AIC/ICL. With covariate_effect = "em" it grows by (k - 1) * p for p covariate columns.

iterations, converged

EM iterations used by the retained fit and whether it met tol.

states

Character vector of state names.

n_sequences

Number of sequences actually fitted (rows with missing covariates are dropped under covariate_effect = "em").

covariates

The post-hoc covariate analysis (list), or NULL when covariates = NULL.

network_method, build_args, htna_partition

Provenance kept from netobject / HTNA input so per-cluster networks can be rebuilt the same way; NULL otherwise.

metadata

The netobject's per-sequence metadata, one row per fitted sequence in the row order of posterior, so session_ids can name each sequence; NULL for other input.

In print.net_mmm() and print.net_mmm_clustering(): The input object, invisibly.

In summary.net_mmm(): A per-component summary data.frame. The class and visibility depend on whether the model was fitted with covariates:

No covariates

A plain data.frame with one row per component and columns component, prior, n_assigned, mean_posterior, avepp, returned visibly (so it auto-prints after the printed summary block).

With covariates

A tidy_covariates/data.frame (the tidied covariate table, with the per-component stats attached), returned invisibly.

In both cases the printed summary (model fit, per-cluster transition matrices, optional covariate profiles) is emitted as a side effect.

In plot.net_mmm(): A ggplot object, invisibly; for type = "covariates" with combined = FALSE, a list of ggplot objects named by cluster (invisibly).

In plot.net_mmm_clustering(): A ggplot object, invisibly; for type = "covariates" / "predictors" with combined = FALSE, a list of ggplot objects named by cluster (invisibly).

Initial states

The first sequence column has special status: it is read directly as the per-sequence initial state (init_state[i] <- match(raw_data[i, state_cols[1L]], states)). The function does not scan forward to the first non-missing position, and it does not apply any na_syms-style symbol conversion (unlike build_clusters). The state vocabulary is built from the unique non-NA values across all columns, so if your data uses a sentinel character such as "*" or "%" for missing cells, that sentinel becomes a real state and the first column reads it as a valid initial state. If you want padded leading missings to be treated as missing, recode them to NA before calling build_mmm() (then match() returns NA, which the EM treats as an uninformative initial distribution), or left-trim the leading missings so each sequence's first column carries an observed state.

Methods

  • plot.net_mmm_clustering(): Plot routines for the MMM clustering metadata attached to the netobject_group that build_network materializes from a cluster_mmm fit (or that cluster_network returns directly with cluster_by = "mmm"). Mirrors the type-driven surface of plot.net_clustering but covers only the metrics the EM fit produces – there is no distance matrix on an MMM clustering, so "silhouette" / "mds" / "heatmap" aren't defined here and the dispatcher raises a clear error if you ask for one of those on an MMM result.

  • print.net_mmm(): Compact summary of a Mixed Markov Model fit. Header carries dimensions and information criteria; cluster table carries N, mixing share, and per-cluster average posterior probability (AvePP). Layout matches print.net_clustering so distance- and model-based clusterings can be compared at a glance.

  • print.net_mmm_clustering(): Prints the clustering metadata attached to the netobject_group that build_network materializes from a cluster_mmm fit (attr(grp, "clustering")). Layout mirrors print.net_clustering: a one-line dimension header, a quality line with AvePP / entropy / classification error, information criteria, and a per-cluster table.

Examples

seqs <- data.frame(V1 = sample(c("A","B","C"), 30, TRUE),
                   V2 = sample(c("A","B","C"), 30, TRUE))
mmm <- build_mmm(seqs, k = 2, n_starts = 1, max_iter = 10, seed = 1)
mmm
#> Mixed Markov Model
#>   Sequences: 30  |  Clusters: 2  |  States: 3
#>   ICs: LL = -62.330  |  BIC = 182.480  |  AIC = 158.659  |  ICL = 184.652
#>   Quality: AvePP = 0.965  |  Entropy = 0.212  |  Class.Err = 0.0%
#> 
#>   Cluster  N           Mix%   AvePP
#>   1        24 (80.0%)  78.1%  0.966
#>   2        6 (20.0%)   21.9%  0.961
# \donttest{
seqs <- data.frame(
  V1 = sample(LETTERS[1:3], 30, TRUE), V2 = sample(LETTERS[1:3], 30, TRUE),
  V3 = sample(LETTERS[1:3], 30, TRUE), V4 = sample(LETTERS[1:3], 30, TRUE)
)
mmm <- build_mmm(seqs, k = 2, seed = 42)
print(mmm)
#> Mixed Markov Model
#>   Sequences: 30  |  Clusters: 2  |  States: 3
#>   ICs: LL = -124.400  |  BIC = 306.621  |  AIC = 282.800  |  ICL = 313.645
#>   Quality: AvePP = 0.904  |  Entropy = 0.304  |  Class.Err = 0.0%
#> 
#>   Cluster  N           Mix%   AvePP
#>   1        20 (66.7%)  62.4%  0.896
#>   2        10 (33.3%)  37.6%  0.919
summary(mmm)
#> Mixed Markov Model
#>   Sequences: 30  |  Clusters: 2  |  States: 3
#>   ICs: LL = -124.400  |  BIC = 306.621  |  AIC = 282.800  |  ICL = 313.645
#>   Quality: AvePP = 0.904  |  Entropy = 0.304  |  Class.Err = 0.0%
#> 
#>   Cluster  N           Mix%   AvePP
#>   1        20 (66.7%)  62.4%  0.896
#>   2        10 (33.3%)  37.6%  0.919
#> 
#> --- Cluster 1 (62.4%, n=20) ---
#>       A     B     C
#> A 0.001 0.568 0.431
#> B 0.485 0.203 0.313
#> C 0.189 0.629 0.181
#> 
#> --- Cluster 2 (37.6%, n=10) ---
#>       A     B     C
#> A 0.593 0.261 0.145
#> B 0.149 0.397 0.454
#> C 0.394 0.002 0.604
#> 
#>   component     prior n_assigned mean_posterior     avepp
#> 1         1 0.6241614         20      0.8959195 0.8959195
#> 2         2 0.3758386         10      0.9193350 0.9193350
# }