Skip to contents

Builds a Multi-Cluster Multi-Level (MCML) model from raw transition data (edge lists or sequences) by recoding node labels to cluster labels and counting actual transitions. Unlike cluster_summary which aggregates a pre-computed weight matrix, this function works from the original transition data to produce the TRUE Markov chain over cluster states.

Usage

build_mcml(
  x,
  clusters = NULL,
  method = c("sum", "mean", "median", "max", "min", "density", "geomean"),
  type = c("tna", "frequency", "cooccurrence", "raw"),
  directed = TRUE,
  compute_within = TRUE,
  actor = NULL,
  action = NULL,
  time = NULL,
  order = NULL,
  session = NULL,
  time_threshold = 900,
  exclude = NULL,
  trim = NULL,
  end = FALSE,
  end_by = NULL,
  labels = NULL,
  combine = NULL,
  expand = NULL
)

# S3 method for class 'mcml_layer'
print(x, ...)

# S3 method for class 'mcml'
print(x, ...)

# S3 method for class 'mcml'
summary(object, ...)

Arguments

x

Input data. Accepts multiple formats:

data.frame with from/to columns

Edge list. Columns named from/source/src/v1/node1/i and to/target/tgt/v2/node2/j are auto-detected. Optional weight column (weight/w/value/strength).

data.frame without from/to columns

Sequence data. Each row is a sequence, columns are time steps. Consecutive pairs (t, t+1) become transitions.

tna object

If x$data is non-NULL, uses sequence path on the raw data. Otherwise falls back to cluster_summary.

netobject

If x$data is non-NULL, detects edge list vs sequence data. Otherwise falls back to cluster_summary.

mcml

An mcml object is returned unchanged, unless clusters, combine or expand changes its partition: it is then re-estimated from the sequences it carries (see combine).

square numeric matrix

Falls back to cluster_summary.

non-square or character matrix

Treated as sequence data.

For the print() method: an object of class mcml_layer or mcml.

clusters

Cluster/group assignments. Accepts:

named list

Direct mapping. List names = cluster names, values = character vectors of node labels. Example: list(A = c("N1","N2"), B = c("N3","N4"))

data.frame

A data frame where the first column contains node names and the second column contains group/cluster names. Example: data.frame(node = c("N1","N2","N3"), group = c("A","A","B"))

membership vector

Character or numeric vector. Node names are extracted from the data. Example: c("A","A","B","B")

column name string

For edge list data.frames, the name of a column containing cluster labels. The mapping is built from unique (node, group) pairs in both from and to columns. Limitation: this mode assigns the row's group label to both endpoints, so it only makes sense for edge lists where source and target nodes always share the same group (within-group edges only). For general edge lists where a single node may be source in some rows and target in others, or where source and target belong to different groups, pass an explicit named list (list(G1 = c("N1","N2"), ...)) or a two-column data frame data.frame(node, group) instead.

NULL

Auto-detect from a netobject's node table (a clusters/cluster/groups/group column) or from its node_groups. Only netobject input can auto-detect; data frames and matrices require an explicit assignment.

method

Aggregation method for combining edge weights: "sum", "mean", "median", "max", "min", "density", "geomean". Default "sum". For raw sequence/event-log inputs the function is counting observed transitions, so "sum" is the only interpretation that preserves the count semantics – the other methods are useful when aggregating weighted edge lists or pre-existing weight matrices, where each row already represents a measurement rather than a single observation.

type

Post-processing of the aggregated count matrix. One of:

"tna"

(default) Row-normalize so each row sums to 1 (first-order Markov transition probabilities).

"raw"

Return the un-normalized count matrix.

"frequency"

Explicit alias of "raw" – identical raw count construction (kept as a synonym for callers using frequency-network terminology).

"cooccurrence"

Symmetrize the matrix (undirected co-occurrence).

"semi_markov" is not accepted: the package does not implement a semi-Markov / holding-time construction, so passing it errors rather than silently aliasing "tna".

directed

Logical. If TRUE (default), treat transitions as directed. If FALSE, symmetrize sequence- and edge-derived weights before returning raw/frequency weights or before row-normalizing transition probabilities.

compute_within

Logical. Compute within-cluster matrices? Default TRUE.

actor, action, time, order, session, time_threshold

Long-format event-log shortcut. When action is supplied on a data.frame input, the data is passed through prepare() to derive a wide sequence, which is then routed to the existing sequence path. Behaves identically to prepare(...) |> build_network() |> build_mcml().

exclude

Optional character vector of state labels to drop before the network is built (e.g. a technical-void marker). On long-format input the matching events are removed before sequences are formed, so a dropped state never occupies a sequence position; on wide sequence input the matching cells are set to NA.

trim

Optional truncation of each sequence. NULL (default) keeps every time point. A fraction in (0, 1) keeps the columns covering that quantile of sequence lengths (trim = 0.95 keeps the shortest 95%); a value >= 1 is an absolute cut (trim = 10 keeps the first 10 time points). Same semantics as sequence_plot's trim.

end

Terminal state appended after each sequence's last observed state. FALSE (default) adds none, TRUE adds one labelled "End", and a string supplies the label. Applied after trim, so trimming can never remove the marker.

end_by

Optional column name(s) grouping the terminal marker at a coarser unit than the sequence. NULL (default) marks every sequence. When supplied, only the last sequence of each group is marked – e.g. session = "AttemptID", end = "Conclude", end_by = "SkillID" closes each skill once, not each attempt. Requires long-format input, since the grouping column lives there.

labels

Optional name -> label remap applied to within-cluster nodes (the macro layer is left untouched because its labels are cluster names). Accepts a 2-column data.frame (name, label), a named character vector c(name = "label"), or a named list. Unmapped names pass through unchanged.

combine, expand

Change the partition. On new input they apply to clusters (in any of its accepted forms, including auto-detection) before estimation, so build_mcml(data, clusters = cl, combine = c("A", "B")) equals a build with A and B merged in cl (the merged cluster lists its states in cluster-name order); this works for every input type, matrices included. On an existing mcml they re-partition it (see below). combine merges clusters into one: a character vector merges one group, a list merges several and its names label them (default label "A + B"). expand then splits the named clusters (or "all"/TRUE) into one cluster per member state, named by the state. For an existing mcml the model is re-estimated – macro network, within-cluster networks and stored sequences – from the sequences the mcml carries, with its original type, method and directed unless passed explicitly. Any session split, exclude, trim or end applied when it was built is already in those sequences, so passing any of these (or actor, action, time, session, labels) together with a re-partition is an error. Passing a new clusters list instead re-estimates under that partition; with the partition unchanged the result equals the input. Errors with class nestimate_mcml_no_sequences when the mcml was built from a matrix, from an edge list (only within-cluster edges are kept), or with compute_within = FALSE. sequence_plot accepts the same two arguments as display options: its combine draws exactly what it draws for build_mcml(x, combine = ), while its expand opens clusters in the Summary panel only and keeps one panel per cluster.

...

In print.mcml(), print.mcml_layer() and summary.mcml(): Unsupported. Supplying unused arguments raises an error.

object

For the summary() method: an object of class mcml.

Value

An mcml object with the same layout as the return value of cluster_summary (macro, clusters, cluster_members, edges, meta). On the sequence and edge-list paths meta$source is "transitions", meta$type records the type post-processing, and edges is a tidy data frame with one row per observed node-level transition and columns from, to, weight, cluster_from, cluster_to, type ("within"/"between"). Matrix input falls through to cluster_summary, so meta$source is "matrix" and edges is NULL. Works with print(), summary(), as_tna, as_htna and macro_network.

In print.mcml_layer(): The input mcml_layer, invisibly.

In print.mcml(): The input object, invisibly.

In summary.mcml(): A tidy data frame with one row per cluster and columns cluster, size, within_total, between_out, between_in. For undirected macro networks the in/out split is not meaningful, so between_out reports total incident weight and between_in is NA. The data frame is returned silently without printing the full object – call print(object) explicitly if you want the verbose dump.

Methods

  • print.mcml_layer(): Compact view of one mcml layer (the macro layer or a single within-cluster network): a header line with the node and non-zero edge counts and the weight range, the rounded weight matrix, the initial probabilities as a bar chart, and the dimensions of any attached data – rather than the raw list contents.

See also

cluster_summary for matrix-based aggregation, as_tna to promote the layers to netobjects, macro_network for the cluster-level network with one cluster expanded

Examples

# Edge list with clusters
edges <- data.frame(
  from = c("A", "A", "B", "C", "C", "D"),
  to   = c("B", "C", "A", "D", "D", "A"),
  weight = c(1, 2, 1, 3, 1, 2)
)
clusters <- list(G1 = c("A", "B"), G2 = c("C", "D"))
build_mcml(edges, clusters)
#> MCML Network
#> ============
#> Type: tna  | Method: sum 
#> Nodes: 4  | Clusters: 2 
#> Transitions: 6 
#>   Macro: 2  | Per-cluster: 4 
#> 
#> Clusters:
#>   G1 (2): A, B
#>   G2 (2): C, D
#> 
#> Macro (cluster-level) weights:
#>        G1     G2
#> G1 0.5000 0.5000
#> G2 0.3333 0.6667

# Sequence data with clusters
seqs <- data.frame(
  T1 = c("A", "C", "B"),
  T2 = c("B", "D", "A"),
  T3 = c("C", "C", "D"),
  T4 = c("D", "A", "C")
)
cs <- build_mcml(seqs, clusters, type = "raw")
cs
#> MCML Network
#> ============
#> Type: raw  | Method: sum 
#> Nodes: 4  | Clusters: 2 
#> Transitions: 9 
#>   Macro: 3  | Per-cluster: 6 
#> 
#> Clusters:
#>   G1 (2): A, B
#>   G2 (2): C, D
#> 
#> Macro (cluster-level) weights:
#>    G1 G2
#> G1  2  2
#> G2  1  4
summary(cs)
#>   cluster size within_total between_out between_in
#> 1      G1    2            2           2          1
#> 2      G2    2            4           1          2

# Change the partition while building ...
three <- build_mcml(seqs, list(G1 = "A", G2 = "B", G3 = c("C", "D")))
build_mcml(seqs, list(G1 = "A", G2 = "B", G3 = c("C", "D")),
           combine = c("G1", "G2"))
#> MCML Network
#> ============
#> Type: tna  | Method: sum 
#> Nodes: 4  | Clusters: 2 
#> Transitions: 9 
#>   Macro: 3  | Per-cluster: 6 
#> 
#> Clusters:
#>   G1 + G2 (2): A, B
#>   G3 (2): C, D
#> 
#> Macro (cluster-level) weights:
#>         G1 + G2  G3
#> G1 + G2     0.5 0.5
#> G3          0.2 0.8

# ... or re-partition an existing mcml: the model is re-estimated
build_mcml(three, combine = c("G1", "G2"))           # G1 + G2 as one cluster
#> MCML Network
#> ============
#> Type: tna  | Method: sum 
#> Nodes: 4  | Clusters: 2 
#> Transitions: 9 
#>   Macro: 3  | Per-cluster: 6 
#> 
#> Clusters:
#>   G1 + G2 (2): A, B
#>   G3 (2): C, D
#> 
#> Macro (cluster-level) weights:
#>         G1 + G2  G3
#> G1 + G2     0.5 0.5
#> G3          0.2 0.8
build_mcml(three, combine = list(AB = c("G1", "G2"))) # named merge
#> MCML Network
#> ============
#> Type: tna  | Method: sum 
#> Nodes: 4  | Clusters: 2 
#> Transitions: 9 
#>   Macro: 3  | Per-cluster: 6 
#> 
#> Clusters:
#>   AB (2): A, B
#>   G3 (2): C, D
#> 
#> Macro (cluster-level) weights:
#>     AB  G3
#> AB 0.5 0.5
#> G3 0.2 0.8
build_mcml(three, expand = "G3")                      # C and D as clusters
#> MCML Network
#> ============
#> Type: tna  | Method: sum 
#> Nodes: 4  | Clusters: 4 
#> Transitions: 9 
#>   Macro: 9  | Per-cluster: 0 
#> 
#> Clusters:
#>   C (1): C
#>   D (1): D
#>   G1 (1): A
#>   G2 (1): B
#> 
#> Macro (cluster-level) weights:
#>      C      D     G1  G2
#> C  0.0 0.6667 0.3333 0.0
#> D  1.0 0.0000 0.0000 0.0
#> G1 0.0 0.5000 0.0000 0.5
#> G2 0.5 0.0000 0.5000 0.0