Skip to contents

Universal network estimation function that supports both transition networks (relative, frequency, co-occurrence) and association networks (correlation, partial correlation, graphical lasso). Uses the global estimator registry, so custom estimators can also be used.

Usage

build_network(
  data,
  method,
  actor = NULL,
  action = NULL,
  time = NULL,
  session = NULL,
  order = NULL,
  codes = NULL,
  group = NULL,
  format = "auto",
  window_size = 3L,
  mode = c("non-overlapping", "overlapping"),
  scaling = NULL,
  threshold = 0,
  level = NULL,
  time_threshold = 900,
  timezone = "UTC",
  predictability = TRUE,
  state_cols = NULL,
  metadata_cols = NULL,
  start = FALSE,
  end = FALSE,
  params = list(),
  labels = NULL,
  ...
)

# S3 method for class 'netobject'
print(x, ...)

# S3 method for class 'netobject_group'
print(x, digits = 3L, ...)

# S3 method for class 'netobject_ml'
print(x, ...)

# S3 method for class 'netobject'
summary(object, ...)

# S3 method for class 'netobject_group'
summary(object, combined = TRUE, ...)

# S3 method for class 'summary.netobject'
print(x, ...)

# S3 method for class 'summary.netobject_group'
print(x, ...)

Arguments

data

Data frame (sequences or per-observation frequencies) or a square symmetric matrix (correlation or covariance). A fitted net_clustering or net_mmm object is also accepted: the per-cluster networks are (re)built and a netobject_group is returned.

method

Character. Required, except when data is a net_clustering or net_mmm object, where the fitted object's own network method is used. Name of a registered estimator. Built-in methods: "relative", "frequency", "co_occurrence", "cor", "pcor", "glasso", "ising", "mgm", "attention", "wtna", "wtna_cooccurrence", "ngram", "gap", "reverse". The last three mirror tna::build_model() types "n-gram" (adjacent pairs counted once per n-gram window containing them; params = list(n_gram = 2)), "gap" (pairs up to max_gap + 1 positions apart, weighted by 1 / distance; params = list(max_gap = 1)) and "reverse" (reply network: the transpose of "frequency"; params = list(weighted = FALSE)). All three return raw weights; add scaling = "normalize" for row probabilities. Aliases: "tna" and "transition" map to "relative"; "ftna" and "counts" map to "frequency"; "cna" and "wcna" map to "co_occurrence"; "corr" and "correlation" map to "cor"; "partial" maps to "pcor"; "ebicglasso" and "regularized" map to "glasso"; "isingfit" maps to "ising"; "atna" maps to "attention"; "mixed" and "mixed_graphical" map to "mgm"; "wtna_transition" maps to "wtna"; "co-occurrence" maps to "co_occurrence"; "n-gram" and "n_gram" map to "ngram".

actor

Character. Name of the actor/person ID column for sequence grouping. Default: NULL.

action

Character. Name of the action/state column (long format). Default: NULL.

time

Character. Name of the time column (long format). Default: NULL.

session

Character. Name of the session column. Default: NULL.

order

Character. Name of the ordering column. Default: NULL.

codes

Character vector. Column names of one-hot encoded states (for onehot format). Default: NULL.

group

Character. Name of a grouping column for per-group networks. Returns a netobject_group (named list of netobjects). Default: NULL.

format

Character. Input format: "auto", "wide", "long", or "onehot". Default: "auto".

window_size

Integer. Window size for one-hot windowing. Default: 3L.

mode

Character. Windowing mode for one-hot input only: "non-overlapping" or "overlapping". Has no effect on wide or long sequence data (only the one-hot/wtna path reads it). Default: "non-overlapping".

scaling

Character vector or NULL. Post-estimation scaling to apply (in order). Options: "minmax", "max", "rank", "normalize". Can combine: c("rank", "minmax"). Default: NULL (no scaling).

threshold

Numeric. Absolute values below this are set to zero in the result matrix. Default: 0 (no thresholding).

level

Character or NULL. Multilevel decomposition for the undirected association methods (cor, pcor, glasso); a directed estimator errors. One of NULL, "between", "within", "both". Requires an id column, supplied either as actor or as params$id / params$id_col. Default: NULL.

time_threshold

Numeric or FALSE. Maximum time gap (seconds) for long format session splitting. Set to FALSE to switch session-interval splitting off, so each actor (or actor-session) forms a single sequence. Default: 900.

timezone

Character. Olson time zone used to interpret naive timestamps in long-format data (offset-bearing timestamps such as ...Z or +02:00 are converted from their offset). Passed to prepare. Default: "UTC".

predictability

Logical. If TRUE (default), compute and store node predictability (R-squared) for undirected association methods (glasso, pcor, cor). Stored in $predictability and auto-displayed as donuts by cograph::splot().

state_cols

Character vector or NULL. Explicit names of columns to classify as state columns in the returned netobject's $data slot. When provided, all other columns of the cleaned input go to $metadata. Auto-detection (values-in-nodes heuristic) is bypassed. Use this when a metadata column happens to contain values that overlap with node names (e.g. condition labels "A","B","C" and nodes "A","B","C") and auto-detection would misclassify it. Default: NULL (auto-detect).

metadata_cols

Character vector or NULL. Explicit names of columns to force into the $metadata slot. The remaining columns are auto-detected as state via the values-in-nodes rule. Cannot overlap with state_cols. Default: NULL.

start

Boundary marker prepended to every sequence as an explicit start state (a pure source: no incoming edges, every sequence's first transition is start -> first_observed). FALSE (default) adds nothing; TRUE uses the label "Start"; a single string uses that string as the label. Only valid for the transition methods (relative, frequency, co_occurrence, attention, ngram, gap, reverse); errors otherwise (wtna included).

end

Boundary marker placed in the single cell after each sequence's last observed (non-NA) state, as an explicit terminal state (a pure sink: no outgoing edges, no self-loop – distinct from mark_terminal_state, which fills all trailing NAs into an absorbing state). FALSE (default) adds nothing; TRUE uses the label "End"; a single string uses that string as the label. Same method restriction as start.

params

Named list. Method-specific parameters passed to the estimator function (e.g. list(gamma = 0.5) for glasso, or list(format = "wide") for transition methods). This is the key composability feature: downstream functions like bootstrap or grid search can store and replay the full params list without knowing method internals. Transition estimators accept tna-style sequence options such as weighted and concat (and the low-level begin_state / end_state, of which start / end are the public form – see those arguments). Column-like entries in params (action, id, id_col, actor, time, session, order, cols, codes, and group) are resolved before format detection and must name existing columns. If the same column role is supplied both directly and through params, the names must agree.

labels

Optional name -> label remap applied after construction. Accepts a 2-column data.frame (name, label), a named character vector c(name = "label"), or a named list. Rewrites $nodes$label and dimnames(weights). Unmapped names pass through unchanged.

...

Additional arguments passed to the estimator function. In print.netobject(), print.netobject_group() and print.netobject_ml(): Additional arguments (ignored). In print.summary.netobject(), print.summary.netobject_group(), summary.netobject() and summary.netobject_group(): Ignored.

x

For the print() method: an object of class netobject, netobject_group or netobject_ml (or its summary()).

digits

Integer. Decimal places for the weight summary. Default 3. Non-breaking: print(x) keeps the same shape as before, with the addition of a weight-range column.

object

For the summary() method: an object of class netobject or netobject_group.

combined

Logical. Combine into one wide data.frame? Default TRUE.

Value

An object of class c("netobject", "cograph_network") containing:

data

The state columns of the cleaned input data, as a data frame.

metadata

Data frame of the non-state columns of the cleaned input (and, for long input, the per-sequence metadata), or NULL.

weights

The estimated network weight matrix.

nodes

Data frame with columns id, label, name, x, y. Node labels are in $nodes$label.

edges

Data frame of non-zero edges with integer from/to (node IDs) and numeric weight.

directed

Logical. Whether the network is directed.

method

The resolved method name.

params

The params list used (for reproducibility).

scaling

The scaling applied (or NULL).

threshold

The threshold applied.

n_nodes

Number of nodes.

n_edges

Number of non-zero edges.

level

Decomposition level used (or NULL).

build_args

The resolved column/format arguments (actor, action, time, session, order, codes, format, window_size, mode) used for this build.

meta

List with source, layout, and tna metadata (cograph-compatible).

node_groups

Node groupings data frame, or NULL.

predictability

Named numeric vector of R-squared predictability values per node (for undirected association methods when predictability = TRUE). NULL for directed methods.

Method-specific extras (e.g. precision_matrix, cor_matrix, frequency_matrix, initial, lambda_selected, etc.) are preserved from the estimator output.

When level = "both", returns an object of class "netobject_ml" with $between and $within sub-networks and a $method field. level = "between" or "within" returns a single netobject estimated on the decomposed data.

When group is supplied (or data is a net_clustering / net_mmm object), returns an object of class "netobject_group": a named list of netobjects, one per group, carrying the grouping column in attr(x, "group_col").

In print.netobject(), print.netobject_group() and print.netobject_ml(): The input object, invisibly.

In summary.netobject(): A data.frame with columns metric and value, of class c("summary.netobject", "data.frame").

In summary.netobject_group(): Either a data.frame (one column per group) or a named list of summary.netobject objects, of class c("summary.netobject_group", ...).

In print.summary.netobject(): x, invisibly.

In print.summary.netobject_group(): x, invisibly.

Details

The function works as follows:

  1. Resolves method aliases to canonical names.

  2. Validates explicit column arguments before any format guessing.

  3. Retrieves the estimator function from the global registry.

  4. For association methods with level specified, decomposes the data (between-person means or within-person centering).

  5. Calls the estimator: do.call(fn, c(list(data = data), params)).

  6. Applies scaling and thresholding to the result matrix.

  7. Extracts edges and constructs the netobject.

For long-format transition data, supplying action without actor is allowed and treats all rows as one sequence in row/time order. The function warns because a one-sequence transition network is not recommended and cannot be validated by bootstrap or other confirmatory tests.

Methods

  • print.netobject_group(): Compact summary of a netobject_group. Header surfaces the source (a clustering attached by cluster_network or cluster_mmm, or a plain split by group_col). The per-group table carries node and edge counts, weight range, and – when a clustering attribute is present – N and percentage of sequences per cluster (matching the layout used by print.net_clustering and print.net_mmm).

  • summary.netobject(): Computes node count, edge count, density, mean shortest-path distance, mean and SD of in/out strength, mean and SD of in/out degree, in/out degree centralization (Freeman), and reciprocity. Mirrors the metric set returned by tna::summary.tna() so a Nestimate netobject and the equivalent tna model report numerically identical descriptive metrics.

  • summary.netobject_group(): Returns one summary per constituent network. With combined = TRUE (default) the per-group tables are joined into a single wide data.frame with one column per group; with combined = FALSE returns a named list.

Examples

seqs <- data.frame(V1 = c("A","B","C","A"), V2 = c("B","C","A","B"))
net <- build_network(seqs, method = "relative")
net
#> Transition Network (relative probabilities) [directed]
#>   Weights: [1.000, 1.000]  |  mean: 1.000
#> 
#>   Weight matrix:
#>     A B C
#>   A 0 1 0
#>   B 0 0 1
#>   C 1 0 0 
#> 
#>   Initial probabilities:
#>   A             0.500  ████████████████████████████████████████
#>   B             0.250  ████████████████████
#>   C             0.250  ████████████████████
# \donttest{
# Transition network (relative probabilities)
seqs <- data.frame(
  V1 = sample(LETTERS[1:4], 30, TRUE), V2 = sample(LETTERS[1:4], 30, TRUE),
  V3 = sample(LETTERS[1:4], 30, TRUE), V4 = sample(LETTERS[1:4], 30, TRUE)
)
net <- build_network(seqs, method = "relative")
print(net)
#> Transition Network (relative probabilities) [directed]
#>   Weights: [0.111, 0.556]  |  mean: 0.250
#> 
#>   Weight matrix:
#>         A     B     C     D
#>   A 0.423 0.192 0.192 0.192
#>   B 0.304 0.304 0.217 0.174
#>   C 0.217 0.304 0.304 0.174
#>   D 0.222 0.111 0.556 0.111 
#> 
#>   Initial probabilities:
#>   A             0.333  ████████████████████████████████████████
#>   B             0.300  ████████████████████████████████████
#>   D             0.200  ████████████████████████
#>   C             0.167  ████████████████████

# Association network (glasso)
freq_data <- convert_sequence_format(seqs, format = "frequency")
net_glasso <- build_network(freq_data, method = "glasso",
                             params = list(gamma = 0.5, nlambda = 50))

# With scaling
net_scaled <- build_network(seqs, method = "relative",
                             scaling = c("rank", "minmax"))
# }