Universal network estimation function that supports both transition networks (relative, frequency, co-occurrence) and association networks (correlation, partial correlation, graphical lasso). Uses the global estimator registry, so custom estimators can also be used.
Usage
build_network(
data,
method,
actor = NULL,
action = NULL,
time = NULL,
session = NULL,
order = NULL,
codes = NULL,
group = NULL,
format = "auto",
window_size = 3L,
mode = c("non-overlapping", "overlapping"),
scaling = NULL,
threshold = 0,
level = NULL,
time_threshold = 900,
timezone = "UTC",
predictability = TRUE,
state_cols = NULL,
metadata_cols = NULL,
start = FALSE,
end = FALSE,
params = list(),
labels = NULL,
...
)
# S3 method for class 'netobject'
print(x, ...)
# S3 method for class 'netobject_group'
print(x, digits = 3L, ...)
# S3 method for class 'netobject_ml'
print(x, ...)
# S3 method for class 'netobject'
summary(object, ...)
# S3 method for class 'netobject_group'
summary(object, combined = TRUE, ...)
# S3 method for class 'summary.netobject'
print(x, ...)
# S3 method for class 'summary.netobject_group'
print(x, ...)Arguments
- data
Data frame (sequences or per-observation frequencies) or a square symmetric matrix (correlation or covariance). A fitted
net_clusteringornet_mmmobject is also accepted: the per-cluster networks are (re)built and anetobject_groupis returned.- method
Character. Required, except when
datais anet_clusteringornet_mmmobject, where the fitted object's own network method is used. Name of a registered estimator. Built-in methods:"relative","frequency","co_occurrence","cor","pcor","glasso","ising","mgm","attention","wtna","wtna_cooccurrence","ngram","gap","reverse". The last three mirrortna::build_model()types"n-gram"(adjacent pairs counted once per n-gram window containing them;params = list(n_gram = 2)),"gap"(pairs up tomax_gap + 1positions apart, weighted by1 / distance;params = list(max_gap = 1)) and"reverse"(reply network: the transpose of"frequency";params = list(weighted = FALSE)). All three return raw weights; addscaling = "normalize"for row probabilities. Aliases:"tna"and"transition"map to"relative";"ftna"and"counts"map to"frequency";"cna"and"wcna"map to"co_occurrence";"corr"and"correlation"map to"cor";"partial"maps to"pcor";"ebicglasso"and"regularized"map to"glasso";"isingfit"maps to"ising";"atna"maps to"attention";"mixed"and"mixed_graphical"map to"mgm";"wtna_transition"maps to"wtna";"co-occurrence"maps to"co_occurrence";"n-gram"and"n_gram"map to"ngram".- actor
Character. Name of the actor/person ID column for sequence grouping. Default:
NULL.- action
Character. Name of the action/state column (long format). Default:
NULL.- time
Character. Name of the time column (long format). Default:
NULL.- session
Character. Name of the session column. Default:
NULL.- order
Character. Name of the ordering column. Default:
NULL.- codes
Character vector. Column names of one-hot encoded states (for onehot format). Default:
NULL.- group
Character. Name of a grouping column for per-group networks. Returns a
netobject_group(named list of netobjects). Default:NULL.- format
Character. Input format:
"auto","wide","long", or"onehot". Default:"auto".- window_size
Integer. Window size for one-hot windowing. Default:
3L.- mode
Character. Windowing mode for one-hot input only:
"non-overlapping"or"overlapping". Has no effect on wide or long sequence data (only the one-hot/wtna path reads it). Default:"non-overlapping".- scaling
Character vector or NULL. Post-estimation scaling to apply (in order). Options:
"minmax","max","rank","normalize". Can combine:c("rank", "minmax"). Default:NULL(no scaling).- threshold
Numeric. Absolute values below this are set to zero in the result matrix. Default: 0 (no thresholding).
- level
Character or NULL. Multilevel decomposition for the undirected association methods (
cor,pcor,glasso); a directed estimator errors. One ofNULL,"between","within","both". Requires an id column, supplied either asactoror asparams$id/params$id_col. Default:NULL.- time_threshold
Numeric or FALSE. Maximum time gap (seconds) for long format session splitting. Set to
FALSEto switch session-interval splitting off, so each actor (or actor-session) forms a single sequence. Default:900.- timezone
Character. Olson time zone used to interpret naive timestamps in long-format data (offset-bearing timestamps such as
...Zor+02:00are converted from their offset). Passed toprepare. Default:"UTC".- predictability
Logical. If
TRUE(default), compute and store node predictability (R-squared) for undirected association methods (glasso, pcor, cor). Stored in$predictabilityand auto-displayed as donuts bycograph::splot().- state_cols
Character vector or
NULL. Explicit names of columns to classify as state columns in the returned netobject's$dataslot. When provided, all other columns of the cleaned input go to$metadata. Auto-detection (values-in-nodes heuristic) is bypassed. Use this when a metadata column happens to contain values that overlap with node names (e.g. condition labels"A","B","C"and nodes"A","B","C") and auto-detection would misclassify it. Default:NULL(auto-detect).- metadata_cols
Character vector or
NULL. Explicit names of columns to force into the$metadataslot. The remaining columns are auto-detected as state via the values-in-nodes rule. Cannot overlap withstate_cols. Default:NULL.- start
Boundary marker prepended to every sequence as an explicit start state (a pure source: no incoming edges, every sequence's first transition is
start -> first_observed).FALSE(default) adds nothing;TRUEuses the label"Start"; a single string uses that string as the label. Only valid for the transition methods (relative,frequency,co_occurrence,attention,ngram,gap,reverse); errors otherwise (wtnaincluded).- end
Boundary marker placed in the single cell after each sequence's last observed (non-
NA) state, as an explicit terminal state (a pure sink: no outgoing edges, no self-loop – distinct frommark_terminal_state, which fills all trailing NAs into an absorbing state).FALSE(default) adds nothing;TRUEuses the label"End"; a single string uses that string as the label. Same method restriction asstart.- params
Named list. Method-specific parameters passed to the estimator function (e.g.
list(gamma = 0.5)for glasso, orlist(format = "wide")for transition methods). This is the key composability feature: downstream functions like bootstrap or grid search can store and replay the full params list without knowing method internals. Transition estimators accept tna-style sequence options such asweightedandconcat(and the low-levelbegin_state/end_state, of whichstart/endare the public form – see those arguments). Column-like entries inparams(action,id,id_col,actor,time,session,order,cols,codes, andgroup) are resolved before format detection and must name existing columns. If the same column role is supplied both directly and throughparams, the names must agree.- labels
Optional name -> label remap applied after construction. Accepts a 2-column data.frame
(name, label), a named character vectorc(name = "label"), or a named list. Rewrites$nodes$labelanddimnames(weights). Unmapped names pass through unchanged.- ...
Additional arguments passed to the estimator function. In
print.netobject(),print.netobject_group()andprint.netobject_ml(): Additional arguments (ignored). Inprint.summary.netobject(),print.summary.netobject_group(),summary.netobject()andsummary.netobject_group(): Ignored.- x
For the
print()method: an object of classnetobject,netobject_groupornetobject_ml(or itssummary()).- digits
Integer. Decimal places for the weight summary. Default
3. Non-breaking:print(x)keeps the same shape as before, with the addition of a weight-range column.- object
For the
summary()method: an object of classnetobjectornetobject_group.- combined
Logical. Combine into one wide data.frame? Default
TRUE.
Value
An object of class c("netobject", "cograph_network") containing:
- data
The state columns of the cleaned input data, as a data frame.
- metadata
Data frame of the non-state columns of the cleaned input (and, for long input, the per-sequence metadata), or NULL.
- weights
The estimated network weight matrix.
- nodes
Data frame with columns
id,label,name,x,y. Node labels are in$nodes$label.- edges
Data frame of non-zero edges with integer
from/to(node IDs) and numericweight.- directed
Logical. Whether the network is directed.
- method
The resolved method name.
- params
The params list used (for reproducibility).
- scaling
The scaling applied (or NULL).
- threshold
The threshold applied.
- n_nodes
Number of nodes.
- n_edges
Number of non-zero edges.
- level
Decomposition level used (or NULL).
- build_args
The resolved column/format arguments (
actor,action,time,session,order,codes,format,window_size,mode) used for this build.- meta
List with
source,layout, andtnametadata (cograph-compatible).- node_groups
Node groupings data frame, or NULL.
- predictability
Named numeric vector of R-squared predictability values per node (for undirected association methods when
predictability = TRUE). NULL for directed methods.
Method-specific extras (e.g. precision_matrix, cor_matrix,
frequency_matrix, initial, lambda_selected, etc.) are
preserved from the estimator output.
When level = "both", returns an object of class
"netobject_ml" with $between and $within
sub-networks and a $method field. level = "between" or
"within" returns a single netobject estimated on the
decomposed data.
When group is supplied (or data is a net_clustering /
net_mmm object), returns an object of class
"netobject_group": a named list of netobjects, one per group,
carrying the grouping column in attr(x, "group_col").
In print.netobject(), print.netobject_group() and print.netobject_ml(): The input object, invisibly.
In summary.netobject(): A data.frame with columns metric and value, of class c("summary.netobject", "data.frame").
In summary.netobject_group(): Either a data.frame (one column per group) or a named list of summary.netobject objects, of class c("summary.netobject_group", ...).
In print.summary.netobject(): x, invisibly.
In print.summary.netobject_group(): x, invisibly.
Details
The function works as follows:
Resolves method aliases to canonical names.
Validates explicit column arguments before any format guessing.
Retrieves the estimator function from the global registry.
For association methods with
levelspecified, decomposes the data (between-person means or within-person centering).Calls the estimator:
do.call(fn, c(list(data = data), params)).Applies scaling and thresholding to the result matrix.
Extracts edges and constructs the
netobject.
For long-format transition data, supplying action without
actor is allowed and treats all rows as one sequence in row/time
order. The function warns because a one-sequence transition network is not
recommended and cannot be validated by bootstrap or other confirmatory
tests.
Methods
print.netobject_group(): Compact summary of anetobject_group. Header surfaces the source (a clustering attached bycluster_networkorcluster_mmm, or a plain split bygroup_col). The per-group table carries node and edge counts, weight range, and – when a clustering attribute is present – N and percentage of sequences per cluster (matching the layout used byprint.net_clusteringandprint.net_mmm).summary.netobject(): Computes node count, edge count, density, mean shortest-path distance, mean and SD of in/out strength, mean and SD of in/out degree, in/out degree centralization (Freeman), and reciprocity. Mirrors the metric set returned bytna::summary.tna()so a Nestimate netobject and the equivalent tna model report numerically identical descriptive metrics.summary.netobject_group(): Returns one summary per constituent network. Withcombined = TRUE(default) the per-group tables are joined into a single widedata.framewith one column per group; withcombined = FALSEreturns a named list.
Examples
seqs <- data.frame(V1 = c("A","B","C","A"), V2 = c("B","C","A","B"))
net <- build_network(seqs, method = "relative")
net
#> Transition Network (relative probabilities) [directed]
#> Weights: [1.000, 1.000] | mean: 1.000
#>
#> Weight matrix:
#> A B C
#> A 0 1 0
#> B 0 0 1
#> C 1 0 0
#>
#> Initial probabilities:
#> A 0.500 ████████████████████████████████████████
#> B 0.250 ████████████████████
#> C 0.250 ████████████████████
# \donttest{
# Transition network (relative probabilities)
seqs <- data.frame(
V1 = sample(LETTERS[1:4], 30, TRUE), V2 = sample(LETTERS[1:4], 30, TRUE),
V3 = sample(LETTERS[1:4], 30, TRUE), V4 = sample(LETTERS[1:4], 30, TRUE)
)
net <- build_network(seqs, method = "relative")
print(net)
#> Transition Network (relative probabilities) [directed]
#> Weights: [0.111, 0.556] | mean: 0.250
#>
#> Weight matrix:
#> A B C D
#> A 0.423 0.192 0.192 0.192
#> B 0.304 0.304 0.217 0.174
#> C 0.217 0.304 0.304 0.174
#> D 0.222 0.111 0.556 0.111
#>
#> Initial probabilities:
#> A 0.333 ████████████████████████████████████████
#> B 0.300 ████████████████████████████████████
#> D 0.200 ████████████████████████
#> C 0.167 ████████████████████
# Association network (glasso)
freq_data <- convert_sequence_format(seqs, format = "frequency")
net_glasso <- build_network(freq_data, method = "glasso",
params = list(gamma = 0.5, nlambda = 50))
# With scaling
net_scaled <- build_network(seqs, method = "relative",
scaling = c("rank", "minmax"))
# }