Transductive label spreading on a hypergraph
Source:R/hypergraph_laplacian.R
hypergraph_transduction.RdSemi-supervised classification of hypergraph nodes by the regularization
framework of Zhou et al. (2006): given labels for a subset of nodes, the
scores F = (1 - xi) * (I - xi * S)^{-1} Y spread the labels over the
hypergraph, where S = I - L is the normalized similarity operator of
the chosen Laplacian and Y is the label indicator matrix. Each node is
assigned the class with the highest score. This is the non-neural
ancestor of hypergraph-attention text classifiers: with documents as
hyperedges over words (or vice versa) it classifies unlabeled nodes from
a handful of labeled ones.
Usage
hypergraph_transduction(
hg,
labels,
xi = 0.99,
type = c("zhou", "random_walk"),
edge_weights = NULL
)
# S3 method for class 'net_hypergraph_transduction'
print(x, ...)
# S3 method for class 'net_hypergraph_transduction'
summary(object, ...)
# S3 method for class 'net_hypergraph_transduction'
as.data.frame(
x,
row.names = NULL,
optional = FALSE,
what = c("predictions", "scores"),
...
)
# S3 method for class 'net_hypergraph_transduction'
plot(x, ...)Arguments
- hg
A connected
net_hypergraph.- labels
Node labels. Either a named vector (names = node names, values = class labels) covering a subset of nodes, or a full-length vector aligned with
hg$nodeswithNAfor unlabeled nodes. At least two distinct classes must be labeled.- xi
Numeric in
(0, 1). Spreading coefficient (default0.99); larger values weight the hypergraph structure more relative to the initial labels.- type, edge_weights
Passed to
hypergraph_laplacian().- x
For the
print(),as.data.frame()andplot()methods: an object of classnet_hypergraph_transduction.- ...
In
as.data.frame.net_hypergraph_transduction(),plot.net_hypergraph_transduction(),print.net_hypergraph_transduction()andsummary.net_hypergraph_transduction(): Additional arguments (ignored).- object
For the
summary()method: an object of classnet_hypergraph_transduction.- row.names
NULL(default) or a character vector of row names for the returned data frame.- optional
Ignored; present so the method matches the signature of the
as.data.frame()generic.- what
Character.
"predictions"(default) for the one-row-per-node table,"scores"for the tidy long score table (one row per node x class:node,class,score).
Value
An object of class net_hypergraph_transduction: a list with
$predictions (data.frame, one row per node: node, label (given,
NA if unlabeled), predicted, score (winning class score),
margin (winning minus runner-up score)), $classes, $scores
(node x class score matrix), $xi, $type, $n_labeled, $n_nodes
and $params (the edge_weights used).
Has print, summary, plot and as.data.frame methods;
as.data.frame(x, what = "scores") returns the tidy long score table.
In print.net_hypergraph_transduction(): The input object, invisibly.
In summary.net_hypergraph_transduction(): A data.frame, one row per class: class, n_labeled, n_predicted, mean_margin (mean winning margin among the nodes predicted into the class).
In as.data.frame.net_hypergraph_transduction(): A data.frame selected by what: for "predictions", one row per node with columns node, label (the given label, NA if unlabeled), predicted, score and margin; for "scores", one row per node x class with columns node, class and score.
In plot.net_hypergraph_transduction(): A ggplot object (the score heatmap), returned visibly so that plot(x) draws it.
Methods
plot.net_hypergraph_transduction(): Heatmap of the full node-by-class score matrix: rows are nodes (grouped by predicted class), columns are classes, tile shading and printed values are the spreading scores. Seed nodes (given labels) carry a black tile border, and each node's winning class is marked with a dot, so agreement between seeds, scores, and decisions is visible in one panel. Rows whose winning and runner-up scores are close (smallmargin) are the assignments to distrust.
References
Zhou, D., Huang, J., & Scholkopf, B. (2006). Learning with hypergraphs: Clustering, classification, and embedding. NeurIPS 19.
Examples
events <- data.frame(
person = c("a", "b", "c", "a", "b", "c", "d", "e", "f",
"d", "e", "f", "c", "d"),
meeting = c("m1", "m1", "m1", "m2", "m2", "m2", "m3", "m3", "m3",
"m4", "m4", "m4", "m5", "m5")
)
hg <- bipartite_groups(events, player = "person", group = "meeting")
tr <- hypergraph_transduction(hg, labels = c(a = "x", d = "y"))
tr
#> Hypergraph transductive label spreading (zhou Laplacian, xi = 0.99)
#> Nodes: 6 (2 labeled) | Classes: x, y
#> Predicted: x = 1, y = 5
as.data.frame(tr)
#> node label predicted score margin
#> 1 a x x 0.1654808 0.003987167
#> 2 b <NA> y 0.1614936 0.006012833
#> 3 c <NA> y 0.2037821 0.019834793
#> 4 d y y 0.2321154 0.070621782
#> 5 e <NA> y 0.1839473 0.055966489
#> 6 f <NA> y 0.1839473 0.055966489