Guide to Idiographic Statistical Workflows
Source:vignettes/statistical-workflows.Rmd
statistical-workflows.RmdThis document indexes the non-network statistical workflows in
idiographic. The package contains 16 major functions in
this part of its interface. They address different estimands: data
quality, within-person association, prediction, coefficient synthesis,
subgroup structure, and treatment-effect heterogeneity. Choosing a
function begins with the research question, not with the preferred
algorithm.
Start with the analytical question
| Question | Primary function | Detailed vignette |
|---|---|---|
| Is each person’s series usable and sufficiently variable? | describe_persons() |
Person-level description |
| How are variables associated within each person? | correlate_persons() |
Within-person correlation |
| How much variation lies within versus between people? | variance_components() |
Variance decomposition |
| How should a panel be ordered, lagged, centred, and screened? | preprocess_panel() |
Panel preparation |
| What are the pooled and person-specific linear coefficients? | fit_lm() |
Linear models |
| What are the pooled and person-specific binary or count associations? | fit_glm() |
Generalized linear models |
| Which algorithm predicts later observations most accurately? | fit_ml() |
Machine learning |
| Does predictive performance persist across forecast origins? | fit_rolling() |
Rolling-origin validation |
| Do within-person and between-person effects differ? | fit_within_between() |
Within-between models |
| What is the uncertainty-weighted average of person-specific coefficients? | pool_coefs() |
Coefficient pooling |
| How should noisy person-specific coefficients be stabilised? | shrink_coefs() |
Coefficient shrinkage |
| Is a discrete subgroup representation supported at all? | test_subgroups() |
Testing subgroup existence |
| Which people share a reproducible coefficient profile? | find_subgroups() |
Subgroup discovery |
| Does subgroup-specific modelling improve held-out performance? | fit_subgroups() |
Subgroup-specific models |
| What is the treatment effect, and who benefits more? | fit_effects() |
Treatment effects |
| Is a person-varying target predictably heterogeneous? | fit_heterogeneity() |
Repeated-split heterogeneity |
Recommended sequence
Most analyses should not begin with model fitting. The first four functions establish whether the data support the intended estimand.
- Use
describe_persons()to inspect usable observations, gaps, dispersion, successive change, autocorrelation, floor or ceiling concentration, and trends by person. - Use
variance_components()to determine whether the target variation is primarily within people or between people. This decision changes the meaning of a pooled coefficient. - Use
correlate_persons()for an initial description of contemporaneous within-person association. Do not interpret these correlations as lagged or causal effects. - Use
preprocess_panel()to make chronology, lags, centring, scaling, and missing-data consequences explicit before fitting a model.
The next function depends on the estimand. For a prespecified
coefficient, use fit_lm(), fit_glm(), or
fit_within_between(). For future predictive performance,
use fit_ml() and confirm temporal stability with
fit_rolling(). For an intervention contrast, use
fit_effects(); prediction accuracy alone cannot identify a
treatment effect.
Coefficients: estimate, pool, or shrink
fit_lm() and fit_glm() can estimate one
model for the pooled panel, one per subgroup, one per person, or several
scopes together. The resulting person-specific coefficients contain both
genuine heterogeneity and sampling error. Two downstream functions
answer different questions.
-
pool_coefs()estimates the uncertainty-weighted average coefficient and quantifies residual between-person heterogeneity. Use it for population synthesis while retainingtauand I-squared. -
shrink_coefs()estimates a more stable coefficient for each person by moving imprecise values towards the pooled distribution. Use it for ranking, description, or downstream prediction when raw individual slopes are noisy.
Pooling does not assert that every person has the pooled effect. Shrinkage does not erase heterogeneity. Both procedures require comparable coefficients from models with the same outcome, predictors, coding, and scale.
Prediction: algorithm, scope, and time
fit_ml() separates three decisions that are often
conflated.
- Algorithm: linear, regularised, neighbour-based, tree-based, or an optional backend.
- Scope: pooled, subgroup, or individual.
- Validation design: which later observations form the validation and test blocks.
The machine-learning vignette compares
models with a mean baseline, audits hyperparameter selection, compares
pooled and individual scopes, identifies people with poor predictions,
inspects their trajectories, and computes model-agnostic permutation
importance. fit_rolling() extends the same logic across
several temporal origins. A model should not be called personally useful
from an aggregate metric alone; report the distribution of person-level
errors.
Subgroups: evidence before assignment
Subgroup work has three distinct stages.
-
test_subgroups()compares a one-population coefficient distribution with candidate finite mixtures. A one-group result is a valid conclusion. -
find_subgroups()assigns people only after a multi-group representation is defensible and reports resampling stability for every assignment. -
fit_subgroups()tests whether the resulting partition improves held-out modelling relative to pooled and individual alternatives.
A clustering algorithm always partitions the data when asked. This does not prove that latent classes exist. Skewness, heavy tails, and continuous heterogeneity can resemble discrete groups, so the BIC comparison, shape diagnostics, stability, and out-of-sample utility should be considered together.
Effects and heterogeneity
fit_effects() estimates an average treatment effect for
a binary or multi-arm treatment, or an average partial effect for a
continuous dose. Its sorted effect groups and heterogeneity slope ask
whether the effect differs systematically. These quantities require
intervention-specific identification assumptions, including treatment
variation, overlap, consistency, and adequate adjustment for
confounding.
fit_heterogeneity() generalises repeated-split
heterogeneity analysis to a treatment effect, prediction error, or gain
from individualisation. It learns a proxy in one sample and estimates
grouped and linear heterogeneity in another. Repeated splitting reduces
dependence on one random partition but does not create additional
independent people.
Reporting standard
A complete idiographic analysis should report the person identifier and time ordering, observation counts and gaps by person, the exact outcome and predictors, the model scope, preprocessing learned from training data, the validation design, failed units, aggregate and person-level results, and the assumptions that determine interpretation. Figures should answer a named question: show uncertainty for coefficients and effects, show a reference model for prediction, show chronology for forecasts, and show assignment stability for subgroups.
The 16 linked vignettes provide executable examples using the bundled
srl data. Each example prints the public result object
before interpreting it, uses public accessors for detailed tables, and
states when a neighbouring function answers a more appropriate
question.