Package {semrulesid}


Type: Package
Title: Evaluate Structural Equation Model Identification Rules
Version: 0.4.1
Description: Evaluates selected necessary and sufficient identification conditions in structural equation models (SEMs), including latent-variable scaling constraints. Output reports rule status and applicability and provides diagnostic messages to support model specification and respecification. The package is intended as a diagnostic aid and does not implement a universal identification algorithm. For more details, see Bollen (2026, ISBN:978-1009312820).
License: GPL (≥ 3)
Encoding: UTF-8
Imports: lavaan
Suggests: knitr, rmarkdown, testthat (≥ 3.0.0)
Config/testthat/edition: 3
VignetteBuilder: knitr
Language: en-US
Config/roxygen2/version: 8.1.0
URL: https://github.com/zacharyvig/semrulesid
BugReports: https://github.com/zacharyvig/semrulesid/issues
NeedsCompilation: no
Packaged: 2026-09-15 13:13:05 UTC; ZACHMAC
Author: Zach Vig [aut, cre, cph]
Maintainer: Zach Vig <zachvig@rocketmail.com>
Repository: CRAN
Date/Publication: 2026-09-26 16:20:07 UTC

semrulesid package outline

Description

Evaluates selected necessary and sufficient identification conditions in structural equation models (SEMs), including latent-variable scaling constraints. Output reports rule status and applicability and provides diagnostic messages to support model specification and respecification. The package is intended as a diagnostic aid and does not implement a universal identification algorithm. For more details, see Bollen (2026, ISBN:978-1009312820).

Purpose

'semrulesid' allows the user to input a Structural Equation Model (SEM) in ['lavaan'](https://lavaan.ugent.be/) (Rosseel, 2012) syntax and check it against a number of identification rules from the literature. Rules are specified as being necessary and/or sufficient and specific reasons are given when a rule is broken. Users should not rely solely on the output of the package as the sole determinant of model identification – instead, it should be used as a "quick check" for any outstanding issues with the model.

Main functions

The 'id()' function is the workhorse function of the package which evaluates a user-supplied model against a number of identification rules. Calling the function in-line on a model string will print a table to the console with the results of the rule checks. Assigning the output of 'id()' to an object will save the results and can be printed later.

The 'scaling()' function lets the user verify if each latent variable is scaled in a model with latent variables. The output includes whether the variable is scaled, the scaling indicator (if applicable), the method used to scale it (if applicable), whether the model has a mean structure, and the reason why it is not scaled (if applicable).

The 'id2()' function uses the two-step rule for SEM identification in which (a) all structural paths are converted to covariances, making the model a confirmatory factor analysis (CFA) model, (b) the CFA model is checked for identification, (c) if the CFA is identified, the original SEM is converted into a simultaneous equations model (SEM) where all latent variables are treated as observed, and (d) the resulting model is checked for identification. If both the CFA and SEM are identified, the original SEM is identified. The 'id2()' function prints the results of both steps to the console.

Development notes

The core of the package was developed by Zach Vig, informed by Ken Bollen's 'Elements of Structural Equation Models (SEMs)' (2026), with style inspiration from the 'lavaan' package. Additionally, OpenAI large-language models (https://openai.com) were used to (a) edit functions and documentation for clarity, (b) implement the Depth-First Search algorithm for checking recursion, (c) help fix bugs, and (d) help write the test suite.

Author(s)

Maintainer: Zach Vig zachvig@rocketmail.com [copyright holder]

Authors:

See Also

Useful links:


Rules for confirmatory factor analysis models

Description

Confirmatory Factor Analysis (CFA) models are models with latent variables but no structural paths between latent variables. CFA rules assume each latent variable is correctly scaled (see scaling).

Usage

rule_cfa_three_indicator(partable)

rule_cfa_two_indicator(partable)

Arguments

partable

A lavaan parameter table

Details

Two Indicator Rule

In a model with more than one latent variable, each latent variable can have just two indicators if each indicator loads on exactly one variable, none of the errors of the indicators are correlated, and each latent variable correlates with at least one other latent variable. Sufficient but not necessary, unless higher-order factors are present in which case sufficiency cannot be established.

Three Indicator Rule

In a model with one or more latent variables, each latent variable can have just three indicators if each indicator loads on exactly one variable and none of the errors of the indicators are correlated. Sufficient but not necessary, unless higher-order factors are present in which case sufficiency cannot be established.

References

Bollen (2026). Elements of Structural Equation Models (SEMs).

Kenny, D. A. (1979). Correlation and causality.


Check recursion using Depth-First Search algorithm

Description

Check recursion using Depth-First Search algorithm

Usage

check_recursion(partable, start)

Arguments

partable

A lavaan parameter table

start

A character vector of variable names from which to start the algorithm.

Value

A logical value indicating whether the model is recursive (TRUE) or not (FALSE).


Classify a model using the parameter table for internal use

Description

Classify a model using the parameter table for internal use

Usage

classify_model(partable = NULL)

Format a function call for printing for internal use

Description

Format a function call for printing for internal use

Usage

format_lavaan_fun(fun)

Arguments

fun

A character string of the function call

Value

A formatted character string of the function call


Extracts the cmd object from a lavaan fitted object for internal use

Description

Extracts the cmd object from a lavaan fitted object for internal use

Usage

get_lavaan_cmd(obj)

Arguments

obj

A fitted lavaan object

Value

The function/command used to fit the model


Gather identification rules as a list

Description

The 'semrulesid' package stores identification rule functions internally. This function makes them available to the user as a list.

Usage

get_rules(rule = "*", model_type = "*")

Arguments

rule

A character vector specifying the name of the rule as it's defined in the package. Use "*" to get all rules (or all rules of the defined model type). Partial matches are acceptable.

model_type

A character vector specifying the model sub-type from which to get rules. Sub-types include "reg" (simultaneous equations models/regression models) and "cfa" (confirmatory factor analysis models). Use "sem" to get rules that apply to all structural equation models. Use "*" to get all rules in the package.

Value

A list object with the rule function(s) specified by the user.

Examples

# Get all rules for a CFA model
rules <- get_rules(rule = "*", model_type = "cfa")
# Get a specific rule for a SEM model
rules <- get_rules(rule = "latent_scaling", model_type = "sem")
latent_scaling_rule <- rules[[1]]

Evaluate common Structural Equation Model (SEM) identification rules

Description

This is the "workhorse" function of the semrulesid package. The user supplies a model string in lavaan syntax (see model.syntax for more details) and the function prints an informative table to the console about the status of the model on a variety of common identification rules.

Usage

id(x, print_msgs = TRUE, lav_fun = "sem", twostep = FALSE, ...)

id2(x, print_msgs = TRUE, lav_fun = "sem", ...)

Arguments

x

A character string model in lavaan syntax, a lavaan parameter table, or a fitted lavaan object.

print_msgs

Logical. If TRUE, the output will include why a rule does not pass or is not applicable, along with any other helpful information. Default: TRUE.

lav_fun

A character string specifying the lavaan function you intend to use to fit the model. This will ensure the correct model defaults are specified. Options currently include "lavaan", "sem", or "cfa". If a parameter table or fit object are supplied, this argument is ignored. Default: "sem".

twostep

A logical indicating whether to use the two-step identification rule instead of the usual one-step. See details. Default: FALSE.

...

Additional arguments passed to the lavaanify function from lavaan. See lavaanify for more information. If parameter tables or fitted model objects are supplied, these arguments are ignored.

Details

The primary output of id() is a table printed to the console, where rows correspond to rules, and columns include "Pass" (did the rule pass?), "Necessary" (is the rule necessary for identification?), and "Sufficient" (is the rule sufficient for identification?). These columns can take values "Yes", "No", or be left blank in the case of NA values.

If the user set the print_msgs argument to "TRUE" (which is the default), a column labeled "Messages" is appended to the table, and an output section called "Messages" is printed below the table. Messages include why a rule failed, why a rule is not applicable to the current model, or why the necessary/sufficient conditions may not apply as usual.

Messages are identified by a number, and corresponding message numbers are listed in the "Messages" column of the table.

lav_fun takes character values "lavaan", "sem", or "cfa", specifying which lavaan function the user intends to call (and thus which defaults should be used) when fitting the model in the case a model string is supplied. Supplying a parameter table or fitted model object ignores the lav_fun argument since defaults will have already been implemented. In these cases, you can set lav_fun = NA to avoid warnings about the argument being ignored. For id_mplus(), lav_fun is only needed if the user supplies an Mplus model string due to how lavaan::lav_mplus_syntax_model() works. If an Mplus input file is supplied, lav_fun is ignored since lavaan::lav_mplus_lavaan() always uses sem() defaults.

id2 is a wrapper function for calling id with argument twostep set to TRUE. The two-step method parses an SEM into a CFA model and a latent variable/structural model, and evaluates the identification rules on each. If both parts are identified, the whole model is identified.

Value

An object of class semid or semid2 (if twostep = TRUE or id2 is called. See details.)

Examples

my_model <- ' L1 =~ x1 + x2 + x3
              L2 =~ x4 + x5 + x6
              L3 =~ x7 + x8 + x9
              L2 ~ L1
              L3 ~ L2 '
id(my_model, print_msgs = TRUE, lav_fun = "sem",
   meanstructure = FALSE)
id2(my_model, print_msgs = TRUE, lav_fun = "sem",
   meanstructure = FALSE)

Evaluate identification rules for an Mplus model

Description

This is a wrapper function for id that handles Mplus model syntax (as a string) or Mplus input files (with extension ".inp"). It internally converts the Mplus model to a lavaan model, then calls id using the specified arguments.

Usage

id_mplus(x, print_msgs = TRUE, lav_fun = "sem", twostep = FALSE, ...)

Arguments

x

A character string model in Mplus syntax, or a path to an Mplus input file.

print_msgs

Logical. If TRUE, the output will include why a rule does not pass or is not applicable, along with any other helpful information. Default: TRUE.

lav_fun

A character string specifying the lavaan function you intend to use to fit the model. This will ensure the correct model defaults are specified. Options currently include "lavaan", "sem", or "cfa". If a parameter table or fit object are supplied, this argument is ignored. Default: "sem".

twostep

A logical indicating whether to use the two-step identification rule instead of the usual one-step. See details. Default: FALSE.

...

Additional arguments passed to the lavaanify function from lavaan. See lavaanify for more information. If parameter tables or fitted model objects are supplied, these arguments are ignored.

Value

An object of class semid or semid2 (if twostep = TRUE).

Examples

my_model <- ' L1 BY x1 x2 x3;
              L2 BY x4 x5 x6;
              L3 BY x7 x8 x9;
              L2 ON L1;
              L3 ON L2; '
id_mplus(my_model, print_msgs = TRUE, lav_fun = "sem",
   meanstructure = FALSE)


Printing function for rules list

Description

Printing function for rules list

Usage

## S3 method for class 'semid'
print(
  x,
  col_names = c("", "Pass", "Necessary", "Sufficient"),
  print_msgs = NULL,
  msgs_name = "Message",
  window = 56L,
  pos_lab = "Yes",
  neg_lab = "No",
  na_lab = "-",
  print_version = TRUE,
  print_lav_fun = TRUE,
  print_model_type = TRUE,
  name_value_sep = ":",
  na_rule_policy = c("footnote", "hide", "show"),
  ...
)

Arguments

x

A semid object

col_names

Character vector. The names of the columns of the main table

print_msgs

Logical. If TRUE messages are printed

msgs_name

Character. The name of the messages index column

window

Integer. The width of the output window

pos_lab

Character. The label for positive cells, e.g., "Yes".

neg_lab

Character. The label for negative cells, e.g., "No".

na_lab

Character. The label for NA/blank cells.

print_version

Logical. If TRUE, the version of the package is printed in a header before the rules output.

print_lav_fun

Logical. If TRUE, the lavaan function is printed in a header before the rules output.

print_model_type

Logical. If TRUE, the model type is printed in a header before the rules output.

name_value_sep

Character. The separator between the meta labels and values.

na_rule_policy

Character. How to handle rules that are not applicable to the model. Options are "hide" (default), "footnote", or "show". Option "hide" does not print them at all; option "footnote", prints their names only in a footnote after the main table; option "show" prints them in the main table with a message indicating they are not applicable.

...

Arguments passed to print.semscale if applicable.

Value

The function prints the rules list to the console and returns the semid object invisibly.


Printing function for two-step rules lists

Description

Printing function for two-step rules lists

Usage

## S3 method for class 'semid2'
print(
  x,
  ...,
  step_names = c("Measurement Model", "Latent Variable/Structural Model"),
  step_titles = c("Step 1", "Step 2"),
  print_version = TRUE,
  print_lav_fun = TRUE,
  print_model_type = TRUE,
  window = 56L,
  name_value_sep = ":"
)

Arguments

x

A semid2 object

...

Arguments to be passed to print.semid.

step_names

Character. The names of each step/block.

step_titles

Character. The labels to be given to each step.

print_version

Logical. If TRUE, the version of the package is printed in a header before the rules output.

print_lav_fun

Logical. If TRUE, the lavaan function is printed in a header before the rules output.

print_model_type

Logical. If TRUE, the model type is printed in a header before the rules output.

window

Integer. The width of the output window.

name_value_sep

Character. The separator between the meta labels and values.

Value

The function prints the two-step rules list to the console and returns the semid2 object invisibly.


Printing function for scaling table

Description

Printing function for scaling table

Usage

## S3 method for class 'semscale'
print(
  x,
  ...,
  print_msgs = TRUE,
  window = 56L,
  sep_spaces = 3L,
  indent_lengths = c(0L, 2L, 2L),
  na_lab = "na",
  pos_lab = "Yes",
  neg_lab = "No",
  empty_lab = "None",
  bullet = "-",
  name_value_sep = ":",
  print_version = TRUE,
  print_lav_fun = TRUE
)

Arguments

x

A semscale object, i.e., a list of scaling information for each latent variable in the model.

...

Not currently used.

print_msgs

Logical. If TRUE messages are printed.

window

Integer. The width of the output window.

sep_spaces

Integer. The number of spaces to separate the row names from the row values.

indent_lengths

Integer vector of length 3. The number of spaces to indent for each level of information (currently there are three supported).

na_lab

Character. The label for NA/blank cells.

pos_lab

Character. The label for positive cells, e.g., "Yes".

neg_lab

Character. The label for negative cells, e.g., "No".

empty_lab

Character. The label for empty cells, e.g., "None".

bullet

Character. The bullet symbol for messages, e.g., "-".

name_value_sep

Character. The separator between the row names and row values.

print_version

Logical. If TRUE, the version of the package is printed in a header before the rules output.

print_lav_fun

Logical. If TRUE, the lavaan function is printed in a header before the rules output.

Value

The function prints the scaling table to the console and returns the semscale object invisibly.


Rules for simultaneous equations/regression models

Description

Simultaneous equations models are structural models without any latent variables. Regression models are a subset of these models with a single outcome, while the general simultaneous equations model can handle multiple outcomes

Usage

rule_reg_null_byy(partable)

rule_reg_fully_recursive(partable)

rule_reg_recursive_corr_err(partable)

Arguments

partable

A lavaan parameter table

Details

Null B_YY Rule

No endogenous variable is the predictor of another endogenous variable. Sufficient but not necessary.

Fully Recursive Model Rule

The model has no feedback loops (i.e., is recursive) and there are no correlated errors ("fully" recursive). Sufficient but not necessary.

Recursive with Correlated Errors Rule

The model has no feedback loops (i.e., is recursive) but can have correlated errors as long as the errors are not for terms between which exists a direct structural/directional path. Sufficient but not necessary.

References

Bollen (2026). Elements of Structural Equation Models (SEMs).

Brito, C., & Pearl, J. (2002). A new identification condition for recursive models with correlated errors.


Check if latent variables are correctly scaled

Description

This function checks whether latent variables in the model are scaled, which is a necessary condition for model identification. A latent variable is scaled if it has a fixed loading on a scaling indicator or the latent variable has a fixed variance. When mean structure is present, the scaling indicator must also have a fixed intercept or the latent variable must also have a fixed mean. If any latent variable is not scaled, the model is not identified.

Usage

scaling(
  x,
  lav_fun = "sem",
  print_msgs = TRUE,
  lv = NULL,
  return_type = c("object", "logical"),
  ...
)

## S3 method for class 'lavaan'
scaling(
  x,
  lav_fun = "sem",
  print_msgs = TRUE,
  lv = NULL,
  return_type = c("object", "logical"),
  ...
)

Arguments

x

lavaan model syntax, a lavaan parameter table or a semrulesid ID object.

lav_fun

A character string specifying the function you intend to use to fit the model. This will ensure the correct model defaults are specified. Options currently include "lavaan", "sem", or "cfa". If a parameter table or fitted model object are supplied, this argument is ignored.

print_msgs

Logical. If TRUE the output will print why a latent variable is not scaled when applicable and the method or methods by which it is scaled when it is scaled. If FALSE, the output will only print whether the latent variable is scaled or not and a few other descriptives.

lv

An optional character vector of the specific latent variables you would like to check. Otherwise, all latent variables are extracted from the parameter table

return_type

A character string specifying the type of output. Options include "logical" (a logical vector specifying whether each latent variable is correctly scaled) or "object" (an object with additional information about the scaling of each latent variable). The latter is of type semscale and has a custom print method for additional diagnosis of identification issues.

...

Additional arguments passed to the lavaanify function from lavaan. See lavaanify for more information. These are only used when a model string is supplied. Otherwise, they are ignored.

Value

An object of class semscale with all scaling information (if return_type = "object"), or a logical vector (if return_type = "logical") indicating whether each latent variable is correctly scaled.

Examples

my_model <- ' L1 =~ x1 + x2 + x3
              L2 =~ x4 + x5 + x6
              L3 =~ x7 + x8 + x9
              L2 ~ L1
              L3 ~ L2 '
scaling(my_model, print_msgs = TRUE, lav_fun = "sem",
        meanstructure = FALSE)


Rules for all structural equation models

Description

Structural Equation Models (SEM) refer to the general class of models consisting of structural/directional relations, latent variables, or both.

Usage

rule_sem_ntheta(partable)

rule_sem_latent_scaling(partable)

rule_sem_two_emitted_paths(partable)

rule_sem_exogenous_x(partable)

Arguments

partable

A lavaan parameter table

Details

N_theta Rule (a.k.a. t-Rule)

The number of free parameters must be less than or equal to the number of sample means, variances, and covariances. Necessary but not sufficient

Latent Scaling Rule

In a model with latent variables, all latent variables must be correctly scaled (see scaling). Necessary but not sufficient.

Two Emitted Paths Rule

In a model with latent variables, all latent variables must emit two paths, either to other latent variables or to observed variables. This rule only applies to latent variables that have free variances and whose downstream variables have free error/disturbance variance. Necessary but not sufficient.

Exogenous X Rule/MIMIC Rules

These rules apply to models in which one or more latent variables have a causal indicator (or generally, are downstream of an observed variable), in addition to effect indicators. In such a model with a single latent variable, there only need to be one (or more) causal indicators as long as there are at least two effect indicators. This rule is not currently implemented for models with multiple latent variables. Sufficient but not necessary.

References

Bollen, K. A., Lilly, A. G., & Luo, L. (2024). Selecting scaling indicators in structural equation models (SEMs).

Bollen (2026). Elements of Structural Equation Models (SEMs).

Bollen & Davis (2009). Two Rules of Identification for Structural Equation Models.