| Type: | Package |
| Title: | Evaluate Structural Equation Model Identification Rules |
| Version: | 0.4.1 |
| Description: | Evaluates selected necessary and sufficient identification conditions in structural equation models (SEMs), including latent-variable scaling constraints. Output reports rule status and applicability and provides diagnostic messages to support model specification and respecification. The package is intended as a diagnostic aid and does not implement a universal identification algorithm. For more details, see Bollen (2026, ISBN:978-1009312820). |
| License: | GPL (≥ 3) |
| Encoding: | UTF-8 |
| Imports: | lavaan |
| Suggests: | knitr, rmarkdown, testthat (≥ 3.0.0) |
| Config/testthat/edition: | 3 |
| VignetteBuilder: | knitr |
| Language: | en-US |
| Config/roxygen2/version: | 8.1.0 |
| URL: | https://github.com/zacharyvig/semrulesid |
| BugReports: | https://github.com/zacharyvig/semrulesid/issues |
| NeedsCompilation: | no |
| Packaged: | 2026-09-15 13:13:05 UTC; ZACHMAC |
| Author: | Zach Vig [aut, cre, cph] |
| Maintainer: | Zach Vig <zachvig@rocketmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-26 16:20:07 UTC |
semrulesid package outline
Description
Evaluates selected necessary and sufficient identification conditions in structural equation models (SEMs), including latent-variable scaling constraints. Output reports rule status and applicability and provides diagnostic messages to support model specification and respecification. The package is intended as a diagnostic aid and does not implement a universal identification algorithm. For more details, see Bollen (2026, ISBN:978-1009312820).
Purpose
'semrulesid' allows the user to input a Structural Equation Model (SEM) in ['lavaan'](https://lavaan.ugent.be/) (Rosseel, 2012) syntax and check it against a number of identification rules from the literature. Rules are specified as being necessary and/or sufficient and specific reasons are given when a rule is broken. Users should not rely solely on the output of the package as the sole determinant of model identification – instead, it should be used as a "quick check" for any outstanding issues with the model.
Main functions
The 'id()' function is the workhorse function of the package which evaluates a user-supplied model against a number of identification rules. Calling the function in-line on a model string will print a table to the console with the results of the rule checks. Assigning the output of 'id()' to an object will save the results and can be printed later.
The 'scaling()' function lets the user verify if each latent variable is scaled in a model with latent variables. The output includes whether the variable is scaled, the scaling indicator (if applicable), the method used to scale it (if applicable), whether the model has a mean structure, and the reason why it is not scaled (if applicable).
The 'id2()' function uses the two-step rule for SEM identification in which (a) all structural paths are converted to covariances, making the model a confirmatory factor analysis (CFA) model, (b) the CFA model is checked for identification, (c) if the CFA is identified, the original SEM is converted into a simultaneous equations model (SEM) where all latent variables are treated as observed, and (d) the resulting model is checked for identification. If both the CFA and SEM are identified, the original SEM is identified. The 'id2()' function prints the results of both steps to the console.
Development notes
The core of the package was developed by Zach Vig, informed by Ken Bollen's 'Elements of Structural Equation Models (SEMs)' (2026), with style inspiration from the 'lavaan' package. Additionally, OpenAI large-language models (https://openai.com) were used to (a) edit functions and documentation for clarity, (b) implement the Depth-First Search algorithm for checking recursion, (c) help fix bugs, and (d) help write the test suite.
Author(s)
Maintainer: Zach Vig zachvig@rocketmail.com [copyright holder]
Authors:
Zach Vig zachvig@rocketmail.com [copyright holder]
See Also
Useful links:
Report bugs at https://github.com/zacharyvig/semrulesid/issues
Rules for confirmatory factor analysis models
Description
Confirmatory Factor Analysis (CFA) models are models with latent variables but no structural paths between latent variables. CFA rules assume each latent variable is correctly scaled (see scaling).
Usage
rule_cfa_three_indicator(partable)
rule_cfa_two_indicator(partable)
Arguments
partable |
A |
Details
- Two Indicator Rule
In a model with more than one latent variable, each latent variable can have just two indicators if each indicator loads on exactly one variable, none of the errors of the indicators are correlated, and each latent variable correlates with at least one other latent variable. Sufficient but not necessary, unless higher-order factors are present in which case sufficiency cannot be established.
- Three Indicator Rule
In a model with one or more latent variables, each latent variable can have just three indicators if each indicator loads on exactly one variable and none of the errors of the indicators are correlated. Sufficient but not necessary, unless higher-order factors are present in which case sufficiency cannot be established.
References
Bollen (2026). Elements of Structural Equation Models (SEMs).
Kenny, D. A. (1979). Correlation and causality.
Check recursion using Depth-First Search algorithm
Description
Check recursion using Depth-First Search algorithm
Usage
check_recursion(partable, start)
Arguments
partable |
A lavaan parameter table |
start |
A character vector of variable names from which to start the algorithm. |
Value
A logical value indicating whether the model is recursive (TRUE) or not (FALSE).
Classify a model using the parameter table for internal use
Description
Classify a model using the parameter table for internal use
Usage
classify_model(partable = NULL)
Format a function call for printing for internal use
Description
Format a function call for printing for internal use
Usage
format_lavaan_fun(fun)
Arguments
fun |
A character string of the function call |
Value
A formatted character string of the function call
Extracts the cmd object from a lavaan fitted object for internal use
Description
Extracts the cmd object from a lavaan fitted object for internal use
Usage
get_lavaan_cmd(obj)
Arguments
obj |
A fitted lavaan object |
Value
The function/command used to fit the model
Gather identification rules as a list
Description
The 'semrulesid' package stores identification rule functions internally. This function makes them available to the user as a list.
Usage
get_rules(rule = "*", model_type = "*")
Arguments
rule |
A character vector specifying the name of the rule as it's defined in the package. Use "*" to get all rules (or all rules of the defined model type). Partial matches are acceptable. |
model_type |
A character vector specifying the model sub-type from which to get rules. Sub-types include "reg" (simultaneous equations models/regression models) and "cfa" (confirmatory factor analysis models). Use "sem" to get rules that apply to all structural equation models. Use "*" to get all rules in the package. |
Value
A list object with the rule function(s) specified by the user.
Examples
# Get all rules for a CFA model
rules <- get_rules(rule = "*", model_type = "cfa")
# Get a specific rule for a SEM model
rules <- get_rules(rule = "latent_scaling", model_type = "sem")
latent_scaling_rule <- rules[[1]]
Evaluate common Structural Equation Model (SEM) identification rules
Description
This is the "workhorse" function of the semrulesid package. The user
supplies a model string in lavaan syntax (see
model.syntax for more details) and the function prints an
informative table to the console about the status of the model on a variety
of common identification rules.
Usage
id(x, print_msgs = TRUE, lav_fun = "sem", twostep = FALSE, ...)
id2(x, print_msgs = TRUE, lav_fun = "sem", ...)
Arguments
x |
A character string model in |
print_msgs |
Logical. If |
lav_fun |
A character string specifying the lavaan function you intend to use to fit the model. This will ensure the correct model defaults are specified. Options currently include "lavaan", "sem", or "cfa". If a parameter table or fit object are supplied, this argument is ignored. Default: "sem". |
twostep |
A logical indicating whether to use the two-step
identification rule instead of the usual one-step. See details.
Default: |
... |
Additional arguments passed to the |
Details
The primary output of id() is a table printed to the console, where
rows correspond to rules, and columns include "Pass" (did the rule pass?),
"Necessary" (is the rule necessary for identification?), and "Sufficient" (is
the rule sufficient for identification?). These columns can take values
"Yes", "No", or be left blank in the case of NA values.
If the user set the print_msgs argument to "TRUE" (which is the
default), a column labeled "Messages" is appended to the table, and an output
section called "Messages" is printed below the table. Messages include why a
rule failed, why a rule is not applicable to the current model, or why the
necessary/sufficient conditions may not apply as usual.
Messages are identified by a number, and corresponding message numbers are listed in the "Messages" column of the table.
lav_fun takes character values "lavaan", "sem", or "cfa", specifying
which lavaan function the user intends to call (and thus which
defaults should be used) when fitting the model in the case a model string
is supplied. Supplying a parameter table or fitted model object ignores the
lav_fun argument since defaults will have already been implemented. In
these cases, you can set lav_fun = NA to avoid warnings about the
argument being ignored. For id_mplus(), lav_fun is only needed
if the user supplies an Mplus model string due to how
lavaan::lav_mplus_syntax_model() works. If an Mplus input file is
supplied, lav_fun is ignored since lavaan::lav_mplus_lavaan()
always uses sem() defaults.
id2 is a wrapper function for calling id with argument
twostep set to TRUE. The two-step method parses an SEM into a
CFA model and a latent variable/structural model, and evaluates the
identification rules on each. If both parts are identified, the whole model
is identified.
Value
An object of class semid or semid2 (if twostep =
TRUE or id2 is called. See details.)
Examples
my_model <- ' L1 =~ x1 + x2 + x3
L2 =~ x4 + x5 + x6
L3 =~ x7 + x8 + x9
L2 ~ L1
L3 ~ L2 '
id(my_model, print_msgs = TRUE, lav_fun = "sem",
meanstructure = FALSE)
id2(my_model, print_msgs = TRUE, lav_fun = "sem",
meanstructure = FALSE)
Evaluate identification rules for an Mplus model
Description
This is a wrapper function for id that handles Mplus model
syntax (as a string) or Mplus input files (with extension ".inp"). It
internally converts the Mplus model to a lavaan model, then calls
id using the specified arguments.
Usage
id_mplus(x, print_msgs = TRUE, lav_fun = "sem", twostep = FALSE, ...)
Arguments
x |
A character string model in Mplus syntax, or a path to an Mplus input file. |
print_msgs |
Logical. If |
lav_fun |
A character string specifying the lavaan function you intend to use to fit the model. This will ensure the correct model defaults are specified. Options currently include "lavaan", "sem", or "cfa". If a parameter table or fit object are supplied, this argument is ignored. Default: "sem". |
twostep |
A logical indicating whether to use the two-step
identification rule instead of the usual one-step. See details.
Default: |
... |
Additional arguments passed to the |
Value
An object of class semid or semid2 (if twostep =
TRUE).
Examples
my_model <- ' L1 BY x1 x2 x3;
L2 BY x4 x5 x6;
L3 BY x7 x8 x9;
L2 ON L1;
L3 ON L2; '
id_mplus(my_model, print_msgs = TRUE, lav_fun = "sem",
meanstructure = FALSE)
Printing function for rules list
Description
Printing function for rules list
Usage
## S3 method for class 'semid'
print(
x,
col_names = c("", "Pass", "Necessary", "Sufficient"),
print_msgs = NULL,
msgs_name = "Message",
window = 56L,
pos_lab = "Yes",
neg_lab = "No",
na_lab = "-",
print_version = TRUE,
print_lav_fun = TRUE,
print_model_type = TRUE,
name_value_sep = ":",
na_rule_policy = c("footnote", "hide", "show"),
...
)
Arguments
x |
A |
col_names |
Character vector. The names of the columns of the main table |
print_msgs |
Logical. If |
msgs_name |
Character. The name of the messages index column |
window |
Integer. The width of the output window |
pos_lab |
Character. The label for positive cells, e.g., "Yes". |
neg_lab |
Character. The label for negative cells, e.g., "No". |
na_lab |
Character. The label for |
print_version |
Logical. If |
print_lav_fun |
Logical. If |
print_model_type |
Logical. If |
name_value_sep |
Character. The separator between the meta labels and values. |
na_rule_policy |
Character. How to handle rules that are not applicable to the model. Options are "hide" (default), "footnote", or "show". Option "hide" does not print them at all; option "footnote", prints their names only in a footnote after the main table; option "show" prints them in the main table with a message indicating they are not applicable. |
... |
Arguments passed to print.semscale if applicable. |
Value
The function prints the rules list to the console and returns the
semid object invisibly.
Printing function for two-step rules lists
Description
Printing function for two-step rules lists
Usage
## S3 method for class 'semid2'
print(
x,
...,
step_names = c("Measurement Model", "Latent Variable/Structural Model"),
step_titles = c("Step 1", "Step 2"),
print_version = TRUE,
print_lav_fun = TRUE,
print_model_type = TRUE,
window = 56L,
name_value_sep = ":"
)
Arguments
x |
A |
... |
Arguments to be passed to print.semid. |
step_names |
Character. The names of each step/block. |
step_titles |
Character. The labels to be given to each step. |
print_version |
Logical. If |
print_lav_fun |
Logical. If |
print_model_type |
Logical. If |
window |
Integer. The width of the output window. |
name_value_sep |
Character. The separator between the meta labels and values. |
Value
The function prints the two-step rules list to the console and
returns the semid2 object invisibly.
Printing function for scaling table
Description
Printing function for scaling table
Usage
## S3 method for class 'semscale'
print(
x,
...,
print_msgs = TRUE,
window = 56L,
sep_spaces = 3L,
indent_lengths = c(0L, 2L, 2L),
na_lab = "na",
pos_lab = "Yes",
neg_lab = "No",
empty_lab = "None",
bullet = "-",
name_value_sep = ":",
print_version = TRUE,
print_lav_fun = TRUE
)
Arguments
x |
A |
... |
Not currently used. |
print_msgs |
Logical. If |
window |
Integer. The width of the output window. |
sep_spaces |
Integer. The number of spaces to separate the row names from the row values. |
indent_lengths |
Integer vector of length 3. The number of spaces to indent for each level of information (currently there are three supported). |
na_lab |
Character. The label for NA/blank cells. |
pos_lab |
Character. The label for positive cells, e.g., "Yes". |
neg_lab |
Character. The label for negative cells, e.g., "No". |
empty_lab |
Character. The label for empty cells, e.g., "None". |
bullet |
Character. The bullet symbol for messages, e.g., "-". |
name_value_sep |
Character. The separator between the row names and row values. |
print_version |
Logical. If |
print_lav_fun |
Logical. If |
Value
The function prints the scaling table to the console and returns the
semscale object invisibly.
Rules for simultaneous equations/regression models
Description
Simultaneous equations models are structural models without any latent variables. Regression models are a subset of these models with a single outcome, while the general simultaneous equations model can handle multiple outcomes
Usage
rule_reg_null_byy(partable)
rule_reg_fully_recursive(partable)
rule_reg_recursive_corr_err(partable)
Arguments
partable |
A |
Details
- Null B_YY Rule
No endogenous variable is the predictor of another endogenous variable. Sufficient but not necessary.
- Fully Recursive Model Rule
The model has no feedback loops (i.e., is recursive) and there are no correlated errors ("fully" recursive). Sufficient but not necessary.
- Recursive with Correlated Errors Rule
The model has no feedback loops (i.e., is recursive) but can have correlated errors as long as the errors are not for terms between which exists a direct structural/directional path. Sufficient but not necessary.
References
Bollen (2026). Elements of Structural Equation Models (SEMs).
Brito, C., & Pearl, J. (2002). A new identification condition for recursive models with correlated errors.
Check if latent variables are correctly scaled
Description
This function checks whether latent variables in the model are scaled, which is a necessary condition for model identification. A latent variable is scaled if it has a fixed loading on a scaling indicator or the latent variable has a fixed variance. When mean structure is present, the scaling indicator must also have a fixed intercept or the latent variable must also have a fixed mean. If any latent variable is not scaled, the model is not identified.
Usage
scaling(
x,
lav_fun = "sem",
print_msgs = TRUE,
lv = NULL,
return_type = c("object", "logical"),
...
)
## S3 method for class 'lavaan'
scaling(
x,
lav_fun = "sem",
print_msgs = TRUE,
lv = NULL,
return_type = c("object", "logical"),
...
)
Arguments
x |
|
lav_fun |
A character string specifying the function you intend to use to fit the model. This will ensure the correct model defaults are specified. Options currently include "lavaan", "sem", or "cfa". If a parameter table or fitted model object are supplied, this argument is ignored. |
print_msgs |
Logical. If |
lv |
An optional character vector of the specific latent variables you would like to check. Otherwise, all latent variables are extracted from the parameter table |
return_type |
A character string specifying the type of output. Options
include "logical" (a logical vector specifying whether each latent
variable is correctly scaled) or "object" (an object with additional
information about the scaling of each latent variable). The latter is
of type |
... |
Additional arguments passed to the |
Value
An object of class semscale with all scaling information (if
return_type = "object"), or a logical vector (if return_type =
"logical") indicating whether each latent variable is correctly scaled.
Examples
my_model <- ' L1 =~ x1 + x2 + x3
L2 =~ x4 + x5 + x6
L3 =~ x7 + x8 + x9
L2 ~ L1
L3 ~ L2 '
scaling(my_model, print_msgs = TRUE, lav_fun = "sem",
meanstructure = FALSE)
Rules for all structural equation models
Description
Structural Equation Models (SEM) refer to the general class of models consisting of structural/directional relations, latent variables, or both.
Usage
rule_sem_ntheta(partable)
rule_sem_latent_scaling(partable)
rule_sem_two_emitted_paths(partable)
rule_sem_exogenous_x(partable)
Arguments
partable |
A |
Details
- N_theta Rule (a.k.a. t-Rule)
The number of free parameters must be less than or equal to the number of sample means, variances, and covariances. Necessary but not sufficient
- Latent Scaling Rule
In a model with latent variables, all latent variables must be correctly scaled (see scaling). Necessary but not sufficient.
- Two Emitted Paths Rule
In a model with latent variables, all latent variables must emit two paths, either to other latent variables or to observed variables. This rule only applies to latent variables that have free variances and whose downstream variables have free error/disturbance variance. Necessary but not sufficient.
- Exogenous X Rule/MIMIC Rules
These rules apply to models in which one or more latent variables have a causal indicator (or generally, are downstream of an observed variable), in addition to effect indicators. In such a model with a single latent variable, there only need to be one (or more) causal indicators as long as there are at least two effect indicators. This rule is not currently implemented for models with multiple latent variables. Sufficient but not necessary.
References
Bollen, K. A., Lilly, A. G., & Luo, L. (2024). Selecting scaling indicators in structural equation models (SEMs).
Bollen (2026). Elements of Structural Equation Models (SEMs).
Bollen & Davis (2009). Two Rules of Identification for Structural Equation Models.