Links to specific topics

(See also under "Labels" at the bottom-left area of this blog)
[ Welcome post ] [ Installation issues ] [ WarpPLS.com ] [ Posts with YouTube links ] [ Model-driven data analytics ] [ PLS-SEM email list ]
Showing posts with label resampling. Show all posts
Showing posts with label resampling. Show all posts

Monday, August 29, 2011

Using WarpPLS in E-Collaboration Studies: Mediating Effects, Control and Second Order Variables, and Algorithm Choices

A new article discussing WarpPLS is available. The article is titled “Using WarpPLS in E-Collaboration Studies: Mediating Effects, Control and Second Order Variables, and Algorithm Choices”. It has been recently published in the International Journal of e-Collaboration. A full text version of the article is available here as a PDF file. Below is the abstract of the article.

This is a follow-up on two previous articles on WarpPLS and e-collaboration. The first discussed the five main steps through which a variance-based nonlinear structural equation modeling analysis could be conducted with the software WarpPLS (Kock, 2010b). The second covered specific features related to grouped descriptive statistics, viewing and changing analysis algorithm and resampling settings, and viewing and saving various results (Kock, 2011). This and the previous articles use data from the same e-collaboration study as a basis for the discussion of important WarpPLS features. Unlike the previous articles, the focus here is on a brief discussion of more advanced issues, such as: testing the significance of mediating effects, including control variables in an analysis, using second order latent variables, choosing the right warping algorithm, and using bootstrapping and jackknifing in combination.

Thursday, January 28, 2010

Bootstrapping or jackknifing (or both) in WarpPLS?

The blog post below refers to resampling methods, notably bootstrapping and jackknifing, which are often used for the generation of estimates employed and hypothesis testing. Even though they are widely used, resampling methods are inherently unstable. More recent versions of WarpPLS employ "stable" methods for the same purpose, with various advantages. See the most recent version of the WarpPLS User Manual (linked below) for more details.

http://www.scriptwarp.com/warppls/#User_Manual

***

Arguably jackknifing does a better job at addressing problems associated with the presence of outliers due to errors in data collection. Generally speaking, jackknifing tends to generate more stable resample path coefficients (and thus more reliable P values) with small sample sizes (lower than 100), and with samples containing outliers. In these cases, outlier data points do not appear more than once in the set of resamples, which accounts for the better performance of jackknifing (see, e.g., Chiquoine & Hjalmarsson, 2009).

Bootstrapping tends to generate more stable resample path coefficients (and thus more reliable P values) with larger samples and with samples where the data points are evenly distributed on a scatter plot. The use of bootstrapping with small sample sizes (lower than 100) has been discouraged (Nevitt & Hancock, 2001).

Since the warping algorithms are also sensitive to the presence of outliers, in many cases it is a good idea to estimate P values with both bootstrapping and jackknifing, and use the P values associated with the most stable coefficients. An indication of instability is a high P value (i.e., statistically insignificant) associated with path coefficients that could be reasonably expected to have low P values. For example, with a sample size of 100, a path coefficient of .2 could be reasonably expected to yield a P value that is statistically significant at the .05 level. If that is not the case, there may be a stability problem. Another indication of instability is a marked difference between the P values estimated through bootstrapping and jackknifing.

P values can be easily estimated using both resampling methods, bootstrapping and jackknifing, by following this simple procedure. Run an SEM analysis of the desired model, using one of the resampling methods, and save the project. Then save the project again, this time with a different name, change the resampling method, and run the SEM analysis again. Then save the second project again. Each project file will now have results that refer to one of the two resampling methods. The P values can then be compared, and the most stable ones used in a research report on the SEM analysis.

References:

Chiquoine, B., & Hjalmarsson, E. (2009). Jackknifing stock return predictions. Journal of Empirical Finance, 16(5), 793-803.

Nevitt, J., & Hancock, G.R. (2001). Performance of bootstrapping approaches to model test statistics and parameter standard error estimation in structural equation modeling. Structural Equation Modeling, 8(3), 353-377.

How many resamples to use in bootstrapping?


The default number of resamples is 100 for bootstrapping in WarpPLS. This setting can be modified by entering a different number in the appropriate edit box. (Please note that we are talking about the number of resamples here, not the original data sample size.)

Leaving the number of resamples for bootstrapping as 100 is recommended because it has been shown that higher numbers of resamples lead to negligible improvements in the reliability of P values; in fact, even setting the number of resamples at 50 is likely to lead to fairly reliable P value estimates (Efron et al., 2004).

Conversely, increasing the number of resamples well beyond 100 leads to a higher computation load on the software, making the software look like it is having a hard time coming up with the results. In very complex models, a high number of resamples may make the software run very slowly.

Some researchers have suggested in the past that a large number of resamples can address problems with the data, such as the presence of outliers due to errors in data collection. This opinion is not shared by the original developer of the bootstrapping method, Bradley Efron (see, e.g., Efron et al., 2004).

Reference:

Efron, B., Rogosa, D., & Tibshirani, R. (2004). Resampling methods of estimation. In N.J. Smelser, & P.B. Baltes (Eds.). International Encyclopedia of the Social & Behavioral Sciences (pp. 13216-13220). New York, NY: Elsevier.

Saturday, January 23, 2010

How are the model fit indices calculated by WarpPLS?


WarpPLS is unique among software that implement PLS-SEM algorithms in that it provides users with a number of model-wide fit indices; arguably more than any other SEM software. Three of the main model fit indices calculated by WarpPLS are the following: average path coefficient (APC), average R-squared (ARS), and average variance inflation factor (AFVIF).

They are discussed in the WarpPLS User Manual, which is available separately from the software, as a standalone document, on the WarpPLS web site.

These fit indices (there are several others) are calculated as their name implies, that is, as averages of: the (absolute values of the ) path coefficients in the model, the R-squared values in the model, and the variance inflation factors in the model. All of these are also provided individually by the software.

The P values for APC and ARS are calculated through re-sampling. A correction is made to account for the fact that these indices are calculated based on other parameters, which leads to a biasing effect – a variance reduction effect associated with the central limit theorem.

Typically the addition of new latent variables into a model will increase the ARS, even if those latent variables are weakly associated with the existing latent variables in the model. However, that will generally lead to a decrease in APC, since the path coefficients associated with the new latent variables will be low. Thus, the APC and ARS will counterbalance each other, and will only increase together if the latent variables that are added to the model enhance the overall predictive and explanatory quality of the model.

The AFVIF index will increase if new latent variables are added to the model in such a way as to add multicolinearity to the model, which may result from the inclusion of new latent variables that overlap in meaning with existing latent variables. It is generally undesirable to have different latent variables in the same model that measure the same thing; those should be combined into one single latent variable. Thus, the AFVIF brings in a new dimension that adds to a comprehensive assessment of a model’s overall predictive and explanatory quality.

Starting in version 6.0 of WarpPLS, new indices are available that allow investigators to assess the fit between the model-implied and empirical indicator correlation matrices. These new indices are available from the "Explore" menu option, after Step 5 is completed. They are discussed on page 26 of the WarpPLS User Manual for version 6.0, and on a video clip (links below).

http://www.scriptwarp.com/warppls/UserManual_v_6_0.pdf#page=26

https://youtu.be/YutkhEPW-CE

Also, these new indices are discussed in the following article (see link below for PDF): Kock, N. (2020). Using indicator correlation fit indices in PLS-SEM: Selecting the algorithm with the best fit. Data Analysis Perspectives Journal, 1(4), 1-4.

https://scriptwarp.com/dapj

As a final note, I would like to point out that the interpretation of the model fit indices depends on the goal of the SEM analysis. If the goal is to test hypotheses, where each arrow represents a hypothesis, then the model fit indices are of some importance, but in a limited way. However, if the goal is to find out whether one model has a better fit with the original data than another, then the model fit indices become more important, and are a useful set of measures related to model quality.

Wednesday, December 23, 2009

Change resampling method in WarpPLS: YouTube video


The blog post below refers to resampling methods, bootstrapping and jackknifing, which are often used for the generation of estimates employed and hypothesis testing. Even though they are widely used, resampling methods are inherently unstable, as illustrated in the post. More recent versions of WarpPLS employ "stable" methods for the same purpose, with various advantages. See the most recent version of the WarpPLS User Manual (linked below) for more details.

http://www.scriptwarp.com/warppls/#User_Manual

***

A new Youtube video is available for WarpPLS:

http://www.youtube.com/watch?v=Hf-t70r7NKo

This video shows how one can conduct an SEM analysis using WarpPLS, save that analysis with a different project name, change the resampling method (from bootstrapping to jackknifing), and then redo the analysis.

At the end, the user has two project files, one with all of the P values calculated through bootstrapping, and the other with all of the P values calculated through jackknifing.

As noted in the WarpPLS User Manual, bootstrapping and jackknifing provide a good complement to each other in the context of warped PLS-based SEM.

Thus, some users may want to run two analyses of the same model, one with each resampling method,  and use the results that are associated with the most stable resample path coefficients. These will typically be the ones with the lowest P values, since P values go up as the standard errors in the resample set go up. High resample standard errors are associated with instability. The instability itself often comes from outliers, which may drastically change the shape of a warped relationship in each resample.

Well, moving from statspeach to plain English, there are good theoretical reasons to recommend that users choose the most stable results (i.e., with the lowest P values) as the results that they will use in research reports, whether they are obtained with bootstrapping or jackknifing. The choice may be made individually, for each path coefficient. This should be disclosed to the readers of the report; a sentence like this would probably be enough: "Both bootstrapping or jackknifing were used in the analyses. The results reported here are those associated with the most stable resample estimates."