Data analysis

Data and statistical analysis

Upload your data. The studio cleans it step by step, shows every change, and gives you a cleaned file and an audit report.

1Upload your data

The first row must hold the variable names. For Excel, the first sheet is read.

Or drop your file here, or click to choose

If the buttons above do not open your files, copy the cells from Excel (include the header row) or paste CSV text here, then press Load.

Your file is processed in your browser and is not uploaded anywhere. Do not rely on this tool alone to remove patient identifiers. De-identification is not part of this phase.
Independent checks run on the analyses

Each statistic was computed by this page and by Python on the same random test data. Each row used one random data set, not many. R was not available. Differences are the largest absolute difference found.

AnalysisCompared withLargest difference
Welch t-test, Mann-Whitney, ANOVA, Kruskal-Wallis, chi-square, Fisher exact, Pearson, Spearmanscipy 1.183.3e-9 (17 values)
Paired t-test (t, p, 95% CI)scipy1.2e-13
Wilcoxon signed-rank (exact for up to 50 pairs without ties or zeros; normal approximation otherwise)scipy0 (exact), 6.7e-15 (approximation)
Kaplan-Meier survival, Greenwood standard error, median and its 95% CIstatsmodels 0.151.1e-16 (survival), 1.4e-17 (SE), medians and CIs identical
Log-rank teststatsmodels2.7e-15
Linear regression (coefficients, SE, p, CI, R-squared, F test)statsmodels5e-14
Variance inflation factorstatsmodelsagrees to 6 decimals
Logistic regression (coefficients, SE, p, log-likelihood, likelihood ratio test)statsmodels8.4e-9
AUCscikit-learn 1.91.1e-16
Cox regression with Breslow ties (coefficients, SE, p, log-likelihood)statsmodels1.5e-14
Concordance indexIndependent pair-counting code in Python0
Newcombe 95% CI for a risk difference (3 tables)statsmodels2.1e-9
Decision curve net benefit (6 thresholds)Direct calculation in Python1.1e-16
Principal components (variance explained, eigenvalues, loadings)scikit-learn1.7e-16, 1.8e-15, 2.9e-15
k-means sum of squares; silhouette from the same labelsscikit-learnidentical on 3 separated clusters; silhouette 1e-15
Ridge, LASSO and elastic net coefficients at fixed penalty strengthscikit-learn (ridge: closed form)3.9e-11
Cross-validated error across the penalty path (same folds)scikit-learn7.5e-13
Gradient boosting predictions (same settings, regression and classification)scikit-learn3.6e-15 and 2.2e-16

Python scripts (members): each of the 13 script types was run in Python on the 90-row test file. Coefficients, test statistics and p-values matched the studio. Cross-validated figures, the random forest and the penalty choice differ slightly because scikit-learn draws its own folds and random numbers.

Checked end to end: on a 90-row file, the linear, logistic and Cox regression tables produced by this page (including dummy coding of categories) agree with statsmodels to the rounding shown.

Not identical by design: the random forest uses random resampling, so it is checked by performance, not value by value. On the test data the out-of-bag AUC was 0.708 here and 0.707 in scikit-learn, with the same top variable and the same variable ranking. For a numeric outcome, out-of-bag R-squared was 0.633 here and 0.597 in scikit-learn. For k-means on overlapping data with 5 clusters, the best of 20 random starts here was 0.2% above the best of 30 starts in scikit-learn (the two agreed for 2, 3 and 4 clusters).

Not covered: the automatic choice between mean-based and rank-based methods is a rule of thumb. Effect sizes, the Table 1 layout, the bootstrap AUC interval, the calibration table, cross-validation fold assignment, the charts, the cost-effectiveness arithmetic (checked by hand on one example) and the proportional hazards assumption were not compared with other software.

Sources
Kahn MG, et al. A harmonized data quality assessment terminology and framework for the secondary use of electronic health record data. eGEMs. 2016;4(1):1244. doi:10.13063/2327-9214.1244
Van den Broeck J, et al. Data cleaning: detecting, diagnosing, and editing data abnormalities. PLoS Med. 2005;2(10):e267. doi:10.1371/journal.pmed.0020267
Sterne JAC, et al. Multiple imputation for missing data in epidemiological and clinical research. BMJ. 2009;338:b2393. doi:10.1136/bmj.b2393
Kaplan EL, Meier P. Nonparametric estimation from incomplete observations. J Am Stat Assoc. 1958;53(282):457-481. doi:10.1080/01621459.1958.10501452
Cox DR. Regression models and life-tables. J R Stat Soc Series B. 1972;34(2):187-220. doi:10.1111/j.2517-6161.1972.tb00899.x
Wilcoxon F. Individual comparisons by ranking methods. Biometrics Bulletin. 1945;1(6):80-83. doi:10.2307/3001968
Newcombe RG. Interval estimation for the difference between independent proportions: comparison of eleven methods. Stat Med. 1998;17(8):873-890. doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I
Altman DG. Confidence intervals for the number needed to treat. BMJ. 1998;317(7168):1309-1312. doi:10.1136/bmj.317.7168.1309
Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361
Hoerl AE, Kennard RW. Ridge regression: biased estimation for nonorthogonal problems. Technometrics. 1970;12(1):55-67. doi:10.1080/00401706.1970.10488634
Tibshirani R. Regression shrinkage and selection via the lasso. J R Stat Soc Series B. 1996;58(1):267-288. doi:10.1111/j.2517-6161.1996.tb02080.x
Zou H, Hastie T. Regularization and variable selection via the elastic net. J R Stat Soc Series B. 2005;67(2):301-320. doi:10.1111/j.1467-9868.2005.00503.x
Jolliffe IT, Cadima J. Principal component analysis: a review and recent developments. Philos Trans A. 2016;374(2065):20150202. doi:10.1098/rsta.2015.0202
Rousseeuw PJ. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J Comput Appl Math. 1987;20:53-65. doi:10.1016/0377-0427(87)90125-7
Breiman L. Random forests. Mach Learn. 2001;45:5-32. doi:10.1023/A:1010933404324
Friedman JH. Greedy function approximation: a gradient boosting machine. Ann Stat. 2001;29(5):1189-1232. doi:10.1214/aos/1013203451
The plausibility limits are a draft based on clinical judgment and have not been reviewed or sourced. Uniqueness is added to the Kahn categories. Version 0.4. Datathrob. Clinical research made simple.