A GUI-based tool for statistical analysis of experimental data.
Users can import Excel or CSV files, define groups, and let BioMedStatX handle the rest: outlier detection, assumption checks, guided data transformations, automatic test selection, post-hoc analyses, and fully documented HTML reports.
Repository: philippkrumm/BioMedStatX
Releases (Download ready-to-use app): https://github.com/philippkrumm/BioMedStatX/releases
License: MIT License
BioMedStatX is designed for experimental and biomedical research workflows:
-
Intuitive GUI
Load data, select groups and variables, and trigger analyses without writing code. -
Automated statistical pipeline
- Assumption checks (normality, variance homogeneity, etc.)
- Guided data transformations where appropriate
- Automatic selection of parametric vs. nonparametric tests for supported designs
- Guided post-hoc analyses when needed
-
Rich output
- Publication-ready plots, exportable as PNG or SVG from the interactive figure builder
- Self-contained HTML report with all intermediate steps, assumptions, test decisions, and an interactive decision tree
- Clear documentation of which test was selected and why
-
Multi-Dataset Analysis
- Run several measurement columns (for example several genes or markers) through the same factor mapping in one go
- One summary card per column in a combined report, with Benjamini-Hochberg FDR correction applied across the family of p-values
- Columns that could not be analysed are listed with their reason instead of quietly disappearing
- Restricted to ANOVA-capable designs
-
Outlier detection (a separate step, run from Analysis -> Detect Outliers)
- Grubbs' test (selected by default) and/or a modified Z-score, iterative by default
- Works on a copy and writes its findings to an output file of your choosing: the loaded data is never edited and no rows are dropped
- Deliberately not part of the automatic analysis run, so removing a flagged value stays your decision
-
Excel/CSV support
- Direct import of
.xlsxand.csvfiles
- Direct import of
-
Correlation & Regression
- Pearson / Spearman correlation with 95% confidence intervals (auto-selected by sample size and distribution shape, i.e. skewness / kurtosis, not a Shapiro-Wilk gate)
- Simple and multiple linear regression (OLS) with full residual diagnostics (Ramsey RESET, Breusch-Pagan, Shapiro-Wilk on residuals)
- Exploratory correlation matrix across all numeric variables with FDR (Benjamini-Hochberg) or Bonferroni correction, pairwise deletion, and optional stratification
-
Transparent methodology
- Advanced explanations for ANOVA workflows: see Advanced ANOVA Guide
- Correlation and regression methodology: see Correlation & Regression Guide
BioMedStatX is distributed as a standalone application for end users and as source code for developers.
Go to the GitHub Releases page and download the archive for your platform:
-> https://github.com/philippkrumm/BioMedStatX/releases
Windows
- Download the Windows archive.
- Extract it to a folder of your choice. The app uses one-folder packaging: keep
BioMedStatX.exeand the_internalfolder together. - Start the application by double-clicking
BioMedStatX.exe.
macOS (Apple Silicon)
- Download the macOS archive.
- Extract it and move
BioMedStatX.appwhere you want to keep it. - The app is not signed with an Apple Developer certificate, so macOS will refuse to open it on the first attempt. Right-click the app and choose "Open", then confirm.
- If that is not enough, macOS has flagged the download as quarantined. The release notes for your version give the exact
xattrcommand to clear it.
Both builds are self-contained. No Python installation or command-line usage is required.
The app checks GitHub for a newer release a few seconds after it starts, and again on demand via Help -> Check for Updates.... It only reads the public releases page and never uploads anything; if the machine is offline, the check fails quietly.
For a step-by-step walkthrough of the GUI (with screenshots), see:
-> How to use BioMedStatX (User Guide with screenshots)
If you want to inspect or modify the source code, or contribute to the project:
- Clone the repository
git clone https://github.com/philippkrumm/BioMedStatX.git
cd BioMedStatXFor information about the helper scripts included in this repository, including start.sh and run.bat, see: docs/SCRIPTS.md
BioMedStatX provides a GUI-based workflow for statistical analysis.
A detailed, step-by-step User Guide with screenshots and numbered button references is available here:
-> How to use BioMedStatX (User Guide with screenshots)
-
Start BioMedStatX
Launch the main application (see the User Guide for details on entry points and executables). -
Load your dataset
- Import an Excel or CSV file.
-
Define groups and variables
- Select the sheet (for Excel files).
- Choose grouping variables and measurement columns.
- Optionally restrict groups via Select Groups For Analysis.
- Optionally restrict rows via Filter bucket.
-
Configure analysis options (optional)
- Choose plots and statistics to generate.
- Adjust settings as needed (see the User Guide for screenshots).
-
Run the analysis
BioMedStatX will automatically detect outliers, check assumptions, select the appropriate test, and run post-hoc analyses when needed. When prompted, you decide whether to apply a suggested transformation and which post-hoc procedure to use. -
Inspect the output
- Review plots and statistical results.
- Open the generated HTML report for a fully documented analysis pipeline, including the interactive decision tree.
For a complete, screenshot-based walkthrough, including which button to click at each step, see the User Guide.
-
User Guide (GUI, step-by-step with screenshots): -> docs/HowTo.md
Includes:
- first-time orientation (what the app decides vs what the user decides),
- minimum data structure requirements,
- end-to-end analysis order,
- mapping-to-test logic and export interpretation.
-
Advanced ANOVA methodology and interpretation: -> docs/ADVANCED_ANOVA_GUIDE.md
-
Correlation & Regression methodology and interpretation: -> docs/CORRELATION_REGRESSION_GUIDE.md
Additional documentation can be added to the docs/ folder.
- Standard nonparametric workflows are available for common one-factor designs such as Mann-Whitney, Wilcoxon, Kruskal-Wallis, Friedman, and related post-hoc analyses.
- Advanced parametric workflows are available for Two-Way ANOVA, Repeated Measures ANOVA, and Mixed ANOVA.
- Nonparametric fallbacks for advanced ANOVA designs are fully implemented: Friedman test (Repeated Measures fallback), Freedman-Lane permutation test (Two-Way ANOVA fallback), and Brunner-Langer ATS (Mixed ANOVA fallback), each with appropriate post-hoc comparisons.
- ANCOVA (One-Way and Two-Way) is supported when continuous covariates are placed in the Covariates bucket alongside a categorical Factor 1.
- Linear Mixed Models (LMM) and Logistic Regression are available for longitudinal and binary outcome designs via the Auto-pilot.
- Correlation (Pearson/Spearman) and linear regression (OLS) are supported via the Auto-pilot when a continuous variable is assigned to the Factor 1 bucket.
- Exploratory correlation matrices are available via Analysis -> Exploratory Correlation Matrix.
- Multi-Dataset Analysis takes two or more measurement columns under one factor mapping and analyses them in sequence. Because that is a family of simultaneous tests, Benjamini-Hochberg FDR correction is applied across the p-values and both the adjusted values and the family size are shown in the combined report. The family covers the columns that produced a p-value. With fewer than two, no correction is applied, and the report leaves out the FDR note rather than showing an adjustment that did not happen.
- For two independent groups the pipeline always uses Welch's t-test rather than Student's, and for one-way designs Welch's ANOVA. This is deliberate: Welch is valid whether or not variances are equal, so the choice does not depend on a variance pre-test that is itself unreliable at small n.
- For Windows and macOS end users, the recommended path is to use the packaged application from the GitHub Releases page. The repository launcher scripts are mainly intended for source-based usage.
A brief overview of the repository layout:
BioMedStatX/
├─ README.md # Landing page (this file)
├─ LICENSE # MIT License
├─ CONTRIBUTING.md # Detailed contributing guidelines
├─ CODE_OF_CONDUCT.md # Contributor Covenant Code of Conduct
├─ start.sh # Launcher for Linux/macOS source/binary startup
├─ run.bat # Launcher for Windows source/binary startup
├─ src/ # Main application source code
├─ tests/ # Unit and regression tests
├─ validation/ # Numerical validation against R and published results
├─ fuzzing/ # Fuzzers that drive the real pipeline and check its output
├─ tools/ # Consistency validator and maintenance scripts
├─ docs/ # User-facing documentation
│ ├─ HowTo.md # Screenshot-based user guide (GUI)
│ ├─ ADVANCED_ANOVA_GUIDE.md # Advanced ANOVA explanations
│ └─ CORRELATION_REGRESSION_GUIDE.md # Correlation & regression methodology
└─ .github/
└─ ISSUE_TEMPLATE/
├─ bug_report.yml # Bug report issue template
└─ feature_request.yml # Feature request issue template
BioMedStatX is developed as an open-source academic project.
We welcome bug reports, feature requests, and code contributions.
- Contribution workflow and coding guidelines: CONTRIBUTING.md
- Community rules: CODE_OF_CONDUCT.md
- License: LICENSE
Please use the structured issue templates on GitHub:
This will guide you through our predefined templates for:
- Bug reports
- Feature requests
- (Any additional templates you may add in the future)
Using the templates helps us reproduce issues more easily and keep the project maintainable.
If you plan to contribute code:
- Fork the repository and create a feature branch.
- Implement your changes and add tests where applicable.
- Ensure code style and formatting follow the guidelines.
- Open a Pull Request (PR) with a clear description of your changes.
For full details, including branch naming conventions, commit message guidelines, testing expectations, and how to add new statistical functions, please read:
BioMedStatX is released under the MIT License.
-> See the full license text in LICENSE.
If you use BioMedStatX in a scientific publication, please cite the software repository or the published paper once a formal citation is available. This section will be updated when the final reference and DOI are available.
A direct publication link will be added here once the corresponding paper is publicly available.
For questions regarding the software, collaboration requests, or feedback, you can reach the maintainer at:
- Email: pkrumm@ukaachen.de
- GitHub: @philippkrumm
Please use GitHub Issues for bug reports and feature requests so that the discussion remains transparent and searchable.
These are features worth adding when resources allow. Contributions are welcome.
-> Contributing & Issue Reporting
- Improve support for running from source on Linux, especially for workflows that depend on Excel-specific behavior.