Distributed vs. centralized ML for stroke prediction using Apache Spark. Benchmarks a per-partition scikit-learn ensemble (mapPartitions) against MLlib Random Forest, Logistic Regression, and Decision Tree models. Includes speed-up, size-up, and scale-up scalability analysis on a 43K-row clinical dataset.
python data-science machine-learning big-data apache-spark random-forest scikit-learn distributed-computing jupyter-notebook pyspark healthcare classification class-imbalance smote spark-mllib feature-importance scalability-analysis stroke-prediction feature-importance-analysis mappartitions
-
Updated
Jun 25, 2026 - Jupyter Notebook