A beginner-friendly Quantitative UX / Customer Insights portfolio project that turns recurring friction in real mobile-app reviews into prioritised actions for a product team.
Read the project story on Medium: What 13,928 Daylio Reviews Taught Me About Product Friction
FrictionMap analyses Google Play reviews for Daylio to identify which product-experience problems are both common and severe, then ranks the top opportunities for product action.
Product teams cannot read thousands of reviews one by one. FrictionMap combines qualitative feedback with quantitative signals:
- Explore the structure and quality of the data (EDA).
- Identify low-rated reviews and explicit problem or feature-request signals.
- Assign one or more friction categories with simple, inspectable rules.
- Combine frequency, severity, and helpfulness in a transparent priority score.
- Sample the automated coding for a blinded second-pass review.
This first version deliberately avoids machine learning, topic modelling, large-scale LLM coding, and dashboards. The goal is to establish sound research reasoning and practical Python fundamentals first.
Which recurring UX frictions in Daylio's user reviews carry the strongest dissatisfaction signals, and which three areas should the product team address first?
Supporting questions:
- Which types of friction appear most often?
- Which frictions are associated with lower ratings?
- Which complaints receive the strongest helpfulness signal from other users?
- Which examples do the automated rules miss or misclassify?
See the complete project brief.
The primary source is the MHARD — Mental Health App Reviews Dataset. It contains 200,972 reviews from 73 Google Play apps, including rating, date, likes, and developer-response fields. This case study uses the 13,928 Daylio reviews only. The source is MIT licensed. No health or clinical outcomes are inferred.
The repository includes the source project's public 100-row sample for a technical smoke test. Portfolio findings require the full dataset filtered to Daylio. See dataset options and selection rationale.
The full raw dataset and row-level analysis exports are intentionally excluded from version control. They can be regenerated locally with the documented commands. This keeps the repository lightweight and avoids republishing thousands of complete review texts.
From the project directory:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"On Windows PowerShell, use .venv\Scripts\Activate.ps1 for the activation step.
python scripts/download_data.py --sample
frictionmap --input data/sample/mhard_sample.csv --output outputs/sampleThis run confirms that the analysis pipeline works. The sample is too small and covers multiple apps, so it must not be used for product decisions.
python scripts/download_data.py --full
frictionmap \
--input data/raw/mhard_full.csv \
--app daylio \
--output outputs/daylioCore outputs:
coded_reviews.csv: each review and its assigned friction codesfriction_summary.csv: category-level priority tablemanual_validation_sample.csv: 50 reviews for blind human codingmanual_validation_key.csv: automated-code key; keep closed until human coding is completeblind_validation_sample.csv: refreshed 50-review set excluding assisted training examplesblind_validation_key.csv: hidden automatic-code key for the refreshed setassistant_validation_sample.csv: completed blinded assistant reviewassistant_validation_metrics.json: completion, agreement, and codebook-coverage resultsassistant_validation_disagreements.csv: cases to investigate in v0.2data_quality.json: row count, date range, and missing-data summaryrating_distribution.pngandfriction_priority.png: initial portfolio chartsfinal_priorities.csvandfinal_priority_areas.png: three product decision areas
The v0.1 analysis identifies three product decision areas:
| Priority | Product area | Unique reviews | Mean rating |
|---|---|---|---|
| 1 | Pricing and paywall | 583 | 2.849 |
| 2 | Feature gaps | 1,048 | 4.031 |
| 3 | Core reliability | 382 | 3.073 |
Pricing and feature gaps are the two robust leading signals. Core reliability groups three nearly tied categories—notifications, stability/performance, and access/account—rather than claiming a precise ordering that the scores do not support. A blinded assistant second pass achieved 70% exact agreement and 80% codebook coverage, meeting both predeclared v0.1 thresholds. This is disclosed as assistant review, not human validation.
Read the full Daylio case study.
Each category receives a score from 0 to 100:
priority = 100 × (
0.50 × frequency
+ 0.40 × severity
+ 0.10 × helpfulness
)
frequency: share of all friction candidates assigned to the categoryseverity: mean rating transformation where one star is most severe and five stars least severehelpfulness: normalised mean of log-transformed likes
This score is not a measure of actual business impact. It is a transparent decision aid built from the available public signals. Helpfulness receives the smallest weight because app-store likes are affected by review visibility and category sample size. If internal data such as revenue loss, task completion, retention, or support volume becomes available, the model should be revised.
A category is marked actionable only when it has at least 30 coded reviews and is not other_unclear. Smaller categories remain visible as signals but are treated as insufficient evidence for a portfolio recommendation.
notebooks/01_eda.ipynb: load a table; inspect rows, columns, missingness, filters, and groupssrc/frictionmap/taxonomy.py: work with text rules, functions, lists, and conditionssrc/frictionmap/priority.py: usegroupby, ratios, normalisation, and weighted scoressrc/frictionmap/pipeline.py: combine small functions into one reproducible workflowtests/: verify expected behaviour with small examples
After completing blind coding, measure agreement:
python scripts/score_validation.py \
--reviewer outputs/daylio/assistant_validation_sample.csv \
--key outputs/daylio/blind_validation_key.csvRead the friction codebook and applied learning path.
frictionmap/
├── data/ # Source sample, raw data, and optional intermediate data
├── docs/ # Research scope, dataset decision, and codebook
├── notebooks/ # Guided exploratory analysis
├── outputs/ # Reproducible tables and charts
├── scripts/ # Data download and validation helpers
├── src/frictionmap/ # Analysis package
└── tests/ # Behaviour checks
- Usernames are excluded from the analysis and exported outputs.
- Reviews are not interpreted as evidence of diagnosis, clinical status, or treatment effectiveness.
- The portfolio report should not reproduce long or identifying quotations.
- Public app-store reviewers are not representative of all users; self-selection bias must be reported.
Wang, Q., Erqsous, M., Khatiwada, P., Karwankar, A., Alhassan, F. M., Chandrasekaran, A., Abraham, B., Lovell, F., Ngo, A. A., & Mauriello, M. L. (2025). Leveraging Large Language Models for Review Classification and Rating Estimation of Mental Health Applications. Proceedings of the International AAAI Conference on Web and Social Media, 19(1), 2017–2029. https://doi.org/10.1609/icwsm.v19i1.35916

