Skip to content

About

Find the files to refactor first: hotspots (change frequency × complexity), bus factor and change coupling from your git history. One npx command, offline, any language.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

code-hotspots

English | ภาษาไทย

Point it at a git repository and it tells you which files to refactor first, who is the only person who understands them, and which files secretly depend on each other. One command, no setup, nothing leaves your machine.

CI npm License: MIT

Treemap of the express repository: box size is lines of code, red boxes are hotspots

The whole history of expressjs/express. Size is lines of code, colour is the hotspot score, the numbers are the top five.

cd your-repo
npx code-hotspots
npx code-hotspots -o report.html   # interactive treemap, open it in a browser

What it looks like

$ npx code-hotspots ~/src/express --all --top 5
express  2009-06-26 to 2026-09-29 · 1718 commits · 240 authors (12 active) · 195 files · 18,349 lines

Hotspots  changed often and complex: refactor these first
  #  File                  Commits  Lines  Complexity  Score  Main author
  1  lib/response.js           394    907         816   1.00  Tj Holowaychuk 52%
  2  test/app.router.js         91    940        3009   0.85  Douglas Christopher Wilson 34%
  3  test/res.sendFile.js       70    749        3172   0.69  Douglas Christopher Wilson 79%
  4  test/res.send.js           70    501        1526   0.33  Douglas Christopher Wilson 40%
  5  lib/application.js        184    530         417   0.24  Tj Holowaychuk 53%

Change coupling  files that keep changing together
  File                                  File                              Shared  Degree
  examples/…/views/users/index.ejs  <>  examples/…/views/users/view.ejs        6    100%
  examples/…/views/posts/index.ejs  <>  examples/…/views/users/index.ejs       5     91%
  examples/…/views/posts/index.ejs  <>  examples/…/views/users/view.ejs        5     91%
  examples/…/views/users/edit.ejs   <>  examples/…/views/users/index.ejs       6     86%
  examples/…/views/users/edit.ejs   <>  examples/…/views/users/view.ejs        6     86%

Knowledge islands  one person made 80%+ of the changes
  File                        Lines  Owner                       Share
  test/express.urlencoded.js    701  Douglas Christopher Wilson    96%  inactive
  test/express.static.js        692  Douglas Christopher Wilson    98%  inactive
  test/express.json.js          643  Douglas Christopher Wilson    97%  inactive
  test/express.text.js          478  Douglas Christopher Wilson    95%  inactive
  test/express.raw.js           432  Douglas Christopher Wilson    94%  inactive

Bus factor by directory  fewest people who made over half of the changes
  Directory  Commits  Bus factor  Top authors
  lib/           962           1  Tj Holowaychuk 60%, Douglas Christopher Wilson 21%, Jonathan Ong 4%
  test/          753           2  Douglas Christopher Wilson 47%, Tj Holowaychuk 26%, Jonathan Ong 3%
  examples/      288           1  Tj Holowaychuk 62%, Douglas Christopher Wilson 10%, Jamie Barton 4%
  .github/        81           3  Chris de Almeida 26%, Jon Church 21%, Ulises Gascón 20%
  ./              14           2  Douglas Christopher Wilson 47%, Yuta Hiroto 16%, Tj Holowaychuk 13%

Whole repository bus factor: 2. Analyzed in 1.1s.

Read it like this: lib/response.js is where most of the work in express has happened and it is not simple code, so a bug or a slow review there costs the most. Five large test files were written almost entirely by someone who no longer commits. If one of them breaks, nobody on the current team wrote it.

Why

Every codebase has more debt than anyone has time to pay. The usual way to pick what to fix is gut feeling, or a static analysis tool that lists thousands of warnings with no sense of which ones matter.

Your version control already knows what matters. Complex code that nobody touches costs little. Complex code that the team changes every week is where bugs, merge conflicts and slow reviews come from. Adam Tornhill's book Your Code as a Crime Scene shows how to find that code from the git log. This tool packages the core of those ideas into a single command that works on any language, so a tech lead can bring real numbers to a planning meeting instead of opinions.

Install

Nothing to install, npx code-hotspots runs the latest version. To keep it around:

npm install -g code-hotspots

Requires Node.js 20 or newer and git on your PATH. There are no runtime dependencies.

Usage

code-hotspots [path] [options]

path can be the repository or any directory inside it; only that directory is analyzed and paths are shown relative to it. That is handy in a monorepo: code-hotspots packages/api.

Option Default What it does
--since <when> 12 months Period to analyze. 6m, 2y, 2025-01-01, "3 weeks", anything git log --since takes
--all Whole history
-o, --output <file> Write a report. .html, .json, .csv or .svg, repeatable
--format <f> text What to print: text, json or csv
--top <n> 10 Rows per table in the terminal
--exclude <glob> Ignore more files, gitignore syntax, repeatable
--no-default-ignores Also analyze lockfiles, build output, docs and data files
--include-bots Keep commits by dependabot, renovate and other bots
--depth <n> 1 Directory depth for the bus factor table
--min-shared <n> 5 Coupling: minimum commits two files must share
--min-coupling <pct> 30 Coupling: minimum degree
--max-commit-size <n> 30 Coupling: skip commits that touch more files
--island <pct> 80 Knowledge island: main author's share of the changes
--inactive <months> 6 An author with no commit for this long counts as inactive

Twelve months is the default on purpose. The question is usually "what hurts now", and code that was a hotspot in 2019 may be stable today. Use --all for the long view.

Reports

code-hotspots -o hotspots.html -o hotspots.json

The HTML report is a single file with no external requests, so it works offline and can be attached to a ticket or a slide deck. It has a zoomable treemap (click a directory to zoom in, Esc to go back), colour modes for hotspots, change frequency, complexity, ownership and inactive owners, and sortable, filterable tables for every metric.

--format json prints the whole report for scripts:

npx code-hotspots --format json | jq -r '.files[:5][] | "\(.score)\t\(.path)"'

Ignoring files

Lockfiles, dependency manifests, build output, vendored code, minified and generated files, snapshots, prose (*.md) and data files (*.csv, *.svg) are skipped by default because they change for reasons that have nothing to do with code quality. Binary files and files over 2 MB are always skipped.

Add your own rules with --exclude or a .hotspotsignore file in the analyzed directory. Both use gitignore syntax, including ! to re-include:

# .hotspotsignore
src/generated/
*.pb.ts
!src/generated/keep-this.ts

As a library

import { analyze } from 'code-hotspots'

const report = await analyze('.', { since: '6 months ago' })
console.log(report.files.slice(0, 5).map((f) => `${f.score} ${f.path}`))

analyze takes the same settings as the CLI in camelCase (minShared, islandShare as 0..1, and so on). Note that its since is unset by default, which means the whole history; the 12 month default belongs to the CLI. renderHtml, renderSvg, renderText and renderCsv turn a report into the other formats.

How it works

Everything comes from two sources: git log --numstat -M for the history, and the current content of each tracked file.

Change frequency is the number of commits that touched a file in the period. Renames are followed, so a file keeps its history after git mv. Merge commits are skipped because they repeat changes already counted.

Complexity is the sum of indentation levels over all non-blank lines. It works for any language because nesting shows up as whitespace, and Hindle, Godfrey and Holt (2008) found it correlates well with traditional complexity metrics, which is all a ranking needs. The tool detects whether a file uses 2 or 4 spaces and counts tabs as one level. It is a ranking signal, not a quality grade.

Hotspot score is normalized change frequency times normalized complexity, scaled so the top file is 1.00. A score of 0.50 means "half as much of a hotspot as the worst file in this repository". Scores are not comparable between repositories.

Main author and knowledge islands. For each file, the tool adds up the lines each author added or removed in the period. The main author is whoever changed the most. A file is a knowledge island when the main author made 80% or more of the changes. If that person has not committed anything for 6 months, they are flagged as inactive. Line counts reward big diffs over careful small ones, so treat this as "who has the most context", not "who did the most work".

Bus factor for a directory is the smallest number of authors who together made more than half of its changes. A bus factor of 1 means one person did most of the work there.

Change coupling looks for files that keep showing up in the same commits. The degree is the number of shared commits divided by the average number of commits of the two files. Pairs need at least 5 shared commits and a 30% degree to show up, and commits that touch more than 30 files (formatting runs, dependency bumps, renames) are ignored because they couple everything with everything. High coupling between files in different modules usually means a missing abstraction or a boundary that only exists on paper.

Authors are matched through .mailmap first. On top of that, names that differ only in case ("TJ Holowaychuk" and "Tj Holowaychuk") and commits that share a real email address are treated as one person. Placeholder addresses like you@example.com are never used to merge people. Commits by bots are dropped unless you pass --include-bots.

Speed

Measured on an Apple M2 with node dist/cli.js <repo> --format json:

Repository Period Commits analyzed Files Time
expressjs/express 12 months 27 195 0.1 s
expressjs/express whole history 1,718 195 1.2 s
pallets/flask whole history 1,752 124 1.3 s
nestjs/nest 12 months 549 2,322 1.5 s
nestjs/nest whole history 4,121 2,322 12 s

On nest's full history, about 10 of the 12 seconds are git log itself working out renames across 22,000 commits. The analysis on top is around a second.

Treemap of the nest repository

The whole history of nestjs/nest, 2,322 files. packages/core/injector/injector.ts and the Fastify adapter stand out right away.

Limitations and FAQ

Is a high score bad code? No. It means the code is both complex and busy, which makes it the most expensive place for a problem to be. Some hotspots are fine, like a big router table that everyone appends to. The point is to know where to look first.

Tests and examples show up as hotspots. They do, because they are code you maintain. Exclude them with --exclude 'test/' if you want to focus on production code, or keep them: a test file that needs constant edits is often a sign of a brittle design underneath.

Shallow clones. CI checkouts usually have a single commit, so every number is too low. The tool warns when it sees a shallow clone. Fetch the history first with git fetch --unshallow, or use fetch-depth: 0 in actions/checkout.

Squash merges turn a whole pull request into one commit. Change frequency still works, but authorship goes to whoever merged and coupling gets stronger than it really is.

Generated files you commit (API clients, protobuf output) will rank high. Put them in .hotspotsignore.

Moved code is not tracked across files. If you split a big file into three, the new files start with a fresh history from the commit that created them.

Very large histories are read into memory in one go. Repositories with hundreds of thousands of commits work but take a while; use --since to keep it fast.

Acknowledgements

The ideas here come from Adam Tornhill: the book Your Code as a Crime Scene (and Software Design X-Rays), his open source tool code-maat, and the commercial product CodeScene, which goes much further than this tool does. If you find this useful, read the book. No code was taken from those projects. The treemap layout is the squarified algorithm by Bruls, Huizing and van Wijk (2000).

Contributing

Bug reports and pull requests are welcome, in English or Thai. See CONTRIBUTING.md.

License

MIT

About

Find the files to refactor first: hotspots (change frequency × complexity), bus factor and change coupling from your git history. One npx command, offline, any language.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages