Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 41 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,29 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.2.0] - 2026-07-08

This release adds a static code-analysis toolkit on top of the tokenizer:
diff/PR analysis, complexity metrics, n-gram "naturalness", and clone/plagiarism
similarity — plus a comprehensive per-language test suite that hardened comment
handling across the board.

### Added
- **Fingerprinting & similarity** (`PyReprism.fingerprints`): winnowing k-gram
fingerprints over the normalized token stream for clone/plagiarism detection —
`fingerprint()`, `similarity()`/`containment()` (rename-invariant by default),
and a `FingerprintIndex` for many-to-many detection over a corpus. CLI:
`pyreprism similarity a b` and `pyreprism clones DIR --threshold`.
- **N-gram analysis & code naturalness** (`PyReprism.ngrams`): token/type n-gram
extraction and frequency counts, plus an `NgramModel` (add-k smoothing,
save/load) that measures cross-entropy / perplexity against a trained corpus
("naturalness of software"). CLI: `pyreprism ngrams` and
`pyreprism perplexity --train`.
- **Complexity metrics** computed from the token stream: `halstead()`
(volume/difficulty/effort/bugs), `cyclomatic_complexity()` (approximate McCabe),
`maintainability_index()` (0–100), `max_nesting_depth()`, and `code_metrics()`
which bundles them with the line/token stats. Exposed on the CLI via
`pyreprism stats --full`.
- **Diff processing** (`PyReprism.diffs`): parse unified/`git` diffs and analyze
the changed code per file in its own language. Includes churn metrics
(`diff_stats`: added/removed split into code, comment and blank), cosmetic
Expand All @@ -17,6 +39,23 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
`old_source` enables accurate full-file classification. New CLI command
`pyreprism diff` (`--json` / `--csv` / `--per-file` / `--cosmetic`).

### Changed
- Adopted the standard `src/` layout with tests at the top level, consolidated
all dependencies into `pyproject.toml` extras, and rewrote the landing page.
- Single-sourced the version from the package and automated releases: pushing a
`vX.Y.Z` tag now publishes to PyPI (Trusted Publishing) and creates a GitHub
Release. Added Python 3.13 to the CI matrix.
- Added a comprehensive, data-driven per-language comment test suite that
requires every registered language to be covered.

### Fixed
- Fixed comment stripping in 8 more languages surfaced by the new test suite:
`smalltalk` (kept the comment and deleted the code) and
`eiffel`/`bro`/`coffeescript`/`io`/`nix`/`gherkin`/`gedcom` (trailing comments
not removed and newlines dropped).
- Fixed a CLI argument-ordering incompatibility on Python ≤ 3.11 (options must
follow positionals).

## [0.1.0] - 2026-07-07

This release turns PyReprism from a comment-removal helper into a full
Expand Down Expand Up @@ -73,6 +112,7 @@ ML-oriented normalization, an optional accurate backend, and batch processing.
- Early beta releases: comment removal for an initial set of languages and the
`Normalizer` whitespace helper.

[Unreleased]: https://github.com/unlv-evol/PyReprism/compare/v0.1.0...HEAD
[Unreleased]: https://github.com/unlv-evol/PyReprism/compare/v0.2.0...HEAD
[0.2.0]: https://github.com/unlv-evol/PyReprism/compare/v0.1.0...v0.2.0
[0.1.0]: https://github.com/unlv-evol/PyReprism/releases/tag/v0.1.0
[0.0.4]: https://github.com/unlv-evol/PyReprism/releases/tag/v0.0.4
50 changes: 50 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,56 @@ s.comment_to_code_ratio, s.comment_density
s.as_dict() # includes per-token-type counts, ready for JSON / dataframes
```

Complexity metrics (computed from the token stream):

```python
pr.halstead(source, lang="python").volume # Halstead volume/difficulty/effort/bugs
pr.cyclomatic_complexity(source, lang="python") # approximate McCabe complexity
pr.maintainability_index(source, lang="python") # 0–100 (higher is better)
pr.code_metrics(source, lang="python") # everything above in one dict
```

On the CLI: `pyreprism stats --full file.py` (add `--json` for machine output).

### Similarity & clone / plagiarism detection

Winnowing k-gram fingerprints over the *normalized* token stream, so matches
survive variable renaming and literal changes (Type-2 clones):

```python
from PyReprism import fingerprints as fp

fp.similarity(code_a, code_b, "python") # 0.0–1.0 (Jaccard of fingerprints)
fp.containment(code_a, code_b, "python") # how much of A appears in B

index = fp.FingerprintIndex() # many-to-many clone detection
index.add_paths("submissions/")
index.similar_pairs(threshold=0.7) # -> [(file_a, file_b, score), ...]
```

On the CLI: `pyreprism similarity a.py b.py` and
`pyreprism clones submissions/ --threshold 0.7`.

### N-grams & code "naturalness"

Token n-grams (over token text or, structurally, over token *types*) and an
n-gram language model that measures how predictable/"natural" code is
(Hindle et al.):

```python
from PyReprism import ngrams

ngrams.ngram_counts(source, "python", n=3).most_common(10)
ngrams.ngrams(source, "python", n=2, types=True) # structural n-grams

model = ngrams.train("corpus/", n=3) # train on a code corpus
model.perplexity(ngrams.token_sequence(source, "python")) # lower = more natural
model.save("model.json")
```

On the CLI: `pyreprism ngrams file.py -n 3 --top 20` and
`pyreprism perplexity --train corpus/ file.py`.

### Normalization for ML / clone detection

Canonicalize code so that only its structure remains — rename identifiers to
Expand Down
27 changes: 27 additions & 0 deletions docs/fingerprints.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
.. _fingerprints_toplevel:

=========================================
Fingerprinting & similarity
=========================================

The :mod:`PyReprism.fingerprints` module detects clones and plagiarism using
**winnowing** k-gram fingerprints (Schleimer, Wilkerson & Aiken) over the
*normalized* token stream, so matches are robust to variable renaming and
literal changes (Type-2 clones)::

from PyReprism import fingerprints as fp

fp.similarity(code_a, code_b, "python") # Jaccard of fingerprints, 0-1
fp.containment(code_a, code_b, "python") # fraction of A found in B

index = fp.FingerprintIndex()
index.add_paths("submissions/")
index.similar_pairs(threshold=0.7)

Fingerprints are deterministic (CRC32-hashed) and comparable across runs.

.. automodule:: PyReprism.fingerprints
:members:
:undoc-members:
:show-inheritance:
:exclude-members: __dict__, __weakref__
2 changes: 2 additions & 0 deletions docs/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@ PyReprism documentation!
cli
tokens
metrics
ngrams
fingerprints
engines
batch
diffs
Expand Down
109 changes: 78 additions & 31 deletions docs/intro.rst
Original file line number Diff line number Diff line change
@@ -1,53 +1,100 @@
.. _intro_toplevel:

==================
Overview / Install
==================

PyReprism is a Python framework that helps researchers and developers the task of source code preprocessing. With PyReprism, you can easily match, extract, count, and remove comments, whitespaces, operators, numbers and other language specific constructs from over 150 programming languages and file extensions.

========
Overview
========

**PyReprism** is a Python framework for source-code preprocessing. It lets you
**match, extract, count, and remove** comments, strings, numbers, operators,
keywords and other language-specific constructs across **145+ programming
languages and file formats** — through a small high-level API, a command-line
tool, code metrics, ML-oriented normalization, and diff/pull-request analysis.

Highlights
==========

* One consistent API for every language: ``remove_*`` / ``extract_*`` /
``count_*`` / ``match_*`` for each construct, plus a lossless ``tokenize()``.
* Code metrics (``stats``) and ML canonicalization (``normalize``) for
clone/plagiarism detection and code-embedding pipelines.
* Optional high-accuracy `Pygments`_ backend — with **zero required
dependencies** by default.
* Batch/corpus analysis and unified/``git`` diff processing.
* A ``pyreprism`` command-line tool (also ``python -m PyReprism``).

Requirements
============

* `Python`_ 3.6 or newer
* `Git`_
* `Python`_ 3.8 or newer

.. _Python: https://www.python.org
.. _Git: https://git-scm.com/
Install
=======

Installing PyReprism
====================
.. code-block:: shell

Installing PyReprism is easily done using `pip`_. Assuming it is installed, just run the following from the command-line:
pip install PyReprism

.. _pip: https://pip.pypa.io/en/latest/installing.html
For the higher-accuracy tokenizer backend:

.. sourcecode:: none
.. code-block:: shell

$ pip install PyReprism
pip install "PyReprism[accurate]"

Source Code
Quick start
===========

PyReprism's git repo is available on GitHub, which can be browsed at:
.. code-block:: python

import PyReprism as pr

source = """
# a comment
x = 5 + 6
print(x) # inline
"""

pr.remove_comments(source, lang="python") # code without comments
pr.extract_comments(source, lang="python") # ['# a comment', '# inline']
pr.count_comments(source, lang="python") # 2

``lang`` accepts a language name (``"python"``), a file extension (``".py"``),
or a language class. Canonicalize code for machine learning, or measure it:

.. code-block:: python

* https://github.com/unlv-evol/PyReprism.git
pr.normalize("total = price * 42", lang="python") # 'VAR1 = VAR2 * 0'
pr.stats(source, lang="python").comment_to_code_ratio

and cloned using::
Command line
============

.. code-block:: shell

pyreprism remove comments file.py
pyreprism scan myproject/ --csv # aggregate code metrics over a tree
git diff | pyreprism diff --cosmetic # flag comment/whitespace-only changes

Documentation
=============

$ git clone https://github.com/unlv-evol/PyReprism.git
$ cd PyReprism
Full documentation, including the API reference and the list of supported
languages, is available at https://pyreprism.readthedocs.io.

Optionally (but suggested), make use of virtual environment. Therefore, before installing the requirements run::

$ python3 -m venv venv
$ source venv/bin/activate
Source code
===========

PyReprism is developed on GitHub at
https://github.com/unlv-evol/PyReprism. To work on it locally:

Install the requirements::

$ pip install -r requirements.txt
.. code-block:: shell

and run the tests using pytest::
git clone https://github.com/unlv-evol/PyReprism.git
cd PyReprism
python3 -m venv venv && source venv/bin/activate
pip install -e ".[dev]"
pytest

$ pytest
See ``CONTRIBUTING.md`` for contribution guidelines.

.. _Python: https://www.python.org
.. _Pygments: https://pygments.org
28 changes: 28 additions & 0 deletions docs/ngrams.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
.. _ngrams_toplevel:

==============================
N-grams & code naturalness
==============================

The :mod:`PyReprism.ngrams` module extracts token n-grams and provides an
n-gram language model for measuring code *naturalness* (cross-entropy /
perplexity), following Hindle et al., "On the Naturalness of Software".

n-grams can be taken over token **text** or over token **types** (structural,
AST-free), which is useful for clone detection and style analysis::

from PyReprism import ngrams

ngrams.ngram_counts(source, "python", n=3).most_common(10)
ngrams.ngrams(source, "python", n=2, types=True)

model = ngrams.train("corpus/", n=3)
seq = ngrams.token_sequence(source, "python")
model.perplexity(seq) # lower = more predictable / "natural"
model.save("model.json")

.. automodule:: PyReprism.ngrams
:members:
:undoc-members:
:show-inheritance:
:exclude-members: __dict__, __weakref__
47 changes: 44 additions & 3 deletions src/PyReprism/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,11 @@
import os
from typing import List, Optional, Sequence, Type, Union

from .metrics import CodeStats
from .metrics import CodeStats, Halstead
from .tokens import Token, TokenType
from .utils.normalizer import Normalizer

__version__ = "0.1.0"
__version__ = "0.2.0"

LanguageLike = Union[str, "type"]

Expand Down Expand Up @@ -274,6 +274,46 @@ def stats(source: str, lang: LanguageLike, engine: str = 'regex') -> CodeStats:
return _tokenops.stats(_engine_tokens(source, lang, engine), source)


def halstead(source: str, lang: LanguageLike, engine: str = 'regex') -> Halstead:
"""Return the :class:`Halstead` complexity measures for ``source``."""
if _is_regex(engine):
return _resolve(lang).halstead(source)
from . import _tokenops
return _tokenops.halstead(_engine_tokens(source, lang, engine))


def cyclomatic_complexity(source: str, lang: LanguageLike, engine: str = 'regex') -> int:
"""Approximate McCabe cyclomatic complexity (token-based)."""
if _is_regex(engine):
return _resolve(lang).cyclomatic_complexity(source)
from . import _tokenops
return _tokenops.cyclomatic(_engine_tokens(source, lang, engine))


def maintainability_index(source: str, lang: LanguageLike, engine: str = 'regex') -> float:
"""SEI-normalized Maintainability Index in ``[0, 100]`` (higher is better)."""
if _is_regex(engine):
return _resolve(lang).maintainability_index(source)
from . import _tokenops
tokens = _engine_tokens(source, lang, engine)
return _tokenops.maintainability_index(tokens, _tokenops.stats(tokens, source).code_lines)


def code_metrics(source: str, lang: LanguageLike, engine: str = 'regex') -> dict:
"""Return a combined metrics dict (line stats + Halstead + complexity + MI)."""
if _is_regex(engine):
return _resolve(lang).code_metrics(source)
from . import _tokenops
tokens = _engine_tokens(source, lang, engine)
stats_obj = _tokenops.stats(tokens, source)
data = stats_obj.as_dict()
data['halstead'] = _tokenops.halstead(tokens).as_dict()
data['cyclomatic_complexity'] = _tokenops.cyclomatic(tokens)
data['max_nesting_depth'] = _tokenops.max_nesting_depth(tokens)
data['maintainability_index'] = _tokenops.maintainability_index(tokens, stats_obj.code_lines)
return data


def normalize(source: str, lang: LanguageLike, engine: str = 'regex', **options) -> str:
"""Return a canonicalized form of ``source`` for ML / clone detection.

Expand Down Expand Up @@ -329,7 +369,7 @@ def preprocess(source: str, lang: LanguageLike, steps: Sequence[str] = ('comment

__all__ = [
'__version__',
'Token', 'TokenType', 'CodeStats', 'Normalizer',
'Token', 'TokenType', 'CodeStats', 'Halstead', 'Normalizer',
'get_language', 'detect_language',
'remove_comments', 'extract_comments', 'count_comments', 'match_comments',
'remove_keywords', 'extract_keywords', 'count_keywords',
Expand All @@ -338,5 +378,6 @@ def preprocess(source: str, lang: LanguageLike, steps: Sequence[str] = ('comment
'remove_strings', 'extract_strings',
'extract_identifiers', 'remove_whitespaces',
'blank_comments', 'stats', 'normalize',
'halstead', 'cyclomatic_complexity', 'maintainability_index', 'code_metrics',
'tokenize', 'preprocess',
]
Loading
Loading