Skip to content
Merged
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,9 @@ Notable changes to JevGate. Versions follow [Semantic Versioning](https://semver

## [Unreleased]

- File organization: a long application file (400 lines or more) whose outline, recheck and kind of file raised no finding is asked about its candidate parts, one request per part with the part's source: whether it does a job of its own that a reader would look for apart from the rest, and what it is within the file. A part of 100 lines or more whose answer reaches 0.65 and whose role leans to a job of its own is a consider naming its members. The kind of file cleared long single-type files wholesale: of 50 such files labeled from the code, 15 were worth splitting (a URL scraper inside lobsters' `Story`, a diff engine inside a renderer, a JSON parser inside a protocol module), and asking each part found 4 of them, at the part the labeler named, and no file to keep. The parts are the outline's groups without the links that merged every method of a large class into one group and without links through a helper most members call, with links for neighbours and for names sharing a distinctive word. On the corpus, 11 considers were added and nothing else changed: 9 right and 2 wrong (a demo page's placeholder table, a class's public API), 5 of 7 right outside the files used for tuning and 1 of 1 on the held-out projects; on 9 projects never used before (click, rich, zod, hono, viper, ripgrep, sinatra, jsoup, guzzle), 2 of 2 (Guzzle's `WWW-Authenticate` parser inside `DigestAuth`, ripgrep's `--hyperlink-format` language). File-organization considers were 63% right before. Benchmarks, examples and `scripts` and `docs` directories are not asked (5 of 5 such findings were wrong), nor is a part holding `main`. About $0.08 on the corpus. JevGate's own large files stay clear: their parts read as one job, and what decided them by hand, which other files use a part, is missing for Rust functions passed by name.
- A request whose answer is already cached fits the provider limit, whatever the token calibration says. The calibration (`.jevgate/token-budget.json`) is replaced after each run by the bytes per token of that run's fresh requests, and a request near the limit, such as a recheck that sends a whole file, was sent in one run and not the next: the file's finding changed with no change to its code. On the corpus, 0.23.1 under three calibrations (the last run's, a snapshot's and a stricter one) differed in one consider and five undecided units, and under the stricter one asked 42 thousand new tokens of questions; with this, only three units whose recheck the provider refused as too long still differ, and nothing new is asked. `--refresh` skips the cache, so it plans by the estimate alone.

## [0.23.1] - 2026-09-27

- Injection: the follow-up that asks a path finding what its paths can hold now also asks a markup finding what its values hold where they enter the markup (already escaped, percent-encoded or serialized as a URL; typed; the program's own; or raw text) and a redirect finding where its targets can lead (a fixed path such as `/admin` first keeps it on the site; `origin + next` with no slash between them does not). Leaning to harmless values, the finding is a note. On the 38 corpus projects with such findings, vaultwarden's `hibp_breach` (a username percent-encoded before the link), its admin login redirect and shiori's login redirect, all labeled wrong, are notes; the 26 markup and 6 redirect findings labeled right put at most 0.22 and 0.44 on the harmless options and stay. About $0.01 on the corpus.
Expand Down
20 changes: 20 additions & 0 deletions docs/classification-cascade.md
Original file line number Diff line number Diff line change
Expand Up @@ -154,6 +154,26 @@ signatures, or one candidate pair.
whole gets no recheck, so its undecided first answer is asked the kind
from the outline alone; large Java classes and their test files otherwise
stayed uncertain.
An application file of 400 lines or more whose outline, recheck and kind
raised no finding is then asked about its candidate parts, one request
per part: the part's members with their source, the file's other members
by signature, whether the part does a job of its own that a reader would
look for apart from the rest, and what it is within the file (a job of
its own, more of what the rest does, helpers the rest uses throughout, or
the file's main job). A part of 100 lines or more whose Noul reaches 0.65
and whose role leans to a job of its own raises a consider naming its
members. Asked of the whole outline, the split and the kind of file read
a URL scraper inside a Rails model and a diff engine inside a renderer as
one feature: of 50 long files the kind cleared, labeled from the code,
15 were worth splitting, and asking each part found 4 of them, each at the
part the labeler named, and no file to keep. The parts are the outline's
groups without the links a type's members share when it has more than
twelve (those merged every method of a large class into one group),
without links through a helper most members call, and with a link for
neighbours and for names sharing a distinctive word. Benchmarks,
examples, `scripts` and `docs` directories are not asked (a script runs
its steps top to bottom and a benchmark is often pinned by hash: 5 of 5
such findings were wrong), nor is a part holding `main`.
A test left undecided on whether it re-implements the code or checks only
its mocks is asked again with the bodies of the functions it calls and its
file's imports, mocks and setup hooks (a part too long is left out, never
Expand Down
20 changes: 20 additions & 0 deletions site/src/how-it-works.md
Original file line number Diff line number Diff line change
Expand Up @@ -172,6 +172,26 @@ signatures, or one candidate pair.
whole gets no recheck, so its undecided first answer is asked the kind
from the outline alone; large Java classes and their test files otherwise
stayed uncertain.
An application file of 400 lines or more whose outline, recheck and kind
raised no finding is then asked about its candidate parts, one request
per part: the part's members with their source, the file's other members
by signature, whether the part does a job of its own that a reader would
look for apart from the rest, and what it is within the file (a job of
its own, more of what the rest does, helpers the rest uses throughout, or
the file's main job). A part of 100 lines or more whose Noul reaches 0.65
and whose role leans to a job of its own raises a consider naming its
members. Asked of the whole outline, the split and the kind of file read
a URL scraper inside a Rails model and a diff engine inside a renderer as
one feature: of 50 long files the kind cleared, labeled from the code,
15 were worth splitting, and asking each part found 4 of them, each at the
part the labeler named, and no file to keep. The parts are the outline's
groups without the links a type's members share when it has more than
twelve (those merged every method of a large class into one group),
without links through a helper most members call, and with a link for
neighbours and for names sharing a distinctive word. Benchmarks,
examples, `scripts` and `docs` directories are not asked (a script runs
its steps top to bottom and a benchmark is often pinned by hash: 5 of 5
such findings were wrong), nor is a part holding `main`.
A test left undecided on whether it re-implements the code or checks only
its mocks is asked again with the bodies of the functions it calls and its
file's imports, mocks and setup hooks (a part too long is left out, never
Expand Down
194 changes: 194 additions & 0 deletions src/analysis/groups.rs
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,145 @@ pub fn groups(units: &[Unit], members: &[usize], imports: &BTreeSet<String>) ->
.collect()
}

/// Candidate parts of a long file, for the part follow-up: `groups` links
/// every method of a class through their owner, so a scraper inside a
/// model or a codec inside a manager never stood apart. Here methods of a
/// type with more than `PART_OWNER_MEMBERS` members link only through calls
/// and shared names, calls to a helper that more than three in ten members
/// call do not link, and members next to each other or sharing a distinctive
/// word of their names link once more. Scored against the parts that
/// labelers named on 23 files, the best matching part rose from 0.45 to
/// 0.72 (F1 over member lines), and from 5 to 9 of the 15 files to split.
pub fn parts(units: &[Unit], members: &[usize], imports: &BTreeSet<String>) -> Vec<Vec<usize>> {
clusters(part_weights(units, members, imports))
.into_iter()
.map(|set| set.into_iter().map(|m| members[m]).collect())
.collect()
}

/// A type with more members than this holds parts of its own.
const PART_OWNER_MEMBERS: usize = 12;
/// The share of members calling a helper above which calls to it do not link.
const HUB_SHARE: f64 = 0.3;
/// The share of members whose names a word may appear in and still link them.
const DISTINCT_WORD_SHARE: f64 = 0.25;

fn part_weights(units: &[Unit], members: &[usize], imports: &BTreeSet<String>) -> Vec<Vec<u32>> {
let n = members.len();
let unit = |i: usize| &units[members[i]];
let names = linking_names(units, members, imports);
let tallies = Tallies::of(units, members);
let calls = |x: &Unit, y: &Unit| {
!tallies.hub(&y.short_name)
&& (x.calls.contains(&y.short_name)
|| !y.owner.is_empty() && x.calls.contains(&y.owner) && !tallies.hub(&y.owner))
};
let mut weights = vec![vec![0u32; n]; n];
for i in 0..n {
for j in i + 1..n {
let (a, b) = (unit(i), unit(j));
let mut weight = names[i].intersection(&names[j]).count().min(2) as u32;
if calls(a, b) || calls(b, a) {
weight += CALL_WEIGHT;
}
if !a.owner.is_empty() && a.owner == b.owner && !tallies.large(&a.owner) {
weight += OWNER_WEIGHT;
}
weight += u32::from(j == i + 1) + u32::from(tallies.share_word(i, j));
weights[i][j] = weight;
weights[j][i] = weight;
}
}
weights
}

/// What the part links weigh against, counted over a file's members: how
/// many call each name, how many each type owns, and the words of each
/// member's name with how many names hold each word.
struct Tallies<'u> {
members: usize,
called: BTreeMap<&'u str, usize>,
owners: BTreeMap<&'u str, usize>,
words: Vec<BTreeSet<String>>,
spread: BTreeMap<String, usize>,
}

impl<'u> Tallies<'u> {
fn of(units: &'u [Unit], members: &[usize]) -> Self {
let mut tallies = Tallies {
members: members.len(),
called: BTreeMap::new(),
owners: BTreeMap::new(),
words: Vec::with_capacity(members.len()),
spread: BTreeMap::new(),
};
for unit in members.iter().map(|&m| &units[m]) {
for call in &unit.calls {
*tallies.called.entry(call.as_str()).or_default() += 1;
}
if !unit.owner.is_empty() {
*tallies.owners.entry(unit.owner.as_str()).or_default() += 1;
}
let words = name_words(&unit.short_name);
for word in &words {
*tallies.spread.entry(word.clone()).or_default() += 1;
}
tallies.words.push(words);
}
tallies
}

/// A helper that more than `HUB_SHARE` of the members call.
fn hub(&self, name: &str) -> bool {
self.called
.get(name)
.is_some_and(|&c| c as f64 > (self.members as f64 * HUB_SHARE).max(3.0))
}

/// A type with more than `PART_OWNER_MEMBERS` members.
fn large(&self, owner: &str) -> bool {
self.owners
.get(owner)
.is_some_and(|&c| c > PART_OWNER_MEMBERS)
}

/// Whether the names of members `i` and `j` share a word that at most
/// `DISTINCT_WORD_SHARE` of the members' names hold.
fn share_word(&self, i: usize, j: usize) -> bool {
let distinct = (self.members as f64 * DISTINCT_WORD_SHARE).max(2.0);
self.words[i]
.intersection(&self.words[j])
.any(|w| self.spread[w] as f64 <= distinct)
}
}

/// The lowercase words of a name longer than two letters, split at
/// underscores, punctuation and camel-case humps: `fetchedAttributesHtml`
/// and `fetched_attributes_pdf` share `fetched` and `attributes`.
fn name_words(name: &str) -> BTreeSet<String> {
let mut words = BTreeSet::new();
let mut word = String::new();
let mut previous: Option<char> = None;
for c in name.trim_start_matches(['#', '_']).chars() {
let hump =
c.is_uppercase() && previous.is_some_and(|p| p.is_lowercase() || p.is_ascii_digit());
if !c.is_alphanumeric() || hump {
if word.chars().count() > 2 {
words.insert(std::mem::take(&mut word));
}
word.clear();
}
if c.is_alphanumeric() {
word.extend(c.to_lowercase());
}
previous = Some(c);
}
if word.chars().count() > 2 {
words.insert(word);
}
words
}

/// Group a test file's cases and the support code they share: cases link by
/// their innermost suite and by the subjects and helpers they share, and a
/// case that calls a helper links to it. Positions `0..cases.len()` are the
Expand Down Expand Up @@ -307,6 +446,61 @@ mod tests {
}
}

#[test]
fn a_large_type_is_parted_by_its_calls_not_its_owner() {
let scraper = [
"fetched_attributes",
"fetched_html",
"fetched_pdf",
"canonical_target",
];
let mut source = String::from("struct Story { title: String }\nimpl Story {\n");
for i in 0..9 {
source.push_str(&format!(
" fn field{i}(&self) -> usize {{ self.title.len() + {i} }}\n"
));
}
for name in scraper {
let calls: Vec<String> = scraper
.iter()
.filter(|other| **other != name)
.map(|other| format!("self.{other}()"))
.collect();
source.push_str(&format!(
" fn {name}(&self) -> usize {{ {} }}\n",
calls.join(" + ")
));
}
source.push_str("}\n");
let file = super::super::units::parse(Path::new("story.rs"), &source).unwrap();
let members: Vec<usize> = (0..file.units.len()).collect();
let name = |m: &usize| file.units[*m].short_name.as_str();
assert!(
groups(&file.units, &members, &file.imports)
.iter()
.any(|g| g.members.iter().any(|m| name(m) == "field0")
&& g.members.iter().any(|m| name(m) == scraper[0])),
"a shared owner links the scraper to the fields"
);
let parts = parts(&file.units, &members, &file.imports);
let scraping = parts
.iter()
.find(|p| p.iter().any(|m| name(m) == scraper[0]))
.unwrap();
assert_eq!(scraping.iter().map(name).collect::<Vec<_>>(), scraper);
}

#[test]
fn name_words_split_humps_underscores_and_private_marks() {
let words = |name: &str| name_words(name).into_iter().collect::<Vec<_>>();
assert_eq!(
words("fetchedAttributesHtml"),
["attributes", "fetched", "html"]
);
assert_eq!(words("#parse_UTF8_value"), ["parse", "utf8", "value"]);
assert_eq!(words("to"), Vec::<String>::new());
}

#[test]
fn constructing_a_class_links_to_its_methods() {
let source = "export class GatewayError extends Error {\n constructor(code: string) {\n super(code)\n }\n}\nexport function requireLive(at: number) {\n if (Date.now() >= at) throw new GatewayError('EXPIRED')\n}\n";
Expand Down
2 changes: 1 addition & 1 deletion src/catalog.rs
Original file line number Diff line number Diff line change
Expand Up @@ -287,7 +287,7 @@ pub fn rules() -> Vec<Rule> {

pub fn rule_version(key: &str) -> &'static str {
match key {
FILE_ORGANIZATION => "21",
FILE_ORGANIZATION => "22",
FUNCTION_SIMPLIFICATION => "16",
SHARED_LOGIC => "22",
TEST_VALUE => "7",
Expand Down
18 changes: 14 additions & 4 deletions src/evaluate.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ use super::{
options::CheckArgs,
schema::{self, FileResult, Report, Status},
storage::Store,
token_budget::TokenBudget,
token_budget::{Limits, TokenBudget},
transport::Evaluator,
};
use crate::config::ConfigContext;
Expand Down Expand Up @@ -129,13 +129,16 @@ fn empty_report(args: &CheckArgs, current: &SnapshotContext<'_>, files: Vec<File
/// blocks) are not known yet.
fn preview(inputs: &[Input], args: &CheckArgs, root: &std::path::Path, report: &mut Report) {
let budget = &TokenBudget::load(root);
let answered =
|request: &serde_json::Value| crate::requests::answered(root, args, request).is_some();
let limits = Limits::new(budget, &answered);
let mut planned = Vec::new();
let mut views = BTreeMap::new();
for (owner, input) in inputs.iter().enumerate() {
if report.files[owner].status != Status::Pending {
continue;
}
match schedule(input, args, budget, &mut report.files[owner]) {
match schedule(input, args, limits, &mut report.files[owner]) {
Ok(Scheduled::None) => {}
Ok(Scheduled::Purpose(request)) => {
match cached_purpose(input, args, root, &request, &mut report.files[owner]) {
Expand Down Expand Up @@ -218,6 +221,7 @@ impl Session<'_> {
// Traces judge where a security concern's values come from; rechecks
// settle uncertain units; a security check still undecided is asked
// where its URL comes from or its output goes, and an outline its kind;
// a long file left without a finding is asked about its parts;
// locate follow-ups then point split findings at a block, and a
// located value is asked what it is. Each depends on the answers
// before it.
Expand All @@ -227,6 +231,7 @@ impl Session<'_> {
crate::units::rechecks,
crate::units::settles,
crate::units::kinds,
crate::units::parts,
crate::units::locates,
crate::units::value_kinds,
] {
Expand Down Expand Up @@ -261,12 +266,17 @@ impl Session<'_> {
) {
let mut purpose = Vec::new();
let mut views = BTreeMap::new();
let root = &self.context.root;
let answered = |request: &serde_json::Value| {
crate::requests::answered(root, self.args, request).is_some()
};
let limits = Limits::new(&self.budget, &answered);
for (owner, file) in report.files.iter_mut().enumerate() {
if file.status != Status::Pending {
continue;
}
file.judgments.clear();
match schedule(&inputs[owner], self.args, &self.budget, file) {
match schedule(&inputs[owner], self.args, limits, file) {
Ok(Scheduled::None) => file.cached = false,
Ok(Scheduled::Purpose(request)) => {
file.cached = true;
Expand Down Expand Up @@ -549,7 +559,7 @@ fn apply_classification(file: &mut FileResult, class: crate::file_kind::Classifi
fn schedule(
input: &Input,
args: &CheckArgs,
budget: &TokenBudget,
budget: Limits<'_>,
file: &mut FileResult,
) -> Result<Scheduled> {
if file.status != Status::Pending {
Expand Down
Loading
Loading