Script-based malware (PowerShell, JavaScript, VBScript) is a primary attack vector. Scripts are easy to obfuscate, use built-in OS tools, and evade signature-based detection. We need a fast, statistical classifier that detects malicious scripts with high accuracy and low false positive rate, without parsing or executing code.
Rejected approaches and why:
| Approach | Reason Rejected |
|---|---|
| Regex-based features | Slow, language-specific, easily evaded |
| AST parsing | Fails on obfuscated code, requires multiple parsers |
| Sandbox execution | Seconds per file, needs VM infrastructure |
| Word TF-IDF | Vocabulary explosion, misses obfuscation |
| Full byte n-gram histograms | 65K-16M features, too sparse |
| Deep learning | Overkill, slow inference, harder to deploy |
Core insight: Malware obfuscation changes byte-level statistical properties. We don't need to understand code — we measure how "weird" the byte distribution looks.
Round 1 — Regex (46 features): Keyword matching, structure counts, network indicators. 30+ minutes for 20K samples. Slow, evadable, language-specific.
Round 2 — Character TF-IDF (12 + 5000): Replaced word tokens with char n-grams. Still needed sklearn fitting, 5000 features mostly redundant.
Round 3 — Byte Histogram (8 + 256): Operated directly on raw bytes. Fast and language-agnostic. Problem: minified benign JS and obfuscated malware have similar byte frequencies. A WordPress auto-generated file was flagged as malware.
Round 4 — Hashed Byte Bigrams (8 + 1024) [CURRENT]: Byte frequencies alone aren't enough — byte transitions carry language-specific structure. Full bigrams would be 65,536 features. Instead, hash each pair into 1024 bins (hashing trick). Collisions are tolerated; LightGBM disentangles them using other features.
Why this fixes the WordPress false positive: /* and */ (comment markers) hash to specific bins. JavaScript pairs (
va, ar, fu) hash differently than PowerShell pairs (ie, ex, -e). The model learns language-specific byte
transitions.
Statistical Features (8):
| # | Feature | How Computed | Signal |
|---|---|---|---|
| 1 | line_count |
Count \n bytes |
Obfuscated code often 1-5 long lines |
| 2 | max_line_length |
np.diff(newline_positions) |
Long lines = minified/packed |
| 3 | uppercase_ratio |
sum(counts[65:91]) / length |
PowerShell PascalCase cmdlets |
| 4 | digit_ratio |
sum(counts[48:58]) / length |
High in hex/base64 payloads |
| 5 | punctuation_ratio |
sum(counts[PUNCT]) / length |
Elevated in obfuscated code |
| 6 | whitespace_ratio |
sum(counts[SPACE]) / length |
Low in packed, high in readable |
| 7 | entropy |
-sum(p * log2(p)) |
High = encrypted/packed/random |
| 8 | unique_byte_count |
count_nonzero(counts) |
Low unique + high entropy = suspicious |
Hashed Bigram Features (1024):
Each byte pair (b0, b1) is packed into (b0 << 8) | b1 and hashed via pair % 1024. The normalized histogram
captures byte transitions without the memory cost of 65,536 features.
Examples: /* → comment, ev + al → eval, ie + ex → iex, -e + nc → -enc, \r + \n → Windows newlines.
- Single-pass extraction: All features in one function call per file
- No Python loops over bytes:
np.bincount()+ masked sums for character counting, vectorized entropy, packed uint32 arrays for bigrams - Precomputed boolean masks: 256-element arrays for O(1) character class lookup
- float32 throughout: Half memory, faster CPU operations
LightGBM chosen over XGBoost (faster, less memory), deep learning (overkill), and Random Forest (slower convergence).
| Parameter | Value | Reason |
|---|---|---|
| n_estimators | 800 | Converges without overfitting |
| learning_rate | 0.03 | Smooth convergence |
| max_depth | 5 | Prevents overfitting to specific bigram bins |
| num_leaves | 16 | Conservative, matches max_depth |
| reg_alpha | 1.0 | L1 for sparse feature usage |
| reg_lambda | 5.0 | Strong L2 regularization |
| min_child_samples | 50 | Prevents noise patterns in small leaves |
| Metric | Train | Val | Gap |
|---|---|---|---|
| Accuracy | 0.9945 | 0.9740 | 0.0205 |
| Precision | 0.9985 | 0.9867 | 0.0118 |
| Recall | 0.9906 | 0.9610 | 0.0296 |
| F1 | 0.9945 | 0.9737 | 0.0209 |
| AUC-ROC | 0.9999 | 0.9958 | 0.0040 |
| AUC-PR | 0.9999 | 0.9962 | 0.0036 |
| MCC | 0.9891 | 0.9483 | 0.0408 |
| FPR | 0.0015 | 0.0130 | 0.0115 |
| FNR | 0.0094 | 0.0390 | 0.0296 |
| Train | Val | |
|---|---|---|
| TN | 9,985 | 987 |
| FP | 15 | 13 |
| FN | 94 | 39 |
| TP | 9,905 | 961 |
Key takeaways:
- 97.4% validation accuracy
- 1.3% false positive rate — 13 per 1000 benign files
- 3.9% false negative rate — 39 per 1000 malware missed
- Training gap (0.02-0.04) indicates mild, acceptable overfitting
| Metric | Time |
|---|---|
| Feature extraction | 0.837 ms/sample |
| Model prediction | 0.049 ms/sample |
| Total | 0.887 ms/sample |
| Throughput | >1,100 files/sec on single CPU core |
- Minified benign JS can trigger false positives (statistical overlap with obfuscated malware)
- Very short files (<50 bytes) lack sufficient signal
- Language-agnostic — no semantic understanding
- No adversarial robustness guarantees
- Add hashed trigram features (3-byte sequences:
var,eval,base64) - File header features (first 100 bytes: shebangs, comments, signatures)
- Ensemble with YARA rules for defense-in-depth
- Per-file-type threshold calibration
- Active learning for low-confidence predictions