Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Script Malware Detection

Problem

Script-based malware (PowerShell, JavaScript, VBScript) is a primary attack vector. Scripts are easy to obfuscate, use built-in OS tools, and evade signature-based detection. We need a fast, statistical classifier that detects malicious scripts with high accuracy and low false positive rate, without parsing or executing code.

Design Decisions

Rejected approaches and why:

Approach Reason Rejected
Regex-based features Slow, language-specific, easily evaded
AST parsing Fails on obfuscated code, requires multiple parsers
Sandbox execution Seconds per file, needs VM infrastructure
Word TF-IDF Vocabulary explosion, misses obfuscation
Full byte n-gram histograms 65K-16M features, too sparse
Deep learning Overkill, slow inference, harder to deploy

Core insight: Malware obfuscation changes byte-level statistical properties. We don't need to understand code — we measure how "weird" the byte distribution looks.

Feature Evolution

Round 1 — Regex (46 features): Keyword matching, structure counts, network indicators. 30+ minutes for 20K samples. Slow, evadable, language-specific.

Round 2 — Character TF-IDF (12 + 5000): Replaced word tokens with char n-grams. Still needed sklearn fitting, 5000 features mostly redundant.

Round 3 — Byte Histogram (8 + 256): Operated directly on raw bytes. Fast and language-agnostic. Problem: minified benign JS and obfuscated malware have similar byte frequencies. A WordPress auto-generated file was flagged as malware.

Round 4 — Hashed Byte Bigrams (8 + 1024) [CURRENT]: Byte frequencies alone aren't enough — byte transitions carry language-specific structure. Full bigrams would be 65,536 features. Instead, hash each pair into 1024 bins (hashing trick). Collisions are tolerated; LightGBM disentangles them using other features.

Why this fixes the WordPress false positive: /* and */ (comment markers) hash to specific bins. JavaScript pairs ( va, ar, fu) hash differently than PowerShell pairs (ie, ex, -e). The model learns language-specific byte transitions.

Final Feature Set

Statistical Features (8):

# Feature How Computed Signal
1 line_count Count \n bytes Obfuscated code often 1-5 long lines
2 max_line_length np.diff(newline_positions) Long lines = minified/packed
3 uppercase_ratio sum(counts[65:91]) / length PowerShell PascalCase cmdlets
4 digit_ratio sum(counts[48:58]) / length High in hex/base64 payloads
5 punctuation_ratio sum(counts[PUNCT]) / length Elevated in obfuscated code
6 whitespace_ratio sum(counts[SPACE]) / length Low in packed, high in readable
7 entropy -sum(p * log2(p)) High = encrypted/packed/random
8 unique_byte_count count_nonzero(counts) Low unique + high entropy = suspicious

Hashed Bigram Features (1024):

Each byte pair (b0, b1) is packed into (b0 << 8) | b1 and hashed via pair % 1024. The normalized histogram captures byte transitions without the memory cost of 65,536 features.

Examples: /* → comment, ev + al → eval, ie + ex → iex, -e + nc → -enc, \r + \n → Windows newlines.

Performance Optimizations

  • Single-pass extraction: All features in one function call per file
  • No Python loops over bytes: np.bincount() + masked sums for character counting, vectorized entropy, packed uint32 arrays for bigrams
  • Precomputed boolean masks: 256-element arrays for O(1) character class lookup
  • float32 throughout: Half memory, faster CPU operations

Model

LightGBM chosen over XGBoost (faster, less memory), deep learning (overkill), and Random Forest (slower convergence).

Parameter Value Reason
n_estimators 800 Converges without overfitting
learning_rate 0.03 Smooth convergence
max_depth 5 Prevents overfitting to specific bigram bins
num_leaves 16 Conservative, matches max_depth
reg_alpha 1.0 L1 for sparse feature usage
reg_lambda 5.0 Strong L2 regularization
min_child_samples 50 Prevents noise patterns in small leaves

Results

Metric Train Val Gap
Accuracy 0.9945 0.9740 0.0205
Precision 0.9985 0.9867 0.0118
Recall 0.9906 0.9610 0.0296
F1 0.9945 0.9737 0.0209
AUC-ROC 0.9999 0.9958 0.0040
AUC-PR 0.9999 0.9962 0.0036
MCC 0.9891 0.9483 0.0408
FPR 0.0015 0.0130 0.0115
FNR 0.0094 0.0390 0.0296
Train Val
TN 9,985 987
FP 15 13
FN 94 39
TP 9,905 961

Key takeaways:

  • 97.4% validation accuracy
  • 1.3% false positive rate — 13 per 1000 benign files
  • 3.9% false negative rate — 39 per 1000 malware missed
  • Training gap (0.02-0.04) indicates mild, acceptable overfitting

Inference Speed

Metric Time
Feature extraction 0.837 ms/sample
Model prediction 0.049 ms/sample
Total 0.887 ms/sample
Throughput >1,100 files/sec on single CPU core

Limitations

  • Minified benign JS can trigger false positives (statistical overlap with obfuscated malware)
  • Very short files (<50 bytes) lack sufficient signal
  • Language-agnostic — no semantic understanding
  • No adversarial robustness guarantees

Future Work

  • Add hashed trigram features (3-byte sequences: var, eval, base64)
  • File header features (first 100 bytes: shebangs, comments, signatures)
  • Ensemble with YARA rules for defense-in-depth
  • Per-file-type threshold calibration
  • Active learning for low-confidence predictions

About

A lightweight, high-throughput malware detector for PowerShell, JavaScript, and VBScript that identifies obfuscated scripts using byte-level statistical features and hashed byte-pair patterns—without parsing or executing code. Built with NumPy and LightGBM, achieving 97.4% validation accuracy with sub-millisecond CPU inference.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages