You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Empirical study of candidate-order instability in non-autoregressive multi-candidate transformers, with cardinality scaling, permutation marginalization, cyclic shifts, calibration, and reproducible statistical evaluation.
Fine-tuned DeBERTa-v3-base for 77-class BANKING77 intent classification with leakage-aware evaluation and statistical comparison against a TF-IDF baseline.
Benchmark and results for Closing the Compliance Gap: open-weights and proprietary LLMs for privacy-constrained contact center automation (under review, NeurIPS 2026 InfPriv workshop)
Fine-tunes Google FLAN-T5-base with LoRA (PEFT) on Banking77 to classify 77 banking intents. Trains only 1.77M params (0.71%) in ~2hrs on free Colab GPU. Served via FastAPI REST API with batch prediction, Docker support & evaluation metrics.
Reproducible Laya-CoreML vs Jev benchmark for zero-shot intent classification on Banking77, ArBanking77, and CLINC150. Includes paired accuracy and per-dataset results.
NLP lab comparing a Hybrid XGBoost against a fine-tuned DistilBERT for 77-way banking intent classification, evaluating F1, latency and structural trade-offs.
Which bank support tickets can a machine safely handle, which must reach a person, and what each approach costs. Benchmarks keyword rules, classic ML and Gemini on BANKING77, scored by harm, not just accuracy.
Frozen encoder + Mahalanobis prototype for class-incremental intent classification. 50+ experiments across BANKING77, CLINC150, HWU64, AG News. Matches fine-tuned baselines at 5MB state with zero forgetting, order-invariance, 455 QPS.
Local Jev-style System One decision models on Apple Silicon Macs: AnyJev, Kev, Laya (Ollaya) and CLM benchmarked on English, Korean and BANKING77 (accuracy, latency, order-flip)
Measured experiments on decision models. Can an open-weight model re-rank search results? Does a token budget explain a benchmark failure? What does one typed decision cost across model families? Every request and response is recorded, and negative results are published as they came out.