[ICML 2026] Inference-time LLM structured pruning via probe-based representation-parameter coupling
-
Updated
Oct 6, 2026 - Python
[ICML 2026] Inference-time LLM structured pruning via probe-based representation-parameter coupling
[🔥 ICML2026 🔥 ]A training-free orchestration framework for building interactive omni-modal assistants by composing off-the-shelf modality experts, explicit LLM routing, text-centric cross-modal memory, and interruption-aware streaming interaction.
To associate your repository with the trainingfree topic, visit your repo's landing page and select "manage topics."