Whose Side Is Your Agent On? PrincipalBench — a multi-turn benchmark and post-training methods (prompt scaffold + per-token-KL distillation) for multi-party principal loyalty in LLM agents.
nlp benchmark reinforcement-learning knowledge-distillation ai-safety ai-alignment large-language-models llm-evaluation llm-agents agent-safety principal-agent
-
Updated
Jul 29, 2026 - Python