Safety Research
Popular repositories Loading
-
persona_vectors
persona_vectors PublicPersona Vectors: Monitoring and Controlling Character Traits in Language Models
-
-
-
assistant-axis
assistant-axis PublicThe Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations—sometimes toward bizarr…
-
safety-tooling
safety-tooling PublicInference API for many LLMs and other useful tools for empirical research
Repositories
Showing 10 of 53 repositories
- self-modeling-eval Public
An eval benchmark suite that tests how well a LLM can predict its own behavior.
- auditing-agents Public
- misalignment-indicators Public
Source code for the paper: Probing the Misaligned Thinking Process of Language Models
- SCONE-bench Public
- aligning-ai-teams Public
- faithful-cot Public
Code for "Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning"
Top languages
Loading…
Most used topics
Loading…