A dual-memory reasoning framework that continuously aligns the evolving diagnostic state with reusable operational experience for reliable failure diagnosis.
- [2026-08-13] OpsMem is accepted by ISSRE 2026 Industry Track!
- [2026-07-06] OpsMem code released!
- Short-Term Memory (STM) maintains the evolving diagnostic state of the current incident.
- Long-Term Memory (LTM) organizes reusable operational experience across incidents.
- Cross-Memory Resonance (CMR) aligns the current diagnostic state with relevant operational experience.
- Long-Term Memory Consolidation distills reusable experience from solved incidents into the LTM.
OpsMem/
README.md
LICENSE # Apache License 2.0
requirements.txt # Python dependencies
assets/
overview.png # overview figure
code/
config.yaml # configuration
main.py # entry point
pipeline.py # orchestration of CMR, diagnosis, and consolidation
eval.py # self-consistent LLM-as-a-Judge evaluation
datasets/case_1/ # sanitized real-case example
agents/ # Central/Expert agents and their LLM action wrappers
stm/ # Short-Term Memory that captures the current diagnostic state
ltm/ # Long-Term Memory that organizes reusable operational experience
cmr/ # Cross-Memory Resonance that aligns STM and LTM
consolidation/ # Long-Term Memory consolidation
tools/ # diagnostic tool interface (metric/log/shell tools)
prompts/ # prompts
utils/ # LLM client, embedding provider, and logging utilities
The original diagnosis dataset and long-term knowledge base are derived from Huawei operational data and internal troubleshooting experience. Due to confidentiality and compliance restrictions, the full dataset and knowledge base cannot be publicly released.
To make the repository runnable and to illustrate the end-to-end workflow, we provide one sanitized case and a small sanitized seed knowledge base distilled from real data.
conda create -n OpsMem python=3.10 -y
conda activate OpsMempip install -r requirements.txt- Edit
code/config.yamland set the OpenAI-compatible LLM endpoint undermodel.models.default.
- The default embedding provider is defined in
code/utils/embedding.py(local_bge_m3by default). ReplaceEMBEDDING_MODEL_NAME_OR_PATHwith your local BGE-M3 model directory. - You can also use another local or cloud embedding model by implementing/changing the provider in the same file, as long as its
encode(texts)method returns an embedding matrix.
cd code
python main.pypython eval.py --input output/opsmem/answers/default.csvOpsMem builds upon GoS, our previous neuro-symbolic reasoning framework for abductive tasks, where an explicit belief state constrains multi-agent reasoning. OpsMem inherits this state-guided reasoning design and extends it into a dual-memory framework for failure diagnosis, where the current diagnostic state and reusable operational experience jointly condition the diagnosis process.
If you find this work helpful, please consider citing us:
@article{sun2026opsmem,
title={OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis},
author={Sun, Yongqian and Gao, Rongchen and Luo, Yu and Gu, Wenwei and Zhang, Shenglin and Guo, Qingyi and Fu, Qiuai and Wu, Yaoliang and Pei, Dan},
journal={arXiv preprint arXiv:2607.11357},
year={2026}
}