Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis


A dual-memory reasoning framework that continuously aligns the evolving diagnostic state with reusable operational experience for reliable failure diagnosis.

🎉 News

  • [2026-08-13] OpsMem is accepted by ISSRE 2026 Industry Track!
  • [2026-07-06] OpsMem code released!

🏴 Overview

OpsMem overview

  • Short-Term Memory (STM) maintains the evolving diagnostic state of the current incident.
  • Long-Term Memory (LTM) organizes reusable operational experience across incidents.
  • Cross-Memory Resonance (CMR) aligns the current diagnostic state with relevant operational experience.
  • Long-Term Memory Consolidation distills reusable experience from solved incidents into the LTM.

📦 Code Structure

OpsMem/
  README.md
  LICENSE                   # Apache License 2.0                
  requirements.txt          # Python dependencies
  assets/
    overview.png            # overview figure
  code/
    config.yaml             # configuration
    main.py                 # entry point
    pipeline.py             # orchestration of CMR, diagnosis, and consolidation
    eval.py                 # self-consistent LLM-as-a-Judge evaluation
    datasets/case_1/        # sanitized real-case example
    agents/                 # Central/Expert agents and their LLM action wrappers
    stm/                    # Short-Term Memory that captures the current diagnostic state
    ltm/                    # Long-Term Memory that organizes reusable operational experience
    cmr/                    # Cross-Memory Resonance that aligns STM and LTM
    consolidation/          # Long-Term Memory consolidation
    tools/                  # diagnostic tool interface (metric/log/shell tools)
    prompts/                # prompts
    utils/                  # LLM client, embedding provider, and logging utilities

🗂️ Dataset and Knowledge Base

The original diagnosis dataset and long-term knowledge base are derived from Huawei operational data and internal troubleshooting experience. Due to confidentiality and compliance restrictions, the full dataset and knowledge base cannot be publicly released.

To make the repository runnable and to illustrate the end-to-end workflow, we provide one sanitized case and a small sanitized seed knowledge base distilled from real data.

🛠 Quick Start

Create environment

conda create -n OpsMem python=3.10 -y
conda activate OpsMem

Install dependencies

pip install -r requirements.txt

Configure LLM

  • Edit code/config.yaml and set the OpenAI-compatible LLM endpoint under model.models.default.

Configure embedding

  • The default embedding provider is defined in code/utils/embedding.py (local_bge_m3 by default). Replace EMBEDDING_MODEL_NAME_OR_PATH with your local BGE-M3 model directory.
  • You can also use another local or cloud embedding model by implementing/changing the provider in the same file, as long as its encode(texts) method returns an embedding matrix.

Run example

cd code
python main.py

Evaluate

python eval.py --input output/opsmem/answers/default.csv

🙏 Acknowledgements

OpsMem builds upon GoS, our previous neuro-symbolic reasoning framework for abductive tasks, where an explicit belief state constrains multi-agent reasoning. OpsMem inherits this state-guided reasoning design and extends it into a dual-memory framework for failure diagnosis, where the current diagnostic state and reusable operational experience jointly condition the diagnosis process.

📬 Citation

If you find this work helpful, please consider citing us:

@article{sun2026opsmem,
  title={OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis},
  author={Sun, Yongqian and Gao, Rongchen and Luo, Yu and Gu, Wenwei and Zhang, Shenglin and Guo, Qingyi and Fu, Qiuai and Wu, Yaoliang and Pei, Dan},
  journal={arXiv preprint arXiv:2607.11357},
  year={2026}
}

About

A dual-memory reasoning framework that continuously aligns the evolving diagnostic state with reusable operational experience for reliable failure diagnosis.

Topics

Resources

Stars

8 stars

Watchers

0 watching

Forks

Contributors

Languages