Built with Python β’ Ollama β’ FAISS β’ Streamlit β’ Sentence Transformers
- Project Overview
- Air-Gap-Capable Operation
- Core Capabilities
- How the Platform Works
- Project Architecture
- Technology Stack
- Cybersecurity Knowledge Base
- Repository Structure
- Installation and Launch
- Example Questions
- Project Screenshots
- Testing and Validation
- Documentation
- Skills Demonstrated
- Future Improvements
- Important Limitations
- Author
- License
The AI SOC Analyst Assistant is a local cybersecurity operations platform that combines two practical analyst workflows:
- Automated phishing email triage
- Retrieval-Augmented Generation cybersecurity assistance
The phishing workflow allows an analyst to paste or upload suspicious email content for structured review. The application can identify indicators, suspicious language, social-engineering characteristics, sender concerns, links, and other details that may require further investigation.
The cybersecurity knowledge assistant allows users to ask natural-language questions about threats, vulnerabilities, defensive frameworks, incident response, phishing, malware, security controls, and Security Operations Center procedures.
Before generating an answer, the assistant searches a locally stored cybersecurity knowledge base using FAISS semantic retrieval. Relevant document sections are then provided to a locally hosted Large Language Model through Ollama.
This produces responses that are more grounded, transparent, and useful than answers generated from the language model alone.
The platform is designed for private, restricted, disconnected, and air-gapped environments.
After the required software, Python packages, Ollama model, vector index, and cybersecurity documents have been installed locally, the application's core workflows can operate without an active internet connection.
During normal offline operation:
- Analyst questions remain on the local workstation.
- Suspicious email content remains on the local workstation.
- Cybersecurity documents are retrieved from local storage.
- Embeddings are processed locally.
- FAISS searches are performed locally.
- AI responses are generated locally through Ollama.
- No OpenAI API key is required.
- No paid cloud AI subscription is required.
- No prompts need to be transmitted to an external AI provider.
Important: Internet access is required during initial provisioning to download the project, Python dependencies, Ollama, the selected language model, and any required knowledge-base documents. After those resources are stored locally, the main analysis workflows can function without external network access.
This design makes the project appropriate for:
- Cybersecurity laboratories
- SOC training environments
- Restricted networks
- Sensitive testing environments
- Privacy-focused deployments
- Offline demonstrations
- Air-gapped research systems
The Retrieval-Augmented Generation interface allows analysts to ask questions about:
- Security operations
- Cybersecurity threats
- Detection engineering
- Incident response
- Phishing
- Malware
- Vulnerabilities
- CVEs
- MITRE ATT&CK
- MITRE ATLAS
- MITRE D3FEND
- NIST guidance
- OWASP guidance
- Authentication security
- Password security
- Security logging
- Defensive controls
The assistant retrieves relevant information from the local knowledge base before generating a response through Ollama.
- Natural-language cybersecurity questions
- Local document retrieval
- Semantic vector search
- Context-aware response generation
- Source-supported answers
- Similarity-ranked retrieval results
- Reduced dependence on model memory
- Reduced hallucination risk
- Local inference without cloud AI services
The phishing workflow supports analyst review of suspicious email content.
Potential analysis areas include:
- Sender information
- Subject line
- Message body
- Suspicious URLs
- Urgent or threatening language
- Credential-harvesting indicators
- Sender impersonation
- Social-engineering techniques
- Payment or account pressure
- Suspicious attachment references
- Indicators of compromise
- Severity observations
- Recommended analyst actions
The application is intended to support human review rather than replace analyst judgment.
The application:
- Loads trusted cybersecurity documents.
- Extracts text from the documents.
- Splits the text into searchable chunks.
- Generates local vector embeddings.
- Stores the embeddings in a FAISS index.
- Converts a user's question into an embedding.
- Searches for the most relevant document chunks.
- Sends the retrieved context to Ollama.
- Generates a grounded response.
- Displays relevant source information.
- Modular Python architecture
- Streamlit web interface
- Local Ollama integration
- Sentence Transformer embeddings
- FAISS vector database
- PDF document processing
- Automatic text chunking
- Semantic similarity search
- Automated testing with Pytest
- Git version control
- Beginner-friendly documentation
- Rebuildable local vector store
- Expandable knowledge-base structure
The project supports two complementary analyst workflows.
Analyst Question
β
βΌ
Streamlit Interface
β
βΌ
RAG Engine
β
βΌ
Sentence Transformer Embedding
β
βΌ
FAISS Vector Search
β
βΌ
Relevant Cybersecurity Documents
β
βΌ
Ollama Local Language Model
β
βΌ
Grounded Answer and Sources
Suspicious Email
β
βΌ
Streamlit Interface
β
βΌ
Email Parsing and Analysis
β
βΌ
Indicator and Pattern Review
β
βΌ
Structured Phishing Findings
β
βΌ
Human Analyst Decision
Analyst
β
βΌ
Streamlit Web Interface
β
ββββββββββββββββββ΄βββββββββββββββββ
β β
βΌ βΌ
Cybersecurity Knowledge Mode Phishing Triage Mode
β β
βΌ βΌ
RAG Engine Email Analysis Engine
β β
βββββββββ΄βββββββββ β
β β β
βΌ βΌ βΌ
Sentence Transformer FAISS Search Structured Findings
β β β
βββββββββ¬βββββββββ β
βΌ β
Retrieved Local Context β
β β
βΌ β
Ollama Local LLM β
β β
ββββββββββββββββββ¬βββββββββββββββββ
βΌ
Analyst Decision Support
Both workflows are designed to execute locally. The RAG workflow searches the local FAISS index before response generation, while the phishing workflow processes suspicious email content and presents structured findings for analyst review.
| Category | Technology |
|---|---|
| Programming Language | Python |
| Web Interface | Streamlit |
| Local AI Runtime | Ollama |
| Large Language Model | Llama 3.2 |
| AI Architecture | Retrieval-Augmented Generation |
| Embedding Model | Sentence Transformers |
| Vector Database | FAISS |
| Document Processing | PyMuPDF |
| Testing | Pytest |
| Version Control | Git |
| Repository Hosting | GitHub |
| Development Environment | Visual Studio Code |
The local knowledge base includes material covering:
- MITRE ATT&CK
- MITRE ATLAS
- MITRE D3FEND
- NIST Cybersecurity Framework
- NIST security guidance
- NIST password guidance
- OWASP Top 10
- CISA phishing guidance
- High-impact CVEs
- Log4Shell and CVE-2021-44228
- Incident-response playbooks
- SOC procedures
- Authentication best practices
- Password security
- Security logging
- Phishing detection
- Threat intelligence concepts
- Defensive cybersecurity controls
MITRE ATT&CK provides structured information about adversary tactics, techniques, and procedures observed in real-world attacks.
The assistant can use indexed ATT&CK documentation to support questions about:
- Initial access
- Execution
- Persistence
- Privilege escalation
- Credential access
- Discovery
- Lateral movement
- Collection
- Command and control
- Exfiltration
- Impact
MITRE ATLAS documents adversarial tactics and techniques affecting artificial-intelligence and machine-learning systems.
The assistant can answer questions involving:
- AI system threats
- Prompt injection
- Model extraction
- Data poisoning
- Adversarial machine learning
- AI-focused mitigations
- Differences between ATLAS and ATT&CK
MITRE D3FEND provides a knowledge graph of defensive cybersecurity techniques.
The assistant can explain:
- Defensive countermeasures
- The relationship between ATT&CK and D3FEND
- Detection and protection concepts
- Defensive technique selection
- Cybersecurity control relationships
The assistant can retrieve locally stored NIST material covering:
- Cybersecurity risk management
- Identify
- Protect
- Detect
- Respond
- Recover
- Govern
- Authentication
- Password guidance
- Security controls
The local knowledge base includes high-impact vulnerability information.
Example topics include:
- CVE identifiers
- Vulnerability impact
- Affected technologies
- Severity
- Exploitation risk
- Recommended mitigations
- Log4Shell
- CVE-2021-44228
AI-SOC-Assistant/
β
βββ backend/
β βββ rag_engine.py
β βββ supporting backend modules
β
βββ frontend/
β βββ streamlit_app.py
β βββ supporting interface files
β
βββ knowledge_base/
β βββ documents/
β β βββ mitre_attack/
β β βββ mitre_atlas/
β β βββ mitre_d3fend/
β β βββ nist/
β β βββ owasp/
β β βββ cve/
β β βββ additional cybersecurity documents
β β
β βββ indexes/
β βββ FAISS index files
β βββ document metadata
β
βββ tests/
β βββ RAG retrieval tests
β βββ application tests
β βββ validation scripts
β
βββ docs/
β βββ Part-A-Install-and-Run.md
β βββ Part-B-Build-From-Scratch.md
β βββ Part-C-Testing-and-Validation.md
β βββ Part-D-GitHub-Deployment-and-Troubleshooting.md
β
βββ screenshots/
β βββ 01-main-dashboard.png
β βββ 02-phishing-answer.png
β βββ 03-mitre-answer.png
β βββ 04-source-citations.png
β βββ 05-pytest-results.png
β βββ 06-vector-store-files.png
β βββ 07-ollama-models.png
β βββ 08-automated-phishing-engine-results.png
β βββ 09-mitre-atlas.png
β βββ 10-d3fend-related-to-mitre.png
β βββ 11-nist-framework.png
β βββ 12-log4shell-cve.png
β
βββ evaluation/
βββ data/
βββ app.py
βββ dashboard.py
βββ requirements.txt
βββ README.md
βββ .gitignore
The exact supporting filenames may change as the project develops. Refer to the repository for the current structure.
Complete beginner installation instructions are available in:
Part A β Install and Run the Finished Project
git clone https://github.com/ericsledge/AI-SOC-Assistant.git
cd AI-SOC-Assistantpython -m venv venv.\venv\Scripts\Activate.ps1python -m pip install --upgrade pip
pip install -r requirements.txtollama --versionollama pull llama3.2:3bollama listFrom the project root:
python -m streamlit run .\frontend\streamlit_app.pyStreamlit should open the application in a local web browser.
A typical local address is:
http://localhost:8501
Depending on the current interface design, phishing triage may appear as a separate page, navigation option, or application mode.
Use the following questions to validate the local knowledge assistant.
What is MITRE ATLAS?
How does MITRE ATLAS differ from MITRE ATT&CK?
What is data poisoning?
What is prompt injection?
What is MITRE D3FEND?
How does MITRE D3FEND relate to MITRE ATT&CK?
How can D3FEND support defensive cybersecurity operations?
What is the NIST Cybersecurity Framework?
What are the core functions of the NIST Cybersecurity Framework?
What password guidance does NIST provide?
What is CVE-2021-44228?
Explain Log4Shell.
What are high-impact CVEs?
What mitigations are associated with Log4Shell?
How should a SOC analyst investigate a phishing email?
What indicators should an analyst review during a phishing investigation?
What is an indicator of compromise?
How should a security team respond to suspected credential theft?
What is the OWASP Top 10?
How can SQL injection be prevented?
How can cross-site scripting be prevented?
The following screenshots demonstrate the completed platform, testing process, local AI model, vector database, phishing workflow, and expanded cybersecurity knowledge base.
If an image does not appear, confirm that the filename and extension exactly match the file stored in the
screenshots/directory.
The main Streamlit interface provides access to the AI SOC Analyst Assistant through a local web application.
The phishing workflow reviews suspicious email content and presents analyst-oriented observations.
The assistant retrieves indexed MITRE documentation before generating a grounded cybersecurity response.
Retrieved source information allows the analyst to identify which local documents contributed to the response.
Automated testing helps validate retrieval, application behavior, and supporting project components.
Cybersecurity document embeddings and supporting metadata are stored locally for semantic retrieval.
Ollama runs the selected language model directly on the local workstation without requiring a paid cloud AI API.
The phishing engine produces structured findings that can support analyst review of suspicious messages, indicators, links, and social-engineering characteristics.
The assistant retrieves locally indexed MITRE ATLAS content to answer questions about adversarial activity affecting AI and machine-learning systems.
The assistant explains how MITRE D3FEND defensive techniques relate to adversary behavior documented through MITRE ATT&CK.
The local RAG system retrieves NIST documentation to explain cybersecurity risk-management functions and framework concepts.
The assistant retrieves local vulnerability documentation to explain Log4Shell, CVE-2021-44228, potential impact, and mitigation considerations.
The project includes automated and manual testing.
Run the test suite from the project root:
pytestFor more detailed output:
pytest -vThe following areas should be validated:
| Test Area | Example Question | Expected Source Type |
|---|---|---|
| MITRE ATT&CK | What is MITRE ATT&CK? | MITRE ATT&CK documentation |
| MITRE ATLAS | What is MITRE ATLAS? | MITRE ATLAS overview |
| MITRE D3FEND | How does D3FEND relate to ATT&CK? | MITRE D3FEND overview |
| NIST | What is the NIST Cybersecurity Framework? | NIST documentation |
| CVE | What is CVE-2021-44228? | High-impact CVE documentation |
| Log4Shell | Explain Log4Shell. | CVE or Log4Shell documentation |
| Phishing | Analyze a suspicious email. | Structured phishing findings |
For each RAG question, confirm that:
- A relevant response is generated.
- Relevant document chunks are retrieved.
- The expected source appears.
- The application does not require a cloud AI API.
- The local model is responding through Ollama.
After all resources are installed:
- Confirm the Ollama model is stored locally.
- Confirm the FAISS index is present.
- Confirm the knowledge-base documents are present.
- Launch and test the application while connected.
- Disconnect the workstation from the internet.
- Restart or refresh the local application.
- Ask an ATLAS, D3FEND, NIST, or CVE question.
- Submit a phishing email for analysis.
- Confirm that both local workflows continue to function.
This validates air-gap-capable operation after initial provisioning.
The repository contains a complete beginner-friendly documentation series.
Part A β Install and Run the Finished Project
Covers:
- Required software
- System requirements
- Git installation
- Python installation
- Visual Studio Code installation
- Ollama installation
- Repository cloning
- Virtual environments
- Dependency installation
- Model download
- Application launch
- Initial verification
Part B β Build the Entire Project from Scratch
Covers:
- Repository organization
- Python modules
- Document loading
- Text chunking
- Embeddings
- FAISS
- RAG engine development
- Ollama integration
- Streamlit development
- Phishing analysis
- Knowledge-base expansion
Part C β Testing and Validation
Covers:
- Automated testing
- Retrieval testing
- Source validation
- ATLAS validation
- D3FEND validation
- NIST validation
- CVE validation
- Phishing workflow validation
- Offline operational testing
Part D β GitHub Deployment, Troubleshooting, and Final Submission
Covers:
- Git status
- Staging changes
- Commits
- GitHub pushes
- README verification
- Screenshot verification
- Common errors
- Troubleshooting
- Final repository review
This project demonstrates how to:
- Build a local AI application.
- Implement Retrieval-Augmented Generation.
- Generate document embeddings.
- Store and search vectors with FAISS.
- Connect a local language model through Ollama.
- Build a Streamlit web interface.
- Process cybersecurity PDF documents.
- Expand a knowledge base without changing the RAG architecture.
- Perform local phishing email triage.
- Validate source retrieval.
- Test an AI-assisted cybersecurity application.
- Document a full capstone project.
- Use Git and GitHub for version control.
- Prepare a system for restricted or disconnected operation.
- Local Large Language Models
- Retrieval-Augmented Generation
- Prompt engineering
- Embedding generation
- Semantic search
- Context retrieval
- Response grounding
- Local inference
- Security Operations Center workflows
- Phishing analysis
- Incident response
- Threat intelligence
- Indicators of compromise
- MITRE ATT&CK
- MITRE ATLAS
- MITRE D3FEND
- NIST Cybersecurity Framework
- OWASP
- Vulnerability analysis
- CVE interpretation
- Log4Shell analysis
- Python development
- Modular architecture
- Streamlit
- FAISS
- PyMuPDF
- Pytest
- Virtual environments
- Dependency management
- Error handling
- Debugging
- Git
- GitHub
- Technical documentation
- Local-first processing
- Air-gap-capable operation
- Data privacy
- Restricted-network deployment
- Human-in-the-loop analyst support
- Expandable knowledge-base design
Many AI chatbot projects depend entirely on cloud-hosted models and the model's internal training data.
The AI SOC Analyst Assistant takes a different approach.
It combines:
- Local language-model inference
- Local cybersecurity documents
- Local semantic vector search
- Source-supported RAG responses
- Automated phishing triage
- Cybersecurity framework retrieval
- CVE knowledge
- Offline operational capability
- Human analyst review
The result is a practical educational platform that demonstrates how AI can support security operations without requiring sensitive information to be sent to an external AI service during normal operation.
Potential future enhancements include:
- Unified Streamlit navigation for both application modes
- Additional MITRE ATLAS techniques
- Expanded MITRE D3FEND mappings
- Additional CVE collections
- Automated document ingestion
- Conversation memory
- Analyst case notes
- Exportable PDF reports
- IOC export
- YARA rule assistance
- Sigma rule assistance
- SIEM integration
- Local threat-intelligence feeds
- Advanced email-header parsing
- Attachment metadata analysis
- Role-based access controls
- Model selection controls
- Performance monitoring
- Improved retrieval evaluation
- Additional local language models
- Containerized offline deployment
Any future internet-based integrations should remain optional so the core local and air-gap-capable workflows continue to function independently.
This application is an educational and analyst-support project.
It should not be treated as:
- A replacement for trained security analysts
- A final authority on whether an email is malicious
- A complete malware-analysis platform
- A production SIEM
- A replacement for enterprise threat-intelligence systems
- A substitute for organizational incident-response procedures
- A source of guaranteed vulnerability remediation advice
AI-generated results and automated findings should be reviewed by a qualified human analyst.
The quality of RAG responses also depends on:
- The quality of the indexed documents
- The relevance of retrieved chunks
- The selected embedding model
- The selected Ollama model
- The wording of the question
- The configuration of the retrieval process
No.
The application uses Ollama for local language-model inference.
No.
The selected model runs locally after it has been downloaded.
Yes, after initial provisioning.
Python dependencies, Ollama, the language model, project files, and knowledge-base documents must first be downloaded and installed. After those resources are stored locally, the main RAG and phishing-analysis workflows can operate without an external internet connection.
The project is designed to be air-gap-capable after all required software, models, dependencies, indexes, and documents have been transferred to and installed on the disconnected system.
The core RAG workflow searches the local FAISS knowledge base rather than the public internet.
Yes.
The knowledge base can be expanded with additional trusted documents. The vector store must be rebuilt so the new documents are embedded and indexed.
No.
The phishing workflow supports analyst review. Human validation is still required.
Yes, although model configuration and system-resource requirements may need to be adjusted.
Constructive feedback, testing, documentation improvements, and feature suggestions are welcome.
A typical contribution workflow is:
- Fork the repository.
- Create a new branch.
- Make and test the changes.
- Commit the changes.
- Push the branch.
- Open a pull request.
This project uses tools and knowledge made available by the open-source and cybersecurity communities.
Special acknowledgment goes to:
- Python
- Ollama
- Streamlit
- FAISS
- Sentence Transformers
- PyMuPDF
- Hugging Face
- Pytest
- Git
- GitHub
- MITRE
- NIST
- OWASP
- CISA
Artificial Intelligence β’ Cybersecurity β’ Python Development
GitHub:
Repository:
https://github.com/ericsledge/AI-SOC-Assistant
This repository is provided for educational, portfolio, and cybersecurity-learning purposes.
Review the repository's license file for the complete terms governing use, modification, and redistribution.
If this project helped you learn about artificial intelligence, cybersecurity, Python, Streamlit, FAISS, Ollama, or Retrieval-Augmented Generation, consider giving the repository a star.
Keep Learning β’ Keep Building β’ Keep Defending











