From 0882d8fdb74d560a8be6a499b5729a7cd9ab7737 Mon Sep 17 00:00:00 2001 From: armstrongttwalker-alt Date: Thu, 23 Jul 2026 06:20:27 +0000 Subject: [PATCH] Auto-update ModelScope documentation [$(TZ='Asia/Shanghai' date +'%Y-%m-%d %H:%M')] --- docs/flagrelease_en/model_list.txt | 41 +++- ...elease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md | 127 ++++++++++ ...lease_AI21-Jamba-1.5-Mini-nvidia-FlagOS.md | 125 ++++++++++ .../FlagRelease_Baguettotron-metax-FlagOS.md | 108 +++++++++ ..._DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS.md | 108 +++++++++ ...pSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md | 120 ++++++++++ ...Seek-R1-Distill-Qwen-1.5B-nvidia-FlagOS.md | 151 ++++++++++++ ...gRelease_ERNIE-4.5-0.3B-PT-hygon-FlagOS.md | 137 +++++++++++ ...Release_ERNIE-4.5-0.3B-PT-nvidia-FlagOS.md | 140 +++++++++++ .../FlagRelease_ERNIE-4.5-0.3B-PT.md | 134 +++++++++++ ...lease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md | 145 ++++++++++++ ...ease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md | 147 ++++++++++++ ...elease_GLM-4-32B-Base-0414-hygon-FlagOS.md | 155 ++++++++++++ ...lease_GLM-4-32B-Base-0414-nvidia-FlagOS.md | 137 +++++++++++ .../FlagRelease_HY-MT2-7B-mthreads-FlagOS.md | 2 +- .../FlagRelease_Hy3-hygon-FlagOS.md | 156 ++++++++++++ .../FlagRelease_Hy3-iluvatar-FlagOS.md | 121 ++++++++++ .../FlagRelease_Hy3-metax-FlagOS.md | 171 +++++++++++++ .../FlagRelease_Hy3-mthreads-FlagOS.md | 150 ++++++++++++ .../FlagRelease_Hy3-nvidia-FlagOS-Express.md | 136 +++++++++++ .../FlagRelease_Hy3-tsingmicro-FlagOS.md | 122 ++++++++++ .../FlagRelease_Hy3-zhenwu-FlagOS.md | 121 ++++++++++ ...mi-Linear-48B-A3B-Instruct-hygon-FlagOS.md | 182 ++++++++++++++ ...i-Linear-48B-A3B-Instruct-nvidia-FlagOS.md | 75 +++--- .../FlagRelease_LFM2-2.6B-Exp-metax-FlagOS.md | 108 +++++++++ ...e_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md | 148 ++++++++++++ ..._Meta-Llama-3-8B-Instruct-nvidia-FlagOS.md | 132 +++++++++++ ...lagRelease_MiniCPM5-1B-kunlunxin-FlagOS.md | 17 +- ...gRelease_Moonlight-16B-A3B-hygon-FlagOS.md | 188 +++++++++++++++ ...Release_Moonlight-16B-A3B-nvidia-FlagOS.md | 156 ++++++++++++ ...lease_Phi-3.5-MoE-instruct-hygon-FlagOS.md | 135 +++++++++++ ...ease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md | 224 ++++++++++++++++++ ...lagRelease_Phi-3.5-mini-instruct-FlagOS.md | 108 +++++++++ ...ease_Phi-3.5-mini-instruct-metax-FlagOS.md | 108 +++++++++ ...elease_Qwen2.5-7B-Instruct-metax-FlagOS.md | 108 +++++++++ ..._Qwen2.5-Coder-7B-Instruct-metax-FlagOS.md | 108 +++++++++ ...elease_Qwen3.6-27B-metax-FlagOS-Express.md | 145 ++++++++++++ ...ease_Seed-OSS-36B-Instruct-hygon-FlagOS.md | 147 ++++++++++++ ...ase_Seed-OSS-36B-Instruct-nvidia-FlagOS.md | 141 +++++++++++ .../FlagRelease_ZCK-Qwen3-8B-metax-FlagOS.md | 48 ++++ .../FlagRelease_ZCK-Qwen3-8B-nvidia-FlagOS.md | 108 +++++++++ .../FlagRelease_aya-23-8B-metax-FlagOS.md | 108 +++++++++ ...lagRelease_gemma-1.1-7b-it-metax-FlagOS.md | 108 +++++++++ 43 files changed, 5312 insertions(+), 44 deletions(-) create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Baguettotron-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-iluvatar-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-mthreads-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-nvidia-FlagOS-Express.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-tsingmicro-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Hy3-zhenwu-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_LFM2-2.6B-Exp-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-7B-Instruct-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-hygon-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-nvidia-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_aya-23-8B-metax-FlagOS.md create mode 100644 docs/flagrelease_en/model_readmes/FlagRelease_gemma-1.1-7b-it-metax-FlagOS.md diff --git a/docs/flagrelease_en/model_list.txt b/docs/flagrelease_en/model_list.txt index 0b2db3263..aa1c6b66f 100644 --- a/docs/flagrelease_en/model_list.txt +++ b/docs/flagrelease_en/model_list.txt @@ -1,6 +1,12 @@ +FlagRelease/AI21-Jamba-1.5-Mini-hygon-FlagOS +FlagRelease/AI21-Jamba-1.5-Mini-nvidia-FlagOS FlagRelease/BAAI-Cardiac-Agent-hygon-FlagOS +FlagRelease/Baguettotron-metax-FlagOS FlagRelease/C2S-Scale-Gemma-2-27B-hygon-FlagOS FlagRelease/C2S-Scale-Gemma-2-27B-nvidia-FlagOS +FlagRelease/DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS +FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS +FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS FlagRelease/DeepSeek-R1-Distill-Qwen-32B-FlagOS-Cambricon FlagRelease/DeepSeek-R1-Distill-Qwen-32B-FlagOS-NVIDIA FlagRelease/DeepSeek-R1-FlagOS-Cambricon-BF16 @@ -25,8 +31,15 @@ FlagRelease/DeepSeek-V4-Pro-hygon-FlagOS FlagRelease/DeepSeek-V4-Pro-metax-FlagOS FlagRelease/DeepSeek-V4-Pro-mthreads-FlagOS FlagRelease/DeepSeek-V4-Pro-nvidia-FlagOS +FlagRelease/ERNIE-4.5-0.3B-PT +FlagRelease/ERNIE-4.5-0.3B-PT-hygon-FlagOS +FlagRelease/ERNIE-4.5-0.3B-PT-nvidia-FlagOS +FlagRelease/ERNIE-4.5-21B-A3B-PT-hygon-FlagOS +FlagRelease/ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS FlagRelease/ERNIE-4.5-300B-A47B-PT-FlagOS FlagRelease/Emu3.5-FlagOS +FlagRelease/GLM-4-32B-Base-0414-hygon-FlagOS +FlagRelease/GLM-4-32B-Base-0414-nvidia-FlagOS FlagRelease/GLM-4.5-FlagOS FlagRelease/GLM-5-FP8-FlagOS FlagRelease/GLM-5-ascend-FlagOS @@ -53,9 +66,20 @@ FlagRelease/HY-MT2-7B-mthreads-FlagOS FlagRelease/HY-MT2-7B-nvidia-FlagOS FlagRelease/HY-MT2-7B-zhenwu-FlagOS FlagRelease/Hunyuan-A13B-Instruct-FlagOS +FlagRelease/Hy3-hygon-FlagOS +FlagRelease/Hy3-iluvatar-FlagOS +FlagRelease/Hy3-metax-FlagOS +FlagRelease/Hy3-mthreads-FlagOS +FlagRelease/Hy3-nvidia-FlagOS-Express +FlagRelease/Hy3-tsingmicro-FlagOS +FlagRelease/Hy3-zhenwu-FlagOS FlagRelease/Kimi-K2-Instruct-FlagOS FlagRelease/Kimi-K2-Thinking-FlagOS +FlagRelease/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS FlagRelease/Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS +FlagRelease/LFM2-2.6B-Exp-metax-FlagOS +FlagRelease/Meta-Llama-3-8B-Instruct-hygon-FlagOS +FlagRelease/Meta-Llama-3-8B-Instruct-nvidia-FlagOS FlagRelease/MiniCPM-V-4-FlagOS FlagRelease/MiniCPM-V-4-metax-FlagOS FlagRelease/MiniCPM-o-4.5-ascend-FlagOS @@ -90,12 +114,20 @@ FlagRelease/MiniMax-M3-metax-FlagOS FlagRelease/MiniMax-M3-mthreads-FlagOS FlagRelease/MiniMax-M3-nvidia-FlagOS FlagRelease/MiniMax-M3-zhenwu-FlagOS +FlagRelease/Moonlight-16B-A3B-hygon-FlagOS +FlagRelease/Moonlight-16B-A3B-nvidia-FlagOS +FlagRelease/Phi-3.5-MoE-instruct-hygon-FlagOS +FlagRelease/Phi-3.5-MoE-instruct-nvidia-FlagOS +FlagRelease/Phi-3.5-mini-instruct-FlagOS +FlagRelease/Phi-3.5-mini-instruct-metax-FlagOS FlagRelease/QwQ-32B-FlagOS-Cambricon FlagRelease/QwQ-32B-FlagOS-Iluvatar FlagRelease/QwQ-32B-FlagOS-Nvidia FlagRelease/Qwen2-7B-FlagOS-Arm FlagRelease/Qwen2-7B-Instruct-FlagOS FlagRelease/Qwen2.5-32B-Instruct-FlagOS-Nvidia +FlagRelease/Qwen2.5-7B-Instruct-metax-FlagOS +FlagRelease/Qwen2.5-Coder-7B-Instruct-metax-FlagOS FlagRelease/Qwen2.5-VL-32B-Instruct-FlagOS-Metax-BF16 FlagRelease/Qwen2.5-VL-32B-Instruct-FlagOS-Nvidia FlagRelease/Qwen3-235B-A22B-FlagOS-nvidia @@ -126,7 +158,7 @@ FlagRelease/Qwen3.5-397B-A17B-metax-FlagOS FlagRelease/Qwen3.5-397B-A17B-nvidia-FlagOS FlagRelease/Qwen3.5-397B-A17B-zhenwu-FlagOS FlagRelease/Qwen3.6-27B-hygon-FlagOS -FlagRelease/Qwen3.6-27B-metax-FlagOS +FlagRelease/Qwen3.6-27B-metax-FlagOS-Express FlagRelease/Qwen3.6-35B-A3B-nomtp-ascend-FlagOS FlagRelease/Qwen3.6-35B-A3B-nomtp-hygon-FlagOS FlagRelease/Qwen3.6-35B-A3B-nomtp-iluvatar-FlagOS @@ -147,11 +179,16 @@ FlagRelease/RoboBrain2.0-7B-metax-FlagOS FlagRelease/RoboBrain2.5-8B-FlagOS FlagRelease/RoboBrain2.5-8B-ascend-FlagOS FlagRelease/Seed-OSS-36B-Instruct-FlagOS -FlagRelease/Seed-OSS-36B-Instruct-iluvatar-FlagOS +FlagRelease/Seed-OSS-36B-Instruct-hygon-FlagOS +FlagRelease/Seed-OSS-36B-Instruct-nvidia-FlagOS FlagRelease/TeleChat3-36B-Thinking-mthreads-FlagOS +FlagRelease/ZCK-Qwen3-8B-metax-FlagOS +FlagRelease/ZCK-Qwen3-8B-nvidia-FlagOS +FlagRelease/aya-23-8B-metax-FlagOS FlagRelease/deepseek-r1-1.5b-nvidia-FlagOS FlagRelease/farm_molecular_representation-hygon-FlagOS FlagRelease/farm_molecular_representation-nvidia-FlagOS +FlagRelease/gemma-1.1-7b-it-metax-FlagOS FlagRelease/gpt-oss-120b-FlagOS FlagRelease/grok-2-FlagOS FlagRelease/materials.smi-ted-hygon-FlagOS diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md new file mode 100644 index 000000000..97964992a --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-hygon-FlagOS.md @@ -0,0 +1,127 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +AI21-Jamba-1.5-Mini is an open-source large language model released by AI21. This release completes adaptation and validation on the Nvidia platform and is published based on the FlagOS software stack. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | AI21-Jamba-1.5-Mini-Nvidia-Origin | AI21-Jamba-1.5-Mini-Hygon-FlagOS | +|---------------------------------|-----------------------------------|----------------------------------| +| GPQA_Diamond | 0.177 | 0.222 | +| GPQA_Generative_CoT | 0.218 | 0.223 | +| LiveBench_New | 0.275 | 0.267 | +| MUSR_Generative | 0.290 | 0.300 | +| MMLU_Pro | 0.401 | 0.406 | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.24, build 297e128 | +| Operating System | Sugon OS 8.9 | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-hygon-tree_0.5.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607101038 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download \ + --model FlagRelease/AI21-Jamba-1.5-Mini-hygon-FlagOS \ + --local_dir /data/models/vllm-plugin-fl/AI21-Jamba-1.5-Mini-hygon-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name ai21-jamba-hygon \ + --network=host \ + --privileged=true \ + --shm-size=16g \ + -v /data/models/vllm-plugin-fl:/data/vllm-plugin-fl \ + -v /opt/hyhal:/opt/hyhal:ro \ + -itd \ + harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-hygon-tree_0.5.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607101038 \ + sleep infinity + +docker exec -it ai21-jamba-hygon bash + +``` +### Start the Server +```bash +nohup env \ + HIP_VISIBLE_DEVICES=0,1 \ + VLLM_PLUGINS=fl \ + TRITON_ALL_BLOCKS_PARALLEL=1 \ + vllm serve \ + --model /data/vllm-plugin-fl/AI21-Jamba-1.5-Mini-hygon-FlagOS \ + --tensor-parallel-size 2 \ + --enforce-eager \ + --max-cudagraph-capture-size 0 \ + --served-model-name ai21_flagos \ + --port 8131 \ + --gpu-memory-utilization 0.85 \ + > /workspace/ai21-flagos.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8131/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ai21_flagos", + "messages": [ + { + "role": "user", + "content": "你好" + } + ] + }' +``` + + +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from AI-ModelScope/AI21-Jamba-1.5-Mini and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-nvidia-FlagOS.md new file mode 100644 index 000000000..ad29b347a --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_AI21-Jamba-1.5-Mini-nvidia-FlagOS.md @@ -0,0 +1,125 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +AI21-Jamba-1.5-Mini is an open-source large language model released by AI21. This release completes adaptation and validation on the Nvidia platform and is published based on the FlagOS software stack. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | AI21-Jamba-1.5-Mini-Nvidia-Origin | AI21-Jamba-1.5-Mini-Nvidia-FlagOS | +|--------------|-----------------------------------|-----------------------------------| +| GPQA_Diamond | 0.1869 | 0.1717 | +| GPQA | 0.2324 | 0.2106 | +| LiveBench | 0.2760 | 0.2642 | +| MMLU_Pro | 0.4033 | 0.4044 | +| MUSR | 0.3082 | 0.2712 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-nvidia-gems_5.0.2-vllm_0.13.0-plugin_0.0.0-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_h20_3e-arc_amd64-driver_570.133.20:20260513152138 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/AI21-Jamba-1.5-Mini-nvidia-FlagOS --local_dir /data/AI21-Jamba-1.5-Mini-nvidia-FlagOS +``` + +### Start the Container +```bash +docker run --init --detach \ + --name ai21_jamba_1_5_mini_flagos \ + --gpus all \ + --shm-size=16g \ + -p 8000:8000 \ + -w /workspace \ + -v /data/AI21-Jamba-1.5-Mini-nvidia-FlagOS:/data/AI21-Jamba-1.5-Mini-nvidia-FlagOS \ + harbor.baai.ac.cn/external-cooperation/ai21-jamba-1.5-mini-nvidia-gems_5.0.2-vllm_0.13.0-plugin_0.0.0-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_h20_3e-arc_amd64-driver_570.133.20:20260513152138 \ + sleep infinity + +docker exec -it ai21_jamba_1_5_mini_flagos bash +``` +### Start the Server +```bash +nohup env CUDA_VISIBLE_DEVICES=0 \ +VLLM_PLUGINS=fl \ +TRITON_ALL_BLOCKS_PARALLEL=1 \ +VLLM_FL_FLAGOS_WHITELIST=exponential_,sort,rand_like,masked_fill_,argmax,lt_scalar,neg,where_self,exp,fill_tensor_,fill_scalar_,gt_scalar,gather,softmax,zeros,zeros_like,where_self,where_self_out,ones,sub,cumsum_out,sort_stable,index,to_copy,copy_,rms_norm \ +vllm serve \ + --model /data/AI21-Jamba-1.5-Mini-nvidia-FlagOS \ + --served-model-name ai21-jamba-1.5-mini-flagos \ + --port 8000 \ + --max-num-batched-tokens 8192 \ + > ai21-jamba-1.5-mini-flagos.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ai21-jamba-1.5-mini-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from AI-ModelScope/AI21-Jamba-1.5-Mini and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Baguettotron-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Baguettotron-metax-FlagOS.md new file mode 100644 index 000000000..9f94e73ff --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Baguettotron-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Baguettotron-metax-FlagOS-Origin | Baguettotron-metax-FlagOS-FlagOS | +|--------------|----------------------------------|----------------------------------| +| GPQA_Diamond | 0 | 0.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/baguettotron-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607121209-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Baguettotron-metax-FlagOS --local_dir /data/Baguettotron-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/baguettotron-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607121209-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Baguettotron-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Baguettotron --tensor-parallel-size 1 --max-model-len 4096 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Baguettotron", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PleIAs/Baguettotron and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS.md new file mode 100644 index 000000000..4a76bd4fd --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS-Origin | DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS-FlagOS | +|--------------|-----------------------------------------------|-----------------------------------------------| +| GPQA_Diamond | 0 | 62.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/deepseek-r1-0528-qwen3-8b-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607231257-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/DeepSeek-R1-0528-Qwen3-8B-metax-FlagOS --local_dir /data/DeepSeek-R1-0528-Qwen3-8B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/deepseek-r1-0528-qwen3-8b-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607231257-v2 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/DeepSeek-R1-0528-Qwen3-8B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name DeepSeek-R1-0528-Qwen3-8B --trust-remote-code --max-model-len 32768 --tensor-parallel-size 1 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "DeepSeek-R1-0528-Qwen3-8B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from deepseek-ai/DeepSeek-R1-0528-Qwen3-8B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md new file mode 100644 index 000000000..993127f63 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS.md @@ -0,0 +1,120 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero is trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as an initial stage, and it delivers outstanding reasoning capabilities. Through RL training, DeepSeek-R1-Zero naturally exhibits numerous powerful and intriguing reasoning behaviors. + +Nevertheless, DeepSeek-R1-Zero suffers from issues such as endless repetition, poor readability, and mixed-language outputs. To address these flaws and further boost reasoning performance, we developed DeepSeek-R1, which incorporates cold-start data prior to the RL phase. DeepSeek-R1 achieves performance comparable to OpenAI o1 on mathematical, coding, and reasoning tasks. + +To support the research community, we have open-sourced DeepSeek-R1-Zero, DeepSeek-R1, as well as six dense models distilled from DeepSeek-R1 based on the Llama and Qwen architectures. DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI o1-mini across a wide range of benchmarks, setting a new state-of-the-art record among dense models. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | DeepSeek-R1-Distill-Qwen-1.5B-Nvidia-Origin | DeepSeek-R1-Distill-Qwen-1.5B-Iluvatar-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| musr_generative | 0.3320 | 0.3475 | +| mmlu_pro | 0.1747 | 0.1781 | +| aime | 0 | 0 | +| livebench_new | 0.1261 | 0.1229 | +| gpqa_generative_cot | 0.0866 | 0.0914 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.25, build 20.10.25-0ubuntu1~20.04.1 | +| Operating System | Ubuntu 24.04.2 LTS (Noble Numbat) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/deepseek-r1-distill-qwen-1.5b-iluvatar-tree_0.5.1_iluvatar3.1-gems_5.0.2-vllm_0.13.0_cu102-plugin_0.1.1-cx_none-python_3.10.18-torch_2.7.1_corex.4.4.0-pcp_cuda10.2-gpu_biv150-arc_amd64-driver_4.4.0:2606251110 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS --local_dir /data/DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS +``` + +### Start the Container +```bash +docker run -it --name flagos --network=host --privileged=true --shm-size=16g -v /data:/data -itd harbor.baai.ac.cn/external-cooperation/deepseek-r1-distill-qwen-1.5b-iluvatar-tree_0.5.1_iluvatar3.1-gems_5.0.2-vllm_0.13.0_cu102-plugin_0.1.1-cx_none-python_3.10.18-torch_2.7.1_corex.4.4.0-pcp_cuda10.2-gpu_biv150-arc_amd64-driver_4.4.0:2606251110 bash +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +CUDA_VISIBLE_DEVICES=0 VLLM_PLUGINS=fl USE_FLAGGEMS=1 VLLM_FL_FLAGOS_WHITELIST=max,index,argmax,where_self,where_self_out,gather,lt,lt_scalar,le,scatter,arange_start,ones,full,fill_scalar_,rand_like,zeros,zero_,exponential_,cat,to_copy vllm serve --model /data/DeepSeek-R1-Distill-Qwen-1.5B-hygon-FlagOS --served-model-name DeepSeek-R1-Distill-Qwen-1.5B --port 8000 --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "DeepSeek-R1-Distill-Qwen-1.5B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS.md new file mode 100644 index 000000000..f9e7c8efe --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS.md @@ -0,0 +1,151 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- + +# Introduction + +We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero is trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as an initial stage, and it delivers outstanding reasoning capabilities. Through RL training, DeepSeek-R1-Zero naturally exhibits numerous powerful and intriguing reasoning behaviors. + +Nevertheless, DeepSeek-R1-Zero suffers from issues such as endless repetition, poor readability, and mixed-language outputs. To address these flaws and further boost reasoning performance, we developed DeepSeek-R1, which incorporates cold-start data prior to the RL phase. DeepSeek-R1 achieves performance comparable to OpenAI o1 on mathematical, coding, and reasoning tasks. + +To support the research community, we have open-sourced DeepSeek-R1-Zero, DeepSeek-R1, as well as six dense models distilled from DeepSeek-R1 based on the Llama and Qwen architectures. DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI o1-mini across a wide range of benchmarks, setting a new state-of-the-art record among dense models. + +### Integrated Deployment + +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes + + ### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + +# Evaluation Results + +## Benchmark Result + +| Metrics | DeepSeek-R1-Distill-Qwen-1.5B-Nvidia-Origin | DeepSeek-R1-Distill-Qwen-1.5B-Nvidia-FlagOS | +| ------------------- | ------------------------------------------- | ------------------------------------------- | +| musr_generative | 0.3320 | 0.3429 | +| mmlu_pro | 0.1747 | 0.1734 | +| aime | 0 | 0 | +| livebench_new | 0.1261 | 0.1308 | +| gpqa_generative_cot | 0.0866 | 0.0901 | + +# User Guide + +Environment Setup + +| Item | Version | +| ---------------- | ------------------------------------ | +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image + +```bash +docker pull harbor.baai.ac.cn/external-cooperation/deepseek-r1-distill-qwen-1.5b-nvidia-tree_0.5.0-gems_0.5.1rc0_vllm_0.13.0-plugin_v0.1.0_vllm0.13.0-cx_none-python_3.12.3-torch_2.9.0.dev20250804_cu128-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.124.06:2607211543 +``` + +### Download Open-source Model Weights + +```bash +pip install modelscope +modelscope download --model DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS --local_dir /data/vllm-plugin-fl/DeepSeek-R1-Distill-Qwen-1.5B-nvidia-FlagOS +``` + +### Start the Container + +```bash +docker run -itd --name=flagos --gpus=all --network=host -v /data:/data harbor.baai.ac.cn/external-cooperation/deepseek-r1-distill-qwen-1.5b-nvidia-tree_0.5.0-gems_0.5.1rc0_vllm_0.13.0-plugin_v0.1.0_vllm0.13.0-cx_none-python_3.12.3-torch_2.9.0.dev20250804_cu128-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.124.06:2607211543 sleep infinity +docker exec -it flagos /bin/bash +``` + +### Start the Server + +```bash +CUDA_VISIBLE_DEVICES=0 VLLM_PLUGINS=fl USE_FLAGGEMS=1 VLLM_FL_ALLOW_VENDORS=cuda VLLM_FL_FLAGOS_WHITELIST=reciprocal,mul,cos,sin,attention_backend,embedding,rms_norm,rms_norm_forward,addmm,rotary_embedding,rope_forward,rotary_pos_embedding,fused_add_rms_norm,silu_and_mul,index,rand_like,full,argmax,sort,sort_stable,gather,lt,le vllm serve --model /data/vllm-plugin-fl/DeepSeek-R1-Distill-Qwen-1.5B --served-model-name DeepSeek-R1-Distill-Qwen-1.5B --port 8000 --enforce-eager +``` + +## Service Invocation + +### Invocation Script + +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "DeepSeek-R1-Distill-Qwen-1.5B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response + + # Technical Overview + + **FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + + ## FlagGems + + FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. + + ## FlagTree + + FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. + + ## FlagScale and vllm-plugin-fl + + Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. + vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. + + ## **FlagCX** + + FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + +- **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. +- **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + + # License + + The model weights are derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-hygon-FlagOS.md new file mode 100644 index 000000000..03a38c9d4 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-hygon-FlagOS.md @@ -0,0 +1,137 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction + Based on the Wenxin ERNIE 4.5 0.3B pre-trained model (ERNIE-4.5-0.3B-PT), this is a lightweight large language model adapted and optimized on the Hygon DCU platform with the FlagOS system. It inherits the powerful semantic understanding capabilities of ERNIE 4.5, while achieving efficient inference deployment on Hygon domestic hardware. It is suitable for various natural language processing tasks such as text generation, question-answering dialogue, and code understanding in resource-constrained scenarios. +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | ERNIE-4.5-0.3B-PT-Nvidia-Origin | ERNIE-4.5-0.3B-PT-Hygon-FlagOS | +|---------------------|---------------------------------|--------------------------------| +| musr_generative | 0.377 | 0.3664 | +| mmlu_pro | 0.1737 | 0.176 | +| gpqa_generative_cot | 0.25 | 0.224 | +| livebench_new | 0.1495 | 0.1377 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.24, build 297e128 | +| Operating System | Sugon OS 8.9 | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-0.3b-pt-hygon-tree_0.5.0-gems_5.0.0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10-torch_2.9.0-pcp_hygon-dtk_26.04-gpu_bw200-arc_hygon:2606301452 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/ERNIE-4.5-0.3B-PT-hygon-FlagOS --local_dir /data/vllm-plugin-fl/ERNIE-4.5-0.3B-PT-hygon-FlagOS +``` + +### Start the Container +```bash +docker run -d \ + --name ernie-4.5-0.3b-flagos \ + --privileged \ + --net host \ + --ipc=host \ + --group-add video \ + --cap-add=SYS_PTRACE \ + --security-opt seccomp=unconfined \ + --device=/dev/kfd \ + --device=/dev/mkfd \ + --device=/dev/dri \ + -v /opt/hyhal:/opt/hyhal:ro \ + -v /data/vllm-plugin-fl:/data/vllm-plugin-fl \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-0.3b-pt-hygon-tree_0.5.0-gems_5.0.0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10-torch_2.9.0-pcp_hygon-dtk_26.04-gpu_bw200-arc_hygon:2606301452 \ + sleep infinity +docker exec -it ernie-4.5-0.3b-flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +nohup vllm serve /data/vllm-plugin-fl/ERNIE-4.5-0.3B-PT-hygon-FlagOS \ +--served-model-name ernie-4.5-0.3b-flagos \ +--port 9010 \ +--tensor-parallel-size 1 \ +--enforce-eager \ +--trust-remote-code \ +> ernie_flagos.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:9010/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ernie-4.5-0.3b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PaddlePaddle/ERNIE-4.5-0.3B-PT and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-nvidia-FlagOS.md new file mode 100644 index 000000000..e834a7337 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT-nvidia-FlagOS.md @@ -0,0 +1,140 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction + The ERNIE 4.5 0.3B series is a highly efficient, lightweight pre-trained language model designed for low-latency inference and edge deployment scenarios. This version has been specifically optimized using the **FlagOS-Nvidia** software stack, integrating high-performance operators via **FlagGems** and a unified compiler via **FlagTree** to maximize throughput on NVIDIA GPUs. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | ERNIE-4.5-0.3B-PT-Nvidia-Origin | ERNIE-4.5-0.3B-PT-Nvidia-FlagOS | +|-----------------------------|---------------------------------|--------------------------------------| +| musr_generative | 0.3056 | 0.2923 | +| mmlu_pro | 0.1248 | 0.1268 | +| gpqa_diamond_generative_cot | 0.2374 | 0.2374 | +| gpqa_generative_cot | 0.2013 | 0.1854 | +| livebench_new | 0.1351 | 0.1436 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-0.3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.2-vllm_0.13.0-plugin_0.0.0-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2605130721 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/ERNIE-4.5-0.3B-PT-nvidia-FlagOS --local_dir /data/models/ERNIE-4.5-0.3B-PT-nvidia-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name ernie4_5_03b_test \ + --network=host \ + --ipc=host \ + --privileged=true \ + --ulimit memlock=-1 \ + --shm-size=32G \ + --gpus all \ + -v /data/models:/data/models \ + -itd \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-0.3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.2-vllm_0.13.0-plugin_0.0.0-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2605130721 \ + sleep infinity + +docker exec -it ernie4_5_03b_test bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +nohup vllm serve /data/models/ERNIE-4.5-0.3B-PT-nvidia-FlagOS \ +--served-model-name ernie-4.5-0.3b-flagos \ +--port 8000 \ +--tensor-parallel-size 1 \ +--trust-remote-code \ +> ernie_flagos.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ernie-4.5-0.3b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PaddlePaddle/ERNIE-4.5-0.3B-PT and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT.md new file mode 100644 index 000000000..fd428031c --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-0.3B-PT.md @@ -0,0 +1,134 @@ +--- +license: apache-2.0 +language: +- en +- zh +pipeline_tag: text-generation +tags: +- ERNIE4.5 +library_name: transformers +--- + + + + + +# ERNIE-4.5-0.3B + +> [!NOTE] +> Note: "**-Paddle**" models use [PaddlePaddle](https://github.com/PaddlePaddle/Paddle) weights, while "**-PT**" models use Transformer-style PyTorch weights. + +## ERNIE 4.5 Highlights + +The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations: + +1. **Multimodal Heterogeneous MoE Pre-Training:** Our models are jointly trained on both textual and visual modalities to better capture the nuances of multimodal information and improve performance on tasks involving text understanding and generation, image understanding, and cross-modal reasoning. To achieve this without one modality hindering the learning of another, we designed a *heterogeneous MoE structure*, incorporated *modality-isolated routing*, and employed *router orthogonal loss* and *multimodal token-balanced loss*. These architectural choices ensure that both modalities are effectively represented, allowing for mutual reinforcement during training. + +2. **Scaling-Efficient Infrastructure:** We propose a novel heterogeneous hybrid parallelism and hierarchical load balancing strategy for efficient training of ERNIE 4.5 models. By using intra-node expert parallelism, memory-efficient pipeline scheduling, FP8 mixed-precision training and finegrained recomputation methods, we achieve remarkable pre-training throughput. For inference, we propose *multi-expert parallel collaboration* method and *convolutional code quantization* algorithm to achieve 4-bit/2-bit lossless quantization. Furthermore, we introduce PD disaggregation with dynamic role switching for effective resource utilization to enhance inference performance for ERNIE 4.5 MoE models. Built on [PaddlePaddle](https://github.com/PaddlePaddle/Paddle), ERNIE 4.5 delivers high-performance inference across a wide range of hardware platforms. + +3. **Modality-Specific Post-Training:** To meet the diverse requirements of real-world applications, we fine-tuned variants of the pre-trained model for specific modalities. Our LLMs are optimized for general-purpose language understanding and generation. The VLMs focuses on visuallanguage understanding and supports both thinking and non-thinking modes. Each model employed a combination of *Supervised Fine-tuning (SFT)*, *Direct Preference Optimization (DPO)* or a modified reinforcement learning method named *Unified Preference Optimization (UPO)* for post-training. + +## Model Overview + +ERNIE-4.5-0.3B is a text dense Post-trained model. The following are the model configuration details: + +| Key | Value | +| -------------- | ------------ | +| Modality | Text | +| Training Stage | Posttraining | +| Params | 0.36B | +| Layers | 18 | +| Heads(Q/KV) | 16 / 2 | +| Context Length | 131072 | + +## Quickstart + +### Using `transformers` library + +**Note**: Before using the model, please ensure you have the `transformers` library installed +(upcoming version 4.54.0 or [the latest version](https://github.com/huggingface/transformers?tab=readme-ov-file#installation)) + +The following contains a code snippet illustrating how to use the model generate content based on given inputs. + +```python +import torch +from transformers import AutoModelForCausalLM, AutoTokenizer + +model_name = "baidu/ERNIE-4.5-0.3B-PT" + +# load the tokenizer and the model +tokenizer = AutoTokenizer.from_pretrained(model_name) +model = AutoModelForCausalLM.from_pretrained( + model_name, + device_map="auto", + torch_dtype=torch.bfloat16, +) + +# prepare the model input +prompt = "Give me a short introduction to large language model." +messages = [ + {"role": "user", "content": prompt} +] +text = tokenizer.apply_chat_template( + messages, + tokenize=False, + add_generation_prompt=True +) +model_inputs = tokenizer([text], add_special_tokens=False, return_tensors="pt").to(model.device) + +# conduct text completion +generated_ids = model.generate( + **model_inputs, + max_new_tokens=1024 +) +output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() + +# decode the generated ids +generate_text = tokenizer.decode(output_ids, skip_special_tokens=True) +print("generate_text:", generate_text) +``` + +### vLLM inference + +[vllm](https://github.com/vllm-project/vllm/tree/main) github library. Python-only [build](https://docs.vllm.ai/en/latest/getting_started/installation/gpu.html#set-up-using-python-only-build-without-compilation). + +```bash +vllm serve baidu/ERNIE-4.5-0.3B-PT +``` + +## License + +The ERNIE 4.5 models are provided under the Apache License 2.0. This license permits commercial use, subject to its terms and conditions. Copyright (c) 2025 Baidu, Inc. All Rights Reserved. + +## Citation + +If you find ERNIE 4.5 useful or wish to use it in your projects, please kindly cite our technical report: + +```bibtex +@misc{ernie2025technicalreport, + title={ERNIE 4.5 Technical Report}, + author={Baidu ERNIE Team}, + year={2025}, + eprint={}, + archivePrefix={arXiv}, + primaryClass={cs.CL}, + url={} +} +``` diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md new file mode 100644 index 000000000..f441eacac --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-hygon-FlagOS.md @@ -0,0 +1,145 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations. + +1. **Multimodal Heterogeneous MoE Pre-Training:** The models are jointly trained on both textual and visual modalities to better capture the nuances of multimodal information and improve performance on tasks involving text understanding and generation, image understanding, and cross-modal reasoning. To achieve this without one modality hindering the learning of another, a *heterogeneous MoE structure* was designed, incorporating *modality-isolated routing*, *router orthogonal loss*, and *multimodal token-balanced loss*. These architectural choices ensure that both modalities are effectively represented, allowing for mutual reinforcement during training. + +2. **Scaling-Efficient Infrastructure:** A novel heterogeneous hybrid parallelism and hierarchical load balancing strategy is introduced for efficient training of ERNIE 4.5 models. By utilizing intra-node expert parallelism, memory-efficient pipeline scheduling, FP8 mixed-precision training, and fine-grained recomputation methods, high pre-training throughput is achieved. For inference, a *multi-expert parallel collaboration* method and a *convolutional code quantization* algorithm are employed to achieve 4-bit/2-bit lossless quantization. Furthermore, PD disaggregation with dynamic role switching is introduced for effective resource utilization, enhancing inference performance for ERNIE 4.5 MoE models. Built on [PaddlePaddle](GitHub - PaddlePaddle/Paddle: PArallel Distributed Deep LEarning: Machine Learning Framework from In), ERNIE 4.5 delivers high-performance inference across a wide range of hardware platforms. + +3. **Modality-Specific Post-Training:** To meet the diverse requirements of real-world applications, variants of the pre-trained model are fine-tuned for specific modalities. The LLMs are optimized for general-purpose language understanding and generation, while the VLMs focus on vision-language understanding and support both thinking and non-thinking modes. Each model employs a combination of *Supervised Fine-tuning (SFT)*, *Direct Preference Optimization (DPO)*, or a modified reinforcement learning method named *Unified Preference Optimization (UPO)* during post-training. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | ERNIE-4.5-21B-A3B-PT-Nvidia-Origin | ERNIE-4.5-21B-A3B-PT-Hygon-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| aime | 0.3667 | 0.4 | +| gpqa_generative_cot | 0.5713 | 0.5831 | +| mmlu_pro | 0.6644 | 0.6618 | +| musr_generative | 0.6296 | 0.6495 | +| livebench_new | 0.5563 | 0.5537 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.5, build 55c4c88 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607021138 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/ERNIE-4.5-21B-A3B-PT-hygon-FlagOS --local_dir /data/ERNIE-4.5-21B-A3B-PT-hygon-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name flagos \ + --network=host \ + --ipc=host \ + --device=/dev/kfd \ + --device=/dev/mkfd \ + --device=/dev/dri \ + -v /opt/hyhal:/opt/hyhal \ + -v /data:/data/models \ + --group-add video \ + --cap-add=SYS_PTRACE \ + --security-opt seccomp=unconfined \ + -itd \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607021138 \ + sleep infinity +docker exec -it flagos bash +``` +### Start the Server +```bash +nohup env GEMS_VENDOR=hygon \ + VLLM_PLUGINS=fl \ + USE_FLAGGEMS=1 \ + vllm serve --model /data/models/ERNIE-4.5-21B-A3B-PT-hygon-FlagOS \ + --tensor-parallel-size 1 \ + --enforce-eager \ + --served-model-name ernie-4.5-21b-a3b-pt-flagos \ + --port 8000 \ + > serve.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ernie-4.5-21b-a3b-pt-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PaddlePaddle/ERNIE-4.5-21B-A3B-PT and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md new file mode 100644 index 000000000..58057d020 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS.md @@ -0,0 +1,147 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations: + +1. **Multimodal Heterogeneous MoE Pre-Training:** The models are jointly trained on both textual and visual modalities to better capture the nuances of multimodal information and improve performance on tasks involving text understanding and generation, image understanding, and cross-modal reasoning. To achieve this without one modality hindering the learning of another, a *heterogeneous MoE structure* was designed, incorporating *modality-isolated routing*, *router orthogonal loss*, and *multimodal token-balanced loss*. These architectural choices ensure that both modalities are effectively represented, allowing for mutual reinforcement during training. + +2. **Scaling-Efficient Infrastructure:** A novel heterogeneous hybrid parallelism and hierarchical load balancing strategy is introduced for efficient training of ERNIE 4.5 models. By utilizing intra-node expert parallelism, memory-efficient pipeline scheduling, FP8 mixed-precision training, and fine-grained recomputation methods, high pre-training throughput is achieved. For inference, a *multi-expert parallel collaboration* method and a *convolutional code quantization* algorithm are employed to achieve 4-bit/2-bit lossless quantization. Furthermore, PD disaggregation with dynamic role switching is introduced for effective resource utilization, enhancing inference performance for ERNIE 4.5 MoE models. Built on [PaddlePaddle](https://github.com/PaddlePaddle/Paddle), ERNIE 4.5 delivers high-performance inference across a wide range of hardware platforms. + +3. **Modality-Specific Post-Training:** To meet the diverse requirements of real-world applications, variants of the pre-trained model are fine-tuned for specific modalities. The LLMs are optimized for general-purpose language understanding and generation, while the VLMs focus on vision-language understanding and support both thinking and non-thinking modes. Each model employs a combination of *Supervised Fine-tuning (SFT)*, *Direct Preference Optimization (DPO)*, or a modified reinforcement learning method named *Unified Preference Optimization (UPO)* during post-training. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | ERNIE-4.5-21B-A3B-PT-Nvidia-Origin | ERNIE-4.5-21B-A3B-PT-Nvidia-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| aime | 0.3667 | 0.4 | +| gpqa_generative_cot | 0.5713 | 0.5881 | +| musr_generative | 0.6296 | 0.6548 | +| mmlu_pro | 0.6644 | 0.6619 | +| livebench_new | 0.5563 | 0.5499 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.158.01:2605111355 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS --local_dir /data/ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name flagos \ + --network=host \ + --privileged \ + --shm-size=32G \ + --gpus all \ + -v /data:/data/models \ + -itd \ + harbor.baai.ac.cn/external-cooperation/ernie-4.5-21b-a3b-pt-nvidia-tree_0.5.0_3.5-gems_5.0.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.158.01:2605111355 \ + sleep infinity +docker exec -it flagos bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 + +vllm serve \ + --model /data/models/ERNIE-4.5-21B-A3B-PT-nvidia-FlagOS \ + --tensor-parallel-size 1 \ + --port 8000 \ + --enforce-eager \ + --served-model-name ernie-4.5-21b-a3b-pt-flagos +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "ernie-4.5-21b-a3b-pt-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from PaddlePaddle/ERNIE-4.5-21B-A3B-PT and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-hygon-FlagOS.md new file mode 100644 index 000000000..1110bf4f6 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-hygon-FlagOS.md @@ -0,0 +1,155 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforcement learning, we enhanced the model’s performance in instruction following, engineering code, and function calling, thus strengthening the atomic capabilities required for agent tasks. GLM-4-32B-0414 achieves good results in engineering code, Artifact generation, function calling, search-based Q&A, and report generation. In particular, on several benchmarks, such as code generation or specific Q&A tasks, GLM-4-32B-Base-0414 achieves comparable performance with those larger models like GPT-4o and DeepSeek-V3-0324 (671B). +GLM-Z1-32B-0414 is a reasoning model with deep thinking capabilities. This was developed based on GLM-4-32B-0414 through cold start, extended reinforcement learning, and further training on tasks including mathematics, code, and logic. Compared to the base model, GLM-Z1-32B-0414 significantly improves mathematical abilities and the capability to solve complex tasks. During training, we also introduced general reinforcement learning based on pairwise ranking feedback, which enhances the model's general capabilities. +GLM-Z1-Rumination-32B-0414 is a deep reasoning model with rumination capabilities (against OpenAI's Deep Research). Unlike typical deep thinking models, the rumination model is capable of deeper and longer thinking to solve more open-ended and complex problems (e.g., writing a comparative analysis of AI development in two cities and their future development plans). Z1-Rumination is trained through scaling end-to-end reinforcement learning with responses graded by the ground truth answers or rubrics and can make use of search tools during its deep thinking process to handle complex tasks. The model shows significant improvements in research-style writing and complex tasks. +Finally, GLM-Z1-9B-0414 is a surprise. We employed all the aforementioned techniques to train a small model (9B). GLM-Z1-9B-0414 exhibits excellent capabilities in mathematical reasoning and general tasks. Its overall performance is top-ranked among all open-source models of the same size. Especially in resource-constrained scenarios, this model achieves an excellent balance between efficiency and effectiveness, providing a powerful option for users seeking lightweight deployment. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-4-32B-Base-0414-Nvidia-Origin | GLM-4-32B-Base-0414-Hygon-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| mmlu | 0.7710 | 0.7705 | +| cmmlu | 0.8325 | 0.8314 | +| gsm8k | 0.8719 | 0.8635 | +| leaderboard_bbh | 0.6574 | 0.6558 | +| hellaswag | 0.6469 | 0.6476 | +| truthfulqa_mc1 | 0.3268 | 0.3293 | +| winogrande | 0.7908 | 0.7877 | +| commonsense_qa | 0.7715 | 0.7699 | +| piqa | 0.8194 | 0.8183 | +| openbookqa | 0.368 | 0.370 | +| boolq | 0.8801 | 0.8789 | +| arc_easy | 0.8561 | 0.8544 | +| arc_challenge | 0.5922 | 0.5922 | +| minerva_math_algebra | 0.6731 | 0.6807 | +| ceval-valid | 0.8076 | 0.8105 | +| pubmedqa | 0.788 | 0.788 | +| medqa_4options | 0.729 | 0.7298 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.5, build 55c4c88 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607030559 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-4-32B-Base-0414-hygon-FlagOS --local_dir /data/GLM-4-32B-Base-0414-hygon-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name flagos \ + --network=host \ + --ipc=host \ + --device=/dev/kfd \ + --device=/dev/mkfd \ + --device=/dev/dri \ + -v /opt/hyhal:/opt/hyhal \ + -v /data:/data/models \ + --group-add video \ + --cap-add=SYS_PTRACE \ + --security-opt seccomp=unconfined \ + -itd \ + harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607030559 \ + sleep infinity +docker exec -it flagos bash +``` +### Start the Server +```bash +nohup env GEMS_VENDOR=hygon \ + VLLM_PLUGINS=fl \ + USE_FLAGGEMS=1 \ + VLLM_FL_FLAGOS_WHITELIST=rms_norm,rotary_embedding,arange_start,rand_like,attention_backend,ones,reciprocal,embedding,cos,index,rsub_scalar,argmax,zeros,softmax,sum_dim,zero_,cumsum,gather,sin,le,lt,lt_scalar,cumsum_out,where_self,scatter,full \ + vllm serve --model /data/models/GLM-4-32B-Base-0414-hygon-FlagOS \ + --tensor-parallel-size 2 \ + --enforce-eager \ + --served-model-name glm-4-32b-base-0414-flagos \ + --port 8000 \ + > serve.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "glm-4-32b-base-0414-flagos", + "prompt": "你好" + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from ZhipuAI/GLM-4-32B-Base-0414 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md new file mode 100644 index 000000000..557641feb --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_GLM-4-32B-Base-0414-nvidia-FlagOS.md @@ -0,0 +1,137 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforcement learning, we enhanced the model’s performance in instruction following, engineering code, and function calling, thus strengthening the atomic capabilities required for agent tasks. GLM-4-32B-0414 achieves good results in engineering code, Artifact generation, function calling, search-based Q&A, and report generation. In particular, on several benchmarks, such as code generation or specific Q&A tasks, GLM-4-32B-Base-0414 achieves comparable performance with those larger models like GPT-4o and DeepSeek-V3-0324 (671B). +GLM-Z1-32B-0414 is a reasoning model with deep thinking capabilities. This was developed based on GLM-4-32B-0414 through cold start, extended reinforcement learning, and further training on tasks including mathematics, code, and logic. Compared to the base model, GLM-Z1-32B-0414 significantly improves mathematical abilities and the capability to solve complex tasks. During training, we also introduced general reinforcement learning based on pairwise ranking feedback, which enhances the model's general capabilities. +GLM-Z1-Rumination-32B-0414 is a deep reasoning model with rumination capabilities (against OpenAI's Deep Research). Unlike typical deep thinking models, the rumination model is capable of deeper and longer thinking to solve more open-ended and complex problems (e.g., writing a comparative analysis of AI development in two cities and their future development plans). Z1-Rumination is trained through scaling end-to-end reinforcement learning with responses graded by the ground truth answers or rubrics and can make use of search tools during its deep thinking process to handle complex tasks. The model shows significant improvements in research-style writing and complex tasks. +Finally, GLM-Z1-9B-0414 is a surprise. We employed all the aforementioned techniques to train a small model (9B). GLM-Z1-9B-0414 exhibits excellent capabilities in mathematical reasoning and general tasks. Its overall performance is top-ranked among all open-source models of the same size. Especially in resource-constrained scenarios, this model achieves an excellent balance between efficiency and effectiveness, providing a powerful option for users seeking lightweight deployment. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | GLM-4-32B-Base-0414-Nvidia-Origin | GLM-4-32B-Base-0414-Nvidia-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| mmlu | 0.7718 | 0.7716 | +| cmmlu | 0.8328 | 0.8325 | +| gsm8k | 0.8643 | 0.8613 | +| leaderboard_bbh | 0.6556 | 0.6579 | +| hellaswag | 0.6481 | 0.6480 | +| truthfulqa_mc1 | 0.328 | 0.3293 | +| winogrande | 0.7916 | 0.7916 | +| commonsense_qa | 0.7731 | 0.7731 | +| piqa | 0.8199 | 0.8194 | +| openbookqa | 0.368 | 0.368 | +| boolq | 0.8798 | 0.8786 | +| arc_easy | 0.8552 | 0.8552 | +| arc_challenge | 0.5922 | 0.5913 | +| minerva_math_algebra | 0.6790 | 0.7035 | +| ceval-valid | 0.8113 | 0.8105 | +| pubmedqa | 0.788 | 0.788 | +| medqa_4options | 0.7337 | 0.7321 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.4.3, build 055a478 | +| Operating System | 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-nvidia-tree_0.5.0-gems_0.5.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2606091434 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/GLM-4-32B-Base-0414-nvidia-FlagOS --local_dir /data/GLM-4-32B-Base-0414-nvidia-FlagOS +``` + +### Start the Container +```bash +docker run -itd \ + --name=flagos \ + --gpus all \ + --network host \ + --ipc host \ + --privileged=true \ + --shm-size=32G \ + -v /data:/data/models \ + harbor.baai.ac.cn/external-cooperation/glm-4-32b-base-0414-nvidia-tree_0.5.0-gems_0.5.1rc0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda13.2-gpu_nvidia003-arc_amd64-driver_570.133.20:2606091434 \ + sleep infinity +docker exec -it flagos bash +``` +### Start the Server +```bash +nohup env VLLM_PLUGINS=fl USE_FLAGGEMS=1 \ + VLLM_FL_FLAGOS_BLACKLIST=mul,to_copy,cat,silu_and_mul,sub,cos,rms_norm,fill_scalar_,sort,attention_backend \ + vllm serve --model /data/models/GLM-4-32B-Base-0414-nvidia-FlagOS \ + --tensor-parallel-size 4 \ + --served-model-name glm-4-32b-base-0414-flagos \ + --port 8000 \ + > serve.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "glm-4-32b-base-0414-flagos", + "prompt": "你好" + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. +# License +The model weights are derived from ZhipuAI/GLM-4-32B-Base-0414 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_HY-MT2-7B-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_HY-MT2-7B-mthreads-FlagOS.md index 3913f1229..f2e9f8505 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_HY-MT2-7B-mthreads-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_HY-MT2-7B-mthreads-FlagOS.md @@ -68,7 +68,7 @@ docker run -d \ -v /lib/x86_64-linux-gnu:/lib/x86_64-linux-gnu \ -v /etc/alternatives:/etc/alternatives \ -v /etc/localtime:/etc/localtime \ - flagtree-mthreads3.6-test_wys:latest \ + harbor.baai.ac.cn/flagrelease-public/flagrelease-hy-mt2-7b-mthreads-tree_0.5.1_mthreads3.6-gems_5.0.2-vllm_0.13.1.dev44_g3d4cc4bc7.d20260310.musa-plugin_0.1.0-cx_0.8.0-python_3.10.12-torch_2.7.1-pcp_musa4.3.5-driver_3.3.6:202605281243 \ bash docker exec -it flagos bash diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-hygon-FlagOS.md new file mode 100644 index 000000000..2d8d454e9 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-hygon-FlagOS.md @@ -0,0 +1,156 @@ +--- +base_model: +- "" +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Hygon-FlagOS | +|--------------|-------------------|------------------| +| GPQA_Diamond | 66.16 | 63.78 | +| arc_challenge_chat | 96.33 | 96.5 | +| math_500 | 94.6 | 94.2 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.5, build 55c4c88 | +| Operating System | Ubuntu 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-hygon001-gems5.4.0-treenone-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-hygon-FlagOS --local_dir /data/Hy3 +``` + +### Start the Container +```bash +docker run --name flagos --network=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri -v /opt/hyhal:/opt/hyhal -v /data:/data --group-add video --cap-add=SYS_PTRACE --security-opt seccomp=unconfined -itd harbor.baai.ac.cn/flagrelease-public/hy3-hygon001-gems5.4.0-treenone-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt210-dtknone-x64-6.3.30-v1.4.1a bin/bash +docker exec -it flagos /bin/bash +``` +### Start the Server +master, fix NCCL/GLOO to your specific configs. +```bash +export VLLM_PLUGINS='fl' +export NCCL_IB_DISABLE=1 +export NCCL_SOCKET_IFNAME=bond1 +export GLOO_SOCKET_IFNAME=bond1 + +vllm serve /data/Hy3 \ + -tp 16 \ + --max-model-len 131072 \ + --served-model-name hygonhyv3 \ + --port 9010 \ + --nnodes 2 \ + --node-rank 0 \ + --master-addr `` \ + --reasoning-parser hy_v3 \ + --enforce-eager +``` + +worker + +```bash +export VLLM_PLUGINS='fl' +export NCCL_IB_DISABLE=1 +export NCCL_SOCKET_IFNAME=bond1 +export GLOO_SOCKET_IFNAME=bond1 + +vllm serve /data/Hy3 \ + -tp 16 \ + --max-model-len 131072 \ + --served-model-name hygonhyv3 \ + --port 9010 \ + --nnodes 2 \ + --node-rank 1 \ + --headless \ + --master-addr `` \ + --reasoning-parser hy_v3 \ + --enforce-eager +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:9010/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "hygonhyv3", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-iluvatar-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-iluvatar-FlagOS.md new file mode 100644 index 000000000..73df9f03d --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-iluvatar-FlagOS.md @@ -0,0 +1,121 @@ +--- +base_model: +- "" +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Iluvatar** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Iluvatar-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 83.33 | 81.82 | +| arc_challenge_chat | 96.33 | 95.82 | +| math_500 | 94.6 | 90.6 | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.25, build 20.10.25-0ubuntu1~20.04.1 | +| Operating System | Ubuntu 20.04.6 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-iluvatar001-gems5.0.2-treenone-cxnone-pluginnone-vllm0.17.0-cp312-pt27-ixml44-x64-4.4.0 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-iluvatar-FlagOS --local_dir /data/Hy3 +``` + +### Start the Container +```bash +docker run -dit -v /mnt:/mnt -v /usr/src:/usr/src -v /lib/modules:/lib/modules -v /dev:/dev -v /data:/data -v /home:/home --network=host --name=flagos --ipc=host --privileged --cap-add=ALL --pid=host harbor.baai.ac.cn/flagrelease-public/hy3-iluvatar001-gems5.0.2-treenone-cxnone-pluginnone-vllm0.17.0-cp312-pt27-ixml44-x64-4.4.0 /bin/bash +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +VLLM_W8A8_MOE_USE_W4A8=1 vllm serve /data/Hy3 -tp 16 -dp 1 --port 9010 --enforce-eager --reasoning-parser hy_v3 --served-model-name hyv3iluvatarw4a8 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:9010/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "hyv3iluvatarw4a8", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-metax-FlagOS.md new file mode 100644 index 000000000..d76bc9a5e --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-metax-FlagOS.md @@ -0,0 +1,171 @@ +--- +base_model: +- "" +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Metax-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 83.33 | 84.85 | +| arc_challenge_chat | 96.33 | 96.59 | +| math_500 | 94.6 | 95.2 | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 28.0.4, build b8034c0 | +| Operating System | Ubuntu 22.04.4 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.8.1:202607061058 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-metax-FlagOS --local_dir /data/Hy3 +``` + +### Start the Container +```bash +docker run -d --name flagos --network host --shm-size 64g --device /dev/dri:/dev/dri:rwm --device /dev/mxcd:/dev/mxcd:rwm -v /data:/data harbor.baai.ac.cn/flagrelease-public/hy3-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.8.1:202607061058 sleep infinity +docker exec -it flagos /bin/bash +``` +### Start the Server +master, fix MCCL/NCCL/GLOO to your specific configs. +```bash +export MCCL_HC_PLUGIN=/opt/maca-3.3.0/lib/libmxccl_21605plugin.so +export MCCL_EXT_CCL_ENABLE=1 +export NCCL_FORCE_USE_MN_MIX=1 +export MCCL_SOCKET_IFNAME=inbond1 +export NCCL_SOCKET_IFNAME=inbond1 +export GLOO_SOCKET_IFNAME=inbond1 +export MCCL_DEBUG=WARN +export MCCL_PROTO=LL +export MCCL_P2P_LL_THRESHOLD=1073741824 +export MCCL_IB_DISABLE=1 +export NCCL_IB_DISABLE=1 +export NCCL_DEBUG=WARN +export MASTER_ADDR=192.168.2.127 +export MASTER_PORT=29501 +export VLLM_FL_FLAGOS_BLACKLIST="flash_attention_forward,mm,mm_out,bmm,addmm,addmm_out,baddbmm" +export VLLM_ENGINE_ITERATION_TIMEOUT_S=72000 +export NCCL_TIMEOUT=7200 +export VLLM_RPC_TIMEOUT=72000000 +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200 + +export GEMS_VENDOR=metax +export VLLM_PLUGINS=fl +vllm serve /data/Hy3 --host 0.0.0.0 --port 9010 --served-model-name metaxhy3 --reasoning-parser hy_v3 --tensor-parallel-size 16 --nnodes 2 --node-rank 0 --master-addr `` +``` + +worker +```bash +export MCCL_HC_PLUGIN=/opt/maca-3.3.0/lib/libmxccl_21605plugin.so +export MCCL_EXT_CCL_ENABLE=1 +export NCCL_FORCE_USE_MN_MIX=1 +export MCCL_SOCKET_IFNAME=inbond1 +export NCCL_SOCKET_IFNAME=inbond1 +export GLOO_SOCKET_IFNAME=inbond1 +export MCCL_DEBUG=WARN +export MCCL_PROTO=LL +export MCCL_P2P_LL_THRESHOLD=1073741824 +export MCCL_IB_DISABLE=1 +export NCCL_IB_DISABLE=1 +export NCCL_DEBUG=WARN +export MASTER_ADDR=192.168.2.127 +export MASTER_PORT=29501 +export VLLM_FL_FLAGOS_BLACKLIST="flash_attention_forward,mm,mm_out,bmm,addmm,addmm_out,baddbmm" +export VLLM_ENGINE_ITERATION_TIMEOUT_S=72000 +export NCCL_TIMEOUT=7200 +export VLLM_RPC_TIMEOUT=72000000 +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=7200 + +export GEMS_VENDOR=metax +export VLLM_PLUGINS=fl +vllm serve /data/Hy3 --host 0.0.0.0 --port 9010 --served-model-name metaxhy3 --reasoning-parser hy_v3 --tensor-parallel-size 16 --nnodes 2 --node-rank 1 --master-addr `` --headless +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:9010/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "metaxhy3", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-mthreads-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-mthreads-FlagOS.md new file mode 100644 index 000000000..d3919befe --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-mthreads-FlagOS.md @@ -0,0 +1,150 @@ +--- +base_model: +- "" +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Mthreads** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Mthreads-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 83.33 | 83.33 | +| arc_challenge_chat | 96.33 | 96.5 | +| math_500 | 94.6 | 91 | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 9f9e405 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-mthreads001-gems5.4.0-treenone-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607021123 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-mthreads-FlagOS --local_dir /data/Hy3 +``` + +### Start the Container +```bash +docker run -itd --privileged --net host \ + --name flagos \ + -w /workspace \ + -v /data/:/data/ \ + -v /mnt/:/mnt/ \ + -v /public-ks3/:/public/ \ + --env MTHREADS_VISIBLE_DEVICES=all \ + --shm-size=560g \ + harbor.baai.ac.cn/flagrelease-public/hy3-mthreads001-gems5.4.0-treenone-cxnone-plugin0.2.0-vllm0.20.2-cp310-pt27-musa43-x64-3.3.6-server:202607021123 +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +# in node1 +export VLLM_FL_FLAGOS_WHITELIST=moe_sum,grouped_topk,moe_align_block_size,invoke_fused_moe_triton_kernel,embeddind,rsqrt,index_select,silu_and_mul,rand_like,cumsum +export VLLM_CONFIGURE_LOGGING=1 +vllm serve /data/Hy3 \ + --tensor-parallel-size 8 --pipeline-parallel-size 2\ + --port 8000 \ + --gpu-memory-utilization 0.95 \ + --served-model-name hy3 \ + --nnodes 2 --node-rank 0 \ + --master-addr --master-port 29500 \ + --reasoning-parser hy_v3 +# in node2 +export VLLM_FL_FLAGOS_WHITELIST=moe_sum,grouped_topk,moe_align_block_size,invoke_fused_moe_triton_kernel,embeddind,rsqrt,index_select,silu_and_mul,rand_like,cumsum +export VLLM_CONFIGURE_LOGGING=1 +vllm serve /data/Hy3 \ + --tensor-parallel-size 8 --pipeline-parallel-size 2 \ + --port 8000 \ + --gpu-memory-utilization 0.95 \ + --served-model-name hy3 \ + --nnodes 2 --node-rank 1 \ + --master-addr --master-port 29500 --headless \ + --reasoning-parser hy_v3 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "flagOS", + "messages": [{"role": "user", "content": "hi!"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-nvidia-FlagOS-Express.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-nvidia-FlagOS-Express.md new file mode 100644 index 000000000..bcb8dcc09 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-nvidia-FlagOS-Express.md @@ -0,0 +1,136 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Nvidia-FlagOS | +|--------------|--------------------------------|-------------------| +| GPQA_Diamond | 83.33 | 85.86 | +| arc_challenge_chat | 96.33 | 95.99 | +| math_500 | 94.6 | 93.6 | + +## Performance Benchmark +| Test Scenario | 4k & 1k 64 Concurrent | 16k & 1k 64 Concurrent | 32k & 1k 64 Concurrent | +|----------------------------------------|-----------------------|------------------------|------------------------| +| Speedup Ratio (NV-flagos / NV-native) | 107.19% | 104.59% | 103.23% | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-nvidia003-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt211-cu130-x64-580.95.05:202607021605 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-nvidia-FlagOS --local_dir /data/Hy3 +``` + +### Start the Container +```bash +docker run -itd \ + --name flagos \ + --entrypoint /bin/bash \ + --gpus all \ + --ipc=host \ + --net host \ + --shm-size 512g \ + -v /data/:/data \ + harbor.baai.ac.cn/flagrelease-public/hy3-nvidia003-gems5.4.0-tree0.6.0-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt211-cu130-x64-580.95.05:202607021605 +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_FL_FLAGOS_WHITELIST=invoke_fused_moe_triton_kernel,exponential_ +vllm serve /data/Hy3/ \ + --tensor-parallel-size 8 \ + --port 8000 \ + --gpu-memory-utilization 0.95 \ + --served-model-name hy3 \ + --reasoning-parser hy_v3 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "hy3", + "messages": [{"role": "user", "content": "hi!"}] +}' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-tsingmicro-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-tsingmicro-FlagOS.md new file mode 100644 index 000000000..67ee90372 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-tsingmicro-FlagOS.md @@ -0,0 +1,122 @@ +--- +license: apache-2.0 +language: +- zh +- en +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Tsingmicro** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Tsingmicro-FlagOS | +|--------------|-------------------|-----------------------| +| GPQA_Diamond | 66.16 | Evaluating | +| arc_challenge_chat | 96.33 | Evaluating | +| math_500 | 94.6 | 92 | + + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 29.1.3, build 29.1.3-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.4 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-tsingmicro001-gems4.2.1-treenone-cx0.1.0-plugin0.1.1-vllm0.20.2-cp310-pt24-raisa260629-x64-v260629145101:202607071144 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-tsingmicro-FlagOS --local_dir /data/Hy3 +``` + +### Start the Container +```bash +docker run -d -name flagos --network host --shm-size=128g --privileged --ipc=host -v /dev:/dev -v /tmp:/tmp -v /lib/modules:/lib/modules -v /sys:/sys -v /data:/data -it harbor.baai.ac.cn/flagrelease-public/hy3-tsingmicro001-gems4.2.1-treenone-cx0.1.0-plugin0.1.1-vllm0.20.2-cp310-pt24-raisa260629-x64-v260629145101:202607071144 bash +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +# Enter the inference-related directory +cd /workspace/hy3_env +# Set environment variables, need to execute every time entering the container +source test_env.sh +# Run the inference serve script on 8 machines +bash ./hy3_server.sh +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "flagOS", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-zhenwu-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-zhenwu-FlagOS.md new file mode 100644 index 000000000..253b816e5 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Hy3-zhenwu-FlagOS.md @@ -0,0 +1,121 @@ +--- +base_model: +- "" +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +Hy3 is Tencent Hunyuan team's next-generation MoE large model with integrated fast and slow thinking: 295B total parameters, 21B active parameters (plus 3.8B MTP layer parameters), 192 experts with top-8 activation, supporting 256K context. Compared to the Preview version released in late April, Hy3 has achieved a comprehensive leap in intelligence through incorporating real-world business feedback, scaling up RL compute, and improving post-training data quality, significantly outperforming open-source models of similar size. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Zhenwu** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Hy3-Nvidia-Origin | Hy3-Zhenwu-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | 83.33 | 85.35 | +| arc_challenge_chat | 96.33 | 96.33 | +| math_500 | 94.6 | 93.8 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 28.1.0, build 4d8c241 | +| Operating System | Ubuntu 24.04.2 LTS | + +## Operation Steps +The image for this task is exported from Alibaba Cloud PAI and can be used on Alibaba Cloud EAS and DSW, both of which are container‑based resource services. For detailed instructions on how to use this image, please contact the PAI platform support team. The task released by BAAI is developed based on the container environment launched via the PAI platform. + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/hy3-pp001-gems5.4.0-treenone-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt210-hggc130-x64-1.3.2-d7f5a2:202607021622 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Hy3-zhenwu-FlagOS --local_dir /data/Hy3 +``` + +### Start the Server +```bash +export VLLM_FL_FLAGOS_WHITELIST=attention_backend,grouped_topk,moe_align_block_size,invoke_fused_moe_triton_kernel,silu_and_mul,moe_sum,argmax,rand_like,arange,cat,to,where +export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=3000 +vllm serve /data/Hy3 \ + --tensor-parallel-size 16 \ + --port 8000 \ + --reasoning-parser hy_v3 \ + --served-model-name hy3 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "hy3", + "messages": [{"role": "user", "content": "hi!"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Tencent-Hunyuan/Hy3 and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md new file mode 100644 index 000000000..9e8a79732 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS.md @@ -0,0 +1,182 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction + +Kimi-Linear-48B-A3B-Instruct is a high-efficiency large language model developed by MoonshotAI. Built with an innovative hybrid linear attention architecture and equipped with 48B total parameters, it is specially optimized for long-context comprehension, multi-turn dialogue and complex reasoning scenarios, supporting an ultra-long context window up to 1 million tokens. + +Adopting a 3:1 structural ratio of Kimi Delta Attention and global MLA, this model greatly cuts down KV cache occupancy and improves inference throughput while maintaining strong comprehensive capability. It achieves outstanding results on multiple authoritative benchmarks, natively compatible with Transformers and vLLM frameworks, and can be quickly deployed for long document parsing, knowledge question answering and industrial intelligent conversation services. + +### Integrated Deployment + +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes + +### Consistency Validation + +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results + +## Benchmark Result + +| Metrics | Kimi-Linear-48B-A3B-Instruct-Nvidia-Origin | Kimi-Linear-48B-A3B-Instruct-Hygon-FlagOS | +| ------------------- | ------------------------------------------ | ----------------------------------------- | +| aime | 0.4667 | 0.5667 | +| musr_generative | 0.5926 | 0.5516 | +| mmlu_pro | 0.515 | 0.5266 | +| gpqa_generative_cot | 0.4295 | 0.4253 | +| livebench_new | 0.5438 | 0.5254 | + + +# User Guide + +Environment Setup + +| Item | Version | +| ---------------- | -------------------------------------- | +| Docker Version | Docker version 20.10.24, build 297e128 | +| Operating System | Sugon OS 8.9 | + +## Operation Steps + +### Download FlagOS Image + +```bash +docker pull harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607011028 +``` + +### Download Open-source Model Weights + +```bash +pip install modelscope +modelscope download --model FlagRelease/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS --local_dir /data/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS +``` + +### Start the Container + +```bash +docker run \ + --name flagos \ + --network=host \ + --ipc=host \ + --device=/dev/kfd \ + --device=/dev/mkfd \ + --device=/dev/dri \ + -v /opt/hyhal:/opt/hyhal \ + -v /data:/data \ + --group-add video \ + --cap-add=SYS_PTRACE \ + --security-opt seccomp=unconfined \ + -itd \ + harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607011028 \ + sleep infinity + +docker exec -it flagos bash +``` + +### Start the Server + +```bash +#建议按实际卡号调整 +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +export HIP_VISIBLE_DEVICES=2,3 +export VLLM_FL_FLAGOS_WHITELIST="arange_start,lt,where_self_out,argmax,zeros_like,bitwise_or_tensor,scatter,rsub_scalar,ones,cumsum,bitwise_and_tensor,resolve_neg,lt_scalar,sum_dim,add,diff,index,le,masked_fill,where_self,bitwise_not,gather,mul,zero_,nonzero,resolve_conj,cumsum_out,gt_scalar,softmax_out,softmax" + +ulimit -n 2048 && nohup vllm serve \ +--model /data/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS \ +--served-model-name kimi-linear-48b-a3b-instruct-flagos \ +--host 0.0.0.0 \ +--port 8000 \ +--gpu-memory-utilization 0.90 \ +--trust-remote-code \ +--tensor-parallel-size 2 \ +> kimi_flagos.log 2>&1 & +``` + +## Service Invocation + +### Invocation Script + +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "kimi-linear-48b-a3b-instruct-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response + +# Technical Overview + +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + +## FlagGems + +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. + +## FlagTree + +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. + +## FlagScale and vllm-plugin-fl + +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. + +## **FlagCX** + +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License + +The model weights are derived from moonshotai/Kimi-Linear-48B-A3B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS.md index 3e1c8e9e9..aff94534f 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS.md @@ -1,64 +1,72 @@ -# Introduction +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +### Introduction Kimi-Linear-48B-A3B-Instruct is a high-efficiency large language model developed by MoonshotAI. Built with an innovative hybrid linear attention architecture and equipped with 48B total parameters, it is specially optimized for long-context comprehension, multi-turn dialogue and complex reasoning scenarios, supporting an ultra-long context window up to 1 million tokens. Adopting a 3:1 structural ratio of Kimi Delta Attention and global MLA, this model greatly cuts down KV cache occupancy and improves inference throughput while maintaining strong comprehensive capability. It achieves outstanding results on multiple authoritative benchmarks, natively compatible with Transformers and vLLM frameworks, and can be quickly deployed for long document parsing, knowledge question answering and industrial intelligent conversation services. + ### Integrated Deployment -- Out-of-the-box inference scripts with pre-configured hardware and software parameters -- Released **FlagOS-Metax** container image supporting deployment within minutes +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes + ### Consistency Validation -- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. # Evaluation Results ## Benchmark Result -| Metrics | Kimi-Linear-48B-A3B-Instruct-metax-FlagOS-Nvidia-Origin | Kimi-Linear-48B-A3B-Instruct-metax-FlagOS-Metax-FlagOS | + +| Metrics | Kimi-Linear-48B-A3B-Instruct-Nvidia-Origin | Kimi-Linear-48B-A3B-Instruct-Nvidia-FlagOS | | ------------------- | -------------------------------------------------------- | -------------------------------------------------------- | -| aime | 0.4667 | 0.4620 | -| musr_generative | 0.5926 | 0.5542 | -| mmlu_pro | 0.515 | 0.4784 | -| gpqa_generative_cot | 0.4295 | 0.3985 | -| livebench_new | 0.5438 | 0.5231 | +| aime | 0.4667 | 0.4667 | +| musr_generative | 0.5926 | 0.5635 | +| mmlu_pro | 0.515 | 0.5315 | +| gpqa_generative_cot | 0.4295 | 0.4295 | +| livebench_new | 0.5438 | 0.5178 | # User Guide Environment Setup -| Item | Version | -|------------------|----------------------| -| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | -| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | +| Item | Version | +| ---------------- | ------------------------------------ | +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | ## Operation Steps ### Download FlagOS Image ```bash -docker pull harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2_cu128-pcp_cuda12.8-gpu_metax_c550-arc_amd64-driver_3.3.12:2606081508 +docker pull harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-nvidia-tree_0.5.0_3.5-gems_5.0.2-vllm_0.13.0-plugin_0.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.158.01:2605110300 ``` ### Download Open-source Model Weights ```bash pip install modelscope -modelscope download --model FlagRelease/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS --local_dir /data/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS +modelscope download --model FlagRelease/Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS --local_dir /data/Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS ``` ### Start the Container ```bash -docker run -itd \ - --name=flagos \ - --privileged \ - --network=host \ - -v /data/vllm-plugin-fl:/data/models \ - harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-metax-tree_0.5.1_metax3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.11-torch_2.8.0_metax3.3.0.2_cu128-pcp_cuda12.8-gpu_metax_c550-arc_amd64-driver_3.3.12:2606081436 \ - sleep infinity +docker run -itd --name=flagos --gpus=all --network=host -v /data:/data harbor.baai.ac.cn/external-cooperation/kimi-linear-48b-a3b-instruct-nvidia-tree_0.5.0_3.5-gems_5.0.2-vllm_0.13.0-plugin_0.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.158.01:2605110300 sleep infinity + docker exec -it flagos bash + ``` ### Start the Server @@ -66,20 +74,16 @@ docker exec -it flagos bash ```bash export VLLM_PLUGINS=fl export TRITON_ALL_BLOCKS_PARALLEL=1 -export USE_FLAGGEMS=1 -export CUDA_VISIBLE_DEVICES=0,1 - -export VLLM_FL_FLAGOS_BLACKLIST="sort,mm,mul,masked_fill_" - -ulimit -n 2048 && nohup vllm serve \ ---model /data/Kimi-Linear-48B-A3B-Instruct-metax-FlagOS \ ---served-model-name kimi-linear \ +nohup vllm serve \ +--model /data/Kimi-Linear-48B-A3B-Instruct-nvidia-FlagOS \ +--served-model-name kimi-linear-48b-a3b-instruct \ --host 0.0.0.0 \ --port 8000 \ --trust-remote-code \ --tensor-parallel-size 2 \ --enforce-eager \ -> kimi_flagos.log 2>&1 & +> kimi-flagos.log 2>&1 & + ``` ## Service Invocation @@ -90,7 +94,7 @@ ulimit -n 2048 && nohup vllm serve \ curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ - "model": "kimi-linear-48b-a3b-instruct-flagos", + "model": "kimi-linear-48b-a3b-instruct", "messages": [{"role": "user", "content": "你好"}] }' ``` @@ -144,7 +148,7 @@ FlagCX is a scalable and adaptive cross-chip communication library. It serves as FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. - - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. # Contributing @@ -156,4 +160,7 @@ We warmly welcome global developers to join us: 4. Expand hardware adaptation support # License + The model weights are derived from moonshotai/Kimi-Linear-48B-A3B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_LFM2-2.6B-Exp-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_LFM2-2.6B-Exp-metax-FlagOS.md new file mode 100644 index 000000000..a8cbc6c1c --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_LFM2-2.6B-Exp-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | LFM2-2.6B-Exp-metax-FlagOS-Origin | LFM2-2.6B-Exp-metax-FlagOS-FlagOS | +|--------------|-----------------------------------|-----------------------------------| +| GPQA_Diamond | 0 | 0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/lfm2-2.6b-exp-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607171312-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/LFM2-2.6B-Exp-metax-FlagOS --local_dir /data/LFM2-2.6B-Exp-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name flagos --network=host --device /dev/dri --device /dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/lfm2-2.6b-exp-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607171312-v2 +``` +### Start the Server +```bash +vllm serve /data/LFM2-2.6B-Exp-FlagOS --host 0.0.0.0 --port 8000 --served-model-name LFM2-2.6B-Exp --tensor-parallel-size 1 --max-num-batched-tokens 16384 --max-num-seqs 256 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "LFM2-2.6B-Exp", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from LiquidAI/LFM2-2.6B-Exp and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md new file mode 100644 index 000000000..11fb6faaf --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-hygon-FlagOS.md @@ -0,0 +1,148 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +Meta developed and released the Meta Llama 3 family of Large Language Models (LLMs), a suite of generative text models available in pre-trained and instruction-tuned variants with parameter sizes of 8B and 70B. The instruction-tuned Llama 3 models are optimized for dialogue scenarios and outperform many existing open-source chat models on mainstream industry benchmarks. Additionally, great emphasis was placed on enhancing model helpfulness and safety throughout the development process. +Model Developer: Meta +Variants: Llama 3 comes in two parameter sizes (8B and 70B), with both pre-trained and instruction-tuned releases available. +Input: The model only accepts text inputs. +Output: The model generates only text and code. +Model Architecture: Llama 3 is an autoregressive language model built on an optimized Transformer architecture. Its instruction-tuned variants leverage Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) to align outputs with human preferences regarding helpfulness and safety. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Meta-Llama-3-8B-Instruct-Nvidia-Origin | Meta-Llama-3-8B-Instruct-Hygon-FlagOS | +|--------------------|----------------------------------------|------------------------------------------| +| musr_generative | 0.4524 | 0.4378 | +| mmlu_pro | 0.2174 | 0.2031 | +| aime | 0 | 0 | +| livebench_new | 0.2835 | 0.2768 | +| gpqa_generative_cot| 0.3154 | 0.3188 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.24, build 297e128 | +| Operating System | Sugon OS 8.9 | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/meta-llama-3-8b-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2606300957 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Meta-Llama-3-8B-Instruct-hygon-FlagOS --local_dir /data/models/Meta-Llama-3-8B-Instruct-hygon-FlagOS +``` + +### Start the Container +```bash +docker run \ + --name llama-3-8b-flagos \ + --network=host \ + --ipc=host \ + --device=/dev/kfd \ + --device=/dev/mkfd \ + --device=/dev/dri \ + -v /opt/hyhal:/opt/hyhal \ + -v /data/models:/data/models \ + --group-add video \ + --cap-add=SYS_PTRACE \ + --security-opt seccomp=unconfined \ + -itd \ + harbor.baai.ac.cn/external-cooperation/meta-llama-3-8b-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2606300957 \ + sleep infinity +docker exec -it llama-3-8b-flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +export VLLM_FL_FLAGOS_WHITELIST="softmax,rms_norm,add,sub,gather,masked_fill_,cumsum,cumsum_out,lt,lt_scalar,where_self,where_self_out,sum_dim,arange_start,zero_,zeros,ones,full,rand_like,index,reciprocal,cos,sin,cat,to_copy,argmax,le,scatter" +nohup vllm serve /data/models/Meta-Llama-3-8B-Instruct-hygon-FlagOS \ + --served-model-name llama-3-8b-flagos \ + --port 8000 \ + --max-num-batched-tokens 2048 \ + --enforce-eager \ + > fl_serve.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "llama-3-8b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configurexport +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from LLM-Research/Meta-Llama-3-8B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-nvidia-FlagOS.md new file mode 100644 index 000000000..ddfdf4f9a --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Meta-Llama-3-8B-Instruct-nvidia-FlagOS.md @@ -0,0 +1,132 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +Meta developed and released the Meta Llama 3 family of Large Language Models (LLMs), a suite of generative text models available in pre-trained and instruction-tuned variants with parameter sizes of 8B and 70B. The instruction-tuned Llama 3 models are optimized for dialogue scenarios and outperform many existing open-source chat models on mainstream industry benchmarks. Additionally, great emphasis was placed on enhancing model helpfulness and safety throughout the development process. +Model Developer: Meta +Variants: Llama 3 comes in two parameter sizes (8B and 70B), with both pre-trained and instruction-tuned releases available. +Input: The model only accepts text inputs. +Output: The model generates only text and code. +Model Architecture: Llama 3 is an autoregressive language model built on an optimized Transformer architecture. Its instruction-tuned variants leverage Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) to align outputs with human preferences regarding helpfulness and safety. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public benchmarks, including AIME, GPQA, LiveBench, MuSR, and MMLU, covering mathematical reasoning, scientific QA, general reasoning, and language understanding. + + +# Evaluation Results +## Benchmark Result +| Metrics | Meta-Llama-3-8B-Instruct-Nvidia-Origin | Meta-Llama-3-8B-Instruct-Nvidia-FlagOS | +| ------------------- | -------------------------------------- | -------------------------------------- | +| aime | 0 | 0 | +| gpqa_generative_cot | 0.3146 | 0.3247 | +| livebench_new | 0.2823 | 0.2975 | +| musr_generative | 0.4524 | 0.4418 | +| mmlu | 0.204 | 0.2178 | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/llama-3-8b-nvidia-tree_0.5.0_3.5-gems_5.0.1.rc.0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0.dev20250804_cu128-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.133.20:2605260901 + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Meta-Llama-3-8B-Instruct-nvidia-FlagOS --local_dir /data/models/Meta-Llama-3-8B-Instruct-nvidia-FlagOS +``` + +### Start the Container +```bash +docker run -d \ + --name meta-llama-3-8b-instruct-nvidia-flagos \ + -v /data/models:/data/models \ + --gpus all \ + --network host \ + harbor.baai.ac.cn/external-cooperation/llama-3-8b-nvidia-tree_0.5.0_3.5-gems_5.0.1.rc.0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0.dev20250804_cu128-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.133.20:2605260901 \ + sleep infinity +docker exec -it meta-llama-3-8b-instruct-nvidia-flagos bash +``` +### Start the Server +```bash +export VLLM_PLUGINS='fl' +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +nohup vllm serve /data/models/Meta-Llama-3-8B-Instruct-nvidia-FlagOS --served-model-name meta-llama-3-8b-instruct-nvidia-flagos --port 8000 --enforce-eager > Llama-3-8B.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "meta-llama-3-8b-instruct-nvidia-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from LLM-Research/Meta-Llama-3-8B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_MiniCPM5-1B-kunlunxin-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_MiniCPM5-1B-kunlunxin-FlagOS.md index 03efdf47b..8c36bdf4d 100644 --- a/docs/flagrelease_en/model_readmes/FlagRelease_MiniCPM5-1B-kunlunxin-FlagOS.md +++ b/docs/flagrelease_en/model_readmes/FlagRelease_MiniCPM5-1B-kunlunxin-FlagOS.md @@ -70,14 +70,17 @@ docker exec -it flagos /bin/bash ``` ### Start the Server ```bash +export VLLM_FL_PLATFORM=kunlunxin +export VLLM_FL_PREFER=flagos +export VLLM_FL_FLAGOS_WHITELIST='rms_norm,silu_and_mul,rotary_embedding' +export FLAGCX_PATH=/env/FlagCX +export CUDA_VISIBLE_DEVICES=0 vllm serve /data/MiniCPM5-1B \ ---trust-remote-code \ ---dtype bfloat16 \ ---enforce-eager \ ---port 8000 \ ---host 0.0.0.0 \ ---served-model-name MiniCPM5-1B \ ---gpu-memory-utilization 0.85 + --served-model-name MiniCPM5-1B \ + --port 8000 \ + --tensor-parallel-size 1 \ + --gpu-memory-utilization 0.95 \ + --enforce-eager ``` ## Service Invocation diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-hygon-FlagOS.md new file mode 100644 index 000000000..1522ae712 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-hygon-FlagOS.md @@ -0,0 +1,188 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- + +# Introduction + +**Moonlight-16B-A3B** is a high-performance large language model optimized for the Hygon DCU platform. Built on a Mixture of Experts (MoE) architecture, the model features a total of 16 billion parameters with approximately 3 billion active parameters per inference, striking an optimal balance between high performance and efficient inference throughput. Deeply optimized for Hygon GPU hardware, Moonlight-16B-A3B supports the vLLM inference framework, making it well-suited for large-scale deployment scenarios. + +### Integrated Deployment + +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes + +### Consistency Validation + +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public benchmarks. + +# Evaluation Results + +## Benchmark Result + +| Metrics | Moonlight-16B-A3B-Nvidia-Origin | Moonlight-16B-A3B-Hygon-FlagOS | +|------------------|--------------------------------|--------------------------------| +| GPQA_Diamond | 0.1384 | 0.1183 | +| LiveBench New | 0.0475 | 0.0512 | +| musr | 0.0172 | 0.0437 | +| mmlu_pro | 0.1986 | 0.3265 | +| aime | 0.0000 | 0.0000 | + +# User Guide + +## Environment Setup + +| Item | Version | +|------|----------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image + +```bash +docker pull harbor.baai.ac.cn/external-cooperation/moonlight-16b-a3b-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:202607071400 +``` + +### Download Open-source Model Weights + +```bash +pip install modelscope + +modelscope download \ + --model FlagRelease/Moonlight-16B-A3B-hygon-FlagOS \ + --local_dir /data/Moonlight-16B-A3B-hygon-FlagOS +``` + +### Start the Container + +```bash +docker run -itd \ + --name flagos \ + --device=/dev/kfd \ + --device=/dev/mkfd \ + --device=/dev/dri \ + --group-add video \ + -v /opt/hyhal:/opt/hyhal \ + --ipc=host \ + --ulimit memlock=-1 \ + --ulimit stack=67108864 \ + --network host \ + -v /data/Moonlight-16B-A3B-hygon-FlagOS:/data/Moonlight-16B-A3B-hygon-FlagOS \ + harbor.baai.ac.cn/external-cooperation/moonlight-16b-a3b-hygon-tree_0.5.0_hcu3.0-gems_5.0.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:202607071400 \ + sleep infinity +``` + +### Enter the Container + +```bash +docker exec -it flagos /bin/bash +``` + +### Start the Server + +```bash +#建议按实际卡号调整 +export VLLM_FL_FLAGOS_BLACKLIST="mul,copy_" +export HIP_VISIBLE_DEVICES=4,5 +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +nohup python3 -m vllm.entrypoints.openai.api_server \ + --model /data/Moonlight-16B-A3B-hygon-FlagOS \ + --served-model-name moonlight-flagos \ + --port 8003 \ + --trust-remote-code \ + --max-model-len 8192 \ + --gpu-memory-utilization 0.9 \ + --tensor-parallel-size 2 \ + --enforce-eager \ + > /workspace/moon-test/flagos_server.log 2>&1 & +``` + +## Service Invocation + +### Invocation Script + +```bash +curl http://localhost:8003/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "moonlight-flagos", + "messages": [{"role": "user", "content": "hello"}] + }' +``` + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters: + - API URL: http://localhost:8003/v1 + - Model: moonlight-flagos +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response + +# Technical Overview + +FlagOS is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. + +With core technologies such as FlagScale, together with vllm-plugin-fl, distributed training/inference framework, FlagGems universal operator library, FlagCX communication library, and FlagTree unified compiler, the FlagRelease platform leverages the FlagOS stack to automatically produce and release various combinations of . + +This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + +## FlagGems + +FlagGems is a high-performance, generic operator library implemented in Triton language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM training and inference across diverse hardware platforms. + +## FlagTree + +FlagTree is an open-source unified compiler for multiple AI chips. It provides unified compilation capabilities across multiple backends and rapidly implements single-repository multi-backend support. + +## FlagScale and vllm-plugin-fl + +FlagScale is a comprehensive toolkit designed to support the entire lifecycle of large models. It integrates capabilities from Megatron-LM and vLLM to provide an end-to-end solution for training and inference. + +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend. + +## FlagCX + +FlagCX is a scalable and adaptive cross-chip communication library for distributed AI workloads. + +## FlagEval Evaluation Framework + +FlagEval is a comprehensive evaluation system and open platform for large models. It supports large-scale benchmark evaluation across NLP, CV, Audio, and Multimodal tasks. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License + +The model weights are derived from moonshotai/Moonlight-16B-A3B and are open-sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-nvidia-FlagOS.md new file mode 100644 index 000000000..b48a413e3 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Moonlight-16B-A3B-nvidia-FlagOS.md @@ -0,0 +1,156 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction + +**Moonlight-16B-A3B** is a high-performance large language model released by Moonshot AI, built on a Mixture of Experts (MoE) architecture. The model leverages an advanced MoE design that significantly expands model capacity while maintaining efficient inference performance. Moonlight-16B-A3B delivers strong results across a wide range of natural language processing tasks, including text generation, code understanding, mathematical reasoning, and knowledge-based question answering, making it a powerful foundation model for both academic research and industrial applications. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes + +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public benchmarks. + +# Evaluation Results +## Benchmark Result +| Task | Moonlight-16B-A3B-Nvidia-Origin | Moonlight-16B-A3B-Nvidia-FlagOS | +| ------------------- | -------------------- | -------------------- | +| aime | 0.0000 | 0.0000 | +| gpqa_generative_cot | 0.1384 | 0.1367 | +| LiveBench New | 0.0475 | 0.0495 | +| musr_generative | 0.0159 | 0.0066 | +| mmlu_pro | 0.1986 | 0.2008 | + +# User Guide + +## Environment Setup + +| Item | Version | +|------------------|--------------------------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | Ubuntu 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/moonlight-16b-a3b-nvidia-tree_0.5.0_3.5-gems_5.0.2-vllm_0.13.0-plugin_0.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.133.20:2606031430 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope + +modelscope download \ + --model FlagRelease/Moonlight-16B-A3B-nvidia-FlagOS \ + --local_dir /data/Moonlight-16B-A3B-nvidia-FlagOS +``` + +### Start the Container +```bash +docker run -itd \ + --name flagos \ + --gpus all \ + --ipc=host \ + --ulimit memlock=-1 \ + --ulimit stack=67108864 \ + --network host \ + -v /data/Moonlight-16B-A3B-nvidia-FlagOS:/data/Moonlight-16B-A3B-nvidia-FlagOS \ + harbor.baai.ac.cn/external-cooperation/moonlight-16b-a3b-nvidia-tree_0.5.0_3.5-gems_5.0.2-vllm_0.13.0-plugin_0.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.133.20:2606031430 +``` + +### Enter the Container +```bash +docker exec -it flagos /bin/bash +``` + +### Start the Server +Inside the container: +```bash +#建议按实际卡号调整 +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +VLLM_USE_MODELSCOPE=true CUDA_VISIBLE_DEVICES=3,4 nohup vllm serve \ + /data/Moonlight-16B-A3B-nvidia-FlagOS \ + --served-model-name moonlight-16b-a3b-flagos \ + --port 8003 \ + --trust-remote-code \ + --max-model-len 8192 \ + --gpu-memory-utilization 0.95 \ + --tensor-parallel-size 2 \ + > /workspace/flagos_server.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8003/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "moonlight-16b-a3b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + +### AnythingLLM Integration Guide + +#### 1. Download & Install +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response + +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. + +With core technologies such as **FlagScale**, together with **vllm-plugin-fl** (distributed training/inference framework), **FlagGems** (universal operator library), **FlagCX** (communication library), and **FlagTree** (unified compiler), the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of ``. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM training and inference across diverse hardware platforms. + +## FlagTree +FlagTree is an open-source, unified compiler for multiple AI chips. It provides unified compilation capabilities across multiple backends and rapidly implements single-repository multi-backend support. + +## FlagScale and vllm-plugin-fl +**FlagScale** is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. + +**vllm-plugin-fl** is a vLLM plugin built on the FlagOS unified multi-chip backend, to help FlagScale support multi-chip on the vLLM framework. + +## FlagCX +**FlagCX** is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## FlagEval Evaluation Framework +**FlagEval** is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: +- **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. +- **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License +The model weights are derived from moonshotai/Moonlight-16B-A3B and are open-sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md new file mode 100644 index 000000000..763079df1 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-hygon-FlagOS.md @@ -0,0 +1,135 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (in tokens). The model underwent a rigorous enhancement process, incorporating supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3.5-MoE-instruct-Nvidia-Origin | Phi-3.5-MoE-instruct-Hygon-FlagOS | +|---------------------|--------------------------------|--------------------------------------| +| aime | 0.0334 | 0.1000 | +| gpqa_generative_cot | 0.3171 | 0.3180 | +| mmlu_pro | 0.5336 | 0.5310 | +| musr_generative | 0.5040 | 0.4987 | +| livebench_new | 0.2863 | 0.2757 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.5, build 55c4c88 | +| Operating System | 22.04.4 LTS | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/phi-3.5-moe-instruct-hygon-tree_0.5.0_hcu3.0-gems_5.0.0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607010920 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3.5-MoE-instruct-hygon-FlagOS --local_dir /data/Phi-3.5-MoE-instruct-hygon-FlagOS +``` + +### Start the Container +```bash +docker run --name phi-3.5-moe-instruct-hygon-flagos --network=host --ipc=host --device=/dev/kfd --device=/dev/mkfd --device=/dev/dri -v /opt/hyhal:/opt/hyhal -v /data:/data --group-add video --cap-add=SYS_PTRACE --security-opt seccomp=unconfined -itd harbor.baai.ac.cn/external-cooperation/phi-3.5-moe-instruct-hygon-tree_0.5.0_hcu3.0-gems_5.0.0-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0-das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607010920 sleep infinity +docker exec -it phi-3.5-moe-instruct-hygon-flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl && \ +export TRITON_ALL_BLOCKS_PARALLEL=1 && \ +export USE_FLAGGEMS=1 && \ +export VLLM_FL_FLAGOS_WHITELIST=zero_,zeros,arange,reciprocal,cos,sin,ge_scalar,lt_scalar,bitwise_and_tensor,bitwise_or_tensor,bitwise_not,embedding,index,rand_like,full,argmax,where_self,where_self_out,lt,cumsum_out,le,scatter,to_copy && \ +nohup vllm serve /data/Phi-3.5-MoE-instruct-hygon-FlagOS \ + --served-model-name phi-3.5-moe-instruct-hygon-flagos \ + --host 0.0.0.0 \ + --port 8230 \ + --trust-remote-code \ + --enforce-eager \ + --tensor-parallel-size 2 \ + --gpu-memory-utilization 0.9 \ + --max-model-len 8192 \ + > Phi-3.5-MoE.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8230/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "phi-3.5-moe-instruct-hygon-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response + +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support + +# License +The model weights are derived from LLM-Research/Phi-3.5-MoE-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + + + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md new file mode 100644 index 000000000..ed64d104a --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-MoE-instruct-nvidia-FlagOS.md @@ -0,0 +1,224 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction + +Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (in tokens). The model underwent a rigorous enhancement process, incorporating supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures. + + +### Integrated Deployment + +- Out-of-the-box inference scripts with pre-configured hardware and software parameters + +- Released FlagOS-Nvidia container image supporting deployment within minutes + +### Consistency Validation + +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + + + + +##tion Results + +## Benchmark Result + +| Metrics | Phi-3.5-MoE-instruct-nvidia-FlagOS-Nvidia-Origin | Phi-3.5-MoE-instruct-nvidia-FlagOS-Nvidia-FlagOS | +|---------|----------------------------------------|------------------------------------------| +| aime | 0.0334 | 0.0334 | +| gpqa_generative_cot | 0.3171 | 0.3456 | +| mmlu_pro | 0.5336 | 0.5330 | +| musr_generative | 0.5040 | 0.5079 | +| livebench_new | 0.2863 | 0.2784 | + +# User Guide + +## Environment Setup + +| Item | Version | +|------|---------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | Operation Steps + + +### Download FlagOS Image + +```bash + +docker pull harbor.baai.ac.cn/external-cooperation/phi-3.5-moe-instruct-nvidia-gems_5.0.2-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0-cu128-driver_570.133.20-arc_amd64:2606091628 + +``` + + +### Download Open-source Model Weights + +```bash + +pip install modelscope + +modelscope download --model FlagRelease/Phi-3.5-MoE-instruct-nvidia-FlagOS --local_dir /data/models/Phi-3.5-MoE-instruct-nvidia-FlagOS + +``` + + +### Start the Container + +```bash + +docker run -itd --name Phi-3.5-MoE-instruct-nvidia-flagOS --gpus all --shm-size="32g" --privileged --cap-add=ALL --pid=host --net=host -w /workspace -v /usr/src:/usr/src -v /data/models/:/data/models -v /lib/modules:/lib/modules -v /dev:/dev harbor.baai.ac.cn/external-cooperation/phi-3.5-moe-instruct-nvidia-gems_5.0.2-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0-cu128-driver_570.133.20-arc_amd64:2606091628 sleep infinity +docker exec -it Phi-3.5-MoE-instruct-nvidia-flagOS /bin/bash +``` + +### Start the Server + +```bash + +export VLLM_PLUGINS=fl + +export TRITON_ALL_BLOCKS_PARALLEL=1 + +export USE_FLAGGEMS=1 + +export CUDA_VISIBLE_DEVICES=3 + +export VLLM_FL_FLAGOS_WHITELIST="min,mean,arange,max,gather,silu_and_mul,moe_sum,moe_align_block_size,softmax,rand_like,where_self_out,where_self,argmax,true_divide_,true_divide,sort,bitwise_not,embedding,cos,sin,std,reciprocal,lt,ge_scalar,abs" + +ulimit -n 2048 && nohup vllm serve \ + +--model /data/models/Phi-3.5-MoE-instruct-nvidia-FlagOS \ + +--served-model-name phi-3.5-moe-instruct-nvidia-flagos \ + +--host 0.0.0.0 \ + +--port 6679 \ + +--max-model-len 10000 \ + +--gpu-memory-utilization 0.95 \ + +--trust-remote-code \ + +--tensor-parallel-size 1 \ + +--enforce-eager \ + +> phi-3.5_flagos.log 2>&1 & + +``` + + +## Service Invocation + +### Invocation Script + +```bash + +curl http://localhost:6679/v1/chat/completions \ + +-H "Content-Type: application/json" \ + +-d '{ + +"model": "phi-3.5-moe-instruct-nvidia-flagos", + +"messages": [{"role": "user", "content": "你好"}] + +}' + +``` + + + +### AnythingLLM Integration Guide + + +#### 1. Download & Install + + +- Visit the official site: https://anythingllm.com/ + +- Choose the appropriate version for your OS (Windows/macOS/Linux) + +- Follow the installation wizard to complete the setup + + +#### 2. Configuration + + +- Launch AnythingLLM + +- Open settings (bottom left, fourth tab) + +- Configure core LLM parameters + +- Click "Save Settings" to apply changes + + +#### 3. Model Interaction + + +- After model loading is complete: + +- Click "New Conversation" + +- Enter your question (e.g., “Explain the basics of quantum computing”) + +- Click the send button to get a response + + +# Technical Overview + +FlagOS is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the FlagScale, together with vllm-plugin-fl, distributed training/inference framework, FlagGems universal operator library, FlagCX communication library, and FlagTree unified compiler, the FlagRelease platform leverages the FlagOS stack to automatically produce and release various combinations of . This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. + +## FlagGems + +FlagGems is a high-performance, generic operator libraryimplemented in Triton language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. + +## FlagTree + +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. + +## FlagScale and vllm-plugin-fl + +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including Megatron-LM and vLLM, to provide a robust, end-to-end solution for managing and scaling large models. + +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. + +## FlagCX + +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + + +## FlagEval Evaluation Framework + +FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + +- Multi-dimensional Evaluation: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + +- Industry-Grade Use Cases: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + + + +# Contributing + + +We warmly welcome global developers to join us: + + +\1. Submit Issues to report problems + +\2. Create Pull Requests to contribute code + +\3. Improve technical documentation + +\4. Expand hardware adaptation support + + +# License + +The model weights are derived from LLM-Research/Phi-3.5-MoE-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-FlagOS.md new file mode 100644 index 000000000..a52450c85 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3.5-mini-instruct-FlagOS-Origin | Phi-3.5-mini-instruct-FlagOS-FlagOS | +|--------------|-------------------------------------|-------------------------------------| +| GPQA_Diamond | 26.0 | 24.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 24.0.0 | +| Operating System | Ubuntu 24.04.3 LTS (Noble Numbat) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/phi-3.5-mini-instruct-nvidia003-gems5.0.2-tree0.6.0-cxnone-plugin0.0.0-vllm0.20.2-cp312-pt211-cu130-x64-570.158.01:202607161715-v5 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3.5-mini-instruct-FlagOS --local_dir /data/Phi-3.5-mini-instruct-FlagOS +``` + +### Start the Container +```bash +docker run -itd --name flagos --gpus=all --network=host -v /data:/data harbor.baai.ac.cn/flagrelease-project/phi-3.5-mini-instruct-nvidia003-gems5.0.2-tree0.6.0-cxnone-plugin0.0.0-vllm0.20.2-cp312-pt211-cu130-x64-570.158.01:202607161715-v5 +``` +### Start the Server +```bash +vllm serve /data/Phi-3.5-mini-instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Phi-3.5-mini-instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Phi-3.5-mini-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from LLM-Research/Phi-3.5-mini-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-metax-FlagOS.md new file mode 100644 index 000000000..28c3a13b4 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Phi-3.5-mini-instruct-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Phi-3.5-mini-instruct-metax-FlagOS-Origin | Phi-3.5-mini-instruct-metax-FlagOS-FlagOS | +|--------------|-------------------------------------------|-------------------------------------------| +| GPQA_Diamond | 0 | 28.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/phi-3.5-mini-instruct-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607162145-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Phi-3.5-mini-instruct-metax-FlagOS --local_dir /data/Phi-3.5-mini-instruct-FlagOS +``` + +### Start the Container +```bash +docker run -v /data:/data -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data/models/Phi-3.5-mini-instruct:/data/models/Phi-3.5-mini-instruct harbor.baai.ac.cn/flagrelease-public/phi-3.5-mini-instruct-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607162145-v2 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/Phi-3.5-mini-instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Phi-3.5-mini-instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Phi-3.5-mini-instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from LLM-Research/Phi-3.5-mini-instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-7B-Instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-7B-Instruct-metax-FlagOS.md new file mode 100644 index 000000000..35b8cac1d --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-7B-Instruct-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen2.5-7B-Instruct-metax-FlagOS-Origin | Qwen2.5-7B-Instruct-metax-FlagOS-FlagOS | +|--------------|-----------------------------------------|-----------------------------------------| +| GPQA_Diamond | 0 | 38.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen2.5-7b-instruct-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607230144-v2 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen2.5-7B-Instruct-metax-FlagOS --local_dir /data/Qwen2.5-7B-Instruct-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/qwen2.5-7b-instruct-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607230144-v2 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Qwen2.5-7B-Instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Qwen2.5-7B-Instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen2.5-7B-Instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen2.5-7B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-metax-FlagOS.md new file mode 100644 index 000000000..41be68aef --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen2.5-Coder-7B-Instruct-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen2.5-Coder-7B-Instruct-metax-FlagOS-Origin | Qwen2.5-Coder-7B-Instruct-metax-FlagOS-FlagOS | +|--------------|-----------------------------------------------|-----------------------------------------------| +| GPQA_Diamond | 0 | 38.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/qwen2.5-coder-7b-instruct-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607230542-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen2.5-Coder-7B-Instruct-metax-FlagOS --local_dir /data/Qwen2.5-Coder-7B-Instruct-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-project/qwen2.5-coder-7b-instruct-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607230542-v3 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/Qwen2.5-Coder-7B-Instruct-FlagOS --host 0.0.0.0 --port 8000 --served-model-name Qwen2.5-Coder-7B-Instruct --tensor-parallel-size 1 --max-model-len 32768 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Qwen2.5-Coder-7B-Instruct", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen2.5-Coder-7B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md new file mode 100644 index 000000000..f9afe717d --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Qwen3.6-27B-metax-FlagOS-Express.md @@ -0,0 +1,145 @@ +--- +language: +- zh +- en +license: apache-2.0 +--- + +# Introduction +The first open-weight release of Qwen3.6 is now available. Building on the Qwen3.5 series released in February and shaped by direct community feedback, Qwen3.6 prioritizes stability and real-world utility to deliver a more intuitive, responsive, and productive coding experience. Key improvements include enhanced agentic coding capabilities for frontend workflows and repository-level reasoning, along with a new thinking preservation option that retains reasoning context from historical messages to streamline iterative development. + + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Qwen3.6-27B-Nvidia-Origin | Qwen3.6-27B-Metax-FlagOS | +|--------------|---------------------------|--------------------------| +| GPQA_Diamond | 85.86 | 84.34 | +| ERQA | 59.25 | 58.25 | + +## Performance Benchmark Result +|Metric| 1k&1k 64 Concurrency| 4k&1k 64 Concurrency| 16k&1k 64 Concurrency|64k&1k 64 Concurrency| +|--------------|---------------------------|--------------------------|---|---| +|Equal Computing Power Ratio (flagos/H100)| 96.35% |97.36%| 83.33%|90.31%| + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2 | +| Operating System | Ubuntu 22.04.5 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/qwen3.6-27b-metax001-gems5.4.0-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.8.1:202607220155 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Qwen3.6-27B-metax-FlagOS-Express --local_dir /data/Qwen3.6-27B +``` + +### Start the Container +```bash +docker run -itd \ + --name flagos \ + --privileged \ + --network=host \ + --security-opt seccomp=unconfined \ + --security-opt apparmor=unconfined \ + --shm-size '100gb' \ + --ulimit memlock=-1 \ + --group-add video \ + --device=/dev/dri \ + --device=/dev/mxcd \ + --device=/dev/mem \ + --device=/dev/infiniband \ + -v /usr/local/:/usr/local/ \ + -v /data/:/data/ \ + harbor.baai.ac.cn/flagrelease-public/qwen3.6-27b-metax001-gems5.4.0-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.8.1:202607220155 \ + /bin/bash +docker exec -it flagos /bin/bash +``` +### Start the Server +```bash +export FLAGGEMS_VENDOR=metax +export CUDA_VISIBLE_DEVICES=6,7 +export VLLM_FL_FLAGOS_WHITELIST=cat,cos,cumsum,fill,full,gather,gt,le,lt,max,mul,sin,softmax,to,where,zeros,zeros_like +export VLLM020_CONTIGUOUS_SINGLE_PREFILL=1 +vllm serve /data/models/Qwen3.6-27B/ \ + --tensor-parallel-size 2 --port 8000 --trust-remote-code --dtype bfloat16 \ + --served-model-name qwen36-27b \ + --max-num-batched-tokens 16384 +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "qwen36-27b", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from Qwen/Qwen3.6-27B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-hygon-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-hygon-FlagOS.md new file mode 100644 index 000000000..fb55c0584 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-hygon-FlagOS.md @@ -0,0 +1,147 @@ +--- +frameworks: +- "" +tasks: [] +--- +# Introduction +Seed-OSS is a series of open-source large language models developed by ByteDance's Seed Team, designed for powerful long-context, reasoning, agent and general capabilities, and versatile developer-friendly features. Although trained with only 12T tokens, Seed-OSS achieves excellent performance on several popular open benchmarks. +We release this series of models to the open-source community under the Apache-2.0 license. +Key Features + Flexible Control of Thinking Budget: Allowing users to flexibly adjust the reasoning length as needed. This capability of dynamically controlling the reasoning length enhances inference efficiency in practical application scenarios. + Enhanced Reasoning Capability: Specifically optimized for reasoning tasks while maintaining balanced and excellent general capabilities. + Agentic Intelligence: Performs exceptionally well in agentic tasks such as tool-using and issue resolving. + Research-Friendly: Given that the inclusion of synthetic instruction data in pre-training may affect the post-training research, we released pre-trained models both with and without instruction data, providing the research community with more diverse options. + Native Long Context: Trained with up-to-512K long context natively. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Hygon** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Seed-OSS-36B-Instruct-Nvidia-Origin | Seed-OSS-36B-Instruct-Hygon-FlagOS | +|---------------------|----------------------------------------------------|---------------------------------------------------| +| gpqa_generative_cot | 0.6149 | 0.6364 | +| aime | 0.6667 | 0.7 | +| musr_generative | 0.4167 | 0.3836 | +| livebench_new | 0.5115 | 0.5075 | +| mmlu_pro | 0.4886 | 0.4841 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 20.10.24, build 297e128 | +| Operating System | Sugon OS 8.9 | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/seed-oss-36b-instruct-hygon-tree_0.5.0-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607091506 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Seed-OSS-36B-Instruct-hygon-FlagOS --local_dir /data/Seed-OSS-36B-Instruct-hygon-FlagOS +``` + +### Start the Container +```bash +sudo docker run -itd \ + --name seed-oss-36b-flagos \ + --network=host \ + --ipc=host \ + --pid=host \ + --device=/dev/infiniband \ + --device=/dev/kfd \ + --device=/dev/dri \ + --device=/dev/mkfd \ + --shm-size 512G \ + --ulimit memlock=-1 \ + -w /workspace \ + -v /opt/hyhal:/opt/hyhal \ + --security-opt seccomp=unconfined \ + --dns 8.8.8.8 \ + --dns 8.8.4.4 \ + -v /data/Seed-OSS-36B-Instruct-hygon-FlagOS:/data/Seed-OSS-36B-Instruct-hygon-FlagOS \ + harbor.baai.ac.cn/external-cooperation/seed-oss-36b-instruct-hygon-tree_0.5.0-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.10.12-torch_2.9.0_das.opt1.dtk2604.20260206.g275d08c2-pcp_hygon-dpu_hygon-x86_64-driver_1.11.0:2607091506 \ + sleep infinity +docker exec -it seed-oss-36b-flagos /bin/bash +``` +### Start the Server +```bash +export HIP_VISIBLE_DEVICES=0,1,2,3 +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export USE_FLAGGEMS=1 +export GEMS_VENDOR=hygon +export VLLM_FL_FLAGOS_BLACKLIST="sort,masked_fill_,mm,mul,addmm" +nohup vllm serve --model /data/Seed-OSS-36B-Instruct-hygon-FlagOS/ --served-model-name seed-oss-36b-flagos --port 8000 --tensor-parallel-size 2 --trust-remote-code --max-model-len 2600 --enforce-eager >eager-seed-oss-gems.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "seed-oss-36b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from unsloth/Seed-OSS-36B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-nvidia-FlagOS.md new file mode 100644 index 000000000..cd3cc8c12 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_Seed-OSS-36B-Instruct-nvidia-FlagOS.md @@ -0,0 +1,141 @@ +--- +frameworks: +- "" +language: +- zh +- en +license: apache-2.0 +tasks: [] +--- +# Introduction +Seed-OSS is a series of open-source large language models developed by ByteDance's Seed Team, designed for powerful long-context, reasoning, agent and general capabilities, and versatile developer-friendly features. Although trained with only 12T tokens, Seed-OSS achieves excellent performance on several popular open benchmarks. +We release this series of models to the open-source community under the Apache-2.0 license. +Key Features + Flexible Control of Thinking Budget: Allowing users to flexibly adjust the reasoning length as needed. This capability of dynamically controlling the reasoning length enhances inference efficiency in practical application scenarios. + Enhanced Reasoning Capability: Specifically optimized for reasoning tasks while maintaining balanced and excellent general capabilities. +Agentic Intelligence: Performs exceptionally well in agentic tasks such as tool-using and issue resolving. + Research-Friendly: Given that the inclusion of synthetic instruction data in pre-training may affect the post-training research, we released pre-trained models both with and without instruction data, providing the research community with more diverse options. + Native Long Context: Trained with up-to-512K long context natively. +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | Seed-OSS-36B-Instruct-Nvidia-Origin | Seed-OSS-36B-Instruct-Nvidia-FlagOS | +|---------------------|-------------------------------------|-------------------------------------| +| gpqa_generative_cot | 0.6099 | 0.6208 | +| aime | 0.6667 | 0.6667 | +| musr_generative | 0.4153 | 0.4061 | +| livebench_new | 0.5103 | 0.5093 | +| mmlu_pro | 0.4875 | 0.4901 | +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/external-cooperation/seed-oss-36b-instruct-nvidia-tree_0.5.1-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.133.20:2606090953 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/Seed-OSS-36B-Instruct-nvidia-FlagOS --local_dir /data/Seed-OSS-36B-Instruct-nvidia-FlagOS +``` + +### Start the Container +```bash +sudo docker run -itd \ + --name seed-oss-36b-flagos \ + --network=host \ + --ipc=host \ + --pid=host \ + --device=/dev/infiniband \ + --shm-size 512G \ + --ulimit memlock=-1 \ + --gpus all \ + -w /workspace \ + -v /data/Seed-OSS-36B-Instruct-nvidia-FlagOS:/data/Seed-OSS-36B-Instruct-nvidia-FlagOS \ + harbor.baai.ac.cn/external-cooperation/seed-oss-36b-instruct-nvidia-tree_0.5.1-gems_0.5.2-vllm_0.13.0-plugin_0.1.1-cx_none-python_3.12.3-torch_2.9.0_cu128-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.133.20:2606090953 sleep infinity +docker exec -it seed-oss-36b-flagos /bin/bash +``` +### Start the Server +```bash +export VLLM_PLUGINS=fl +export TRITON_ALL_BLOCKS_PARALLEL=1 +export VLLM_FL_FLAGOS_BLACKLIST="sort,masked_fill_,mm,mul,addmm" +export USE_FLAGGEMS=1 +nohup vllm serve --model /data/Seed-OSS-36B-Instruct-nvidia-FlagOS/ --served-model-name seed-oss-36b-flagos --port 8000 --tensor-parallel-size 2 --trust-remote-code --max-model-len 8192 >vllm.log 2>&1 & +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "seed-oss-36b-flagos", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from unsloth/Seed-OSS-36B-Instruct and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt + diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-metax-FlagOS.md new file mode 100644 index 000000000..d35e618cc --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-metax-FlagOS.md @@ -0,0 +1,48 @@ +--- +license: Apache License 2.0 +tags: [] + +#model-type: +##如 gpt、phi、llama、chatglm、baichuan 等 +#- gpt + +#domain: +##如 nlp、cv、audio、multi-modal +#- nlp + +#language: +##语言代码列表 https://help.aliyun.com/document_detail/215387.html?spm=a2c4g.11186623.0.0.9f8d7467kni6Aa +#- cn + +#metrics: +##如 CIDEr、Blue、ROUGE 等 +#- CIDEr + +#tags: +##各种自定义,包括 pretrained、fine-tuned、instruction-tuned、RL-tuned 等训练方法和其他 +#- pretrained + +#tools: +##如 vllm、fastchat、llamacpp、AdaSeq 等 +#- vllm +--- +### 当前模型的贡献者未提供更加详细的模型介绍。模型文件和权重,可浏览“模型文件”页面获取。 +#### 您可以通过如下git clone命令,或者ModelScope SDK来下载模型 + +SDK下载 +```bash +#安装ModelScope +pip install modelscope +``` +```python +#SDK模型下载 +from modelscope import snapshot_download +model_dir = snapshot_download('FlagRelease/ZCK-Qwen3-8B-metax-FlagOS') +``` +Git下载 +``` +#Git模型下载 +git clone https://www.modelscope.cn/FlagRelease/ZCK-Qwen3-8B-metax-FlagOS.git +``` + +

如果您是本模型的贡献者,我们邀请您根据模型贡献文档,及时完善模型卡片内容。

\ No newline at end of file diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-nvidia-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-nvidia-FlagOS.md new file mode 100644 index 000000000..5684cc03a --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_ZCK-Qwen3-8B-nvidia-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +AI21-Jamba-1.5-Mini is an open‑source large language model. This release includes full adaptation for the Ascend platform, along with provided images, cache files, and performance results to facilitate rapid deployment and validation. + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Nvidia** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | AI21-Jamba-1.5-Mini-ascend-FlagOS-Nvidia-Origin | AI21-Jamba-1.5-Mini-ascend-FlagOS-Nvidia-FlagOS | +|--------------|--------------------------------|--------------------------------------| +| GPQA_Diamond | - | - | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | Docker version 24.0.0, build 98fdcd7 | +| Operating System | 22.04.4 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash + +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/AI21-Jamba-1.5-Mini-ascend-FlagOS --local_dir /data/AI21-Jamba-1.5-Mini-ascend-FlagOS +``` + +### Start the Container +```bash + +``` +### Start the Server +```bash + +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "flagOS", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., “Explain the basics of quantum computing”) +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a “develop once, run anywhere” workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator libraryimplemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutralkernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multipleAI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. Forupstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to supportthe entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ modelevaluations across NLP, CV, Audio, and Multimodal fields,covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizonta1 evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from AI-ModelScope/AI21-Jamba-1.5-Mini and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_aya-23-8B-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_aya-23-8B-metax-FlagOS.md new file mode 100644 index 000000000..1d5097ff5 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_aya-23-8B-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | aya-23-8B-metax-FlagOS-Origin | aya-23-8B-metax-FlagOS-FlagOS | +|--------------|-------------------------------|-------------------------------| +| GPQA_Diamond | 0 | 30.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-project/aya-23-8b-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607230807-v3 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/aya-23-8B-metax-FlagOS --local_dir /data/aya-23-8B-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/hy3-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.8.1:202607061058 sleep infinity +``` +### Start the Server +```bash +VLLM_PLUGINS=fl vllm serve /data/aya-23-8B-FlagOS --host 0.0.0.0 --port 8000 --served-model-name aya-23-8B --tensor-parallel-size 1 --max-model-len 8192 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "aya-23-8B", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from CohereLabs/aya-23-8B and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt diff --git a/docs/flagrelease_en/model_readmes/FlagRelease_gemma-1.1-7b-it-metax-FlagOS.md b/docs/flagrelease_en/model_readmes/FlagRelease_gemma-1.1-7b-it-metax-FlagOS.md new file mode 100644 index 000000000..ec08d0228 --- /dev/null +++ b/docs/flagrelease_en/model_readmes/FlagRelease_gemma-1.1-7b-it-metax-FlagOS.md @@ -0,0 +1,108 @@ +# Introduction +新模型介绍,待定.... + +### Integrated Deployment +- Out-of-the-box inference scripts with pre-configured hardware and software parameters +- Released **FlagOS-Metax** container image supporting deployment within minutes +### Consistency Validation +- Rigorously evaluated through benchmark testing: Performance and results from the FlagOS software stack are compared against native stacks on multiple public. + + +# Evaluation Results +## Benchmark Result +| Metrics | gemma-1.1-7b-it-metax-FlagOS-Origin | gemma-1.1-7b-it-metax-FlagOS-FlagOS | +|--------------|-------------------------------------|-------------------------------------| +| GPQA_Diamond | 0 | 36.0 | +| ERQA | - | - | +| Aime24 | - | - | + +# User Guide +Environment Setup + +| Item | Version | +|------------------|----------------------| +| Docker Version | 29.4.0 | +| Operating System | Ubuntu 22.04.3 LTS (Jammy Jellyfish) | + +## Operation Steps + +### Download FlagOS Image +```bash +docker pull harbor.baai.ac.cn/flagrelease-public/gemma-1.1-7b-it-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607211846-v5 +``` + +### Download Open-source Model Weights +```bash +pip install modelscope +modelscope download --model FlagRelease/gemma-1.1-7b-it-metax-FlagOS --local_dir /data/gemma-1.1-7b-it-FlagOS +``` + +### Start the Container +```bash +docker run -d --name flagos --net=host --ipc=host --privileged --shm-size=64g --group-add video --ulimit memlock=-1 --security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri --device=/dev/mxcd -v /data:/data harbor.baai.ac.cn/flagrelease-public/gemma-1.1-7b-it-metax001-gems5.0.2-tree0.5.1-cxnone-plugin0.2.0-vllm0.20.2-cp312-pt28-maca37-x64-3.3.12:202607211846-v5 sleep infinity +``` +### Start the Server +```bash +vllm serve /data/gemma-1.1-7b-it-FlagOS --host 0.0.0.0 --port 8000 --served-model-name gemma-1.1-7b-it --tensor-parallel-size 1 --max-model-len 8192 --trust-remote-code +``` + +## Service Invocation +### Invocation Script +```bash +curl http://localhost:8000/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "gemma-1.1-7b-it", + "messages": [{"role": "user", "content": "你好"}] + }' +``` + + +### AnythingLLM Integration Guide + +#### 1. Download & Install + +- Visit the official site: https://anythingllm.com/ +- Choose the appropriate version for your OS (Windows/macOS/Linux) +- Follow the installation wizard to complete the setup + +#### 2. Configuration + +- Launch AnythingLLM +- Open settings (bottom left, fourth tab) +- Configure core LLM parameters +- Click "Save Settings" to apply changes + +#### 3. Model Interaction + +- After model loading is complete: +- Click **"New Conversation"** +- Enter your question (e.g., "Explain the basics of quantum computing") +- Click the send button to get a response +# Technical Overview +**FlagOS** is a fully open-source system software stack designed to unify the "model–system–chip" layers and foster an open, collaborative ecosystem. It enables a "develop once, run anywhere" workflow across diverse AI accelerators, unlocking hardware performance, eliminating fragmentation among vendor-specific software stacks, and substantially lowering the cost of porting and maintaining AI workloads. With core technologies such as the **FlagScale**, together with vllm-plugin-fl, distributed training/inference framework, **FlagGems** universal operator library, **FlagCX** communication library, and **FlagTree** unified compiler, the **FlagRelease** platform leverages the **FlagOS** stack to automatically produce and release various combinations of \. This enables efficient and automated model migration across diverse chips, opening a new chapter for large model deployment and application. +## FlagGems +FlagGems is a high-performance, generic operator library implemented in [Triton](https://github.com/openai/triton) language. It is built on a collection of backend-neutral kernels that aims to accelerate LLM (Large-Language Models) training and inference across diverse hardware platforms. +## FlagTree +FlagTree is an open source, unified compiler for multiple AI chips project dedicated to developing a diverse ecosystem of AI chip compilers and related tooling platforms, thereby fostering and strengthening the upstream and downstream Triton ecosystem. Currently in its initial phase, the project aims to maintain compatibility with existing adaptation solutions while unifying the codebase to rapidly implement single-repository multi-backend support. For upstream model users, it provides unified compilation capabilities across multiple backends; for downstream chip manufacturers, it offers examples of Triton ecosystem integration. +## FlagScale and vllm-plugin-fl +Flagscale is a comprehensive toolkit designed to support the entire lifecycle of large models. It builds on the strengths of several prominent open-source projects, including [Megatron-LM](https://github.com/NVIDIA/Megatron-LM) and [vLLM](https://github.com/vllm-project/vllm), to provide a robust, end-to-end solution for managing and scaling large models. +vllm-plugin-fl is a vLLM plugin built on the FlagOS unified multi-chip backend, to help flagscale support multi-chip on vllm framework. +## **FlagCX** +FlagCX is a scalable and adaptive cross-chip communication library. It serves as a platform where developers, researchers, and AI engineers can collaborate on various projects, contribute to the development of cutting-edge AI solutions, and share their work with the global community. + +## **FlagEval Evaluation Framework** + FlagEval is a comprehensive evaluation system and open platform for large models launched in 2023. It aims to establish scientific, fair, and open benchmarks, methodologies, and tools to help researchers assess model and training algorithm performance. It features: + - **Multi-dimensional Evaluation**: Supports 800+ model evaluations across NLP, CV, Audio, and Multimodal fields, covering 20+ downstream tasks including language understanding and image-text generation. + - **Industry-Grade Use Cases**: Has completed horizontal evaluations of mainstream large models, providing authoritative benchmarks for chip-model performance validation. + +# Contributing + +We warmly welcome global developers to join us: + +1. Submit Issues to report problems +2. Create Pull Requests to contribute code +3. Improve technical documentation +4. Expand hardware adaptation support +# License +The model weights are derived from google/gemma-1.1-7b-it and are open‑sourced under the Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0.txt