Skip to content

docs: clarify Qwen3.8 benchmark limits - #3104

Open
localai-org-maint-bot wants to merge 3 commits into
mudler:mainfrom
localai-org-maint-bot:docs/refresh-serving-guide
Open

docs: clarify Qwen3.8 benchmark limits#3104
localai-org-maint-bot wants to merge 3 commits into
mudler:mainfrom
localai-org-maint-bot:docs/refresh-serving-guide

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

The Qwen3.8 model guide denies a token comparison that its earlier section records. Clarify that the CUDA shape tests and the later model token gate are separate measurements.

Add the September EXL3 comparison to README News. State that comparator repetitions are incomplete, configurations differ, and the run has no correctness gate. Expose the incomplete repetitions in the benchmark index.

The changes use the current source and committed evidence. No runtime behavior or measured number changes. The scope and evidence anchors are in .agents/specs/qwen38-public-doc-refresh.md, which owns ISSUE-LOCAL-01M221WAVSK2STJ49Y3WTDSDS3.

CPU-only validation passed:

  • python3 scripts/check-readme-structure.py
  • python3 scripts/check-benchmark-index.py
  • python3 scripts/check-site.py
  • python3 scripts/check-agent-record.py
  • Existing README and benchmark-index mutation suites.
  • Commit style, protocol trailers, and git diff --check.

Full preflight was attempted but is not green. The temporary Alpine tool environment caused release-packaging subprocess failures and initially lacked validation dependencies. Focused documentation checks pass after local tool setup. No GPU tests or new benchmarks were run.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]

The public pages omit a recent comparison and disagree about an FP8 gate.
Record the editorial scope before correcting those claims.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]
The benchmark index omits incomplete comparator repetitions. The model
guide also denies a token gate that its preceding section records.
Align both summaries with the evidence and add the September comparison
to README News without a speed or correctness claim.

Issue: ISSUE-LOCAL-01M221WAVSK2STJ49Y3WTDSDS3

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [exec]
Keep the source anchors and validation limits with the pending PR.
The issue remains open until upstream integration.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-6 [Codex]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant