Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ credentials should not prevent the Agent from writing or reviewing code.
|---|---|
| Start a normal integration | [SKILL.md](SKILL.md) |
| Find the relevant official endpoint contract | [Endpoint map](references/api_reference.md) |
| LangChain embeddings, thinking/tool history, Anthropic SDK | [Integration recipes](references/examples.md) |
| Image edits, multimodal retrieval, LangChain embeddings, thinking/tool history, Anthropic SDK | [Integration recipes](references/examples.md) |
| Model choice or current account facts | [Model/account pointers](references/models.md) |
| Diagnose a failure or select a live probe | [Targeted diagnosis](references/workflows.md) |
| Inspect dated evidence | [Known deviations](references/known_deviations.md) |
Expand Down
4 changes: 3 additions & 1 deletion SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ name: ecnu-api
description: >
Connect an application or SDK to the ECNU / ChatECNU LLM Open Platform
at chat.ecnu.edu.cn, or diagnose an ECNU API failure. Covers chat,
Responses, tools, vision, embeddings, rerank, images, TTS, and
Responses, tools, vision, text/image retrieval, image generation/editing, TTS, and
Anthropic-compatible clients. Not for general ECNU information,
unrelated model questions, or generic Agent/prompt design.
---
Expand Down Expand Up @@ -55,6 +55,8 @@ Its full messages URL ends in `/open/api/anthropic/v1/messages`.
|---|---|
| Chat, streaming, vision, Responses, JSON, rerank, images, TTS, or browser integration | [Official endpoint map](references/api_reference.md): open only the matching page |
| LangChain embeddings | [Embedding recipe](references/examples.md#langchain-embeddings) |
| Image editing or generation | [Image recipe](references/examples.md#image-generation-and-editing): distinguish JSON generation from multipart editing |
| Image/text embeddings or rerank | [Multimodal retrieval](references/examples.md#multimodal-retrieval): choose the model and preserve vector-space compatibility |
| Anthropic SDK setup | [Anthropic recipe](references/examples.md#anthropic-sdk) |
| Thinking and tool continuation | [Tool-history recipe](references/examples.md#thinking-and-tool-history) plus the linked ECNU contract |
| Model choice, prices, quotas, or deployment questions | [Model and account pointers](references/models.md) |
Expand Down
6 changes: 6 additions & 0 deletions references/agent_development.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,12 @@ Model facts and protocol rules are linked from [models.md](models.md) and
[api_reference.md](api_reference.md). The design advice and prompt examples
below are **`application-policy`**: starting points to evaluate on your tasks.

**Contract update checked 2026-10-02:** the 2026-09-18 model change made
`ecnu-max` text-only again; use `ecnu-plus` for vision. Plus now supports
`low` / `medium` / `xhigh` effort when thinking is enabled. The older model
choices below describe the September 12 checks, not current execution advice;
follow [current model selection](models.md) and [thinking](examples.md#thinking-and-tool-history).

## Choose a model and reasoning budget

| Workload | Starting point | What to check before escalating |
Expand Down
7 changes: 4 additions & 3 deletions references/api_reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,9 +23,10 @@ in environment variables; ticket URLs are also credentials.
| Chat, streaming, tools, image understanding | [Chat Completions](https://developer.ecnu.edu.cn/vitepress/llm/api/completions.html) | [Thinking/tool history](examples.md#thinking-and-tool-history) |
| Responses-compatible client | [Responses](https://developer.ecnu.edu.cn/vitepress/llm/api/responses.html) | Verify the specific advanced tool/event support; compatibility alone is insufficient |
| JSON Schema or JSON object output | [Structured output](https://developer.ecnu.edu.cn/vitepress/llm/api/structuredoutput.html) | Check completion, parse raw JSON, and validate the supplied schema; do not hide a mismatch by stripping fences |
| Embeddings | [Text vectors](https://developer.ecnu.edu.cn/vitepress/llm/api/embedding.html) | [Raw-string LangChain recipe](examples.md#langchain-embeddings) |
| Rerank | [Rerank](https://developer.ecnu.edu.cn/vitepress/llm/api/rerank.html) | Use the endpoint's contract, not an assumed OpenAI SDK method |
| Image generation | [Images](https://developer.ecnu.edu.cn/vitepress/llm/api/imagegenerate.html) | Observe the current URL lifetime and avoid duplicate paid generations |
| Text or multimodal embeddings | [Vectors](https://developer.ecnu.edu.cn/vitepress/llm/api/embedding.html) | [Text LangChain recipe](examples.md#langchain-embeddings) or [VL payloads](examples.md#multimodal-retrieval); dimensions are model-specific |
| Text or multimodal rerank | [Rerank](https://developer.ecnu.edu.cn/vitepress/llm/api/rerank.html) | [VL payloads](examples.md#multimodal-retrieval); use the endpoint's contract, not an assumed OpenAI SDK method |
| Image generation | [Images](https://developer.ecnu.edu.cn/vitepress/llm/api/imagegenerate.html) | [Generation/edit differences](examples.md#image-generation-and-editing); original and revised prompts can differ |
| Edit an existing image | [Image edits](https://developer.ecnu.edu.cn/vitepress/llm/api/imageedit.html) | [Multipart recipe](examples.md#image-generation-and-editing); one uploaded image, original instruction, no output-size control |
| Text-to-speech | [Audio](https://developer.ecnu.edu.cn/vitepress/llm/api/audio.html) | [Non-JSON errors and PCM headers](workflows.md#start-from-the-symptom) |
| Model discovery | [Models endpoint](https://developer.ecnu.edu.cn/vitepress/llm/api/models.html) | A visible ID does not prove usable capability or valid authentication |
| Anthropic-compatible client | [Anthropic API](https://developer.ecnu.edu.cn/vitepress/llm/api/anthropic.html) | [SDK setup](examples.md#anthropic-sdk); investigate suffix errors only when they occur |
Expand Down
133 changes: 124 additions & 9 deletions references/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,13 @@ an example edit alone does not refresh that evidence.

## LangChain embeddings

The [ECNU embedding contract](https://developer.ecnu.edu.cn/vitepress/llm/api/embedding.html)
For `ecnu-embedding-small`, the [ECNU embedding contract](https://developer.ecnu.edu.cn/vitepress/llm/api/embedding.html)
accepts strings, not OpenAI token-ID arrays. Disable LangChain's token conversion.
The official LangChain example sets `dimensions` to 1024, but the request table
lists only `model` and `input` and does not explain dimension selection.
This recipe omits `dimensions` and validates the documented 1024-value output,
following the dated recipe coverage above. That check does not establish whether
the service accepts or rejects the field. Reject an empty list before making a call.
The current parameter table limits `dimensions` to `ecnu-embedding-vl`, although
the page's older LangChain example still sets it for the small model. This recipe
follows the parameter table: omit `dimensions` and validate the small model's
1024-value output. Reject an empty list before making a call. Multimodal inputs
use the [VL recipe](#multimodal-retrieval), not this text-only adapter.

Standalone example; requires `langchain-openai`:

Expand Down Expand Up @@ -49,13 +49,128 @@ For longer inputs, follow the current endpoint's character limit and split
locally. Do not infer an undocumented batch maximum from per-call billing.
When using direct HTTP, align vectors with inputs by their returned `index`.

## Multimodal retrieval

Use the current [embedding](https://developer.ecnu.edu.cn/vitepress/llm/api/embedding.html)
and [rerank](https://developer.ecnu.edu.cn/vitepress/llm/api/rerank.html) contracts.
Both VL models accept strings or objects with `text`, `image`, or both. An
`image` is a complete PNG/JPEG base64 data URL, not a remote URL or Chat
Completions `image_url` content block. Validate image type, decoded size
(at most 5 MiB) and pixel count (at most 16 million) before encoding/uploading.
Each object holds one image; video and multi-image objects are unsupported.

Request-building fragment for an existing HTTP client; `image_data_url` has
already passed those checks. These dictionaries do not send requests:

```python
item = {"text": "A red flower beside a green leaf", "image": image_data_url}
embedding_payload = {
"model": "ecnu-embedding-vl",
"input": [item, "A red flower"],
"dimensions": 1024,
"encoding_format": "float",
}
rerank_payload = {
"model": "ecnu-rerank-vl",
"query": "A red flower",
"documents": [item, "A blue car"],
"top_n": 2,
"return_documents": False,
}
```

Send `embedding_payload` as JSON to `POST /embeddings`, or `rerank_payload`
as JSON to `POST /rerank`, using the OpenAI-compatible base and Bearer auth.
Keep requests serial when checking both. For embedding, cap batches at 32
items; supported dimensions are 1024/2048/4096, with 4096 as the default.
Check the returned count, unique `index` values and finite vector lengths
against the inputs and requested dimension. `encoding_format=base64` instead
returns little-endian float32 bytes encoded as a string; `usage` may be null.

For rerank, map `results[].index` back to the original candidates and validate
finite scores in [0, 1]. `return_documents=True` can echo entire image data
URLs, so omit that output from logs. The embedding batch limit is not a
documented rerank limit. Both endpoint pages retain an 8192-character text
limit; the model page's 32K context is not permission to exceed it.

Choose the model explicitly. Keep an existing text index on
`ecnu-embedding-small` until a deliberate migration: even a 1024-dimensional
VL vector belongs to a different space. A VL reranker only changes ordering
of candidates; it does not require a new vector index or replace image-text
consistency/education review.

## Image generation and editing

The [model page](https://developer.ecnu.edu.cn/vitepress/llm/model.html) lists
`ecnu-image` as qwen-image-2.1 following the 2026-09-30 upgrade. Use the stable
alias and choose the endpoint according to the task:

- [Generation](https://developer.ecnu.edu.cn/vitepress/llm/api/imagegenerate.html)
uses JSON at `POST /images/generations`. Prompts are automatically expanded;
retain the original prompt separately from a returned `revised_prompt`.
Check this endpoint's allowed `size` values; do not assume arbitrary sizes or
a batch `n` parameter from OpenAI compatibility.
- [Editing](https://developer.ecnu.edu.cn/vitepress/llm/api/imageedit.html)
uses multipart at `POST /images/edits`, with one `image` file and a `prompt`
of at most 1024 characters. Instructions pass through unchanged after safety
review. Only `n=1` is supported; `size` is ignored and not forwarded.

Both support `url` and `b64_json` results. URLs expire after 24 hours; transfer
them promptly when using URL output. Both add an AI watermark. A successful
HTTP status or non-empty `data` alone is insufficient: a rejected request can
have `err_message` and only a masked `revised_prompt`, without an image.
Do not infer mask, multiple references, transparency, fixed edit dimensions,
or guaranteed character consistency from the editing capability. The editing
page does not specify upload size/format limits; do not copy the VL limits.

Direct HTTP editing fragment; `image_path` names a local image approved for
upload and `prompt` is the editing instruction. Requires `requests`; no SDK
change is needed. This makes one request with no automatic retry:

```python
import base64
import os
import requests

if not isinstance(prompt, str) or not prompt.strip() or len(prompt) > 1024:
raise ValueError("Expected a non-empty editing instruction of at most 1024 characters")
with open(image_path, "rb") as source:
response = requests.post(
"https://chat.ecnu.edu.cn/open/api/v1/images/edits",
headers={"Authorization": f"Bearer {os.environ['ECNU_API_KEY']}"},
data={"model": "ecnu-image", "prompt": prompt, "response_format": "b64_json"},
files={"image": source},
timeout=120,
)
response.raise_for_status()
body = response.json()
if not isinstance(body, dict) or body.get("err_message"):
raise RuntimeError("Image edit failed; no output accepted")
items = body.get("data")
if not isinstance(items, list) or len(items) != 1 or not isinstance(items[0], dict):
raise RuntimeError("Expected one edited image")
encoded = items[0].get("b64_json")
if not isinstance(encoded, str) or not encoded:
raise RuntimeError("Image edit returned no image bytes")
edited_bytes = base64.b64decode(encoded, validate=True)
if not edited_bytes:
raise RuntimeError("Image edit returned empty image bytes")
```

Decode with the application's image loader before accepting or saving these
bytes, then save to a new asset path and inspect the result before replacing
the source. The fragment leaves the input file untouched. Actual edit quality,
character retention and output dimensions still require live verification.

## Thinking and tool history

Use the [ECNU thinking contract](https://developer.ecnu.edu.cn/vitepress/llm/thinking.html)
and the wire format of the chosen endpoint. For Chat Completions, enable
thinking through `thinking.type`; direct `reasoning_effort` uses `low`, `high`,
or `max` on `ecnu-max` and is ignored by `ecnu-plus`. Pass ECNU extensions through
`extra_body` with the OpenAI SDK. Do not substitute upstream template switches.
thinking through `thinking.type`; direct `reasoning_effort` uses `low` / `high` /
`max` on `ecnu-max`, and `low` / `medium` / `xhigh` on `ecnu-plus`. Effort applies
when thinking is enabled; thinking defaults to off. Pass ECNU extensions through
`extra_body` with the OpenAI SDK. Do not substitute upstream template switches
or silently map an unsupported tier between models.

For a tool exchange, append the complete actual assistant message before the
matching tool results. ECNU documents preserving `reasoning_content` for
Expand Down
5 changes: 5 additions & 0 deletions references/known_deviations.md
Original file line number Diff line number Diff line change
Expand Up @@ -193,6 +193,11 @@ Only response structure, field-preservation checks, and usage were retained.

## Direct `ecnu-max` image input

Current-contract note (documentation checked 2026-10-02, not a retest): the
[2026-09-18 model change](https://developer.ecnu.edu.cn/vitepress/llm/model.html)
made `ecnu-max` text-only again. Use `ecnu-plus` for image understanding.
The resolved status below describes only the September 12 fixture.

- **Tested at:** 2026-09-12; previous failure on 2026-08-23
- **Environment:** live-2026-09-12-a, direct HTTP; historical live-2026-08-23-a
- **Protocol and endpoint:** OpenAI-compatible `POST /chat/completions`
Expand Down
28 changes: 28 additions & 0 deletions references/models.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,34 @@ a measured quality or latency ranking.
| What changed and when | [Release notes](https://developer.ecnu.edu.cn/vitepress/llm/release.html) |
| Is there a reported incident? | [Service status](https://chat.ecnu.edu.cn/status) |

## Select the capability, not just a new model name

The [2026-09-30 release](https://developer.ecnu.edu.cn/vitepress/llm/release.html)
added multimodal retrieval and image editing. Contract checked 2026-10-02;
these are documented capabilities, not new live-test results.

| Model | Backend in the current model page | Integration choice |
|---|---|---|
| `ecnu-embedding-small` | bge-m3 | Keep for existing text-only, 1024-dimensional indexes |
| `ecnu-embedding-vl` | Qwen3-VL-Embedding-8B | Text, one image per item, or both; default 4096 dimensions, optional 1024/2048/4096 |
| `ecnu-rerank` | bge-reranker-v2-m3 | Text query and text candidates |
| `ecnu-rerank-vl` | Qwen3-VL-Reranker-8B | Text/image query and candidates; rerank an already retrieved candidate set |
| `ecnu-image` | qwen-image-2.1 | Generate from text or edit one existing image; different request formats and prompt handling |

Use the [retrieval and image recipes](examples.md) for the differences that
affect requests. Equal vector dimensions do not imply compatible embedding
spaces: changing embedding models requires re-embedding the indexed items and
using that same model and dimensions for queries. A reranker can change without
replacing stored vectors; its relevance scores are not scientific or educational
quality judgments and should not be compared across models.

For image understanding, the current model page specifies `ecnu-plus`.
`ecnu-max` currently uses DeepSeek-V4-Flash-0731 and is text-only; older
DeepSeek-V4.1 and successful vision observations are historical, not current
capability guarantees. Keep stable ECNU aliases in requests and recheck backend
labels before displaying them. Do not relabel old generated assets as output
from a newly announced backend.

## ECNU-specific boundaries

Use primary ECNU model names for new integrations rather than upstream names
Expand Down
11 changes: 10 additions & 1 deletion references/workflows.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ These links are dated observations, not claims that the issue still reproduces.
| Tool continuation loses state or reasoning fields | Preserve the actual message and inspect serialization: [recipe](examples.md#thinking-and-tool-history), [field variation](known_deviations.md#max-thinking-response-fields) |
| TTS error parsing crashes | Tolerate non-JSON errors: [invalid voice](known_deviations.md#invalid-tts-voice-error-shape) |
| PCM bytes arrive without format metadata | Configure the format explicitly: [missing headers](known_deviations.md#successful-tts-response-headers) |
| Historical max-vision limitation conflicts with current docs | Check the chosen protocol and current contract: [resolved fixture](known_deviations.md#direct-ecnu-max-image-input) |
| Old max-vision success conflicts with current docs | Current max is text-only; the [September 12 fixture](known_deviations.md#direct-ecnu-max-image-input) predates the September 18 contract change |
| `422` | Inspect `detail` and the relevant [request contract](api_reference.md); do not retry the unchanged request |
| `429` | Check [current credits/quota](models.md); stop parallel retries |
| Timeout or dropped connection after POST | Completion and debit may be unknown; do not automatically resubmit |
Expand All @@ -44,6 +44,9 @@ disables POST retries, reserves estimated credits, and redacts reports. Its
50-credit default is a cap, not authorization; use the lower approved cap and
obtain separate authorization before exceeding 50. Estimates are not actual debit.
Keep the same cumulative allowance across reruns, not a new allowance per process.
Token estimates use base/off-peak rates, not the current peak/holiday multiplier;
include that multiplier when checking the approved allowance. Fixed-price
embedding, rerank, image and TTS calls do not use peak pricing.

To inspect options without a network request:

Expand All @@ -65,6 +68,12 @@ and `billable` group other probes. Use `--case` to narrow them. Do not use `all`
for an ordinary integration check. Image/TTS probes need authorization covering
those billable operations and must not be scheduled automatically.

The runner does not yet have live cases for VL retrieval or image editing.
For those tasks, adapt the [minimal recipes](examples.md) to one authorized,
serial check with synthetic or approved media; preserve the same timeout,
no-retry and redaction boundaries. Offline recipe tests validate request
serialization and error handling, not service availability or output quality.

## Interpret the evidence

| Runner result | Meaning |
Expand Down
Loading
Loading