Skip to content

Update ECNU multimodal retrieval and image editing guidance - #6

Merged
JJasonSun merged 1 commit into
mainfrom
feat/multimodal-api-update
Oct 2, 2026
Merged

JJasonSun merged 1 commit into
mainfrom
feat/multimodal-api-update

Conversation

@JJasonSun

Copy link
Copy Markdown
Owner

Summary

ECNU's new multimodal retrieval models were reported as undocumented, and the skill lacked image-edit guidance. Add the relevant endpoint routing and request examples, distinguish vector spaces and generation/edit behavior, and correct thinking and current vision guidance.

- VL models classified as undocumented; text rerank estimated at 0.1 credits
+ VL models recognized; text rerank 0.05, VL retrieval 0.2 credits
- Core vision checks include ecnu-max
+ Core vision uses ecnu-plus; dated observations remain historical
+ Multipart image-edit recipe and mixed text/image retrieval payloads

Evidence

  • Before: Offline inspection of main reports both VL IDs as visible-but-undocumented, rerank at 0.1, and both chat models in core vision.
    After: The same check recognizes both IDs, reports 0.05, and selects only ecnu-plus for core vision.
  • All 76 offline tests pass, including multipart file serialization, image rejection handling, VL JSON preservation, model classification, and pricing checks.
  • Repository validation, Skill Creator validation, skills-ref, compileall, and git diff --check pass.
  • Contracts checked against the official model, image-edit, embedding, rerank, thinking, and pricing pages. No live ECNU requests were made; edit quality and new endpoint availability remain unverified. Token estimates explicitly identify their base/off-peak rates.

Merge Danger

Door: two-way

Reverting restores the previous skill and diagnostics. No application code or stored vectors are changed.

Blast Radius: skill

Agents using this skill receive updated guidance and diagnostic expectations. The new recipes do not run automatically or migrate existing indexes.

@JJasonSun
JJasonSun marked this pull request as ready for review October 2, 2026 14:50
@JJasonSun
JJasonSun merged commit 9772662 into main Oct 2, 2026
2 checks passed
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T14:56:53.444406Z e80b910 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e80b9109e6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/smoke_test.py
+ output_tokens * output_rate
) / 1_000_000
return credits, f"official {model} miss/hit/output token formula"
return credits, f"{model} base/off-peak token estimate; peak multiplier not applied"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Apply peak pricing before enforcing the credit cap

When a live run occurs during peak/holiday pricing, this returns only the base-rate estimate, while run_one reserves the similarly unmultiplied estimated_credits against --max-credits. Because the runner has no multiplier input, it can admit calls and report consumption below the actual peak-priced cumulative cost, exceeding the user's approved cap despite the workflow's instruction to account for the multiplier. Apply the multiplier to both preflight reservations and reported consumption, or require it as an explicit input.

AGENTS.md reference: AGENTS.md:L41-L42

Useful? React with 👍 / 👎.

Comment thread tests/test_smoke_test.py
Comment on lines +968 to +969
for case_id in ("responses_max_vision_compatibility", "anthropic_max_vision_compatibility"):
self.assertEqual(self.case(case_id).profiles, {"compatibility"})

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Accept documented max-vision rejection in compatibility probes

When either compatibility endpoint enforces the newly documented text-only ecnu-max contract by returning 400 or 422, retaining these probes only under the compatibility profile is insufficient: their specs still allow only (200,) and their response matchers require successful output, so the runner reports a correct rejection as a mismatch. Update or remove these probes so documented text-only rejection is an expected observational outcome.

AGENTS.md reference: AGENTS.md:L23-L23

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant