Update ECNU multimodal retrieval and image editing guidance - #6
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e80b9109e6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| + output_tokens * output_rate | ||
| ) / 1_000_000 | ||
| return credits, f"official {model} miss/hit/output token formula" | ||
| return credits, f"{model} base/off-peak token estimate; peak multiplier not applied" |
There was a problem hiding this comment.
Apply peak pricing before enforcing the credit cap
When a live run occurs during peak/holiday pricing, this returns only the base-rate estimate, while run_one reserves the similarly unmultiplied estimated_credits against --max-credits. Because the runner has no multiplier input, it can admit calls and report consumption below the actual peak-priced cumulative cost, exceeding the user's approved cap despite the workflow's instruction to account for the multiplier. Apply the multiplier to both preflight reservations and reported consumption, or require it as an explicit input.
AGENTS.md reference: AGENTS.md:L41-L42
Useful? React with 👍 / 👎.
| for case_id in ("responses_max_vision_compatibility", "anthropic_max_vision_compatibility"): | ||
| self.assertEqual(self.case(case_id).profiles, {"compatibility"}) |
There was a problem hiding this comment.
Accept documented max-vision rejection in compatibility probes
When either compatibility endpoint enforces the newly documented text-only ecnu-max contract by returning 400 or 422, retaining these probes only under the compatibility profile is insufficient: their specs still allow only (200,) and their response matchers require successful output, so the runner reports a correct rejection as a mismatch. Update or remove these probes so documented text-only rejection is an expected observational outcome.
AGENTS.md reference: AGENTS.md:L23-L23
Useful? React with 👍 / 👎.
Summary
ECNU's new multimodal retrieval models were reported as undocumented, and the skill lacked image-edit guidance. Add the relevant endpoint routing and request examples, distinguish vector spaces and generation/edit behavior, and correct thinking and current vision guidance.
Evidence
visible-but-undocumented, rerank at0.1, and both chat models in core vision.After: The same check recognizes both IDs, reports
0.05, and selects onlyecnu-plusfor core vision.skills-ref, compileall, andgit diff --checkpass.Merge Danger
Door: two-way
Reverting restores the previous skill and diagnostics. No application code or stored vectors are changed.
Blast Radius: skill
Agents using this skill receive updated guidance and diagnostic expectations. The new recipes do not run automatically or migrate existing indexes.