Skip to content

feat(tools): a Literal-typed tool argument is sent as a JSON-schema enum - #4575

Merged
kovtcharov-amd merged 1 commit into
amd:mainfrom
kovtcharov:kalin/tool-enum-args
Oct 5, 2026
Merged

kovtcharov-amd merged 1 commit into
amd:mainfrom
kovtcharov:kalin/tool-enum-args

Conversation

@kovtcharov

Copy link
Copy Markdown
Contributor

Asked to look up the latest Lemonade release, local Gemma called fetch_page with extract=True — twice, identically — and burned two steps on "Invalid extract mode 'True'". The tool's schema only said "type": "string", so nothing told the model (or the server's tool-call grammar) which values exist. A Literal[...]-annotated tool argument now carries a JSON-schema enum; fetch_page.extract is the first to use it. Other tools can opt in by changing an annotation.

Draft until an eval runs — this changes the tool schema sent to the model.

Test plan

  • pytest tests/unit/test_tool_enum_args.py — Literal → enum in the registry and in the OpenAI tool schema; fetch_page.extract lists text, html, links, tables
  • pytest tests/unit/test_tool_decorator.py tests/unit/test_browser_tools.py pass (test_tool_admission_order::test_cap_is_restored_at_the_next_turn_boundary fails on main too)
  • gaia eval agent web category on this branch and on main

@github-actions github-actions Bot added tests Test changes agents labels Oct 2, 2026
@kovtcharov

Copy link
Copy Markdown
Contributor Author

gaia_web eval, same machine, back to back, with CI's fixture server and env: a build with this PR (plus the other open walkthrough fixes) passed 4/5, average 8.9/10; main passed 3/5, average 8.1/10. No fetch_page argument errors in either run.

🔍 Technical details

Local Gemma-4-E4B, sandboxed fresh install with its own embedded Lemonade; Agent UI servers on :4301 (integration) and :4302 (origin/main 842d014), GAIA_AUTO_APPROVE_TOOLS=1, GAIA_WEB_ALLOWED_HOSTS=127.0.0.1, tests/fixtures/gaia/serve_fixtures.py on :8765 as in test_gaia_agent_eval.yml. One run per scenario per build.

scenario this build main
web_download_file 4.8 FAIL 4.7 FAIL
web_fetch_honest_404 9.9 9.9
web_fetch_product_fact 9.8 7.4 FAIL
web_live_search_canary 9.8 9.9
web_multi_page_compare 9.8 9.9

web_download_file fails on both for a harness reason: the sandbox home sits under the system temp folder, which the download allowlist refuses.

@kovtcharov
kovtcharov force-pushed the kalin/tool-enum-args branch 3 times, most recently from 15cc93e to 63c4892 Compare October 4, 2026 21:03
Asked to look up a GitHub release, local Gemma called fetch_page with
extract=True twice and got "Invalid extract mode 'True'" both times: the
schema said only "string". Literal[...] annotations now add an enum to
the tool schema, so the server's tool-call grammar can rule such values
out; fetch_page's extract is the first to use it.
@kovtcharov
kovtcharov force-pushed the kalin/tool-enum-args branch from 63c4892 to 92597c1 Compare October 4, 2026 22:09
@kovtcharov-amd
kovtcharov-amd marked this pull request as ready for review October 5, 2026 03:44
@kovtcharov-amd
kovtcharov-amd self-requested a review as a code owner October 5, 2026 03:44
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Verdict: Approve with suggestions. Merge once the web eval in the test plan has run.

A tool argument typed as a fixed set of choices now tells the model which values are allowed. fetch_page is the first to use it, so local Gemma should stop sending extract=True and wasting steps on the error. The change is small and correct, and existing tools are unaffected.

  • The eval is the gate. This changes the tool schema sent to the model, which CLAUDE.md says needs a gaia eval agent run before merge. The PR already tracks this as a draft item. Compare the web category against the flagship baseline, not only against main.
  • Two other tool listings still leave out the allowed values (both optional, see details): the text tool list for models without native tool calling, and the tool definitions the agent server publishes. Neither breaks anything, but those consumers get no benefit yet.

Real-world evidence

N/A: no evidence bundle was produced for this run. The change is an internal tool schema with no CLI, API, or UI output to capture, and its real effect on model behavior only shows up in the pending eval. This verdict rests on reading the code. I couldn't run the unit tests either, because pytest isn't installed on this runner.

🔍 Technical details

Issues

🟡 Eval required before merge. Tool schema changes are on the CLAUDE.md "REQUIRE an eval run" list. Run the web/browser category on the Strix Halo pool (eval_flagship.yml) and compare against tests/fixtures/eval_baselines/gaia-flagship/. List PASS→FAIL and FAIL→PASS separately.

🟢 The text-mode tool prompt drops the enum (src/gaia/agents/base/agent.py:2468-2473). _format_tools_for_prompt renders extract?: string, so a model without native tool calls still has to guess. One option is to render the choices there too, e.g. extract?: "text"|"html"|"links"|"tables". This touches prompt assembly, so it would fall under the same eval.

🟢 The agent server's tool definitions drop the enum (src/gaia/agents/base/server.py:202-213). That builder copies type and description but not enum, so MCP/sidecar clients see a plain string. To match _build_openai_tool_schemas:

                param_description = param_info.get("description", "")
                if param_description:
                    prop["description"] = param_description
                if param_info.get("enum"):
                    prop["enum"] = list(param_info["enum"])

🟢 A mixed-type Literal takes its type from the first choice (src/gaia/agents/base/tools.py:283). Literal["a", 1] becomes type: "string" with an integer in the enum. No current tool does this. A one-line comment, or leaving the type as unknown when the choices' types differ, would head it off.

Strengths

  • The tool keeps its own valid_modes check, so a backend that ignores enum still gets the clear error message.
  • _literal_choices handles Optional[Literal[...]] the same way _infer_param_type handles unions, so it reads like the surrounding code.
  • The second test checks the actual OpenAI schema fetch_page produces, not just the registry entry.

@kovtcharov-amd
kovtcharov-amd added this pull request to the merge queue Oct 5, 2026
Merged via the queue into amd:main with commit 82e294e Oct 5, 2026
84 of 85 checks passed
@kovtcharov-amd
kovtcharov-amd deleted the kalin/tool-enum-args branch October 5, 2026 03:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agents tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants