feat(tools): a Literal-typed tool argument is sent as a JSON-schema enum - #4575
Conversation
|
🔍 Technical detailsLocal Gemma-4-E4B, sandboxed fresh install with its own embedded Lemonade; Agent UI servers on :4301 (integration) and :4302 (
|
15cc93e to
63c4892
Compare
Asked to look up a GitHub release, local Gemma called fetch_page with extract=True twice and got "Invalid extract mode 'True'" both times: the schema said only "string". Literal[...] annotations now add an enum to the tool schema, so the server's tool-call grammar can rule such values out; fetch_page's extract is the first to use it.
63c4892 to
92597c1
Compare
|
Verdict: Approve with suggestions. Merge once the web eval in the test plan has run. A tool argument typed as a fixed set of choices now tells the model which values are allowed.
Real-world evidenceN/A: no evidence bundle was produced for this run. The change is an internal tool schema with no CLI, API, or UI output to capture, and its real effect on model behavior only shows up in the pending eval. This verdict rests on reading the code. I couldn't run the unit tests either, because pytest isn't installed on this runner. 🔍 Technical detailsIssues 🟡 Eval required before merge. Tool schema changes are on the CLAUDE.md "REQUIRE an eval run" list. Run the web/browser category on the Strix Halo pool ( 🟢 The text-mode tool prompt drops the enum ( 🟢 The agent server's tool definitions drop the enum ( 🟢 A mixed-type Strengths
|
Asked to look up the latest Lemonade release, local Gemma called
fetch_pagewithextract=True— twice, identically — and burned two steps on "Invalid extract mode 'True'". The tool's schema only said"type": "string", so nothing told the model (or the server's tool-call grammar) which values exist. ALiteral[...]-annotated tool argument now carries a JSON-schemaenum;fetch_page.extractis the first to use it. Other tools can opt in by changing an annotation.Draft until an eval runs — this changes the tool schema sent to the model.
Test plan
pytest tests/unit/test_tool_enum_args.py—Literal→enumin the registry and in the OpenAI tool schema;fetch_page.extractliststext, html, links, tablespytest tests/unit/test_tool_decorator.py tests/unit/test_browser_tools.pypass (test_tool_admission_order::test_cap_is_restored_at_the_next_turn_boundaryfails on main too)gaia eval agentweb category on this branch and on main