Skip to content

fix(tools): exclude thought text from preloaded memories - #7473

Open
rioyu123 wants to merge 1 commit into
google:mainfrom
rioyu123:codex/preload-memory-visible-text
Open

rioyu123 wants to merge 1 commit into
google:mainfrom
rioyu123:codex/preload-memory-visible-text

Conversation

@rioyu123

@rioyu123 rioyu123 commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Link to Issue or Description of Change

Problem:
PreloadMemoryTool includes parts marked thought=True when it turns recalled memories into a <PAST_CONVERSATIONS> block. This presents the model's earlier reasoning as ordinary conversation text, even when that reasoning considers an option the final answer rejects.

Steps to reproduce:

  1. Store a model event containing a thought part (Consider a zeppelin, but reject that option.) and a visible answer (Take the train to Paris.).
  2. Search for Paris in a later session using PreloadMemoryTool.
  3. Inspect the next model request. Both texts appear inside <PAST_CONVERSATIONS>; only the visible answer should.

This reproduces with both InMemoryMemoryService and SqliteMemoryService at 128faabb6dddacf26e029aa8c28c903c3a7a6b61 (ADK 2.11.0). The checks use an offline recording model, not LiteLLM or a live model API.

Solution:
Skip thought-marked parts when extracting text for preloading, consistent with ADK's other visible-text extractors. Parts with thought=False or thought=None are retained. Stored content, search behavior, structured load_memory results, and timestamp formatting are unchanged.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally.

Added six parameterized cases covering thought flags, unchanged memory content, and empty/non-text/thought-only entries without timestamps. Without the fix, two cases fail (thought=True and thought-only); with the fix, the three focused tool and memory test files pass (98 tests on Windows, Python 3.12.13).

Full-suite validation on Linux uses tox for Python 3.11 and tox -p 3 -e py312,py313,py314 for the remaining versions:

Python Passed Failed Skipped Xfailed / Xpassed
3.11 18,360 4 88 24 / 2
3.12 18,351 4 89 24 / 2
3.13 18,351 4 89 24 / 2
3.14 18,351 4 89 24 / 2

The four failures in every version are the order-dependent LiteLLM logging tests already tracked in #7461. A full run on the unmodified base using the same Python 3.11 environment reproduced exactly those failures (18,354 passed, 4 failed); the patch adds six passing cases. The full-suite checkboxes remain unchecked because the suite is not entirely green.

pre-commit run --files src/google/adk/tools/_memory_entry_utils.py tests/unittests/tools/test_preload_memory_tool.py and uv build pass.

Manual End-to-End (E2E) Tests:

An offline recording BaseLlm, an Agent with PreloadMemoryTool, and a real Runner exercise a new session against each memory backend. The checks verify that the request contains the visible answer but not thought text, and that searching the stored memory still returns its original thought part. No paid model calls are made.

Built the wheel with uv build, installed it with aiosqlite in a clean Linux Python 3.12.13 environment, and ran the script below outside the checkout:

InMemoryMemoryService: PASS - visible answer recalled; thought excluded; stored payload unchanged
SqliteMemoryService: PASS - visible answer recalled; thought excluded; stored payload unchanged
Offline Runner reproduction
import asyncio
from google.adk.agents import Agent
from google.adk.events.event import Event
from google.adk.memory.in_memory_memory_service import InMemoryMemoryService
from google.adk.memory._sqlite_memory_service import SqliteMemoryService
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_response import LlmResponse
from google.adk.runners import Runner
from google.adk.sessions.in_memory_session_service import InMemorySessionService
from google.adk.sessions.session import Session
from google.adk.tools.preload_memory_tool import PreloadMemoryTool
from google.genai import types
from pydantic import PrivateAttr

class RecordingModel(BaseLlm):
    _requests: list = PrivateAttr(default_factory=list)

    async def generate_content_async(self, llm_request, stream=False):
        self._requests.append(llm_request.model_copy(deep=True))
        yield LlmResponse(content=types.ModelContent('Take the train to Paris.'))

async def main():
    thought = 'Consider a zeppelin, but reject that option.'
    answer = 'Take the train to Paris.'
    for backend in (InMemoryMemoryService(), SqliteMemoryService(':memory:', fts='off')):
        historical = Session(app_name='memory_smoke', user_id='alice', id='previous', events=[
            Event(author='assistant', content=types.Content(role='model', parts=[
                types.Part(text=thought, thought=True), types.Part(text=answer)
            ]))
        ])
        await backend.add_session_to_memory(historical)
        before = await backend.search_memory(app_name='memory_smoke', user_id='alice', query='Paris')
        original_contents = [entry.content.model_copy(deep=True) for entry in before.memories]
        sessions = InMemorySessionService()
        session = await sessions.create_session(app_name='memory_smoke', user_id='alice')
        model = RecordingModel(model='offline-recording-model')
        runner = Runner(app_name='memory_smoke', agent=Agent(name='assistant', model=model, tools=[PreloadMemoryTool()]), session_service=sessions, memory_service=backend)
        try:
            events = [event async for event in runner.run_async(user_id='alice', session_id=session.id, new_message=types.UserContent('Paris'))]
            assert events and len(model._requests) == 1
            text = '\n'.join(part.text for content in model._requests[0].contents for part in content.parts or [] if part.text)
            assert '<PAST_CONVERSATIONS>' in text and answer in text
            assert thought not in text
            stored = await backend.search_memory(app_name='memory_smoke', user_id='alice', query='Paris')
            assert [entry.content for entry in stored.memories] == original_contents
            assert any(part.thought and part.text == thought for entry in stored.memories for part in entry.content.parts or [])
            print(type(backend).__name__ + ': PASS - visible answer recalled; thought excluded; stored payload unchanged')
        finally:
            await runner.close()
            if isinstance(backend, SqliteMemoryService):
                await backend.close()

asyncio.run(main())

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules. (No dependent changes.)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants