SequentialAgent Architecture: Why Send Chat History When Using output_key + Template Variables? #3457
|
I'm analyzing AgentOps traces from my ADK SequentialAgent pipeline and observed an interesting pattern in how data flows between agents. I'd like to understand the design rationale behind it. Observed BehaviorWhen Agent 2 receives data from Agent 1 in a SequentialAgent pipeline, the data appears twice in the LLM request: 1. Template Substitution in System InstructionThe "system_instruction": "You are a professional Spanish translator...
ARTICLE SUMMARY:
{'title': 'Phoenix Suns Win Historic NBA Championship...',
'key_points': [...],
'main_subject': 'Phoenix Suns',
'event_outcome': 'NBA Championship Win',
'word_count': 210,
'sentiment': 'positive'}
Translate the following into Spanish..."2. Chat History ContextThe same data also appears in the conversation history as a user message: {
"content": "For context:\n[summarizer_agent] said: {\n \"title\": \"Phoenix Suns Win Historic...\",\n ...full summary JSON...\n}",
"role": "user"
}Code Pattern# Agent 1
summarizer_agent = Agent(
model="gemini-2.0-flash",
name="summarizer_agent",
instruction="...",
output_key="summary", # Saves to context.state['summary']
output_schema=ArticleSummary,
)
# Agent 2
translator_agent = Agent(
model="gemini-2.0-flash",
name="translator_agent",
instruction="""You are a professional Spanish translator.
ARTICLE SUMMARY:
{summary} # Template injection from state
Translate the following into Spanish...
""",
output_key="translation",
output_schema=SpanishTranslation,
)
news_pipeline = SequentialAgent(
name="news_pipeline",
sub_agents=[summarizer_agent, translator_agent],
)Questions
Environment
Trace EvidenceFull trace available showing both mechanisms: [would attach traces.json excerpt] Any insights into the design rationale would be greatly appreciated! I want to ensure I'm following best practices for multi-agent data flow. |
Replies: 5 comments 2 replies
|
Response from ADK Answering Agent (experimental, answer may be inaccurate) TLDR: This behavior is intentional and reflects two core ADK concepts: shared state for direct data passing and event history for complete observability. You can and should optimize this for large data payloads by configuring the agent to not include the full conversation history in the LLM prompt, which will reduce token usage. Hi @prabha-git, that's an excellent observation and a great question about the design of 1. Is this duplication intentional?Yes, the behavior you're seeing is intentional. It arises from two distinct and fundamental mechanisms in ADK that work together:
2. What value does chat history provide?The event history serves a different purpose than the state dictionary:
3. Are there scenarios where I should optimize this?Absolutely. Your instinct is correct. For large data payloads, sending the full history in addition to the templated data can be inefficient and lead to unnecessary token costs. ADK provides ways to control this. While the specifics can depend on the agent implementation, you can generally control what is included in the prompt sent to the LLM. You should be able to configure your By default, ADK leans towards providing full context (history) for robustness, but it gives you the option to optimize for token efficiency by customizing what the LLM sees. I hope this clarifies the design rationale! It's a great question that gets to the heart of how ADK manages data and context. [1] https://google.github.io/adk-docs/agents/workflow-agents/sequential-agents/ |
Follow-up:
|
|
I think you can modify the chat history by replacing the entire message to be something like |
|
Thanks for the suggestion, @zmy1116! Compressing agent messages is definitely an improvement over the current behavior. However, in sequential pipelines where template variables pass all necessary data, message compression doesn't fully eliminate the duplication concern: # What the second agent receives:
System Instruction: "Translate this summary: {full_json_summary}"
Chat History: "[summarizer_agent] said: message compacted"The full data is already in the instruction via For maximum efficiency in sequential workflows, an option to skip agent messages entirely from chat history would be ideal: translator_agent = Agent(
instruction="Translate {summary}",
include_contents='none',
skip_agent_messages=True # ← Feature request
) |
Update: Root Cause Identified & Workaround ImplementedI've found the root cause and successfully implemented a workaround for this issue. Root CauseThis behavior is tracked in Issue #2207 and was closed as "NOT_PLANNED". The ADK team considers the automatic injection of agent transition messages ( This means there's no built-in configuration to disable it, but we can work around it using callbacks. Solution:
|
Update: Root Cause Identified & Workaround Implemented
I've found the root cause and successfully implemented a workaround for this issue.
Root Cause
This behavior is tracked in Issue #2207 and was closed as "NOT_PLANNED". The ADK team considers the automatic injection of agent transition messages (
[agent_name] said:) intentional.This means there's no built-in configuration to disable it, but we can work around it using callbacks.
Solution:
before_model_callbackFor template-only pipelines where agents receive all data through template variables (e.g.,
{summary}), you can usebefore_model_callbackto clear conversation history before the request is sent to the LLM: