fix(scenarios, simulator): red-team floor, rare quiet lines, unique names, coherence, caller STT and confirmations - #140
Open
KarthikAvinashFI wants to merge 2 commits into
Open
KarthikAvinashFI wants to merge 2 commits into
KarthikAvinashFI wants to merge 2 commits into
Conversation
…name uniqueness, coherence and real-use-case wording
…read-backs, stay in the caller role
KarthikAvinashFI
marked this pull request as ready for review
October 2, 2026 00:25
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Scenario and simulated-caller quality fixes, found by reading generated suites end to end and by reviewing 60 calls from a production run (transcripts, recordings and evals). All changes are generic; nothing is specific to one agent.
Scenarios
real_use_casesmust be capabilities the agent's own text states, never stretched from a passing word or a tool name.Simulated caller
Speech-to-text. A caller whose language the multilingual model lacks (for example Mandarin or Korean) was transcribed in that language only, so the agent's English came back empty or garbled and the caller never understood the agent. It now uses the multilingual model. Checked by re-transcribing two affected recordings: the old setting returned nothing, the new one recovers the agent's words.
The simulator transcribes the agent, not the caller (the caller's speech is generated). What each caller now gets with Deepgram nova-3 (the Mandarin row was checked on two real call recordings; the others follow Deepgram's multilingual language list):
Agents that answer in an uncovered language need a different STT provider; out of scope here.
Confirmations. The caller confirms a recap the way people do, a yes or the one wrong detail, instead of reading back the whole address or summary.
Role. The caller never says the agent's lines, such as a recap or a question asking whether to go ahead.
Meta questions. The "ask what a term means" caller move applies only to terms a person in the caller's place would not know, not to ordinary words.
Sub-goal grading is unchanged.
Verification
Scenario generation only, all scenarios read:
Tests: new tests for the red-team floor, the quiet-line cap and the family-name cap; the speech-to-text test updated. The
tests/harnessandtests/test_harness.pyfailures are identical todev.