Skip to content

fix(scenarios, simulator): red-team floor, rare quiet lines, unique names, coherence, caller STT and confirmations - #140

Open
KarthikAvinashFI wants to merge 2 commits into
devfrom
fix/scenario-names-redteam-coherence
Open

KarthikAvinashFI wants to merge 2 commits into
devfrom
fix/scenario-names-redteam-coherence

Conversation

@KarthikAvinashFI

@KarthikAvinashFI KarthikAvinashFI commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

What

Scenario and simulated-caller quality fixes, found by reading generated suites end to end and by reviewing 60 calls from a production run (transcripts, recordings and evals). All changes are generic; nothing is specific to one agent.

Scenarios

  • Red-teaming floor. A writer's slice is refused a plain scenario once only its attack slots remain, so at least a fifth of every slice red-teams the agent whenever the plan deals attack levels. The planner's pre-brief check says the same. Chat suites that grow are exempt.
  • Quiet lines stay rare. The share check capped every level at a third but exempted the interface axis, so one run dealt quiet lines to 40% of the suite. A quiet interface level is now held to a sixth of a slice.
  • Names. Uniqueness moved from first names to full names; requiring a distinct first name across hundreds of callers pushed writers to famous and fictional names. A family name is refused a third time. The writer checklist rules out names close to a famous one.
  • Coherence. New checklist item: who calls, who travels, where they are, the sound around them and what they hold all agree; the caller knows only what someone in their place would; nothing presumes a record the agent's world does not hold. Every address, a home included, is a real place with its city.
  • No invented use cases. real_use_cases must be capabilities the agent's own text states, never stretched from a passing word or a tool name.
  • Run-specific authoring policy now sits right after the discovered skills in the planner and writer prompts instead of at the very end.

Simulated caller

  • Speech-to-text. A caller whose language the multilingual model lacks (for example Mandarin or Korean) was transcribed in that language only, so the agent's English came back empty or garbled and the caller never understood the agent. It now uses the multilingual model. Checked by re-transcribing two affected recordings: the old setting returned nothing, the new one recovers the agent's words.
    The simulator transcribes the agent, not the caller (the caller's speech is generated). What each caller now gets with Deepgram nova-3 (the Mandarin row was checked on two real call recordings; the others follow Deepgram's multilingual language list):

    Caller languages STT setting Agent speaking English Agent speaking the caller's language
    English multi heard n/a
    Hindi + English, or Spanish, French, German, Japanese, Russian, Portuguese, Italian, Dutch (+ English) multi heard heard, including mid-sentence switching
    Mandarin, Turkish, Korean, Arabic (alone or + English); Hindi + Mandarin multi (was the caller's language alone) heard (was empty or garbled) not transcribed: nova-3 multilingual does not cover these and its single-language modes cannot follow a switch to English

    Agents that answer in an uncovered language need a different STT provider; out of scope here.

  • Confirmations. The caller confirms a recap the way people do, a yes or the one wrong detail, instead of reading back the whole address or summary.

  • Role. The caller never says the agent's lines, such as a recap or a question asking whether to go ahead.

  • Meta questions. The "ask what a term means" caller move applies only to terms a person in the caller's place would not know, not to ordinary words.

Sub-goal grading is unchanged.

Verification

Scenario generation only, all scenarios read:

Run Agent Size Attacks Quiet lines Notes
before ride booking (prod) 500 7% - first-name rule, TV-character personas, city-less addresses
r28 ride booking 100 17% 7% quiet cap
r29 ride booking 200 18% 6% famous-name echoes remained
r30 ride booking 200 22% 4% invented use case; surnames repeated
r31 ride booking 500 19% 4% real use cases only, no repeated full names, max 2 per family name
r35 ride booking 100 27% 15% final branch
r35 auto insurance 100 23% 16% final branch on an unrelated agent: 7 real use cases, coherent, agent-specific sub-goals

Tests: new tests for the red-team floor, the quiet-line cap and the family-name cap; the speech-to-text test updated. The tests/harness and tests/test_harness.py failures are identical to dev.

…name uniqueness, coherence and real-use-case wording
@KarthikAvinashFI KarthikAvinashFI self-assigned this Oct 1, 2026
@KarthikAvinashFI KarthikAvinashFI changed the title fix(scenarios): red-team floor, rare quiet lines, unique full names, coherence fix(scenarios, simulator): red-team floor, rare quiet lines, unique names, coherence, caller STT and confirmations Oct 2, 2026
@KarthikAvinashFI
KarthikAvinashFI marked this pull request as ready for review October 2, 2026 00:25

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant