Vercel AI experimental_evaluate with openrouter.evaluationModel('typesafe/jev-1.13') (@openrouter/ai-sdk-provider 3.1.0) gives an evaluate span, but the answers in gen_ai.output.messages lose confidence and legend.
The OpenRouter provider moves these fields into providerMetadata.openrouter.answers[<question id>] ({ confidence, legend }). Our withProviderConfidence in packages/server-utils/src/integrations/vercel-ai/vercel-ai-dc-subscriber.ts only reads providerMetadata.typesafe.confidence, which is the shape of @ai-sdk/typesafe-ai.
Expected
Evaluate answers from the OpenRouter provider keep confidence (choice, score) and legend (score), the same as answers from @ai-sdk/typesafe-ai.
Notes
Vercel AI
experimental_evaluatewithopenrouter.evaluationModel('typesafe/jev-1.13')(@openrouter/ai-sdk-provider3.1.0) gives an evaluate span, but the answers ingen_ai.output.messagesloseconfidenceandlegend.The OpenRouter provider moves these fields into
providerMetadata.openrouter.answers[<question id>]({ confidence, legend }). OurwithProviderConfidenceinpackages/server-utils/src/integrations/vercel-ai/vercel-ai-dc-subscriber.tsonly readsproviderMetadata.typesafe.confidence, which is the shape of@ai-sdk/typesafe-ai.Expected
Evaluate answers from the OpenRouter provider keep
confidence(choice, score) andlegend(score), the same as answers from@ai-sdk/typesafe-ai.Notes
providerMetadata.openrouter.usage.costandprovider. We could record these too, but they are optional.experimental_evaluate(feat(server-utils): Instrument Vercel AI experimental_evaluate #24694).