AI Rudder published self evaluated τ-voice benchmark results for its voice agent: 83.89% pass@1 with the Standard customer simulator (237 of 278 tasks) and 88.82% with the Custom simulator (250 of 278 tasks) across retail, airline and telecom, above the prior published best of 81.72% and 86.23%. The post discloses the evaluated stack (Deepgram nova-3 speech recognition, gemini-3.7-flash, ElevenLabs flash v2.5 speech, no training or fine tuning) and the runtime controls around the model, including an enabled check that blocks a repeated identical tool write within a turn.
Our readLenders running voice agents on collections and payment calls need evidence the agent completes account actions once and correctly, and this gives a stated method and a named model stack to test against. The scores are self reported, exclude the τ-voice Banking domain, and were not on the official leaderboard when checked.