The voice is only the entry point
Wen said customer-support agents need to sound natural enough that people believe they can solve a problem. Trust, in his account, can build over the first two or three exchanges. If the agent handles the task, a customer may decide there is no need to speak to a person.
Otter marketing director Alex Gay sees a more demanding test for meeting assistants. They need to identify speakers and intent, and shape what was said using the organization’s knowledge. Otter is also working on digital twins that could represent people in meetings; Gay said their voices should carry the emotional nuance of speaking with the person themselves.
For meetings, Gay said, disagreement, strategic discussion and relationships between participants matter. An avatar that cannot take part in that kind of exchange is reduced to a question-and-answer chatbot.
Errors compound downstream
Voice models have improved, but assistants still misunderstand users. Meeting-recording services can also mis-transcribe speech or produce an inaccurate summary. Wen said automatic speech recognition systems often miss important keywords, losing context.
Gay said Otter is still improving transcription, and that voice models’ language abilities need more work. Transcription was not the end goal for Otter; it was the foundation for productivity tools. If the transcript is wrong, later actions can be wrong too. Once a system acts on faulty information, users lose trust in the platform.
That makes transcription quality more than a matter of cleaner notes. It sits upstream of everything the product does with a conversation.
Trust needs disclosure, too
The speakers also raised a separate trust problem: people need to know when they are being recorded or speaking with AI. Otter is considering ways to notify everyone in a chat that a meeting is being recorded, even when its bot is not present. Wen said business calls should make clear from the outset that the other party is AI.
I think that disclosure is the more immediate test than a voice that sounds indistinguishable from a person. The conversation described here depends on users trusting both what the system understood and what it is. Better speech and faster reasoning may make an agent easier to use; they do not resolve either condition on their own.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X