Audio reasoning: quality and latency put to the test
In a publication dated October 13, 2025, Artificial Analysis reports a 92% score on an audio benchmark with 1,000 questions and compares time to first token with and without reasoning.
Source: Gemini 2.5 Native Audio Thinking no Big Bench Audio (x.com). Text prepared with AI from this source.
What happened and what to do
On October 13, 2025, Artificial Analysis published results for Gemini 2.5 Native Audio Thinking on Big Bench Audio: 92% on 1,000 audio questions adapted from Big Bench Hard. According to the publication, average time to first token was 3.87 seconds with thinking and 0.63 seconds without it. The test offers a comparison of reasoning and latency for speech-to-speech models. Consult the original Artificial Analysis publication, titled “Gemini 2.5 Native Audio Thinking no Big Bench Audio,” and verify the results and methodology in the source.
For a company evaluating voice-based service, the practical response is to test quality and latency with questions representative of its own process, measure response times, and set acceptable limits before integrating a model. An evaluation and monitoring infrastructure can compare options and track performance in production; the highest-scoring option need not be the best fit for the use case.
How the consultancy can help
Wendelmaques can assess the audio process and its requirements, define an evaluation scope with quality and latency metrics, and implement test, integration, and monitoring pipelines on the client's infrastructure. The proposal can also cover operating the solution.
Next step
Send a short description of your audio use case to receive a scoped proposal for diagnosis, implementation, and operation.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Quoted per project
Request a proposal