Whistle: 16.9 MB speech-to-text model built for on-device use
Cactus-Compute published Whistle, a 16.9 MB speech-to-text model for on-device use. The author claims it rivals Whisper base. For companies handling sensitive audio, the announcement opens a case for local transcription.
Source: huggingface.co (huggingface.co). Text prepared with AI from this source.
What happened and what to do
Cactus-Compute published Whistle, a 16.9 MB speech-to-text model available on a public model hosting platform. According to the publication, the author claims the model rivals Whisper base for on-device use, a setting where size and privacy matter more than server scale.
For a company, the practical decision is to test whether an ASR model of this size meets a real workflow before adopting it. An evaluation can measure word error rate, latency and memory use on recordings representative of the client, in Portuguese and with the business vocabulary, compared against Whisper base. If results are positive, transcription can run inside the app or at service terminals without sending audio to third parties, with a fallback to a private server when confidence is low.
How the consultancy can help
Wendelmaques can diagnose where local transcription would make sense in your product or operation, build a test set of recordings and compare compact models with Whisper base on accuracy, latency and memory. It then implements the chosen pipeline on the client's infrastructure, with monitoring and ongoing operation.
Next step
If your company handles sensitive audio and wants to evaluate on-device transcription, send a short description of the case, with volume, languages and where the audio is generated. You will receive a scoped proposal for diagnosis and implementation.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Quoted per project
Request a proposal