AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute announced Whistle, a 16.9 MB speech-recognition model designed to transcribe audio on-device using its C++ engine. The company reports support for seven languages and favorable results against selected models on several benchmarks, but Whistle leads only on some datasets and the comparisons use different runtimes and settings.

Cactus Compute announced Whistle, a speech-recognition model packaged as a 16.9 MB file that the company says runs on a device’s CPU without external dependencies. It is intended for phones, wearables, robots, smart-home and automotive systems, and supports transcription in seven languages—a compact on-device option where network access, latency or sending audio to a server is a concern.

According to Cactus Compute’s October 2 release, Whistle processes 16 kHz mono audio in clips of up to 30 seconds. It can detect the spoken language automatically or use a language specified by the user. The supported languages are English, German, French, Spanish, Italian, Dutch and Polish. The company says audio stays on the device during transcription; its browser demonstration downloads the model when used.

The release describes two additional outputs beyond a transcript: word-level timestamps, with start and end times and a probability for each word, and speech embeddings that provide an encoder representation for successive audio frames without producing a transcript. Users can also provide keywords to bias recognition. Cactus says the model returns an empty transcript for audio below its loudness threshold rather than starting beam-search decoding.

For performance, Cactus reports that Whistle reaches its first token in 11.1 milliseconds and decodes at 1,319 tokens per second in a test using 10 seconds of audio on an Apple M4 Pro CPU. The company compared it with Whisper base and Moonshine tiny v2, running each through its stated official runtime and default settings. Those are vendor-reported figures, not an independent evaluation, and the release notes that Whisper pads inputs to 30 seconds while the other systems were tested differently.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute released Whistle, a compact speech-to-text model it says can run locally on CPUs and transcribe seven languages.

Local Transcription on Smaller Devices

A model that fits in 16.9 MB could make speech input practical in devices with limited storage or unreliable connectivity. If the stated on-device operation works as described, audio need not be sent to a cloud service for transcription, which may help reduce network dependence and keep processing local. Those advantages can matter for wearables, vehicle systems and embedded devices, where responsiveness, bandwidth and privacy constraints shape product design.

The reported speed is also relevant to interactive uses such as voice commands. Cactus says Whistle can run in the same CPU engine as its Needle model, and that both use the same container and quantization. The company presents this shared implementation as a way to connect speech recognition with tool-calling workflows. The announcement describes that possibility, but does not provide a separate evaluation of end-to-end tool use or demonstrate performance across the range of devices named as targets.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Cactus Describes Whistle

Whistle uses an audio encoder and text decoder, according to the technical account in Cactus Compute’s release. Its audio front end converts sound into log-mel features, then reduces the frame count before passing the result through eight encoder blocks. The decoder uses cross-attention to read the encoded audio while generating text. A five-beam search is used for decoding, and the release says the transcript is capped at 320 tokens.

Cactus says the encoder and decoder blocks are based on components shared with Needle, its existing model, rather than copied into a separate implementation. The release also says users can select decoder depth at load time, while all eight encoder blocks still run. These details describe the model’s design; they do not by themselves establish how it performs across different processors or real-world audio conditions.

The company’s benchmark chart presents a mixed comparison. It reports lower word error rates for Whistle than the compared systems on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base is ahead on TED-LIUM, AMI and the MLS average. The source also flags missing results where a model’s authors did not publish a score, and says Whisper’s AMI result uses a different subset from the AMI figures for the other systems. Results therefore vary by dataset and should not be read as a single overall ranking.

““one 16.9 MB file, runs on the CPU with no dependencies””

— Cactus Compute, in the Whistle announcement

Amazon

multilingual voice recognition device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Testing Is Still Needed

The performance and accuracy figures in the announcement come from Cactus Compute’s own testing. The supplied report does not include an independent reproduction, confidence intervals or a full account of how audio conditions and hardware differences might affect results. Runtime defaults differ, and the company identifies benchmark gaps and a difference in the AMI subset used for Whisper. These qualifications make direct comparisons difficult.

It is also not clear from the release how Whistle performs on devices other than the Apple M4 Pro used for the timing comparison, how recognition changes with accents or noisy recordings, or how much memory is required while the model is running. The company names several target device categories but does not report tests on specific phones, wearables, vehicles or microcontrollers. Its local-processing statement describes the intended operation; users would still need to review the implementation and deployment setup to verify data handling in a particular product.

Amazon

voice command smart home devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Availability and Device Tests

Cactus Compute’s release includes a browser demonstration and describes Whistle as available as a model that can be loaded into its C++ engine. The provided announcement does not set out a broader release schedule, licensing terms, supported device list or a plan for independent benchmark testing. Those details will help determine how readily developers can adopt the model.

The next evidence to watch for is testing beyond the company’s stated Apple M4 Pro comparison, alongside fuller documentation on installation, licensing, memory use and results across devices and audio conditions. Until those details are available, Whistle’s small file size and reported latency are concrete claims from its maker, while its wider compatibility and comparative accuracy remain to be independently established.

Amazon

wearable speech-to-text app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. The company says it runs on-device through a CPU-based C++ engine and is packaged as a 16.9 MB file.

Which languages does it support?

The release lists English, German, French, Spanish, Italian, Dutch and Polish. Cactus says Whistle can detect the language or accept a language specified by the user.

Does Whistle send recordings to the cloud?

Cactus says Whistle processes audio on the device and that audio in its browser demo does not leave the device. The announcement does not independently verify data handling in every possible deployment.

Is Whistle more accurate than Whisper?

Not across every dataset in the company’s comparison. Cactus reports better Whistle word-error-rate results on several benchmarks, while Whisper base scores ahead on TED-LIUM, AMI and the MLS average. The comparison is company-reported, and the release notes differences and missing results.

What devices have been tested?

The performance figures supplied in the announcement use an Apple M4 Pro CPU. Cactus lists phones, wearables, robots, smart-home and automotive systems, and microcontrollers as intended targets, but does not provide test results for those device categories.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Revolutionary Breakthroughs In AI, Mathematics, And Theoretical Computer Science

OpenAI publishes a list of ten recent advances in mathematics and theoretical computer science, highlighting AI’s growing role in formal scientific research.

The 2027 Roundup: 9 Best Wireless Gaming Mice By Play Style

A 2027 buying guide compares nine wireless gaming mouse listings by shape, controls, battery approach and intended play style.

The SSD Squeeze: Why Storage Joined the Party

Record-breaking NAND price increases driven by AI’s rising storage needs and wafer competition, impacting enterprise and consumer markets in 2026.

Samsung’s new wide foldable phone revealed in leaked images

Leaked images reveal Samsung’s upcoming Galaxy Z Fold 8 Ultra, featuring a wide foldable design, triple rear camera, and new case options, set for release this July.