SquadStack Opens Arth V2: Vendor Benchmark Puts 67 ms Latency First, Semantic WER 8.85 on 863 Indian Sales Calls
SquadStack made Arth V2 public on 11 Oct 2026: vendor-reported 67 ms finalize-to-response latency and 8.85% semantic WER on 863 Hindi-English sales calls. Standalone USD API pricing was not published.

SquadStack made Arth V2 public on 11 October 2026. The Noida based voice AI company says its in-house streaming speech recogniser already listens to about 80% of the more than 50 lakh calls its agents place each day.
On a held-out set of 863 real Hindi-English sales calls, SquadStack reports Arth V2 at 8.85% semantic word error rate and 67 ms finalize-to-response latency at the 80th percentile. Those numbers come from SquadStack’s own blog and from the open Conversational Streaming ASR Benchmark on Hugging Face. They are vendor figures, not an independent lab scorecard.
For Indian voice agent builders, the useful part is the test design: 8 kHz phone audio, Hinglish code-mixing, and a score that counts only errors that would change what an agent does next.
What SquadStack announced
CEO and co-founder Apurv Agrawal’s post frames Arth V2 as the first model from SquadStack’s “model factory”. New versions ship only when they beat the one already in production on real Indian sales calls.
Arth V2 runs inside SquadStack’s Humanoid Voice AI Agents. The company says every new agent now starts on Arth by default, and that Arth has already handled more than 1 crore minutes of live customer audio in two months. A separate public API with USD list pricing was not published in the announcement.
Training claims in the same post: more than 60 crore minutes of real Indian sales conversations from more than 10 crore unique speakers, covering over 85% of Indian pincodes. That corpus is proprietary contact-centre audio, not a scrape of public video or podcasts.
Vendor benchmark numbers, with Soniox in the mix
SquadStack scored eleven streaming systems on the same 863 calls through one pipeline. The headline metric is semantic WER: the share of reference words that would change a downstream agent action, not raw string WER.
| System | Vendor | Semantic WER | FTR P80 (ms) |
|---|---|---|---|
| Soniox v5 | Soniox | 8.77% | 89.4 |
| Arth V2 | SquadStack | 8.85% | 67.0 |
| Arth V1 | SquadStack | 9.17% | 67.0 |
| Nova 3 | Deepgram | 9.31% | 93.8 |
| Saaras v4 | Sarvam | 10.84% | 176.7 |
| Gemini 3.5 Transcribe | 11.93% | 264.2 | |
| gpt-live-transcribe | OpenAI | 12.30% | 830.8 |
Source: SquadStack Hugging Face dataset card for the Conversational Streaming ASR Benchmark, version 1.0.0. Latency is finalize-to-response P80 on a locked 50-call set streamed from a Delhi server. Semantic WER is pooled over about 67,000 reference words.
Read carefully: Soniox v5 edges Arth V2 on semantic WER by 0.08 points. Arth V2 is first on latency among the systems SquadStack timed. The marketing line that Arth V2 “beats Sarvam, Deepgram, Google and OpenAI” on accuracy holds on this table. The claim that it “matches Soniox” is fair for practical purposes; the published leaderboard still puts Soniox slightly ahead on the accuracy column.
Relative to Sarvam Saaras v4, Google Gemini 3.5 Transcribe and OpenAI gpt-live-transcribe, SquadStack’s blog cites about 18%, 26% and 28% fewer meaning-changing errors for Arth V2. Those percentages line up with the absolute rates above.
Why semantic WER matters on Indian phone calls
Raw WER on this set sits roughly between 27% and 55% across the eleven systems. Semantic WER compresses that to about 9% to 22%. SquadStack’s card says that for the median system, only about 32% of word errors would change what an agent does.
Two mistakes can look identical on a raw scorecard. Hearing an extra filler before “teen hazar” usually costs nothing. Hearing “chhatteese” when the customer said “chhabees” quotes the wrong EMI date. The benchmark tags meaning-changing errors into classes such as number, negation, name, commitment and content.
Noise moves scores more than code-mixing. On noisy calls, SquadStack reports Arth V2 at 10.5% semantic WER with the smallest climb from clean audio among the engines tested. On heavy Hindi-English mixing it reports 9.7%. Marketplace is the hardest vertical in the set; Education is the easiest, though Education has only 54 calls.
What is confirmed versus still open
| Claim | Status |
|---|---|
| Arth V2 public announcement dated 11 Oct 2026 | Confirmed on SquadStack blog |
| 863-call benchmark on Hugging Face with per-call error profiles | Confirmed dataset release |
| 8.85% semantic WER and 67 ms FTR P80 for Arth V2 | Vendor-reported on that benchmark |
| Live on ~80% of SquadStack daily agent calls | Company claim |
| Standalone public STT API with USD pricing | Not published in launch materials |
| Independent third-party replication of the full 863-call run | Not yet available |
The Hugging Face release is evaluation-only. Training on the audio is forbidden under SquadStack’s Benchmark Evaluation Licence. That is a serious constraint if you hoped to fine-tune your own model on the set, but it is useful if you only want to score an API you already pay for.
What it means for Indian developers
If you ship voice agents for BFSI, marketplace, logistics, travel or education in India, this is one of the few public scorecards built on 8 kHz telephony Hinglish instead of clean read speech. You can load the dataset, run your own engine through the same normalisation rules SquadStack documents, and compare against the published error profiles.
If you buy STT as a commodity API in USD, Arth V2 is not a drop-in SKU yet. It ships as part of SquadStack’s agent stack. Deepgram Nova 3, Sarvam Saaras v4, Google and OpenAI remain the practical buy options with published cloud endpoints. For a related look at how audio and multimodal models are moving on-device, see our note on EmbeddingGemma 2.
Builders already watching India’s speech stack should also keep Sarvam’s India AI event coverage and broader ChatGPT Voice / GPT Live updates in view. Creative voice tooling is moving fast too; ElevenLabs’ ad jingle contest shows the other end of the audio market.
How to treat the latency claim
Finalize-to-response is not the same as time-to-first-speech from the customer’s last phoneme. SquadStack’s method adds a fixed voice-activity wait when converting FTR to TTFS. Latency was measured from one Delhi server on a 50-call set at five-way concurrency. Re-run from your own network before you pick a vendor on milliseconds alone.
OpenAI’s gpt-live-transcribe lands at 830.8 ms FTR P80 on that same method, far behind the sub-100 ms group. That gap matters for barge-in and natural turn taking even if accuracy is acceptable for some scripts.
Bottom line
Arth V2 is a strong vendor release for Indian telephony speech, backed by an unusually transparent evaluation set. Soniox still leads the published semantic WER column by a hair. Arth V2 leads on latency in SquadStack’s timed set. Until a standalone USD API appears, most outside teams will use the benchmark to pressure-test the engines they already buy, not to swap in Arth itself.
FAQ
When did SquadStack announce Arth V2?
The company blog post is dated 11 October 2026 and is written by CEO Apurv Agrawal.
Is Arth V2 better than Soniox on accuracy?
On SquadStack’s published semantic WER table, Soniox v5 scores 8.77% and Arth V2 scores 8.85%. Treat them as effectively tied for many call scripts, with Soniox slightly ahead on that metric.
Can I buy Arth V2 as a standalone STT API in USD?
Launch materials do not list a public STT SKU or USD token price. Arth V2 is described as running inside SquadStack’s Humanoid Voice AI Agents.
Where is the benchmark data?
On Hugging Face under Squadstack/conversational-streaming-asr-benchmark. It includes audio, references and per-system error profiles for 863 calls. Training on the audio is not allowed.
Does this prove Arth V2 wins on every Indian language?
No. The set is Hindi-English code-mixed telesales audio. Other scheduled Indian languages and non-sales domains are out of scope for this release.