AI Services · Speech to Text

Free Speech-to-Text API Credits: Give Your Companion Ears (ASR/STT)

Talking to your companion starts with real-time speech recognition (ASR — also called STT). Unlike picking a chat model, the first thing that matters here isn't price, it's language: if a provider can't transcribe the language you actually speak, the biggest credit pile in the world is worthless. Here's the whole decision — language, latency, cost, free credits, and referral bonuses — in one place.

Three things to weigh before picking a speech-to-text provider

1. Language support — the one that actually disqualifies providers

Confirm the provider supports your language in real-time streaming mode, and pay special attention to multilingual use: if you speak one language, lock it in; if you mix English and another language mid-sentence — or switch between English, Japanese, or Korean — you want a model with automatic language detection and mid-conversation switching, or you'll be changing settings every time. Also note the difference between file transcription and streaming: some providers support more languages for recorded files than for live streams, and desktop chat uses the live stream.

2. Latency

Companion chat feels best when the transcript lands as you finish speaking: the closer the servers, the faster the first words appear and the more natural barge-in feels. Providers with nodes near you win; overseas providers depend on how stable your route to them is.

3. Cost

The same three-tier logic as chat models: sign-up credits (AssemblyAI $50, Deepgram $200), standing monthly free hours (Speechmatics gives 8 hours every month, Qwen's Paraformer about 10), and the price after credits — streaming recognition runs a few dollars per hour at most, so desktop-companion usage barely dents a wallet.

Free credits and language support at a glance

Compiled August 2026, covering every STT provider AniMate supports. Bonus terms shift, so treat the sign-up page as current.

Provider Streaming languages Free tier Card required?
AssemblyAI Chinese, English, Spanish, French, German, Italian, Portuguese $50 credit (hundreds of hours) No
Deepgram Nova-3 Chinese (Simplified/Traditional/Cantonese), English, Japanese, Korean, Spanish, French, German, and 36+ more $200 credit (~750 hours) No
Speechmatics Chinese (Simplified/Traditional), English, Japanese, Korean, and 55+ more 8 free hours every month + $100 credit No
Qwen AI (Alibaba) Mandarin, Cantonese, Sichuanese and other Chinese dialects, English (Chinese-English mixing supported) ~10 free hours monthly on Paraformer, refreshed monthly ID verification, no card
ElevenLabs Scribe v2 Chinese, English, Japanese, and 90+ more (auto language detection, mid-conversation switching) ~30 minutes/month on the free plan No

Language lists follow each provider's documentation and grow with model updates; for mixed-language chat, leave automatic detection on in AniMate's language scope setting.

Provider-by-provider reviews

AssemblyAI: our first pick — easiest sign-up

Free tier: New accounts get $50 in credits — hundreds of hours of streaming at going rates, which is years of desktop-companion chat. AniMate's default preset works out of the box.
Sign-up experience: Email, create a key, paste it into AniMate — no card anywhere in the flow, and the smoothest onboarding of the overseas providers.
Heads-up: Use the Universal streaming model; set your language scope in AniMate's language settings, and keep auto-detection on for mixed Chinese-English chat.

🎁 Invite bonus: sign up through our link

Register for AssemblyAI through our invite link — both sides earn bonus rewards, and the standard $50 credit still applies.

Sign up through the invite link: $50 credit + referral bonus →

Deepgram: the biggest credit pile, strongest multilingual coverage

Free tier: $200 in credits (~750 hours of streaming), no card required — the most generous one-time bonus of the bunch.
Sign-up experience: Email registration, create a key, go.
Heads-up: For Chinese, use the Nova-3 model family — Deepgram's Flux conversational model doesn't cover Chinese. Simplified, Traditional, and Cantonese are all supported, along with Japanese and Korean. There's no friend-referral program; the $200 signup credit is the whole deal.

Speechmatics: 8 free hours, every month

Free tier: The Free plan includes 480 minutes (8 hours) monthly, officially no card required, plus a $100 credit; 55+ languages with Chinese, Japanese, and Korean covered.
Sign-up experience: Register and create an API key in the console.
Heads-up: Monthly hours don't roll over; heavy users can pair it with another provider. No referral program either.

Qwen AI (Alibaba): the Chinese-dialect specialist

Free tier: Paraformer real-time models come with roughly 36,000 seconds (10 hours) of free time refreshed on the first of every month — the only provider that tops you up monthly.
Sign-up experience: Create a key at platform.qianwenai.com; domestic nodes mean the lowest latency for Chinese speech.
Heads-up: Mandarin, Cantonese, Sichuanese, and more dialects are supported — if your companion should understand dialect, this is the one. Enable the spend cap; free-tier terms are confirmed in the console.

ElevenLabs: low latency, small free tier

Free tier: 10,000 credits monthly, and real-time STT burns 330 credits per minute — about 30 minutes. Fine for a taste, not a workhorse.
Sign-up experience: Registration only, no card; Scribe v2 Realtime's end-to-end latency is its standout trait.
Heads-up: If you also use ElevenLabs for TTS, one account covers both directions of the conversation.

Plug the key back into AniMate

  • Open Settings → AI Services in AniMate and click "Add ASR" under speech services.
  • Pick the provider and paste your API key.
  • Choose the recognition model and language scope: lock a single language, or keep auto-detection for mixed speech.
  • Save, say a sentence into your mic, confirm the transcript, and set it as default — your companion can hear you now.

FAQ

How much transcription time does companion chat actually use?

A few minutes a day for most people. At that pace AssemblyAI's $50 or Deepgram's $200 lasts over a year, and Speechmatics' monthly 8 hours or Qwen's monthly 10 are nearly impossible to exhaust.

Why does a provider claim Chinese support but fail in practice?

Because desktop chat uses real-time streaming, and some providers' file-transcription models support more languages than their streaming models. Check the "streaming" language list before signing up — AniMate's ASR presets are already matched to streaming capabilities.

Do I need extra setup for interruptions?

No — interruptions (barge-in) are an AniMate setting that activates with real-time audio output. The STT provider only transcribes.

What happens when the free credits run out?

Claim another provider's credits or top up — at a few dollars per hour and a few minutes of daily use, the ongoing cost is negligible.

Ears done — next, give her a voice

Free credits for the chat model and text-to-speech are in the other two guides; all three together unlock full voice conversation.

Download AniMate

Keep reading