Updated on
Gladia: reviews and analysis
Audio transcription API billed per hour, with free credits to get started.
Affiliate link · no extra cost to you
Our verdict
Gladia is not an app with buttons, it is the plumbing: a transcription API billed per hour of audio, with free credits covering dozens of test hours. If you transcribe at volume and can wire it up, it costs far less than subscription tools. If you want an interface to drop a file into, look elsewhere.
Best for: Technical teams transcribing high volume who want to pay for actual usage.
What changed this week
Sweep of
No substantive change from the previous sweep: what we checked this week confirms what was already there.
- down
- no change
We have tracked it since July 25, 2026: 3 sweeps on record, with the score moving between 4.0 and 4.2. We publish the latest change here; the full series is not published.
What the internet says
It remains a solid player in multilingual transcription with per-hour pricing and no fixed fee, but this week brings no fresh structural news: the OVH Groupe integration, closed 31 July, moves ahead with no public detail on its future independence. What does move is the gap to the independent benchmark leaders, which has widened, while the official SDK keeps piling up issues with no maintainer response.
What the web repeats in favour
- Coverage of nearly a hundred languages and solid behaviour when a speaker switches language mid-sentence, which is where most fail
- Pay per processed hour with no monthly fee, from 0.61 dollars an hour for pre-recorded audio, with a 50-euro welcome credit that never resets monthly, reconfirmed this week on the official site
- Speaker separation and audio analysis come with the base rate, not as extras that inflate the bill
- The ecosystem adopted it on its own: there is a dedicated provider in Vercel's AI SDK, a dedicated service in Pipecat and support in LiveKit, all maintained by third parties rather than the vendor
- It shipped GladiaFlow, an open-source (MIT-licensed) desktop dictation app that adds voice-to-text across any system application with custom vocabulary support, released 15 July 2026
What the web repeats against
- There is no user interface: anyone who wants to upload audio and download text without writing code has to look elsewhere
- The steep per-hour discount requires an upfront volume commitment, so the good price is out of reach for small usage
- The lead spot the vendor claims for its Solaria-3 model does not hold up on the independent benchmark: at a 3.2% error rate it now trails at least six rival models, a wider gap than last week's reading
- The official SDK still has open issues with no maintainer reply, including a new one from 29 July about Python response objects serializing to empty, and the live session lifecycle still causes trouble when wired into third-party frameworks
Sweep sources: Official pricing · Press · Review sites · GitHub · GitHub · Docs
Latest Gladia news
OVH Groupe completes its acquisition of Gladia
The deal announced on 11 June is now closed and is paid in new OVH shares rather than cash, so the release puts no price on the purchase. It also says nothing about the fate of Gladia’s API or the 2,000 enterprise customers it claims, which is exactly what anyone running it in production needs to know. Source
OVHcloud enters exclusive talks to buy Gladia
The deal is not closed and neither the amount nor a closing date has been published. Anyone integrating Gladia's API today should factor in that ownership and the price list can change once it completes. Source
Pros / Cons
Pros
- Pay per hour processed, no monthly fee
- Generous free welcome credits
- Live transcription on top of recorded files
Cons
- It is an API: you need someone to integrate it
- No user interface for manual work
TLDR: Gladia is a transcription API, not an app with buttons. It bills per hour of audio processed instead of a monthly fee, gives welcome credits covering dozens of hours, and offers both file and live transcription. If you transcribe at volume and have someone to integrate it, it costs far less than any subscription. If you want to upload a file and read the text, it is not your tool.
What Gladia is and how it works
Gladia sells infrastructure, not interface. Its product is a speech-to-text API you call from your application, your automation or your internal process, returning a transcript with speaker separation, timestamps and language detection.
That product decision defines who it serves. Consumer transcription tools charge a monthly fee and give you a website to drag files into. Gladia gives the opposite: no website to work in, and a per-hour price that above a certain volume simply does not compare with subscriptions. It is the difference between buying a finished product and buying the engine.
It covers two modes: asynchronous transcription (upload audio, receive text) and live transcription, built for applications needing subtitles or notes while someone speaks. Paid plans also include the compliance certifications companies require when audio contains personal data.
What it is like day to day
For the right profile, day to day means you never see it. Gladia lives inside a process: the meeting recording that transcribes itself when it ends, the video that generates its subtitles on upload, the internal tool that turns calls into searchable text. Well integrated, the tool disappears and only the result remains.
The real work is the initial integration: wiring the API, deciding where the text lands and what happens to it. That is an afternoon or a day for someone who codes, not a marketing task. In exchange, the process is built once and the cost scales with actual usage instead of with the calendar.
The other practical advantage is predictable spend. With a subscription you pay for quiet months the same as busy ones. With per-hour billing, a month without recordings costs zero, and a conference month with twenty sessions costs exactly what those twenty sessions run.
Pricing and plans
The model is pay-as-you-go: you prepay credits and each hour of processed audio draws them down, at different rates for file and live transcription. New users receive welcome credits covering dozens of hours, so a full evaluation costs nothing.
The rate drops with volume on higher plans, and that is the economic argument against subscription tools: for anyone transcribing hundreds of hours a year, the cost difference is an order of magnitude. For anyone transcribing five hours a month, the convenience of a tool with an interface is worth more than the saving.
Who it is for (and who it is not for)
Gladia is for technical teams with volume: products that need transcription as a feature, media outlets processing audio archives, companies systematically turning meetings or calls into text. Also for anyone wanting live transcription inside their own application, a case consumer tools do not cover.
It is not for the individual creator who wants to subtitle their videos: Descript transcribes and edits in the same window there, and Submagic applies styled subtitles without writing a line of code. Nor is it for anyone who needs the result today with nobody available to integrate an API.
Gladia alternatives
The alternatives are not other APIs but another approach: finished tools where transcription is a means rather than the product. Descript transcribes so you can edit by text, and ElevenLabs plays the opposite side of audio, generating voice rather than reading it. The rest of the category is in the best AI audio tools ranking.
Frequently Asked Questions
Do I need to code to use Gladia?
Yes, or have someone who does. It is an API with no user interface: the value comes from wiring it into a process of your own. If you want to drag a file and read the result, a tool with an interface serves you better.
How much does transcribing one hour of audio cost?
You pay per hour processed, at different rates for file versus live transcription, and the rate drops as you move up plans. The free welcome credits cover dozens of hours, enough to measure your real cost before committing.
How accurate is it?
Current speech-to-text models perform well on standard speech and lose precision with strong accents, overlapping voices and very specific jargon. Speaker separation helps considerably in interviews and meetings.
Can I use it to subtitle videos?
As an engine, yes: it returns text with timestamps, and your process builds the subtitles from there. The visual side of subtitling (style, animation, burning into the video) needs another tool.
Guides that use Gladia
More tools in this category
Synthesizer V Studio 2 Pro
One-time-purchase offline vocal synthesis with note editing, a DAW plugin and…
4.7 · ExcellentACE Studio
A lyrics-and-MIDI singing studio with over 160 voices, instruments and custom voice…
4.7 · ExcellentMoises
Stem separation, chords, tempo and music practice across web, mobile and desktop.