Best AI Transcription Tools 2026: Price Per Hour, Diarisation and Languages Compared
Seven speech-to-text options compared on verified per-hour price, whether speaker diarisation costs extra, published language coverage, and real-time support.
Start here
If software is going to read the transcript, buy an API and pay cents per hour: Deepgram when you need accurate text with speaker labels at volume, because diarisation is included rather than billed, AssemblyAI when you also want entities, redaction or summarisation from the same call, and ElevenLabs Scribe when the recording has many voices or when non-speech sound matters. If a person is going to read, correct and share the transcript, buy a product: Otter.ai for live meetings and Descript for anything you will edit afterwards. If a lawyer, clinician or regulator has to rely on it, buy Rev human transcription at $1.99 a minute. If the audio cannot leave your infrastructure, self-host Whisper.
The gap between those two answers is now about two orders of magnitude, and that is the thing to understand before comparing features. A machine hour costs roughly $0.21 to $0.27 across Deepgram, AssemblyAI, ElevenLabs and OpenAI. A human hour from Rev costs $119.40. Seat-based products sit in between and are priced on caps rather than usage, which is why Otter Pro is 1,200 recording minutes a month and not unlimited.
Three things changed since this page was last revised, and each of them invalidated a claim it used to make. Whisper is no longer a single model with no speaker support: OpenAI now publishes a transcription family that includes gpt-4o-transcribe-diarize. Diarisation is no longer uniformly included: AssemblyAI prices it as a stacking add-on while Deepgram includes it for pre-recorded audio. And the accuracy percentages that circulate for these tools are mostly not vendor claims at all, which is why this version quotes only figures the vendor publishes and says where they publish nothing.
This page compares transcription engines and transcription workflows. For tools whose job is summarising and actioning a meeting rather than transcribing it, see the meeting intelligence comparison and the AI note-taking and meeting assistants comparison.
Quick Comparison
| Tool | Type | Best For | Price (checked 18 Sep 2026) | Speaker diarisation | Real-time |
|---|---|---|---|---|---|
| OpenAI transcription models | API and self-hosted model | Cheapest credible per-minute API, or fully offline use | gpt-transcribe $0.0045/min; Whisper and gpt-4o-transcribe $0.006/min; gpt-live-transcribe $0.017/min; open Whisper weights free to self-host | Yes, via the separate gpt-4o-transcribe-diarize model | Yes, via gpt-live-transcribe |
| Deepgram | Developer API | Lowest cost at volume with diarisation included | Nova-3 pre-recorded from $0.0043/min pay-as-you-go; streaming from $0.0048/min | Included at no extra charge for pre-recorded audio | Yes (Nova-3 and Flux streaming) |
| AssemblyAI | Developer API | Transcription plus speech understanding in one call | Universal-2 $0.15/hr; Universal-3.5 Pro $0.21/hr async; streaming from $0.15/hr | Paid add-on: +$0.02/hr async standard, +$0.12/hr streaming | Yes, billed on session duration not audio duration |
| ElevenLabs Scribe | API and app | Diarising many speakers and tagging non-speech audio | Scribe v2 $0.22/hr; Scribe v2 Realtime $0.39/hr | Included, up to 32 speakers | Yes (Scribe v2 Realtime) |
| Otter.ai | End-user meeting product | Live meetings with shareable, attributed notes | Free (300 min/mo) / Pro $16.99/mo, or $8.33/mo billed annually / Business $30/mo, or $19.99 annually | Yes, on every plan | Yes |
| Descript | Editing product | Editing audio and video by editing the transcript | Free / Hobbyist $24/mo, $16 annually / Creator $35/mo, $24 annually / Business $65/mo, $50 annually | Yes, speaker detection on all plans | No |
| Rev | Human plus AI service | Accuracy someone has to sign off on | Human transcription $1.99/min; AI plans Free (45 min/mo), Essentials from $25.49 per seat/mo, Pro from $47.99 | Yes | Yes (AI) |
OpenAI transcription models
- Type
- API and self-hosted model
- Best For
- Cheapest credible per-minute API, or fully offline use
- Price (checked 18 Sep 2026)
- gpt-transcribe $0.0045/min; Whisper and gpt-4o-transcribe $0.006/min; gpt-live-transcribe $0.017/min; open Whisper weights free to self-host
- Speaker diarisation
- Yes, via the separate gpt-4o-transcribe-diarize model
- Real-time
- Yes, via gpt-live-transcribe
Deepgram
- Type
- Developer API
- Best For
- Lowest cost at volume with diarisation included
- Price (checked 18 Sep 2026)
- Nova-3 pre-recorded from $0.0043/min pay-as-you-go; streaming from $0.0048/min
- Speaker diarisation
- Included at no extra charge for pre-recorded audio
- Real-time
- Yes (Nova-3 and Flux streaming)
AssemblyAI
- Type
- Developer API
- Best For
- Transcription plus speech understanding in one call
- Price (checked 18 Sep 2026)
- Universal-2 $0.15/hr; Universal-3.5 Pro $0.21/hr async; streaming from $0.15/hr
- Speaker diarisation
- Paid add-on: +$0.02/hr async standard, +$0.12/hr streaming
- Real-time
- Yes, billed on session duration not audio duration
ElevenLabs Scribe
- Type
- API and app
- Best For
- Diarising many speakers and tagging non-speech audio
- Price (checked 18 Sep 2026)
- Scribe v2 $0.22/hr; Scribe v2 Realtime $0.39/hr
- Speaker diarisation
- Included, up to 32 speakers
- Real-time
- Yes (Scribe v2 Realtime)
Otter.ai
- Type
- End-user meeting product
- Best For
- Live meetings with shareable, attributed notes
- Price (checked 18 Sep 2026)
- Free (300 min/mo) / Pro $16.99/mo, or $8.33/mo billed annually / Business $30/mo, or $19.99 annually
- Speaker diarisation
- Yes, on every plan
- Real-time
- Yes
Descript
- Type
- Editing product
- Best For
- Editing audio and video by editing the transcript
- Price (checked 18 Sep 2026)
- Free / Hobbyist $24/mo, $16 annually / Creator $35/mo, $24 annually / Business $65/mo, $50 annually
- Speaker diarisation
- Yes, speaker detection on all plans
- Real-time
- No
Rev
- Type
- Human plus AI service
- Best For
- Accuracy someone has to sign off on
- Price (checked 18 Sep 2026)
- Human transcription $1.99/min; AI plans Free (45 min/mo), Essentials from $25.49 per seat/mo, Pro from $47.99
- Speaker diarisation
- Yes
- Real-time
- Yes (AI)
OpenAI transcription models
Best OverallBest for: Cheapest credible per-minute API, and the only route to fully offline transcription
“Whisper is no longer the whole story. OpenAI now publishes a family of transcription models at different price points, and the important change for this comparison is that diarisation finally exists as a first-party option rather than a community add-on. The open Whisper weights still run locally for free, which remains the only answer when audio cannot leave your infrastructure.”
Pros
- gpt-transcribe at $0.0045 per minute is the cheapest first-party option in this comparison, and the classic Whisper endpoint is still $0.006 per minute
- gpt-4o-transcribe-diarize gives speaker attribution as a supported model rather than a bolted-on pipeline, which was the single biggest gap in the old Whisper story
- The open Whisper model still runs entirely on your own hardware, which keeps legal, medical and financial audio off third-party servers
Cons
- Self-hosting still needs a capable GPU and Python competence, which most teams cannot resource without engineering time
- There is no end-user product, no upload interface, no editing workflow and no collaboration layer, so this is a component and not a tool
What changed since Whisper
The single model has become a lineup. Alongside the original Whisper endpoint, OpenAI's pricing page now lists gpt-transcribe, gpt-4o-transcribe and gpt-4o-mini-transcribe for batch work. It adds gpt-4o-transcribe-diarize for speaker attribution, plus gpt-live-transcribe and gpt-realtime-whisper for streaming at $0.017 per minute. The newer models are billed on tokens with a per-minute equivalent shown for guidance, which matters when you are forecasting a large workload.
Self-hosting, and when it is worth it
The open Whisper weights are still the reason this entry ranks first for anyone with a compliance constraint. Running large-v3 locally needs roughly 10GB of VRAM and processes audio several times faster than real time on a modern GPU, and smaller variants trade accuracy for lower hardware requirements. For legal depositions, clinical consultations and recorded financial calls, self-hosting means no audio leaves your tenant, and no data processing agreement has to cover it. Pair it with a diarisation library, because the open model does not attribute speakers on its own.
The ecosystem effect
Whisper's release raised the accuracy floor across the whole market, and several products in this comparison were built on it or on models trained in its wake. That is why the differentiation between commercial transcription tools in 2026 is almost never raw word accuracy. It is diarisation quality, streaming latency, language coverage, and whether the billing unit matches how you actually use the service.
gpt-transcribe $0.0045/min; Whisper and gpt-4o-transcribe $0.006/min; gpt-4o-mini-transcribe $0.003/min; gpt-4o-transcribe-diarize $0.006/min; gpt-live-transcribe $0.017/min; open Whisper weights free to self-host
Deepgram
Best ValueBest for: High-volume transcription where diarisation should not be a line item
“Deepgram is the cheapest credible per-minute API in this comparison and the only one that includes speaker diarisation at no extra charge for pre-recorded audio. For a product processing thousands of hours a month, that combination decides the build. Nova-3 is the current flagship, with the Flux models covering lower-latency streaming.”
Pros
- Nova-3 pre-recorded starts at $0.0043 per minute pay-as-you-go, which is roughly $0.26 an hour, and drops further on the Growth plan
- Speaker diarisation is included at no extra charge for pre-recorded audio, where AssemblyAI charges for it as an add-on
- Separate monolingual and multilingual price points let you avoid paying the multilingual premium on English-only workloads
Cons
- The headline rates shown are promotional current prices sitting below a stated regular price, so model the regular rate when you budget
- Deepgram's accuracy claims are relative comparisons against unnamed competitors rather than an absolute word error rate you can independently reproduce
Why diarisation being free matters
Speaker attribution is not a nice-to-have in most commercial speech workloads. Call analytics, meeting products, compliance recording and media subtitling all need to know who spoke. AssemblyAI charges for it, at $0.02 an hour for standard async diarisation and $0.12 an hour on streaming. Deepgram includes it for pre-recorded audio. On a workload of several thousand hours a month that difference compounds into a meaningful number, and it is the clearest reason to shortlist Deepgram for a build.
Model choice and language scope
Nova-3 ships in monolingual and multilingual variants at different prices, and Deepgram's model documentation lists coverage across English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, Dutch, Arabic and Chinese among others, with regional variants. Check the exact language list against your traffic before committing, because the multilingual premium is only worth paying if you actually receive that traffic.
Nova-3 monolingual: $0.0043/min pre-recorded and $0.0048/min streaming on pay-as-you-go; Nova-3 multilingual $0.0052/min pre-recorded, $0.0058/min streaming; Flux English streaming $0.0065/min; Growth plan rates are lower
AssemblyAI
Runner UpBest for: Getting transcription plus structure, entities and summarisation from one integration
“AssemblyAI is still the most feature-complete speech API for developers who want more than text back. The model line is now Universal-2 and Universal-3.5 Pro, and the pricing is unusually legible: a base rate per model, then a published price for every add-on. That transparency is a genuine advantage, and it also makes it obvious how quickly a realistic configuration costs more than the headline.”
Pros
- Universal-2 at $0.15 an hour is competitive, and the per-add-on price list means you can cost a configuration exactly before writing code
- Speech understanding features sit behind the same API: entity detection at $0.08 an hour, PII redaction at $0.08, keyterm prompting at $0.05, medical mode at $0.15
- SDK and documentation quality are ahead of the large cloud speech APIs, which shortens the path from evaluation to production
Cons
- Add-ons stack additively on the base rate, so a diarised, redacted, entity-tagged transcript costs materially more than the advertised model price
- There is no end-user product at all; this is an API for people building something else
Beyond a transcript
The reason to pick AssemblyAI over a cheaper engine is what comes back with the text. Entity detection, PII redaction, keyterm prompting and medical mode are all first-party, priced per hour and callable in the same request. For call-centre analytics, compliance monitoring or any pipeline that would otherwise chain a speech API to a separate language model, that consolidation removes a whole integration and its failure modes.
Costing a realistic configuration
Work the arithmetic before you compare vendors. Universal-3.5 Pro async at $0.21 an hour with standard diarisation at $0.02 and PII redaction at $0.08 is $0.31 an hour, which is materially above Deepgram's pre-recorded Nova-3 rate with diarisation already included. AssemblyAI earns that premium when you are buying the speech understanding features too, and does not when you only need text and speaker labels.
Streaming billing, stated plainly
AssemblyAI bills streaming on session duration, not audio duration, and says so on its pricing page. That is a design decision rather than a trap, but it changes how you architect a client. Close sockets aggressively on silence, or price the idle time into your unit economics.
Async: Universal-2 $0.15/hr, Universal-3.5 Pro $0.21/hr. Streaming: Universal-Streaming English or Multilingual $0.15/hr, Universal-3.5 Pro Realtime $0.45/hr. Add-ons stack: diarisation +$0.02/hr async or +$0.12/hr streaming, entity detection +$0.08/hr, PII redaction +$0.08/hr, medical mode +$0.15/hr
ElevenLabs Scribe
Best for EnterpriseBest for: Multi-speaker audio and recordings where non-speech sound matters
“Scribe is the strongest new entrant on this page since it last ran. Scribe v2 handles diarisation for up to 32 speakers, produces word-level timestamps, and tags non-speech audio events such as laughter and applause, all included in a flat $0.22 per hour. For panel discussions, courtroom-style recordings and documentary or podcast material, that feature set is closer to the job than a bare transcript.”
Pros
- Diarisation for up to 32 speakers is included in the base rate, which is well beyond what most engines attempt
- Audio event tagging captures laughter, applause and similar non-speech sound, which matters for subtitling and for media production
- ElevenLabs publishes per-language accuracy bands rather than a single marketing accuracy figure, which is a more honest disclosure than most of this market offers
Cons
- Real-time costs a clear premium at $0.39 an hour against $0.22 for batch, so latency is a deliberate budget decision
- The included-hours table is plan-bound, so paying monthly for a plan you underuse is worse value than a pure pay-as-you-go API
Diarisation depth
Most engines degrade badly past five or six speakers. Scribe v2 documents support for up to 32, which is the difference between a usable transcript of a board meeting or a conference panel and a wall of unattributed text. Combined with precise word-level timestamps, it is also the most practical option here for generating subtitle files that have to carry speaker labels.
Language coverage, read properly
Scribe v2 supports more than 90 languages, and ElevenLabs groups them into published accuracy bands rather than claiming one number across all of them. English is in the top band at 5% word error rate or better. Languages in the lower bands can exceed 25%. Use the band your audio falls into as the planning number. A vendor that publishes the distribution instead of the best case is giving you something useful, so take it at face value and test accordingly.
Medical variant
Scribe v2 Medical is priced identically to the general model at $0.22 an hour. If your workload is clinical, that removes the usual premium other vendors attach to domain vocabulary. AssemblyAI, by contrast, prices medical mode as a $0.15 per hour add-on on top of the base model rate.
Scribe v2 and Scribe v2 Medical $0.22/hr; Scribe v2 Realtime $0.39/hr. Included hours by plan: Starter $6/mo (4.5 hrs), Creator $22/mo (27 hrs), Pro $99/mo (100 hrs), Scale $299/mo (450 hrs), Business $990/mo (1,359 hrs). Entity detection +$0.070/hr, keyterm prompting +$0.050/hr
Otter.ai
Honorable MentionBest for: Live meetings that need attributed, shareable notes without any engineering
“Otter is still the most polished ready-to-use meeting product here, and it is the right answer when the buyer is an operations lead rather than a developer. What has changed is the shape of the plans. The free tier and the Pro tier are both minute-capped in ways that catch people out, and only Business removes the cap on meetings.”
Pros
- Joins Zoom, Google Meet and Microsoft Teams automatically and produces speaker-attributed live captions plus a shareable transcript, with no setup work
- Speaker identification and AI summaries are available on every plan including the free one, which is unusual in this market
- Annual billing more than halves the Pro price, from $16.99 a month to $8.33 a month per user
Cons
- Pro is capped at 1,200 in-app recording minutes a month with a 90 minute per-meeting limit, which is not unlimited and is easy to hit
- File imports are tightly rationed: three lifetime imports on Free and ten a month on Pro, so bulk transcription of existing recordings is a Business-tier activity
Live meeting integration
Otter joins calls as a participant, records, and produces a real-time transcript with live captions. Participants get a shareable link afterwards. Speaker identification works by matching voice patterns to participants and improves for recurring attendees. Accuracy degrades with crosstalk, poor microphones and heavily accented speech in the same way every engine on this page does, so the meeting hygiene matters more than the vendor choice.
Collaboration and search
Transcripts behave like documents. Team members highlight passages, comment and tag action items on the transcript itself, and search indexes every transcript so a decision can be found by keyword months later. For organisations with a recurring who-agreed-to-what problem, that searchable record is the actual product, and the transcription is the mechanism.
Where the caps bite
The two limits that generate complaints are the per-meeting length ceiling and the file import allowance. A 90 minute cap on Pro cuts off long workshops and all-hands calls. A ten-per-month import allowance rules out digesting an archive. Both are removed on Business, and neither is obvious from the plan headline, so check them against your real meeting pattern before buying seats.
Free (300 min/mo, 30 min per conversation, 3 lifetime file imports) / Pro $16.99/mo per user, or $8.33/mo billed annually (1,200 recording min/mo, 90 min per meeting, 10 imports/mo) / Business $30/mo per user, or $19.99 annually (unlimited meetings, 4 hr limit, unlimited imports) / Enterprise custom with SSO, SCIM and Domain Capture
Descript
Best Free OptionBest for: Editing recorded audio and video by editing the transcript
“Descript is still the clearest example of a good idea shipped well: transcribe the recording, edit the text, and the media edits itself to match. It is an editing product with excellent transcription attached, not a transcription service. The 2026 change that matters is that Descript now meters AI separately from media hours, so both limits apply at once.”
Pros
- Deleting a sentence of text removes the corresponding audio and video, which turns tightening a recording into a text-editing task
- Filler word removal and silence trimming run across a whole recording in one action, which is the feature that actually saves hours per episode
- Speaker detection is available on every plan including Free, and filler word removal, Studio Sound and voice cloning are all on the entry paid tier
Cons
- Every plan is bounded twice, once by media hours and once by monthly AI credits, and the credit cost per action is not published as a table
- No real-time transcription at all, so live captioning and meeting attendance are out of scope
Transcript-based editing
The editor presents a recording as a document. Selecting and deleting text removes the matching media. Reordering paragraphs reorders the timeline. Removing a tangent, resequencing a conversation or cutting a mistake becomes the same action as editing a doc. For anyone who has spent evenings trimming silences in a waveform editor, that is the whole pitch and it holds up.
Filler words, silence and Studio Sound
Filler word detection finds verbal fillers across the transcript and removes them in one pass, with the option to review each instance. Silence removal shortens long pauses on the same basis. Studio Sound applies enhancement that reduces background noise and evens out levels, which is the practical fix for remote recordings where participants had mismatched microphones. On interview-style content these three features are where the time saving actually comes from.
Two meters, not one
Descript bounds every plan by media hours and by monthly AI credits at the same time, and does not publish a per-action credit table. That means a plan can run out of either resource first depending on how heavily you lean on the AI features. Watch which meter you exhaust in the first month and size the plan from that, rather than from the media hours alone.
Free (60 min media/mo, 100 one-time AI credits) / Hobbyist $24/mo, $16 billed annually (10 media hrs, 400 credits/mo) / Creator $35/mo, $24 annually (30 media hrs, 800 credits/mo) / Business $65/mo, $50 annually (40 media hrs, 1,500 credits/mo) / Enterprise custom
Rev
Best for EnterpriseBest for: Transcripts a person will stand behind
“Rev remains the only service here that puts a human in the loop and attaches an accuracy claim to it. Its own site states 99% or better accuracy on human transcription delivered in 12 hours or less, and 99.7% or better on the legal premium transcript. Both the price and the structure have moved since this page last ran: human transcription is now $1.99 a minute, and the AI side is sold as seat-based subscriptions rather than per-minute.”
Pros
- Human transcription carries a stated 99%+ accuracy claim with a 12 hour or faster turnaround, which is a commitment rather than a benchmark
- Legal-specific products are priced per page rather than per minute, at $2.25 for a rough draft and $2.50 for a premium transcript with a 99.7%+ claim
- Subscription plans include large AI transcription allowances, at 5,000 minutes per user per month on Essentials and 10,000 on Pro
Cons
- At $1.99 a minute, an hour of human transcription is $119.40, which is several hundred times the per-hour cost of the APIs on this page
- Global subtitles are priced steeply by language, from $6.49 a minute for Spanish or Hindi up to $15.99 a minute for Japanese or Korean
When a human is the requirement
Some contexts do not accept a probabilistic transcript. Legal transcription needs verbatim capture including false starts and crosstalk. Clinical records need correct terminology and correct attribution. Regulated recordings need a defensible record. In those settings the gap between a machine transcript and a reviewed one is not a quality preference, it is whether the artefact is usable at all, and $1.99 a minute is cheap against the cost of an error.
Reading the subscription against the a la carte price
Rev now sells AI transcription through seats rather than per minute, which makes cross-vendor comparison harder. Essentials includes 5,000 AI minutes per user per month for $25.49 per seat monthly. That is about 83 hours, which works out around $0.31 an hour if you use all of it, and considerably worse if you do not. The plans also carry discounts on human work, 10% on Essentials and 15% on Pro, which is the main reason to hold a seat at all.
Subtitles and localisation
Global subtitles are the line item that surprises buyers. Rev prices them per minute by target language, from $6.49 for Spanish and Hindi, through $10.49 for a group including French, German, Arabic and Portuguese, to $15.99 for Japanese and Korean. For a regular multilingual publishing schedule, price that against a machine translation pipeline with human review before committing.
Human transcription $1.99/min (99%+ accuracy, 12 hours or less); legal rough draft $2.25/page; legal premium transcript $2.50/page (99.7%+); global subtitles $6.49 to $15.99/min by language. AI plans: Free 45 min/mo; Essentials $25.49 per seat/mo (5,000 AI min per user); Pro $47.99 per seat/mo (10,000 AI min per user); Unlimited custom
Which One Should You Pick?
| Use Case | Our Recommendation |
|---|---|
| Transcribing team meetings with shared, searchable notes | Otter.ai, and check the caps first. Free covers 300 minutes a month with a 30 minute limit per conversation. Pro at $8.33 a month billed annually covers 1,200 recording minutes with a 90 minute per-meeting limit. Only Business at $19.99 a month billed annually removes the meeting cap and the file import limit. |
| Editing a podcast or video without a timeline editor | Descript. Transcribe, edit the text, and the media follows. Filler word removal and silence trimming alone justify the subscription for a regular show. Size the plan by media hours and AI credits together, because Creator's 30 media hours and 800 credits are separate ceilings that can run out independently. |
| Processing confidential audio that cannot leave your infrastructure | Self-hosted Whisper, still. The open weights run entirely on your own hardware with a capable GPU, so no audio reaches a third party and no data processing agreement is needed. Pair it with a diarisation library, or use OpenAI's gpt-4o-transcribe-diarize if the API is acceptable for your compliance posture. |
| Transcript that a lawyer, clinician or regulator will rely on | Rev human transcription at $1.99 a minute, with a stated 99%+ accuracy and a 12 hour or faster turnaround. For litigation work specifically, the legal premium transcript is priced per page at $2.50 with a 99.7%+ claim, which is usually the right product rather than the per-minute service. |
| Building speech-to-text into a product at volume | Deepgram if you mainly need accurate text with speaker labels, because Nova-3 pre-recorded starts at $0.0043 a minute and includes diarisation at no extra charge. AssemblyAI if you also need entities, redaction or summarisation in the same call, accepting that add-ons stack on the base rate. |
| Transcribing a panel, board meeting or documentary with many voices | ElevenLabs Scribe v2 at $0.22 an hour. Diarisation for up to 32 speakers is included, word-level timestamps come as standard, and non-speech audio events such as laughter and applause are tagged, which matters when the recording has to become subtitles. |
| Generating subtitle files for a video library on a budget | OpenAI gpt-transcribe at $0.0045 a minute or Deepgram Nova-3 at $0.0043 a minute are the cheapest credible routes to timestamped text. Choose ElevenLabs Scribe instead when the subtitles need speaker labels, and Rev's per-language subtitle service only when a human has to certify the translation. |
| Live captioning or real-time voice applications | Compare the billing unit, not just the rate. Deepgram Nova-3 streaming starts at $0.0048 a minute. ElevenLabs Scribe v2 Realtime is $0.39 an hour. AssemblyAI streaming starts at $0.15 an hour but is billed on session duration rather than audio duration, so idle sockets cost money. OpenAI's gpt-live-transcribe is $0.017 a minute. |
How we evaluated
Speech-to-text splits into three products that are routinely compared as if they were one: engines sold per hour to developers, seat-based workflow tools sold to teams, and human services sold per minute. This comparison keeps them in one table because buyers shortlist across all three, and labels the type of each so the prices can be read against each other honestly.
Each option was assessed on the dimensions in the comparison table above:
- Type: API, end-user product, or human service. This is the first filter, and it decides more than any feature does.
- Verified price: the rate published on the vendor's own pricing page on the date below, with the billing unit stated. Per-minute, per-hour, per-seat and per-page rates are not interchangeable and are not silently converted here.
- Speaker diarisation: whether it is included, priced as an add-on, or absent. This is the most common cause of a speech budget coming in over estimate.
- Real-time support: whether streaming exists, at what price, and on what billing unit. AssemblyAI bills streaming on session duration rather than audio duration and says so, which changes client design.
- Accuracy and language claims: only where the vendor publishes them, quoted as the vendor states them, with the nature of the claim made explicit. A relative word error rate reduction against unnamed competitors is not the same kind of statement as a published per-language error band, and neither is the same as a contractual accuracy commitment on human work.
What we reviewed
Every price, limit and capability on this page came from the vendor's own pricing page or documentation, read on the verification date below:
- OpenAI API pricing for the transcription model family, including the diarisation and live-transcription models
- Deepgram pricing and Deepgram's models and languages documentation
- AssemblyAI pricing, including the published add-on rates and the streaming billing note
- ElevenLabs API pricing and the Scribe speech-to-text documentation for speaker limits, timestamps, audio tagging and the per-language accuracy bands
- Otter.ai pricing for minute caps, per-meeting limits and file import allowances
- Descript pricing for media hours and monthly AI credits per plan
- Rev pricing for human per-minute rates, per-page legal products, subtitle rates by language and the current AI subscription tiers
Where a vendor does not publish something, this page says so rather than filling the gap from elsewhere. Descript does not publish a per-action AI credit table. Deepgram's accuracy claims do not name the competitor set or corpus. Those absences are reported because they are part of the buying decision.
This is a research comparison, not a hands-on test. No accuracy figure here comes from private benchmarking, and every claim traces to a vendor document you can open from the links above.
Last verified: 18 September 2026. This pass removed accuracy percentages that were not attributable to any vendor, corrected Otter's free and Pro tiers to the published minute caps and import limits, replaced Descript's single price with the current four-tier structure and its dual media-hour and AI-credit meters, corrected Rev human transcription from $1.50 to $1.99 a minute and added its new seat-based AI plans, corrected AssemblyAI from a single hourly rate to the current model rates with diarisation priced as an add-on, expanded the Whisper entry to the current OpenAI transcription family including the diarisation model, and added Deepgram and ElevenLabs Scribe as entrants.
Editorial independence: this is a vendor-neutral comparison with no paid placements, sponsorships, or affiliate links. Rankings reflect fit for the stated use cases, not commercial relationships.
Frequently Asked Questions
What does AI transcription actually cost per hour in 2026?
Does speaker diarisation cost extra?
How accurate are these tools, really?
Which languages do these actually handle well?
Is it safe to upload confidential audio to a transcription service?
Should I buy a transcription API or a transcription product?
What audio quality do I need for a usable transcript?
Related Comparisons
GPU Cloud and AI Compute
Top 7 GPU Cloud and AI Compute Providers 2026: Price Per GPU-Hour by Chip Class, Verified From Each Provider's Own Pricing Page
7 tools compared
LLM Evaluation and Prompt Management
Top 7 LLM Evaluation and Prompt Management Platforms 2026: How to Stop Shipping Prompt Changes Blind
7 tools compared
Fine-Tuning and Model Customization
Top 8 Fine-Tuning and Model Customization Platforms 2026: What It Costs, When It Wins, and Why OpenAI Is Shutting Its Own Down
8 tools compared
AI Legal / Contract
Top 5 AI Legal and Contract Tools 2026: Harvey vs Spellbook vs Ironclad vs LegalOn vs Luminance
5 tools compared