Skip to content
AI Tools · AI Transcription

Best AI Transcription Tools 2026: Price Per Hour, Diarisation and Languages Compared

Seven speech-to-text options compared on verified per-hour price, whether speaker diarisation costs extra, published language coverage, and real-time support.

By ·Apr 1, 2026·Updated Sep 18, 2026·17 min·7 tools compared
AI TranscriptionWhisperSpeech-to-TextDeepgramAssemblyAISubtitles

Start here

If software is going to read the transcript, buy an API and pay cents per hour: Deepgram when you need accurate text with speaker labels at volume, because diarisation is included rather than billed, AssemblyAI when you also want entities, redaction or summarisation from the same call, and ElevenLabs Scribe when the recording has many voices or when non-speech sound matters. If a person is going to read, correct and share the transcript, buy a product: Otter.ai for live meetings and Descript for anything you will edit afterwards. If a lawyer, clinician or regulator has to rely on it, buy Rev human transcription at $1.99 a minute. If the audio cannot leave your infrastructure, self-host Whisper.

The gap between those two answers is now about two orders of magnitude, and that is the thing to understand before comparing features. A machine hour costs roughly $0.21 to $0.27 across Deepgram, AssemblyAI, ElevenLabs and OpenAI. A human hour from Rev costs $119.40. Seat-based products sit in between and are priced on caps rather than usage, which is why Otter Pro is 1,200 recording minutes a month and not unlimited.

Three things changed since this page was last revised, and each of them invalidated a claim it used to make. Whisper is no longer a single model with no speaker support: OpenAI now publishes a transcription family that includes gpt-4o-transcribe-diarize. Diarisation is no longer uniformly included: AssemblyAI prices it as a stacking add-on while Deepgram includes it for pre-recorded audio. And the accuracy percentages that circulate for these tools are mostly not vendor claims at all, which is why this version quotes only figures the vendor publishes and says where they publish nothing.

This page compares transcription engines and transcription workflows. For tools whose job is summarising and actioning a meeting rather than transcribing it, see the meeting intelligence comparison and the AI note-taking and meeting assistants comparison.

Quick Comparison

ToolTypeBest ForPrice (checked 18 Sep 2026)Speaker diarisationReal-time
OpenAI transcription modelsAPI and self-hosted modelCheapest credible per-minute API, or fully offline usegpt-transcribe $0.0045/min; Whisper and gpt-4o-transcribe $0.006/min; gpt-live-transcribe $0.017/min; open Whisper weights free to self-hostYes, via the separate gpt-4o-transcribe-diarize modelYes, via gpt-live-transcribe
DeepgramDeveloper APILowest cost at volume with diarisation includedNova-3 pre-recorded from $0.0043/min pay-as-you-go; streaming from $0.0048/minIncluded at no extra charge for pre-recorded audioYes (Nova-3 and Flux streaming)
AssemblyAIDeveloper APITranscription plus speech understanding in one callUniversal-2 $0.15/hr; Universal-3.5 Pro $0.21/hr async; streaming from $0.15/hrPaid add-on: +$0.02/hr async standard, +$0.12/hr streamingYes, billed on session duration not audio duration
ElevenLabs ScribeAPI and appDiarising many speakers and tagging non-speech audioScribe v2 $0.22/hr; Scribe v2 Realtime $0.39/hrIncluded, up to 32 speakersYes (Scribe v2 Realtime)
Otter.aiEnd-user meeting productLive meetings with shareable, attributed notesFree (300 min/mo) / Pro $16.99/mo, or $8.33/mo billed annually / Business $30/mo, or $19.99 annuallyYes, on every planYes
DescriptEditing productEditing audio and video by editing the transcriptFree / Hobbyist $24/mo, $16 annually / Creator $35/mo, $24 annually / Business $65/mo, $50 annuallyYes, speaker detection on all plansNo
RevHuman plus AI serviceAccuracy someone has to sign off onHuman transcription $1.99/min; AI plans Free (45 min/mo), Essentials from $25.49 per seat/mo, Pro from $47.99YesYes (AI)

OpenAI transcription models

Type
API and self-hosted model
Best For
Cheapest credible per-minute API, or fully offline use
Price (checked 18 Sep 2026)
gpt-transcribe $0.0045/min; Whisper and gpt-4o-transcribe $0.006/min; gpt-live-transcribe $0.017/min; open Whisper weights free to self-host
Speaker diarisation
Yes, via the separate gpt-4o-transcribe-diarize model
Real-time
Yes, via gpt-live-transcribe

Deepgram

Type
Developer API
Best For
Lowest cost at volume with diarisation included
Price (checked 18 Sep 2026)
Nova-3 pre-recorded from $0.0043/min pay-as-you-go; streaming from $0.0048/min
Speaker diarisation
Included at no extra charge for pre-recorded audio
Real-time
Yes (Nova-3 and Flux streaming)

AssemblyAI

Type
Developer API
Best For
Transcription plus speech understanding in one call
Price (checked 18 Sep 2026)
Universal-2 $0.15/hr; Universal-3.5 Pro $0.21/hr async; streaming from $0.15/hr
Speaker diarisation
Paid add-on: +$0.02/hr async standard, +$0.12/hr streaming
Real-time
Yes, billed on session duration not audio duration

ElevenLabs Scribe

Type
API and app
Best For
Diarising many speakers and tagging non-speech audio
Price (checked 18 Sep 2026)
Scribe v2 $0.22/hr; Scribe v2 Realtime $0.39/hr
Speaker diarisation
Included, up to 32 speakers
Real-time
Yes (Scribe v2 Realtime)

Otter.ai

Type
End-user meeting product
Best For
Live meetings with shareable, attributed notes
Price (checked 18 Sep 2026)
Free (300 min/mo) / Pro $16.99/mo, or $8.33/mo billed annually / Business $30/mo, or $19.99 annually
Speaker diarisation
Yes, on every plan
Real-time
Yes

Descript

Type
Editing product
Best For
Editing audio and video by editing the transcript
Price (checked 18 Sep 2026)
Free / Hobbyist $24/mo, $16 annually / Creator $35/mo, $24 annually / Business $65/mo, $50 annually
Speaker diarisation
Yes, speaker detection on all plans
Real-time
No

Rev

Type
Human plus AI service
Best For
Accuracy someone has to sign off on
Price (checked 18 Sep 2026)
Human transcription $1.99/min; AI plans Free (45 min/mo), Essentials from $25.49 per seat/mo, Pro from $47.99
Speaker diarisation
Yes
Real-time
Yes (AI)
1

OpenAI transcription models

Best Overall

Best for: Cheapest credible per-minute API, and the only route to fully offline transcription

“Whisper is no longer the whole story. OpenAI now publishes a family of transcription models at different price points, and the important change for this comparison is that diarisation finally exists as a first-party option rather than a community add-on. The open Whisper weights still run locally for free, which remains the only answer when audio cannot leave your infrastructure.”

Pros

  • gpt-transcribe at $0.0045 per minute is the cheapest first-party option in this comparison, and the classic Whisper endpoint is still $0.006 per minute
  • gpt-4o-transcribe-diarize gives speaker attribution as a supported model rather than a bolted-on pipeline, which was the single biggest gap in the old Whisper story
  • The open Whisper model still runs entirely on your own hardware, which keeps legal, medical and financial audio off third-party servers

Cons

  • Self-hosting still needs a capable GPU and Python competence, which most teams cannot resource without engineering time
  • There is no end-user product, no upload interface, no editing workflow and no collaboration layer, so this is a component and not a tool
Honest Weakness: This is a model family, not a transcription product. Everything past raw text with timestamps, including formatting, filler-word cleanup, editing and sharing, is your problem to solve. The published per-minute prices are also estimates rather than the billing unit for the newer models. gpt-4o-transcribe and gpt-4o-mini-transcribe are billed per million input and output tokens, with the per-minute figure quoted alongside for guidance. If you are comparing to a flat per-hour API like Deepgram or ElevenLabs, model your own audio rather than trusting the headline conversion.

What changed since Whisper

The single model has become a lineup. Alongside the original Whisper endpoint, OpenAI's pricing page now lists gpt-transcribe, gpt-4o-transcribe and gpt-4o-mini-transcribe for batch work. It adds gpt-4o-transcribe-diarize for speaker attribution, plus gpt-live-transcribe and gpt-realtime-whisper for streaming at $0.017 per minute. The newer models are billed on tokens with a per-minute equivalent shown for guidance, which matters when you are forecasting a large workload.

Self-hosting, and when it is worth it

The open Whisper weights are still the reason this entry ranks first for anyone with a compliance constraint. Running large-v3 locally needs roughly 10GB of VRAM and processes audio several times faster than real time on a modern GPU, and smaller variants trade accuracy for lower hardware requirements. For legal depositions, clinical consultations and recorded financial calls, self-hosting means no audio leaves your tenant, and no data processing agreement has to cover it. Pair it with a diarisation library, because the open model does not attribute speakers on its own.

The ecosystem effect

Whisper's release raised the accuracy floor across the whole market, and several products in this comparison were built on it or on models trained in its wake. That is why the differentiation between commercial transcription tools in 2026 is almost never raw word accuracy. It is diarisation quality, streaming latency, language coverage, and whether the billing unit matches how you actually use the service.

gpt-transcribe $0.0045/min; Whisper and gpt-4o-transcribe $0.006/min; gpt-4o-mini-transcribe $0.003/min; gpt-4o-transcribe-diarize $0.006/min; gpt-live-transcribe $0.017/min; open Whisper weights free to self-host

Visit OpenAI transcription models
2

Deepgram

Best Value

Best for: High-volume transcription where diarisation should not be a line item

“Deepgram is the cheapest credible per-minute API in this comparison and the only one that includes speaker diarisation at no extra charge for pre-recorded audio. For a product processing thousands of hours a month, that combination decides the build. Nova-3 is the current flagship, with the Flux models covering lower-latency streaming.”

Pros

  • Nova-3 pre-recorded starts at $0.0043 per minute pay-as-you-go, which is roughly $0.26 an hour, and drops further on the Growth plan
  • Speaker diarisation is included at no extra charge for pre-recorded audio, where AssemblyAI charges for it as an add-on
  • Separate monolingual and multilingual price points let you avoid paying the multilingual premium on English-only workloads

Cons

  • The headline rates shown are promotional current prices sitting below a stated regular price, so model the regular rate when you budget
  • Deepgram's accuracy claims are relative comparisons against unnamed competitors rather than an absolute word error rate you can independently reproduce
Honest Weakness: Deepgram's published accuracy claim is a marketing comparison, not a benchmark you can verify. The company states a 54.2% reduction in word error rate for streaming and 47.4% for batch against competitors, and reports preference ratios up to 8:1 in multilingual testing. Neither figure names the competitor set or the evaluation corpus. Treat those numbers as a reason to run your own sample, not as a specification. The pricing also carries promotional rates displayed against higher regular prices, which is fine until the promotion ends mid-contract.

Why diarisation being free matters

Speaker attribution is not a nice-to-have in most commercial speech workloads. Call analytics, meeting products, compliance recording and media subtitling all need to know who spoke. AssemblyAI charges for it, at $0.02 an hour for standard async diarisation and $0.12 an hour on streaming. Deepgram includes it for pre-recorded audio. On a workload of several thousand hours a month that difference compounds into a meaningful number, and it is the clearest reason to shortlist Deepgram for a build.

Model choice and language scope

Nova-3 ships in monolingual and multilingual variants at different prices, and Deepgram's model documentation lists coverage across English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, Dutch, Arabic and Chinese among others, with regional variants. Check the exact language list against your traffic before committing, because the multilingual premium is only worth paying if you actually receive that traffic.

Nova-3 monolingual: $0.0043/min pre-recorded and $0.0048/min streaming on pay-as-you-go; Nova-3 multilingual $0.0052/min pre-recorded, $0.0058/min streaming; Flux English streaming $0.0065/min; Growth plan rates are lower

Visit Deepgram
3

AssemblyAI

Runner Up

Best for: Getting transcription plus structure, entities and summarisation from one integration

“AssemblyAI is still the most feature-complete speech API for developers who want more than text back. The model line is now Universal-2 and Universal-3.5 Pro, and the pricing is unusually legible: a base rate per model, then a published price for every add-on. That transparency is a genuine advantage, and it also makes it obvious how quickly a realistic configuration costs more than the headline.”

Pros

  • Universal-2 at $0.15 an hour is competitive, and the per-add-on price list means you can cost a configuration exactly before writing code
  • Speech understanding features sit behind the same API: entity detection at $0.08 an hour, PII redaction at $0.08, keyterm prompting at $0.05, medical mode at $0.15
  • SDK and documentation quality are ahead of the large cloud speech APIs, which shortens the path from evaluation to production

Cons

  • Add-ons stack additively on the base rate, so a diarised, redacted, entity-tagged transcript costs materially more than the advertised model price
  • There is no end-user product at all; this is an API for people building something else
Honest Weakness: Two billing details will surprise you if you skim the pricing page. First, diarisation is an add-on rather than a default, at $0.02 an hour on standard async, $0.065 on the experimental async option and $0.12 on streaming, where Deepgram simply includes it for pre-recorded audio. Second, and more consequential, streaming is billed on session duration rather than audio duration: AssemblyAI states that a WebSocket held open for 60 minutes carrying 30 minutes of audio is billed for 60 minutes. Any application that keeps a socket open across a quiet call will pay for the silence.

Beyond a transcript

The reason to pick AssemblyAI over a cheaper engine is what comes back with the text. Entity detection, PII redaction, keyterm prompting and medical mode are all first-party, priced per hour and callable in the same request. For call-centre analytics, compliance monitoring or any pipeline that would otherwise chain a speech API to a separate language model, that consolidation removes a whole integration and its failure modes.

Costing a realistic configuration

Work the arithmetic before you compare vendors. Universal-3.5 Pro async at $0.21 an hour with standard diarisation at $0.02 and PII redaction at $0.08 is $0.31 an hour, which is materially above Deepgram's pre-recorded Nova-3 rate with diarisation already included. AssemblyAI earns that premium when you are buying the speech understanding features too, and does not when you only need text and speaker labels.

Streaming billing, stated plainly

AssemblyAI bills streaming on session duration, not audio duration, and says so on its pricing page. That is a design decision rather than a trap, but it changes how you architect a client. Close sockets aggressively on silence, or price the idle time into your unit economics.

Async: Universal-2 $0.15/hr, Universal-3.5 Pro $0.21/hr. Streaming: Universal-Streaming English or Multilingual $0.15/hr, Universal-3.5 Pro Realtime $0.45/hr. Add-ons stack: diarisation +$0.02/hr async or +$0.12/hr streaming, entity detection +$0.08/hr, PII redaction +$0.08/hr, medical mode +$0.15/hr

Visit AssemblyAI
4

ElevenLabs Scribe

Best for Enterprise

Best for: Multi-speaker audio and recordings where non-speech sound matters

“Scribe is the strongest new entrant on this page since it last ran. Scribe v2 handles diarisation for up to 32 speakers, produces word-level timestamps, and tags non-speech audio events such as laughter and applause, all included in a flat $0.22 per hour. For panel discussions, courtroom-style recordings and documentary or podcast material, that feature set is closer to the job than a bare transcript.”

Pros

  • Diarisation for up to 32 speakers is included in the base rate, which is well beyond what most engines attempt
  • Audio event tagging captures laughter, applause and similar non-speech sound, which matters for subtitling and for media production
  • ElevenLabs publishes per-language accuracy bands rather than a single marketing accuracy figure, which is a more honest disclosure than most of this market offers

Cons

  • Real-time costs a clear premium at $0.39 an hour against $0.22 for batch, so latency is a deliberate budget decision
  • The included-hours table is plan-bound, so paying monthly for a plan you underuse is worse value than a pure pay-as-you-go API
Honest Weakness: The accuracy story is honest but uneven, and you need to read it per language rather than as a headline. ElevenLabs publishes word error rate bands across four tiers, from an excellent tier at 5% or better down to a moderate tier above 25%, across 90-plus supported languages. English sits in the excellent band. Many of the 90 do not. If your audio is not in a top-tier language, the supported-language count is close to meaningless and you should test your own material before committing. The plan structure also bundles Scribe hours into general ElevenLabs tiers, so the effective rate depends on whether you also use the voice products.

Diarisation depth

Most engines degrade badly past five or six speakers. Scribe v2 documents support for up to 32, which is the difference between a usable transcript of a board meeting or a conference panel and a wall of unattributed text. Combined with precise word-level timestamps, it is also the most practical option here for generating subtitle files that have to carry speaker labels.

Language coverage, read properly

Scribe v2 supports more than 90 languages, and ElevenLabs groups them into published accuracy bands rather than claiming one number across all of them. English is in the top band at 5% word error rate or better. Languages in the lower bands can exceed 25%. Use the band your audio falls into as the planning number. A vendor that publishes the distribution instead of the best case is giving you something useful, so take it at face value and test accordingly.

Medical variant

Scribe v2 Medical is priced identically to the general model at $0.22 an hour. If your workload is clinical, that removes the usual premium other vendors attach to domain vocabulary. AssemblyAI, by contrast, prices medical mode as a $0.15 per hour add-on on top of the base model rate.

Scribe v2 and Scribe v2 Medical $0.22/hr; Scribe v2 Realtime $0.39/hr. Included hours by plan: Starter $6/mo (4.5 hrs), Creator $22/mo (27 hrs), Pro $99/mo (100 hrs), Scale $299/mo (450 hrs), Business $990/mo (1,359 hrs). Entity detection +$0.070/hr, keyterm prompting +$0.050/hr

Visit ElevenLabs Scribe
5

Otter.ai

Honorable Mention

Best for: Live meetings that need attributed, shareable notes without any engineering

“Otter is still the most polished ready-to-use meeting product here, and it is the right answer when the buyer is an operations lead rather than a developer. What has changed is the shape of the plans. The free tier and the Pro tier are both minute-capped in ways that catch people out, and only Business removes the cap on meetings.”

Pros

  • Joins Zoom, Google Meet and Microsoft Teams automatically and produces speaker-attributed live captions plus a shareable transcript, with no setup work
  • Speaker identification and AI summaries are available on every plan including the free one, which is unusual in this market
  • Annual billing more than halves the Pro price, from $16.99 a month to $8.33 a month per user

Cons

  • Pro is capped at 1,200 in-app recording minutes a month with a 90 minute per-meeting limit, which is not unlimited and is easy to hit
  • File imports are tightly rationed: three lifetime imports on Free and ten a month on Pro, so bulk transcription of existing recordings is a Business-tier activity
Honest Weakness: The free tier is smaller than it appears and the paid tier is more capped than most buyers assume. Free is 300 transcription minutes a month with a 30 minute limit per conversation and only three lifetime audio or video file imports. Pro is 1,200 in-app recording minutes a month, up to 90 minutes per meeting, with ten file imports a month. Only Business gives unlimited meetings and in-app recordings, up to four hours each, with unlimited imports, at $30 a month per user or $19.99 billed annually. If the plan is to feed Otter a back catalogue of recordings, price Business from the start.

Live meeting integration

Otter joins calls as a participant, records, and produces a real-time transcript with live captions. Participants get a shareable link afterwards. Speaker identification works by matching voice patterns to participants and improves for recurring attendees. Accuracy degrades with crosstalk, poor microphones and heavily accented speech in the same way every engine on this page does, so the meeting hygiene matters more than the vendor choice.

Collaboration and search

Transcripts behave like documents. Team members highlight passages, comment and tag action items on the transcript itself, and search indexes every transcript so a decision can be found by keyword months later. For organisations with a recurring who-agreed-to-what problem, that searchable record is the actual product, and the transcription is the mechanism.

Where the caps bite

The two limits that generate complaints are the per-meeting length ceiling and the file import allowance. A 90 minute cap on Pro cuts off long workshops and all-hands calls. A ten-per-month import allowance rules out digesting an archive. Both are removed on Business, and neither is obvious from the plan headline, so check them against your real meeting pattern before buying seats.

Free (300 min/mo, 30 min per conversation, 3 lifetime file imports) / Pro $16.99/mo per user, or $8.33/mo billed annually (1,200 recording min/mo, 90 min per meeting, 10 imports/mo) / Business $30/mo per user, or $19.99 annually (unlimited meetings, 4 hr limit, unlimited imports) / Enterprise custom with SSO, SCIM and Domain Capture

Visit Otter.ai
6

Descript

Best Free Option

Best for: Editing recorded audio and video by editing the transcript

“Descript is still the clearest example of a good idea shipped well: transcribe the recording, edit the text, and the media edits itself to match. It is an editing product with excellent transcription attached, not a transcription service. The 2026 change that matters is that Descript now meters AI separately from media hours, so both limits apply at once.”

Pros

  • Deleting a sentence of text removes the corresponding audio and video, which turns tightening a recording into a text-editing task
  • Filler word removal and silence trimming run across a whole recording in one action, which is the feature that actually saves hours per episode
  • Speaker detection is available on every plan including Free, and filler word removal, Studio Sound and voice cloning are all on the entry paid tier

Cons

  • Every plan is bounded twice, once by media hours and once by monthly AI credits, and the credit cost per action is not published as a table
  • No real-time transcription at all, so live captioning and meeting attendance are out of scope
Honest Weakness: You are paying for an editor, and the quotas reflect that. Free is 60 minutes of media a month with 100 one-time AI credits. Hobbyist at $24 a month, or $16 billed annually, is 10 media hours and 400 AI credits a month. Creator at $35 a month, or $24 annually, is 30 media hours and 800 credits. Business at $65 a month, or $50 annually, is 40 media hours and 1,500 credits. If you only want text out of an hour of audio, you are buying an editing suite you will not open, and a per-hour API on this page will do the same job for a few cents.

Transcript-based editing

The editor presents a recording as a document. Selecting and deleting text removes the matching media. Reordering paragraphs reorders the timeline. Removing a tangent, resequencing a conversation or cutting a mistake becomes the same action as editing a doc. For anyone who has spent evenings trimming silences in a waveform editor, that is the whole pitch and it holds up.

Filler words, silence and Studio Sound

Filler word detection finds verbal fillers across the transcript and removes them in one pass, with the option to review each instance. Silence removal shortens long pauses on the same basis. Studio Sound applies enhancement that reduces background noise and evens out levels, which is the practical fix for remote recordings where participants had mismatched microphones. On interview-style content these three features are where the time saving actually comes from.

Two meters, not one

Descript bounds every plan by media hours and by monthly AI credits at the same time, and does not publish a per-action credit table. That means a plan can run out of either resource first depending on how heavily you lean on the AI features. Watch which meter you exhaust in the first month and size the plan from that, rather than from the media hours alone.

Free (60 min media/mo, 100 one-time AI credits) / Hobbyist $24/mo, $16 billed annually (10 media hrs, 400 credits/mo) / Creator $35/mo, $24 annually (30 media hrs, 800 credits/mo) / Business $65/mo, $50 annually (40 media hrs, 1,500 credits/mo) / Enterprise custom

Visit Descript
7

Rev

Best for Enterprise

Best for: Transcripts a person will stand behind

“Rev remains the only service here that puts a human in the loop and attaches an accuracy claim to it. Its own site states 99% or better accuracy on human transcription delivered in 12 hours or less, and 99.7% or better on the legal premium transcript. Both the price and the structure have moved since this page last ran: human transcription is now $1.99 a minute, and the AI side is sold as seat-based subscriptions rather than per-minute.”

Pros

  • Human transcription carries a stated 99%+ accuracy claim with a 12 hour or faster turnaround, which is a commitment rather than a benchmark
  • Legal-specific products are priced per page rather than per minute, at $2.25 for a rough draft and $2.50 for a premium transcript with a 99.7%+ claim
  • Subscription plans include large AI transcription allowances, at 5,000 minutes per user per month on Essentials and 10,000 on Pro

Cons

  • At $1.99 a minute, an hour of human transcription is $119.40, which is several hundred times the per-hour cost of the APIs on this page
  • Global subtitles are priced steeply by language, from $6.49 a minute for Spanish or Hindi up to $15.99 a minute for Japanese or Korean
Honest Weakness: Rev's AI transcription is unremarkable in a market where the engines above it cost cents per hour, so you are not buying Rev for its model. You are buying the human review and the defensibility that comes with it, and that costs $1.99 a minute. The subscription plans muddy the comparison. Essentials at $25.49 per seat per month, or $29.99 per seat on the annual option, bundles 5,000 AI minutes and a 10% discount on human work. That is good value only if you consume both. For occasional human transcription, buy it a la carte and skip the seat.

When a human is the requirement

Some contexts do not accept a probabilistic transcript. Legal transcription needs verbatim capture including false starts and crosstalk. Clinical records need correct terminology and correct attribution. Regulated recordings need a defensible record. In those settings the gap between a machine transcript and a reviewed one is not a quality preference, it is whether the artefact is usable at all, and $1.99 a minute is cheap against the cost of an error.

Reading the subscription against the a la carte price

Rev now sells AI transcription through seats rather than per minute, which makes cross-vendor comparison harder. Essentials includes 5,000 AI minutes per user per month for $25.49 per seat monthly. That is about 83 hours, which works out around $0.31 an hour if you use all of it, and considerably worse if you do not. The plans also carry discounts on human work, 10% on Essentials and 15% on Pro, which is the main reason to hold a seat at all.

Subtitles and localisation

Global subtitles are the line item that surprises buyers. Rev prices them per minute by target language, from $6.49 for Spanish and Hindi, through $10.49 for a group including French, German, Arabic and Portuguese, to $15.99 for Japanese and Korean. For a regular multilingual publishing schedule, price that against a machine translation pipeline with human review before committing.

Human transcription $1.99/min (99%+ accuracy, 12 hours or less); legal rough draft $2.25/page; legal premium transcript $2.50/page (99.7%+); global subtitles $6.49 to $15.99/min by language. AI plans: Free 45 min/mo; Essentials $25.49 per seat/mo (5,000 AI min per user); Pro $47.99 per seat/mo (10,000 AI min per user); Unlimited custom

Visit Rev

Which One Should You Pick?

Use CaseOur Recommendation
Transcribing team meetings with shared, searchable notesOtter.ai, and check the caps first. Free covers 300 minutes a month with a 30 minute limit per conversation. Pro at $8.33 a month billed annually covers 1,200 recording minutes with a 90 minute per-meeting limit. Only Business at $19.99 a month billed annually removes the meeting cap and the file import limit.
Editing a podcast or video without a timeline editorDescript. Transcribe, edit the text, and the media follows. Filler word removal and silence trimming alone justify the subscription for a regular show. Size the plan by media hours and AI credits together, because Creator's 30 media hours and 800 credits are separate ceilings that can run out independently.
Processing confidential audio that cannot leave your infrastructureSelf-hosted Whisper, still. The open weights run entirely on your own hardware with a capable GPU, so no audio reaches a third party and no data processing agreement is needed. Pair it with a diarisation library, or use OpenAI's gpt-4o-transcribe-diarize if the API is acceptable for your compliance posture.
Transcript that a lawyer, clinician or regulator will rely onRev human transcription at $1.99 a minute, with a stated 99%+ accuracy and a 12 hour or faster turnaround. For litigation work specifically, the legal premium transcript is priced per page at $2.50 with a 99.7%+ claim, which is usually the right product rather than the per-minute service.
Building speech-to-text into a product at volumeDeepgram if you mainly need accurate text with speaker labels, because Nova-3 pre-recorded starts at $0.0043 a minute and includes diarisation at no extra charge. AssemblyAI if you also need entities, redaction or summarisation in the same call, accepting that add-ons stack on the base rate.
Transcribing a panel, board meeting or documentary with many voicesElevenLabs Scribe v2 at $0.22 an hour. Diarisation for up to 32 speakers is included, word-level timestamps come as standard, and non-speech audio events such as laughter and applause are tagged, which matters when the recording has to become subtitles.
Generating subtitle files for a video library on a budgetOpenAI gpt-transcribe at $0.0045 a minute or Deepgram Nova-3 at $0.0043 a minute are the cheapest credible routes to timestamped text. Choose ElevenLabs Scribe instead when the subtitles need speaker labels, and Rev's per-language subtitle service only when a human has to certify the translation.
Live captioning or real-time voice applicationsCompare the billing unit, not just the rate. Deepgram Nova-3 streaming starts at $0.0048 a minute. ElevenLabs Scribe v2 Realtime is $0.39 an hour. AssemblyAI streaming starts at $0.15 an hour but is billed on session duration rather than audio duration, so idle sockets cost money. OpenAI's gpt-live-transcribe is $0.017 a minute.

How we evaluated

Speech-to-text splits into three products that are routinely compared as if they were one: engines sold per hour to developers, seat-based workflow tools sold to teams, and human services sold per minute. This comparison keeps them in one table because buyers shortlist across all three, and labels the type of each so the prices can be read against each other honestly.

Each option was assessed on the dimensions in the comparison table above:

  • Type: API, end-user product, or human service. This is the first filter, and it decides more than any feature does.
  • Verified price: the rate published on the vendor's own pricing page on the date below, with the billing unit stated. Per-minute, per-hour, per-seat and per-page rates are not interchangeable and are not silently converted here.
  • Speaker diarisation: whether it is included, priced as an add-on, or absent. This is the most common cause of a speech budget coming in over estimate.
  • Real-time support: whether streaming exists, at what price, and on what billing unit. AssemblyAI bills streaming on session duration rather than audio duration and says so, which changes client design.
  • Accuracy and language claims: only where the vendor publishes them, quoted as the vendor states them, with the nature of the claim made explicit. A relative word error rate reduction against unnamed competitors is not the same kind of statement as a published per-language error band, and neither is the same as a contractual accuracy commitment on human work.

What we reviewed

Every price, limit and capability on this page came from the vendor's own pricing page or documentation, read on the verification date below:

Where a vendor does not publish something, this page says so rather than filling the gap from elsewhere. Descript does not publish a per-action AI credit table. Deepgram's accuracy claims do not name the competitor set or corpus. Those absences are reported because they are part of the buying decision.

This is a research comparison, not a hands-on test. No accuracy figure here comes from private benchmarking, and every claim traces to a vendor document you can open from the links above.

Last verified: 18 September 2026. This pass removed accuracy percentages that were not attributable to any vendor, corrected Otter's free and Pro tiers to the published minute caps and import limits, replaced Descript's single price with the current four-tier structure and its dual media-hour and AI-credit meters, corrected Rev human transcription from $1.50 to $1.99 a minute and added its new seat-based AI plans, corrected AssemblyAI from a single hourly rate to the current model rates with diarisation priced as an add-on, expanded the Whisper entry to the current OpenAI transcription family including the diarisation model, and added Deepgram and ElevenLabs Scribe as entrants.

Note

Editorial independence: this is a vendor-neutral comparison with no paid placements, sponsorships, or affiliate links. Rankings reflect fit for the stated use cases, not commercial relationships.

Frequently Asked Questions

What does AI transcription actually cost per hour in 2026?
Far less than most buyers expect, if you use an API. Deepgram Nova-3 pre-recorded starts at $0.0043 a minute pay-as-you-go, which is about $0.26 an hour. AssemblyAI Universal-2 is $0.15 an hour and Universal-3.5 Pro is $0.21. ElevenLabs Scribe v2 is $0.22 an hour. OpenAI's gpt-transcribe is $0.0045 a minute, about $0.27 an hour. Human transcription from Rev is $1.99 a minute, which is $119.40 an hour. End-user products such as Otter and Descript are sold as seats with monthly minute or media-hour caps instead, so their effective hourly cost depends entirely on how much of the allowance you use.
Does speaker diarisation cost extra?
It depends on the vendor, and this is the most common budgeting mistake in speech projects. Deepgram includes speaker diarisation at no extra charge for pre-recorded audio. ElevenLabs Scribe v2 includes it for up to 32 speakers in the base $0.22 an hour rate. AssemblyAI charges for it as an add-on that stacks on the model rate: $0.02 an hour for standard async, $0.065 for the experimental async option and $0.12 for streaming. OpenAI provides it through a separate model, gpt-4o-transcribe-diarize, at $0.006 a minute. The open Whisper weights do not diarise at all and need a separate library.
How accurate are these tools, really?
Be careful with accuracy percentages, including the ones vendors publish. ElevenLabs is the most honest disclosure in this comparison: it publishes word error rate bands across its 90-plus supported languages, with English in the top band at 5% or better and some languages above 25%. Deepgram publishes relative claims, a 54.2% reduction in word error rate for streaming and 47.4% for batch against unnamed competitors. Rev attaches an accuracy commitment to human work rather than to a model, at 99%+ for standard human transcription and 99.7%+ for the legal premium transcript. Nobody publishes a single number that transfers to your audio, because audio quality, accent, crosstalk and domain vocabulary dominate the result. Run a sample of your own material.
Which languages do these actually handle well?
Read the distribution rather than the count. ElevenLabs Scribe v2 documents more than 90 languages grouped into accuracy bands, which is the only coverage disclosure here that tells you where the weak spots are. Deepgram sells Nova-3 in separate monolingual and multilingual variants at different prices, with published language lists including regional variants, so you can avoid paying a multilingual premium on English-only traffic. The open Whisper model is multilingual with performance varying widely by language. Meeting products such as Otter are strongest in English and should be checked against your actual meeting languages before a rollout.
Is it safe to upload confidential audio to a transcription service?
Treat every cloud service as a processor that needs a contract. Otter, Descript, Rev, AssemblyAI, Deepgram and ElevenLabs all process audio on their infrastructure, so read the retention and training terms and get an agreement in place before sending anything sensitive. For legal, clinical, financial or HR audio, the defensible option is still self-hosted Whisper on your own hardware, where the audio never leaves your tenant. Where an enterprise plan offers custom retention terms and a compliance attestation, that is the minimum bar for regulated material, not an upgrade.
Should I buy a transcription API or a transcription product?
Ask who is going to use the output. If a person needs to read, correct, share and search the transcript, buy a product: Otter for meetings, Descript for anything you will edit. If software consumes the output, buy an API, because you will pay cents per hour instead of tens of dollars per seat. The expensive mistake is buying a seat-based product to do bulk batch transcription, which is what the file-import limits on those products exist to discourage.
What audio quality do I need for a usable transcript?
Audio quality still dominates model choice. Use a dedicated microphone rather than a laptop microphone, record somewhere quiet, and in meetings prefer a central conference microphone over several open laptop mics. Overlapping speech is the single hardest condition for every engine on this page, including the ones that diarise well. If you are stuck with a poor recording, run enhancement first, for example Descript's Studio Sound, before transcribing. No vendor on this list can recover words that were never captured cleanly.

About the author

is the founder and creator of LoginRadius, a customer identity platform he built and scaled to over a billion users. He is now the founder of GrackerAI, a GEO platform for B2B SaaS and cybersecurity teams, and has spent more than 15 years building identity and security products.

Related Comparisons