FA-Bench Leaderboard for ASR with Word Timestamps

Research

FA-Bench Leaderboard for ASR with Word Timestamps

We ranked 28 speech recognition pipelines, 19 built from open models and 9 that call a commercial API, by how closely their word timestamps land on the boundaries human linguists placed.

FA-Bench Leaderboard for ASR with Word Timestamps

github.com/olewave/fa-bench ยท Paper on arXiv

Today we're publishing the FA-Bench leaderboard for ASR with word timestamps.

A subtitle has to know when each word starts and ends. Most speech recognition APIs return word times, and few of them say how accurate those times are. FA-Bench now ranks 28 pipelines on exactly that. Each one hears only the audio and returns its own words, with a start and an end time for each.

How the score works

Each bar is the word boundary F1 at 20 ms. A boundary is a hit when the words on both sides of it match the reference and its time falls within 20 ms of where a human annotator placed it. A misheard word costs a system the boundaries beside it, so the score charges recognition errors and timing errors together.

The default view averages four test cells. They are TIMIT core test and Buckeye test, each on clean audio and on the mean of four degraded versions with reverb, noise, music or babble added. Choose a single cell from the menu to rank the systems by one corpus or one condition.

Choose Open models to compare only the pipelines you can run on your own hardware. A two-step pipeline counts as commercial when either step calls a paid API.

Each bar carries the colour and logo of the organization behind it. For a two-step pipeline that is the maker of the aligner, the step that sets the times. Olign 0.9, the version in the FA-Bench paper, needs a transcript, so here it aligns the words Qwen3 recognizes, as most two-step pipelines do. The Olign 1.0 API also takes audio alone: it recognizes the words itself, then aligns them.

What we found

  • Recognize first, then align. The four best pipelines all pass an ASR transcript to a forced aligner. Qwen3 followed by Olign 0.9 scores 0.645, the highest of the 28. Its numbers are exactly the ones the FA-Bench paper reports for Olign; numbers for Olign 1.0 are coming soon.
  • The best open pipeline beats every one-step API. Qwen3 followed by MFA 3.4 scores 0.585. The best one-step commercial API, Google Chirp 2, scores 0.436.
  • A low word error rate says little about timing. ElevenLabs Scribe v2 has the lowest word error rate of any one-step system on TIMIT, 2.0%, and ranks 26th of 28 here.

Reproduce every number

The numbers come from FA-Bench's published records, snapshot 202609, scored with release v1.1.0. The chart keeps one pipeline per aligner, and the records hold all 32. Download the CSV under the chart, or clone the repository and rescore every system yourself.