FA-Bench: A Benchmark for Evaluating Phone- and Word-Level Timestamp Accuracy in Forced Aligners

Research

FA-Bench: A Benchmark for Evaluating Phone- and Word-Level Timestamp Accuracy in Forced Aligners

Most forced aligner results are hard to reproduce. FA-Bench is an open, standardized baseline. We scored 30 systems on phone- and word-level timestamp accuracy against boundaries that human linguists placed, on clean audio and on four kinds of degradation.

FA-Bench: A Benchmark for Evaluating Phone- and Word-Level Timestamp Accuracy in Forced Aligners

github.com/olewave/fa-bench · Paper on arXiv

We're releasing FA-Bench, an open benchmark for the accuracy of phone- and word-level timestamps.

Most forced aligner results are hard to reproduce. Papers use different phoneme sets, their own test splits and their own evaluation metrics, so a number in one paper rarely measures the same thing as a number in the next. That leaves a speech researcher unable to run a clean ablation study, and a voice AI engineer unable to tell which system to deploy.

FA-Bench publishes the whole evaluation as code, from how labels are folded to how each boundary is scored. Every system runs through the same pipeline, so its numbers compare directly with every other system's.

What we tested

We scored 30 systems, 21 open models and 9 commercial APIs. Each one runs on clean audio and on four degraded versions of it, with reverb, noise, music or babble added. Every results table puts the clean score and the four degraded scores on one row, so you can see how much accuracy a system keeps as the audio gets harder.

Scoring uses read speech from TIMIT and spontaneous speech from Buckeye. Both corpora ship phone and word boundaries that human linguists corrected by hand, and those boundaries are the reference. FA-Bench adds no annotation of its own. It keeps the boundary times exactly as the corpora give them and folds the phone labels into the shared TIMIT-39 set.

Results are reported in two tracks. In the gold-transcript track, each system is handed the reference words, so its score measures timing alone. In the ASR track, each system recognizes the words itself, so recognition errors add to its timing errors. The two tracks are ranked separately.

What we found

Olign 1.0 has the highest word boundary F1 of every system given the reference transcript, on both corpora, clean and with noise added.

Word boundary F1 at 20 ms, gold-transcript track. Higher is better.

System TIMIT core-test Buckeye test
clean / noisy clean / noisy
Olign 1.0 0.782 / 0.669 0.744 / 0.659
MFA 3.4 0.644 / 0.624 0.682 / 0.583
Qwen3-ForcedAligner-0.6B 0.399 / 0.377 0.432 / 0.391
WhisperX 0.180 / 0.170 0.165 / 0.165

Noisy is the mean of the four degraded conditions. With noise added, Olign's lead over Montreal Forced Aligner 3.4 narrows on TIMIT, from 0.138 to 0.045, and widens on Buckeye, from 0.062 to 0.076.

The systems

Forced aligners, given the reference transcript. BFA · Charsiu · CrisperWhisper · FALCON · MAPS · Montreal Forced Aligner 2.0 and 3.4 · MMS-FA · NeMo Forced Aligner · NeuFA · Qwen3-ForcedAligner · stable-ts · TorchAudio-FA · UnitY2 · WhisperX

Open ASR with timestamps, recognizing their own words. Parakeet-TDT · Qwen3-ASR · TorchAudio wav2vec2 · Whisper large-v3 · Whisper-timestamped

Commercial APIs. Olign · Amazon Transcribe · AssemblyAI Universal 3.5 · Azure AI Speech · Deepgram Nova-3 · ElevenLabs Scribe v2 · Google Chirp 2 · IBM Watson · Speechmatics

We also score 19 two-step cascades, in which one of the aligners re-times the words an ASR produced.

How we score

Before anything can be measured, each unit a system produced has to be paired with a unit in the reference. FA-Bench pairs them in two ways, because each way has a blind spot that the other covers.

Boundary MAE, median and signed error pair label first. A Levenshtein alignment runs over the labels, and the times of the pairs it finds are compared. Every figure carries a bootstrap 95 % confidence interval, with threshold accuracy at 10, 50 and 100 ms beside it. An unpaired phone drops out of this average, so skipping a hard phone costs a system nothing here.

S, D, I and PER show why a reference phone went unpaired. S counts phones the system relabelled, D counts phones it never emitted, and I counts phones it invented. A system that skips hard phones to lower its MAE shows up in D, so a lower MAE with a higher deletion rate is not necessarily better.

Boundary F1 at 20 ms pairs by label and time together. A boundary counts only when the units on both sides of it match the reference and it lands within 20 ms, with the two utterance edges included. This is the headline number. A skipped phone costs the boundaries beside it, and a substitution at the right moment scores nothing.

Time-only precision, recall, over-segmentation and R-value are on the Details pages. They ignore labels, so a boundary at the right moment on the wrong word counts as found there. The headline F1 gives it no credit.

The word tier reports MAE and the same label-checked F1 at 20 ms. It does not depend on phones, so it is the only tier where a system that outputs words alone can be scored.

Run it yourself

FA-Bench ships manifests and recipes. You obtain the audio yourself, TIMIT from the LDC and Buckeye through OSU registration, under their own licences. Run fabench init and it asks where each corpus is.

You can check the measurement chain before you touch any licensed audio. The self-test builds a synthetic corpus and runs the whole chain on it. Each gate compares a result against an answer fixed in advance, so a broken chain fails the gate.

git clone https://github.com/olewave/fa-bench && cd fa-bench
uv venv --python 3.12 .venv && . .venv/bin/activate
uv pip install -e ".[test]"
fabench selftest
fabench gates

Full results are on GitHub for TIMIT words and Buckeye words, with the methodology.

FA-Bench is released under PolyForm Noncommercial 1.0.0. It is free for research, teaching, personal study, and work by charitable, educational, public safety, environmental and government organisations. Commercial use needs a separate licence from Olewave.