FA-Bench: A Benchmark for Evaluating Phone- and Word-Level Timestamp Accuracy in Forced Aligners
Most forced aligner results are hard to reproduce. FA-Bench is an open, standardized baseline. We scored 30 systems on phone- and word-level timestamp accuracy against boundaries that human linguists placed, on clean audio and on four kinds of degradation.
github.com/olewave/fa-bench · Paper on arXiv
We're releasing FA-Bench, an open benchmark for the accuracy of phone- and word-level timestamps.
Most forced aligner results are hard to reproduce. Papers use different phoneme sets, their own test splits and their own evaluation metrics, so a number in one paper rarely measures the same thing as a number in the next. That leaves a speech researcher unable to run a clean ablation study, and a voice AI engineer unable to tell which system to deploy.
FA-Bench publishes the whole evaluation as code, from how labels are folded to how each boundary is scored. Every system runs through the same pipeline, so its numbers compare directly with every other system's.
What we tested
We scored 30 systems, 21 open models and 9 commercial APIs. Each one runs on clean audio and on four degraded versions of it, with reverb, noise, music or babble added. Every results table puts the clean score and the four degraded scores on one row, so you can see how much accuracy a system keeps as the audio gets harder.
Scoring uses read speech from TIMIT and spontaneous speech from Buckeye. Both corpora ship phone and word boundaries that human linguists corrected by hand, and those boundaries are the reference. FA-Bench adds no annotation of its own. It keeps the boundary times exactly as the corpora give them and folds the phone labels into the shared TIMIT-39 set.
Results are reported in two tracks. In the gold-transcript track, each system is handed the reference words, so its score measures timing alone. In the ASR track, each system recognizes the words itself, so recognition errors add to its timing errors. The two tracks are ranked separately.
What we found
Olign 1.0 has the highest word boundary F1 of every system given the reference transcript, on both corpora, clean and with noise added.
Word boundary F1 at 20 ms, gold-transcript track. Higher is better.
| System | TIMIT core-test | Buckeye test |
|---|---|---|
| clean / noisy | clean / noisy | |
| Olign 1.0 | 0.782 / 0.669 | 0.744 / 0.659 |
| MFA 3.4 | 0.644 / 0.624 | 0.682 / 0.583 |
| Qwen3-ForcedAligner-0.6B | 0.399 / 0.377 | 0.432 / 0.391 |
| WhisperX | 0.180 / 0.170 | 0.165 / 0.165 |
Noisy is the mean of the four degraded conditions. With noise added, Olign's lead over Montreal Forced Aligner 3.4 narrows on TIMIT, from 0.138 to 0.045, and widens on Buckeye, from 0.062 to 0.076.
The systems
Forced aligners, given the reference transcript. BFA · Charsiu · CrisperWhisper · FALCON · MAPS · Montreal Forced Aligner 2.0 and 3.4 · MMS-FA · NeMo Forced Aligner · NeuFA · Qwen3-ForcedAligner · stable-ts · TorchAudio-FA · UnitY2 · WhisperX
Open ASR with timestamps, recognizing their own words. Parakeet-TDT · Qwen3-ASR · TorchAudio wav2vec2 · Whisper large-v3 · Whisper-timestamped
Commercial APIs. Olign · Amazon Transcribe · AssemblyAI Universal 3.5 · Azure AI Speech · Deepgram Nova-3 · ElevenLabs Scribe v2 · Google Chirp 2 · IBM Watson · Speechmatics
We also score 19 two-step cascades, in which one of the aligners re-times the words an ASR produced.
How we score
Before anything can be measured, each unit a system produced has to be paired with a unit in the reference. FA-Bench pairs them in two ways, because each way has a blind spot that the other covers.
Boundary MAE, median and signed error pair label first. A Levenshtein alignment runs over the labels, and the times of the pairs it finds are compared. Every figure carries a bootstrap 95 % confidence interval, with threshold accuracy at 10, 50 and 100 ms beside it. An unpaired phone drops out of this average, so skipping a hard phone costs a system nothing here.
S, D, I and PER show why a reference phone went unpaired. S counts phones the system relabelled, D counts phones it never emitted, and I counts phones it invented. A system that skips hard phones to lower its MAE shows up in D, so a lower MAE with a higher deletion rate is not necessarily better.
Boundary F1 at 20 ms pairs by label and time together. A boundary counts only when the units on both sides of it match the reference and it lands within 20 ms, with the two utterance edges included. This is the headline number. A skipped phone costs the boundaries beside it, and a substitution at the right moment scores nothing.
Time-only precision, recall, over-segmentation and R-value are on the Details pages. They ignore labels, so a boundary at the right moment on the wrong word counts as found there. The headline F1 gives it no credit.
The word tier reports MAE and the same label-checked F1 at 20 ms. It does not depend on phones, so it is the only tier where a system that outputs words alone can be scored.
Run it yourself
FA-Bench ships manifests and recipes. You obtain the audio yourself, TIMIT from
the LDC and Buckeye through OSU registration, under their own licences. Run
fabench init and it asks where each corpus is.
You can check the measurement chain before you touch any licensed audio. The self-test builds a synthetic corpus and runs the whole chain on it. Each gate compares a result against an answer fixed in advance, so a broken chain fails the gate.
git clone https://github.com/olewave/fa-bench && cd fa-bench
uv venv --python 3.12 .venv && . .venv/bin/activate
uv pip install -e ".[test]"
fabench selftest
fabench gates
Full results are on GitHub for TIMIT words and Buckeye words, with the methodology.
FA-Bench is released under PolyForm Noncommercial 1.0.0. It is free for research, teaching, personal study, and work by charitable, educational, public safety, environmental and government organisations. Commercial use needs a separate licence from Olewave.
