We proudly offer

Olign — Olewave’s Lancet-Accurate Speech-To-Text Forced Alignment Service

Send a recording with its transcript, or with an ASR result. Without either, Olign transcribes the audio itself. Every word and phone comes back with its start and end time and a score from 0 to 100. Its word timestamps are more accurate than those from Montreal Forced Aligner, Qwen3-ForcedAligner and WhisperX.

REST · BETA https://api.olewave.com/olign/v1

Free beta access is open. Request an invitation code and create your own API key.

Word timestamps are often less accurate than they look. In the FA-Bench paper, Whisper places words about 150 ms early, and seven of the eight commercial speech-to-text APIs start their words late, three of them by more than 50 ms. Forced aligners fail differently. With background noise added to the Buckeye test split, Montreal Forced Aligner 3.4 returned no alignment for 354 of the 4,513 utterances.

We built Olign to place word boundaries accurately on real recordings, including noisy ones.

The highest boundary F1 of any forced aligner

Given the same transcript, Olign places word boundaries more accurately than every other forced aligner FA-Bench scores, on both corpora, clean and with noise added. On TIMIT core-test its word boundary F1 is 0.782, against 0.644 for Montreal Forced Aligner 3.4. Its word starts land 3.1 ms early on average. F1 is the benchmark's primary metric because it counts every word a system fails to place. Mean absolute error leaves those words out.

Better timing for the ASR you already use

Keep your speech recognizer and let Olign time its words. Handed the words Google Chirp 2 recognized on TIMIT core-test, Olign raises their boundary F1 from 0.443 to 0.734. The recognized words are unchanged, so the whole gain comes from timing. Run this way, Olign beats the timestamps of all eight commercial APIs in the benchmark on every split.

Alignments that hold up on noisy audio

Across the Buckeye test split, clean and under four kinds of degradation, Olign returned an alignment for all but 3 of 22,565 utterances. Montreal Forced Aligner 3.4 returned nothing for 573 of them.

Numbers you can check

Every result on this page comes from FA-Bench, which is open source and described in a paper on arXiv. The v0.9 API that produced the published numbers stays available, so you can reproduce them and run the same benchmark against your current stack.

Self-hosted when audio cannot leave your network

If your recordings have to stay inside your own VPC, we can discuss a self-hosted deployment.

Word Timestamp Accuracy on FA-Bench

FA-Bench
0.0 0.2 0.4 0.6 0.8 1.0 F1 0.763 Olign 0.663 MFA 0.415 Qwen3 0.172 WhisperX Quiet 0.664 Olign 0.603 MFA 0.384 Qwen3 0.168 WhisperX Noisy
OlignMFAQwen3-FAWhisperX
Word boundary F1 at 20 ms, and higher is better. A boundary counts as correct only when the words on both sides of it match the reference and it lands within 20 ms of the hand-labelled boundary. The splits are TIMIT core-test and Buckeye test as defined in FA-Bench, September 2026 snapshot, scored under its fabench protocol on a 10 ms grid with bootstrap 95 % confidence intervals. All four systems run in the gold-transcript track, so each is handed the reference words and only its timing is measured. No recognition error enters these numbers. Both splits are held out and speaker-disjoint. TIMIT core-test excludes the SA sentences as the corpus documentation requires, and the Buckeye test split is stratified on that corpus's own sex by age design over 24, 8 and 8 speakers, with per-utterance membership committed in the repo. Noisy is the mean of the four degradations, which are reverb, noise, music and babble. Baselines are Montreal Forced Aligner 3.4.1, Qwen3-ForcedAligner-0.6B and WhisperX 3.8.6. Full results for TIMIT and Buckeye.

Free beta access is open now for US and UK English, with more languages planned. Request an invitation code, create your API key and send whole recordings of up to 60 minutes as jobs. The beta covers 60 minutes of audio per account each day, and a job takes about an eighth of the recording's length to run. The accuracy figures on this page are measured on American English.

Try Olign free during the beta.

Request an invitation code and create your API key. The beta gives you 60 minutes of audio a day, which is enough to measure Olign against the aligner or ASR you use today.

Request an invitation code →