We proudly offer

Olign — Olewave’s Lancet-Accurate Speech-To-Text Forced Alignment Service

A service that aligns speech to a transcript, or to an ASR result, and returns word- and phone-level timings with confidence scores. Its word timestamps are more accurate than those from Montreal Forced Aligner, Qwen3, and WhisperX.

REST ยท BETA https://api.olewave.com/olign/v1

The public REST API is in closed beta. Contact us to request a testing token and access.

Word Timestamp Accuracy on FA-Bench

FA-Bench
0.0 0.2 0.4 0.6 0.8 1.0 F1 0.802 Olign 0.730 MFA 0.436 Qwen3 0.149 WhisperX Quiet 0.737 Olign 0.659 MFA 0.400 Qwen3 0.152 WhisperX Noisy
OlignMFAQwen3-FAWhisperX
Word-boundary F1 @ 20 ms tolerance — a boundary is counted correct if it lands within 20 ms of the gold-standard boundary. Higher is better. Evaluated on the TIMIT core-test and Buckeye test splits as defined in FA-Bench, under its fabench protocol: boundaries are paired to the gold standard by time alone, labels ignored, on a 10 ms grid, with bootstrap 95 % confidence intervals. All four systems are scored in track 1 — forced alignment on the reference transcript — so only timing is measured and no recognition error enters the number. The splits are held out and speaker-disjoint: TIMIT core-test with the SA sentences excluded as the corpus documentation requires, and a Buckeye test split stratified on that corpus's own sex × age design (24/8/8 speakers), with per-utterance membership committed in the repo. Noisy is the mean of the four degradations — reverb, noise, music and babble. Full methodology. Baselines: Montreal FA (MFA) 3.4, Qwen3-ForcedAligner-0.6B and WhisperX v3.1.

Request beta access to Olign.

The REST API is in closed beta. Get a testing token, run Olign against your workload โ€” batch cleaning, real-time captioning, pronunciation assessment, dataset labeling โ€” and see the numbers on your own data.

Contact Olewave for a beta token โ†’