Skip to main content

The highlight that follows the singer, not the line.

Word-by-Word Lyric Sync

Line-level sync puts a whole lyric on screen at once and leaves the audience guessing where in it the singer is. Word-level sync highlights each word as it arrives — the behaviour people recognise from commercial karaoke. Voxmin generates word timings automatically and keeps them editable, so a word that lands early is a correction rather than a re-render.

  • Editing is free and unlimited
  • LRC and ASS export
  • Nothing to install
Cue sheet
1 issue
00:12.0000:12.0000:25.00

the city lights are fading out

00:12.48 – 00:15.12

and I'm still counting every mile

00:15.12 – 00:17.96

you said we'd never look behind

00:17.96 – 00:21.60 Overlaps the next line

but here I am again tonight

00:21.04 – 00:24.60

Line 3 runs 0.56s into line 4.

Line 3 ends at 00:21.60 and overlaps line 4.

A real cue sheet, mid-correction. Line 3 is the one the AI missed.

What actually goes wrong

Automatic timing is right far more often than it is wrong. These are the three ways it misses, and each has a fix that takes seconds.

00:27.60

The whole line lights up at once

Nothing tells the singer which word they are on, which defeats the point of a karaoke track.

01:04.28

Held notes break the highlight

A word sung across four beats gets the same duration as a quick one, so the highlight races ahead of the voice.

01:52.04

Fast sections lose the singer

Dense phrasing outruns timings that were estimated rather than heard.

Four keys do most of the work

Timing a line is tapping a key on the beat, the way a stopwatch works. You never drag a handle to a waveform.

  1. Press I. Play the song and tap where a line starts

    Press I on the beat the line comes in. The cue takes that timestamp — no dragging a handle to a waveform you cannot read.

  2. Press O. Tap again where it ends

    Press O on the last syllable. If you were slightly late, nudge the whole line by 10ms steps rather than re-tapping it.

  3. Press S. Split a line the AI ran together

    Put the playhead where the break belongs and press S. Long transcribed lines become two cues that each hold the screen long enough to read.

  4. Press R. Hear just that line back

    Press R to replay the selected cue on its own. Checking one line no longer means scrubbing the whole song to find it again.

Everything else in the sheet

Nudge the whole sheet at once

When every line is late by the same amount — usually an intro the AI misjudged — move all of them together instead of fixing 40 cues one at a time.

Problems are flagged, not hunted

Lines that overlap the next one, end before they start, or are too short to read are marked in the sheet as you work.

Undo everything, save nothing

Full undo and redo history with ⌘Z. Your work saves on its own, so an experiment costs you nothing.

Export LRC and ASS

Take the corrected timings out as an enhanced LRC with word timings, or an ASS subtitle file — or render the video from the same sheet.

Bring your own lyrics

Paste the correct words when the transcription mishears them, and keep the timings the AI already got right.

100 languages transcribed

Hindi, Korean, Japanese, Spanish, Tamil, Arabic and more, with translation into 24 languages alongside the original.

the difference

Most AI lyric tools render a video and stop. When a line lands late, your only move is to run the job again with different settings and hope. Voxmin keeps the timings in a sheet you can open — so the fix is a keypress, not a second render.

Open the editor

Questions

Keep reading