Live caption tool specifically made for interviews.
  • Python 67%
  • JavaScript 13.2%
  • CSS 7.6%
  • Shell 6.7%
  • Nix 3.1%
  • Other 2.4%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
[email protected] f2146189b9 Review recording, desktop launcher, MIT license
Record both sides of the call next to the session log for afterwards:
remote.wav (the exact stream the captions were made from) and mic.wav
(you, recorded only - never captioned, never translated). Pause stops
both, so the control for "this minute should not exist" already exists.

scripts/launch.sh + a .desktop entry make it a one-click start that runs
preflight and opens the captions tab itself, reading LLAMA_API_KEY from
~/.config/live-captions/env since a menu-launched process inherits none
of the shell's exports.

Docs: frame the tool as what it is (Japanese interviews), document the
jmdict build and desktop entry as first-run setup, fix the stale test
count and the ../flake.nix path that is now ../whisper/flake.nix.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-10-05 13:31:08 +09:00
app Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
data init 2026-10-04 21:34:02 +09:00
scripts Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
tests Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
web Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
.gitignore init 2026-10-04 21:34:02 +09:00
config.toml Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
flake.lock init 2026-10-04 21:34:02 +09:00
flake.nix Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
LICENSE Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00
README.md Review recording, desktop launcher, MIT license 2026-10-05 13:31:08 +09:00

live-captions — live JA→EN captions for Japanese interviews

A single-purpose tool: sit in a Japanese-language interview conducted over a video call and read what is being said, fast enough to answer. Everything below is shaped by that one job — the latency budget, the per-word glosses, the glossary, the EN→JA phrase helper, the fact that your own mic is recorded but never captioned.

Captures the call audio on hachi (the monitor of the default sink), segments it with Silero VAD, transcribes Japanese with a ReazonSpeech zipformer on the CPU, and sends each finished sentence to Qwen3-30B-A3B on fauna. Captions render in a browser tab on hachi — or on a phone at http://hachi:8765.

Japanese is the main line, shown large with furigana. Under each word sits a short English gloss for the meaning that word has in this sentence — particles included (は = topic, た = past). The full sentence translation sits below in smaller text as a safety net, since Japanese word order doesn't assemble into English from glosses alone. Tap any word for its full JMdict entry.

 らいしゅう                        たつ    よてい
 来週      ヨーロッパ   に    立つ   予定   です
 next week  Europe     to    depart  plan   is
 │ I plan to depart for Europe next week.

Monday runbook

Launch Live Captions from rofi (rofi -show drun) or your app menu: it opens a terminal, runs preflight, starts the app and opens the captions tab itself. Ctrl-C in that terminal stops everything.

If the key is not already in your environment, put it in ~/.config/live-captions/env (LLAMA_API_KEY=...) — a menu-launched process inherits nothing from your shell.

By hand, one terminal plus a browser:

cd live-captions
nix develop                      # sets LLAMA_API_KEY from ../whisper/flake.nix

./scripts/preflight.sh           # check everything before the call
python -m app.main

ASR is in-process now, so there is no server to start first. If you switch asr.engine back to whispercpp you need the old third terminal again: ./scripts/run_whisper_server.sh, and preflight will check it instead.

One-time setup

You need Nix (the flake carries everything else), PipeWire for capture, and an OpenAI-compatible LLM endpoint for translation — here llama-swap on another box on the LAN, set in config.toml under [translation]. The two large data files are gitignored and fetched by the steps below.

The ReazonSpeech weights are 700MB and not in git:

mkdir -p data/reazonspeech && curl -L \
  https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-zipformer-ja-reazonspeech-2024-08-01.tar.bz2 \
  | tar -xj --strip-components=1 -C data/reazonspeech

Build the dictionary (downloads the latest jmdict-simplified release, ~1 min):

nix develop -c python scripts/build_jmdict.py

For the rofi/app-menu entry, drop this in ~/.local/share/applications/live-captions.desktop with the path and terminal adjusted:

[Desktop Entry]
Type=Application
Name=Live Captions
Comment=Live JA-EN captions for Japanese interviews
Exec=kitty --title live-captions -e /path/to/live-captions/scripts/launch.sh
Icon=audio-input-microphone
Terminal=false
Categories=AudioVideo;

Open http://localhost:8765, click ⚙, paste the interviewer's name, the company and any role jargon into the glossary, and hit Apply glossary. Then join the Teams call in the same browser.

Before you join

  • Pick your output device first. Capture follows the default sink; plugging in headphones mid-call changes it and the feed goes quiet. If you do switch, re-pick the device under ⚙.
  • Prefer Chromium for Teams. Teams web has historically been degraded on Firefox. Capture itself is browser-agnostic — it taps the sink, not the tab — so either browser works for the audio.
  • Mute other apps. Notification dings land in the monitor stream and can produce a junk caption.
  • Fill in the glossary. It feeds both the ASR (as ReazonSpeech hotwords) and the translation, and it is the cheapest accuracy win available — more so than on whisper, since hotwords also pull kana spellings back to kanji.

Measured performance (hachi, Radeon 780M)

Stage Measured Doc target
VAD silence → finalize 0.60 s 0.5–0.7 s
ASR final (ReazonSpeech, CPU) 0.06–0.10 s < 1.0 s
Tokenize + JMdict 1.8 ms < 50 ms
Translation + all word glosses 0.62–1.08 s < 2.0 s
End of speech → English ≈1.4 s < 3.5 s

Was ≈2.2 s on whisper/Vulkan, which cost 0.72–0.82 s for the same sentences. Worst case is the 15 s cap: a 33-word segment glosses in 1.8 s (349 output tokens, no truncation), so ~2.5 s end to end.

Verified end to end by playing Japanese speech through the speakers and reading it back off the sink monitor: 6 sentences in, 6 correct captions and translations out. python -m app.replay tests/fixtures/interview.wav re-runs that check against the fixtures at any time.

Saying something yourself (EN → JA)

Click Say…, type English, press Enter. The phrase comes back as Japanese in the same interlinear format, plus a romaji line to read out loud and a back-translation so you can check it says what you meant:

you want to say: Could you repeat that more slowly please?
  もう一度、ゆっくり言っていただけますか?
  mou ichido yukkuri itte itadakemasu ka
  │ Could you please say it again, slowly?

Two calls (~1.7 s): EN→JA, then the same gloss pass the captions use. Always ですます form — it exists to be said out loud in an interview. These phrases never enter the caption translation context; they are your words, not theirs.

The romaji line is less trivial than it looks and is covered by tests: a trailing small っ is merged with the next word (言っ+て → itte, not "ixtsu te"), auxiliaries are merged back onto their stem (すみ/ませ/ん → sumimasen), particles use pronunciation (は → wa, へ → e) while everything else uses spelling (経験 → keiken, not phonetic "keeken").

Controls

Control What it does
⏸ Pause / ▶ Resume Stops processing but keeps the capture device open, so resuming is instant and can't fail to reacquire the sink. Captions, VAD and ASR all go idle, and the review recording stops too.
Clear Clears this screen only. The JSONL log on disk and the server's backlog are untouched — it's for getting a clean view mid-session.
⚙ Furigana Kana above kanji.
⚙ Word glosses The per-word English under each word. Off = plain Japanese line.
⚙ Full sentence translation The dim line underneath. Off if you'd rather work for it.
⚙ Font size 80–200%.
Say… EN → JA for something you need to say. See above.

⚙ and Say… both toggle; Escape closes whichever is open.

Pause is also the answer to non-Japanese audio: an English video playing will produce Japanese-looking nonsense. ReazonSpeech is Japanese-only by construction — there is no language parameter to set and no detection to turn on, so this is now structural rather than a tuning choice. (On the whisper engine it was a measured rejection: auto-detect needs language=auto + verbose_json, which cost +1.3 s on every line.) Pausing is free.

Single-language views

URL Shows
/ Japanese with furigana + word glosses, English underneath. The operator's view.
/ja Japanese only — no glosses either, since those are English.
/en English sentences only.

Same page and same WebSocket behind all three — a body class picks which half gets painted — so a second screen on http://hachi:8765/en and the laptop on / stay in lockstep. /ja and /en hide the operator controls (Pause, Clear, Say…, ⚙) and never write the stored view prefs, so putting one on a projector cannot reconfigure the window you are driving. Font size carries over from /. The connection dot and the fauna readout stay on all three: on a screen nobody is driving, they are the only way to tell a broken feed from a quiet one.

Why these choices

  • ReazonSpeech on the CPU, not whisper on the GPU. A zipformer transducer decodes a sentence in 54–64 ms where whisper's encoder needs 759 ms, because whisper pads every segment to 30 s and this doesn't. Same content on all six fixtures. onnxruntime was already a dependency for Silero VAD, so sherpa-onnx adds a decode loop, not a second runtime — and the 780M now sits idle instead of being the bottleneck for every segment. Full table in app/asr.py.
  • "Vulkan is mandatory" was true and is now obsolete. It was true of whisper: 1.09 s encode on the 780M against 14.1 s on the CPU, which is why the architecture doc's SenseVoice CPU fallback was dropped instead of written. That was a fact about whisper's encoder, not about CPUs. The whispercpp engine is still there (asr.engine = "whispercpp") and still needs Vulkan.
  • int8 encoder, modified_beam_search. fp32 is 4x the size for a byte-identical transcript on every fixture. Beam search costs ~10 ms over greedy and is non-negotiable: hotwords are silently ignored under greedy_search, so greedy would quietly discard the whole glossary.
  • The glossary matters more than it did. ReazonSpeech prefers kana where whisper picked kanji (もの for 物, たつ for 発つ), and a kana verb loses its furigana ruby and makes the JMdict lookup ambiguous. Hotwords fix it exactly: 発つ in the glossary turns …にたつ into …に発つ. They are a biasing FST rather than a prompt prefix, so unlike whisper's 224-token decoder prompt a long glossary costs nothing and can't get echoed into the transcript.
  • It is not strictly more accurate, just much faster. On the fixtures ReazonSpeech read 支払いにカードは使いますか where whisper got 使えますか — "do you use a card" instead of "can I use a card", an error that propagates into the translation. One sentence in six, and six sentences prove nothing either way. The trade bought is 12x latency and a free GPU, not better transcription.
  • No ASR partials — but the reason expired. They were ruled out because each one cost a full 1.09 s whisper encode (whisper pads every segment to 30 s), which made 500 ms partials arithmetically impossible and meant any partial in flight delayed the final. ReazonSpeech decodes in 54 ms, so re-decoding the in-progress segment every 300 ms now costs ~18% of one core and delays nothing. It is still off because nobody has tried it on a real call and unstable text that rewrites itself mid-sentence may read worse than the listening… 3.2 s pill, which costs nothing. This is now a UX question, not a budget one. vad.partial_interval_ms is the switch; the engine would need to re-decode segmenter._cur rather than a finished segment.
  • qwen3-30b-a3b-translate. 694 ms median against the 27b's 789 ms on the fixture set, full gloss coverage on both. The bigger win is operational: it cold-loads in ~16 s where the vLLM-backed 27b took ~141 s, so a mid-call swap is survivable rather than fatal. It is also Instruct-2507, the non-thinking variant, so no <think> block can eat the token budget. Preflight checks residency and warms it.
  • No cache_prompt. This model is llama.cpp-backed, where the field already defaults to true — measured, an identical prefix came back cache_n=30 with prompt eval dropping 201 ms → 33 ms. The system prompt is a constant and the message layout byte-stable so that cache hits. Sending the field buys nothing and would break a vLLM-backed model, which prefix-caches on its own and can reject unknown request fields.
  • Glosses are keyed by word index, not matched by string. fugashi has already split the sentence, so the model receives a numbered word list and replies with {"<index>": "<gloss>"}. It never re-segments Japanese, which is where this kind of feature usually breaks. A malformed gloss payload degrades to translation-only rather than dropping the line.
  • Streaming was dropped. response_format: json_object can't be consumed incrementally, and one call now returns translation and glosses in ~0.7 s.

Traps worth knowing

Both of these fail silently — they look like working code and produce a plausible-looking wrong result:

  1. pw-record needs stream.capture.sink=true. With --target <sink> alone it attaches to the default source and records the microphone: you get your own room instead of the call. A file still appears, with audio in it.

  2. This Silero ONNX export needs 576 samples, not 512 — 64 samples of carried context plus 512 new. Fed a bare 512 it returns ~0.0 for speech at 0.94 peak, so the VAD never fires and the feed stays empty.

  3. Autoscroll must not infer "the user scrolled up" from scroll position. A hidden tab, or a visible window sitting behind another one, stops running Chrome's rendering pipeline: scrollTo does not move scrollY, while rows arriving keep growing scrollHeight. That reads exactly like someone having scrolled up to re-read, so the old code latched its atBottom flag off — and since nothing reset it, the feed stayed frozen on one row even after you tabbed back. It now takes a real gesture (wheel, touchmove, keydown, mousedown) to leave the bottom; positions are only trusted for 400 ms after one. Tagging our own scrollTo and ignoring the event it queues looks simpler and is wrong: nothing orders that event against a setTimeout, and in a tab Chrome is not rendering it can arrive long after. Not in tests/ — it needs a DOM. The three cases to re-check by hand are tab away and back, scroll up to re-read (must stay put), and scroll back to the bottom (must resume).

  4. An author display: rule beats the browser's [hidden]. #settings{display:grid} and .listening{display:flex} silently outranked the UA stylesheet's [hidden]{display:none}, so toggling the attribute did nothing and both panels were stuck permanently open. There is now a [hidden]{display:none!important} rule near the top of style.css; any new panel that uses the attribute depends on it.

  5. Every word reserves a furigana slot, carrying &nbsp; when it has no reading. Without it only kanji words get a <ruby>, those columns are taller, and with vertical-align:top the glosses under particles ride up out of line with the rest of the row.

  6. app.replay must build the same tokenizer as the live app. It didn't, so it sent an empty word list and produced no glosses while the real app was fine. Both now load the NLP stack through app/nlp/__init__.py so they cannot drift — replay exists precisely to rule this class of thing out.

All covered by tests/test_pipeline.py.

Layout

app/
  main.py        FastAPI, /ws fan-out, control messages, device + dict routes
  pipeline.py    frames -> VAD -> ASR -> tokens -> translation (shared with replay)
  audio.py       pw-record capture, PipeWire device listing
  vad.py         Silero VAD segmenter (CPU, 0.067ms/frame)
  asr.py         ReazonSpeech (CPU) and whisper-server (GPU) engines + filter
  translate.py   streaming fauna client, FIFO queue, rolling context, glossary
  nlp/           fugashi+UniDic tokenizer, JMdict SQLite lookup
  session.py     ring buffer for reconnects + JSONL log
  replay.py      python -m app.replay <file.wav>
  config.py      config.toml loading
  events.py      the wire format shared by /ws and the JSONL log
  wavio.py       wav read/wrap/stream-write (fixtures, replay, recording)
scripts/         launch.sh, preflight.sh, run_whisper_server.sh, build_jmdict.py
web/             single static page, vanilla JS

Other commands

python -m app.replay tests/fixtures/interview.wav      # full pipeline on a file
python -m app.replay somefile.wav --no-translate       # ASR only, no fauna
python -m pytest tests/ -q                             # 22 tests, ~4s
python scripts/build_jmdict.py                         # rebuild data/jmdict.sqlite

Session transcripts land in data/sessions/<timestamp>/session.jsonl with per-stage timings and the per-word glosses, and download from the ⚙ panel — useful to read back over afterwards. Clear never touches it.

Recording the call for review

session.save_audio = true keeps the audio next to that transcript: remote.wav (the call, exactly the stream the captions were made from) and, when audio.mic_node is "default" or a node name, mic.wav (you). Both are mono 16kHz s16 — about 1.9MB per side per minute. Your mic is only recorded, never captioned.

Both sides of the call end up on disk, so whether that is okay is your call to make before the call, not after.

Starting and stopping it is the app itself: launching starts the recording, Ctrl-C ends it, and each run gets its own session directory — there is no separate record button to forget. ⏸ Pause stops the recording too, both sides of it, which is the control to hit for a minute that should not exist. Resume continues the same two files, with the paused span simply absent (so wav time drifts from the session.jsonl timestamps across a pause).

The two files are separate, which is what you want for review — you can hear one side without the other. To hear them as a conversation:

ffmpeg -i remote.wav -i mic.wav -filter_complex amix=inputs=2 call.wav

Recording costs nothing measurable: it is a writeframes on frames the capture loop already has in hand, with no second pw-record on the call side. The wav header is only finalised on a clean exit — if the app is SIGKILLed, the bytes are all there but the header claims zero frames, so recover with ffmpeg -f s16le -ar 16000 -ac 1 -i remote.wav fixed.wav.

Config

Everything is in config.toml. The values most likely to need touching: audio.loopback_node (pin a device instead of following the default sink), vad.silence_ms (how long a pause ends a sentence), vad.max_segment_s (hard cap; long interviewer turns get split and the rolling context carries meaning across), and translation.model.

License

MIT — see LICENSE.