From a3b323f6685ec02edcd94fa9510d34df2be2e337 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Miguel=20=C3=81ngel?= Date: Mon, 5 Oct 2026 22:25:27 -0400 Subject: [PATCH 1/2] docs(cli): record transcription output contract --- ...26-10-05-transcribe-language-progress.html | 66 +++++++++++++++++++ 1 file changed, 66 insertions(+) create mode 100644 docs/contracts/2026-10-05-transcribe-language-progress.html diff --git a/docs/contracts/2026-10-05-transcribe-language-progress.html b/docs/contracts/2026-10-05-transcribe-language-progress.html new file mode 100644 index 0000000000..c78d61dff3 --- /dev/null +++ b/docs/contracts/2026-10-05-transcribe-language-progress.html @@ -0,0 +1,66 @@ +Transcription language and progress contract

Transcription language and progress

Carry automatic language detection in the JSON result and emit progress separately so stdout stays parseable.

2026-10-05 · hyperframes @ d69eac0cb · 4 boundaries audited · native execution not exercised · decisions approved
+
In one breath. stdout is one result, stderr carries typed progress records plus diagnostics, and detectedLanguage reports inference rather than a requested language.
+

Where the value travels

Whisper JSON result.language → transcription result → CLI stdout → JSON consumer
+HTTP response bytes → download pipeline → model callback → CLI stderr → progress consumer
+Whisper words → normalization → transcript.json (word array, unchanged)
+

The contract, per boundary

+

Language

The native result identifies the language used to decode. Only automatic selection qualifies as detection.

result.language: string
+Proposed terminal addition: detectedLanguage: string | null

Upstream JSON writes result.language from whisper_full_lang_id. Explicit language also sets that state. Existing detection incorrectly searches stdout although the native logger writes stderr; missing detection then omits --language and falls back to native English default.

whisper.cpp 60c0be6ac: examples/cli/cli.cpp:84,735-736; src/whisper.cpp:6993-7013,7151-7153,9357-9366. CLI: whisper/transcribe.ts:17-23,466-504,534,560.
exists: native JSONreachable: dropped by wrapperwritten: no terminal field
+

Result and artifact

Add metadata to the terminal result and preserve the existing word-array file.

{ok:true, engine, model, wordCount, durationSeconds,
+ speechOnsetSeconds, transcriptPath, detectedLanguage}

The existing model field currently reports the requested model even when the wrapper selects a multilingual variant. The wrapper will own the resolved model and return it with the result.

commands/transcribe.ts:397 writes Word[]; :400-410 emits stdout. Existing file consumer: commands/transcribe.ts:384 calls loadTranscript. External new-field consumption remains unverified.
exists: result and wordsreachable: stdout and fileread: additions pending
+

Download progress

Count bytes in the existing transfer pipeline. Unknown size stays unknown.

{"type":"progress","phase":"download","model":"small",
+ "receivedBytes":65536,"totalBytes":null}

DownloadOptions currently has timeoutMs and maxBytes only. Add a byte callback at the transfer owner, thread it through ensureModel, and serialize it only at the command boundary. Cache hits emit no download events. A known total is bytes, never an invented percentage.

utils/download.ts:10-15,113-122; whisper/manager.ts:215-223; commands/transcribe.ts:341,347-354.
exists: callback absentreachable: JSON disables callbackswritten: no progress records
+

Transcription progress

Indicate work started and completed without inventing numerical progress.

{"type":"progress","phase":"transcription","model":"small",
+ "status":"started"}
+{"type":"progress","phase":"transcription","model":"small",
+ "status":"completed"}

Progress records are individual valid JSON lines on stderr. Other stderr diagnostics remain allowed. The terminal stdout result is emitted once. Failure never emits a completed transcription event.

whisper/transcribe.ts:488,512-515; whisper/parakeet.ts:294-304; whisper/sherpa.ts:397-413,445. Plain fallback diagnostics: commands/transcribe.ts:370.
exists: text callbacksreachable: JSON disabledwritten: no typed records
+

How it says no

SignalMeaning and response
detectedLanguage:nullNo actual automatic inference was reported, including explicit language, English-only model, or an engine without detection. Never infer language from a model label.
totalBytes:nullResponse has no trustworthy size. Show bytes or indeterminate progress.
ok:false / nonzero exitOperation failed. Diagnostic stderr may explain why. No completion event is evidence of success without the terminal result.
Progress line without typeNot part of this progress contract. Consumers filter diagnostics separately.
+

Decisions approved

  1. detectedLanguage means real automatic detection. Explicit --language and English-only models return null.
  2. stderr permits diagnostics alongside progress; every progress line is valid JSON with type.
  3. model reports the actual resolved model. The file stays Word[].
  4. Transcription progress is phase-only. Download progress reports real bytes with a nullable total.
+

Coverage, honestly

Audited: CLI producer, transfer pipeline, native JSON writer and logger, current terminal/file paths. Installed CLI help executed.

Trusting: arbitrary environment-selected Whisper versions are not pinned to the inspected upstream checkout.

Not exercised: native audio inference, new progress writer and external live consumer. These are contract findings, not execution proof.

+

What would prove this page wrong

  • A public transcribe invocation with a native fixture that emits detection only on stderr must decode automatically and return its language.
  • Explicit language and English-only models must return null even if native JSON reports a language.
  • Download callbacks must observe actual chunks, including unknown totals, and stdout must parse as one result while stderr records parse independently.
  • A switched multilingual model must be the model reported in the final result.
\ No newline at end of file From 366f1accba6f14a3ddbd22fe1668887f56b5cfe6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Miguel=20=C3=81ngel?= Date: Mon, 5 Oct 2026 23:00:57 -0400 Subject: [PATCH 2/2] fix(cli): report transcription language and progress --- ...26-10-05-transcribe-language-progress.html | 2 +- packages/cli/src/commands/transcribe.test.ts | 65 ++++++++- packages/cli/src/commands/transcribe.ts | 56 ++++++-- packages/cli/src/utils/download.test.ts | 42 ++++++ packages/cli/src/utils/download.ts | 36 +++-- .../cli/src/whisper/manager.install.test.ts | 40 ++++++ packages/cli/src/whisper/manager.ts | 13 +- packages/cli/src/whisper/parakeet.ts | 27 +++- packages/cli/src/whisper/progress.test.ts | 37 +++++ packages/cli/src/whisper/progress.ts | 30 +++++ packages/cli/src/whisper/sherpa.ts | 29 +++- .../src/whisper/transcribe.metadata.test.ts | 127 ++++++++++++++++++ packages/cli/src/whisper/transcribe.ts | 95 ++++++------- 13 files changed, 511 insertions(+), 88 deletions(-) create mode 100644 packages/cli/src/whisper/manager.install.test.ts create mode 100644 packages/cli/src/whisper/progress.test.ts create mode 100644 packages/cli/src/whisper/progress.ts create mode 100644 packages/cli/src/whisper/transcribe.metadata.test.ts diff --git a/docs/contracts/2026-10-05-transcribe-language-progress.html b/docs/contracts/2026-10-05-transcribe-language-progress.html index c78d61dff3..27caf59991 100644 --- a/docs/contracts/2026-10-05-transcribe-language-progress.html +++ b/docs/contracts/2026-10-05-transcribe-language-progress.html @@ -61,6 +61,6 @@ {"type":"progress","phase":"transcription","model":"small", "status":"completed"}

Progress records are individual valid JSON lines on stderr. Other stderr diagnostics remain allowed. The terminal stdout result is emitted once. Failure never emits a completed transcription event.

whisper/transcribe.ts:488,512-515; whisper/parakeet.ts:294-304; whisper/sherpa.ts:397-413,445. Plain fallback diagnostics: commands/transcribe.ts:370.
exists: text callbacksreachable: JSON disabledwritten: no typed records

How it says no

SignalMeaning and response
detectedLanguage:nullNo actual automatic inference was reported, including explicit language, English-only model, or an engine without detection. Never infer language from a model label.
totalBytes:nullResponse has no trustworthy size. Show bytes or indeterminate progress.
ok:false / nonzero exitOperation failed. Diagnostic stderr may explain why. No completion event is evidence of success without the terminal result.
Progress line without typeNot part of this progress contract. Consumers filter diagnostics separately.
-

Decisions approved

  1. detectedLanguage means real automatic detection. Explicit --language and English-only models return null.
  2. stderr permits diagnostics alongside progress; every progress line is valid JSON with type.
  3. model reports the actual resolved model. The file stays Word[].
  4. Transcription progress is phase-only. Download progress reports real bytes with a nullable total.
+

Decisions approved

  1. detectedLanguage means real automatic detection. Explicit --language and English-only models return null.
  2. stderr permits diagnostics alongside progress; every progress line is valid JSON with type.
  3. model reports the actual resolved model. The file stays Word[].
  4. Transcription progress is phase-only. Download progress reports real bytes with a nullable total.
  5. --no-runtime-install requires an existing Whisper runtime on all paths, including auto fallback. Model downloads remain allowed. No prior opt-out exists in the inspected CLI source; ensureWhisper owns this policy.

Coverage, honestly

Audited: CLI producer, transfer pipeline, native JSON writer and logger, current terminal/file paths. Installed CLI help executed.

Trusting: arbitrary environment-selected Whisper versions are not pinned to the inspected upstream checkout.

Not exercised: native audio inference, new progress writer and external live consumer. These are contract findings, not execution proof.

What would prove this page wrong

  • A public transcribe invocation with a native fixture that emits detection only on stderr must decode automatically and return its language.
  • Explicit language and English-only models must return null even if native JSON reports a language.
  • Download callbacks must observe actual chunks, including unknown totals, and stdout must parse as one result while stderr records parse independently.
  • A switched multilingual model must be the model reported in the final result.
\ No newline at end of file diff --git a/packages/cli/src/commands/transcribe.test.ts b/packages/cli/src/commands/transcribe.test.ts index 307765cf00..e572d1c175 100644 --- a/packages/cli/src/commands/transcribe.test.ts +++ b/packages/cli/src/commands/transcribe.test.ts @@ -1,3 +1,4 @@ +import { runCommand } from "citty"; import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; import { chmodSync, existsSync, writeFileSync, readFileSync, mkdtempSync, rmSync } from "node:fs"; import { join } from "node:path"; @@ -44,7 +45,14 @@ import transcribeCmd from "./transcribe.js"; function fakeTranscript(dir: string, text: string) { const transcriptPath = join(dir, "transcript.json"); writeFileSync(transcriptPath, JSON.stringify([{ text, start: 0, end: 1 }])); - return { transcriptPath, wordCount: 1, durationSeconds: 1, speechOnsetSeconds: null }; + return { + model: text === "whisper" ? "small.en" : "parakeet-tdt-0.6b-v3", + detectedLanguage: null, + transcriptPath, + wordCount: 1, + durationSeconds: 1, + speechOnsetSeconds: null, + }; } function lastJson(): Record { @@ -81,6 +89,47 @@ describe("transcribe command", () => { vi.unstubAllEnvs(); }); + it("keeps typed progress on stderr and reports resolved metadata in one stdout result", async () => { + const { dir, input } = dummyAudio(); + dirs.push(dir); + const stderr = vi.spyOn(process.stderr, "write").mockImplementation(() => true); + transcribeMock.mockImplementation(async (_input, outputDir, options) => { + options.onEvent?.({ + type: "progress", + phase: "download", + model: "small", + receivedBytes: 5, + totalBytes: null, + }); + options.onEvent?.({ + type: "progress", + phase: "transcription", + model: "small", + status: "started", + }); + options.onEvent?.({ + type: "progress", + phase: "transcription", + model: "small", + status: "completed", + }); + return { ...fakeTranscript(outputDir, "hola"), model: "small", detectedLanguage: "es" }; + }); + await transcribeCmd.run!({ + args: { input, dir, json: true, engine: "whisper", model: "small.en" }, + } as never); + expect(console.log).toHaveBeenCalledTimes(1); + expect(lastJson()).toMatchObject({ ok: true, model: "small", detectedLanguage: "es" }); + expect(stderr.mock.calls.map(([line]) => JSON.parse(String(line)))).toEqual([ + { type: "progress", phase: "download", model: "small", receivedBytes: 5, totalBytes: null }, + { type: "progress", phase: "transcription", model: "small", status: "started" }, + { type: "progress", phase: "transcription", model: "small", status: "completed" }, + ]); + expect(JSON.parse(readFileSync(join(dir, "transcript.json"), "utf8"))).toEqual([ + { id: "w0", text: "hola", start: 0, end: 1 }, + ]); + }); + it("explicit run exits non-zero and is NOT reported as a command failure", async () => { const { dir, input } = dummyAudio(); dirs.push(dir); @@ -252,6 +301,20 @@ describe("transcribe command", () => { return { exitCode: exitCode || consumeCommandResult().exitCode, out: lastJson() }; } + it.each(["whisper", "auto"])( + "--no-runtime-install reaches %s including the fallback", + async (engine) => { + crashChild("SIGABRT"); + Object.assign(runners, { sherpa: true, mlx: false }); + const { dir, input } = dummyAudio(); + dirs.push(dir); + await runCommand(transcribeCmd, { + rawArgs: [input, "--engine", engine, "--json", "--no-runtime-install"], + }); + expect(transcribeMock.mock.calls.at(-1)?.[2]).toMatchObject({ installRuntime: false }); + }, + ); + it("auto falls back to whisper with one line naming the Parakeet error and the repair", async () => { crashChild("SIGABRT", "terminate called after throwing an instance of 'Ort::Exception'\n"); expect(await transcribeWith("auto", { sherpa: true, mlx: false })).toMatchObject({ diff --git a/packages/cli/src/commands/transcribe.ts b/packages/cli/src/commands/transcribe.ts index 2ef336a48a..cd278d7c07 100644 --- a/packages/cli/src/commands/transcribe.ts +++ b/packages/cli/src/commands/transcribe.ts @@ -1,3 +1,4 @@ +import { createProgressWriter } from "../whisper/progress.js"; import { failCommand, setCommandExitCode } from "../utils/commandResult.js"; import { normalizeErrorMessage } from "../utils/errorMessage.js"; // fallow-ignore-file code-duplication @@ -7,7 +8,6 @@ import { existsSync, rmSync, writeFileSync } from "node:fs"; import { findParakeet, PARAKEET_LANGUAGES, - PARAKEET_MODEL_LABEL, parakeetSpeaks, transcribeWithParakeet, } from "../whisper/parakeet.js"; @@ -77,7 +77,7 @@ export default defineCommand({ }, json: { type: "boolean", - description: "Output result as JSON", + description: "Output result as JSON; progress JSON lines go to stderr", default: false, }, to: { @@ -95,6 +95,12 @@ export default defineCommand({ "Keep each transcript entry as its own caption cue (skip word-level grouping). Use when exporting an already-cued transcript whose entries have no internal spaces, e.g. single-word or CJK captions.", default: false, }, + "runtime-install": { + type: "boolean", + default: true, + description: + "Allow installing the Whisper runtime. Use --no-runtime-install to require an existing runtime; model downloads remain allowed.", + }, optional: { type: "boolean", description: @@ -153,6 +159,7 @@ export default defineCommand({ language: args.language, json: args.json, optional: args.optional, + installRuntime: args["runtime-install"], timeoutMs, }); }, @@ -294,6 +301,7 @@ async function transcribeAudio( language?: string; json?: boolean; optional?: boolean; + installRuntime?: boolean; timeoutMs?: number; }, ): Promise { @@ -339,20 +347,39 @@ async function transcribeAudio( const spin = opts.json ? null : clack.spinner(); spin?.start(`Transcribing with ${label(runner)}...`); const onProgress = spin ? (msg: string) => spin.message(msg) : undefined; + const onEvent = opts.json ? createProgressWriter(process.stderr) : undefined; let wavPath = inputPath; // Before audio prep: under --json no spinner listens for SIGINT, so Ctrl-C would kill Node. const cancellation = runner === "sherpa" ? createRenderCancellationScope() : null; - const run = (r: Runner) => - r === "sherpa" - ? transcribeWithSherpa(wavPath, dir, { onProgress, signal: cancellation!.signal }) - : r === "parakeet-mlx" - ? transcribeWithParakeet(wavPath, dir, { language: opts.language, onProgress }) - : transcribe(wavPath, dir, { - model, - language: opts.language, - onProgress, - timeoutMs: opts.timeoutMs, - }); + const run = (r: Runner) => { + switch (r) { + case "sherpa": + return transcribeWithSherpa(wavPath, dir, { + onProgress, + onEvent, + signal: cancellation!.signal, + }); + case "parakeet-mlx": + return transcribeWithParakeet(wavPath, dir, { + language: opts.language, + onProgress, + onEvent, + }); + case "whisper": + return transcribe(wavPath, dir, { + model, + language: opts.language, + onProgress, + onEvent, + timeoutMs: opts.timeoutMs, + installRuntime: opts.installRuntime, + }); + default: { + const unreachable: never = r; + throw new Error(`Unknown transcription runner: ${unreachable}`); + } + } + }; try { // Outside the fallback: an unreadable input is not a Parakeet failure. The fallback reuses it. @@ -402,7 +429,8 @@ async function transcribeAudio( JSON.stringify({ ok: true, engine: runner === "whisper" ? "whisper" : "parakeet", - model: runner === "whisper" ? model : PARAKEET_MODEL_LABEL, + model: result.model, + detectedLanguage: result.detectedLanguage, wordCount: words.length, durationSeconds: result.durationSeconds, speechOnsetSeconds: result.speechOnsetSeconds, diff --git a/packages/cli/src/utils/download.test.ts b/packages/cli/src/utils/download.test.ts index 3e8b362ce5..adf0753ffd 100644 --- a/packages/cli/src/utils/download.test.ts +++ b/packages/cli/src/utils/download.test.ts @@ -1,3 +1,5 @@ +import { randomUUID } from "node:crypto"; +import { ensureModel } from "../whisper/manager.js"; import { EventEmitter } from "node:events"; import { existsSync, mkdtempSync, readFileSync, readdirSync, rmSync } from "node:fs"; import { tmpdir } from "node:os"; @@ -50,6 +52,46 @@ afterEach(() => { }); describe("downloadFile", () => { + it("carries model download progress through the real pipeline and stays quiet on cache hits", async () => { + mockGet.mockImplementation(httpsResponse(200, { "content-length": "5" }, "hello")); + const events: unknown[] = []; + const model = `test-${randomUUID()}`; + const options = { + onDownloadProgress: (receivedBytes: number, totalBytes: number | null) => + events.push({ receivedBytes, totalBytes }), + }; + const modelPath = await ensureModel(model, options); + try { + expect(readFileSync(modelPath, "utf8")).toBe("hello"); + expect(events).toEqual([ + { receivedBytes: 0, totalBytes: 5 }, + { receivedBytes: 5, totalBytes: 5 }, + ]); + events.length = 0; + expect(await ensureModel(model, options)).toBe(modelPath); + expect(events).toEqual([]); + expect(mockGet).toHaveBeenCalledTimes(1); + } finally { + rmSync(modelPath); + } + }); + + it.each([ + [{ "content-length": "5" }, 5], + [{}, null], + [{ "content-length": "invalid" }, null], + ])("reports actual response bytes with total %j", async (headers, totalBytes) => { + mockGet.mockImplementation(httpsResponse(200, headers as Record, "hello")); + const dir = mkdtempSync(join(tmpdir(), "hyperframes-download-")); + tempDirs.push(dir); + const events: unknown[] = []; + await downloadFile("https://example.test/model", join(dir, "model"), { + onProgress: (receivedBytes, total) => events.push({ receivedBytes, totalBytes: total }), + }); + expect(events).toContainEqual({ receivedBytes: 5, totalBytes }); + expect(readFileSync(join(dir, "model"), "utf8")).toBe("hello"); + }); + it("keeps concurrent partial downloads separate", async () => { mockGet.mockImplementation(httpsResponse(200, {}, (url) => url)); diff --git a/packages/cli/src/utils/download.ts b/packages/cli/src/utils/download.ts index ae6d34af78..d73c5eaacf 100644 --- a/packages/cli/src/utils/download.ts +++ b/packages/cli/src/utils/download.ts @@ -12,6 +12,7 @@ export interface DownloadOptions { timeoutMs?: number; /** Reject before writing more than this many response bytes. */ maxBytes?: number; + onProgress?: (receivedBytes: number, totalBytes: number | null) => void; } /** Every redirect a host may reasonably answer with, not just the two we saw first. */ @@ -40,17 +41,28 @@ function removePartialFile(path: string): void { } } -function enforceByteLimit(maxBytes: number): Transform { - let received = 0; - return new Transform({ +function measureDownload(options: DownloadOptions, contentLength: string | undefined) { + const length = Number(contentLength); + const totalBytes = Number.isSafeInteger(length) && length >= 0 ? length : null; + let receivedBytes = 0; + let reportedAt = Date.now(); + options.onProgress?.(0, totalBytes); + const stream = new Transform({ transform(chunk: Buffer, _encoding, callback) { - received += chunk.byteLength; - callback( - received > maxBytes ? new Error(`Download exceeded ${maxBytes} bytes`) : null, - chunk, - ); + receivedBytes += chunk.byteLength; + if (options.maxBytes !== undefined && receivedBytes > options.maxBytes) { + callback(new Error(`Download exceeded ${options.maxBytes} bytes`)); + return; + } + const now = Date.now(); + if (now - reportedAt >= 100) { + options.onProgress?.(receivedBytes, totalBytes); + reportedAt = now; + } + callback(null, chunk); }, }); + return { stream, complete: () => options.onProgress?.(receivedBytes, totalBytes) }; } /** @@ -73,7 +85,6 @@ export function downloadFile( ): Promise { const tmp = `${dest}.${process.pid}.${randomUUID()}.tmp`; const timeoutMs = options.timeoutMs ?? DEFAULT_DOWNLOAD_TIMEOUT_MS; - const maxBytes = options.maxBytes; return new Promise((resolve, reject) => { const follow = (u: string, hops = 0) => { let activeResponse: IncomingMessage | undefined; @@ -110,15 +121,14 @@ export function downloadFile( reject(new Error(`Download failed: HTTP ${res.statusCode}`)); return; } + const meter = measureDownload(options, res.headers["content-length"]); const file = createWriteStream(tmp); responsePipelineStarted = true; - const transfer = - maxBytes === undefined - ? pipeline(res, file) - : pipeline(res, enforceByteLimit(maxBytes), file); + const transfer = pipeline(res, meter.stream, file); transfer .then(() => { renameSync(tmp, dest); + meter.complete(); resolve(); }) .catch((err) => { diff --git a/packages/cli/src/whisper/manager.install.test.ts b/packages/cli/src/whisper/manager.install.test.ts new file mode 100644 index 0000000000..f959e98450 --- /dev/null +++ b/packages/cli/src/whisper/manager.install.test.ts @@ -0,0 +1,40 @@ +import { mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { afterEach, expect, it, vi } from "vitest"; +import { ensureWhisper, findWhisper, WhisperUnavailableError } from "./manager.js"; +const calls = vi.hoisted(() => ({ exec: vi.fn() })); +vi.mock("node:child_process", () => ({ execFileSync: calls.exec })); +vi.mock("node:os", async (original) => ({ + ...(await original()), + platform: () => "darwin", +})); +afterEach(() => { + vi.unstubAllEnvs(); + calls.exec.mockReset(); +}); +it("never installs when the runtime vanishes after discovery and installation is disabled", async () => { + const dir = mkdtempSync(join(tmpdir(), "hf-missing-runtime-")); + const binary = join(dir, "whisper-cli"); + writeFileSync(binary, "runtime"); + vi.stubEnv("HYPERFRAMES_WHISPER_PATH", binary); + calls.exec.mockImplementation((command: string, args: string[]) => { + if ((command === "which" || command === "where") && args[0] !== "whisper-cli") + return `/bin/${args[0]}\n`; + throw new Error("not available"); + }); + try { + expect(findWhisper()?.executablePath).toBe(binary); + rmSync(binary); + await expect(ensureWhisper({ installRuntime: false })).rejects.toBeInstanceOf( + WhisperUnavailableError, + ); + expect( + calls.exec.mock.calls.filter( + ([command]) => command === "brew" || command === "git" || command === "cmake", + ), + ).toEqual([]); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}); diff --git a/packages/cli/src/whisper/manager.ts b/packages/cli/src/whisper/manager.ts index 55d4e3b0de..e9cf61304f 100644 --- a/packages/cli/src/whisper/manager.ts +++ b/packages/cli/src/whisper/manager.ts @@ -176,11 +176,17 @@ function hasCmake(): boolean { } export async function ensureWhisper(options?: { + installRuntime?: boolean; onProgress?: (msg: string) => void; }): Promise { // 1. Already installed? const existing = findWhisper(); if (existing) return existing; + if (options?.installRuntime === false) { + throw new WhisperUnavailableError( + "whisper-cpp not found; runtime installation is disabled by --no-runtime-install.", + ); + } // 2. Try brew (macOS, fastest — pre-built bottle) if (platform() === "darwin" && hasBrew()) { @@ -212,7 +218,10 @@ export async function ensureWhisper(options?: { export async function ensureModel( model: string = DEFAULT_MODEL, - options?: { onProgress?: (message: string) => void }, + options?: { + onProgress?: (message: string) => void; + onDownloadProgress?: (receivedBytes: number, totalBytes: number | null) => void; + }, ): Promise { const modelPath = join(MODELS_DIR, `ggml-${model}.bin`); if (existsSync(modelPath)) return modelPath; @@ -220,7 +229,7 @@ export async function ensureModel( mkdirSync(MODELS_DIR, { recursive: true }); options?.onProgress?.(`Downloading model ${model}...`); - await downloadFile(getModelUrl(model), modelPath); + await downloadFile(getModelUrl(model), modelPath, { onProgress: options?.onDownloadProgress }); if (!existsSync(modelPath)) { throw new Error(`Model download failed: ${model}`); diff --git a/packages/cli/src/whisper/parakeet.ts b/packages/cli/src/whisper/parakeet.ts index ea805210aa..6a969ed639 100644 --- a/packages/cli/src/whisper/parakeet.ts +++ b/packages/cli/src/whisper/parakeet.ts @@ -18,7 +18,7 @@ import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync import { homedir, tmpdir } from "node:os"; import { basename, extname, join } from "node:path"; import type { Word } from "./normalize.js"; -import type { TranscribeResult } from "./transcribe.js"; +import type { TranscribeProgress, TranscribeResult } from "./transcribe.js"; /** The model name `transcribe --json` reports for every Parakeet runner. */ export const PARAKEET_MODEL_LABEL = "parakeet-tdt-0.6b-v3"; @@ -273,6 +273,7 @@ interface ParakeetOptions { language?: string; model?: string; onProgress?: (message: string) => void; + onEvent?: (event: TranscribeProgress) => void; } /** Transcribe with Parakeet and write `transcript.json` (Word[]) into `dir`. */ @@ -301,14 +302,27 @@ export function transcribeWithParakeet( try { const argv = [inputPath, "--model", model, "--output-format", "json", "--output-dir", workDir]; if (options?.language) argv.push("--language", options.language); + options?.onEvent?.({ + type: "progress", + phase: "transcription", + model: PARAKEET_MODEL_LABEL, + status: "started", + }); execFileSync(runner, argv, { stdio: ["ignore", "pipe", "pipe"], timeout: 1_800_000 }); const produced = join(workDir, `${basename(inputPath, extname(inputPath))}.json`); if (!existsSync(produced)) throw new Error("Parakeet did not produce output."); - return writeParakeetTranscript( + const result = writeParakeetTranscript( dir, mergeTokensToWords(JSON.parse(readFileSync(produced, "utf-8")) as ParakeetJson), ); + options?.onEvent?.({ + type: "progress", + phase: "transcription", + model: result.model, + status: "completed", + }); + return result; } finally { rmSync(workDir, { recursive: true, force: true }); } @@ -319,5 +333,12 @@ export function writeParakeetTranscript(dir: string, words: Word[]): TranscribeR const transcriptPath = join(dir, "transcript.json"); writeFileSync(transcriptPath, JSON.stringify(words, null, 2)); const durationSeconds = words.length > 0 ? words[words.length - 1]!.end : 0; - return { transcriptPath, wordCount: words.length, durationSeconds, speechOnsetSeconds: null }; + return { + model: PARAKEET_MODEL_LABEL, + detectedLanguage: null, + transcriptPath, + wordCount: words.length, + durationSeconds, + speechOnsetSeconds: null, + }; } diff --git a/packages/cli/src/whisper/progress.test.ts b/packages/cli/src/whisper/progress.test.ts new file mode 100644 index 0000000000..5fd8864148 --- /dev/null +++ b/packages/cli/src/whisper/progress.test.ts @@ -0,0 +1,37 @@ +import { Writable } from "node:stream"; +import { setImmediate } from "node:timers/promises"; +import { expect, it } from "vitest"; +import { createProgressWriter } from "./progress.js"; + +it("coalesces byte updates behind a slow reader while preserving phase records", async () => { + const lines: string[] = []; + let release: (() => void) | undefined; + const stream = new Writable({ + highWaterMark: 1, + write(chunk, _encoding, callback) { + lines.push(String(chunk)); + release = callback; + }, + }); + const emit = createProgressWriter(stream); + for (let receivedBytes = 0; receivedBytes < 10_000; receivedBytes++) { + emit({ type: "progress", phase: "download", model: "tiny", receivedBytes, totalBytes: null }); + } + emit({ type: "progress", phase: "transcription", model: "tiny", status: "started" }); + emit({ type: "progress", phase: "transcription", model: "tiny", status: "completed" }); + expect(lines).toHaveLength(1); + while (release) { + const next = release; + release = undefined; + next(); + await setImmediate(); + } + expect(lines.map((line) => JSON.parse(line))).toEqual([ + { type: "progress", phase: "download", model: "tiny", receivedBytes: 0, totalBytes: null }, + { type: "progress", phase: "download", model: "tiny", receivedBytes: 9999, totalBytes: null }, + { type: "progress", phase: "transcription", model: "tiny", status: "started" }, + { type: "progress", phase: "transcription", model: "tiny", status: "completed" }, + ]); + expect(stream.listenerCount("drain")).toBe(0); + stream.destroy(); +}); diff --git a/packages/cli/src/whisper/progress.ts b/packages/cli/src/whisper/progress.ts new file mode 100644 index 0000000000..773668cc80 --- /dev/null +++ b/packages/cli/src/whisper/progress.ts @@ -0,0 +1,30 @@ +import type { Writable } from "node:stream"; +import type { TranscribeProgress } from "./transcribe.js"; + +export function createProgressWriter(stream: Writable): (event: TranscribeProgress) => void { + let blocked = false; + const pending: TranscribeProgress[] = []; + const write = (event: TranscribeProgress): void => { + if (blocked) { + const previous = pending.at(-1); + if ( + event.phase === "download" && + previous?.phase === "download" && + previous.model === event.model + ) { + pending[pending.length - 1] = event; + } else { + pending.push(event); + } + return; + } + blocked = !stream.write(`${JSON.stringify(event)}\n`); + if (blocked) { + stream.once("drain", () => { + blocked = false; + while (!blocked && pending.length > 0) write(pending.shift()!); + }); + } + }; + return write; +} diff --git a/packages/cli/src/whisper/sherpa.ts b/packages/cli/src/whisper/sherpa.ts index 35dfab44ee..ce8436c8f0 100644 --- a/packages/cli/src/whisper/sherpa.ts +++ b/packages/cli/src/whisper/sherpa.ts @@ -17,10 +17,16 @@ import { mergeWindowsToWords, SHERPA_ERROR_PREFIX, SHERPA_RESULT_PREFIX, + PARAKEET_MODEL_LABEL, writeParakeetTranscript, type SherpaWindow, } from "./parakeet.js"; -import { getPreparedWavDurationSeconds, prepareWav, type TranscribeResult } from "./transcribe.js"; +import { + getPreparedWavDurationSeconds, + prepareWav, + type TranscribeProgress, + type TranscribeResult, +} from "./transcribe.js"; const RUNTIME = "sherpa-onnx-node"; const RUNTIME_VERSION = "1.13.8"; @@ -347,9 +353,26 @@ export function prepareSherpaWav( export async function transcribeWithSherpa( wavPath: string, dir: string, - options: { signal: AbortSignal; onProgress?: (message: string) => void }, + options: { + signal: AbortSignal; + onProgress?: (message: string) => void; + onEvent?: (event: TranscribeProgress) => void; + }, ): Promise { options.onProgress?.("Transcribing with Parakeet..."); + options.onEvent?.({ + type: "progress", + phase: "transcription", + model: PARAKEET_MODEL_LABEL, + status: "started", + }); const windows = await decode(wavPath, options.signal); - return writeParakeetTranscript(dir, mergeWindowsToWords(windows)); + const result = writeParakeetTranscript(dir, mergeWindowsToWords(windows)); + options.onEvent?.({ + type: "progress", + phase: "transcription", + model: result.model, + status: "completed", + }); + return result; } diff --git a/packages/cli/src/whisper/transcribe.metadata.test.ts b/packages/cli/src/whisper/transcribe.metadata.test.ts new file mode 100644 index 0000000000..3e84afd507 --- /dev/null +++ b/packages/cli/src/whisper/transcribe.metadata.test.ts @@ -0,0 +1,127 @@ +import { mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { afterEach, beforeEach, expect, it, vi } from "vitest"; +import { transcribe } from "./transcribe.js"; + +const native = vi.hoisted(() => ({ exec: vi.fn(), missingLanguage: false, runtime: vi.fn() })); +vi.mock("node:child_process", () => ({ execFileSync: native.exec })); +vi.mock("./manager.js", () => ({ + DEFAULT_MODEL: "small.en", + ensureWhisper: native.runtime, + ensureModel: async ( + model: string, + options: { onDownloadProgress?: (received: number, total: number | null) => void }, + ) => { + options.onDownloadProgress?.(5, null); + return `ggml-${model}.bin`; + }, + hasFFmpeg: () => true, +})); +vi.mock("../browser/ffmpeg.js", () => ({ + findFFprobe: () => "ffprobe", + findFFmpeg: () => "ffmpeg", + getFFmpegInstallHint: () => "ffmpeg", +})); +let dir: string; +beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), "hf-language-")); + writeFileSync(join(dir, "audio.wav"), Buffer.alloc(44)); + native.missingLanguage = false; + native.runtime.mockReset().mockResolvedValue({ executablePath: "whisper-cli", source: "env" }); + native.exec.mockReset().mockImplementation((command: string, args: string[]) => { + if (command === "ffprobe") + return JSON.stringify({ + streams: [ + { codec_type: "audio", codec_name: "pcm_s16le", sample_rate: "16000", channels: 1 }, + ], + }); + if (command !== "whisper-cli") throw new Error(`Unexpected executable: ${command}`); + if (args.includes("--detect-language")) return ""; + const selected = args[args.indexOf("--language") + 1]; + let language = args.includes("--language") ? selected : "en"; + if (language === "auto") language = "es"; + const output = args[args.indexOf("--output-file") + 1]; + writeFileSync( + `${output}.json`, + JSON.stringify({ + result: native.missingLanguage ? {} : { language }, + transcription: [ + { + tokens: [ + { text: language === "es" ? "Hola" : "Hello", offsets: { from: 0, to: 1000 } }, + ], + }, + ], + }), + ); + return ""; + }); +}); +afterEach(() => rmSync(dir, { recursive: true, force: true })); + +it("decodes multilingual audio automatically and returns the native detection", async () => { + const result = await transcribe(join(dir, "audio.wav"), dir, { model: "small" }); + expect(result).toMatchObject({ detectedLanguage: "es", model: "small", wordCount: 1 }); + expect(native.exec.mock.calls.filter(([command]) => command === "whisper-cli")).toHaveLength(1); +}); +it.each([ + ["small.en", undefined, "small.en"], + ["small.en", "de", "small"], + ["small", "en", "small"], +])("does not label requested %s / %s as detection", async (model, language, resolved) => { + const result = await transcribe(join(dir, "audio.wav"), dir, { model, language }); + expect(result).toMatchObject({ detectedLanguage: null, model: resolved }); +}); +it("keeps an absent native language unknown", async () => { + native.missingLanguage = true; + expect(await transcribe(join(dir, "audio.wav"), dir, { model: "small" })).toMatchObject({ + detectedLanguage: null, + }); +}); +it("reports download bytes and truthful transcription boundaries", async () => { + const events: unknown[] = []; + await transcribe(join(dir, "audio.wav"), dir, { + model: "small", + onEvent: (event) => events.push(event), + }); + expect(events).toEqual([ + { type: "progress", phase: "download", model: "small", receivedBytes: 5, totalBytes: null }, + { type: "progress", phase: "transcription", model: "small", status: "started" }, + { type: "progress", phase: "transcription", model: "small", status: "completed" }, + ]); +}); + +it("carries the runtime-install policy to discovery while leaving model downloads enabled", async () => { + const events: unknown[] = []; + await transcribe(join(dir, "audio.wav"), dir, { + installRuntime: false, + onEvent: (event) => events.push(event), + }); + expect(native.runtime).toHaveBeenCalledWith(expect.objectContaining({ installRuntime: false })); + expect(events).toContainEqual({ + type: "progress", + phase: "download", + model: "small.en", + receivedBytes: 5, + totalBytes: null, + }); +}); +it("does not claim transcription completed when decoding fails", async () => { + const events: unknown[] = []; + const previous = native.exec.getMockImplementation()!; + native.exec.mockImplementation((command: string, args: string[]) => { + if (command === "whisper-cli") throw new Error("decoder failed"); + return previous(command, args); + }); + await expect( + transcribe(join(dir, "audio.wav"), dir, { + model: "small", + onEvent: (event) => events.push(event), + }), + ).rejects.toThrow("decoder failed"); + expect(events).toEqual([ + { type: "progress", phase: "download", model: "small", receivedBytes: 5, totalBytes: null }, + { type: "progress", phase: "transcription", model: "small", status: "started" }, + ]); +}); diff --git a/packages/cli/src/whisper/transcribe.ts b/packages/cli/src/whisper/transcribe.ts index 1c890a0292..a69992d334 100644 --- a/packages/cli/src/whisper/transcribe.ts +++ b/packages/cli/src/whisper/transcribe.ts @@ -8,24 +8,6 @@ import { findFFmpeg, findFFprobe, getFFmpegInstallHint } from "../browser/ffmpeg import { stoppedByCancelSignal } from "../utils/renderCancellation.js"; import { ensureWhisper, ensureModel, hasFFmpeg, DEFAULT_MODEL } from "./manager.js"; -/** - * Detect the language of a WAV file using whisper's built-in language detection. - * Returns an ISO 639-1 code (e.g. "en", "es", "hi") or null if detection fails. - */ -function detectLanguage(whisperPath: string, modelPath: string, wavPath: string): string | null { - try { - const output = execFileSync(whisperPath, ["--model", modelPath, "--detect-language", wavPath], { - encoding: "utf-8", - timeout: 30_000, - stdio: ["ignore", "pipe", "pipe"], - }); - const match = output.match(/auto-detected language:\s*(\w+)/); - return match?.[1] ?? null; - } catch { - return null; - } -} - function findWavDataChunk(buf: Buffer): { offset: number; size: number } | null { if (buf.length < 12) return null; let pos = 12; // skip RIFF header @@ -258,10 +240,22 @@ export function detectSpeechOnset(wavPath: string): number | null { const AUDIO_EXTENSIONS = new Set([".mp3", ".wav", ".m4a", ".aac", ".ogg", ".flac"]); const VIDEO_EXTENSIONS = new Set([".mp4", ".webm", ".mov", ".mkv", ".avi"]); +export type TranscribeProgress = + | { + type: "progress"; + phase: "download"; + model: string; + receivedBytes: number; + totalBytes: number | null; + } + | { type: "progress"; phase: "transcription"; model: string; status: "started" | "completed" }; + export interface TranscribeOptions { + installRuntime?: boolean; model?: string; language?: string; onProgress?: (message: string) => void; + onEvent?: (event: TranscribeProgress) => void; /** * Explicit whisper spawn timeout in ms. Overrides the duration+model auto- * scaled default. Callers that leave this undefined get the auto-scaled @@ -271,6 +265,8 @@ export interface TranscribeOptions { } export interface TranscribeResult { + model: string; + detectedLanguage: string | null; transcriptPath: string; wordCount: number; durationSeconds: number; @@ -449,63 +445,52 @@ export async function transcribe( // 1. Ensure whisper binary options?.onProgress?.("Checking whisper..."); - const whisper = await ensureWhisper({ onProgress: options?.onProgress }); + const whisper = await ensureWhisper({ + onProgress: options?.onProgress, + installRuntime: options?.installRuntime, + }); // 2. Ensure model options?.onProgress?.("Checking model..."); const modelPath = await ensureModel(model, { onProgress: options?.onProgress, + onDownloadProgress: options?.onEvent + ? (receivedBytes, totalBytes) => + options.onEvent?.({ + type: "progress", + phase: "download", + model, + receivedBytes, + totalBytes, + }) + : undefined, }); // 3. Prepare audio const wavPath = prepareWav(inputPath, options?.onProgress); - // 4. Detect language and ensure correct model - let effectiveModel = model; - let effectiveModelPath = modelPath; - let detectedLanguage = options?.language ?? null; - - // Only auto-detect language when using a multilingual model. - // .en models always report "en" regardless of actual language, so detection - // would be a no-op. If the user chose .en, they want English. - if (!detectedLanguage && !effectiveModel.endsWith(".en")) { - options?.onProgress?.("Detecting language..."); - detectedLanguage = detectLanguage(whisper.executablePath, effectiveModelPath, wavPath); - } - - if (detectedLanguage && detectedLanguage !== "en" && effectiveModel.endsWith(".en")) { - const multilingualModel = effectiveModel.replace(/\.en$/, ""); - options?.onProgress?.( - `Detected ${detectedLanguage} — switching to ${multilingualModel} model...`, - ); - effectiveModelPath = await ensureModel(multilingualModel, { - onProgress: options?.onProgress, - }); - effectiveModel = multilingualModel; - } - - // 5. Run whisper + const automaticLanguage = options?.language === undefined && !model.endsWith(".en"); + const language = options?.language ?? (automaticLanguage ? "auto" : "en"); options?.onProgress?.("Transcribing..."); + options?.onEvent?.({ type: "progress", phase: "transcription", model, status: "started" }); const outputBase = join(outputDir, "transcript"); mkdirSync(outputDir, { recursive: true }); const whisperArgs = [ "--model", - effectiveModelPath, + modelPath, "--output-json-full", "--output-file", outputBase, "--dtw", - dtwPresetForModel(effectiveModel), + dtwPresetForModel(model), "--suppress-nst", ]; - if (detectedLanguage) { - whisperArgs.push("--language", detectedLanguage); - } + whisperArgs.push("--language", language); whisperArgs.push(wavPath); const whisperTimeoutMs = resolveWhisperTimeoutMs(getPreparedWavDurationSeconds(wavPath), { - model: effectiveModel, + model, overrideMs: options?.timeoutMs, }); try { @@ -520,7 +505,7 @@ export async function transcribe( // existing stderr-tail handling in `transcribeAudio` still applies. throw wrapWhisperTimeoutError(err, { effectiveTimeoutMs: whisperTimeoutMs, - model: effectiveModel, + model, wasOverride: options?.timeoutMs != null, }); } @@ -532,6 +517,11 @@ export async function transcribe( } const transcript = JSON.parse(readFileSync(transcriptPath, "utf-8")); + const reportedLanguage: unknown = transcript.result?.language; + const detectedLanguage = + automaticLanguage && typeof reportedLanguage === "string" && reportedLanguage.length > 0 + ? reportedLanguage + : null; const segments = transcript.transcription ?? []; let wordCount = 0; @@ -557,7 +547,10 @@ export async function transcribe( } } + options?.onEvent?.({ type: "progress", phase: "transcription", model, status: "completed" }); return { + model, + detectedLanguage, transcriptPath, wordCount, durationSeconds: maxEnd / 1000,