Transcription language and progress
Carry automatic language detection in the JSON result and emit progress separately so stdout stays parseable.
Where the value travels
Whisper JSON result.language → transcription result → CLI stdout → JSON consumer +HTTP response bytes → download pipeline → model callback → CLI stderr → progress consumer +Whisper words → normalization → transcript.json (word array, unchanged)
The contract, per boundary
+Language
The native result identifies the language used to decode. Only automatic selection qualifies as detection.
result.language: string +Proposed terminal addition: detectedLanguage: string | null
Upstream JSON writes result.language from whisper_full_lang_id. Explicit language also sets that state. Existing detection incorrectly searches stdout although the native logger writes stderr; missing detection then omits --language and falls back to native English default.
whisper.cpp 60c0be6ac: examples/cli/cli.cpp:84,735-736; src/whisper.cpp:6993-7013,7151-7153,9357-9366. CLI: whisper/transcribe.ts:17-23,466-504,534,560.Result and artifact
Add metadata to the terminal result and preserve the existing word-array file.
{ok:true, engine, model, wordCount, durationSeconds,
+ speechOnsetSeconds, transcriptPath, detectedLanguage}The existing model field currently reports the requested model even when the wrapper selects a multilingual variant. The wrapper will own the resolved model and return it with the result.
commands/transcribe.ts:397 writes Word[]; :400-410 emits stdout. Existing file consumer: commands/transcribe.ts:384 calls loadTranscript. External new-field consumption remains unverified.Download progress
Count bytes in the existing transfer pipeline. Unknown size stays unknown.
{"type":"progress","phase":"download","model":"small",
+ "receivedBytes":65536,"totalBytes":null}DownloadOptions currently has timeoutMs and maxBytes only. Add a byte callback at the transfer owner, thread it through ensureModel, and serialize it only at the command boundary. Cache hits emit no download events. A known total is bytes, never an invented percentage.
utils/download.ts:10-15,113-122; whisper/manager.ts:215-223; commands/transcribe.ts:341,347-354.Transcription progress
Indicate work started and completed without inventing numerical progress.
{"type":"progress","phase":"transcription","model":"small",
+ "status":"started"}
+{"type":"progress","phase":"transcription","model":"small",
+ "status":"completed"}Progress records are individual valid JSON lines on stderr. Other stderr diagnostics remain allowed. The terminal stdout result is emitted once. Failure never emits a completed transcription event.
whisper/transcribe.ts:488,512-515; whisper/parakeet.ts:294-304; whisper/sherpa.ts:397-413,445. Plain fallback diagnostics: commands/transcribe.ts:370.How it says no
| Signal | Meaning and response |
|---|---|
| detectedLanguage:null | No actual automatic inference was reported, including explicit language, English-only model, or an engine without detection. Never infer language from a model label. |
| totalBytes:null | Response has no trustworthy size. Show bytes or indeterminate progress. |
| ok:false / nonzero exit | Operation failed. Diagnostic stderr may explain why. No completion event is evidence of success without the terminal result. |
| Progress line without type | Not part of this progress contract. Consumers filter diagnostics separately. |
Decisions approved
- detectedLanguage means real automatic detection. Explicit --language and English-only models return null.
- stderr permits diagnostics alongside progress; every progress line is valid JSON with type.
- model reports the actual resolved model. The file stays Word[].
- Transcription progress is phase-only. Download progress reports real bytes with a nullable total.
- --no-runtime-install requires an existing Whisper runtime on all paths, including auto fallback. Model downloads remain allowed. No prior opt-out exists in the inspected CLI source; ensureWhisper owns this policy.
Coverage, honestly
Audited: CLI producer, transfer pipeline, native JSON writer and logger, current terminal/file paths. Installed CLI help executed.
Trusting: arbitrary environment-selected Whisper versions are not pinned to the inspected upstream checkout.
Not exercised: native audio inference, new progress writer and external live consumer. These are contract findings, not execution proof.
What would prove this page wrong
- A public transcribe invocation with a native fixture that emits detection only on stderr must decode automatically and return its language.
- Explicit language and English-only models must return null even if native JSON reports a language.
- Download callbacks must observe actual chunks, including unknown totals, and stdout must parse as one result while stderr records parse independently.
- A switched multilingual model must be the model reported in the final result.