Skip to content

Audio Capture And Transcription ​

ISA Warden owns microphone capture and transcription so extensions do not need direct media-device, filesystem, model, or credential access. This capability is implemented for the Tauri desktop host. The current API is batch-only: record first, stop the recording, and then start transcription.

Host and dashboard setup ​

Before an extension can list a transcription provider, the dashboard must make one available to its exact launch group:

  1. Add a model to workspace inventory with purpose speechToText.
  2. Link that model explicitly to the group from which the extension is launched.
  3. Link a dashboard-installed extension to the same group.

A distributed extension may declare a speechToTextmodelRecommendations item to help an administrator create the workspace model during Add Extension. That recommendation does not perform steps 2 or 3, does not download local model bytes, and does not replace runtime transcription_list_providers discovery.

For local Whisper, the dashboard Hugging Face manager's transcription filter searches automatic-speech-recognition repositories and accepts whisper.cpp GGML files named like ggml-small.bin. These models use providerKind: "localWhisper"; an OpenAI-compatible model uses providerKind: "openAiCompatible" and keeps its endpoint and API key in the host.

A speech-to-text model does not need to appear in the host's normal chat-model selector. That selector and its vision indicator describe chat capabilities; transcription_list_providers is the authoritative extension inventory.

Dashboard-installed extensions must be linked to the launch group. An explicitly mounted developer extension is treated as locally trusted during development and may use that exact group's models without a dashboard extension record. This exception does not widen workspace or group model visibility.

Manifest and capability discovery ​

A record-and-transcribe extension normally declares:

json
{
  "permissions": [
    "get_app_info",
    "audio_capture",
    "transcribe_audio"
  ]
}

Add filesystem_read and filesystem_write only when derived transcript or project files must be shared through dashboard-backed group storage.

Check capabilities at runtime:

js
const bridge = window.isaExtensionBridge;
const appInfo = await bridge.getAppInfo();

if (!appInfo.audioCapture?.enabled) {
  throw new Error(appInfo.audioCapture?.disabledReason || 'Audio capture is unavailable');
}
if (!appInfo.transcription?.enabled) {
  throw new Error(appInfo.transcription?.disabledReason || 'Transcription is unavailable');
}

The host injects and overwrites the extension ID and frozen workspace/group context for every media command. Never send credentials, consent flags, workspace IDs, or group IDs from extension code.

Privacy boundary ​

Raw audio is local and ephemeral. The host stores a WAV recording below private app data, gives the extension only an opaque recordingId, and never uploads audio to dashboard workspace files. There is no audio publish/share command and no API for audio bytes, base64, blob URLs, or local paths.

An extension may share derived text, transcript segments, minutes, and project metadata through the normal group-scoped filesystem API. Delete the local recording only after those derived files have been committed durably to the host sync store. Abandoned recordings expire after 24 hours by default.

For an OpenAI-compatible transcription model, the host asks the user on every job before sending audio. Approval creates a short-lived, single-use host token bound to the exact extension, workspace, group, recording, and model. The host resolves the endpoint and API key; iframe arguments cannot override them or forge consent. Local Whisper does not send audio off-device.

Record a meeting ​

js
const devices = await bridge.listAudioInputDevices();
const preferredDevice = devices.find((device) => device.isDefault) ?? devices[0];

if (!preferredDevice) throw new Error('No microphone is available');

const recording = await bridge.startRecording({
  deviceId: preferredDevice.deviceId,
  maxDurationMs: 4 * 60 * 60 * 1000
});

// Keep recording.recordingId in extension state immediately.
// Later, after the user presses Stop:
const stopped = await bridge.stopRecording({
  recordingId: recording.recordingId
});

The first start opens a host-owned microphone consent dialog. The grant is stored locally per extension and can be revoked in Settings. Revocation stops an active recording and blocks future starts.

Only one recording can be active across ISA Warden. Capture continues when the extension iframe is minimized, and the global recording control lets the user pause, resume, or stop it. Recording states are starting, recording, paused, stopped, cancelled, and failed.

Subscribe before starting when possible, and always remove listeners when the view is disposed:

js
const unlistenAudio = bridge.on('audio-capture-state-changed', ({ state }) => {
  if (state.recordingId === recording.recordingId) renderRecordingState(state);
});

// Retain this function and call it during component teardown.
const disposeAudioListener = () => unlistenAudio();

Events are not a state store. On mount, call getActiveRecording(), listRecordings(), or getRecording({ recordingId }) to reconcile anything that happened while the iframe was absent.

List providers ​

js
const providers = await bridge.listTranscriptionProviders();

// Example descriptor:
// {
//   id: 'model:...',
//   name: 'ggerganov/whisper.cpp · ggml-small',
//   providerKind: 'localWhisper',
//   location: 'local',
//   modelType: 'speechToText',
//   supportsSpeakerLabels: false
// }

An empty list means the frozen launch group currently has no authorized speech-to-text model. Refresh after dashboard/group changes, on window focus or visibility restoration, and immediately before starting a job. Do not fall back to a workspace-wide model or another group's model.

Provider listing is authorization discovery, not local download status. A local model may be listed before its model file is ready.

supportsSpeakerLabels is model-specific. The host transcription capability can transport labels, but local Whisper and ordinary remote transcription models still return unlabeled segments.

Configure speaker-aware remote transcription ​

Speaker separation is provider behavior selected through trusted model configuration, not an extension request flag. Configure a compatible remote speech-to-text model with:

json
{
  "transcription": {
    "responseFormat": "diarized_json",
    "chunkingStrategy": "auto"
  }
}

For OpenAI's diarization model, the host sends diarized_json with automatic server chunking and omits prompt and timestamp-granularity fields that the model does not support. The normalized result keeps provider labels on segments[].speaker.

Speaker labels are opaque identities within provider output. Extensions may offer explicit mapping to participant names, but must not guess those names. When the host creates multiple independent uploads, it scopes labels per upload part so identical provider-local labels are not silently merged across parts.

Start and recover transcription ​

The first start for a local model may initiate a host-managed download. In that case startTranscription() rejects before a job is created with the message transcription_model_downloading. Keep the stopped recording, show a waiting state, and retry only that sentinel after a short delay. There is currently no public model-download progress event.

js
const delay = (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds));

function bridgeErrorMessage(error) {
  return String(error?.message ?? error ?? 'unknown_error');
}

async function startTranscriptionWhenReady(request, shouldContinue = () => true) {
  while (shouldContinue()) {
    try {
      return await bridge.startTranscription(request);
    } catch (error) {
      const message = bridgeErrorMessage(error);
      if (!message.includes('transcription_model_downloading')) throw error;
      await delay(3000);
    }
  }
  throw new Error('transcription_start_cancelled');
}

const provider = providers[0];
const started = await startTranscriptionWhenReady({
  recordingId: stopped.recordingId,
  modelId: provider.id,
  language: 'nl',
  prompt: 'Optional names or domain vocabulary'
});

// Persist this in extension state/storage immediately; there is no list-jobs API.
const jobId = started.jobId;

transcription_model_not_downloaded is not a waiting signal: the host could not establish a downloaded or downloadable asset. Show a setup error. transcription_model_unavailable means the selected model is no longer available to the exact launch group, so refresh providers. For a remote model, declining the host dialog rejects with external_transcription_cancelled and does not create a job.

After a job exists, events can reduce polling latency, but transcription_get remains authoritative:

js
const terminalStatuses = new Set(['completed', 'failed', 'cancelled', 'interrupted']);
let wakePoll = () => {};

const unlistenProgress = bridge.on('transcription-progress', (event) => {
  if (event.jobId === jobId) wakePoll();
});
const unlistenSegment = bridge.on('transcription-segment', (event) => {
  if (event.jobId === jobId && event.segment) renderPartialSegment(event.segment);
});
const unlistenCompleted = bridge.on('transcription-completed', (event) => {
  if (event.jobId === jobId) wakePoll();
});
const unlistenFailed = bridge.on('transcription-failed', (event) => {
  if (event.jobId === jobId) wakePoll();
});

async function waitForTranscription(id) {
  for (;;) {
    const query = await bridge.getTranscription({ jobId: id });
    renderTranscriptionJob(query.job);

    if (query.job.status === 'completed' && query.result) return query.result;
    if (terminalStatuses.has(query.job.status)) {
      throw new Error(query.job.errorMessage || query.job.errorCode || query.job.status);
    }

    await Promise.race([
      delay(1000),
      new Promise((resolve) => {
        wakePoll = resolve;
      })
    ]);
  }
}

let result;
try {
  result = await waitForTranscription(jobId);
} finally {
  unlistenProgress();
  unlistenSegment();
  unlistenCompleted();
  unlistenFailed();
}

Persisted jobs survive iframe remounts. Store jobId, then call getTranscription({ jobId }) on mount. An unfinished job becomes interrupted after a full host restart; ISA Warden preserves its state but does not resume the inference automatically.

The full response and event schemas, status values, phases, limits, and error contract are in the host command reference.

Share derived text, then clean up ​

Raw audio cannot be shared. To make a transcript or project available through the dashboard, write the derived content below the extension's group folder as a syncable file:

js
const folder = await bridge.ensureGroupFolder();
const safeName = `meeting-${new Date().toISOString().replaceAll(':', '-')}.json`;

await bridge.uploadFiles({
  files: [{
    path: `${folder.path}/${safeName}`,
    text: JSON.stringify({
      title: 'Weekly project meeting',
      transcript: result.text,
      language: result.language,
      durationMs: result.durationMs,
      segments: result.segments
    }),
    contentType: 'application/json',
    metadata: { syncable: 'true' }
  }]
});

A successful local or pending syncable upload means the host has durably stored the complete bytes and can reconcile them with dashboard storage later. Only after that success should the extension remove local media state:

js
await bridge.deleteTranscription({ jobId });
await bridge.deleteRecording({ recordingId: stopped.recordingId });

deleteTranscription accepts terminal jobs only. A recording cannot be deleted while an active transcription still leases it; after a completion event, a very short retry may be needed while the worker releases that lease. Never delete a recording merely because an event arrived before the transcript/project write has succeeded.

Lifecycle summary ​

Local Whisper normalizes the host WAV input to mono 16 kHz and runs one local job at a time. Long remote inputs are split into bounded overlapping WAV chunks and merged with adjusted timestamps. Configured remote providers may return speaker-aware segments; local Whisper does not. The current implementation does not provide system/loopback audio or live transcription.

ISA Warden extension specification