Skip to main content

Voice

Voice in Cobalt is three separate things, each independently available:

The widget experience

In the web widget, dictation and voice mode sit in the composer, and each answer offers read-aloud. Read-aloud uses the channel’s preferred voice, which you pick in the channel’s branding settings. Voice mode needs microphone permission from the browser — on embedded intranets (SharePoint), the page must allow microphone access; see the SharePoint guide’s troubleshooting if the mic prompt never appears.

Slack voice messages

Slack’s own voice clips are supported on the Slack channel: when transcription is enabled in the channel settings, an employee can send a voice message and the agent answers it as a normal text turn. Long clips over the size cap are declined with a short explanation rather than silently dropped.

Anonymous visitors and abuse control

On anonymous widget channels (no sign-in), voice is rate-limited so a public embed can’t be abused:
  • Per-visitor limits on dictation, uploads, and voice-mode sessions.
  • A per-tenant daily cap on anonymous voice-mode usage — when the day’s cap is reached, voice mode pauses until the next day while text chat continues normally.
Signed-in channels (Embedded token, Provider token, Cobalt sign-in) aren’t subject to the anonymous caps.

Who the agent is talking to

A voice conversation uses the same identity as the text chat it starts from. Someone signed in on the widget or the hosted page is the same person when they switch to voice, and a voice session is refused whenever the text session would be. Capabilities limited to Allowed groups behave the same in voice as in text: someone outside the groups never hears about a restricted tool, skill, handoff or knowledge collection, and the agent can’t use it for them.

Usage

Voice turns are metered like any other conversation turn — one credit per answered request, regardless of whether it was typed or spoken. See Billing & usage.