Skip to main content

Overview

AzureVoiceLiveLLMService provides real-time speech-to-speech conversation using Azure’s Voice Live API over WebSocket. Voice Live combines speech recognition, a generative model, and Azure text-to-speech behind a single realtime session, and adds noise suppression, echo cancellation, and semantic end-of-turn detection. The service supports function calling, caller transcription, and conversation management. Voice Live is a separate product from Azure OpenAI Realtime. The two speak different event names and are not interchangeable; Azure OpenAI Realtime deployments are served by AzureRealtimeLLMService.

Azure Voice Live API Reference

Pipecat’s API methods for Azure Voice Live integration

Example Implementation

Complete Azure Voice Live conversation example

Azure Documentation

Official Voice Live documentation

Azure AI Foundry

Create a Foundry resource and manage API keys

Installation

To use Azure Voice Live, install the required dependencies:

Prerequisites

Azure Account Setup

Before using Azure Voice Live, you need:
  1. Azure Account: An Azure subscription
  2. Foundry Resource: An Azure AI Foundry resource with Voice Live available
  3. Credentials: An API key for the resource, or a Microsoft Entra ID token provider

Required Environment Variables

  • AZURE_VOICE_LIVE_API_KEY: API key for the resource
  • AZURE_VOICE_LIVE_ENDPOINT: Resource endpoint (for example https://<resource>.services.ai.azure.com)

Key Features

  • Speech-to-Speech: Audio in and audio out through a single realtime session
  • Azure Voices: Standard, custom, personal, and realtime-native voices
  • Semantic VAD: Server-side turn detection, including a multilingual variant
  • Audio Enhancements: Noise suppression and echo cancellation
  • Function Calling: Register functions as for any other LLM service
  • Caller Transcription: User turns are added to the context as transcripts

Configuration

AzureVoiceLiveLLMService

str
required
Voice Live endpoint for the Foundry resource. Accepts the resource endpoint as shown in the Azure portal (https://<resource>.services.ai.azure.com) or a full WebSocket URL; the scheme and /voice-live/realtime path are filled in when absent.
str
default:"None"
API key for the resource. Required unless token_provider is given.
AzureTokenProvider
default:"None"
Async callable supplying a Microsoft Entra ID bearer token, used instead of api_key when given. Build one with azure.identity.aio.get_bearer_token_provider and the https://ai.azure.com/.default scope.
str
default:"2026-07-15"
Voice Live API version to request.
AzureVoiceLiveLLMService.Settings
default:"None"
Runtime-updatable settings. See Settings below.
bool
default:"False"
Whether to start with audio input paused.
Any
Additional arguments passed to parent LLMService.
Either api_key or token_provider is required; a ValueError is raised when both are missing.

Settings

Settings passed via the settings constructor argument using AzureVoiceLiveLLMService.Settings(...). See Service Settings for details. The default session_properties uses modalities=["text", "audio"], the en-US-Ava:DragonHDLatestNeural Azure standard voice, azure_semantic_vad turn detection, azure-speech input transcription, and input noise reduction. For azure-realtime models the default voice is left unset so the model picks one of its own native voices.
Because session_properties replaces all defaults, provide a complete SessionProperties when setting it. A change sent with LLMUpdateSettingsFrame updates the session, except model, which cannot change once connected.

SessionProperties

SessionProperties and the types below are imported from pipecat.services.azure.voice_live.events.

Usage

Basic Setup

Custom Session Properties

Pipeline

Notes

  • Turn taking: The service proposes turn boundaries from Voice Live’s server-side VAD events, which the recommended external user turn strategies resolve into UserStartedSpeakingFrame and UserStoppedSpeakingFrame. LLMContextAggregatorPair detects this realtime service automatically.
  • Local VAD: To drive turns from a local VAD (LLMUserAggregatorParams.vad_analyzer), pass turn_detection=None in session_properties. While server-side turn detection is on, the strategies this service recommends replace those the analyzer installs, leaving the analyzer with no say in turn-taking.
  • Caller transcription: The caller’s turns reach the context as transcripts. If session_properties omits input_audio_transcription, a warning is logged and those turns are missing from the history, including the history reset_conversation sends to a new session. Set it to None explicitly to opt out of the warning.
  • Model selection: model is selected by the connection URL, so it is fixed for the session and a runtime change is reported as unsupported.
  • Azure OpenAI Realtime: Use AzureRealtimeLLMService for Azure OpenAI Realtime deployments; Voice Live uses different event names.

Event Handlers