> ## Documentation Index
> Fetch the complete documentation index at: https://daily-main.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Azure Voice Live

> AzureVoiceLiveLLMService connects Pipecat to Azure's Voice Live API for real-time speech-to-speech conversation with Azure text-to-speech voices.

## Overview

`AzureVoiceLiveLLMService` provides real-time speech-to-speech conversation using Azure's Voice Live API over WebSocket. Voice Live combines speech recognition, a generative model, and Azure text-to-speech behind a single realtime session, and adds noise suppression, echo cancellation, and semantic end-of-turn detection. The service supports function calling, caller transcription, and conversation management.

Voice Live is a separate product from Azure OpenAI Realtime. The two speak different event names and are not interchangeable; Azure OpenAI Realtime deployments are served by `AzureRealtimeLLMService`.

<CardGroup cols={2}>
  <Card title="Azure Voice Live API Reference" icon="code" href="https://reference-server.pipecat.ai/en/latest/api/pipecat.services.azure.voice_live.llm.html">
    Pipecat's API methods for Azure Voice Live integration
  </Card>

  <Card title="Example Implementation" icon="play" href="https://github.com/pipecat-ai/pipecat/blob/main/examples/realtime/realtime-azure-voice-live.py">
    Complete Azure Voice Live conversation example
  </Card>

  <Card title="Azure Documentation" icon="book" href="https://learn.microsoft.com/azure/ai-services/speech-service/voice-live">
    Official Voice Live documentation
  </Card>

  <Card title="Azure AI Foundry" icon="external-link" href="https://ai.azure.com/">
    Create a Foundry resource and manage API keys
  </Card>
</CardGroup>

## Installation

To use Azure Voice Live, install the required dependencies:

```bash theme={null}
uv add "pipecat-ai[azure]"
```

## Prerequisites

### Azure Account Setup

Before using Azure Voice Live, you need:

1. **Azure Account**: An Azure subscription
2. **Foundry Resource**: An Azure AI Foundry resource with Voice Live available
3. **Credentials**: An API key for the resource, or a Microsoft Entra ID token provider

### Required Environment Variables

* `AZURE_VOICE_LIVE_API_KEY`: API key for the resource
* `AZURE_VOICE_LIVE_ENDPOINT`: Resource endpoint (for example `https://<resource>.services.ai.azure.com`)

### Key Features

* **Speech-to-Speech**: Audio in and audio out through a single realtime session
* **Azure Voices**: Standard, custom, personal, and realtime-native voices
* **Semantic VAD**: Server-side turn detection, including a multilingual variant
* **Audio Enhancements**: Noise suppression and echo cancellation
* **Function Calling**: Register functions as for any other LLM service
* **Caller Transcription**: User turns are added to the context as transcripts

## Configuration

### AzureVoiceLiveLLMService

<ParamField path="endpoint" type="str" required>
  Voice Live endpoint for the Foundry resource. Accepts the resource endpoint as
  shown in the Azure portal (`https://<resource>.services.ai.azure.com`) or a
  full WebSocket URL; the scheme and `/voice-live/realtime` path are filled in
  when absent.
</ParamField>

<ParamField path="api_key" type="str" default="None">
  API key for the resource. Required unless `token_provider` is given.
</ParamField>

<ParamField path="token_provider" type="AzureTokenProvider" default="None">
  Async callable supplying a Microsoft Entra ID bearer token, used instead of
  `api_key` when given. Build one with
  `azure.identity.aio.get_bearer_token_provider` and the
  `https://ai.azure.com/.default` scope.
</ParamField>

<ParamField path="api_version" type="str" default="2026-07-15">
  Voice Live API version to request.
</ParamField>

<ParamField path="settings" type="AzureVoiceLiveLLMService.Settings" default="None">
  Runtime-updatable settings. See [Settings](#settings) below.
</ParamField>

<ParamField path="start_audio_paused" type="bool" default="False">
  Whether to start with audio input paused.
</ParamField>

<ParamField path="**kwargs" type="Any">
  Additional arguments passed to parent LLMService.
</ParamField>

<Note>
  Either `api_key` or `token_provider` is required; a `ValueError` is raised
  when both are missing.
</Note>

### Settings

Settings passed via the `settings` constructor argument using `AzureVoiceLiveLLMService.Settings(...)`. See [Service Settings](/pipecat/fundamentals/service-settings) for details.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `model` | `str` | `"gpt-4o-mini"` | Model for the session. Sent as a query parameter on the connection, so it is fixed for the session. *(Inherited from base settings.)* |
| `system_instruction` | `str` | `None` | System instructions for the model. Kept in sync with `session_properties.instructions`. *(Inherited from base settings.)* |
| `temperature` | `float` | `None` | Sampling temperature, copied into `session_properties`. *(Inherited from base settings.)* |
| `session_properties` | `SessionProperties` | see below | Voice Live session configuration: voice, turn detection, transcription, noise reduction, echo cancellation, and tools. Replaces the defaults wholesale when provided. |

The default `session_properties` uses `modalities=["text", "audio"]`, the `en-US-Ava:DragonHDLatestNeural` Azure standard voice, `azure_semantic_vad` turn detection, `azure-speech` input transcription, and input noise reduction. For `azure-realtime` models the default voice is left unset so the model picks one of its own native voices.

<Note>
  Because `session_properties` **replaces** all defaults, provide a complete
  `SessionProperties` when setting it. A change sent with
  `LLMUpdateSettingsFrame` updates the session, except `model`, which cannot
  change once connected.
</Note>

### SessionProperties

`SessionProperties` and the types below are imported from `pipecat.services.azure.voice_live.events`.

| Field | Type | Description |
| - | - | - |
| `modalities` | `list["text" \| "audio"]` | Output modalities. |
| `voice` | `AzureStandardVoice \| AzureCustomVoice \| AzurePersonalVoice \| AzureRealtimeNativeVoice \| OpenAIVoice` | Voice used for speech output. |
| `turn_detection` | `TurnDetection \| None` | Server-side turn detection. `None` disables it (see [Notes](#notes)). |
| `input_audio_transcription` | `InputAudioTranscription \| None` | Caller transcription. Without it, the caller's turns are missing from the context. |
| `input_audio_noise_reduction` | `InputAudioNoiseReduction \| None` | Noise suppression: `near_field`, `far_field`, or `azure_deep_noise_suppression`. |
| `input_audio_echo_cancellation` | `InputAudioEchoCancellation \| None` | Server echo cancellation. |
| `instructions` | `str` | System instructions; synced with `Settings.system_instruction`. |
| `temperature` | `float` | Sampling temperature. |
| `max_response_output_tokens` | `int \| "inf"` | Maximum tokens per response. |

## Usage

### Basic Setup

```python theme={null}
import os
from pipecat.services.azure.voice_live.llm import AzureVoiceLiveLLMService

llm = AzureVoiceLiveLLMService(
    api_key=os.getenv("AZURE_VOICE_LIVE_API_KEY"),
    endpoint=os.getenv("AZURE_VOICE_LIVE_ENDPOINT"),
    settings=AzureVoiceLiveLLMService.Settings(
        model="gpt-4o-mini",
        system_instruction="You are a helpful voice assistant.",
    ),
)
```

### Custom Session Properties

```python theme={null}
from pipecat.services.azure.voice_live.events import (
    AzureStandardVoice,
    InputAudioNoiseReduction,
    InputAudioTranscription,
    SessionProperties,
    TurnDetection,
)

llm = AzureVoiceLiveLLMService(
    api_key=os.getenv("AZURE_VOICE_LIVE_API_KEY"),
    endpoint=os.getenv("AZURE_VOICE_LIVE_ENDPOINT"),
    settings=AzureVoiceLiveLLMService.Settings(
        session_properties=SessionProperties(
            modalities=["text", "audio"],
            voice=AzureStandardVoice(name="en-US-Ava:DragonHDLatestNeural"),
            turn_detection=TurnDetection(
                type="azure_semantic_vad",
                silence_duration_ms=500,
                remove_filler_words=True,
            ),
            input_audio_transcription=InputAudioTranscription(model="azure-speech"),
            input_audio_noise_reduction=InputAudioNoiseReduction(),
        ),
    ),
)
```

### Pipeline

```python theme={null}
from pipecat.pipeline.pipeline import Pipeline
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import LLMContextAggregatorPair

context = LLMContext(
    [{"role": "user", "content": "Say hello."}],
)
user_aggregator, assistant_aggregator = LLMContextAggregatorPair(context)

pipeline = Pipeline(
    [
        transport.input(),
        user_aggregator,
        llm,
        transport.output(),
        assistant_aggregator,
    ]
)
```

## Notes

* **Turn taking**: The service proposes turn boundaries from Voice Live's server-side VAD events, which the recommended external user turn strategies resolve into `UserStartedSpeakingFrame` and `UserStoppedSpeakingFrame`. `LLMContextAggregatorPair` detects this realtime service automatically.
* **Local VAD**: To drive turns from a local VAD (`LLMUserAggregatorParams.vad_analyzer`), pass `turn_detection=None` in `session_properties`. While server-side turn detection is on, the strategies this service recommends replace those the analyzer installs, leaving the analyzer with no say in turn-taking.
* **Caller transcription**: The caller's turns reach the context as transcripts. If `session_properties` omits `input_audio_transcription`, a warning is logged and those turns are missing from the history, including the history `reset_conversation` sends to a new session. Set it to `None` explicitly to opt out of the warning.
* **Model selection**: `model` is selected by the connection URL, so it is fixed for the session and a runtime change is reported as unsupported.
* **Azure OpenAI Realtime**: Use `AzureRealtimeLLMService` for Azure OpenAI Realtime deployments; Voice Live uses different event names.

## Event Handlers

| Event | Description |
| - | - |
| `on_conversation_item_created` | Called with the item ID and item when a conversation item is created |
| `on_conversation_item_updated` | Called with the item ID and item (or `None`) when an item is updated |

```python theme={null}
@llm.event_handler("on_conversation_item_created")
async def on_conversation_item_created(service, item_id, item):
    print(f"Item created: {item_id}")

@llm.event_handler("on_conversation_item_updated")
async def on_conversation_item_updated(service, item_id, item):
    print(f"Item updated: {item_id}")
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.